Given
two molecules
(SMILES + study metadata) and
molecule A's measured oral bioavailability
,
this model predicts whether oral-bioavailability behavior
transfers
from A to B — i.e. whether the
two molecules behave similarly under the given study context. It is
self-contained
: the frozen
encoders are bundled with the trained head, so it runs end-to-end on raw inputs.
Architecture
Per molecule (siamese — the same encoders + projections are applied to A and B; only the head is
position-aware):
Molecule encoder
—
ibm-research/MoLFormer-XL-both-10pct
(MolFormer-XL),
frozen
: SMILES → mean-pooled token
embedding →
768-d
, then a 2-layer MLP (768→1024→768).
Metadata encoder
—
sentence-transformers/all-MiniLM-L6-v2
(MiniLM),
frozen
: each of the
7 metadata
fields
is embedded
separately
(mean-pooled, L2-normalized) →
384-d
, then a learned
per-field projection → 64-d (7×64 =
448-d
total).
A
missing/empty field uses a learned per-field "missing" embedding
instead of the text embedding,
so absent metadata is handled gracefully and distinctly from any real value.
Pass a dict per molecule keyed by these names.
Omit a key, or pass
None
/
""
, for a missing field
— the model then uses its learned per-field "missing" embedding.
Usage
from transformers import AutoModel
m = AutoModel.from_pretrained("jiosephlee/starling-transfer-ssv2-srcval", trust_remote_code=True).eval()
out = m(
smiles_a=["CC(=O)Oc1ccccc1C(=O)O"], # molecule A (bioavailability known)
smiles_b=["CCO"], # molecule B (candidate)
metadata_a=[{"species_or_population": "human", "dose": "325 mg", "oral_exposure_mode": "tablet"}],
metadata_b=[{"species_or_population": "human"}], # missing fields are fine
source_value=[68.0], # molecule A's RAW oral_bioavailability_value (e.g. percent)
)
p_transfer = out.logits.sigmoid() # batched: pass parallel lists for many pairs
source_value
is molecule A's
raw
oral_bioavailability_value
; the model scales it internally by
100. Inputs are batched lists of equal length.
Training & performance
Trained on the
same_species_v2
oral-bioavailability transfer split (~338M molecule pairs; the frozen
embeddings are precomputed once and the head is trained on top). The label is
|value_A - value_B|
thresholded, so the model uses A's known value as an
anchor
and learns to estimate B's
bioavailability from its structure + metadata.
starling-transfer-ssv2-srcval huggingface.co is an AI model on huggingface.co that provides starling-transfer-ssv2-srcval's model effect (), which can be used instantly with this jiosephlee starling-transfer-ssv2-srcval model. huggingface.co supports a free trial of the starling-transfer-ssv2-srcval model, and also provides paid use of the starling-transfer-ssv2-srcval. Support call starling-transfer-ssv2-srcval model through api, including Node.js, Python, http.
starling-transfer-ssv2-srcval huggingface.co is an online trial and call api platform, which integrates starling-transfer-ssv2-srcval's modeling effects, including api services, and provides a free online trial of starling-transfer-ssv2-srcval, you can try starling-transfer-ssv2-srcval online for free by clicking the link below.
jiosephlee starling-transfer-ssv2-srcval online free url in huggingface.co:
starling-transfer-ssv2-srcval is an open source model from GitHub that offers a free installation service, and any user can find starling-transfer-ssv2-srcval on GitHub to install. At the same time, huggingface.co provides the effect of starling-transfer-ssv2-srcval install, users can directly use starling-transfer-ssv2-srcval installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
starling-transfer-ssv2-srcval install url in huggingface.co: