The DecoderTCR models are designed for the following primary use cases:
TCR-pMHC Binding Prediction
: Predict the interaction between T-cell receptors (TCRs) and peptide-MHC complexes
Interaction Scoring
: Calculate interface energy scores for TCR-pMHC interactions
Sequence Analysis
: Analyze TCR sequences and their interactions with specific peptides
Immunology Research
: Support research in adaptive immunity, T-cell recognition, and antigen presentation
The models are particularly useful for:
Identifying potential TCR-peptide binding pairs
Screening TCR sequences for specific antigen recognition
Understanding the molecular basis of T-cell recognition
Supporting vaccine design and immunotherapy development
Out-of-Scope or Unauthorized Use Cases
Do not use the model for the following purposes:
Use that violates applicable laws, regulations (including trade compliance laws), or third party rights such as privacy or intellectual property rights
Peptide-MHC sequences
: MHC Motif Atlas for peptide-MHC ligandomes and high confidence synthetic interactions via MixMHCpred predictions.
Paired TCR-pMHC Interactions
: VDJdb for paired TCR-pMHC interaction data.
Continual Pre-training Strategy
This model is trained using a
continual pre-training curriculum
that adapts a pretrained ESM2 backbone to new protein domains while preserving previously learned representations.
Overview
Continual pre-training proceeds in
multiple stages
, each leveraging different data regimes and masking strategies:
Stage 2 incorporate
scarcer, structured, or interaction-rich data
, refining conditional dependencies without overwriting earlier knowledge.
The architecture, tokenizer, and objective remain unchanged throughout training; only the data distribution and masking strategy evolve.
Stage 1: Component-Level Adaptation
In the first stage, the model is further pretrained on large collections of unpaired or weakly structured protein sequences relevant to the target domain.
Objective:
Masked Language Modeling (MLM)
Masking:
Component- or region-aware masking schedules that upweight functionally relevant positions
Purpose:
Adapt the pretrained ESM2 representations to the target protein subspace
Learn domain-specific sequence statistics while retaining general protein knowledge
This stage acts as a regularizer, anchoring learning in large-scale marginal data before introducing more complex dependencies.
In subsequent stages, the model is continually trained on
structured or paired sequences
that encode higher-order dependencies (e.g., interactions between protein regions or components).
Objective:
Masked Language Modeling (MLM)
Masking:
Joint masking across interacting regions to encourage cross-context conditioning
Purpose:
Refine conditional relationships learned from limited paired data
Align representations across components without degrading Stage 1 task performance
Biases, Risks, and Limitations
Potential Biases
The model may reflect biases present in the training data, including:
Overrepresentation of certain HLA alleles or peptide types
Limited diversity in TCR sequences from specific populations
Bias toward well-studied antigen systems
Certain TCR clonotypes or peptide types may be underrepresented in training data
Risks
Areas of risk may include but are not limited to:
Inaccurate predictions
: The model may produce incorrect binding predictions, especially for novel sequences or rare HLA-peptide combinations
Overconfidence
: The model may assign high confidence to predictions that are actually uncertain
Biological misinterpretation
: Users may misinterpret model outputs as definitive biological facts rather than predictions
Clinical misuse
: Use in clinical settings without proper validation could lead to incorrect treatment decisions
Limitations
Sequence length
: The model has limitations on maximum sequence length (typically ~1024 tokens)
Novel sequences
: Performance may degrade on sequences very different from training data
HLA diversity
: Limited training data for rare HLA alleles may affect prediction accuracy
Context dependency
: The model may not capture all biological context (e.g., post-translational modifications, cellular environment)
Computational requirements
: GPU is recommended for optimal performance
Caveats and Recommendations
Review and validate outputs
: Always review and validate model predictions, especially for critical applications
Experimental validation
: Model predictions should be validated experimentally before use in research or clinical contexts
Uncertainty awareness
: Be aware that predictions are probabilistic and may have uncertainty
Domain expertise
: Use the model in conjunction with domain expertise in immunology and T-cell biology
Version tracking
: Keep track of which model version and checkpoint you are using
We are committed to advancing the responsible development and use of artificial intelligence. Please follow our
Acceptable Use Policy
when engaging with our services.
Should you have any security or privacy issues or questions related to the services, please reach out to our team at
[email protected]
or
[email protected]
respectively.
Acknowledgements
This model builds upon:
ESM2
by Meta AI (Facebook Research) for the base protein language model
The broader computational biology and immunology research communities
Special thanks to the developers and contributors of the ESM models and the open-source tools that made this work possible.
Runs of biohub DecoderTCR on huggingface.co
0
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs
More Information About DecoderTCR huggingface.co Model
DecoderTCR huggingface.co is an AI model on huggingface.co that provides DecoderTCR's model effect (), which can be used instantly with this biohub DecoderTCR model. huggingface.co supports a free trial of the DecoderTCR model, and also provides paid use of the DecoderTCR. Support call DecoderTCR model through api, including Node.js, Python, http.
DecoderTCR huggingface.co is an online trial and call api platform, which integrates DecoderTCR's modeling effects, including api services, and provides a free online trial of DecoderTCR, you can try DecoderTCR online for free by clicking the link below.
biohub DecoderTCR online free url in huggingface.co:
DecoderTCR is an open source model from GitHub that offers a free installation service, and any user can find DecoderTCR on GitHub to install. At the same time, huggingface.co provides the effect of DecoderTCR install, users can directly use DecoderTCR installed effect in huggingface.co for debugging and trial. It also supports api for free installation.