Introduction of GENERanno-prokaryote-0.5b-cds-annotator
Model Details of GENERanno-prokaryote-0.5b-cds-annotator
GENERanno-prokaryote-0.5b-cds-annotator model
Abouts
In this repository, we present GENERanno-cds-annotator, which is meticulously finetuned on
GENERanno-prokaryote-0.5b-base
for metagenomic annotation tasks. Through comprehensive evaluations, GENERanno-cds-annotator achieves superior accuracy compared to traditional HMM-based methods (e.g.,
GLIMMER3
,
GeneMarkS2
,
Prodigal
) and recent LLM-based approaches (e.g.,
GeneLM
), while demonstrating exceptional generalization ability on archaeal genomes. The detailed annotation results are provided
here
.
How to use
Simple example: CDS annotation
import torch
from transformers import AutoTokenizer, AutoModelForTokenClassification
# Load the tokenizer and model
model_name = "GenerTeam/GENERanno-prokaryote-0.5b-cds-annotator"
tokenizer = AutoTokenizer.from_pretrained(model_name, trust_remote_code=True)
model = AutoModelForTokenClassification.from_pretrained(model_name, trust_remote_code=True)
model.eval() # Set the model to evaluation mode# Prepare the input sequence. Let's use a sample sequence.
sequence = "ATGAGGTGGCAAGAAATGGGCTACGAATTCCATGAGGCTATAGAATAATCTAAGAGAAAT"# Tokenize the sequence. It's crucial to use `add_special_tokens=False`.
inputs = tokenizer(sequence, add_special_tokens=False, return_tensors="pt")
input_ids = inputs["input_ids"]
# The number of tokens should be equal to the length of the sequence.
sequence_length = len(sequence)
assert sequence_length == input_ids.shape[1]
with torch.inference_mode():
logits = model(**inputs).logits
raw_predictions = logits.argmax(dim=-1).cpu()
# Post-process the predictions# This model features multiple prediction heads (for the positive and negative strands).# The predictions for all heads are concatenated into a single output tensor.# Get the model's configuration for processing the output
id2label = model.config.id2label
num_heads = getattr(model, "num_prediction_heads", 1) # Defaults to 1 if not specified# Define the mapping from model labels to annotation characters:# "CDS" -> "+" (Indicates a Coding DNA Sequence)# "NON_CODING" -> "-" (Indicates a non-coding region)
label_to_char = {"CDS": "+", "NON_CODING": "-"}
print(f"Model has {num_heads} prediction head(s). Processing results...")
# The `raw_predictions` tensor has a shape of (batch_size, sequence_length * num_heads).# We need to de-concatenate the predictions for each head.
all_head_annotations = []
preds_for_sequence = raw_predictions[0] # Get predictions for the first sequence in the batchfor h inrange(num_heads):
# Extract the slice of predictions corresponding to the current head
start_idx = h * sequence_length
end_idx = (h + 1) * sequence_length
head_preds_ids = preds_for_sequence[start_idx:end_idx]
# Map the numeric prediction IDs to their string labels (e.g., 1 -> 'CDS')
head_preds_labels = [id2label[pred_id.item()] for pred_id in head_preds_ids]
# Convert the string labels into the final annotation string (e.g., '+'/'-')
annotation_string = "".join([label_to_char[label] for label in head_preds_labels])
all_head_annotations.append(annotation_string)
# Display the final annotations# For this model, the two heads correspond to the positive and negative DNA strands.
head_names = ["Positive Strand", "Negative Strand"] if num_heads == 2else [f"Head {i+1}"for i inrange(num_heads)]
print("\n--- Annotation Results ---")
print(f"Input Sequence: {sequence}")
for i, annotation inenumerate(all_head_annotations):
print(f"Annotation ({head_names[i]}): {annotation}")
print("--------------------------\n")
# How to interpret the output:# - A '+' at a position for the "Positive Strand" annotation means the model predicts that base# is part of a coding sequence on the forward (5' to 3') strand.# - A '+' at a position for the "Negative Strand" annotation means the model predicts that base# is part of a coding sequence on the reverse complementary strand.# - A '-' indicates a non-coding region for that respective strand.
Citation
@article{li2025generanno,
author = {Li, Qiuyi and Wu, Wei and Zhu, Yiheng and Feng, Fuli and Ye, Jieping and Wang, Zheng},
title = {GENERanno: A Genomic Foundation Model for Metagenomic Annotation},
elocation-id = {2025.06.04.656517},
year = {2025},
doi = {10.1101/2025.06.04.656517},
publisher = {Cold Spring Harbor Laboratory},
URL = {https://www.biorxiv.org/content/early/2025/06/05/2025.06.04.656517},
journal = {bioRxiv}
}
Runs of GenerTeam GENERanno-prokaryote-0.5b-cds-annotator on huggingface.co
114
Total runs
1
24-hour runs
3
3-day runs
39
7-day runs
58
30-day runs
More Information About GENERanno-prokaryote-0.5b-cds-annotator huggingface.co Model
More GENERanno-prokaryote-0.5b-cds-annotator license Visit here:
GENERanno-prokaryote-0.5b-cds-annotator huggingface.co is an AI model on huggingface.co that provides GENERanno-prokaryote-0.5b-cds-annotator's model effect (), which can be used instantly with this GenerTeam GENERanno-prokaryote-0.5b-cds-annotator model. huggingface.co supports a free trial of the GENERanno-prokaryote-0.5b-cds-annotator model, and also provides paid use of the GENERanno-prokaryote-0.5b-cds-annotator. Support call GENERanno-prokaryote-0.5b-cds-annotator model through api, including Node.js, Python, http.
GENERanno-prokaryote-0.5b-cds-annotator huggingface.co is an online trial and call api platform, which integrates GENERanno-prokaryote-0.5b-cds-annotator's modeling effects, including api services, and provides a free online trial of GENERanno-prokaryote-0.5b-cds-annotator, you can try GENERanno-prokaryote-0.5b-cds-annotator online for free by clicking the link below.
GenerTeam GENERanno-prokaryote-0.5b-cds-annotator online free url in huggingface.co:
GENERanno-prokaryote-0.5b-cds-annotator is an open source model from GitHub that offers a free installation service, and any user can find GENERanno-prokaryote-0.5b-cds-annotator on GitHub to install. At the same time, huggingface.co provides the effect of GENERanno-prokaryote-0.5b-cds-annotator install, users can directly use GENERanno-prokaryote-0.5b-cds-annotator installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
GENERanno-prokaryote-0.5b-cds-annotator install url in huggingface.co: