If you are using
GENERator
for sequence generation, please ensure that the length of each input sequence is a multiple of
6
. This can be achieved by either:
Padding the sequence on the left with
'A'
(
left padding
);
Truncating the sequence from the left (
left truncation
).
This requirement arises because
GENERator
employs a 6-mer tokenizer. If the input sequence length is not a multiple of
6
, the tokenizer will append an
'<oov>'
(out-of-vocabulary) token to the end of the token sequence. This can result in uninformative subsequent generations, such as repeated
'AAAAAA'
.
We apologize for any inconvenience this may cause and recommend adhering to the above guidelines to ensure accurate and meaningful generation results.
Abouts
In this repository, we present GENERator-v2, a generative genomic foundation with enhanced performance in eukaryotic domain. More technical details are coming soon...
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
# Load the tokenizer and model.
tokenizer = AutoTokenizer.from_pretrained("GenerTeam/GENERator-eukaryote-1.2b-base", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained("GenerTeam/GENERator-eukaryote-1.2b-base")
config = model.config
max_length = config.max_position_embeddings
# Define input sequences.
sequences = [
"ATGAGGTGGCAAGAAATGGGCTAC",
"GAATTCCATGAGGCTATAGAATAATCTAAGAGAAAT"
]
defleft_padding(sequence, padding_char='A', multiple=6):
remainder = len(sequence) % multiple
if remainder != 0:
padding_length = multiple - remainder
return padding_char * padding_length + sequence
return sequence
defleft_truncation(sequence, multiple=6):
remainder = len(sequence) % multiple
if remainder != 0:
return sequence[remainder:]
return sequence
# Apply left_padding to all sequences# padded_sequences = [left_padding(seq) for seq in sequences]# Apply left_truncation to all sequences
truncated_sequences = [left_truncation(seq) for seq in sequences]
# Process the sequences
sequences = [tokenizer.bos_token + sequence for sequence in truncated_sequences]
# Tokenize the sequences
tokenizer.padding_side = "left"
inputs = tokenizer(
sequences,
add_special_tokens=False,
return_tensors="pt",
padding=True,
truncation=True,
max_length=max_length
)
# Generate the sequenceswith torch.inference_mode():
outputs = model.generate(**inputs, max_new_tokens=32, temperature=0.00001, top_k=1)
# Decode the generated sequences
decoded_sequences = tokenizer.batch_decode(outputs, skip_special_tokens=True)
# Print the decoded sequencesprint(decoded_sequences)
# It is expected to observe non-sense decoded sequences (e.g., 'AAAAAA')# The input sequences are too short to provide sufficient context.
Simple example2: embedding
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
# Load the tokenizer and model.
tokenizer = AutoTokenizer.from_pretrained("GENERator-eukaryote-1.2b-base", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained("GENERator-eukaryote-1.2b-base")
config = model.config
max_length = config.max_position_embeddings
# Define input sequences.
sequences = [
"ATGAGGTGGCAAGAAATGGGCTAC",
"GAATTCCATGAGGCTATAGAATAATCTAAGAGAAAT"
]
# Tokenize the sequences with add_special_tokens=True to automatically add special tokens,# such as the BOS EOS token, at the appropriate positions.
tokenizer.padding_side = "right"
inputs = tokenizer(
sequences,
add_special_tokens=True,
return_tensors="pt",
padding=True,
truncation=True,
max_length=max_length
)
# Perform a forward pass through the model to obtain the outputs, including hidden states.with torch.inference_mode():
outputs = model(**inputs, output_hidden_states=True)
# Retrieve the hidden states from the last layer.
hidden_states = outputs.hidden_states[-1] # Shape: (batch_size, sequence_length, hidden_size)# Use the attention_mask to determine the index of the last token in each sequence.# Since add_special_tokens=True is used, the last token is typically the EOS token.
attention_mask = inputs["attention_mask"]
last_token_indices = attention_mask.sum(dim=1) - 1# Index of the last token for each sequence# Extract the embedding corresponding to the EOS token for each sequence.
seq_embeddings = []
for i, token_index inenumerate(last_token_indices):
# Fetch the embedding for the last token (EOS token).
seq_embedding = hidden_states[i, token_index, :]
seq_embeddings.append(seq_embedding)
# Stack the embeddings into a tensor with shape (batch_size, hidden_size)
seq_embeddings = torch.stack(seq_embeddings)
print("Sequence Embeddings:", seq_embeddings)
Citation
@misc{wu2025generator,
title={GENERator: A Long-Context Generative Genomic Foundation Model},
author={Wei Wu and Qiuyi Li and Mingyang Li and Kun Fu and Fuli Feng and Jieping Ye and Hui Xiong and Zheng Wang},
year={2025},
eprint={2502.07272},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2502.07272},
}
Runs of GenerTeam GENERator-v2-eukaryote-1.2b-base on huggingface.co
569
Total runs
-5
24-hour runs
12
3-day runs
13
7-day runs
-38
30-day runs
More Information About GENERator-v2-eukaryote-1.2b-base huggingface.co Model
More GENERator-v2-eukaryote-1.2b-base license Visit here:
GENERator-v2-eukaryote-1.2b-base huggingface.co is an AI model on huggingface.co that provides GENERator-v2-eukaryote-1.2b-base's model effect (), which can be used instantly with this GenerTeam GENERator-v2-eukaryote-1.2b-base model. huggingface.co supports a free trial of the GENERator-v2-eukaryote-1.2b-base model, and also provides paid use of the GENERator-v2-eukaryote-1.2b-base. Support call GENERator-v2-eukaryote-1.2b-base model through api, including Node.js, Python, http.
GENERator-v2-eukaryote-1.2b-base huggingface.co is an online trial and call api platform, which integrates GENERator-v2-eukaryote-1.2b-base's modeling effects, including api services, and provides a free online trial of GENERator-v2-eukaryote-1.2b-base, you can try GENERator-v2-eukaryote-1.2b-base online for free by clicking the link below.
GenerTeam GENERator-v2-eukaryote-1.2b-base online free url in huggingface.co:
GENERator-v2-eukaryote-1.2b-base is an open source model from GitHub that offers a free installation service, and any user can find GENERator-v2-eukaryote-1.2b-base on GitHub to install. At the same time, huggingface.co provides the effect of GENERator-v2-eukaryote-1.2b-base install, users can directly use GENERator-v2-eukaryote-1.2b-base installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
GENERator-v2-eukaryote-1.2b-base install url in huggingface.co: