In this repository, we present GENERator, a generative genomic foundation model featuring a context length of 98k base pairs and 1.2B parameters, trained on an expansive dataset comprising 386 billion base pairs of eukaryotic DNA. Our evaluations demonstrate that the GENERator consistently achieves state-of-the-art performance across a wide spectrum of benchmarks, including
Genomic Benchmarks
,
NT tasks
, and our newly proposed
Gener tasks
.
Beyond benchmark performance, the GENERator adheres to the central dogma of molecular biology, accurately generating protein-coding DNA sequences that produce proteins structurally analogous to known families. Moreover, the GENERator showcases significant promise in sequence optimization, particularly in the design of promoter sequences that regulate gene activity during various biological stages, highlighting its potential for a series of biologically significant tasks. Our findings position the GENERator as a vital resource for genomic research and biotechnological advancement.
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
# Load the tokenizer and model.
tokenizer = AutoTokenizer.from_pretrained("GenerTeam/GENERator-eukaryote-1.2b-base", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained("GenerTeam/GENERator-eukaryote-1.2b-base")
config = model.config
max_length = config.max_position_embeddings
# Define input sequences.
sequences = [
"ATGAGGTGGCAAGAAATGGGCTAC",
"GAATTCCATGAGGCTATAGAATAATCTAAGAGAAAT"
]
# Process the sequences
sequences = [tokenizer.bos_token + sequence for sequence in sequences]
# Tokenize the sequences
tokenizer.padding_side = "left"
inputs = tokenizer(
sequences,
add_special_tokens=False,
return_tensors="pt",
padding=True,
truncation=True,
max_length=max_length
)
# Generate the sequenceswith torch.inference_mode():
outputs = model.generate(**inputs, max_new_tokens=32)
# Decode the generated sequences
decoded_sequences = tokenizer.batch_decode(outputs, skip_special_tokens=True)
# Print the decoded sequencesprint(decoded_sequences)
Simple example2: embedding
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
# Load the tokenizer and model.
tokenizer = AutoTokenizer.from_pretrained("GENERator-eukaryote-1.2b-base", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained("GENERator-eukaryote-1.2b-base")
config = model.config
max_length = config.max_position_embeddings
# Define input sequences.
sequences = [
"ATGAGGTGGCAAGAAATGGGCTAC",
"GAATTCCATGAGGCTATAGAATAATCTAAGAGAAAT"
]
# Tokenize the sequences with add_special_tokens=True to automatically add special tokens,# such as the BOS EOS token, at the appropriate positions.
tokenizer.padding_side = "right"
inputs = tokenizer(
sequences,
add_special_tokens=True,
return_tensors="pt",
padding=True,
truncation=True,
max_length=max_length
)
# Perform a forward pass through the model to obtain the outputs, including hidden states.with torch.inference_mode():
outputs = model(**inputs, output_hidden_states=True)
# Retrieve the hidden states from the last layer.
hidden_states = outputs.hidden_states[-1] # Shape: (batch_size, sequence_length, hidden_size)# Use the attention_mask to determine the index of the last token in each sequence.# Since add_special_tokens=True is used, the last token is typically the EOS token.
attention_mask = inputs["attention_mask"]
last_token_indices = attention_mask.sum(dim=1) - 1# Index of the last token for each sequence# Extract the embedding corresponding to the EOS token for each sequence.
seq_embeddings = []
for i, token_index inenumerate(last_token_indices):
# Fetch the embedding for the last token (EOS token).
seq_embedding = hidden_states[i, token_index, :]
seq_embeddings.append(seq_embedding)
# Stack the embeddings into a tensor with shape (batch_size, hidden_size)
seq_embeddings = torch.stack(seq_embeddings)
print("Sequence Embeddings:", seq_embeddings)
Citation
TBD
Runs of GenerTeam GENERator-eukaryote-1.2b-base on huggingface.co
747
Total runs
0
24-hour runs
-9
3-day runs
10
7-day runs
-293
30-day runs
More Information About GENERator-eukaryote-1.2b-base huggingface.co Model
More GENERator-eukaryote-1.2b-base license Visit here:
GENERator-eukaryote-1.2b-base huggingface.co is an AI model on huggingface.co that provides GENERator-eukaryote-1.2b-base's model effect (), which can be used instantly with this GenerTeam GENERator-eukaryote-1.2b-base model. huggingface.co supports a free trial of the GENERator-eukaryote-1.2b-base model, and also provides paid use of the GENERator-eukaryote-1.2b-base. Support call GENERator-eukaryote-1.2b-base model through api, including Node.js, Python, http.
GENERator-eukaryote-1.2b-base huggingface.co is an online trial and call api platform, which integrates GENERator-eukaryote-1.2b-base's modeling effects, including api services, and provides a free online trial of GENERator-eukaryote-1.2b-base, you can try GENERator-eukaryote-1.2b-base online for free by clicking the link below.
GenerTeam GENERator-eukaryote-1.2b-base online free url in huggingface.co:
GENERator-eukaryote-1.2b-base is an open source model from GitHub that offers a free installation service, and any user can find GENERator-eukaryote-1.2b-base on GitHub to install. At the same time, huggingface.co provides the effect of GENERator-eukaryote-1.2b-base install, users can directly use GENERator-eukaryote-1.2b-base installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
GENERator-eukaryote-1.2b-base install url in huggingface.co: