GenerTeam / GENERator-eukaryote-3b-base

huggingface.co
Total runs: 711
24-hour runs: 0
7-day runs: 243
30-day runs: 572
Model's Last Updated: June 16 2026
text-generation

Introduction of GENERator-eukaryote-3b-base

Model Details of GENERator-eukaryote-3b-base

GENERator-eukaryote-3b-base model

Abouts

In this repository, we present GENERator, a generative genomic foundation model featuring a context length of 98k base pairs and 3B parameters, trained on an expansive dataset comprising 386 billion base pairs of eukaryotic DNA. The extensive and diverse pre-training data endow the GENERator with enhanced understanding and generation capabilities across various organisms.

For more technical details, please refer to our paper GENERator: A Long-Context Generative Genomic Foundation Model .

How to use
Simple example1: generation

import torch
from transformers import AutoTokenizer, AutoModelForCausalLM

# Load the tokenizer and model.
tokenizer = AutoTokenizer.from_pretrained("GENERator-eukaryote-3b-base", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained("GENERator-eukaryote-3b-base")
config = model.config

max_length = config.max_position_embeddings

# Define input sequences.
sequences = [
    "ATGAGGTGGCAAGAAATGGGCTAC",
    "GAATTCCATGAGGCTATAGAATAATCTAAGAGAAAT"
]

# Process the sequences
sequences = [tokenizer.bos_token + sequence for sequence in sequences]

# Tokenize the sequences
tokenizer.padding_side = "left"
inputs = tokenizer(
    sequences,
    add_special_tokens=False,
    return_tensors="pt",
    padding=True,
    truncation=True,
    max_length=max_length
)

# Generate the sequences
with torch.inference_mode():
    outputs = model.generate(**inputs, max_new_tokens=32)

# Decode the generated sequences
decoded_sequences = tokenizer.batch_decode(outputs, skip_special_tokens=True)

# Print the decoded sequences
print(decoded_sequences)
Simple example2: embedding

import torch
from transformers import AutoTokenizer, AutoModelForCausalLM

# Load the tokenizer and model.
tokenizer = AutoTokenizer.from_pretrained("GENERator-eukaryote-3b-base", trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained("GENERator-eukaryote-3b-base")

config = model.config
max_length = config.max_position_embeddings

# Define input sequences.
sequences = [
    "ATGAGGTGGCAAGAAATGGGCTAC",
    "GAATTCCATGAGGCTATAGAATAATCTAAGAGAAAT"
]

# Tokenize the sequences with add_special_tokens=True to automatically add special tokens,
# such as the BOS EOS token, at the appropriate positions.
tokenizer.padding_side = "right"
inputs = tokenizer(
    sequences,
    add_special_tokens=True,
    return_tensors="pt",
    padding=True,
    truncation=True,
    max_length=max_length
)

# Perform a forward pass through the model to obtain the outputs, including hidden states.
with torch.inference_mode():
    outputs = model(**inputs, output_hidden_states=True)

# Retrieve the hidden states from the last layer.
hidden_states = outputs.hidden_states[-1]  # Shape: (batch_size, sequence_length, hidden_size)

# Use the attention_mask to determine the index of the last token in each sequence.
# Since add_special_tokens=True is used, the last token is typically the EOS token.
attention_mask = inputs["attention_mask"]
last_token_indices = attention_mask.sum(dim=1) - 1  # Index of the last token for each sequence

# Extract the embedding corresponding to the EOS token for each sequence.
seq_embeddings = []
for i, token_index in enumerate(last_token_indices):
    # Fetch the embedding for the last token (EOS token).
    seq_embedding = hidden_states[i, token_index, :]
    seq_embeddings.append(seq_embedding)

# Stack the embeddings into a tensor with shape (batch_size, hidden_size)
seq_embeddings = torch.stack(seq_embeddings)

print("Sequence Embeddings:", seq_embeddings)
Citation

TBD

Runs of GenerTeam GENERator-eukaryote-3b-base on huggingface.co

711
Total runs
0
24-hour runs
163
3-day runs
243
7-day runs
572
30-day runs

More Information About GENERator-eukaryote-3b-base huggingface.co Model

More GENERator-eukaryote-3b-base license Visit here:

https://choosealicense.com/licenses/mit

GENERator-eukaryote-3b-base huggingface.co

GENERator-eukaryote-3b-base huggingface.co is an AI model on huggingface.co that provides GENERator-eukaryote-3b-base's model effect (), which can be used instantly with this GenerTeam GENERator-eukaryote-3b-base model. huggingface.co supports a free trial of the GENERator-eukaryote-3b-base model, and also provides paid use of the GENERator-eukaryote-3b-base. Support call GENERator-eukaryote-3b-base model through api, including Node.js, Python, http.

GENERator-eukaryote-3b-base huggingface.co Url

https://huggingface.co/GenerTeam/GENERator-eukaryote-3b-base

GenerTeam GENERator-eukaryote-3b-base online free

GENERator-eukaryote-3b-base huggingface.co is an online trial and call api platform, which integrates GENERator-eukaryote-3b-base's modeling effects, including api services, and provides a free online trial of GENERator-eukaryote-3b-base, you can try GENERator-eukaryote-3b-base online for free by clicking the link below.

GenerTeam GENERator-eukaryote-3b-base online free url in huggingface.co:

https://huggingface.co/GenerTeam/GENERator-eukaryote-3b-base

GENERator-eukaryote-3b-base install

GENERator-eukaryote-3b-base is an open source model from GitHub that offers a free installation service, and any user can find GENERator-eukaryote-3b-base on GitHub to install. At the same time, huggingface.co provides the effect of GENERator-eukaryote-3b-base install, users can directly use GENERator-eukaryote-3b-base installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

GENERator-eukaryote-3b-base install url in huggingface.co:

https://huggingface.co/GenerTeam/GENERator-eukaryote-3b-base

Url of GENERator-eukaryote-3b-base

GENERator-eukaryote-3b-base huggingface.co Url

Provider of GENERator-eukaryote-3b-base huggingface.co

GenerTeam
ORGANIZATIONS

Other API from GenerTeam