state-spaces / mamba-2.8b-hf

huggingface.co
Total runs: 8.6K
24-hour runs: 0
7-day runs: -61
30-day runs: -2.6K
Model's Last Updated: March 06 2024
text-generation

Introduction of mamba-2.8b-hf

Model Details of mamba-2.8b-hf

Mamba

This repository contains the transfromers compatible mamba-2.8b . The checkpoints are untouched, but the full config.json and tokenizer are pushed to this repo.

Usage

You need to install transformers from main until transformers=4.39.0 is released.

pip install git+https://github.com/huggingface/transformers@main

We also recommend you to install both causal_conv_1d and mamba-ssm using:

pip install causal-conv1d>=1.2.0
pip install mamba-ssm

If any of these two is not installed, the "eager" implementation will be used. Otherwise the more optimised cuda kernels will be used.

Generation

You can use the classic generate API:

>>> from transformers import MambaConfig, MambaForCausalLM, AutoTokenizer
>>> import torch

>>> tokenizer = AutoTokenizer.from_pretrained("state-spaces/mamba-2.8b-hf")
>>> model = MambaForCausalLM.from_pretrained("state-spaces/mamba-2.8b-hf")
>>> input_ids = tokenizer("Hey how are you doing?", return_tensors="pt")["input_ids"]

>>> out = model.generate(input_ids, max_new_tokens=10)
>>> print(tokenizer.batch_decode(out))
["Hey how are you doing?\n\nI'm doing great.\n\nI"]
PEFT finetuning example

In order to finetune using the peft library, we recommend keeping the model in float32!

from datasets import load_dataset
from trl import SFTTrainer
from peft import LoraConfig
from transformers import AutoTokenizer, AutoModelForCausalLM, TrainingArguments
tokenizer = AutoTokenizer.from_pretrained("state-spaces/mamba-2.8b-hf")
model = AutoModelForCausalLM.from_pretrained("state-spaces/mamba-2.8b-hf")
dataset = load_dataset("Abirate/english_quotes", split="train")
training_args = TrainingArguments(
    output_dir="./results",
    num_train_epochs=3,
    per_device_train_batch_size=4,
    logging_dir='./logs',
    logging_steps=10,
    learning_rate=2e-3
)
lora_config =  LoraConfig(
        r=8,
        target_modules=["x_proj", "embeddings", "in_proj", "out_proj"],
        task_type="CAUSAL_LM",
        bias="none"
)
trainer = SFTTrainer(
    model=model,
    tokenizer=tokenizer,
    args=training_args,
    peft_config=lora_config,
    train_dataset=dataset,
    dataset_text_field="quote",
)
trainer.train()

Runs of state-spaces mamba-2.8b-hf on huggingface.co

8.6K
Total runs
0
24-hour runs
-42
3-day runs
-61
7-day runs
-2.6K
30-day runs

More Information About mamba-2.8b-hf huggingface.co Model

mamba-2.8b-hf huggingface.co

mamba-2.8b-hf huggingface.co is an AI model on huggingface.co that provides mamba-2.8b-hf's model effect (), which can be used instantly with this state-spaces mamba-2.8b-hf model. huggingface.co supports a free trial of the mamba-2.8b-hf model, and also provides paid use of the mamba-2.8b-hf. Support call mamba-2.8b-hf model through api, including Node.js, Python, http.

state-spaces mamba-2.8b-hf online free

mamba-2.8b-hf huggingface.co is an online trial and call api platform, which integrates mamba-2.8b-hf's modeling effects, including api services, and provides a free online trial of mamba-2.8b-hf, you can try mamba-2.8b-hf online for free by clicking the link below.

state-spaces mamba-2.8b-hf online free url in huggingface.co:

https://huggingface.co/state-spaces/mamba-2.8b-hf

mamba-2.8b-hf install

mamba-2.8b-hf is an open source model from GitHub that offers a free installation service, and any user can find mamba-2.8b-hf on GitHub to install. At the same time, huggingface.co provides the effect of mamba-2.8b-hf install, users can directly use mamba-2.8b-hf installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

mamba-2.8b-hf install url in huggingface.co:

https://huggingface.co/state-spaces/mamba-2.8b-hf

Url of mamba-2.8b-hf

Provider of mamba-2.8b-hf huggingface.co

state-spaces
ORGANIZATIONS

Other API from state-spaces