This is a set of 3 Nano
BERT
models with a modified embeddings layer. The embeddings layer is the same BERT vocabulary (30,522 tokens) projected to a smaller dimensional space then re-encoded to the hidden size. This method is inspired by
MUVERA: Multi-Vector Retrieval via Fixed Dimensional Encodings
.
The number of projections is like a hash. Setting the projections parameter to 5 is like generating a 160-bit hash (5 x float32) for each token. That hash is then projected to the hidden size.
This significantly reduces the number of parameters necessary for token embeddings.
For example:
Standard token embeddings:
30,522 (vocab size) x 768 (hidden size) = 23,440,896 parameters
23,440,896 x 4 (float32) = 93,763,584 bytes
Hash token embeddings:
30,522 (vocab size) x 5 (hash buckets) + 5 x 768 (projection matrix)= 156,450 parameters
These models can be loaded using Hugging Face Transformers as follows. Note that given that this is a custom architecture,
trust_remote_code
needs to be set.
from transformers import AutoModel
model = AutoModel.from_pretrained("neuml/bert-hash-femto", trust_remote_code=True)
Training
Training your own Nano model is simple. All you need is a Hugging Face dataset and the code below using
txtai
.
from datasets import concatenate_datasets, load_dataset
from transformers import AutoTokenizer
from txtai.pipeline import HFTrainer
from configuration_bert_hash import *
from modeling_bert_hash import *
dataset = load_dataset("path to target HF dataset")
tokenizer = AutoTokenizer.from_pretrained("bert-base-uncased")
config = BertHashConfig(
hidden_size=128,
num_hidden_layers=2,
num_attention_heads=2,
intermediate_size=512,
projections=16
)
model = BertHashForMaskedLM(config)
print(config)
print("Total parameters:", sum(p.numel() for p in model.bert.parameters()))
train = HFTrainer()
# Train using MLM
train((model, tokenizer), dataset, task="language-modeling", output_dir="model",
fp16=True, learning_rate=1e-3, per_device_train_batch_size=64, num_train_epochs=3,
warmup_steps=2500, weight_decay=0.01, adam_epsilon=1e-6,
tokenizers=True, dataloader_num_workers=20,
save_strategy="steps", save_steps=5000, logging_steps=500,
)
Future Work
This model demonstrates that smaller models can still be productive models.
The hope is that this work opens the door to many in building small encoder models that pack a punch. Models can be trained in a matter of hours using consumer GPUs.
Imagine more specialized models like this for medical, legal, science and more.
Runs of NeuML bert-hash-femto on huggingface.co
97
Total runs
0
24-hour runs
9
3-day runs
11
7-day runs
79
30-day runs
More Information About bert-hash-femto huggingface.co Model
bert-hash-femto huggingface.co is an AI model on huggingface.co that provides bert-hash-femto's model effect (), which can be used instantly with this NeuML bert-hash-femto model. huggingface.co supports a free trial of the bert-hash-femto model, and also provides paid use of the bert-hash-femto. Support call bert-hash-femto model through api, including Node.js, Python, http.
bert-hash-femto huggingface.co is an online trial and call api platform, which integrates bert-hash-femto's modeling effects, including api services, and provides a free online trial of bert-hash-femto, you can try bert-hash-femto online for free by clicking the link below.
NeuML bert-hash-femto online free url in huggingface.co:
bert-hash-femto is an open source model from GitHub that offers a free installation service, and any user can find bert-hash-femto on GitHub to install. At the same time, huggingface.co provides the effect of bert-hash-femto install, users can directly use bert-hash-femto installed effect in huggingface.co for debugging and trial. It also supports api for free installation.