This is a
9-layer decoder-only Transformer
trained in the
fully classical way
:
monolithic end-to-end training from scratch
, with
no constructive / layer-wise growth
and
no frozen embeddings
. The token embedding matrix is
randomly initialized and fully trainable
, so semantic structure can be learned directly in the input embeddings (as in standard GPT-like training).
This repo exists as a
clean baseline
for controlled comparisons against the constructive-growth models and against models with frozen embedding substrates.
monolithic end-to-end training vs constructive (layer-wise) growth
the role of embedding learning vs frozen embedding substrates
controlled comparisons across embedding types (trainable vs UNICODE frozen vs 16-bit frozen)
Not intended as a general-purpose assistant model. Outputs may be unreliable and the model may reflect biases present in the training data.
How to use (Transformers)
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
tokenizer = AutoTokenizer.from_pretrained("Bochkov/growing-transformers-model-unfrozen-baseline-monolyth-247m")
model = AutoModelForCausalLM.from_pretrained("Bochkov/growing-transformers-model-unfrozen-baseline-monolyth-247m", trust_remote_code=True).to('cuda')
inputs = torch.tensor([tokenizer.encode("Write a short poem about the ocean. ")], dtype=torch.long, device='cuda')
outputs = model.generate(
inputs,
max_new_tokens=50,
do_sample=False
)
print(tokenizer.decode(outputs[0].tolist()))
#Write a short poem about the ocean. The poem is a poem about the sea and the sea and the sea and the sea. The poem is ab
inputs = torch.tensor([tokenizer.encode("Question: What is the capital of India?\nAnswer:")], dtype=torch.long, device='cuda')
outputs = model.generate(
inputs,
max_new_tokens=10,
do_sample=False
)
print(tokenizer.decode(outputs[0].tolist()))
#Question: What is the capital of India?#Answer:Chennai# </s><
🧑🔬 Citation & Concept
If you use this model or the underlying concepts in your research, please cite our work:
@article{
bochkov2025emergent,
title={Emergent Semantics Beyond Token Embeddings: Transformer {LM}s with Frozen Visual Unicode Representations},
author={Andrey Bochkov},
journal={Transactions on Machine Learning Research},
issn={2835-8856},
year={2025},
url={https://openreview.net/forum?id=Odh8IynO1o},
note={}
}
@misc{bochkov2025growingtransformersmodularcomposition,
title={Growing Transformers: Modular Composition and Layer-wise Expansion on a Frozen Substrate},
author={A. Bochkov},
year={2025},
eprint={2507.07129},
archivePrefix={arXiv},
primaryClass={cs.LG},
url={https://arxiv.org/abs/2507.07129},
}
Runs of Bochkov growing-transformers-model-unfrozen-baseline-monolyth-247m on huggingface.co
20
Total runs
0
24-hour runs
0
3-day runs
-2
7-day runs
5
30-day runs
More Information About growing-transformers-model-unfrozen-baseline-monolyth-247m huggingface.co Model
More growing-transformers-model-unfrozen-baseline-monolyth-247m license Visit here:
growing-transformers-model-unfrozen-baseline-monolyth-247m huggingface.co is an AI model on huggingface.co that provides growing-transformers-model-unfrozen-baseline-monolyth-247m's model effect (), which can be used instantly with this Bochkov growing-transformers-model-unfrozen-baseline-monolyth-247m model. huggingface.co supports a free trial of the growing-transformers-model-unfrozen-baseline-monolyth-247m model, and also provides paid use of the growing-transformers-model-unfrozen-baseline-monolyth-247m. Support call growing-transformers-model-unfrozen-baseline-monolyth-247m model through api, including Node.js, Python, http.
growing-transformers-model-unfrozen-baseline-monolyth-247m huggingface.co is an online trial and call api platform, which integrates growing-transformers-model-unfrozen-baseline-monolyth-247m's modeling effects, including api services, and provides a free online trial of growing-transformers-model-unfrozen-baseline-monolyth-247m, you can try growing-transformers-model-unfrozen-baseline-monolyth-247m online for free by clicking the link below.
Bochkov growing-transformers-model-unfrozen-baseline-monolyth-247m online free url in huggingface.co:
growing-transformers-model-unfrozen-baseline-monolyth-247m is an open source model from GitHub that offers a free installation service, and any user can find growing-transformers-model-unfrozen-baseline-monolyth-247m on GitHub to install. At the same time, huggingface.co provides the effect of growing-transformers-model-unfrozen-baseline-monolyth-247m install, users can directly use growing-transformers-model-unfrozen-baseline-monolyth-247m installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
growing-transformers-model-unfrozen-baseline-monolyth-247m install url in huggingface.co: