Bochkov / growing-transformers-model-unfrozen-baseline-monolyth-247m

huggingface.co
Total runs: 20
24-hour runs: 0
7-day runs: -2
30-day runs: 5
Model's Last Updated: January 10 2026
text-generation

Introduction of growing-transformers-model-unfrozen-baseline-monolyth-247m

Model Details of growing-transformers-model-unfrozen-baseline-monolyth-247m

Growing Transformers — Unfrozen Baseline (Monolithic, 247M)

This repository contains growing-transformers-model-unfrozen-baseline-monolyth-247m , a classic monolithic baseline model from the paper:

📚 Paper (Growing Transformers: Modular Composition and Layer-wise Expansion on a Frozen Substrate) -

📚 Paper (Emergent Semantics Beyond Token Embeddings: Transformer LMs with Frozen Visual Unicode Representations) -

It is part of the comparative-study collection:
https://huggingface.co/collections/Bochkov/growing-transformers-layer-wise-expansion-comparative-study

Code:
https://github.com/AVBochkov/PGT


What this model is (in one paragraph)

This is a 9-layer decoder-only Transformer trained in the fully classical way : monolithic end-to-end training from scratch , with no constructive / layer-wise growth and no frozen embeddings . The token embedding matrix is randomly initialized and fully trainable , so semantic structure can be learned directly in the input embeddings (as in standard GPT-like training).

This repo exists as a clean baseline for controlled comparisons against the constructive-growth models and against models with frozen embedding substrates.


Primary comparison (why this repo exists)

This model is intended to be compared to:

  • Bochkov/growing-transformers-model-16-bit-1-9-181m
    (constructive, layer-wise growth; frozen 16-bit embeddings)
What is identical
  • Same controlled-study Transformer stack architecture ( 9 layers , d_model=1024 , n_head=32 )
  • Same tokenizer family / vocabulary size ( 65,536 )
  • Same context length used in training ( 1024 )
What differs
  • Training procedure
    • This repo: monolithic end-to-end training (all layers trained together from scratch)
    • 16-bit constructive repo: trained in stages (1–3, then 4–6, then 7–9), freezing previously trained layers
  • Embedding layer
    • This repo: standard trainable embedding matrix ( vocab_size × d_model )
    • 16-bit repo: extremely small frozen embedding substrate (16-dim binary signal expanded to d_model )
Important note on parameter count (why this model is larger)

This model has more parameters than the 16-bit models because it includes a full trainable embedding matrix:

  • Trainable embedding size here: 65,536 × 1,024 ≈ 67.1M parameters
  • In the 16-bit setup, the embedding-related parameters are ~ 1.0M (as reported in the paper)

So even with the same Transformer block stack , the total parameter count differs primarily due to the embedding matrix.


Model architecture (controlled study)
  • Type: decoder-only Transformer (GPT-like)
  • Layers: 9
  • Hidden size: d_model = 1024
  • Heads: n_head = 32
  • Vocabulary size: 65,536
  • Context length used in training: 1024
  • Embedding: standard trainable token embedding matrix ( vocab_size × d_model )
Parameter count
  • Total: ≈247.6M
  • Trainable: ≈247.6M
  • Frozen: 0.0M

(Counts follow the paper’s controlled-study table.)


Tokenizer

Canonical tokenizer repository:


Intended use

Research / analysis of:

  • monolithic end-to-end training vs constructive (layer-wise) growth
  • the role of embedding learning vs frozen embedding substrates
  • controlled comparisons across embedding types (trainable vs UNICODE frozen vs 16-bit frozen)

Not intended as a general-purpose assistant model. Outputs may be unreliable and the model may reflect biases present in the training data.


How to use (Transformers)

import torch
from transformers import AutoTokenizer, AutoModelForCausalLM

tokenizer = AutoTokenizer.from_pretrained("Bochkov/growing-transformers-model-unfrozen-baseline-monolyth-247m")
model = AutoModelForCausalLM.from_pretrained("Bochkov/growing-transformers-model-unfrozen-baseline-monolyth-247m", trust_remote_code=True).to('cuda')

inputs = torch.tensor([tokenizer.encode("Write a short poem about the ocean. ")], dtype=torch.long, device='cuda')

outputs = model.generate(
    inputs, 
    max_new_tokens=50,
    do_sample=False
)
print(tokenizer.decode(outputs[0].tolist()))
#Write a short poem about the ocean. The poem is a poem about the sea and the sea and the sea and the sea. The poem is ab

inputs = torch.tensor([tokenizer.encode("Question: What is the capital of India?\nAnswer:")], dtype=torch.long, device='cuda')

outputs = model.generate(
    inputs, 
    max_new_tokens=10,
    do_sample=False
)
print(tokenizer.decode(outputs[0].tolist()))
#Question: What is the capital of India?
#Answer:Chennai
#    </s><

🧑‍🔬 Citation & Concept

If you use this model or the underlying concepts in your research, please cite our work:

@article{
      bochkov2025emergent,
      title={Emergent Semantics Beyond Token Embeddings: Transformer {LM}s with Frozen Visual Unicode Representations},
      author={Andrey Bochkov},
      journal={Transactions on Machine Learning Research},
      issn={2835-8856},
      year={2025},
      url={https://openreview.net/forum?id=Odh8IynO1o},
      note={}
}

@misc{bochkov2025growingtransformersmodularcomposition,
      title={Growing Transformers: Modular Composition and Layer-wise Expansion on a Frozen Substrate}, 
      author={A. Bochkov},
      year={2025},
      eprint={2507.07129},
      archivePrefix={arXiv},
      primaryClass={cs.LG},
      url={https://arxiv.org/abs/2507.07129}, 
}

Runs of Bochkov growing-transformers-model-unfrozen-baseline-monolyth-247m on huggingface.co

20
Total runs
0
24-hour runs
0
3-day runs
-2
7-day runs
5
30-day runs

More Information About growing-transformers-model-unfrozen-baseline-monolyth-247m huggingface.co Model

More growing-transformers-model-unfrozen-baseline-monolyth-247m license Visit here:

https://choosealicense.com/licenses/apache-2.0

growing-transformers-model-unfrozen-baseline-monolyth-247m huggingface.co

growing-transformers-model-unfrozen-baseline-monolyth-247m huggingface.co is an AI model on huggingface.co that provides growing-transformers-model-unfrozen-baseline-monolyth-247m's model effect (), which can be used instantly with this Bochkov growing-transformers-model-unfrozen-baseline-monolyth-247m model. huggingface.co supports a free trial of the growing-transformers-model-unfrozen-baseline-monolyth-247m model, and also provides paid use of the growing-transformers-model-unfrozen-baseline-monolyth-247m. Support call growing-transformers-model-unfrozen-baseline-monolyth-247m model through api, including Node.js, Python, http.

growing-transformers-model-unfrozen-baseline-monolyth-247m huggingface.co Url

https://huggingface.co/Bochkov/growing-transformers-model-unfrozen-baseline-monolyth-247m

Bochkov growing-transformers-model-unfrozen-baseline-monolyth-247m online free

growing-transformers-model-unfrozen-baseline-monolyth-247m huggingface.co is an online trial and call api platform, which integrates growing-transformers-model-unfrozen-baseline-monolyth-247m's modeling effects, including api services, and provides a free online trial of growing-transformers-model-unfrozen-baseline-monolyth-247m, you can try growing-transformers-model-unfrozen-baseline-monolyth-247m online for free by clicking the link below.

Bochkov growing-transformers-model-unfrozen-baseline-monolyth-247m online free url in huggingface.co:

https://huggingface.co/Bochkov/growing-transformers-model-unfrozen-baseline-monolyth-247m

growing-transformers-model-unfrozen-baseline-monolyth-247m install

growing-transformers-model-unfrozen-baseline-monolyth-247m is an open source model from GitHub that offers a free installation service, and any user can find growing-transformers-model-unfrozen-baseline-monolyth-247m on GitHub to install. At the same time, huggingface.co provides the effect of growing-transformers-model-unfrozen-baseline-monolyth-247m install, users can directly use growing-transformers-model-unfrozen-baseline-monolyth-247m installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

growing-transformers-model-unfrozen-baseline-monolyth-247m install url in huggingface.co:

https://huggingface.co/Bochkov/growing-transformers-model-unfrozen-baseline-monolyth-247m

Url of growing-transformers-model-unfrozen-baseline-monolyth-247m

growing-transformers-model-unfrozen-baseline-monolyth-247m huggingface.co Url

Provider of growing-transformers-model-unfrozen-baseline-monolyth-247m huggingface.co

Bochkov
ORGANIZATIONS

Other API from Bochkov

huggingface.co

Total runs: 18
Run Growth: 11
Growth Rate: 57.89%
Updated:January 07 2026
huggingface.co

Total runs: 17
Run Growth: 1
Growth Rate: 5.88%
Updated:July 15 2025
huggingface.co

Total runs: 17
Run Growth: 0
Growth Rate: 0.00%
Updated:July 15 2025
huggingface.co

Total runs: 17
Run Growth: 0
Growth Rate: 0.00%
Updated:July 15 2025
huggingface.co

Total runs: 16
Run Growth: -1
Growth Rate: -6.25%
Updated:July 15 2025
huggingface.co

Total runs: 16
Run Growth: 0
Growth Rate: 0.00%
Updated:July 15 2025
huggingface.co

Total runs: 16
Run Growth: 0
Growth Rate: 0.00%
Updated:July 15 2025
huggingface.co

Total runs: 16
Run Growth: 0
Growth Rate: 0.00%
Updated:July 15 2025
huggingface.co

Total runs: 15
Run Growth: -3
Growth Rate: -20.00%
Updated:July 15 2025
huggingface.co

Total runs: 15
Run Growth: 6
Growth Rate: 40.00%
Updated:October 21 2025
huggingface.co

Total runs: 14
Run Growth: 8
Growth Rate: 53.33%
Updated:October 21 2025
huggingface.co

Total runs: 13
Run Growth: 3
Growth Rate: 23.08%
Updated:October 21 2025
huggingface.co

Total runs: 12
Run Growth: -3
Growth Rate: -25.00%
Updated:January 01 2026
huggingface.co

Total runs: 11
Run Growth: 7
Growth Rate: 63.64%
Updated:October 21 2025
huggingface.co

Total runs: 11
Run Growth: 3
Growth Rate: 27.27%
Updated:October 21 2025
huggingface.co

Total runs: 10
Run Growth: 2
Growth Rate: 20.00%
Updated:October 21 2025
huggingface.co

Total runs: 8
Run Growth: 7
Growth Rate: 87.50%
Updated:January 01 2026
huggingface.co

Total runs: 7
Run Growth: 4
Growth Rate: 57.14%
Updated:October 21 2025
huggingface.co

Total runs: 7
Run Growth: 4
Growth Rate: 57.14%
Updated:October 21 2025
huggingface.co

Total runs: 7
Run Growth: 3
Growth Rate: 42.86%
Updated:October 21 2025
huggingface.co

Total runs: 7
Run Growth: 3
Growth Rate: 42.86%
Updated:October 21 2025
huggingface.co

Total runs: 5
Run Growth: -1
Growth Rate: -20.00%
Updated:October 21 2025
huggingface.co

Total runs: 5
Run Growth: -1
Growth Rate: -20.00%
Updated:October 21 2025
huggingface.co

Total runs: 4
Run Growth: -1
Growth Rate: -25.00%
Updated:January 01 2026
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:October 02 2026