This checkpoint is designed to study the effect of
normalization / PCA-style processing
in a
minimal
frozen embedding setting.
Unlike
Model_UNI_GLYPH
, this model does
not
use glyph-based embeddings. Instead, it uses a
frozen 16-dimensional float embedding
per token.
Key idea (what this ablation tests)
This model isolates the impact of having
float
frozen embeddings (with
PCA + normalization
) versus the strictly
binary token-ID
variant (
Model_16_BIT
):
n_embed = 16
per token (
float components
, not binary)
Embedding vectors are
precomputed
(PCA + L2 normalization) and then
frozen
The embedding layer is never updated (
requires_grad=False
)
To match the Transformer hidden size, the 16-dim embedding is expanded to 1024 via a
non-trainable repetition
:
repeat_interleave(64)
→
16 * 64 = 1024
This lets you test whether the model’s behavior changes when the frozen token “identifier” is:
You may load the tokenizer either from this model repo (if included) or from the standalone tokenizer repo. The key requirement is
exact vocab alignment
.
How to use (Transformers)
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
tokenizer = AutoTokenizer.from_pretrained("Bochkov/emergent-semantics-model-16-float-269m")
model = AutoModelForCausalLM.from_pretrained("Bochkov/emergent-semantics-model-16-float-269m", trust_remote_code=True).to('cuda')
inputs = torch.tensor([tokenizer.encode("Question: What is the capital of Japan?\nAnswer:")], dtype=torch.long, device='cuda')
outputs = model.generate(
inputs,
max_new_tokens=10,
do_sample=False
)
print(tokenizer.decode(outputs[0].tolist()))
#Question: What is the capital of Japan?#Answer:A temperature in
Intended use
Research only, especially for:
Comparing
Model_16_FLOAT
vs
Model_16_BIT
(effect of continuous normalized vectors vs binary ID)
Comparing
Model_16_FLOAT
vs
Model_UNI_GLYPH
(effect of glyph-derived structure vs minimal vectors)
Studying emergent semantics when embeddings are
frozen and non-semantic
emergent-semantics-model-16-float-269m huggingface.co is an AI model on huggingface.co that provides emergent-semantics-model-16-float-269m's model effect (), which can be used instantly with this Bochkov emergent-semantics-model-16-float-269m model. huggingface.co supports a free trial of the emergent-semantics-model-16-float-269m model, and also provides paid use of the emergent-semantics-model-16-float-269m. Support call emergent-semantics-model-16-float-269m model through api, including Node.js, Python, http.
emergent-semantics-model-16-float-269m huggingface.co is an online trial and call api platform, which integrates emergent-semantics-model-16-float-269m's modeling effects, including api services, and provides a free online trial of emergent-semantics-model-16-float-269m, you can try emergent-semantics-model-16-float-269m online for free by clicking the link below.
Bochkov emergent-semantics-model-16-float-269m online free url in huggingface.co:
emergent-semantics-model-16-float-269m is an open source model from GitHub that offers a free installation service, and any user can find emergent-semantics-model-16-float-269m on GitHub to install. At the same time, huggingface.co provides the effect of emergent-semantics-model-16-float-269m install, users can directly use emergent-semantics-model-16-float-269m installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
emergent-semantics-model-16-float-269m install url in huggingface.co: