A 300M-parameter language model trained from scratch on
FineWeb-Edu 10BT
(~9.4B tokens, 1 epoch) as part of the
Convergent Evolution
project, which investigates how Fourier features emerge in LLM number embeddings.
Intermediate checkpoints are saved as branches:
tokens-200M
,
tokens-400M
, ...,
tokens-9.6B
.
from transformers import AutoModelForCausalLM
# Load final checkpoint
model = AutoModelForCausalLM.from_pretrained("deqing/convergent-llama-300M-muon-window_4")
# Load intermediate checkpoint (e.g., at 1B tokens)
model = AutoModelForCausalLM.from_pretrained("deqing/convergent-llama-300M-muon-window_4", revision="tokens-1B")
Citation
Paper forthcoming.
Runs of deqing llama-window-4-old on huggingface.co
7
Total runs
1
24-hour runs
1
3-day runs
1
7-day runs
-3
30-day runs
More Information About llama-window-4-old huggingface.co Model
llama-window-4-old huggingface.co is an AI model on huggingface.co that provides llama-window-4-old's model effect (), which can be used instantly with this deqing llama-window-4-old model. huggingface.co supports a free trial of the llama-window-4-old model, and also provides paid use of the llama-window-4-old. Support call llama-window-4-old model through api, including Node.js, Python, http.
llama-window-4-old huggingface.co is an online trial and call api platform, which integrates llama-window-4-old's modeling effects, including api services, and provides a free online trial of llama-window-4-old, you can try llama-window-4-old online for free by clicking the link below.
deqing llama-window-4-old online free url in huggingface.co:
llama-window-4-old is an open source model from GitHub that offers a free installation service, and any user can find llama-window-4-old on GitHub to install. At the same time, huggingface.co provides the effect of llama-window-4-old install, users can directly use llama-window-4-old installed effect in huggingface.co for debugging and trial. It also supports api for free installation.