1.5B-parameter Llama-2-architecture language models pretrained
from scratch
for the paper
Know Your Sources: Data Selection Matters when Rewriting for Data-Constrained Pretraining
.
This repo holds all
9 checkpoints
for this setting: 3 seeds × 3 epoch boundaries.
⚠️ These are
base models
trained on 30B tokens. They are research artifacts for studying
data selection, not instruction-tuned assistants.
What this setting does
The next 5B tokens down the DCLM fastText ranking after the shared anchor is removed from the pool, used
verbatim
. No rewriting anywhere in this mixture.
This is the control arm every other setting is compared against.
Every one of the six settings trains on a 10B-token mixture built the same way: a
shared 5B-token
anchor
of top-ranked DCLM fastText documents, identical across all six settings and never
rewritten, plus
5B tokens contributed by this setting's selection strategy
. Only that second
half differs between settings, so downstream differences isolate source selection alone.
Quick start
The repo root is a copy of
seed 42 at the end of epoch 3
, so
from_pretrained
works with no
subfolder
argument:
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("blab-jhu/KYS-1.5B-Quality-Base") # = seed42/epoch3
tok = AutoTokenizer.from_pretrained("blab-jhu/KYS-1.5B-Quality-Base")
Any other checkpoint is addressed by subfolder:
model = AutoModelForCausalLM.from_pretrained("blab-jhu/KYS-1.5B-Quality-Base", subfolder="seed43/epoch1")
Layout
. <- copy of seed42/epoch3 (seed 42, end of epoch 3)
seed42/epoch1 epoch2 epoch3
seed43/epoch1 epoch2 epoch3
seed44/epoch1 epoch2 epoch3
Each directory is a complete HF checkpoint:
config.json
,
generation_config.json
,
model.safetensors
(bf16, 3,008,627,352 B),
tokenizer.json
,
tokenizer_config.json
,
special_tokens_map.json
.
Optimizer state is not included.
Nanotron wrote an 18.05 GB AdamW state next to every
checkpoint (fp32 master weights + both moments, 974.80 GB across the 54 released checkpoints);
it is omitted from the release. These checkpoints are for inference and evaluation, not for
resuming training.
peak 5e-4, 500-step linear warmup, WSD with linear decay to 0 over the final 10%
Batching
seqlen 2048, global batch 1024 sequences = 2.10M tokens/step
Precision
bf16 parameters, fp32 gradients and optimizer state
Hardware
4 × H100, data-parallel degree 4, ~55 h / ~220 GPU-hours per run
Seeds 42/43/44 control
initialization only
— there is no dropout and the data order is fixed.
All six settings at a given seed start from bit-identical weights, released once as
KYS-Configs/nanotron/init/
.
How these were produced
A 100M-document pool was reservoir-sampled from DCLM-RefinedWeb and annotated with three
quality scorers and a 24-way WebOrganizer topic classifier →
KYS-DCLM-Refinedweb-100M-Scored
.
Checkpoints were converted from Nanotron to HF format with the same converter used for the
paper's own evaluations. The conversion was verified by re-running LightEval on a converted
checkpoint and reproducing the paper's stored numbers
exactly
(
rw_piqa
acc_norm 0.710555,
rw_hellaswag
acc_norm 0.486158, delta 0.000e+00).
Evaluation
LightEval, 0-shot,
acc_norm
with continuation-token-length-normalized log-likelihood, on
ARC-Easy, HellaSwag, PIQA, SIQA, OpenBookQA, CommonsenseQA and MMLU (57 subsets). Task
definitions, launchers and all 276 raw result JSONs are in
KYS-Configs/eval/
.
prompts, vLLM, Nanotron and eval configs + shared init weights
Citation
@misc{kys2026,
title = {Know Your Sources: Data Selection Matters when Rewriting for Data-Constrained Pretraining},
author = {TODO},
year = {2026},
note = {TODO: fill in venue / arXiv id / URL}
}
Runs of blab-jhu KYS-1.5B-Quality-Base on huggingface.co
2
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs
More Information About KYS-1.5B-Quality-Base huggingface.co Model
KYS-1.5B-Quality-Base huggingface.co is an AI model on huggingface.co that provides KYS-1.5B-Quality-Base's model effect (), which can be used instantly with this blab-jhu KYS-1.5B-Quality-Base model. huggingface.co supports a free trial of the KYS-1.5B-Quality-Base model, and also provides paid use of the KYS-1.5B-Quality-Base. Support call KYS-1.5B-Quality-Base model through api, including Node.js, Python, http.
KYS-1.5B-Quality-Base huggingface.co is an online trial and call api platform, which integrates KYS-1.5B-Quality-Base's modeling effects, including api services, and provides a free online trial of KYS-1.5B-Quality-Base, you can try KYS-1.5B-Quality-Base online for free by clicking the link below.
blab-jhu KYS-1.5B-Quality-Base online free url in huggingface.co:
KYS-1.5B-Quality-Base is an open source model from GitHub that offers a free installation service, and any user can find KYS-1.5B-Quality-Base on GitHub to install. At the same time, huggingface.co provides the effect of KYS-1.5B-Quality-Base install, users can directly use KYS-1.5B-Quality-Base installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
KYS-1.5B-Quality-Base install url in huggingface.co: