AIGym / tokenizer-65606

huggingface.co
Total runs: 0
24-hour runs: 0
7-day runs: 0
30-day runs: 0
Model's Last Updated: August 22 2025

Introduction of tokenizer-65606

Model Details of tokenizer-65606

Model Card: OED-Biased BPE Tokenizer

image/png

This model card describes a custom-built Byte-Pair Encoding (BPE) tokenizer with a vocabulary intentionally biased towards the Oxford English Dictionary (OED).

Model Details

This tokenizer was developed to create a vocabulary that is robust for general text while ensuring the inclusion of a specific, high-value set of words from the OED. It is a subword tokenizer based on the BPE algorithm.

  • Model Type: Byte-Pair Encoding (BPE)
  • Vocabulary Size (Target): 50,000
  • Vocabulary Size (Final): 65,606
  • Special Tokens:
    • <|UNK|> (Unknown Token)
    • <|BOS|> (Beginning of Sequence)
    • <|EOS|> (End of Sequence)
    • <|PAD|> (Padding Token)
    • <|MASK|> (Mask Token)
  • Normalization: Unicode NFKC
  • Pre-tokenization: Regex-based splitting to handle contractions, words, numbers, and whitespace.
    • Regex: (?i:'s|'t|'re|'ve|'m|'ll|'d)|[^\r\n\p{L}\p{N}]?\p{L}+|\p{N}{1,3}| ?[^\s\p{L}\p{N}]+[\r\n]*|\s*[\r\n]+|\s+(?!\S)|\s+
Training

The tokenizer was trained in a two-stage process designed to bias the vocabulary.

Stage 1: Initial BPE Training

The tokenizer was first trained on a corpus of text from the following Hugging Face datasets:

  • motionlabs/fineweb-ultra-mini (10,000 samples)
  • suolyer/pile_arxiv (10,000 samples)

The training aimed for a vocabulary size of 50,000 with a minimum token frequency of 2.

Stage 2: Vocabulary Augmentation

After the initial training, the vocabulary was augmented by adding a predefined list of 20,000 words from the Oxford English Dictionary ( 20k.txt ). This step ensures that these specific words are included as complete tokens in the final vocabulary, regardless of their frequency in the training data. This "bias_by_post_add_tokens" method resulted in:

  • Initial Vocabulary Size: 50,000
  • OED Words Added: 20,000
  • Final Vocabulary Size: 65,606

The final vocabulary size exceeds the initial target due to the forced addition of the OED word list.

Intended Use

This tokenizer is intended for natural language processing tasks where a strong representation of the English language, including a rich and potentially less common vocabulary, is important. It is well-suited for models that will process literary, academic, or other forms of text where the OED vocabulary is likely to be present.

Limitations and Bias

The primary bias of this tokenizer is its explicit enrichment with words from the Oxford English Dictionary. While this is an intentional feature, it means the tokenizer may be less efficient for texts that are significantly different in character from the training data and the OED wordlist (e.g., code, specialized technical jargon, or informal social media content). The vocabulary is heavily skewed towards the English language.

Runs of AIGym tokenizer-65606 on huggingface.co

0
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs

More Information About tokenizer-65606 huggingface.co Model

tokenizer-65606 huggingface.co

tokenizer-65606 huggingface.co is an AI model on huggingface.co that provides tokenizer-65606's model effect (), which can be used instantly with this AIGym tokenizer-65606 model. huggingface.co supports a free trial of the tokenizer-65606 model, and also provides paid use of the tokenizer-65606. Support call tokenizer-65606 model through api, including Node.js, Python, http.

tokenizer-65606 huggingface.co Url

https://huggingface.co/AIGym/tokenizer-65606

AIGym tokenizer-65606 online free

tokenizer-65606 huggingface.co is an online trial and call api platform, which integrates tokenizer-65606's modeling effects, including api services, and provides a free online trial of tokenizer-65606, you can try tokenizer-65606 online for free by clicking the link below.

AIGym tokenizer-65606 online free url in huggingface.co:

https://huggingface.co/AIGym/tokenizer-65606

tokenizer-65606 install

tokenizer-65606 is an open source model from GitHub that offers a free installation service, and any user can find tokenizer-65606 on GitHub to install. At the same time, huggingface.co provides the effect of tokenizer-65606 install, users can directly use tokenizer-65606 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

tokenizer-65606 install url in huggingface.co:

https://huggingface.co/AIGym/tokenizer-65606

Url of tokenizer-65606

tokenizer-65606 huggingface.co Url

Provider of tokenizer-65606 huggingface.co

AIGym
ORGANIZATIONS

Other API from AIGym

huggingface.co

Total runs: 648
Run Growth: 0
Growth Rate: 0.00%
Updated:February 25 2024
huggingface.co

Total runs: 570
Run Growth: -8
Growth Rate: -1.39%
Updated:September 02 2025
huggingface.co

Total runs: 24
Run Growth: 0
Growth Rate: 0.00%
Updated:February 09 2025
huggingface.co

Total runs: 16
Run Growth: 0
Growth Rate: 0.00%
Updated:April 06 2025
huggingface.co

Total runs: 14
Run Growth: 2
Growth Rate: 14.29%
Updated:September 04 2025
huggingface.co

Total runs: 10
Run Growth: 2
Growth Rate: 20.00%
Updated:September 04 2025
huggingface.co

Total runs: 9
Run Growth: 0
Growth Rate: 0.00%
Updated:January 09 2025
huggingface.co

Total runs: 5
Run Growth: 0
Growth Rate: 0.00%
Updated:April 10 2025
huggingface.co

Total runs: 3
Run Growth: 0
Growth Rate: 0.00%
Updated:June 14 2024
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:July 25 2025
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:August 21 2025
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:August 16 2025
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:March 14 2025
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:February 11 2025
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:March 14 2025