aloobun / IN-Llama-3-Tokenizer

huggingface.co
Total runs: 0
24-hour runs: 0
7-day runs: 0
30-day runs: 0
Model's Last Updated: March 08 2025

Introduction of IN-Llama-3-Tokenizer

Model Details of IN-Llama-3-Tokenizer

In this experiment, I trained a tokenizer that supports multiple Indian languages and merged it with the Llama-3 tokenizer.

STEP 1:

I sampled data from the multilingual(7 Indic languages) aloobun/dhpileIN dataset and trained a SentencePiece tokenizer.

STEP 2:

I evaluated the tokenizer's performance on:

  • Unicode coverage.
  • Token distribution.
  • Tokenization complexity across different scripts.
  • Encoding and decoding capabilities &
  • Edge cases e.g., special characters, numbers, etc.
STEP 2.1:

The first test gives detailed results of the tokenizer's performance on unicode coverage, token distribution visualiztion and complexity across scripts.

Step 2.2:

The second script tests the encoding and decoding capabilities. Example output:

Bengali Analysis:
Original Text Length: 48 characters
Token IDs Count: 11
Token Strings: ['▁আমি', '▁বাংলাদেশ', '▁থেকে', '▁এসে', 'ছি', '।', '▁কলকাতা', '▁একটি', '▁সুন্দর', '▁শহর', '।']
Text Reconstruction: True

Hindi Analysis:
Original Text Length: 49 characters
Token IDs Count: 15
Token Strings: ['▁नम', 'स्ते', ',', '▁मैं', '▁भारत', '▁से', '▁हू', 'ँ', '।', '▁दिल्ली', '▁बहुत', '▁बड़ा', '▁शहर', '▁है', '।']
Text Reconstruction: True

Kannada Analysis:
Original Text Length: 53 characters
Token IDs Count: 13
Token Strings: ['▁ನಾನು', '▁ಬೆಂಗಳೂರಿ', 'ನಿಂದ', '▁ಬಂದ', 'ಿದ್ದೇನೆ', '।', '▁ಕನ್ನಡ', '▁ಒಂದು', '▁ಸೋ', 'ಂಪ', 'ಿನ', '▁ಭಾಷೆ', '।']
Text Reconstruction: True

Malayalam Analysis:
Original Text Length: 47 characters
Token IDs Count: 15
Token Strings: ['▁ഞ', 'ാ', 'ൻ', '▁കേരള', 'ത്തി', 'ൽ', '▁നിന്നാണ്', '.', '▁കൊച്ചി', '▁ഒരു', '▁സുന്ദ', 'ര', '▁നഗ', 'രം', '.']
Text Reconstruction: True

Telugu Analysis:
Original Text Length: 53 characters
Token IDs Count: 10
Token Strings: ['▁నేను', '▁తెలంగాణ', '▁నుంచి', '▁వచ్చ', 'ాను', '.', '▁హైదరాబాద్', '▁అద్భుతమైన', '▁నగరం', '.']
Text Reconstruction: True

Tamil Analysis:
Original Text Length: 54 characters
Token IDs Count: 13
Token Strings: ['▁நான்', '▁தமிழ்நா', 'ட்டை', 'ச்', '▁சேர்ந்த', 'வன்', '.', '▁சென்னை', '▁ஒரு', '▁பெரிய', '▁நக', 'ரம்', '.']
Text Reconstruction: True

Gujarati Analysis:
Original Text Length: 50 characters
Token IDs Count: 12
Token Strings: ['▁હું', '▁ગુજરાત', '▁થી', '▁આવ્યો', '▁છું', '।', '▁અમદાવાદ', '▁એક', '▁સુંદર', '▁શહેર', '▁છે', '।']
Text Reconstruction: True
STEP 3:

This script is used to merge and extend the tokenizer for the Llama3 tokenizer.

Script ensures:

  • No duplicate tokens are added.
  • Tokens arent excessively long.
  • New tokens are correctly integrated.
  • Token mappings, etc

I feel there are some unecessary bloat like token validation and redundant test methods in the script. I'm still working on how to improve things and will update as soon as I have any progress.

Here's a comparison of sub word fertility scores between sarvam-1 and this model.

sarvam-1 IN-Llama-3-Tokenizer
Bengali 1.7 3.52
Gujrati 2.784313 3.588235
Hindi 1.583333 2.933333
Kannada 2.571428 3.976190
Malayalam 3.487804 4.365853
Tamil 2.767441 3.860465
Telugu 2.372093 3.511627

Runs of aloobun IN-Llama-3-Tokenizer on huggingface.co

0
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs

More Information About IN-Llama-3-Tokenizer huggingface.co Model

More IN-Llama-3-Tokenizer license Visit here:

https://choosealicense.com/licenses/llama3

IN-Llama-3-Tokenizer huggingface.co

IN-Llama-3-Tokenizer huggingface.co is an AI model on huggingface.co that provides IN-Llama-3-Tokenizer's model effect (), which can be used instantly with this aloobun IN-Llama-3-Tokenizer model. huggingface.co supports a free trial of the IN-Llama-3-Tokenizer model, and also provides paid use of the IN-Llama-3-Tokenizer. Support call IN-Llama-3-Tokenizer model through api, including Node.js, Python, http.

IN-Llama-3-Tokenizer huggingface.co Url

https://huggingface.co/aloobun/IN-Llama-3-Tokenizer

aloobun IN-Llama-3-Tokenizer online free

IN-Llama-3-Tokenizer huggingface.co is an online trial and call api platform, which integrates IN-Llama-3-Tokenizer's modeling effects, including api services, and provides a free online trial of IN-Llama-3-Tokenizer, you can try IN-Llama-3-Tokenizer online for free by clicking the link below.

aloobun IN-Llama-3-Tokenizer online free url in huggingface.co:

https://huggingface.co/aloobun/IN-Llama-3-Tokenizer

IN-Llama-3-Tokenizer install

IN-Llama-3-Tokenizer is an open source model from GitHub that offers a free installation service, and any user can find IN-Llama-3-Tokenizer on GitHub to install. At the same time, huggingface.co provides the effect of IN-Llama-3-Tokenizer install, users can directly use IN-Llama-3-Tokenizer installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

IN-Llama-3-Tokenizer install url in huggingface.co:

https://huggingface.co/aloobun/IN-Llama-3-Tokenizer

Url of IN-Llama-3-Tokenizer

IN-Llama-3-Tokenizer huggingface.co Url

Provider of IN-Llama-3-Tokenizer huggingface.co

aloobun
ORGANIZATIONS

Other API from aloobun

huggingface.co

Total runs: 67
Run Growth: 13
Growth Rate: 19.40%
Updated:May 16 2024
huggingface.co

Total runs: 23
Run Growth: 7
Growth Rate: 30.43%
Updated:December 14 2024
huggingface.co

Total runs: 16
Run Growth: 6
Growth Rate: 37.50%
Updated:May 03 2024
huggingface.co

Total runs: 14
Run Growth: 4
Growth Rate: 28.57%
Updated:November 30 2023
huggingface.co

Total runs: 14
Run Growth: 6
Growth Rate: 42.86%
Updated:November 30 2023
huggingface.co

Total runs: 11
Run Growth: 11
Growth Rate: 100.00%
Updated:November 27 2025
huggingface.co

Total runs: 4
Run Growth: 0
Growth Rate: 0.00%
Updated:March 06 2025
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:May 21 2024
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:February 21 2026
huggingface.co

Total runs: 0
Run Growth: 0
Growth Rate: 0.00%
Updated:March 06 2025