-
Added Tool calls tags (
"<|tool_call_start|>",
"<|tool_call_end|>",
"<|tool_result_start|>",
"<|tool_result_end|>",
)
-
Size {
Vocab size: 248191
pad_token_id: 248044
eos_token_id: 248046
}
-
Compatible with Qwen 2.5 and DeekSeek R1 • Optimized for Code . Use resize function without adaptation, see examples below
-
It needs 100k example to fully adapt routing features for RAG Routing model . So check Qwen modification tokenizer rules
-
A DeekSeek R1-based tokenizer enhanced with FIM markers from Microsoft datasets and other tags for ML (
<|fim_prefix|>,
<|fim_middle|>,
<f|im_suffix|>
)
-
Added Robotics & Embodiment tags (
"<|action_start|>",
"<|action_end|>",
"<|trajectory_start|>",
"<|trajectory_end|>",
"<|joint_start|>",
"<|joint_end|>",
"<|sensor_start|>",
"<|sensor_end|>",
"<|command_start|>",
"<|command_end|>",
"<|state_start|>",
"<|state_end|>",
"<|pose|>",
"<|velocity|>",
"<|force|>",
"<|torque|>",
"<|gripper|>",
"<|navigation|>",
"<|obstacle|>",
"<|task_start|>",
"<|task_end|>",
"<|plan_start|>",
"<|plan_end|>",
"<|behavior_start|>",
"<|behavior_end|>",
"<|skill_start|>",
"<|skill_end|>",
"<|motor|>",
"<|servo|>",
"<|imu|>",
"<|lidar|>",
"<|camera|>",
"<|depth|>",
"<|waypoint|>",
"<|path|>",
"<|collision|>",
"<|grasp|>",
"<|release|>",
"<|homing|>",
"<|emergency_stop|>",
"<|calibration|>",
"<|manipulation|>",
"<|locomotion|>",
"<|feedback|>",
"<|control_loop|>",)
-
Added Multi models support (
"<|image|>",
"<|video|>",
"<|sound|>",
"<|voice|>",
"<|listening|>",
"<|vision|>",)
-
Added Human mood tags (
"<|mood_happy|>",
"<|mood_sad|>",
"<|mood_angry|>",
"<|mood_neutral|>",
)
-
Added Tool calls tags (
"<|tool_call_start|>",
"<|tool_call_end|>",
"<|tool_result_start|>",
"<|tool_result_end|>",
)
-
Added RAG routing tags for RAG MoE Systems (
"
SCIENCE
",
"
CODING
",
"
STOCK_EXCHANGE
",
"
MEDICINE
",
"
GOVERNMENT
",
"
NEWS
",
"
GENERAL
",
"
MATERIAL_SCIENCE
",
"
ELECTRONICS
",
"
MICROELECTRONICS
",
"
ENGINEERING
",
"
ROBOTICS
",
"
ENERGY
",
"
AUTOMOTIVE
",
"
AVIATION
",
"
MATH
",
"
PYTHON
",
"
C
",
"
CPP
",
"
C_SHARP
",
"
JAVA
",
"
JAVASCRIPT
",
"
TYPESCRIPT
",
"
RUST
",
"
GO
",
"
RUBY
",
"
PHP
",
"
SWIFT
",
"
KOTLIN
",
"
BASH
",
"
SQL
",
"
ASSEMBLY
",
"
PHILOSOPHY
",
"
LITERATURE
",
"
SOCIOLOGY
",
"
PSYCHOLOGY
",
"
POLITICAL_SCIENCE
",
"
CULTURAL_STUDIES
",
"
ETHNOGRAPHY
",
"
HUMAN_RIGHTS
",
"
COMPLIANCE
",
"
MILITARY
",
"
BANKING
",
"
OIL_INDUSTRY
",
"
LIGHT_INDUSTRY
",
"
NATURE
",
"
OCEAN
",
"
SPORT
",
"
CULINARY
",
"
TRAVEL
",
"
HOBBY
"
)
-
Fully compatible with Microsoft BigCode datasets including The Stack, StarCoder, and NextCoder.
-
Enables efficient training on large-scale coding data for superior code generation and understanding.
(venv_ji) root@jirack2:# python -c '
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("./QwenRoboticsTokenizer")
print("Vocab size:", len(tok))
print("pad_token_id:", tok.pad_token_id)
print("eos_token_id:", tok.eos_token_id)
'