Activated Parameters
: ~0.06B (2 of 8 experts active per token)
Configuration Changes
The following parameters were reduced from the original model:
Parameter
Original
Tiny
num_hidden_layers
93
8
hidden_size
7168
1024
intermediate_size
33792
2048
moe_intermediate_size
3072
256
num_attention_heads
96
8
num_key_value_heads
96
8
q_lora_rank
1536
256
kv_lora_rank
512
128
qk_nope_head_dim
128
64
qk_rope_head_dim
64
32
v_head_dim
128
64
num_experts
896
8
num_shared_experts
2
1
num_experts_per_token
16
2
routed_expert_hidden_size
3584
512
attn_res_block_size
12
4
linear_attn head_dim
74
32
linear_attn num_heads
96
8
vt_num_hidden_layers
27
2
vt_hidden_size
1024
256
Architecture Preserved
Mixed attention
: KDA (linear/delta attention) on layers 0–2, 4–6 and MLA (full multi-latent attention) on layers 3, 7 — same 3:1 KDA:MLA ratio as the original
Attention + MLP residuals
:
attn_res_block_size=4
enabled on all layers
MoE
: layers 1–7 use sparse MoE with latent expert projection; layer 0 is dense MLP
Vision tower
: included but reduced
Checkpoint Structure
Single-shard checkpoint (
model.safetensors
). Key prefix:
language_model.model.layers.{i}.*
, matching the original sharded checkpoint structure.
Usage
import sys
sys.path.insert(0, "/path/to/llm-compressor/src")
from transformers import AutoTokenizer
from llmcompressor.modeling.kimi_k3 import KimiK3ForConditionalGeneration
model = KimiK3ForConditionalGeneration.from_pretrained(
"inference-optimization/Kimi-K3-0.40B", device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained(
"inference-optimization/Kimi-K3-0.40B", trust_remote_code=True
)
inputs = tokenizer("According to all known laws", return_tensors="pt").to(model.device)
output = model.language_model.generate(**inputs, max_new_tokens=20)
print(tokenizer.decode(output[0], skip_special_tokens=True))
Validation Output
Total parameters: 0.396B
Number of layers: 8
Layer 0 attn type: KimiDeltaAttention (KDA/linear)
Layer 3 attn type: KimiMLAAttention (MLA/full)
Layer 7 attn type: KimiMLAAttention (MLA/full)
Layer 0 has MLP: True
Layer 1 has MoE: True
Attention residuals enabled: True
Attn res block size: 4
Forward pass loss: 0.0013
Generated: The FitnessGram Pacer Test is a multistage aerobic capacity test that progressively gets more difficult as it continues. The 20 meter pacer test will begin in 30 seconds
Creation Process
This model was created using the llm-compressor
create-tiny-model
claude skill.
Inspected the original
moonshotai/Kimi-K3
config to identify key architecture parameters
Reduced layer count, hidden dimensions, and expert counts to target ~0.4B total parameters
Preserved the 3:1 KDA:MLA attention pattern and attention/MLP residual connections
Initialized all weights from scratch (normal distribution, std=0.02; norms → 1.0; biases → 0.0)
Fine-tuned on a toy copypasta dataset until perplexity < 3.0 (achieved in ~57 epochs)
Notes
The model uses custom modeling code from
llmcompressor.modeling.kimi_k3
— it cannot be loaded with
AutoModelForCausalLM
without that module on the path
KimiK3ForConditionalGeneration
does not inherit from
GenerationMixin
; use
model.language_model.generate(...)
for text generation
The vision tower is present but untrained for vision tasks
Runs of inference-optimization Kimi-K3-0.40B on huggingface.co
1.1K
Total runs
0
24-hour runs
-32
3-day runs
-20.2K
7-day runs
-84.8K
30-day runs
More Information About Kimi-K3-0.40B huggingface.co Model
Kimi-K3-0.40B huggingface.co is an AI model on huggingface.co that provides Kimi-K3-0.40B's model effect (), which can be used instantly with this inference-optimization Kimi-K3-0.40B model. huggingface.co supports a free trial of the Kimi-K3-0.40B model, and also provides paid use of the Kimi-K3-0.40B. Support call Kimi-K3-0.40B model through api, including Node.js, Python, http.
Kimi-K3-0.40B huggingface.co is an online trial and call api platform, which integrates Kimi-K3-0.40B's modeling effects, including api services, and provides a free online trial of Kimi-K3-0.40B, you can try Kimi-K3-0.40B online for free by clicking the link below.
inference-optimization Kimi-K3-0.40B online free url in huggingface.co:
Kimi-K3-0.40B is an open source model from GitHub that offers a free installation service, and any user can find Kimi-K3-0.40B on GitHub to install. At the same time, huggingface.co provides the effect of Kimi-K3-0.40B install, users can directly use Kimi-K3-0.40B installed effect in huggingface.co for debugging and trial. It also supports api for free installation.