inference-optimization / Kimi-K3-0.40B

huggingface.co
Total runs: 1.1K
24-hour runs: 0
7-day runs: -20.2K
30-day runs: -84.8K
Model's Last Updated: July 28 2026
feature-extraction

Introduction of Kimi-K3-0.40B

Model Details of Kimi-K3-0.40B

Kimi-K3-0.40B

This is a tiny version of moonshotai/Kimi-K3 created for testing and development.

Model Details
  • Base Model : moonshotai/Kimi-K3
  • Architecture : kimi_linear (KimiK3ForConditionalGeneration)
  • Total Parameters : 0.40B
  • Activated Parameters : ~0.06B (2 of 8 experts active per token)
Configuration Changes

The following parameters were reduced from the original model:

Parameter Original Tiny
num_hidden_layers 93 8
hidden_size 7168 1024
intermediate_size 33792 2048
moe_intermediate_size 3072 256
num_attention_heads 96 8
num_key_value_heads 96 8
q_lora_rank 1536 256
kv_lora_rank 512 128
qk_nope_head_dim 128 64
qk_rope_head_dim 64 32
v_head_dim 128 64
num_experts 896 8
num_shared_experts 2 1
num_experts_per_token 16 2
routed_expert_hidden_size 3584 512
attn_res_block_size 12 4
linear_attn head_dim 74 32
linear_attn num_heads 96 8
vt_num_hidden_layers 27 2
vt_hidden_size 1024 256
Architecture Preserved
  • Mixed attention : KDA (linear/delta attention) on layers 0–2, 4–6 and MLA (full multi-latent attention) on layers 3, 7 — same 3:1 KDA:MLA ratio as the original
  • Attention + MLP residuals : attn_res_block_size=4 enabled on all layers
  • MoE : layers 1–7 use sparse MoE with latent expert projection; layer 0 is dense MLP
  • Vision tower : included but reduced
Checkpoint Structure

Single-shard checkpoint ( model.safetensors ). Key prefix: language_model.model.layers.{i}.* , matching the original sharded checkpoint structure.

Usage
import sys
sys.path.insert(0, "/path/to/llm-compressor/src")

from transformers import AutoTokenizer
from llmcompressor.modeling.kimi_k3 import KimiK3ForConditionalGeneration

model = KimiK3ForConditionalGeneration.from_pretrained(
    "inference-optimization/Kimi-K3-0.40B", device_map="auto"
)
tokenizer = AutoTokenizer.from_pretrained(
    "inference-optimization/Kimi-K3-0.40B", trust_remote_code=True
)

inputs = tokenizer("According to all known laws", return_tensors="pt").to(model.device)
output = model.language_model.generate(**inputs, max_new_tokens=20)
print(tokenizer.decode(output[0], skip_special_tokens=True))
Validation Output
Total parameters: 0.396B
Number of layers: 8
Layer 0 attn type: KimiDeltaAttention (KDA/linear)
Layer 3 attn type: KimiMLAAttention (MLA/full)
Layer 7 attn type: KimiMLAAttention (MLA/full)
Layer 0 has MLP: True
Layer 1 has MoE: True
Attention residuals enabled: True
Attn res block size: 4

Forward pass loss: 0.0013
Generated: The FitnessGram Pacer Test is a multistage aerobic capacity test that progressively gets more difficult as it continues. The 20 meter pacer test will begin in 30 seconds
Creation Process

This model was created using the llm-compressor create-tiny-model claude skill.

  1. Inspected the original moonshotai/Kimi-K3 config to identify key architecture parameters
  2. Reduced layer count, hidden dimensions, and expert counts to target ~0.4B total parameters
  3. Preserved the 3:1 KDA:MLA attention pattern and attention/MLP residual connections
  4. Initialized all weights from scratch (normal distribution, std=0.02; norms → 1.0; biases → 0.0)
  5. Fine-tuned on a toy copypasta dataset until perplexity < 3.0 (achieved in ~57 epochs)
Notes
  • The model uses custom modeling code from llmcompressor.modeling.kimi_k3 — it cannot be loaded with AutoModelForCausalLM without that module on the path
  • KimiK3ForConditionalGeneration does not inherit from GenerationMixin ; use model.language_model.generate(...) for text generation
  • The vision tower is present but untrained for vision tasks

Runs of inference-optimization Kimi-K3-0.40B on huggingface.co

1.1K
Total runs
0
24-hour runs
-32
3-day runs
-20.2K
7-day runs
-84.8K
30-day runs

More Information About Kimi-K3-0.40B huggingface.co Model

More Kimi-K3-0.40B license Visit here:

https://choosealicense.com/licenses/mit

Kimi-K3-0.40B huggingface.co

Kimi-K3-0.40B huggingface.co is an AI model on huggingface.co that provides Kimi-K3-0.40B's model effect (), which can be used instantly with this inference-optimization Kimi-K3-0.40B model. huggingface.co supports a free trial of the Kimi-K3-0.40B model, and also provides paid use of the Kimi-K3-0.40B. Support call Kimi-K3-0.40B model through api, including Node.js, Python, http.

inference-optimization Kimi-K3-0.40B online free

Kimi-K3-0.40B huggingface.co is an online trial and call api platform, which integrates Kimi-K3-0.40B's modeling effects, including api services, and provides a free online trial of Kimi-K3-0.40B, you can try Kimi-K3-0.40B online for free by clicking the link below.

inference-optimization Kimi-K3-0.40B online free url in huggingface.co:

https://huggingface.co/inference-optimization/Kimi-K3-0.40B

Kimi-K3-0.40B install

Kimi-K3-0.40B is an open source model from GitHub that offers a free installation service, and any user can find Kimi-K3-0.40B on GitHub to install. At the same time, huggingface.co provides the effect of Kimi-K3-0.40B install, users can directly use Kimi-K3-0.40B installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

Kimi-K3-0.40B install url in huggingface.co:

https://huggingface.co/inference-optimization/Kimi-K3-0.40B

Url of Kimi-K3-0.40B

Provider of Kimi-K3-0.40B huggingface.co

inference-optimization
ORGANIZATIONS

Other API from inference-optimization