inference-optimization / Inkling-0.6B-A0.6B

huggingface.co
Total runs: 733
24-hour runs: 0
7-day runs: 190
30-day runs: 644
Model's Last Updated: July 21 2026
image-text-to-text

Introduction of Inkling-0.6B-A0.6B

Model Details of Inkling-0.6B-A0.6B

Inkling-0.6B-A0.6B

This is a tiny version of thinkingmachines/Inkling created for testing and development.

Model Details
  • Base Model : thinkingmachines/Inkling
  • Architecture : inkling_mm_model (InklingForConditionalGeneration)
  • Total Parameters : 0.644B
  • Activated Parameters : 0.602B
Configuration Changes

The following parameters were reduced from the original model:

Parameter Original Tiny
text_config.num_hidden_layers 66 12
text_config.hidden_size 6144 1024
text_config.intermediate_size 24576 4096
text_config.num_attention_heads 64 8
text_config.num_key_value_heads 8 2
text_config.swa_num_attention_heads 64 8
text_config.swa_num_key_value_heads 16 4
text_config.n_routed_experts 256 8
text_config.num_experts_per_tok 6 4
text_config.moe_intermediate_size 3072 512
text_config.num_mtp_layers 8 1
vision_config.n_layers 4 1
vision_config.hidden_size 1024 256
vision_config.decoder_dmodel 6144 1024
audio_config.decoder_dmodel 6144 1024

Layer type patterns are preserved: 2 repetitions of [5× hybrid_sliding + 1× hybrid ], with the first 2 MLP layers as dense and the rest as sparse (MoE).

Checkpoint Structure

Single safetensors file ( model.safetensors ). Key naming matches the original checkpoint format ( model.llm.* , model.audio.* , model.visual.* ).

Usage
from transformers.models.inkling import InklingForConditionalGeneration
from transformers import AutoTokenizer

model = InklingForConditionalGeneration.from_pretrained("inference-optimization/Inkling-0.6B-A0.6B", device_map="auto")
tokenizer = AutoTokenizer.from_pretrained("inference-optimization/Inkling-0.6B-A0.6B")

input_ids = tokenizer("According to all known laws", return_tensors="pt").input_ids.to(model.device)
output = model.generate(input_ids, max_new_tokens=20)
print(tokenizer.decode(output[0]))
Creation Process

This model was created using the llm-compressor create-tiny-model claude skill.

  1. Config inspected via inspect_config.py
  2. Tiny model created via modified save_tiny_model.py — all-zero params fixed post init_weights
  3. Fine-tuned on copypasta dataset; reached perplexity 1.45 (target: ≤3.0) at lr=5e-4
  4. Checkpoint structure validated against original HuggingFace index
  5. Inference validated via validate_tiny_model.py
Notes
  • The embed_tokens weights require explicit re-initialization after init_weights() (they initialize to zero in this architecture). The save script applies a fixup: any all-zero, non-finite, or extreme-valued parameter is re-initialized with kaiming_uniform / normal / ones as appropriate.
  • MTP (Multi-Token Prediction) layers present in the original checkpoint ( model.mtp.* ) are not included, as InklingForConditionalGeneration does not expose them through its standard interface.
  • Validation output: Success: 1.4451 <= 10.0

Runs of inference-optimization Inkling-0.6B-A0.6B on huggingface.co

733
Total runs
0
24-hour runs
82
3-day runs
190
7-day runs
644
30-day runs

More Information About Inkling-0.6B-A0.6B huggingface.co Model

More Inkling-0.6B-A0.6B license Visit here:

https://choosealicense.com/licenses/mit

Inkling-0.6B-A0.6B huggingface.co

Inkling-0.6B-A0.6B huggingface.co is an AI model on huggingface.co that provides Inkling-0.6B-A0.6B's model effect (), which can be used instantly with this inference-optimization Inkling-0.6B-A0.6B model. huggingface.co supports a free trial of the Inkling-0.6B-A0.6B model, and also provides paid use of the Inkling-0.6B-A0.6B. Support call Inkling-0.6B-A0.6B model through api, including Node.js, Python, http.

inference-optimization Inkling-0.6B-A0.6B online free

Inkling-0.6B-A0.6B huggingface.co is an online trial and call api platform, which integrates Inkling-0.6B-A0.6B's modeling effects, including api services, and provides a free online trial of Inkling-0.6B-A0.6B, you can try Inkling-0.6B-A0.6B online for free by clicking the link below.

inference-optimization Inkling-0.6B-A0.6B online free url in huggingface.co:

https://huggingface.co/inference-optimization/Inkling-0.6B-A0.6B

Inkling-0.6B-A0.6B install

Inkling-0.6B-A0.6B is an open source model from GitHub that offers a free installation service, and any user can find Inkling-0.6B-A0.6B on GitHub to install. At the same time, huggingface.co provides the effect of Inkling-0.6B-A0.6B install, users can directly use Inkling-0.6B-A0.6B installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

Inkling-0.6B-A0.6B install url in huggingface.co:

https://huggingface.co/inference-optimization/Inkling-0.6B-A0.6B

Url of Inkling-0.6B-A0.6B

Provider of Inkling-0.6B-A0.6B huggingface.co

inference-optimization
ORGANIZATIONS

Other API from inference-optimization