I am currently looking for open positions!
🤗
If you find this model useful or are looking for a talented AI/LLM Engineer, please reach out to me on LinkedIn:
Aaryan Kapoor
.
Experimental Build Required
🚧
This model utilizes the
Kimi Delta Attention (KDA)
architecture, which is not yet supported in the main branch of
llama.cpp
.
To run this GGUF, you
must
compile
llama.cpp
from
PR #18381
.
Attempting to run this on a standard build will result in errors.
Kimi Linear
is a hybrid linear attention architecture designed to outperform traditional full attention methods in long-context and scaling regimes. It uses
Kimi Delta Attention (KDA)
and a hybrid architecture (3:1 KDA-to-MLA ratio) to reduce memory usage and boost throughput by up to 6x on long sequences.
Performance & Architecture.
This model is currently quantized to
Q2_K
(and others) to fit on consumer hardware while testing the architecture's correctness. Despite the aggressive quantization, initial tests show the logic and reasoning capabilities remain intact.
Feature
Kimi Linear Specification
Architecture
Hybrid Linear Attention (MoE + MLA + KDA)
Context Length
1M Tokens (Supported by architecture)
Params
48B Total / 3B Activated
Throughput
~6.3x faster TPOT compared to MLA at 1M context
MMLU-Pro
51.0 (4k context)
RULER
84.3 (128k context, Pareto-optimal)
How to Run (llama.cpp)
Prerequisite:
You must clone and build the specific PR branch:
git clone https://github.com/ggml-org/llama.cpp
cd llama.cpp
git fetch origin pull/18381/head:kimi-linear-support
git checkout kimi-linear-support
make -j
1. CLI Inference (Interactive Chat)
./llama-cli -m Kimi-Linear-48B-Instruct.Q2_K.gguf \
-n 2048 \ # Adjust generation limit
-c 8192 \ # Context window (Model supports up to 1M)
--temp 0.8 \ # Recommended temperature
--top-p 0.9 \
-ngl 99 \ # Offload all layers to GPU
-p "<|im_start|>user\nHello, who are you?<|im_end|>\n<|im_start|>assistant\n" \
-cnv
Note:
The current GGUF implementation successfully mitigates previous "state collapse" issues found in early development.
2. Server Mode (API)
Running a persistent server is recommended for this size model to avoid reloading times.
Kimi-Linear-48B-A3B-Instruct-GGUF huggingface.co is an AI model on huggingface.co that provides Kimi-Linear-48B-A3B-Instruct-GGUF's model effect (), which can be used instantly with this AaryanK Kimi-Linear-48B-A3B-Instruct-GGUF model. huggingface.co supports a free trial of the Kimi-Linear-48B-A3B-Instruct-GGUF model, and also provides paid use of the Kimi-Linear-48B-A3B-Instruct-GGUF. Support call Kimi-Linear-48B-A3B-Instruct-GGUF model through api, including Node.js, Python, http.
Kimi-Linear-48B-A3B-Instruct-GGUF huggingface.co is an online trial and call api platform, which integrates Kimi-Linear-48B-A3B-Instruct-GGUF's modeling effects, including api services, and provides a free online trial of Kimi-Linear-48B-A3B-Instruct-GGUF, you can try Kimi-Linear-48B-A3B-Instruct-GGUF online for free by clicking the link below.
AaryanK Kimi-Linear-48B-A3B-Instruct-GGUF online free url in huggingface.co:
Kimi-Linear-48B-A3B-Instruct-GGUF is an open source model from GitHub that offers a free installation service, and any user can find Kimi-Linear-48B-A3B-Instruct-GGUF on GitHub to install. At the same time, huggingface.co provides the effect of Kimi-Linear-48B-A3B-Instruct-GGUF install, users can directly use Kimi-Linear-48B-A3B-Instruct-GGUF installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
Kimi-Linear-48B-A3B-Instruct-GGUF install url in huggingface.co: