kev is a System 1 decision model. It answers typed
choice
,
noul
and
score
questions about a text
state
with one scoring pass per question. It
does not generate text. kev is a frozen
Qwen/Qwen3.5-0.8B-Base
backbone, a
rank-16 LoRA adapter, and a PointerHead readout.
This checkpoint only works with vllm.cpp
(directly, or through LocalAI's
vllm-cpp
backend). It is not a general-purpose checkpoint:
transformers
,
vLLM and llama.cpp do not know the
KevModel
architecture or
head.safetensors
.
Redistribution and license
This is a redistribution of upstream weights, converted for vllm.cpp and
LocalAI. No weights were trained or changed here, other than the LoRA merge
described below.
kev adapter and PointerHead:
jaredpalmer/kev-0.8b
@
9a45d25eb2ab761841196625383fa1dff0e56c1e
, by Jared Palmer, Apache-2.0.
Reference code:
jaredpalmer/kev
.
Base backbone and tokenizer:
Qwen/Qwen3.5-0.8B-Base
@
dc7cdfe2ee4154fa7e30f5b51ca41bfa40174e68
, by the Qwen team, Apache-2.0.
The upstream licenses apply to these files. Credit for the model belongs to
the original authors.
Conversion
The converter is
scripts/convert-kev.py
from vllm.cpp at commit
96788348627b6a079fcc3ef7fc6676a970965d7b
. It merges the LoRA once
(
W' = W + 2.0 * (B @ A)
, r=16, alpha=32, math in F32, stored in the base
dtype BF16), drops the base's vision and MTP tensors, converts
head.pt
to
head.safetensors
, and writes a
config.json
with
architectures: ["KevModel"]
.
Then send the same request body to LocalAI's
POST /v1/systemone
with
"model": "kev-0.8b"
.
Example
Request (served on CPU from this exact directory, 2026-09-30):
curl http://localhost:8000/v1/systemone -H 'Content-Type: application/json' -d '{ "model": "kev", "state": "I was charged twice for my subscription this month. Please refund the duplicate payment.", "questions": { "refund": {"type": "noul", "instructions": "Does the user request a refund?"}, "department": {"type": "choice", "instructions": "Which department should handle this?", "criteria": {"billing": "Payments and refunds", "technical": "Software bugs", "sales": "New purchases"}}, "urgency": {"type": "score", "instructions": "How urgent is this?", "criteria": ["Routine", "Urgent", "Emergency"]}}}'
kev's answer semantics follow the kev reference:
choice
confidence is
(max(p) - 1/K) / (1 - 1/K)
,
score
is the expected level index with
confidence
1 - E|level - mode| / (L - 1)
,
noul
is
p(true)
, and
probabilities are rounded to 4 decimals in this response.
What was verified and what was not
Verified for this upload:
The conversion above ran at the stated vllm.cpp commit and merged all 186
LoRA modules.
The directory loads in the vllm.cpp CPU server built from the same commit,
and the one request above returns the answers shown. The answers are
plausible for the input. This is a smoke test, not an accuracy measurement.
Recorded by the vllm.cpp project, not re-measured for this upload:
PointerHead golden-vector tests (25 cases) and LoRA-merge / hidden-state
goldens against the kev Python reference at
jaredpalmer/kev@19dcae9b6e3e1a48200c5825aad9fc200d31e20a
.
A 5-case end-to-end comparison (choice 2 and 3 options, score 5 levels,
noul, multi-question) against the kev reference server, recorded as equal.
Not verified:
No GPU (CUDA, ROCm, Metal) run of this checkpoint.
No comparison of this upload against the kev reference server. No vLLM
gate exists for kev, because vLLM does not serve this architecture.
No accuracy benchmark.
No GGUF or quantized variant.
Runs of mudler kev-0.8b-vllm-cpp on huggingface.co
12
Total runs
12
24-hour runs
12
3-day runs
12
7-day runs
12
30-day runs
More Information About kev-0.8b-vllm-cpp huggingface.co Model
kev-0.8b-vllm-cpp huggingface.co is an AI model on huggingface.co that provides kev-0.8b-vllm-cpp's model effect (), which can be used instantly with this mudler kev-0.8b-vllm-cpp model. huggingface.co supports a free trial of the kev-0.8b-vllm-cpp model, and also provides paid use of the kev-0.8b-vllm-cpp. Support call kev-0.8b-vllm-cpp model through api, including Node.js, Python, http.
kev-0.8b-vllm-cpp huggingface.co is an online trial and call api platform, which integrates kev-0.8b-vllm-cpp's modeling effects, including api services, and provides a free online trial of kev-0.8b-vllm-cpp, you can try kev-0.8b-vllm-cpp online for free by clicking the link below.
mudler kev-0.8b-vllm-cpp online free url in huggingface.co:
kev-0.8b-vllm-cpp is an open source model from GitHub that offers a free installation service, and any user can find kev-0.8b-vllm-cpp on GitHub to install. At the same time, huggingface.co provides the effect of kev-0.8b-vllm-cpp install, users can directly use kev-0.8b-vllm-cpp installed effect in huggingface.co for debugging and trial. It also supports api for free installation.