Core ML export of
internlm/Intern-Decision-0.8B
(Shanghai AI
Laboratory, Apache-2.0, snapshot
85a0cc5a
): a typed-decision model fine-tuned from Qwen3.5-0.8B that answers a set of
choice
/
score
/
noul
(yes/no) questions about a JSON state in one prefill pass. The text path is exported here;
the vision tower is not (text-only requests).
Files
Path
What
L320_F8/DecisionRow_fp16.mlpackage
Requests up to 320 tokens and 8 fields (the model card's three-field shape is 319 tokens).
L512_F8/DecisionRow_fp16.mlpackage
Up to 512 tokens, 8 fields.
L512_F8/DecisionRow_w8.mlpackage
Same, int8 per-channel weights (480 MB; 2 of 240 answers differ from the reference).
L1024_F16/DecisionRow_fp16.mlpackage
Up to 1,024 tokens, 16 fields.
*/config.json
Bucket dimensions, marker / pad / answer-symbol ids, temperature, system prompt.
embeddings.f16
Token embeddings (fp16, 248,320 × 1,024), gathered on the host.
tokenizer.json
The checkpoint's tokenizer (Qwen3.5 plus the
<decision>
token).
Inputs:
hidden
[1, L, 1024] (embedding rows of the right-padded prompt),
cos
/
sin
[L, 64] (RoPE tables for
positions 0…L−1),
field_onehot
[F, L] (row
i
selects the token before field
i
's
<decision>
). Output:
logits
[F, 62] over the answer symbols; take the first
n
entries for a field with
n
options, softmax, divide log-probs by
the temperature. fp16, GPU (
cpuAndGPU
), iOS 17 / macOS 14. Swift runtime:
InternDecisionManager
in
FluidUse
; conversion scripts in
mobius
(
models/computer-use/intern-decision-0.8b/coreml
).
What the package computes
Intern-Decision renders the request as one chat prompt: a fixed system prompt, a user turn with the state as JSON and
the decision schema (one line per field, its options mapped to the answer symbols
A
–
Z
,
a
–
z
,
0
–
9
), and an
assistant turn that is a JSON skeleton with one
<decision>
token per field. The logits at the position immediately
before each marker, restricted to that field's first
n
symbols, are its answer; the published temperature
(2.7478 for 0.8B) rescales the restricted softmax without changing the argmax. Nothing is generated.
DecisionRow
(
decision_export.py
) is one fixed-length request: the Qwen3.5 decoder from
qwen35_export.py
(the
Kev-0.8B / Cua-S1-4B export), the final norm at the pre-marker positions selected with a one-hot map per field, and
the 62 tied-embedding rows of the answer symbols. Token embeddings are gathered on the host (
embeddings.f16
); the
host applies the restricted softmax and temperature. Prompt compilation and the chat template come from the
checkpoint's own
inference.py
and tokenizer, so the token stream is the reference's.
Package
Tokens
Fields
Fits
L256_F4
256
4
Jevbench easy/original (~220 tokens), AG News (p95 299)
The fixed prompt (system prompt, headings, skeleton) is about 250 tokens, so the smallest three-field request is
319 tokens; 40% of Jevbench-Hard exceeds 1,024 tokens (max 4,074).
Fidelity
Reference: the checkpoint's
DecisionEngine
in fp32 on the Apple GPU (MPS), probabilities after temperature scaling,
on the bundled accuracy suites from the
Intern-Decision repo
(
benchmarks/accuracy-v1
, shuffled with seed 0). The fp32 PyTorch wrapper matches the reference to 3.6e-6
(50 Typed Decision fields).
Package
Suites
Decisions
Top answer differs
Max |Δp|
L512_F8
fp16
Jevbench (3), ToolACE, AG News, WildJailBreak
240
0
0.006
L1024_F16
fp16
Typed Decision (5 fields), Jevbench-Hard, ToolACE
280
0
0.007
L320_F8
fp16
Jevbench easy/original, AG News, WildJailBreak
100
0
0.006
L256_F4
fp16
Jevbench easy/original, AG News, WildJailBreak
100
0
0.006
L512_F8
int8 (per-channel)
same as
L512_F8
fp16
240
2
0.047
Reference and Core ML accuracy against the suite labels are identical on every subset (reports in the mobius directory).
int4 per-block compression needs an iOS 18 deployment target and was not built.
Latency
One request, 319 input tokens, three fields (choice, yes/no, score), the model card's RTX 4090 shape. M5 Pro (24 GB),
warmed,
bench.py
; Core ML on
CPU_AND_GPU
, times include the host embedding gather and RoPE tables.
Runtime
p50
p95
Core ML fp16
L320_F8
57 ms
59 ms
Core ML fp16
L384_F8
66 ms
67 ms
Core ML fp16
L512_F8
88 ms
96 ms
Core ML fp16
L768_F16
133 ms
145 ms
Core ML fp16
L1024_F16
182 ms
190 ms
PyTorch MPS bf16 (checkpoint
inference.py
)
150 ms
170 ms
PyTorch MPS fp32
177 ms
191 ms
The pass is compute-bound and scales with the bucket, not the request, so ship the smallest bucket that fits.
int8 weights leave GPU time unchanged (88 ms at
L512_F8
) and halve the package (955 → 480 MB).
ComputeUnit.ALL
matches
CPU_AND_GPU
: the Gated DeltaNet backbone does not run on the Neural Engine (see the Kev-0.8B notes).
The model card reports 34 ms for this request on an RTX 4090.
Runs of FluidInference intern-decision-0.8b-coreml on huggingface.co
0
Total runs
0
24-hour runs
0
3-day runs
0
7-day runs
0
30-day runs
More Information About intern-decision-0.8b-coreml huggingface.co Model
More intern-decision-0.8b-coreml license Visit here:
intern-decision-0.8b-coreml huggingface.co is an AI model on huggingface.co that provides intern-decision-0.8b-coreml's model effect (), which can be used instantly with this FluidInference intern-decision-0.8b-coreml model. huggingface.co supports a free trial of the intern-decision-0.8b-coreml model, and also provides paid use of the intern-decision-0.8b-coreml. Support call intern-decision-0.8b-coreml model through api, including Node.js, Python, http.
intern-decision-0.8b-coreml huggingface.co is an online trial and call api platform, which integrates intern-decision-0.8b-coreml's modeling effects, including api services, and provides a free online trial of intern-decision-0.8b-coreml, you can try intern-decision-0.8b-coreml online for free by clicking the link below.
FluidInference intern-decision-0.8b-coreml online free url in huggingface.co:
intern-decision-0.8b-coreml is an open source model from GitHub that offers a free installation service, and any user can find intern-decision-0.8b-coreml on GitHub to install. At the same time, huggingface.co provides the effect of intern-decision-0.8b-coreml install, users can directly use intern-decision-0.8b-coreml installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
intern-decision-0.8b-coreml install url in huggingface.co: