Core ML conversion of
Cua-S1-4B-0.2
(Cua, Apache-2.0):
LoRA adapters on
Qwen/Qwen3.5-4B
(Apache-2.0) that pick one
(element, action)
option for a GUI screen state. Both adapters are converted, each merged into its own
copy of the base:
text/
(accessibility tree) and
multimodal/
(screenshot).
The model is a one-pass chooser, not a generator: one forward pass over the prompt, then a softmax over
the answer-letter logits
A..Z
at the last position (Cua's
FourBModel
readout). The Core ML graphs keep
exactly that: prefill only, no KV cache, no vocabulary head.
Runs through
FluidUse
(
CuaS1FourBManager
) on macOS 14+.
tied embedding table, row-major
[248320, 2560]
fp16, gathered on the host
1.27 GB
text/L1024/
text decoder, fp16, 4 parts x 8 layers, prompts up to 1024 tokens
6.8 GB
text/L1024-w8/
int8 linears
3.4 GB
text/L1024-gptq/
GPTQ: MLP int4 (block 16) + other linears int8
2.6 GB
multimodal/L2048/
multimodal decoder, fp16, prompts up to 2048 tokens
6.8 GB
multimodal/L2048-w8/
int8 linears
3.4 GB
multimodal/vision/
vision tower + merger (up to 4096 patches), position table, config
0.66 GB
Each decoder part takes
hidden [1, L, 2560]
,
cos
/
sin [L, 64]
(interleaved M-RoPE, host computed)
and returns the next hidden state; the last part also takes
last_onehot [1, L]
and returns
letter_logits [1, 26]
. Prompts are right-padded; the model is causal, so padding is exact.
Results
Apple M5 Pro (24 GB), macOS 27, GPU (
cpuAndGPU
). Parity reference: fp32 transformers decoder layers on
the merged weights and the unpadded prompt, over 38 tasks from Cua's own
cua_bench_s1
generator.
Accuracy: GUI-360 test split (613 text tasks) rebuilt with Cua's
gui360
converter; Cua's own frozen
615-task split is published by hash only, so this is protocol-adjacent, not the same task set.
Build
Size
Argmax vs fp32 (38)
Max |Δp|
GUI-360 text (613)
text fp16
6.8 GB
38/38
0.005
85.5%
text gptq
2.6 GB
38/38
0.133
85.5%
text w8
3.4 GB
37/38
0.019
—
multimodal fp16 (+ vision)
7.4 GB
38/38
0.153
—
multimodal w8 (+ vision)
4.0 GB
38/38
0.163
—
One decision on the GPU: text ~0.7 s (1024-token bucket) and multimodal ~1.65 s (vision 0.2 s +
2048-token decoder) on an idle machine; about 1.1 s and 3 s while other GPU work was running.
On the same first 100 GUI-360 tasks, Cua's PyTorch runtime (bf16) scores 84%, Core ML fp16 88% and gptq
87%; the Core ML task outcome matches PyTorch on 96-97 of 100. Post-training int4 without calibration and 4/3/2-bit palettes are not
published: they flip 10-32 of the 38 decisions.
Limits
GPU only: the Neural Engine path falls back to the CPU for most of the graph (19 s / decision).
At most 26 options per decision (one letter each). Longer prompts than the loaded bucket are rejected.
Screenshots are smart-resized as in the reference processor, capped at 4096 patches (about 1280x800
px); larger screenshots are downscaled further than the PyTorch reference would.
fp16 vision features move multimodal probabilities by up to 0.15 (argmax unchanged on the fixtures).
Provenance
Conversion code, parity and benchmark scripts:
mobius
models/computer-use/cua-s1-4b/coreml
. Base weights:
Qwen/Qwen3.5-4B
; adapters:
cua-ai/cua-s1-4b-0.2
(
text/
,
multimodal/
). Both Apache-2.0; see
NOTICE
.
Runs of FluidInference cua-s1-4b-coreml on huggingface.co
10
Total runs
0
24-hour runs
1
3-day runs
6
7-day runs
10
30-day runs
More Information About cua-s1-4b-coreml huggingface.co Model
cua-s1-4b-coreml huggingface.co is an AI model on huggingface.co that provides cua-s1-4b-coreml's model effect (), which can be used instantly with this FluidInference cua-s1-4b-coreml model. huggingface.co supports a free trial of the cua-s1-4b-coreml model, and also provides paid use of the cua-s1-4b-coreml. Support call cua-s1-4b-coreml model through api, including Node.js, Python, http.
cua-s1-4b-coreml huggingface.co is an online trial and call api platform, which integrates cua-s1-4b-coreml's modeling effects, including api services, and provides a free online trial of cua-s1-4b-coreml, you can try cua-s1-4b-coreml online for free by clicking the link below.
FluidInference cua-s1-4b-coreml online free url in huggingface.co:
cua-s1-4b-coreml is an open source model from GitHub that offers a free installation service, and any user can find cua-s1-4b-coreml on GitHub to install. At the same time, huggingface.co provides the effect of cua-s1-4b-coreml install, users can directly use cua-s1-4b-coreml installed effect in huggingface.co for debugging and trial. It also supports api for free installation.