Core ML conversion of
laya-multilingual
(Convai Innovations, Apache-2.0): a 322M-parameter
mmBERT-base encoder with a typed decision head that answers
choice
,
score
, and
noul
questions about a text state in one forward pass, returning calibrated probabilities and no
generated tokens. Weights are unchanged from
convaiinnovations/laya
multilingual/
at
revision
1c5edc17a7acd8701df6fc341c0d179f1c62c982
.
Runs through
FluidUse
(
LayaManager
) on macOS 14+.
let laya =tryawaitLayaManager.load() // downloads the 128 + 512 buckets and tokenizer.jsonlet answer =tryawait laya.answer(
state: "The T piece dropped at column 3 leaves one hole under it.",
question: .noul("Is this a clean placement?"))
print(answer.noul!) // P(true)
swift run -c release FluidUseLaya answer --state "…" --type choice \
--instructions "What does the customer want?" --options "refund|order status|technical help"
swift run -c release FluidUseLaya tetris # headless Tetris played by laya decisions
swift run -c release LayaTetrisDemo # SwiftUI demo
Same buckets with an int8 embedding table: 448–453 MB each, accuracy within 0.5 points of fp16 on the full benchmark
tokenizer.json
mmBERT / Gemma vocabulary (256k), byte fallback
Each fp16 bucket is a complete model (614 MB, 393 MB of which is the embedding table) with
32 option slots; the
e8
buckets store that table as int8 per-channel. Encoder-weight int8 and
6-/4-bit palettes fail the parity gates (the ANE in particular), so they are not published.
FluidUse
picks the smallest loaded bucket that fits a prompt and truncates
the state on the right for the largest one, exactly like laya's
max_len
.
Apple M5 Pro, macOS 27.0, 16 fixture questions vs. the unmodified PyTorch FP32 runtime:
16/16 argmax agreement on every bucket and compute-unit setting, max probability error 0.0021
(
ALL
) / 0.0126 (
CPU_AND_NE
). Per-question latency, warm:
Bucket
CPU + ANE
All units
L128
3.6 ms
3.9 ms
L256
9.9 ms
5.2 ms
L512
27.5 ms
9.0 ms
L1024
80.1 ms
17.9 ms
On laya's published application suites (3,899 questions, seed 13, rebuilt from upstream's scripts),
the Core ML buckets answered from Swift match the PyTorch reference's accuracy on every suite at
5.2 ms median per question (p95 18 ms):
Apache-2.0, following the upstream weights and code by Convai Innovations
(
NandhaKishorM/laya
). Independent conversion; not an
official Convai Innovations release.
Runs of FluidInference laya-coreml on huggingface.co
323
Total runs
12
24-hour runs
122
3-day runs
323
7-day runs
323
30-day runs
More Information About laya-coreml huggingface.co Model
laya-coreml huggingface.co is an AI model on huggingface.co that provides laya-coreml's model effect (), which can be used instantly with this FluidInference laya-coreml model. huggingface.co supports a free trial of the laya-coreml model, and also provides paid use of the laya-coreml. Support call laya-coreml model through api, including Node.js, Python, http.
laya-coreml huggingface.co is an online trial and call api platform, which integrates laya-coreml's modeling effects, including api services, and provides a free online trial of laya-coreml, you can try laya-coreml online for free by clicking the link below.
FluidInference laya-coreml online free url in huggingface.co:
laya-coreml is an open source model from GitHub that offers a free installation service, and any user can find laya-coreml on GitHub to install. At the same time, huggingface.co provides the effect of laya-coreml install, users can directly use laya-coreml installed effect in huggingface.co for debugging and trial. It also supports api for free installation.