Ready-to-run deployment package for the Laya typed-decision model family on AX650 / NPU3.
Runtime: packaged
axllm run
Target: AX650 / AX8850, aarch64, NPU3
Graph shape: batch 1, sequence length 256, up to 4 options
Decision primitives:
choice
,
score
, and
noul
Included checkpoints: English, multilingual, and typed decisions
Included assets: AXModel files, Hugging Face tokenizers, runtime configs, sample requests, sample outputs, and
bin/axllm
Laya is a bidirectional decision model. It evaluates user-defined questions over text or structured JSON state without generating text.
Supported Platform
AX650 / AX8850
NPU3
Checkpoint Selection
Directory
Backbone
Recommended use
english/
ModernBERT-large
English routing, guardrails, moderation, and support triage
multilingual/
mmBERT-base
Chinese and other multilingual inputs
typed-decisions/
ModernBERT-large
Invoice, security, customer-service, and agent-trace workflows
The packaged runtime does not automatically route between checkpoints. Select the directory that matches the input language or workflow.
Model Inputs and Outputs
User-Level Input
Users submit a JSON object with two fields:
state
: the text or structured record to evaluate, such as a support ticket, security incident, email, or agent trace.
questions
: named decision definitions. Each question contains a decision
type
, an instruction, and the candidate criteria when applicable.
The runtime supports three decision types:
Type
Meaning
Returned value
choice
Select one label from 2 to 4 candidates
Selected label and probability distribution
score
Evaluate an ordered scale containing 2 to 4 levels
Expected score, probability distribution, and level legend
noul
Estimate whether a statement is true
Probability of
true
For every question,
axllm
serializes the request into this token sequence:
[CLS] <question type and instruction> [SEP]
[MASK] <option 0> [MASK] <option 1> ... [SEP]
<state> [SEP]
Each
[MASK]
position represents one candidate answer. A request containing four questions runs four forwards while keeping the same AXModel resident.
AXModel Tensor Interface
All three packaged AXModels use the same fixed tensor interface.
Tensor
Direction
Dtype
Shape
Meaning
input_ids
Input
S32
[1, 256]
Token IDs for the instruction, candidates, and state
attention_mask
Input
S32
[1, 256]
1
for valid tokens and
0
for padding
marker_pos
Input
S32
[1, 4]
Token position of each candidate's
[MASK]
marker
marker_mask
Input
S32
[1, 4]
Marks which of the four candidate slots are active
qtype
Input
S32
[1]
0
= choice,
1
= score,
2
= noul
logits
Output
FP32
[1, 4]
Candidate decision logits; only active candidate slots are used
act_logits
Output
FP32
[1, 2]
Auxiliary act-versus-escalate logits
The CPU postprocessor applies the packaged temperature values and returns:
choice
: the highest-probability label plus all active option probabilities.
score
: the expected ordinal value, computed as the probability-weighted level index.
noul
: the probability that the statement is true.
confidence
: distribution concentration for choice/score, or the larger side of the binary probability for noul.
action.act_probability
: the auxiliary head's tendency to act automatically. Validate this field on application data before using it as an automation gate.
npu_latency_ms
: NPU forward time for that question.
Performance
Measured on AX650 / NPU3 with
/opt/bin/ax_run_model
, two warmup runs and ten measured runs. Each AXModel invocation evaluates one question. Latency excludes model loading and CPU tokenization.
Checkpoint
Fixed input
Average NPU latency
english
B1, S256, K4
69.991 ms
multilingual
B1, S256, K4
27.722 ms
typed-decisions
B1, S256, K4
69.990 ms
Representative four-question
axllm
requests measured 280.4 ms for English, 110.9 ms for multilingual, and 280.3 ms for typed decisions. These totals are the sum of NPU forward time reported by the runtime.
Why the Multilingual AXModel Is Larger but Faster
AXModel file size and inference latency measure different costs.
Checkpoint
Vocabulary
Token-embedding parameters
Hidden size
Encoder layers
FFN intermediate size
english
50,368
51.6M
1,024
28
2,624
multilingual
256,000
196.6M
768
22
1,152
typed-decisions
50,368
51.6M
1,024
28
2,624
The multilingual checkpoint needs a much larger vocabulary to cover many languages. Its 256,000-by-768 token-embedding table is stored as 16-bit data in the compiled graph and accounts for a large part of the AXModel file and CMM footprint. During one inference, the embedding operator gathers only the rows referenced by the 256 input token IDs; it does not compute over all 256,000 vocabulary entries.
Most NPU time is spent in the encoder layers after the embedding lookup. The multilingual backbone has fewer layers, a smaller hidden dimension, and a much smaller feed-forward dimension. A rough projection-plus-feed-forward compute proxy is about 3.1 times larger for the English backbone, while the measured latency ratio is 2.52 times (69.991 / 27.722). Kernel scheduling and fixed operator overhead account for the difference between the rough compute ratio and measured time.
The larger multilingual file therefore reflects stored vocabulary capacity, while its lower latency reflects a lighter encoder computation.
Runtime Footprint
Checkpoint
AXModel file
Runtime CMM
english
481.13 MiB
481.63 MiB
multilingual
508.61 MiB
508.98 MiB
typed-decisions
481.13 MiB
481.63 MiB
Only one checkpoint needs to be loaded for a single
axllm
process. The complete package is approximately 1.48 GiB before repository metadata.
Every checkpoint directory is self-contained. Keep its config, AXModel, and tokenizer together.
Download the Package
mkdir -p AXERA-TECH/Laya
cd AXERA-TECH/Laya
hf download AXERA-TECH/Laya --local-dir .
Verify the downloaded runtime artifacts:
sha256sum -c SHA256SUMS
Run on the Board
The packaged binary uses the AX650 on-chip backend and the system AXERA runtime libraries.
chmod +x ./bin/axllm
Detailed Example: English Customer-Support Triage
This example evaluates a customer-support ticket. The customer reports a duplicate invoice charge, requests a refund today, and threatens to cancel the service.
Input
The
state
field contains the business record. The four entries under
questions
ask the model to select a handling department, assign an urgency score, detect refund intent, and detect cancellation risk.
{"state":{"body":"Invoice 4411 was charged twice. Please refund the duplicate today or we will cancel our plan."},"questions":{"department":{"type":"choice","instructions":"Which team should handle this request?","criteria":{"billing":"payments, invoices, charges, refunds","technical":"bugs, outages, API or account access failures","sales":"pricing, demos, contracts or purchases","other":"general requests or resolved issues"}},"urgency":{"type":"score","instructions":"How urgent is this request?","criteria":["routine, no deadline","needs attention soon","service blocked or deadline today"]},"refund":{"type":"noul","instructions":"Does the customer explicitly request a refund?"},"churn":{"type":"noul","instructions":"Does the customer threaten to cancel or leave?"}}}
Run the packaged English checkpoint:
./bin/axllm run ./english --input ./english/sample_request.json
AX650 Output
The following is the complete output recorded on the AX650 board:
{"model":"laya-english","answers":{"department":{"type":"choice","choice":"billing","probabilities":{"billing":0.9689645773705189,"technical":0.008899344656499457,"sales":0.010587676002932096,"other":0.011548401970049494},"confidence":0.8757531196892554,"action":{"act_probability":1.0},"npu_latency_ms":70.223844},"urgency":{"type":"score","probabilities":{"0":0.03266581800541651,"1":0.5424619734608962,"2":0.4248722085336873},"legend":{"0":"routine, no deadline","1":"needs attention soon","2":"service blocked or deadline today"},"score":1.3922063905282709,"confidence":0.2652274294222525,"action":{"act_probability":1.0},"npu_latency_ms":70.111539},"refund":{"type":"noul","noul":0.8548107451739729,"confidence":0.8548107451739729,"action":{"act_probability":1.0},"npu_latency_ms":70.021244},"churn":{"type":"noul","noul":0.844980400947017,"confidence":0.844980400947017,"action":{"act_probability":1.0},"npu_latency_ms":70.019742}},"total_npu_latency_ms":280.37636899999995}
How to Read the Output
department
is a
choice
decision. The selected label is
billing
with probability 0.9690. The remaining probabilities show how strongly the model rejected the technical, sales, and other queues.
department.confidence
is 0.8758. For choice and score questions, confidence measures how concentrated the probability distribution is.
urgency
is a
score
decision with levels 0, 1, and 2. The result 1.3922 is the probability-weighted expected level, between "needs attention soon" and "service blocked or deadline today."
refund.noul
is 0.8548, meaning an 85.48% estimated probability that the customer explicitly requested a refund.
churn.noul
is 0.8450, meaning an 84.50% estimated probability that the customer threatened to cancel or leave.
action.act_probability
comes from the auxiliary act-versus-escalate head. Validate this value on application data before using it to authorize an automatic action.
npu_latency_ms
is the NPU time for one question.
total_npu_latency_ms
is the sum for all four questions and excludes model loading and CPU tokenization.
A support system can use this result to route the ticket to billing, raise its priority, start a refund-review workflow, and notify a retention team. Production automation thresholds should be calibrated with representative application data.
Other Checkpoints
The other packaged checkpoints use the same request schema and output fields.
Checkpoint
Sample scenario
Run command
Representative result
multilingual
The same duplicate-charge request and questions written in Chinese
./bin/axllm run ./multilingual --input ./multilingual/sample_request.json
A request contains one
state
value and a named
questions
object.
{"state":{"body":"Invoice 4411 was charged twice. Please refund the duplicate today."},"questions":{"department":{"type":"choice","instructions":"Which team should handle this request?","criteria":{"billing":"payments, invoices, charges, refunds","technical":"bugs, outages, API or account failures","sales":"pricing, demos, contracts or purchases","other":"general requests"}},"urgency":{"type":"score","instructions":"How urgent is this request?","criteria":["routine","needs attention soon","service blocked or deadline today"]},"refund":{"type":"noul","instructions":"Does the customer explicitly request a refund?"}}}
Decision Primitives
Primitive
Input criteria
Output
choice
Object or list containing 2 to 4 options
Selected label and probability for each option
score
Ordered list containing 2 to 4 levels
Expected score, level probabilities, and legend
noul
Optional descriptions for false and true
Probability that the statement is true
Validated Examples
The package includes the exact requests and complete JSON outputs used for board validation.
The AX650 graphs in this release have fixed input shapes:
batch size: 1
sequence length: 256 tokens
maximum options: 4
input dtype: signed 32-bit integer
outputs: four decision logits and two action logits
The runtime performs one NPU forward per question. The upstream checkpoints support longer contexts and batched questions, but those layouts are outside this release.
axllm serve
is not exposed for these models; use
axllm run
.
Probabilities can shift after quantization or when the deployment domain differs from the checkpoint calibration data. Validate thresholds on representative application data before automating high-impact actions.
Conversion References
If you need the original model files or want to rebuild the deployment artifacts, start with:
The Laya source and model family are distributed under Apache License 2.0. The packaged
axllm
binary is based on BSD-3-Clause licensed code; its license text is included in
LICENSES/AXLLM-BSD-3-Clause.txt
.
Laya huggingface.co is an AI model on huggingface.co that provides Laya's model effect (), which can be used instantly with this AXERA-TECH Laya model. huggingface.co supports a free trial of the Laya model, and also provides paid use of the Laya. Support call Laya model through api, including Node.js, Python, http.
Laya huggingface.co is an online trial and call api platform, which integrates Laya's modeling effects, including api services, and provides a free online trial of Laya, you can try Laya online for free by clicking the link below.
AXERA-TECH Laya online free url in huggingface.co:
Laya is an open source model from GitHub that offers a free installation service, and any user can find Laya on GitHub to install. At the same time, huggingface.co provides the effect of Laya install, users can directly use Laya installed effect in huggingface.co for debugging and trial. It also supports api for free installation.