BrainboxAI / nitzotz

huggingface.co
Total runs: 149
24-hour runs: 0
7-day runs: 147
30-day runs: 147
Model's Last Updated: September 29 2026
zero-shot-classification

Introduction of nitzotz

Model Details of nitzotz

Nitzotz: Hebrew decision model

Scam detection accuracy 82.5%, 0.05 s per question, 413 MB, Apache-2.0

What it is. Nitzotz reads a Hebrew message and answers questions you type about it: pick one of several options, give a score on a scale, or say yes or no to a claim. For every answer it gives a probability you can trust, so you know when it is sure and when it is guessing. It does not write text, so it cannot make things up. It runs on a normal laptop, with no internet connection and no cost per question.

What it is for. Deciding what to do with incoming messages: is this a scam, what kind of message is it, which department should get it, how urgent is it. It is not a chatbot and it is not built for long documents (see Limitations ).

In numbers. On 298 Hebrew messages it says correctly whether a message is a scam 82.5% of the time. For context: 70% of those messages are not scams, so a model that always says "not a scam" would score 70.5%. The number that matters more is the ranking: it gives real scams a higher probability than legitimate messages 88% of the time (AUC 0.88).

Try it

Python (the laya library, pip install laya ):

import laya
agent = laya.load("BrainboxAI/nitzotz")
q = {"scam": {"type": "noul",
              "instructions": "ההודעה מנסה לגרום לנמען ללחוץ על קישור, לשלם או למסור פרטים בלי סיבה לגיטימית.",
              "criteria": {"true": "כן, זה ניסיון מרמה",
                           "false": "לא, זו הודעה לגיטימית, גם אם יש בה קישור או בקשת תשלום"}}}
print(agent.predict("החבילה שלך מעוכבת. לשחרור שלם 12.90 בקישור", q)["answers"]["scam"]["noul"])

Output: 0.8613 , the probability that the claim ("the message tries to make the reader click a link, pay or hand over details for no legitimate reason") is true.

The wording of the question matters. This is the exact question the scam test used, and the numbers on this card are for it. In our tries a shorter wording ("the message is a scam attempt") gave clearly worse answers, for example a high scam probability for a plain verification code. If you change the wording, check it on your own messages first.

laya.exe (the standalone binary from ggmlc releases , no Python needed). Download the Q8 file, start the daemon, then send one JSON request per line; each answer comes back on one line:

hf download BrainboxAI/nitzotz nitzotz-q8_0.gguf --local-dir .
laya.exe daemon nitzotz-q8_0.gguf --device vulkan
{"id":"1","state":"היי, הפגישה מחר ב-10 עדיין בתוקף?","questions":{"scam":{"type":"noul","instructions":"ההודעה מנסה לגרום לנמען ללחוץ על קישור, לשלם או למסור פרטים בלי סיבה לגיטימית.","criteria":{"true":"כן, זה ניסיון מרמה","false":"לא, זו הודעה לגיטימית, גם אם יש בה קישור או בקשת תשלום"}}}}
{"status":"ready","model":"laya"}
{"model": "nitzotz", "family": "nitzotz", "route": "forced nitzotz", "answers": {"scam": {"type": "noul", "action": {"act_probability": 0.1511}, "confidence": 0.8785, "noul": 0.1215}}, "usage": {"input_tokens": 66, "output_tokens": 0, "latency_ms": 117.02}, "id": "1"}

noul is the probability that the claim is true. Several questions in one request are answered together in one pass. Use --device cpu on a machine without a GPU.

Benchmarks

8 frozen test sets, 3,990 questions in total, locked by checksum before any training data existed. All three models answered exactly the same questions. The two reference points are the other open laya models that read Hebrew: RoeiG/laya-hebrew (licence CC-BY-NC-SA-4.0, non-commercial use only ) and laya-multilingual (Apache-2.0).

Accuracy per test

Accuracy shape per model

Full table. Accuracy, the best in each row in bold. The last two columns say whether Nitzotz's difference from that model is real or could be luck (a paired exact McNemar test on the same questions): "better" or "worse" means p < 0.05, "tie" means the difference could be chance.

Test (questions) Nitzotz RoeiG (non-commercial) laya-multilingual Chance Nitzotz vs RoeiG Nitzotz vs laya-multilingual
Scam or not? (298 messages) 82.5% 32.2% 50.3% 50.0% better
p<0.001
better
p<0.001
same, hard cases only (65) 63.1% 32.3% 35.4% 50.0% better
p<0.001
better
p=0.008
Message type, 6 options (298) 75.5% 52.7% 20.8% 16.7% better
p<0.001
better
p<0.001
same, hard cases only (65) 53.8% 44.6% 24.6% 16.7% tie
p=0.34
better
p<0.001
Support ticket type, 5 options (30) 83.3% 80.0% 56.7% 20.0% tie
p=1.00
tie
p=0.06
Ticket urgency, 5 levels (30) 60.0% 36.7% 33.3% 20.0% tie
p=0.17
tie
p=0.12
Paying customer? yes/no (30) 63.3% 56.7% 50.0% 50.0% tie
p=0.50
tie
p=0.54
Voice command intent, 20 options (500) 88.8% 72.6% 47.4% 5.0% better
p<0.001
better
p<0.001
Voice command intent, 4 options (500) 96.2% 90.8% 69.4% 25.0% better
p<0.001
better
p<0.001
News topic, 7 options (204) 79.4% 82.3% 66.2% 14.3% tie
p=0.42
better
p<0.001
Does the passage support this answer? (600) 93.8% 94.2% 48.8% 50.0% tie
p=0.90
better
p<0.001
Plausible answer the passage does not give (600) 91.0% 53.0% 53.7% 50.0% better
p<0.001
better
p<0.001
Reading comprehension, 4 options (900) 52.7% 75.4% 31.4% 25.0% worse
p<0.001
better
p<0.001

On Belebele reading comprehension Nitzotz scores 52.7%, below RoeiG/laya-hebrew (75.4%); Nitzotz is built for message decisions, not long-passage comprehension.

How to read it:

  • MASSIVE (voice commands): Nitzotz was trained on MASSIVE's training commands. The test commands are different, but written by the same people in the same style, so this test is easier for Nitzotz than for the others.
  • HeQ : Nitzotz was trained on HeQ's training passages. The test passages are different ones, and training passages that overlapped a test passage were removed.
  • The support-ticket rows have only 30 questions. No difference there is reliable.
  • The spam and scam test has a lean : 70% of its messages are not scams. The "chance" line (50%) is a coin flip, not the best blind strategy.
Why you can trust it

Calibration: predicted confidence against actual accuracy

The probabilities mean something. Take all the answers where Nitzotz said it was about 70 to 80% sure, and count how many were right: the chart above does that for every confidence level, over 3,990 test questions. Where the dots sit above the line, Nitzotz is more often right than it claims (it is modest); where they sit below, it is over-confident. This is what lets you set thresholds (see the next section). The per-question-type temperatures were fitted on 2,901 held-out training items, never on the test sets ( calibration.json ).

It reads the text. With the message removed and only the question left, its accuracy falls to 70.5% on the scam question and 7.7% on message type. So the answers come from the message, not from the wording of the question.

Speed against file size

It is fast on ordinary hardware. On one laptop (Intel Core Ultra 9 285H laptop, built-in Arc 140T GPU, Windows 11), one question at a time:

Runtime Median over 50 check questions (about 135 tokens) Short message (38 tokens) Long input (430 tokens)
laya.exe, Q8 file, GPU (Vulkan) 49 ms 35 ms 135 ms
laya.exe, F16 file, GPU (Vulkan) 52 ms 41 ms 141 ms
Python (laya), GPU (PyTorch XPU) 54 ms 44 ms 177 ms
laya.exe, Q8 file, CPU only (16 threads) 335 ms 177 ms 1233 ms
Python (laya), CPU only 387 ms 130 ms 1260 ms

The short and long columns repeat one fixed input 30 times after 5 warm-up calls. Timings on this laptop change a lot from one session to another (an earlier measurement of the same setup was several times slower), so treat these as rough.

The GGUF files give almost the same answers as the Python model. Compared with the full-precision Python model on the CPU:

File, device Same top answer, 50 questions Largest probability gap Same top answer, 596 spam questions Largest gap Average gap
Q8, GPU 50/50 0.0079 593/596 0.0152 0.00217
F16, GPU 50/50 0.0026 595/596 0.0023 0.00032
Q8, CPU 50/50 0.0094 590/596 0.0275 0.00321
F16, CPU 50/50 0.0017 595/596 0.0014 0.00027

A gap of 0.01 means, for example, 0.83 against 0.84. The few questions where the top answer changes are ones where the two best answers were almost tied. Over both spam questions the Q8 file on the GPU is right 78.9% of the time, against 79.0% for the Python model; the F16 file is closer (79.2%).

Use it in your business

Ramzor: message to Nitzotz to green, yellow or red

The idea ("Ramzor", traffic light). Every incoming message gets one or more quick questions. The probability decides what happens next:

  • Green (confident it is fine): handle automatically, file it, tag it, route it.
  • Yellow (not sure): send it to a person, or to a large language model if you use one.
  • Red (confident it is a scam): block or quarantine it.

The drawing is a worked example on the 298 test messages. The upper threshold, 0.52, is the one chosen for the scam question on 700 held-out training messages (never on the test); the lower one, 0.35, was picked by hand. With them, 183 messages go to green (16 of them are in fact scams), 48 to yellow (21 scams), and 67 to red (16 of them are in fact legitimate). So red should mean "quarantine and check", not "delete". Pick your own thresholds on a sample of your own messages, and decide how many mistakes in green and red you can live with. At the 0.52 threshold alone, the scam answer is right 82.2% of the time overall and 64.6% on the 65 hard cases (at 0.50: 82.9% and 64.6%).

The model is cheap enough to run on every message. The person (or the LLM) only sees the yellow part. Do not use it as the only line of defence for decisions that can hurt someone.

Training data and transparency

Two training stages, on a rented GPU:

  1. Learning to read. HalleluBERT-large was first trained to find the answer to a question inside a passage, on 27,085 HeQ training questions (CC BY 4.0). Passages that overlapped a test passage were removed.
  2. Learning to decide. A laya decision head was put on top and the whole model was trained on 58,000 items (55,099 for training, 2,901 held out to pick the best of 2 passes and to fit the temperatures).
Source Items Licence How it was made
Synthetic: Message type, 6 classes 14,000 teacher output, project-owned (DeepSeek MIT; Gemma Apache-2.0) written by DeepSeek V4.1 Flash, labelled independently by both models
Synthetic: Routing to a department (3 to 6 options) 8,750 teacher output, project-owned (DeepSeek MIT; Gemma Apache-2.0) written by DeepSeek V4.1 Flash, labelled independently by both models
Synthetic: Yes/no claims about a message 8,750 teacher output, project-owned (DeepSeek MIT; Gemma Apache-2.0) written by DeepSeek V4.1 Flash, labelled independently by both models
Synthetic: Urgency, 5 levels 3,500 teacher output, project-owned (DeepSeek MIT; Gemma Apache-2.0) written by DeepSeek V4.1 Flash, labelled independently by both models
HeQ: is the proposed answer supported (yes/no) 6,000 CC BY 4.0 the dataset's gold labels, train split only
HeQ: plausible answer to an unanswerable question 5,000 CC BY 4.0 the dataset's gold labels, train split only
HeQ: 4-option reading 5,000 CC BY 4.0 the dataset's gold labels, train split only
MASSIVE he-IL: voice command intent 7,000 CC BY 4.0 the dataset's gold labels, train split only
  • Synthetic messages (35,000 items). DeepSeek V4.1 Flash (MIT) wrote Israeli-style SMS, WhatsApp and email messages from a plan (intended label, topic, tone, varied fake phone numbers and links). DeepSeek and Gemma 4 31B (Apache-2.0) then each labelled every message on their own. A message was kept only if both agreed; the two models agreed on 94% of them, and 93.7% of the 37,429 written messages were kept. Both ran through DeepInfra (via OpenRouter), with data retention off and "thinking" mode off. The synthetic data is not published.
  • Open data. HeQ v1.1 (CC BY 4.0; about half of its questions are on Geektime articles, shared by the HeQ authors under the same licence) and MASSIVE he-IL (CC BY 4.0). Training splits only.
  • No leaks from the tests. Every training item was compared with every test question; anything sharing an 8-word run with a test text was dropped, and so were MASSIVE commands equal to a test command.
  • Not used: no output of closed commercial chatbots, no DICTA model, no non-commercial or share-alike data.
  • Size: encoder 357.1M parameters (HalleluBERT-large, MIT, fine-tuned); decision head 26.5M parameters, trained from scratch.

About the tests.

Test Questions Source and licence
Spam and business messages 596 (298 messages, 2 questions each) in-house, Israeli SMS, WhatsApp and email style, 65 hard cases
Support tickets (triage30) 90 (30 tickets, 3 questions each) in-house
MASSIVE he-IL, 20 and 4 options 500 + 500 MASSIVE test split, CC BY 4.0
SIB-200 news topic 204 CC BY-SA 4.0, used for testing only
Belebele reading 900 CC BY-SA 4.0, used for testing only
HeQ verify and unanswerable 600 + 600 HeQ v1.1 test split, CC BY 4.0

Caveats that change how much to trust the numbers:

  • The spam and business test was written by an AI model and checked by an AI model, not by a person. Real inboxes will look different. It was frozen after that review (4 labels changed, 2 messages removed).
  • Two training runs, two random seeds; this is one of them. It was picked on held-out training items (94.8% against 94.6%, a small gap), not on the tests. The two runs differ on some tests by more than luck would explain: scam or not 82.5% here against 68.8% in the other run, news topic 79.4% against 74.0%. Expect a new training run to move these two numbers by several points.
  • The training labels come from two AI models. Where both are wrong in the same way, Nitzotz learned their mistake.
  • HeQ's wrong answers in the test were picked by code, not checked by a person.
Limitations
  • It is weak at reading comprehension of longer passages. Belebele: 52.7%, where a blind guess gets 25%. When the right answer is written word for word in the passage it gets 72% (252 questions); when the answer is said in other words it gets 45% (648 questions). Do not ask it whether a long document supports a claim.
  • Checking an answer against a passage (HeQ) works. 93.8% with the passage; with the passage removed it falls to 50.0%, a coin flip. So on this kind of question it really reads the passage.
  • Numbers, dates, amounts and rules : not trained and not measured. Compute them in code and pass the result in.
  • Hard cases are the weak spot : scams written to look legitimate (a "supplier" changing bank details, the "CEO" asking for a transfer) and real messages that look like scams (a real bank alert with a link, a real verification code). On the 65 hard cases the scam question is right 63.1% of the time, which is below the 70.8% you would get by always answering "not a scam" there. Its ranking on those cases is also weak (AUC 0.63).
  • Urgency is subjective. Even the two teacher models matched the intended urgency only about 62 to 64% of the time.
  • Sarcasm and irony are probably read literally. Not measured.
  • 512 tokens (roughly 300 to 400 Hebrew words) per question. A longer message is cut from the end without a warning.
  • Hebrew only. Not trained or tested on English or Arabic.
  • Not a safety system on its own. It makes mistakes in both directions. Keep a person in the loop for anything that can hurt someone.
Licence and attribution

Apache-2.0 for the weights, the GGUF files and the code. Commercial use is allowed. Built on:

  • HalleluBERT-large : the encoder, MIT.
  • laya (NandhaKishorM, Convai Innovations): the decision-head architecture and the runtime, Apache-2.0. No laya weights are used.
  • HeQ : CC BY 4.0, by Webiks for MAFAT and the Israeli National NLP Program (NNLP-IL); includes Geektime passages.
  • MASSIVE : CC BY 4.0, Amazon (FitzGerald et al., 2022).
  • DeepSeek V4.1 Flash (MIT) and Gemma 4 31B (Apache-2.0), used through DeepInfra as data writer and labellers.
  • ggmlc for the GGUF files.
  • Belebele and SIB-200 (CC BY-SA 4.0) were used only to test, never to train.

The full notice is in NOTICE .


ניצוץ, בעברית

ניצוץ: מודל החלטות בעברית

דיוק בזיהוי הונאות 82.5%, 0.05 שניות לשאלה, 413 MB, Apache-2.0

מה זה. ניצוץ קורא הודעה בעברית ועונה על שאלות שאתם מקלידים עליה: לבחור אחת מכמה אפשרויות, לתת ציון בסולם, או לענות כן או לא על טענה. על כל תשובה הוא נותן הסתברות שאפשר לסמוך עליה, כך שיודעים מתי הוא בטוח ומתי הוא מנחש. הוא לא כותב טקסט, ולכן הוא לא יכול להמציא דברים. הוא רץ על מחשב נייד רגיל, בלי אינטרנט ובלי תשלום על כל שאלה.

בשביל מה. להחליט מה עושים עם הודעות נכנסות: האם זו הונאה, איזה סוג הודעה זו, לאיזו מחלקה להעביר, כמה זה דחוף. זה לא צ'אטבוט, והוא לא בנוי למסמכים ארוכים (ראו מגבלות ).

במספרים. על 298 הודעות בעברית הוא קובע נכון אם ההודעה היא הונאה ב-82.5% מהמקרים. בשביל פרופורציה: 70% מההודעות האלה הן לא הונאה, כך שמודל שתמיד עונה "לא הונאה" היה מקבל 70.5%. המספר שחשוב יותר הוא הדירוג: הוא נותן להונאה אמיתית הסתברות גבוהה יותר מאשר להודעה תקינה ב-88% מהמקרים (AUC 0.88).

לנסות

בפייתון (הספרייה laya , pip install laya ):

import laya
agent = laya.load("BrainboxAI/nitzotz")
q = {"scam": {"type": "noul",
              "instructions": "ההודעה מנסה לגרום לנמען ללחוץ על קישור, לשלם או למסור פרטים בלי סיבה לגיטימית.",
              "criteria": {"true": "כן, זה ניסיון מרמה",
                           "false": "לא, זו הודעה לגיטימית, גם אם יש בה קישור או בקשת תשלום"}}}
print(agent.predict("החבילה שלך מעוכבת. לשחרור שלם 12.90 בקישור", q)["answers"]["scam"]["noul"])

הפלט: 0.8613 , ההסתברות שהטענה ("ההודעה מנסה לגרום לנמען ללחוץ על קישור, לשלם או למסור פרטים בלי סיבה לגיטימית") נכונה.

הניסוח של השאלה משנה. זו בדיוק השאלה שבה השתמש מבחן ההונאות, והמספרים בכרטיס הזה הם עליה. בניסיונות שלנו ניסוח קצר יותר ("ההודעה היא ניסיון הונאה") נתן תשובות גרועות בהרבה, למשל הסתברות גבוהה להונאה לקוד אימות רגיל. אם משנים את הניסוח, בודקים אותו קודם על ההודעות שלכם.

בלי פייתון, עם laya.exe (תוכנה עצמאית מ- ggmlc ). מורידים את קובץ Q8, מפעילים, ושולחים בקשת JSON אחת בכל שורה. כל תשובה חוזרת בשורה אחת:

hf download BrainboxAI/nitzotz nitzotz-q8_0.gguf --local-dir .
laya.exe daemon nitzotz-q8_0.gguf --device vulkan
{"id":"1","state":"היי, הפגישה מחר ב-10 עדיין בתוקף?","questions":{"scam":{"type":"noul","instructions":"ההודעה מנסה לגרום לנמען ללחוץ על קישור, לשלם או למסור פרטים בלי סיבה לגיטימית.","criteria":{"true":"כן, זה ניסיון מרמה","false":"לא, זו הודעה לגיטימית, גם אם יש בה קישור או בקשת תשלום"}}}}
{"status":"ready","model":"laya"}
{"model": "nitzotz", "family": "nitzotz", "route": "forced nitzotz", "answers": {"scam": {"type": "noul", "action": {"act_probability": 0.1511}, "confidence": 0.8785, "noul": 0.1215}}, "usage": {"input_tokens": 66, "output_tokens": 0, "latency_ms": 117.02}, "id": "1"}

noul היא ההסתברות שהטענה נכונה. כמה שאלות בבקשה אחת נענות יחד, במעבר אחד. על מחשב בלי כרטיס מסך משתמשים ב- --device cpu .

מבחנים

8 סטים של מבחן, 3,990 שאלות בסך הכול, שננעלו בטביעת אצבע לפני שנוצר פריט אימון אחד. שלושת המודלים ענו על אותן שאלות בדיוק. שתי נקודות ההשוואה הן מודלי laya הפתוחים האחרים שקוראים עברית: RoeiG/laya-hebrew (רישיון CC-BY-NC-SA-4.0, לשימוש לא מסחרי בלבד ) ו- laya-multilingual (Apache-2.0).

דיוק לפי מבחן

הצורה של כל מודל

הטבלה המלאה. אחוז התשובות הנכונות, הטוב ביותר בכל שורה מודגש. שתי העמודות האחרונות אומרות אם ההבדל של ניצוץ מהמודל הזה אמיתי או שאולי זה מזל (מבחן McNemar מדויק על אותן שאלות בדיוק): "טוב יותר" או "חלש יותר" פירושו p קטן מ-0.05, "תיקו" פירושו שההבדל יכול להיות מקרי.

מבחן (מספר שאלות) ניצוץ RoeiG (לא מסחרי) laya-multilingual ניחוש ניצוץ מול RoeiG ניצוץ מול laya-multilingual
הונאה או לא? (298 הודעות) 82.5% 32.2% 50.3% 50.0% טוב יותר
p<0.001
טוב יותר
p<0.001
אותו דבר, רק המקרים הקשים (65) 63.1% 32.3% 35.4% 50.0% טוב יותר
p<0.001
טוב יותר
p=0.008
סוג ההודעה, 6 אפשרויות (298) 75.5% 52.7% 20.8% 16.7% טוב יותר
p<0.001
טוב יותר
p<0.001
אותו דבר, רק המקרים הקשים (65) 53.8% 44.6% 24.6% 16.7% תיקו
p=0.34
טוב יותר
p<0.001
סוג פניית תמיכה, 5 אפשרויות (30) 83.3% 80.0% 56.7% 20.0% תיקו
p=1.00
תיקו
p=0.06
דחיפות הפנייה, 5 רמות (30) 60.0% 36.7% 33.3% 20.0% תיקו
p=0.17
תיקו
p=0.12
לקוח משלם? כן/לא (30) 63.3% 56.7% 50.0% 50.0% תיקו
p=0.50
תיקו
p=0.54
כוונת פקודה קולית, 20 אפשרויות (500) 88.8% 72.6% 47.4% 5.0% טוב יותר
p<0.001
טוב יותר
p<0.001
כוונת פקודה קולית, 4 אפשרויות (500) 96.2% 90.8% 69.4% 25.0% טוב יותר
p<0.001
טוב יותר
p<0.001
נושא של ידיעה, 7 אפשרויות (204) 79.4% 82.3% 66.2% 14.3% תיקו
p=0.42
טוב יותר
p<0.001
האם הקטע תומך בתשובה? (600) 93.8% 94.2% 48.8% 50.0% תיקו
p=0.90
טוב יותר
p<0.001
תשובה סבירה שהקטע לא נותן (600) 91.0% 53.0% 53.7% 50.0% טוב יותר
p<0.001
טוב יותר
p<0.001
הבנת הנקרא, 4 אפשרויות (900) 52.7% 75.4% 31.4% 25.0% חלש יותר
p<0.001
טוב יותר
p<0.001

בהבנת הנקרא של Belebele ניצוץ מקבל 52.7%, פחות מ-RoeiG/laya-hebrew (75.4%). ניצוץ בנוי להחלטות על הודעות, לא להבנה של קטעים ארוכים.

איך לקרוא את זה:

  • MASSIVE (פקודות קוליות): ניצוץ אומן על פקודות האימון של MASSIVE. פקודות המבחן אחרות, אבל נכתבו בידי אותם אנשים ובאותו סגנון, ולכן המבחן הזה קל יותר לניצוץ מאשר לאחרים.
  • HeQ : ניצוץ אומן על קטעי האימון של HeQ. קטעי המבחן אחרים, וקטעי אימון שחפפו לקטע מבחן הוסרו.
  • בשורות של פניות התמיכה יש רק 30 שאלות. שום הבדל שם לא אמין.
  • מבחן הספאם וההונאות לא מאוזן : 70% מההודעות בו הן לא הונאה. קו ה"ניחוש" (50%) הוא הטלת מטבע, לא האסטרטגיה העיוורת הטובה ביותר.
למה אפשר לסמוך עליו

כיול: הביטחון שהוא מצהיר מול כמה פעמים הוא צדק

להסתברויות יש משמעות. קחו את כל התשובות שבהן ניצוץ אמר שהוא בטוח בערך ב-70 עד 80%, וספרו כמה מהן היו נכונות. הגרף עושה את זה לכל רמת ביטחון, על 3,990 שאלות מבחן. כשהנקודות מעל הקו, ניצוץ צודק יותר ממה שהוא אומר (הוא צנוע). כשהן מתחת, הוא בטוח בעצמו יותר מדי. זה מה שמאפשר לקבוע ספים (בפרק הבא). הכיול נעשה על 2,901 פריטי אימון שהופרדו מראש, אף פעם לא על המבחן ( calibration.json ).

הוא באמת קורא את הטקסט. כשמוחקים את ההודעה ומשאירים רק את השאלה, הדיוק יורד ל-70.5% בשאלת ההונאה ול-7.7% בסוג ההודעה. כלומר התשובות באות מההודעה, לא מהניסוח של השאלה.

מהירות מול גודל

הוא מהיר על חומרה רגילה. על מחשב נייד עם מעבד Intel Core Ultra 9 285H וכרטיס המסך המובנה Arc 140T, ווינדוס 11, שאלה אחת בכל פעם:

איך מריצים חציון על 50 שאלות בדיקה (בממוצע 135 טוקנים) הודעה קצרה (38 טוקנים) קלט ארוך (430 טוקנים)
laya.exe, קובץ Q8, כרטיס מסך (Vulkan) 49 ms 35 ms 135 ms
laya.exe, קובץ F16, כרטיס מסך (Vulkan) 52 ms 41 ms 141 ms
פייתון (laya), כרטיס מסך (PyTorch XPU) 54 ms 44 ms 177 ms
laya.exe, קובץ Q8, מעבד בלבד (16 תהליכונים) 335 ms 177 ms 1233 ms
פייתון (laya), מעבד בלבד 387 ms 130 ms 1260 ms

בעמודות של ההודעה הקצרה והקלט הארוך אותה שאלה רצה 30 פעמים, אחרי 5 הרצות חימום. הזמנים על המחשב הזה משתנים הרבה בין הפעלה להפעלה (מדידה קודמת של אותה הגדרה יצאה איטית פי כמה), אז אלה מספרים בקירוב.

קובצי ה-GGUF נותנים כמעט את אותן תשובות כמו מודל הפייתון. בהשוואה למודל הפייתון המלא על המעבד:

קובץ, מכשיר אותה תשובה מובילה, 50 שאלות הפרש הסתברות מרבי אותה תשובה מובילה, 596 שאלות ספאם הפרש מרבי הפרש ממוצע
Q8, כרטיס מסך 50/50 0.0079 593/596 0.0152 0.00217
F16, כרטיס מסך 50/50 0.0026 595/596 0.0023 0.00032
Q8, מעבד 50/50 0.0094 590/596 0.0275 0.00321
F16, מעבד 50/50 0.0017 595/596 0.0014 0.00027

פער של 0.01 פירושו, למשל, 0.83 מול 0.84. השאלות המעטות שבהן התשובה המובילה משתנה הן כאלה שבהן שתי התשובות הטובות היו כמעט שוות. בשתי שאלות הספאם יחד קובץ Q8 על כרטיס המסך צודק ב-78.9%, מול 79.0% למודל הפייתון. קובץ F16 קרוב יותר (79.2%).

שימוש בעסק

רמזור: הודעה, ניצוץ, ירוק צהוב או אדום

הרעיון ("רמזור"). כל הודעה נכנסת מקבלת שאלה מהירה אחת או כמה. ההסתברות מחליטה מה קורה הלאה:

  • ירוק (בטוח שזה בסדר): טיפול אוטומטי, תיוק, תיוג, ניתוב.
  • צהוב (לא בטוח): לבדיקה של אדם, או של מודל שפה גדול אם אתם משתמשים בו.
  • אדום (בטוח שזו הונאה): חסימה או הסגר.

השרטוט הוא דוגמה על 298 הודעות המבחן. הסף העליון, 0.52, הוא הסף שנבחר לשאלת ההונאה על 700 הודעות אימון שהופרדו מראש (אף פעם לא על המבחן). הסף התחתון, 0.35, נבחר ביד. איתם 183 הודעות הולכות לירוק (16 מהן הן בעצם הונאה), 48 לצהוב (21 הונאות), ו-67 לאדום (16 מהן בעצם תקינות). כלומר אדום צריך להיות "הסגר ובדיקה", לא "מחיקה". בחרו ספים משלכם על מדגם של ההודעות שלכם, והחליטו כמה טעויות בירוק ובאדום אתם מוכנים לקבל. בסף 0.52 לבדו, תשובת ההונאה נכונה ב-82.2% מהמקרים בסך הכול וב-64.6% על 65 המקרים הקשים (בסף 0.50: 82.9% ו-64.6%).

המודל זול מספיק כדי להריץ אותו על כל הודעה. האדם (או מודל השפה) רואה רק את החלק הצהוב. אל תשתמשו בו כקו הגנה יחיד בהחלטות שיכולות לפגוע במישהו.

נתוני האימון ושקיפות

שני שלבי אימון, על כרטיס מסך שכור:

  1. ללמוד לקרוא. HalleluBERT-large אומן קודם למצוא את התשובה לשאלה בתוך קטע, על 27,085 שאלות אימון של HeQ (CC BY 4.0). קטעים שחפפו לקטע מבחן הוסרו.
  2. ללמוד להחליט. מעליו הונח ראש החלטות של laya, וכל המודל אומן על 58,000 פריטים (55,099 לאימון, ו-2,901 הופרדו מראש כדי לבחור את הטוב מבין 2 מעברים ולכייל את הטמפרטורות).
מקור פריטים רישיון איך נוצר
סינתטי: סוג הודעה, 6 סוגים 14,000 הפלט של המורה שייך לנו (DeepSeek, MIT; Gemma, Apache-2.0) נכתב בידי DeepSeek V4.1 Flash, סומן בנפרד בידי שני המודלים
סינתטי: ניתוב למחלקה (3 עד 6 אפשרויות) 8,750 הפלט של המורה שייך לנו (DeepSeek, MIT; Gemma, Apache-2.0) נכתב בידי DeepSeek V4.1 Flash, סומן בנפרד בידי שני המודלים
סינתטי: טענות כן/לא על הודעה 8,750 הפלט של המורה שייך לנו (DeepSeek, MIT; Gemma, Apache-2.0) נכתב בידי DeepSeek V4.1 Flash, סומן בנפרד בידי שני המודלים
סינתטי: דחיפות, 5 רמות 3,500 הפלט של המורה שייך לנו (DeepSeek, MIT; Gemma, Apache-2.0) נכתב בידי DeepSeek V4.1 Flash, סומן בנפרד בידי שני המודלים
HeQ: האם התשובה המוצעת נתמכת (כן/לא) 6,000 CC BY 4.0 תוויות הזהב של המאגר, מפיצול האימון בלבד
HeQ: תשובה סבירה לשאלה שאין לה תשובה 5,000 CC BY 4.0 תוויות הזהב של המאגר, מפיצול האימון בלבד
HeQ: קריאה, 4 אפשרויות 5,000 CC BY 4.0 תוויות הזהב של המאגר, מפיצול האימון בלבד
MASSIVE he-IL: כוונת פקודה קולית 7,000 CC BY 4.0 תוויות הזהב של המאגר, מפיצול האימון בלבד
  • הודעות סינתטיות (35,000 פריטים). DeepSeek V4.1 Flash (MIT) כתב הודעות SMS, וואטסאפ ומייל בסגנון ישראלי לפי תוכנית (התווית המתוכננת, נושא, משלב, מספרי טלפון וקישורים מזויפים ומגוונים). אחר כך DeepSeek ו-Gemma 4 31B (Apache-2.0) סימנו כל הודעה, כל אחד לבד. הודעה נשמרה רק אם שניהם הסכימו. הם הסכימו על 94% מההודעות, ונשמרו 93.7% מתוך 37,429 ההודעות שנכתבו. שניהם רצו דרך DeepInfra (דרך OpenRouter), בלי שמירת נתונים ובלי מצב "חשיבה". הנתונים הסינתטיים לא מפורסמים.
  • נתונים פתוחים. HeQ v1.1 (CC BY 4.0. בערך חצי מהשאלות שלו על כתבות של Geektime, ומחברי HeQ משתפים אותם באותו רישיון) ו-MASSIVE he-IL (CC BY 4.0). רק מפיצולי האימון.
  • בלי דליפה מהמבחנים. כל פריט אימון הושווה לכל שאלת מבחן. כל מה שחולק רצף של 8 מילים עם טקסט מבחן נזרק, וגם פקודות MASSIVE שזהות לפקודת מבחן.
  • מה לא שימש: שום פלט של צ'אטבוט מסחרי סגור, שום מודל של DICTA, שום נתונים לא מסחריים או ברישיון "שיתוף זהה".
  • גודל: המקודד 357.1 מיליון פרמטרים (HalleluBERT-large, MIT, אומן מחדש). ראש ההחלטות 26.5 מיליון פרמטרים, אומן מאפס.

על המבחנים.

מבחן שאלות מקור ורישיון
הודעות ספאם ועסקים 596 (298 הודעות, 2 שאלות לכל אחת) פנימי, בסגנון SMS, וואטסאפ ומייל ישראלי, 65 מקרים קשים
פניות תמיכה (triage30) 90 (30 פניות, 3 שאלות לכל אחת) פנימי
MASSIVE he-IL, 20 ו-4 אפשרויות 500 + 500 פיצול המבחן של MASSIVE, CC BY 4.0
SIB-200, נושא ידיעה 204 CC BY-SA 4.0, רק לבדיקה
Belebele, הבנת הנקרא 900 CC BY-SA 4.0, רק לבדיקה
HeQ, אימות ו"אין תשובה" 600 + 600 פיצול המבחן של HeQ v1.1, CC BY 4.0

הסתייגויות שמשנות כמה לסמוך על המספרים:

  • מבחן הספאם והעסקים נכתב בידי מודל AI ונבדק בידי מודל AI, לא בידי אדם. תיבות דואר אמיתיות ייראו אחרת. הוא ננעל אחרי הבדיקה הזו (4 תוויות שונו, 2 הודעות הוסרו).
  • שתי ריצות אימון, עם שני זרעים אקראיים, וזו אחת מהן. היא נבחרה לפי פריטי אימון שהופרדו מראש (94.8% מול 94.6%, פער קטן), לא לפי המבחנים. בכמה מבחנים שתי הריצות שונות יותר ממה שמזל מסביר: הונאה או לא 82.5% כאן מול 68.8% בריצה השנייה, נושא ידיעה 79.4% מול 74.0%. ריצת אימון חדשה יכולה להזיז את שני המספרים האלה בכמה נקודות.
  • תוויות האימון באות משני מודלי AI. איפה ששניהם טועים באותו אופן, ניצוץ למד את הטעות שלהם.
  • התשובות השגויות של HeQ במבחן נבחרו בקוד, ולא נבדקו בידי אדם.
מגבלות
  • הוא חלש בהבנת הנקרא של קטעים ארוכים. Belebele: 52.7%, כשניחוש עיוור מקבל 25%. כשהתשובה הנכונה כתובה בקטע מילה במילה הוא מקבל 72% (252 שאלות). כשהתשובה נאמרת במילים אחרות הוא מקבל 45% (648 שאלות). אל תשאלו אותו אם מסמך ארוך תומך בטענה.
  • בדיקה אם קטע תומך בתשובה (HeQ) עובדת. 93.8% עם הקטע. כשמוחקים את הקטע זה יורד ל-50.0%, הטלת מטבע. כלומר בשאלות מהסוג הזה הוא באמת קורא את הקטע.
  • מספרים, תאריכים, סכומים וכללים : לא אומן עליהם ולא נמדד. חשבו אותם בקוד והעבירו את התוצאה.
  • המקרים הקשים הם נקודת התורפה : הונאות שנכתבו כדי להיראות לגיטימיות ("ספק" שמחליף פרטי חשבון בנק, "המנכ"ל" שמבקש העברה), והודעות אמיתיות שנראות כמו הונאה (התראה אמיתית מהבנק עם קישור, קוד אימות אמיתי). על 65 המקרים הקשים שאלת ההונאה צודקת ב-63.1% מהמקרים. זה פחות מה-70.8% שהיה יוצא אם תמיד עונים שם "לא הונאה", וגם הדירוג שלו במקרים האלה חלש (AUC 0.63).
  • דחיפות היא עניין סובייקטיבי. אפילו שני מודלי המורה הסכימו עם הדחיפות המתוכננת רק בכ-62 עד 64% מהמקרים.
  • ציניות ואירוניה כנראה נקראות מילולית. לא נמדד.
  • 512 טוקנים (בערך 300 עד 400 מילים בעברית) לשאלה. הודעה ארוכה יותר נחתכת מהסוף בלי אזהרה.
  • עברית בלבד. לא אומן ולא נבדק על אנגלית או ערבית.
  • זו לא מערכת הגנה לבד. הוא טועה לשני הכיוונים. השאירו אדם בתהליך בכל דבר שיכול לפגוע במישהו.
רישיון וקרדיטים

Apache-2.0 למשקולות, לקובצי ה-GGUF ולקוד. מותר לשימוש מסחרי. בנוי על:

  • HalleluBERT-large : המקודד, MIT.
  • laya (NandhaKishorM, Convai Innovations): הארכיטקטורה של ראש ההחלטות וסביבת ההרצה, Apache-2.0. לא נעשה שימוש במשקולות של laya.
  • HeQ : CC BY 4.0, של Webiks עבור מפא"ת ותוכנית ה-NLP הלאומית (NNLP-IL). כולל קטעים מ-Geektime.
  • MASSIVE : CC BY 4.0, אמזון (FitzGerald ואחרים, 2022).
  • DeepSeek V4.1 Flash (MIT) ו- Gemma 4 31B (Apache-2.0), דרך DeepInfra, ככותב הנתונים וכמסמנים.
  • ggmlc לקובצי ה-GGUF.
  • Belebele ו-SIB-200 (CC BY-SA 4.0) שימשו רק לבדיקה, אף פעם לא לאימון.

ההודעה המלאה בקובץ NOTICE .

Runs of BrainboxAI nitzotz on huggingface.co

149
Total runs
0
24-hour runs
10
3-day runs
147
7-day runs
147
30-day runs

More Information About nitzotz huggingface.co Model

More nitzotz license Visit here:

https://choosealicense.com/licenses/apache-2.0

nitzotz huggingface.co

nitzotz huggingface.co is an AI model on huggingface.co that provides nitzotz's model effect (), which can be used instantly with this BrainboxAI nitzotz model. huggingface.co supports a free trial of the nitzotz model, and also provides paid use of the nitzotz. Support call nitzotz model through api, including Node.js, Python, http.

BrainboxAI nitzotz online free

nitzotz huggingface.co is an online trial and call api platform, which integrates nitzotz's modeling effects, including api services, and provides a free online trial of nitzotz, you can try nitzotz online for free by clicking the link below.

BrainboxAI nitzotz online free url in huggingface.co:

https://huggingface.co/BrainboxAI/nitzotz

nitzotz install

nitzotz is an open source model from GitHub that offers a free installation service, and any user can find nitzotz on GitHub to install. At the same time, huggingface.co provides the effect of nitzotz install, users can directly use nitzotz installed effect in huggingface.co for debugging and trial. It also supports api for free installation.

nitzotz install url in huggingface.co:

https://huggingface.co/BrainboxAI/nitzotz

Url of nitzotz

Provider of nitzotz huggingface.co

BrainboxAI
ORGANIZATIONS

Other API from BrainboxAI