A foundation model for mobiles, wearables, robots, smart home, automotive and microcontrollers. The whole model is a single 9-35 MB file, and we trade general chat capacity to beat models 10x its size on mobile tool calls and match 2-3x bigger models on extraction.
Needle does three jobs, all of them on the device:
Tool calls
: given the functions your app exposes, Needle picks the right ones and fills every argument from what the user said. Ask for two things and you get two calls in order; ask for something no tool covers and you get an empty list, not a guess.
Structured extraction
: declare a shape, hand over messy text, get typed fields back: an invoice, a booking, a notification, a form. The decode grammar guarantees the output parses, and extraction generalises to classification.
Text embedding
: the same model returns a vector for a sentence, so an app can search, match and route locally.
Model
Needle 3 is a Laddered Simple Attention Network, our small-model recipe: a Monarch Hadamard MLP in place of the FFN, GQA attention with causal conv taps, engram n-gram memory read by gather, and multi-lane hyper-connections, trained so that every depth from 2 to 20 layers is a deployable model. Most of its parameters sit in the engram, so the 121M model does the arithmetic of a 50M one. The weights are compressed to CQ2-bit with Cactus Quants; a byte-level grammar compiled from your schemas constrains every token, and every response carries a calibrated confidence score from a learned head. The architecture diagram is on the
release page
. The repo holds the 20-layer
needle3.cact
, the
needle3.safetensors
checkpoint to fine-tune, and an engine per platform.
Benchmarks
Tool calling is exact-match accuracy on the full test splits, extraction is field micro-F1 on the full test splits.
The interactive frontier plot, the architecture and the fine-tuning results are at
cactuscompute.com/needle
.
import needle
@needle.tooldefget_weather(city: str):
"Get the current weather for a city."return {"city": city, "temp_c": 27, "sky": "clear"}
agent = needle.Needle(tools=[get_weather])
print(agent.run("what's it like in Lagos right now?")["results"])
# [{'city': 'Lagos', 'temp_c': 27, 'sky': 'clear'}]
Every turn returns one JSON object with
function_calls
, the model's
reasoning
and a calibrated
confidence
; an off-topic request returns an empty list rather than a guess. The engine and the weights are fetched from this repo once and cached.
Guides
How to design tools for Needle 3
: one tool per action, names users would say, formats in descriptions, constraints in the grammar, triggers.
The .cact format
: the file the engine maps and reads in place, Cactus Quants at 2.125 bits per weight, and how to parse it yourself.
Customisation
Needle was designed to be customised. Its capacity is a ladder, and a subnetwork as small as 2 layers, fine-tuned on one product's tools, runs optimally on devices far smaller than the full model needs. Fine-tuning on DroidCall lifts every subnetwork by 18 to 36 points, and from 4 layers up the tuned subnetwork passes DeepSeek V4 Flash, starting at 29M parameters.
The Python package fine-tunes with LoRA on the frozen base at the full 20 layers, then
needle build [--layers N]
merges the adapter, slices any subnetwork from 2 to 20 layers and exports a 4-bit
.cact
that runs on the same engine. The 2-bit post-training and quantisation behind the shipped model, enriched with Cactus proprietary datasets, run on the
Cactus Platform
.
Deploy
Every platform folder in this repo holds an engine under 1 MB that loads
needle3.cact
at start.
needle build --platform <folder> [--layers N]
fetches the engine and header and puts the weights beside them at any depth, or download the folder here:
./needle --model needle3.cact --tools tools.json --prompt "dim the living room to 30"
./needle --model needle3.cact --tools tools.json --serve
The
devices guide
covers the runner flags, the C API, the WASI component and air-gapped setup.
Citation
Needle 3 is built by the Cactus Compute team. If you use it in your work, please cite:
@misc{needle3_2026,
title = {Needle: Foundation Tool-Calling Model for Tiny Devices},
author = {Ndubuaku, Henry and Mosoyan, Karen and Mroz, Jakub and Cylich, Noah and
Kumar, Satyajit and Sandhu, Parkirat and Shemet, Roman and Lee, Justin H.},
year = {2026},
organization = {Cactus Compute, Inc.},
howpublished = {\url{https://github.com/cactus-compute/needle}}
}
Reach out on
[email protected]
for partnerships, collaborations, synergies and deploying Needle in your product.
Runs of Cactus-Compute needle3 on huggingface.co
69.7K
Total runs
0
24-hour runs
23.3K
3-day runs
69.6K
7-day runs
69.6K
30-day runs
More Information About needle3 huggingface.co Model
needle3 huggingface.co is an AI model on huggingface.co that provides needle3's model effect (), which can be used instantly with this Cactus-Compute needle3 model. huggingface.co supports a free trial of the needle3 model, and also provides paid use of the needle3. Support call needle3 model through api, including Node.js, Python, http.
needle3 huggingface.co is an online trial and call api platform, which integrates needle3's modeling effects, including api services, and provides a free online trial of needle3, you can try needle3 online for free by clicking the link below.
Cactus-Compute needle3 online free url in huggingface.co:
needle3 is an open source model from GitHub that offers a free installation service, and any user can find needle3 on GitHub to install. At the same time, huggingface.co provides the effect of needle3 install, users can directly use needle3 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.