Introduction of FunctionGemma-270M-CoreML-INT8-FP32
Model Details of FunctionGemma-270M-CoreML-INT8-FP32
FunctionGemma 270M Mobile Exports
FunctionGemma is a Gemma 3 270M variant trained for local function calling. It
is intended to translate user text into structured tool calls, then optionally
turn the tool result into a short user-facing response.
Setup
Accept the
google/functiongemma-270m-it
license on Hugging Face, then
authenticate before running conversion:
export HF_TOKEN=hf_...
cd models/functiongemma/export
poetry env use /opt/homebrew/bin/python3.11
poetry install --with convert
LiteRT-LM Export
poetry run python convert.py --output-dir ./functiongemma-litert
The default export uses LiteRT Torch's
dynamic_wi8_afp32
quantization recipe,
prefill lengths
128,512,1024
, and a
1024
token KV cache. For a larger
mobile prompt budget:
The CoreML artifact is a fixed 128-token last-logits model. It uses int8
weights with float32 compute because the float16 compute export produced NaN
logits in local validation. This CoreML path does full-context recompute for
each generated token; LiteRT-LM remains the preferred production path for tool
calling latency.
poetry run python benchmark.py --backend litert --runs 5 --warmup 1
poetry run python benchmark.py --backend coreml --coreml-compute-units cpu --runs 5 --warmup 1
poetry run python benchmark.py --backend coreml --coreml-compute-units cpu_and_ne --runs 5 --warmup 1
Local results on this machine:
Backend
Quantization
Load RSS Δ
Peak RSS Δ
Mean tok/s
LiteRT-LM CPU
dynamic int8
551.1 MB
865.3 MB
148.54
CoreML CPU
int8 weights, fp32 compute
658.0 MB
1690.4 MB
31.49
CoreML CPU+NE
int8 weights, fp32 compute
86.7 MB
1129.8 MB
32.82
Runtime Loop
The model should be used in two passes:
Build a prompt with
format_tool_call_prompt(...)
and stop on
<end_function_call>
or
<start_function_response>
.
Parse the returned call with
parse_function_calls(...)
, validate it against
an allowlist, and execute the tool.
Build a second prompt with
format_final_response_prompt(...)
and stop on
<end_of_turn>
to get the final user-facing answer.
For command-only actions, the app can skip the second pass and present its own
deterministic UI response after the tool succeeds.
FunctionGemma is trained for single-turn and parallel tool calls. Do not rely on
it for multi-step dependency chains without app-side orchestration or fine-tuning.
The LiteRT-LM Python runtime currently returns FunctionGemma calls as raw text,
for example:
Use
parse_function_calls(...)
to validate and dispatch the call. After the
tool response is sent back as a
tool_response
turn, the same exported model can
produce the final user-facing answer.
Mobile Artifacts
Ship these files:
functiongemma-litert/model.litertlm
functiongemma-litert/config.json
Do not ship local runtime caches such as
model.litertlm.xnnpack_cache_*
; they are regenerated by LiteRT.
Runs of aufklarer FunctionGemma-270M-CoreML-INT8-FP32 on huggingface.co
9
Total runs
0
24-hour runs
2
3-day runs
2
7-day runs
7
30-day runs
More Information About FunctionGemma-270M-CoreML-INT8-FP32 huggingface.co Model
FunctionGemma-270M-CoreML-INT8-FP32 huggingface.co is an AI model on huggingface.co that provides FunctionGemma-270M-CoreML-INT8-FP32's model effect (), which can be used instantly with this aufklarer FunctionGemma-270M-CoreML-INT8-FP32 model. huggingface.co supports a free trial of the FunctionGemma-270M-CoreML-INT8-FP32 model, and also provides paid use of the FunctionGemma-270M-CoreML-INT8-FP32. Support call FunctionGemma-270M-CoreML-INT8-FP32 model through api, including Node.js, Python, http.
FunctionGemma-270M-CoreML-INT8-FP32 huggingface.co is an online trial and call api platform, which integrates FunctionGemma-270M-CoreML-INT8-FP32's modeling effects, including api services, and provides a free online trial of FunctionGemma-270M-CoreML-INT8-FP32, you can try FunctionGemma-270M-CoreML-INT8-FP32 online for free by clicking the link below.
aufklarer FunctionGemma-270M-CoreML-INT8-FP32 online free url in huggingface.co:
FunctionGemma-270M-CoreML-INT8-FP32 is an open source model from GitHub that offers a free installation service, and any user can find FunctionGemma-270M-CoreML-INT8-FP32 on GitHub to install. At the same time, huggingface.co provides the effect of FunctionGemma-270M-CoreML-INT8-FP32 install, users can directly use FunctionGemma-270M-CoreML-INT8-FP32 installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
FunctionGemma-270M-CoreML-INT8-FP32 install url in huggingface.co: