Core ML conversion of Google's
SigLIP 2
google/siglip2-base-patch16-256
, with both towers:
an image encoder and a text encoder, each returning an L2-normalized 768-d embedding. Labels are plain text chosen
at run time, so it classifies images zero-shot. Google authored the Apache-2.0 source model; Fluid Inference
converted it. The image encoder runs entirely on the Apple Neural Engine.
Packages
Package
Input
Output
Size
siglip2-base-patch16-256-image-fp16.mlpackage
pixel_values
float32
[1, 3, 256, 256]
image_embeds
float32
[1, 768]
176 MB
siglip2-base-patch16-256-text-fp16.mlpackage
input_ids
int32
[1, 64]
text_embeds
float32
[1, 768]
539 MB
fp16 weights, macOS 14 / iOS 17 or newer.
config.json
holds the preprocessing and scoring constants:
Image:
resize to 256 × 256 (bilinear, antialiased like PIL), scale to [0, 1], then
(x − 0.5) / 0.5
.
Text:
lowercase, Gemma tokenizer (
tokenizer.json
), append
<eos>
, pad with
<pad>
to 64 tokens. SigLIP was
trained without an attention mask, so the padding must match.
Scores:
cos = image_embeds · text_embeds
.
sigmoid(logit_scale · cos + logit_bias)
(112.90, −16.77) gives
an independent probability per label; softmax or argmax over
cos
picks one label. Embed the labels once and
reuse them for every image.
Accuracy
ImageNet-1k zero-shot, all 50,000 test images
(
clip-benchmark/wds_imagenet1k
, prompt
this is a photo of {class}.
), fp16 on CPU + Neural Engine against the fp32 PyTorch model:
PyTorch fp32
Core ML fp16
Top-1
76.79%
76.76%
Same prediction as PyTorch
—
99.32%
Image embedding cosine, mean / min
—
0.99990 / 0.975
Google reports 79.1% with its own class names and prompts; the single-prompt protocol here is lower for both
backends.
Oxford-IIIT Pets test (3,669 photos, 37 breeds):
Core ML 94.77%, PyTorch 94.74%, 99.89% identical.
Speed and memory
Apple M5 Pro (24 GB), macOS 27. Image encoder alone, one call at a time:
5.2 ms
on CPU + Neural Engine
(100% of ops on the ANE), 3.5 ms CPU + GPU, versus 19.1 ms for PyTorch fp32 on MPS.
End to end on 7,349 Oxford-IIIT Pets photos (JPEG decode, resize, encoder, scoring against 37 cached breed
prompts), same photos and prompts, measured after model load:
FluidUse
wraps both packages (
SigLIP2Manager
, Swift Gemma tokenizer
with token ids identical to the Python tokenizer) and includes the photo-sorting demo
ImageSortDemo
.
Conversion scripts:
FluidInference/mobius
models/emb/siglip2/coreml
.
Runs of FluidInference siglip2-base-patch16-256-coreml on huggingface.co
92
Total runs
10
24-hour runs
32
3-day runs
74
7-day runs
74
30-day runs
More Information About siglip2-base-patch16-256-coreml huggingface.co Model
More siglip2-base-patch16-256-coreml license Visit here:
siglip2-base-patch16-256-coreml huggingface.co is an AI model on huggingface.co that provides siglip2-base-patch16-256-coreml's model effect (), which can be used instantly with this FluidInference siglip2-base-patch16-256-coreml model. huggingface.co supports a free trial of the siglip2-base-patch16-256-coreml model, and also provides paid use of the siglip2-base-patch16-256-coreml. Support call siglip2-base-patch16-256-coreml model through api, including Node.js, Python, http.
siglip2-base-patch16-256-coreml huggingface.co is an online trial and call api platform, which integrates siglip2-base-patch16-256-coreml's modeling effects, including api services, and provides a free online trial of siglip2-base-patch16-256-coreml, you can try siglip2-base-patch16-256-coreml online for free by clicking the link below.
FluidInference siglip2-base-patch16-256-coreml online free url in huggingface.co:
siglip2-base-patch16-256-coreml is an open source model from GitHub that offers a free installation service, and any user can find siglip2-base-patch16-256-coreml on GitHub to install. At the same time, huggingface.co provides the effect of siglip2-base-patch16-256-coreml install, users can directly use siglip2-base-patch16-256-coreml installed effect in huggingface.co for debugging and trial. It also supports api for free installation.
siglip2-base-patch16-256-coreml install url in huggingface.co: