SigLIP2-so400m image encoder, Core ML, batch of 8

The image tower of SigLIP2-so400m converted to Core ML (fp16) with a fixed batch of 8 images per call, so it runs on Apple GPUs at full utilisation. The weights are Google's, untouched โ€” only the format changed. Built for Retriever, a local semantic search over personal video archives on macOS.

Converted from the ONNX export onnx-community/siglip2-so400m-patch14-384-ONNX (vision_model_fp16.onnx) with onnx2torch and coremltools 9.

Input and output

Input pixel_values float32 tensor [8, 3, 384, 384], normalised as (x/255 โˆ’ 0.5)/0.5 (the model keeps the top-left 378ร—378 patch grid, like the ONNX export)
Output embedding pooled image features, [8, 1152], not normalised

Pad a partial batch with zeros and discard those rows.

Fidelity and speed (M2 Pro, GPU)

Against the ONNX export on 31 frames from four videos: cosine 1.0000 mean, 0.9999 minimum; identical top-3 and top-10 in text-to-image retrieval. 166 ms per image in batches of 8, versus 204 ms one image at a time with the batch-1 package.

Downloads last month
4
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for antonlnz/siglip2-so400m-image-coreml-batch8

Quantized
(11)
this model