AI SeedbankHelp preserve open and free AI for humanity's future

← All models

google_siglip2-base-patch16-256

google · View on Hugging Face ↗

Get this model

Download TorrentMagnet Link

Seeders: · Leechers:

Model card

The complete upstream card, rendered from this payload's README.md — the same hash-verified bytes the torrent distributes. Images and off-site links are removed; the original card on Hugging Face carries them.


license: apache-2.0 tags:


SigLIP 2 Base

SigLIP 2 extends the pretraining objective of SigLIP with prior, independently developed techniques into a unified recipe, for improved semantic understanding, localization, and dense features.

Intended uses

You can use the raw model for tasks like zero-shot image classification and image-text retrieval, or as a vision encoder for VLMs (and other vision tasks).

Here is how to use this model to perform zero-shot image classification:

from transformers import pipeline

# load pipeline
ckpt = "google/siglip2-base-patch16-256"
image_classifier = pipeline(model=ckpt, task="zero-shot-image-classification")

# load image and candidate labels
url = "http://images.cocodataset.org/val2017/000000039769.jpg"
candidate_labels = ["2 cats", "a plane", "a remote"]

# run inference
outputs = image_classifier(image, candidate_labels)
print(outputs)

You can encode an image using the Vision Tower like so:

import torch
from transformers import AutoModel, AutoProcessor
from transformers.image_utils import load_image

# load the model and processor
ckpt = "google/siglip2-base-patch16-256"
model = AutoModel.from_pretrained(ckpt, device_map="auto").eval()
processor = AutoProcessor.from_pretrained(ckpt)

# load the image
image = load_image("https://huggingface.co/datasets/merve/coco/resolve/main/val2017/000000000285.jpg")
inputs = processor(images=[image], return_tensors="pt").to(model.device)

# run infernece
with torch.no_grad():
    image_embeddings = model.get_image_features(**inputs)    

print(image_embeddings.shape)

For more code examples, we refer to the siglip documentation.

Training procedure

SigLIP 2 adds some clever training objectives on top of SigLIP:

  1. Decoder loss
  2. Global-local and masked prediction loss
  3. Aspect ratio and resolution adaptibility

Training data

SigLIP 2 is pre-trained on the WebLI dataset (Chen et al., 2023).

Compute

The model was trained on up to 2048 TPU-v5e chips.

Evaluation results

Evaluation of SigLIP 2 is shown below (taken from the paper).

BibTeX entry and citation info

@misc{tschannen2025siglip2multilingualvisionlanguage,
      title={SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features}, 
      author={Michael Tschannen and Alexey Gritsenko and Xiao Wang and Muhammad Ferjad Naeem and Ibrahim Alabdulmohsin and Nikhil Parthasarathy and Talfan Evans and Lucas Beyer and Ye Xia and Basil Mustafa and Olivier Hénaff and Jeremiah Harmsen and Andreas Steiner and Xiaohua Zhai},
      year={2025},
      eprint={2502.14786},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2502.14786}, 
}

Magnet link

Opens the swarm directly in your torrent client — no file download needed. Copy-paste works too:

magnet:?xt=urn:btih:4e5fc11e5b66f4f4cd8577cefb6b22cfc97e02d7&dn=google_siglip2-base-patch16-256

Open magnet in torrent client · infohash 4e5fc11e5b66f4f4cd8577cefb6b22cfc97e02d7

Files & hashes

PathSizesha1sha256
README.md3.3 KB (3,375 B)806ddccbbd58fb973834f317fb38e7584957d9c53f093559c3a18d94d9336249cf2e4e42ec8752413f07eb6634e2820e4ea120ad
config.json276 B (276 B)b0a6b7804332a660ed5d6c2707b39bed801920107b5aedcb8893e31376e129c1ffd7a5392f1a806dbc793ce53eda220c2ec59edf
model.safetensors1.40 GB (1,500,985,224 B)9e1b588822e81e82590e595b3487feb45304b1cf6125cacc01fa93bdc98a0c5101cefcd69b2ed1f8ab4f38d86f4ad5984f5dc863
preprocessor_config.json394 B (394 B)130986149eee6a6fd9a2eb53da12fbc1a415c0e6d14ba2ee3fd816f3de8abaddc31953565128eaf37c73ad4bed32101a98465aff
special_tokens_map.json636 B (636 B)8d6368f7e735fbe4781bf6e956b7c6ad0586df80baec30ea10906f16adb8c18af7a34023002c1746542612b8b41c9f09e1351351
tokenizer.json32.8 MB (34,363,039 B)c43ca480742b934e0ecfadb6fa365f5b1ea64612cb9140fae3ac5122c972d37adf83e1248471a38147ad76f8215c8872c6fd8322
tokenizer.model4.0 MB (4,241,003 B)bbd7e417640374364c78eb9686d9bdc2fec4da9361a7b147390c64585d6c3543dd6fc636906c9af3865a5548f27f31aee1d4c8e2
tokenizer_config.json46.1 KB (47,164 B)d97c5412159422c3b56fbc99076b1dcc25dd785614afe629fe4959b9e0d51e1852b8d9f7ad074f90a1a7125a4fcdd17f06e78fc8

Cite this release

Canonical URL
https://aiseedbank.org/models/google_siglip2-base-patch16-256/
Slug
google_siglip2-base-patch16-256
Infohash
4e5fc11e5b66f4f4cd8577cefb6b22cfc97e02d7
License
apache-2.0
Signing key fingerprint
85a3b32c3712427b

Every file carries a locally computed sha256 — verify a download against the signed sums: google_siglip2-base-patch16-256.SHA256SUMS (+ minisign signature).

Provenance

Upstream repositorygoogle/siglip2-base-patch16-256
Revision (pinned)3f9f96cb90da5dbc758b01813f2f6f1aee24c1ab
Fetched at2026-09-04T00:25:42Z
License at fetchapache-2.0
Snapshot toolhuggingface · seedbank 0.1.0

Trackers

✓ verified · rehash-vs-hf-metadata at 2026-09-04T00:26:38Z

apache-2.01.43 GB (1,539,641,111 bytes)transformerssafetensorssiglipvisionzero-shot-image-classificationendpoints_compatiblepaper: 2502.14786paper: 2303.15343paper: 2209.06794