AI SeedbankHelp preserve open and free AI for humanity's future

← All models

patrickjohncyh_fashion-clip

patrickjohncyh · View on Hugging Face ↗

Get this model

Download TorrentMagnet Link

Seeders: · Leechers:

Model card

The complete upstream card, rendered from this payload's README.md — the same hash-verified bytes the torrent distributes. Images and off-site links are removed; the original card on Hugging Face carries them.


license: mit tags:

  • vision
  • language
  • fashion
  • ecommerce library_name: transformers language:
  • en widget:
    • src: https://cdn-images.farfetch-contents.com/19/76/05/56/19760556_44221665_1000.jpg candidate_labels: black shoe, red shoe, a cat example_title: Black Shoe

Model Card: Fashion CLIP

Disclaimer: The model card adapts the model card from here.

Model Details

UPDATE (10/03/23): We have updated the model! We found that laion/CLIP-ViT-B-32-laion2B-s34B-b79K checkpoint (thanks Bin!) worked better than original OpenAI CLIP on Fashion. We thus fine-tune a newer (and better!) version of FashionCLIP (henceforth FashionCLIP 2.0), while keeping the architecture the same. We postulate that the perofrmance gains afforded by laion/CLIP-ViT-B-32-laion2B-s34B-b79K are due to the increased training data (5x OpenAI CLIP data). Our thesis, however, remains the same -- fine-tuning laion/CLIP on our fashion dataset improved zero-shot perofrmance across our benchmarks. See the below table comparing weighted macro F1 score across models.

Model FMNIST KAGL DEEP
OpenAI CLIP 0.66 0.63 0.45
FashionCLIP 0.74 0.67 0.48
Laion CLIP 0.78 0.71 0.58
FashionCLIP 2.0 0.83 0.73 0.62

FashionCLIP is a CLIP-based model developed to produce general product representations for fashion concepts. Leveraging the pre-trained checkpoint (ViT-B/32) released by OpenAI, we train FashionCLIP on a large, high-quality novel fashion dataset to study whether domain specific fine-tuning of CLIP-like models is sufficient to produce product representations that are zero-shot transferable to entirely new datasets and tasks. FashionCLIP was not developed for model deplyoment - to do so, researchers will first need to carefully study their capabilities in relation to the specific context they’re being deployed within.

Model Date

March 2023

Model Type

The model uses a ViT-B/32 Transformer architecture as an image encoder and uses a masked self-attention Transformer as a text encoder. These encoders are trained, starting from a pre-trained checkpoint, to maximize the similarity of (image, text) pairs via a contrastive loss on a fashion dataset containing 800K products.

Documents

  • FashionCLIP Github Repo
  • FashionCLIP Paper

Data

The model was trained on (image, text) pairs obtained from the Farfecth dataset[^1 Awaiting official release.], an English dataset comprising over 800K fashion products, with more than 3K brands across dozens of object types. The image used for encoding is the standard product image, which is a picture of the item over a white background, with no humans. The text used is a concatenation of the highlight (e.g., “stripes”, “long sleeves”, “Armani”) and short description (“80s styled t-shirt”)) available in the Farfetch dataset.

Limitations, Bias and Fiarness

We acknowledge certain limitations of FashionCLIP and expect that it inherits certain limitations and biases present in the original CLIP model. We do not expect our fine-tuning to significantly augment these limitations: we acknowledge that the fashion data we use makes explicit assumptions about the notion of gender as in "blue shoes for a woman" that inevitably associate aspects of clothing with specific people.

Our investigations also suggest that the data used introduces certain limitations in FashionCLIP. From the textual modality, given that most captions derived from the Farfetch dataset are long, we observe that FashionCLIP may be more performant in longer queries than shorter ones. From the image modality, FashionCLIP is also biased towards standard product images (centered, white background).

Model selection, i.e. selecting an appropariate stopping critera during fine-tuning, remains an open challenge. We observed that using loss on an in-domain (i.e. same distribution as test) validation dataset is a poor selection critera when out-of-domain generalization (i.e. across different datasets) is desired, even when the dataset used is relatively diverse and large.

Citation

@Article{Chia2022,
    title="Contrastive language and vision learning of general fashion concepts",
    author="Chia, Patrick John
            and Attanasio, Giuseppe
            and Bianchi, Federico
            and Terragni, Silvia
            and Magalh{\~a}es, Ana Rita
            and Goncalves, Diogo
            and Greco, Ciro
            and Tagliabue, Jacopo",
    journal="Scientific Reports",
    year="2022",
    month="Nov",
    day="08",
    volume="12",
    number="1",
    abstract="The steady rise of online shopping goes hand in hand with the development of increasingly complex ML and NLP models. While most use cases are cast as specialized supervised learning problems, we argue that practitioners would greatly benefit from general and transferable representations of products. In this work, we build on recent developments in contrastive learning to train FashionCLIP, a CLIP-like model adapted for the fashion industry. We demonstrate the effectiveness of the representations learned by FashionCLIP with extensive tests across a variety of tasks, datasets and generalization probes. We argue that adaptations of large pre-trained models such as CLIP offer new perspectives in terms of scalability and sustainability for certain types of players in the industry. Finally, we detail the costs and environmental impact of training, and release the model weights and code as open source contribution to the community.",
    issn="2045-2322",
    doi="10.1038/s41598-022-23052-9",
    url="https://doi.org/10.1038/s41598-022-23052-9"
}

Magnet link

Opens the swarm directly in your torrent client — no file download needed. Copy-paste works too:

magnet:?xt=urn:btih:1e8a2821a343bef1ac9bf1fc2dbc2a314ba7ab0b&dn=patrickjohncyh_fashion-clip

Open magnet in torrent client · infohash 1e8a2821a343bef1ac9bf1fc2dbc2a314ba7ab0b

Files & hashes

PathSizesha1sha256
README.md6.9 KB (7,020 B)67ce180512c13478e72676ffdedb57f7af5859767991e5d764cbff735a1c7c6f4ca7aac04e2533251c9ec5017ce994b51b0d3873
config.json4.4 KB (4,463 B)78237b6befcc388724e77ef62cbdb3137b3fa7fa79595b5c6be8867e3ba42c34b1e6b9a0f957e549c1e7f43b5c6945f4f6f59471
merges.txt512.4 KB (524,657 B)bbfec752c9a675946c6dce106def6f35c882dcc2f526393189112391ce6f9795d4695f704121ce452c3aad1f5335cc41337eba85
model.safetensors577.1 MB (605,157,890 B)14159894d6ee2e67ae6fb23815650be5b1f530054977e3a54929eccf065ce449aeaf296f0e5cb6b28e8798c3c97d67cb2f6dafc9
onnx/config.json455 B (455 B)7b247ea2df7b3b4c7b57787e4147e552fb7a9d9c00dba32f3053aac366851c20cc36c168001cdffba0e97651c1e92d7c486b6a5f
onnx/merges.txt512.3 KB (524,619 B)76e821f1b6f0a9709293c3b6b51ed90980b3166b9fd691f7c8039210e0fced15865466c65820d09b63988b0174bfe25de299051a
onnx/preprocessor_config.json468 B (468 B)36598b309c6f7a665f53b7cfff7fb7d69fecc43a5df7e578c37e907a431daf47fd592fc49fa50d23ed4c41285a0a34a58a9d2e06
onnx/special_tokens_map.json588 B (588 B)cf0682d6de72c1547f41b4f6d7c59f62deffef942cdb3b8331a60c92fc1e55a13e9fd61fd2293c5a51275fdcccd62b780052530e
onnx/tokenizer.json2.1 MB (2,224,119 B)bc1f77d20440541dd073ebae6f6c401087c7d34ef7f3b7af117d467b58374797691a6438d3e6b9e9cef800dfd5dced7f697a90cd
onnx/tokenizer_config.json772 B (772 B)41e7b7f0900fce2daec404ca68385fbdfc23602054ab04fbbb70952e10f0497b2412ccb3d0accaab9891ead4e9d3bb6154164b2d
onnx/vocab.json842.1 KB (862,328 B)182766ce89b439768edadda342519f33802f53645047b556ce86ccaf6aa22b3ffccfc52d391ea4accdab9c2f2407da5b742d4363
preprocessor_config.json316 B (316 B)5a12a1eb250987a4eee0e3e7d7338c4b22724be1910e70b3956ac9879ebc90b22fb3bc8a75b6a0677814500101a4c072bd7857bd
pytorch_model.bin577.2 MB (605,239,073 B)997a06134b1279966d71ea7bac3fbbe0b4abcd655adfac18a5eda0d68c975b9ddebc219836ca0280b37a1d0dd4e44725193a10b8
special_tokens_map.json389 B (389 B)9bfb42aa97dcd61e89f279ccaee988bccb4fabaef8c0d6c39aee3f8431078ef6646567b0aba7f2246e9c54b8b99d55c22b707cbf
tokenizer.json2.1 MB (2,224,041 B)564c0ebd5ce29c4ee4864004aee693deadd3128cb556ac8c99757ffb677208af34bc8c6721572114111a6e0aaf5fa69ff0b8d842
tokenizer_config.json568 B (568 B)ab0ca294e9ef2a950015496bc13b84e1bb462b09cbcdcf14cede6b5c3a6841f37961364357d82bef3079beb115e95f27252526a0
vocab.json842.1 KB (862,328 B)182766ce89b439768edadda342519f33802f53645047b556ce86ccaf6aa22b3ffccfc52d391ea4accdab9c2f2407da5b742d4363

Cite this release

Canonical URL
https://aiseedbank.org/models/patrickjohncyh_fashion-clip/
Slug
patrickjohncyh_fashion-clip
Infohash
1e8a2821a343bef1ac9bf1fc2dbc2a314ba7ab0b
License
mit
Signing key fingerprint
85a3b32c3712427b

Every file carries a locally computed sha256 — verify a download against the signed sums: patrickjohncyh_fashion-clip.SHA256SUMS (+ minisign signature).

Provenance

Upstream repositorypatrickjohncyh/fashion-clip
Revision (pinned)7e3ba62ce16b379a1ab479346b66f192e76f51b7
Fetched at2026-09-04T05:18:06Z
License at fetchmit
Snapshot toolhuggingface · seedbank 0.1.0

Trackers

✓ verified · rehash-vs-hf-metadata at 2026-09-04T05:18:17Z

mit1.13 GB (1,217,634,094 bytes)transformerspytorchonnxsafetensorsclipzero-shot-image-classificationvisionlanguagefashionecommerceendpoints_compatible1 language (en)