CIR Datasets Browser

← i-CIR project pagei-CIR dataset browser →
One browser for the composed image retrieval benchmarks: real queries, their ground truth, and what text-only, image-only and composed retrieval actually return.
dataset query

How to read a card

Every card is one composed query of the chosen dataset: the image query (reference image) on the left and the text query exactly as the dataset provides it on the right. Below the line:

Each row states at which rank the first ground-truth positive appears. Candidates are ranked against the dataset's official database (or the query's own gallery where the dataset defines one), excluding the reference image. Numbers in the "about this dataset" panel are computed on the sampled queries only.

image queryground-truth positiveretrieved, not labelled positive

About this browser

Composed image retrieval (CIR) is evaluated on a handful of benchmarks whose annotations are rarely looked at. This browser shows up to 200 queries per dataset, sampled at random (stratified where the dataset has categories, fixed seed), with their complete ground truth and with what three simple retrievers return, so that anyone can judge for themselves whether a query is solvable from the text alone, whether the positives are complete, and how hard the negatives really are.

Thumbnails are 320px reductions of the original images and are shown for research and commentary; each dataset's images remain under their original licenses and belong to their respective owners. Datasets are listed with their official source. To have an image removed or to report an issue, use the report channel of the i-CIR project page.

Retrievers: text-only and image-only use SigLIP2 ViT-SO400M-16-384 (webli) features; composed uses BASIC (CLIP ViT-L/14, full preset) from Psomas et al., Instance-Level Composed Image Retrieval, NeurIPS 2025. All rankings are zero-shot; no method was tuned on any of these datasets.

loading…