object instances landmarks, products, fictional characters, tech, artβ¦
composed queries an image query paired with a modification text
database images 99.5% of them curated hard negatives
curated hard negatives median per instance database
The 202 instances of i-CIR, one query photograph each β click a photo to open that instance in the dataset browser.
Composed image retrieval at the instance level: a photo says which object, a text says how it should look.
image query
positivesame temple, in an old archival photo
visual hard negativelook-alike temple, wrong modification
textual hard negativean old archival photo, but of something else
composed hard negativeold archival photo of a look-alike templeThe query has two parts. An image query shows a particular object instance, here the Temple of Hephaestus in Athens, and a modification text says how that object should appear in the images we are looking for, here βin an old archival photoβ. Neither part is enough on its own: the photo does not say what should change, and the text does not say which temple.
What counts as correct. A database image is a positive only if it shows the same instance under the requested modification. Everything else is a negative, and a method is scored by how high it ranks the positives (mean Average Precision).
What makes it hard. Each instance database is filled with near-misses that satisfy only part of the query: visual hard negatives show the same or a look-alike object without the modification (a colour photo of the Parthenon), textual hard negatives match the text but show a different object (some other old archival photo), and composed hard negatives almost satisfy both (an old archival photo of a look-alike temple). A method that leans on one modality alone falls for one of these three.
Explore it yourself. Every one of the 202 instances comes with its own image queries, modification texts, positives and thousands of such curated negatives. Open this instance in the dataset browser, or browse all 202.
A retrieval benchmark that asks a precise question: which images show this very object, changed in this way?
object instances landmarks, products, fictional characters, tech, artβ¦
composed queries an image query paired with a modification text
database images 99.5% of them curated hard negatives
curated hard negatives median per instance database
The 202 instances by visual category (three levels). Hover a slice for its name and count.
i-CIR follows an instance-level class definition: two images belong to the same class only if they show the same particular object, the Temple of Hephaestus or Batman rather than βa Greek templeβ or βa superheroβ. A composed query pairs a photo of an instance with a short text describing a modification, and a method must rank the images that show that instance under that modification above everything else.
The 518 (instance, modification) pairs by modification type. Hover a segment for examples.
1 to 25 photos, median 3
1 to 5 texts, median 2
951 to 10,045, median 3,421
1 to 127, median 5
The plain Text Γ Image baseline reaches 17.48 mAP on i-CIR. Swap the curated negatives for random LAION images and it takes more than 40 million of them to drag the same baseline down to that level: four orders of magnitude more than the 3.7K images an i-CIR instance database holds on average, and 1.5 orders of magnitude more than the 750K images of the whole benchmark. The random-distractor curve is even a lower bound, since unlabeled LAION images inevitably contain false negatives.
Sweep a weight Ξ» from text-only (Ξ» = 0) to image-only (Ξ» = 1) similarity for three simple fusion rules. On i-CIR every rule peaks strictly inside the interval, with a composition gain of +14.9 mAP (+490%) over the best single modality, averaged over the three rules.
The gain shrinks to +3.0 mAP on CIRR, +5.0 on FashionIQ and +6.8 on CIRCO, and CIRR and FashionIQ are best served by the text alone. i-CIR rewards methods that genuinely combine both modalities.
BASIC, a Baseline Approach for Surprisingly strong Composition: training-free, built on frozen VLM features, and it never touches the stored database index.
A composed query is treated as a logical AND: a database image must be similar to the image query and match the text query. BASIC scores the two modalities separately, cleans each similarity of modality-specific noise, and fuses them multiplicatively so that images strong in only one modality are pushed down.
Subtract a mean image feature (computed on LAION) and a mean text feature (computed on the object corpus). This strips generic βimagenessβ and βtextnessβ and leaves the semantic content.
Two LLM-generated corpora define what to keep and what to drop: Cβ, object names such as dog or building, and Cβ, styles, viewpoints and settings such as cartoon, aerial view or in a cloudy day. The top-k eigenvectors of (1βΞ±)Cβ β Ξ±Cβ, built from text-feature covariances, form a projection P that keeps object identity and suppresses style (k = 250, Ξ± = 0.2).
Short queries like during sunset are out of distribution for CLIP, which was trained on captions. Each query is wrapped with random object terms from Cβ (dog during sunset, sculpture dog), the phrases are embedded, centred and averaged.
The image query is enriched with a similarity-weighted combination of its top-ranked database features. It helps on class-level datasets and slightly hurts on i-CIR, so both variants are reported.
Each similarity is rescaled by its empirical minimum, sΜ = (s β smin) / |smin|, so that the two modalities live on comparable ranges. The final score is sΜf = sΜvΒ·sΜt β Ξ» (sΜv + sΜt)2: the product rewards images relevant to both queries, the Harris-style penalty pushes down images that are strong in one modality only (Ξ» = 0.1).
Centering and projection fold into the query: sv = β¨xv, PPβ€(qv β ΞΌv)β© β c(q), where the last term is a query-dependent constant. The database keeps its plain CLIP features, any FAISS index works unchanged, and the corpora, and therefore the projection, can be swapped per application without re-indexing.
mAP (%) while components are switched on one by one (top) and switched off individually (bottom). Bold: best per column.
| Centering | Min norm. | Harris | Context. | Projection | Q. exp. | ImageNet-R | NICO++ | MiniDN | LTLL | i-CIR | |
|---|---|---|---|---|---|---|---|---|---|---|---|
| β | β | β | β | β | β | 7.66 | 9.26 | 9.48 | 19.78 | 17.48 | Text Γ Image |
| β | β | β | β | β | β | 12.16 | 9.95 | 12.16 | 16.93 | 28.33 | |
| β | β | β | β | β | β | 12.06 | 17.20 | 17.72 | 22.20 | 27.30 | |
| β | β | β | β | β | β | 16.21 | 15.06 | 17.79 | 29.70 | 28.42 | |
| β | β | β | β | β | β | 18.61 | 15.34 | 21.01 | 33.74 | 33.48 | |
| β | β | β | β | β | β | 27.54 | 28.90 | 35.75 | 38.22 | 34.35 | BASIC β |
| β | β | β | β | β | β | 32.13 | 31.65 | 39.58 | 41.38 | 31.64 | BASIC |
| β | β | β | β | β | β | 17.31 | 13.96 | 21.22 | 22.42 | 31.78 | no projection |
| β | β | β | β | β | β | 26.18 | 30.61 | 33.64 | 34.50 | 25.85 | no contextualization |
| β | β | β | β | β | β | 30.75 | 29.82 | 38.85 | 40.65 | 31.61 | no Harris |
| β | β | β | β | β | β | 24.50 | 22.74 | 29.65 | 19.36 | 30.75 | plain product |
β without query expansion, the configuration that works best on i-CIR. All rows use CLIP ViT-L/14; hyper-parameters were fixed once on a small private development set.
Mean Average Precision on i-CIR and on four domain-conversion CIR datasets. On i-CIR, mAP is computed per instance and averaged over instances (macro-mAP), so every object counts equally.
| Method | ImageNet-R | NICO++ | MiniDN | LTLL | i-CIR |
|---|---|---|---|---|---|
| Text | 0.74 | 1.09 | 0.57 | 5.72 | 3.01 |
| Image | 3.84 | 6.32 | 6.66 | 16.49 | 3.04 |
| Text + Image | 6.21 | 9.30 | 9.33 | 17.86 | 8.20 |
| Text Γ Image | 7.83 | 9.79 | 9.86 | 23.16 | 17.48 |
| WeiCom | 10.47 | 10.54 | 8.52 | 26.60 | 18.03 |
| Pic2Word | 7.88 | 9.76 | 12.00 | 21.27 | 19.36 |
| CompoDiff | 12.88 | 10.32 | 22.95 | 21.61 | 9.63 |
| CIReVL | 18.11 | 17.80 | 26.20 | 32.60 | 18.66 |
| SEARLE | 14.04 | 15.13 | 21.78 | 25.46 | 19.90 |
| MCL | 8.13 | 19.09 | 18.41 | 16.67 | 19.89 |
| MagicLens | 9.13 | 19.66 | 20.06 | 24.21 | 27.35 |
| CoVR-2 | 11.52 | 24.93 | 27.76 | 24.68 | 28.50 |
| FreeDom | 29.91 | 26.10 | 37.27 | 33.24 | 17.24 |
| FreeDom β | 25.81 | 23.24 | 32.14 | 30.82 | 15.76 |
| BASIC | 32.13 | 31.65 | 39.58 | 41.38 | 31.64 |
| BASIC β | 27.54 | 28.90 | 35.75 | 38.22 | 34.35 |
Average mAP (%). β without query expansion. On i-CIR the reported number is macro-mAP: mAP is computed per instance and then averaged over the 202 instances. All methods use CLIP ViT-L/14 except CompoDiff (ViT-G/14). Click a column header to sort.
Leading on class-level datasets does not carry over to i-CIR: FreeDom is second best on all four domain-conversion datasets but drops below the Text Γ Image baseline here, while BASIC without query expansion is the best-performing configuration.
mAP (%) averaged over the instances of each primary visual category. Hover the legend to isolate a method.
mAP (%) averaged over the queries of each primary modification type.
BASIC ranks first in six of eight visual and five of seven textual categories. MagicLens is stronger on fashion, household, addition and context.
| Method | ImageNet-R | NICO++ | MiniDN | LTLL |
|---|---|---|---|---|
| Text | 0.83 | 1.12 | 0.74 | 4.43 |
| Image | 5.02 | 6.25 | 5.61 | 19.20 |
| Text + Image | 8.60 | 8.95 | 9.74 | 18.44 |
| Text Γ Image | 5.94 | 3.03 | 3.28 | 4.18 |
| FreeDom | 41.82 | 31.81 | 53.63 | 37.22 |
| BASIC | 46.92 | 29.68 | 51.94 | 42.05 |
Average mAP (%) with a SigLIP backbone. Pic2Word, SEARLE, MagicLens and similar methods train heads on CLIP features and cannot be moved to SigLIP, so the comparison is against the training-free FreeDom.
| Method | fict. | land. | mobi. | hous. | tech | fash. | prod. | art |
|---|---|---|---|---|---|---|---|---|
| Text Γ Image | 25.56 | 21.42 | 12.93 | 32.86 | 28.53 | 18.30 | 12.18 | 22.67 |
| FreeDom | 27.75 | 27.10 | 19.32 | 43.02 | 31.36 | 35.31 | 20.27 | 33.11 |
| BASIC | 50.65 | 53.45 | 39.56 | 48.87 | 50.84 | 52.83 | 45.11 | 44.43 |
| Method | proj. | doma. | attr. | appe. | view. | addi. | cont. |
|---|---|---|---|---|---|---|---|
| Text Γ Image | 21.61 | 17.89 | 20.19 | 26.47 | 11.15 | 28.24 | 25.13 |
| FreeDom | 22.94 | 25.25 | 20.03 | 31.52 | 38.34 | 39.17 | 25.59 |
| BASIC | 53.41 | 51.17 | 42.19 | 50.62 | 63.73 | 47.42 | 50.90 |
mAP (%). With SigLIP, BASIC leads FreeDom in every i-CIR category.
These datasets have a different objective: image pairs are picked automatically and their difference is described afterwards, so the text alone often solves the query (the text-only baseline beats the image-only baseline on all three). Methods that shine there, such as MagicLens or CompoDiff, do not on i-CIR, and vice versa. No single approach is best everywhere; we argue that i-CIR is closer to real use.
| Method | R@1 | R@5 | R@10 | R@50 |
|---|---|---|---|---|
| Text | 20.96 | 44.89 | 56.80 | 79.16 |
| Image | 7.42 | 23.61 | 34.07 | 57.40 |
| Text + Image | 12.41 | 36.15 | 49.18 | 78.27 |
| Text Γ Image | 22.55 | 50.36 | 62.84 | 86.02 |
| Pic2Word | 23.90 | 51.70 | 65.30 | 87.80 |
| SEARLE | 24.20 | 52.50 | 66.30 | 88.80 |
| CompoDiff | 18.20 | 53.10 | 70.80 | 90.30 |
| FreeDom | 21.00 | 48.70 | 61.90 | 88.10 |
| CIReVL | 24.60 | 52.30 | 64.90 | 86.30 |
| MagicLens | 30.10 | 61.70 | 74.40 | 92.60 |
| BASIC | 15.83 | 40.89 | 53.90 | 82.27 |
| BASIC β | 17.98 | 44.92 | 58.80 | 86.51 |
| Method | mAP@5 | mAP@10 | mAP@25 | mAP@50 |
|---|---|---|---|---|
| Text | 3.09 | 3.25 | 3.76 | 4.01 |
| Image | 1.60 | 2.02 | 2.76 | 3.13 |
| Text + Image | 4.06 | 5.20 | 6.29 | 6.85 |
| Text Γ Image | 11.64 | 12.29 | 13.64 | 14.28 |
| Pic2Word | 8.70 | 9.50 | 10.70 | 11.30 |
| SEARLE | 11.70 | 12.70 | 14.30 | 15.10 |
| CompoDiff | 12.60 | 13.40 | 15.80 | 16.40 |
| FreeDom | 14.00 | 14.80 | 16.40 | 17.20 |
| CIReVL | 18.60 | 19.00 | 20.90 | 21.80 |
| MagicLens | 29.60 | 30.80 | 33.40 | 34.40 |
| BASIC | 15.95 | 16.77 | 18.19 | 18.94 |
| BASIC β | 15.95 | 16.77 | 18.21 | 19.00 |
| Method | R@10 | R@50 |
|---|---|---|
| Text | 18.95 | 35.73 |
| Image | 7.71 | 16.39 |
| Text + Image | 20.85 | 37.08 |
| Text Γ Image | 25.95 | 43.42 |
| Pic2Word | 24.70 | 43.70 |
| SEARLE | 25.60 | 46.20 |
| CompoDiff | 36.00 | 48.60 |
| FreeDom | 21.60 | 39.50 |
| CIReVL | 28.60 | 48.60 |
| MagicLens | 30.70 | 52.50 |
| BASIC | 22.94 | 41.14 |
| BASIC β | 25.36 | 43.83 |
β BASIC without the Harris penalty, which helps when one modality dominates. CLIP ViT-L/14 throughout.
Everything is a dot product against stored CLIP features plus a few query-side vector operations, so retrieval scales like plain nearest-neighbour search and any FAISS index applies. The one real overhead is contextualization, which embeds 100 caption-like phrases per text query: with it BASIC takes 71 ms per query on ImageNet-R at 32.1 mAP, without it 33 ms at 26.2 mAP, on par with FreeDom (35 ms), SEARLE-XL (36 ms) and Pic2Word (34 ms).
CompoDiff (a diffusion model at inference) and CIReVL (a captioner plus an LLM) are orders of magnitude heavier and are not shown.
Want your method listed here? Evaluate it with the code on GitHub, report macro-mAP on i-CIR, and get in touch.
i-CIR is released as an evaluation-only benchmark under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 license. Please cite the paper when you use it. Every image keeps the license of its original source; the dataset does not alter upstream terms.
Annotators steered clear of faces, licence plates and private premises whenever the task allowed, and visible faces were pixelated where people are part of an instance. The license explicitly prohibits using i-CIR, or models evaluated with it, to identify, profile or track people, directly or through their belongings, and any surveillance, biometric or otherwise privacy-invasive application.
Found an image that reveals personal information, that you own and want removed, or that is otherwise inappropriate? Suspect the dataset is being misused? Tell us. Reports are acknowledged promptly and can lead to content removal or, for license violations, revocation of access.
Report an issueEvaluation-only release under CC BY-NC-SA 4.0. Please read the terms of responsible use before downloading.
Download i-CIR π€ Hugging Face Dataset browserIf you find our project useful, please consider citing us:
@inproceedings{icir2025,
title={Instance-Level Composed Image Retrieval},
author={Psomas, Bill and Retsinas, George and Efthymiadis, Nikos and Filntisis, Panagiotis and Avrithis, Yannis and Maragos, Petros and Chum, Ondrej and Tolias, Giorgos},
booktitle={Advances in Neural Information Processing Systems (NeurIPS)},
year={2025},
}
Questions about the dataset, the code or the benchmark? Reach out to Bill Psomas at vasileios.psomas@fel.cvut.cz.
To report misuse of the dataset or an image that should not be there, use the report channel. Reports are acknowledged promptly and can lead to content removal.