i-CIR:
Instance-Level Composed Image Retrieval

Bill Psomas1* George Retsinas2* Nikos Efthymiadis1 Panagiotis Filntisis2,4 Yannis Avrithis5 Petros Maragos2,3,4 Ondrej Chum1 Giorgos Tolias1
1 VRG, FEE, CTU in Prague · 2 Athena Research Center · 3 NTUA · 4 HERON · 5 IARAI
Composed image retrieval, as it should be: a specific object, a requested change, and thousands of look-alikes standing in the way.

The 202 instances of i-CIR, one query photograph each β€” click a photo to open that instance in the dataset browser.

Task

Composed image retrieval at the instance level: a photo says which object, a text says how it should look.

image queryimage query
+
β€œin an old archival photo”
modification text
β†’
positivepositivesame temple, in an old archival photo
visual hard negativevisual hard negativelook-alike temple, wrong modification
textual hard negativetextual hard negativean old archival photo, but of something else
composed hard negativecomposed hard negativeold archival photo of a look-alike temple

The query has two parts. An image query shows a particular object instance, here the Temple of Hephaestus in Athens, and a modification text says how that object should appear in the images we are looking for, here β€œin an old archival photo”. Neither part is enough on its own: the photo does not say what should change, and the text does not say which temple.

What counts as correct. A database image is a positive only if it shows the same instance under the requested modification. Everything else is a negative, and a method is scored by how high it ranks the positives (mean Average Precision).

What makes it hard. Each instance database is filled with near-misses that satisfy only part of the query: visual hard negatives show the same or a look-alike object without the modification (a colour photo of the Parthenon), textual hard negatives match the text but show a different object (some other old archival photo), and composed hard negatives almost satisfy both (an old archival photo of a look-alike temple). A method that leans on one modality alone falls for one of these three.

Explore it yourself. Every one of the 202 instances comes with its own image queries, modification texts, positives and thousands of such curated negatives. Open this instance in the dataset browser, or browse all 202.

Dataset

A retrieval benchmark that asks a precise question: which images show this very object, changed in this way?

object instances landmarks, products, fictional characters, tech, art…

composed queries an image query paired with a modification text

database images 99.5% of them curated hard negatives

curated hard negatives median per instance database

landmark: 59 instanceslandmarkhousehold: 42 instanceshouseholdfictional: 30 instancesfictionalfashion: 24 instancesfashionproduct: 16 instancesproductmobility: 15 instancesmobilitytech: 8 instancestechart: 6 instancesartanimal: 2 instanceslandmark β€Ί architecture: 52 instancesarchitecturelandmark β€Ί nature: 6 instanceslandmark β€Ί sculpture: 1 instancehousehold β€Ί appliance: 1 instancehousehold β€Ί decor: 9 instancesdecorhousehold β€Ί furniture: 9 instancesfurniturehousehold β€Ί kitchenware: 11 instanceskitchenwarehousehold β€Ί lighting: 2 instanceshousehold β€Ί pet: 1 instancehousehold β€Ί storage: 1 instancehousehold β€Ί textile: 7 instanceshousehold β€Ί travel: 1 instancefictional β€Ί car: 1 instancefictional β€Ί character: 29 instancescharacterfashion β€Ί accessory: 1 instancefashion β€Ί accessory: 10 instancesaccessoryfashion β€Ί bag: 2 instancesfashion β€Ί clothing: 3 instancesfashion β€Ί jewelry: 3 instancesfashion β€Ί perfume: 1 instancefashion β€Ί shoe: 4 instancesproduct β€Ί accessory: 1 instanceproduct β€Ί drink: 4 instancesproduct β€Ί instrument: 2 instancesproduct β€Ί pet: 1 instanceproduct β€Ί snack: 2 instancesproduct β€Ί souvenir: 1 instanceproduct β€Ί stationery: 2 instancesproduct β€Ί tool: 1 instanceproduct β€Ί toy: 2 instancesmobility β€Ί airplane: 3 instancesmobility β€Ί boat: 1 instancemobility β€Ί car: 7 instancesmobility β€Ί motorcycle: 2 instancesmobility β€Ί ship: 1 instancemobility β€Ί spacecraft: 1 instancetech β€Ί audio: 1 instancetech β€Ί camera: 3 instancestech β€Ί gaming: 3 instancestech β€Ί peripheral: 1 instanceart β€Ί coin: 1 instanceart β€Ί inscription: 1 instanceart β€Ί mask: 1 instanceart β€Ί mural: 3 instancesanimal β€Ί pet: 2 instanceslandmark β€Ί amphitheater: 1 instancelandmark β€Ί bell tower: 1 instancelandmark β€Ί bridge: 3 instanceslandmark β€Ί building: 2 instanceslandmark β€Ί castle: 3 instanceslandmark β€Ί church: 1 instancelandmark β€Ί house: 2 instanceslandmark β€Ί library: 1 instancelandmark β€Ί lighthouse: 2 instanceslandmark β€Ί mausoleum: 1 instancelandmark β€Ί memorial: 1 instancelandmark β€Ί museum: 1 instancelandmark β€Ί observation tower: 1 instancelandmark β€Ί pagoda: 2 instanceslandmark β€Ί public market: 1 instancelandmark β€Ί ruins: 1 instancelandmark β€Ί scene: 2 instanceslandmark β€Ί square: 7 instanceslandmark β€Ί tavern: 1 instancelandmark β€Ί temple: 9 instanceslandmark β€Ί village: 9 instanceslandmark β€Ί bay: 1 instancelandmark β€Ί beach: 1 instancelandmark β€Ί lagoon: 1 instancelandmark β€Ί mountain: 3 instanceslandmark β€Ί rock relief: 1 instancehousehold β€Ί hair dryer: 1 instancehousehold β€Ί box: 1 instancehousehold β€Ί candle: 2 instanceshousehold β€Ί coaster: 1 instancehousehold β€Ί curtain tieback: 1 instancehousehold β€Ί figurine: 1 instancehousehold β€Ί photo frame: 1 instancehousehold β€Ί plate: 2 instanceshousehold β€Ί beach chair: 1 instancehousehold β€Ί chair: 1 instancehousehold β€Ί design chair: 1 instancehousehold β€Ί desk: 1 instancehousehold β€Ί gaming chair: 1 instancehousehold β€Ί sofa: 1 instancehousehold β€Ί stool: 1 instancehousehold β€Ί storage box: 1 instancehousehold β€Ί valet stand: 1 instancehousehold β€Ί bowl: 2 instanceshousehold β€Ί dish: 1 instancehousehold β€Ί grater: 1 instancehousehold β€Ί jar: 3 instanceshousehold β€Ί measuring cup: 1 instancehousehold β€Ί mug: 1 instancehousehold β€Ί plate: 1 instancehousehold β€Ί spoon: 1 instancehousehold β€Ί night light: 2 instanceshousehold β€Ί carrier: 1 instancehousehold β€Ί basket: 1 instancehousehold β€Ί bedsheet: 1 instancehousehold β€Ί carpet: 1 instancehousehold β€Ί cushion: 1 instancehousehold β€Ί pillow: 1 instancehousehold β€Ί table runner: 1 instancehousehold β€Ί towel: 2 instanceshousehold β€Ί pillow: 1 instancefictional β€Ί sports: 1 instancefictional β€Ί animated movie: 9 instancesfictional β€Ί anime: 6 instancesfictional β€Ί cartoon: 5 instancesfictional β€Ί comic: 3 instancesfictional β€Ί superhero: 2 instancesfictional β€Ί video game: 4 instancesfashion β€Ί hair tie: 1 instancefashion β€Ί bandana: 1 instancefashion β€Ί carnival mask: 1 instancefashion β€Ί eyeglasses: 1 instancefashion β€Ί hair clip: 2 instancesfashion β€Ί hand fan: 1 instancefashion β€Ί pouch: 1 instancefashion β€Ί ring: 1 instancefashion β€Ί sunglasses: 1 instancefashion β€Ί tote bag: 1 instancefashion β€Ί clutch: 1 instancefashion β€Ί straw bag: 1 instancefashion β€Ί skirt: 1 instancefashion β€Ί swimsuit: 1 instancefashion β€Ί vest: 1 instancefashion β€Ί bracelet: 1 instancefashion β€Ί necklace: 1 instancefashion β€Ί ring: 1 instancefashion β€Ί fragrance: 1 instancefashion β€Ί skate shoe: 1 instancefashion β€Ί slippers: 1 instancefashion β€Ί sneaker: 2 instancesproduct β€Ί case: 1 instanceproduct β€Ί beer: 2 instancesproduct β€Ί energy drink: 1 instanceproduct β€Ί soda: 1 instanceproduct β€Ί electric guitar: 2 instancesproduct β€Ί toy: 1 instanceproduct β€Ί chocolate: 2 instancesproduct β€Ί keychain: 1 instanceproduct β€Ί notebook: 1 instanceproduct β€Ί pencil: 1 instanceproduct β€Ί multitool: 1 instanceproduct β€Ί plush: 1 instanceproduct β€Ί puzzle: 1 instancemobility β€Ί commercial: 1 instancemobility β€Ί military jet: 1 instancemobility β€Ί supersonic jet: 1 instancemobility β€Ί fishing boat: 1 instancemobility β€Ί classic: 1 instancemobility β€Ί compact: 1 instancemobility β€Ί electric pickup: 1 instancemobility β€Ί muscle: 1 instancemobility β€Ί sports: 3 instancesmobility β€Ί adventure: 1 instancemobility β€Ί scooter: 1 instancemobility β€Ί ferry: 1 instancemobility β€Ί rocket: 1 instancetech β€Ί earbuds: 1 instancetech β€Ί film: 1 instancetech β€Ί instant: 1 instancetech β€Ί medium format: 1 instancetech β€Ί console: 2 instancestech β€Ί controller: 1 instancetech β€Ί keyboard: 1 instanceart β€Ί ancient: 1 instanceart β€Ί undeciphered: 1 instanceart β€Ί funerary: 1 instanceart β€Ί information sign: 1 instanceart β€Ί street art: 2 instancesanimal β€Ί dog: 2 instances202instances

The 202 instances by visual category (three levels). Hover a slice for its name and count.

i-CIR follows an instance-level class definition: two images belong to the same class only if they show the same particular object, the Temple of Hephaestus or Batman rather than β€œa Greek temple” or β€œa superhero”. A composed query pairs a photo of an instance with a short text describing a modification, and a method must rank the images that show that instance under that modification above everything else.

  • Curated hard negatives. Each instance database is built from three kinds of near-misses: visual (the same or a look-alike object, wrong modification), textual (right modification, different object) and composed (almost both). They make up 99.5% of the 752,092 database images.
  • Per-instance databases, exhaustive labels. All composed queries of an instance share one database and other instances have their own, so every database image is verified as positive or negative for every query. Positives of one query are negatives of the others.
  • Semi-automated collection, human-verified. Seed images and seed sentences drive CLIP image-to-image and text-to-image search over LAION; candidates are filtered for resolution, watermarks and duplicates, then annotators mark positives, pick the image queries and verify the rest. Images come from LAION, from Creative-Commons results of Google image search, and from photos taken by our annotators. Seed images never enter the dataset and seed sentences never reuse the wording of a text query, to avoid a bias toward CLIP-based methods.
  • Privacy-first. Annotators avoided faces, licence plates and private premises wherever the task allowed; where people are intrinsic to an instance (e.g. apparel) faces are pixelated. i-CIR is released for evaluation only, under CC BY-NC-SA 4.0.

What the modifications ask for

The 518 (instance, modification) pairs by modification type. Hover a segment for examples.

addition 27%domain 21%context 16%appearance 13%attribute 11%projection 8%
addition 27%domain 21%context 16%appearance 13%attribute 11%projection 8%viewpoint 4%

By the numbers

Image queries per instance
02040601: 414112: 555523: 464634: 151545: 131356: 8867: 7778: 5589: 22910: 551011–15: 3311–1516–25: 2216–25image queries per instanceinstances

1 to 25 photos, median 3

Modification texts per instance
0204060801: 303012: 747423: 585834: 343445: 665modification texts per instanceinstances

1 to 5 texts, median 2

Curated hard negatives per instance
0204060800–1K: 220–1K1–2K: 19191–2K2–3K: 51512–3K3–4K: 62623–4K4–5K: 34344–5K5–6K: 15155–6K6–7K: 886–7K7–8K: 667–8K8–9K: 448–9K9–10K: 09–10K10K+: 1110K+curated hard negatives per instanceinstances

951 to 10,045, median 3,421

Positives per composed query
02004006001: 848412: 15515523: 23823834: 26726745: 20220256–10: 4754756–1011–20: 32232211–2021–50: 13213221–5051+: 8851+positives per composed querycomposed queries

1 to 127, median 5

Why i-CIR?

01020304050600M10M20M30M40M50Mi-CIR with its curated hard negatives: 17.48 mAPsame difficulty only after> 40M random distractors0.5M random LAION distractors: 52.75 mAP0.75M random LAION distractors: 49.85 mAP1M random LAION distractors: 47.18 mAP2M random LAION distractors: 39.43 mAP5M random LAION distractors: 32.93 mAP7.5M random LAION distractors: 29.83 mAP10M random LAION distractors: 27.85 mAP15M random LAION distractors: 25.14 mAP20M random LAION distractors: 23.01 mAP30M random LAION distractors: 21.01 mAP40M random LAION distractors: 18.88 mAP50M random LAION distractors: 17.07 mAPi-CIR with random LAION images insteadrandom distractors added to the database (millions)Text Γ— Image mAP (%)

Compact, yet as hard as tens of millions of distractors

The plain Text Γ— Image baseline reaches 17.48 mAP on i-CIR. Swap the curated negatives for random LAION images and it takes more than 40 million of them to drag the same baseline down to that level: four orders of magnitude more than the 3.7K images an i-CIR instance database holds on average, and 1.5 orders of magnitude more than the 750K images of the whole benchmark. The random-distractor curve is even a lower bound, since unlabeled LAION images inevitably contain false negatives.

Truly compositional

WeiCom0510152025303500.250.50.751CIRCO β€” WeiComCIRCO, Ξ»=0: 4 mAPCIRCO, Ξ»=0.1: 10.84 mAPCIRCO, Ξ»=0.2: 10.84 mAPCIRCO, Ξ»=0.3: 11.07 mAPCIRCO, Ξ»=0.4: 11.1 mAPCIRCO, Ξ»=0.5: 10.95 mAPCIRCO, Ξ»=0.6: 10.99 mAPCIRCO, Ξ»=0.7: 10.52 mAPCIRCO, Ξ»=0.8: 10.77 mAPCIRCO, Ξ»=0.9: 10.55 mAPCIRCO, Ξ»=1: 3.26 mAPFashionIQ β€” WeiComFashionIQ, Ξ»=0: 18.94 mAPFashionIQ, Ξ»=0.1: 18.47 mAPFashionIQ, Ξ»=0.2: 17.8 mAPFashionIQ, Ξ»=0.3: 17.79 mAPFashionIQ, Ξ»=0.4: 17.42 mAPFashionIQ, Ξ»=0.5: 16.91 mAPFashionIQ, Ξ»=0.6: 16.28 mAPFashionIQ, Ξ»=0.7: 15.57 mAPFashionIQ, Ξ»=0.8: 14.49 mAPFashionIQ, Ξ»=0.9: 12.7 mAPFashionIQ, Ξ»=1: 7.69 mAPCIRR β€” WeiComCIRR, Ξ»=0: 26.09 mAPCIRR, Ξ»=0.1: 20.86 mAPCIRR, Ξ»=0.2: 19.98 mAPCIRR, Ξ»=0.3: 19.6 mAPCIRR, Ξ»=0.4: 19.51 mAPCIRR, Ξ»=0.5: 18.86 mAPCIRR, Ξ»=0.6: 18.31 mAPCIRR, Ξ»=0.7: 17.99 mAPCIRR, Ξ»=0.8: 17.22 mAPCIRR, Ξ»=0.9: 16.62 mAPCIRR, Ξ»=1: 11.56 mAPi-CIR β€” WeiComi-CIR, Ξ»=0: 3.01 mAPi-CIR, Ξ»=0.1: 15.16 mAPi-CIR, Ξ»=0.2: 18.27 mAPi-CIR, Ξ»=0.3: 18.97 mAPi-CIR, Ξ»=0.4: 18.64 mAPi-CIR, Ξ»=0.5: 18.03 mAPi-CIR, Ξ»=0.6: 16.94 mAPi-CIR, Ξ»=0.7: 15.08 mAPi-CIR, Ξ»=0.8: 13.4 mAPi-CIR, Ξ»=0.9: 9.81 mAPi-CIR, Ξ»=1: 3.04 mAPΞ» (0 = text only, 1 = image only)Text + Image0510152025303500.250.50.751CIRCO β€” Text + ImageCIRCO, Ξ»=0: 4.1 mAPCIRCO, Ξ»=0.1: 7 mAPCIRCO, Ξ»=0.2: 10.3 mAPCIRCO, Ξ»=0.3: 8.14 mAPCIRCO, Ξ»=0.4: 6.93 mAPCIRCO, Ξ»=0.5: 5.61 mAPCIRCO, Ξ»=0.6: 4.57 mAPCIRCO, Ξ»=0.7: 3.82 mAPCIRCO, Ξ»=0.8: 3.16 mAPCIRCO, Ξ»=0.9: 2.86 mAPCIRCO, Ξ»=1: 2.61 mAPFashionIQ β€” Text + ImageFashionIQ, Ξ»=0: 18.95 mAPFashionIQ, Ξ»=0.1: 22.52 mAPFashionIQ, Ξ»=0.2: 25.69 mAPFashionIQ, Ξ»=0.3: 26.95 mAPFashionIQ, Ξ»=0.4: 24.71 mAPFashionIQ, Ξ»=0.5: 20.83 mAPFashionIQ, Ξ»=0.6: 16.22 mAPFashionIQ, Ξ»=0.7: 12.54 mAPFashionIQ, Ξ»=0.8: 10.24 mAPFashionIQ, Ξ»=0.9: 8.66 mAPFashionIQ, Ξ»=1: 7.69 mAPCIRR β€” Text + ImageCIRR, Ξ»=0: 26.65 mAPCIRR, Ξ»=0.1: 30.7 mAPCIRR, Ξ»=0.2: 31.23 mAPCIRR, Ξ»=0.3: 27.4 mAPCIRR, Ξ»=0.4: 22.48 mAPCIRR, Ξ»=0.5: 19.02 mAPCIRR, Ξ»=0.6: 16.26 mAPCIRR, Ξ»=0.7: 14.61 mAPCIRR, Ξ»=0.8: 13.46 mAPCIRR, Ξ»=0.9: 12.49 mAPCIRR, Ξ»=1: 11.73 mAPi-CIR β€” Text + Imagei-CIR, Ξ»=0: 3.01 mAPi-CIR, Ξ»=0.1: 7.6 mAPi-CIR, Ξ»=0.2: 15.78 mAPi-CIR, Ξ»=0.3: 16.6 mAPi-CIR, Ξ»=0.4: 12.21 mAPi-CIR, Ξ»=0.5: 8.2 mAPi-CIR, Ξ»=0.6: 5.6 mAPi-CIR, Ξ»=0.7: 4.51 mAPi-CIR, Ξ»=0.8: 3.79 mAPi-CIR, Ξ»=0.9: 3.43 mAPi-CIR, Ξ»=1: 3.04 mAPΞ» (0 = text only, 1 = image only)Text Γ— Image0510152025303500.250.50.751CIRCO β€” Text Γ— ImageCIRCO, Ξ»=0: 4.1 mAPCIRCO, Ξ»=0.1: 5.3 mAPCIRCO, Ξ»=0.2: 6.49 mAPCIRCO, Ξ»=0.3: 9.11 mAPCIRCO, Ξ»=0.4: 10.35 mAPCIRCO, Ξ»=0.5: 11.12 mAPCIRCO, Ξ»=0.6: 10.5 mAPCIRCO, Ξ»=0.7: 7.61 mAPCIRCO, Ξ»=0.8: 5.75 mAPCIRCO, Ξ»=0.9: 3.93 mAPCIRCO, Ξ»=1: 2.61 mAPFashionIQ β€” Text Γ— ImageFashionIQ, Ξ»=0: 18.95 mAPFashionIQ, Ξ»=0.1: 20.29 mAPFashionIQ, Ξ»=0.2: 21.39 mAPFashionIQ, Ξ»=0.3: 22.86 mAPFashionIQ, Ξ»=0.4: 24.63 mAPFashionIQ, Ξ»=0.5: 25.95 mAPFashionIQ, Ξ»=0.6: 25.94 mAPFashionIQ, Ξ»=0.7: 24.06 mAPFashionIQ, Ξ»=0.8: 19.48 mAPFashionIQ, Ξ»=0.9: 12.6 mAPFashionIQ, Ξ»=1: 7.72 mAPCIRR β€” Text Γ— ImageCIRR, Ξ»=0: 26.66 mAPCIRR, Ξ»=0.1: 28.36 mAPCIRR, Ξ»=0.2: 30.08 mAPCIRR, Ξ»=0.3: 31.19 mAPCIRR, Ξ»=0.4: 30.73 mAPCIRR, Ξ»=0.5: 29.37 mAPCIRR, Ξ»=0.6: 26.33 mAPCIRR, Ξ»=0.7: 22.18 mAPCIRR, Ξ»=0.8: 18.05 mAPCIRR, Ξ»=0.9: 14.45 mAPCIRR, Ξ»=1: 11.73 mAPi-CIR β€” Text Γ— Imagei-CIR, Ξ»=0: 3.01 mAPi-CIR, Ξ»=0.1: 4.03 mAPi-CIR, Ξ»=0.2: 5.98 mAPi-CIR, Ξ»=0.3: 9.52 mAPi-CIR, Ξ»=0.4: 14.11 mAPi-CIR, Ξ»=0.5: 17.48 mAPi-CIR, Ξ»=0.6: 18.27 mAPi-CIR, Ξ»=0.7: 15.65 mAPi-CIR, Ξ»=0.8: 10.27 mAPi-CIR, Ξ»=0.9: 5.34 mAPi-CIR, Ξ»=1: 3.04 mAPΞ» (0 = text only, 1 = image only)mAP (%)i-CIRCIRRFashionIQCIRCO

Sweep a weight Ξ» from text-only (Ξ» = 0) to image-only (Ξ» = 1) similarity for three simple fusion rules. On i-CIR every rule peaks strictly inside the interval, with a composition gain of +14.9 mAP (+490%) over the best single modality, averaged over the three rules.

The gain shrinks to +3.0 mAP on CIRR, +5.0 on FashionIQ and +6.8 on CIRCO, and CIRR and FashionIQ are best served by the text alone. i-CIR rewards methods that genuinely combine both modalities.

Method

BASIC, a Baseline Approach for Surprisingly strong Composition: training-free, built on frozen VLM features, and it never touches the stored database index.

IMAGE QUERY PATHTEXT QUERY PATHobject corpus C+dog Β· building Β· man Β· car Β· treeprojection matrix Ptop-k eigenvectors of(1βˆ’Ξ±)Cβ‚Š βˆ’ Ξ±Cβ‚‹stylistic corpus Cβˆ’sci-fi Β· oil painting Β· in winter Β· aerial viewimage database Xvisual encoder β†’ centred features xΜ„vistored once as plain CLIP features Β· never re-indexedimage queryvisualencoderβˆ’ΞΌvcenteringsemantic projectionP⊀ qΜ„vquery expansionoptionalβŠ™sviin an oldarchivalphototext querycontextualizationdog + text querybuilding + text queryman + text querytext encoderβ†’ meanβˆ’ΞΌtqΜ„tβ€œsomething in anold archival photoβ€βŠ™dot productstimin normalizationsvminstminβŠ—AND+ HarriscriterionsΜƒftop-1 ranked imagesame temple, archival photo

A composed query is treated as a logical AND: a database image must be similar to the image query and match the text query. BASIC scores the two modalities separately, cleans each similarity of modality-specific noise, and fuses them multiplicatively so that images strong in only one modality are pushed down.

The recipe, step by step

Centering

Subtract a mean image feature (computed on LAION) and a mean text feature (computed on the object corpus). This strips generic β€œimageness” and β€œtextness” and leaves the semantic content.

Semantic projection (image side)

Two LLM-generated corpora define what to keep and what to drop: Cβ‚Š, object names such as dog or building, and Cβ‚‹, styles, viewpoints and settings such as cartoon, aerial view or in a cloudy day. The top-k eigenvectors of (1βˆ’Ξ±)Cβ‚Š βˆ’ Ξ±Cβ‚‹, built from text-feature covariances, form a projection P that keeps object identity and suppresses style (k = 250, Ξ± = 0.2).

Contextualization (text side)

Short queries like during sunset are out of distribution for CLIP, which was trained on captions. Each query is wrapped with random object terms from Cβ‚Š (dog during sunset, sculpture dog), the phrases are embedded, centred and averaged.

Query expansion (image side, optional)

The image query is enriched with a similarity-weighted combination of its top-ranked database features. It helps on class-level datasets and slightly hurts on i-CIR, so both variants are reported.

Min-normalization and Harris fusion

Each similarity is rescaled by its empirical minimum, sΜƒ = (s βˆ’ smin) / |smin|, so that the two modalities live on comparable ranges. The final score is sΜƒf = sΜƒvΒ·sΜƒt βˆ’ Ξ» (sΜƒv + sΜƒt)2: the product rewards images relevant to both queries, the Harris-style penalty pushes down images that are strong in one modality only (Ξ» = 0.1).

Query-side only, index untouched

Centering and projection fold into the query: sv = ⟨xv, PP⊀(qv βˆ’ ΞΌv)⟩ βˆ’ c(q), where the last term is a query-dependent constant. The database keeps its plain CLIP features, any FAISS index works unchanged, and the corpora, and therefore the projection, can be swapped per application without re-indexing.

What each component adds

mAP (%) while components are switched on one by one (top) and switched off individually (bottom). Bold: best per column.

CenteringMin norm.HarrisContext.ProjectionQ. exp.ImageNet-RNICO++MiniDNLTLLi-CIR
βœ—βœ—βœ—βœ—βœ—βœ—7.669.269.4819.7817.48Text Γ— Image
βœ“βœ—βœ—βœ—βœ—βœ—12.169.9512.1616.9328.33
βœ“βœ“βœ—βœ—βœ—βœ—12.0617.2017.7222.2027.30
βœ“βœ“βœ“βœ—βœ—βœ—16.2115.0617.7929.7028.42
βœ“βœ“βœ“βœ“βœ—βœ—18.6115.3421.0133.7433.48
βœ“βœ“βœ“βœ“βœ“βœ—27.5428.9035.7538.2234.35BASIC †
βœ“βœ“βœ“βœ“βœ“βœ“32.1331.6539.5841.3831.64BASIC
βœ“βœ“βœ“βœ“βœ—βœ“17.3113.9621.2222.4231.78no projection
βœ“βœ“βœ“βœ—βœ“βœ“26.1830.6133.6434.5025.85no contextualization
βœ“βœ“βœ—βœ“βœ“βœ“30.7529.8238.8540.6531.61no Harris
βœ“βœ—βœ—βœ“βœ“βœ“24.5022.7429.6519.3630.75plain product

† without query expansion, the configuration that works best on i-CIR. All rows use CLIP ViT-L/14; hyper-parameters were fixed once on a small private development set.

Benchmark

Mean Average Precision on i-CIR and on four domain-conversion CIR datasets. On i-CIR, mAP is computed per instance and averaged over instances (macro-mAP), so every object counts equally.

MethodImageNet-RNICO++MiniDNLTLLi-CIR
Text0.741.090.575.723.01
Image3.846.326.6616.493.04
Text + Image6.219.309.3317.868.20
Text Γ— Image7.839.799.8623.1617.48
WeiCom10.4710.548.5226.6018.03
Pic2Word7.889.7612.0021.2719.36
CompoDiff12.8810.3222.9521.619.63
CIReVL18.1117.8026.2032.6018.66
SEARLE14.0415.1321.7825.4619.90
MCL8.1319.0918.4116.6719.89
MagicLens9.1319.6620.0624.2127.35
CoVR-211.5224.9327.7624.6828.50
FreeDom29.9126.1037.2733.2417.24
FreeDom †25.8123.2432.1430.8215.76
BASIC32.1331.6539.5841.3831.64
BASIC †27.5428.9035.7538.2234.35

Average mAP (%). † without query expansion. On i-CIR the reported number is macro-mAP: mAP is computed per instance and then averaged over the 202 instances. All methods use CLIP ViT-L/14 except CompoDiff (ViT-G/14). Click a column header to sort.

i-CIR macro-mAP

010203040BASIC †BASIC †: 34.35 macro-mAP34.35BASICBASIC: 31.64 macro-mAP31.64CoVR-2CoVR-2: 28.50 macro-mAP28.50MagicLensMagicLens: 27.35 macro-mAP27.35SEARLESEARLE: 19.90 macro-mAP19.90MCLMCL: 19.89 macro-mAP19.89Pic2WordPic2Word: 19.36 macro-mAP19.36CIReVLCIReVL: 18.66 macro-mAP18.66WeiComWeiCom: 18.03 macro-mAP18.03Text Γ— ImageText Γ— Image: 17.48 macro-mAP17.48FreeDomFreeDom: 17.24 macro-mAP17.24FreeDom †FreeDom †: 15.76 macro-mAP15.76CompoDiffCompoDiff: 9.63 macro-mAP9.63Text + ImageText + Image: 8.20 macro-mAP8.20ImageImage: 3.04 macro-mAP3.04TextText: 3.01 macro-mAP3.01

Leading on class-level datasets does not carry over to i-CIR: FreeDom is second best on all four domain-conversion datasets but drops below the Text Γ— Image baseline here, while BASIC without query expansion is the best-performing configuration.

Visual categories

mAP (%) averaged over the instances of each primary visual category. Hover the legend to isolate a method.

0102030405060Text Γ— Image β€” product: 18.118.1Pic2Word β€” product: 15.415.4SEARLE β€” product: 17.217.2FreeDom β€” product: 25.525.5MagicLens β€” product: 26.726.7BASIC β€” product: 33.733.7productText Γ— Image β€” tech: 23.023.0Pic2Word β€” tech: 14.214.2SEARLE β€” tech: 17.017.0FreeDom β€” tech: 14.514.5MagicLens β€” tech: 18.718.7BASIC β€” tech: 30.630.6techText Γ— Image β€” fictional: 21.021.0Pic2Word β€” fictional: 19.419.4SEARLE β€” fictional: 31.131.1FreeDom β€” fictional: 30.730.7MagicLens β€” fictional: 28.828.8BASIC β€” fictional: 47.847.8fictionalText Γ— Image β€” fashion: 11.911.9Pic2Word β€” fashion: 14.614.6SEARLE β€” fashion: 14.414.4FreeDom β€” fashion: 12.812.8MagicLens β€” fashion: 25.625.6BASIC β€” fashion: 22.022.0fashionText Γ— Image β€” landmark: 27.927.9Pic2Word β€” landmark: 32.832.8SEARLE β€” landmark: 31.131.1FreeDom β€” landmark: 17.617.6MagicLens β€” landmark: 35.035.0BASIC β€” landmark: 39.339.3landmarkText Γ— Image β€” household: 11.211.2Pic2Word β€” household: 13.913.9SEARLE β€” household: 11.711.7FreeDom β€” household: 15.615.6MagicLens β€” household: 29.129.1BASIC β€” household: 22.422.4householdText Γ— Image β€” mobility: 25.525.5Pic2Word β€” mobility: 21.321.3SEARLE β€” mobility: 18.518.5FreeDom β€” mobility: 28.928.9MagicLens β€” mobility: 29.329.3BASIC β€” mobility: 45.845.8mobilityText Γ— Image β€” art: 26.826.8Pic2Word β€” art: 33.533.5SEARLE β€” art: 22.722.7FreeDom β€” art: 23.523.5MagicLens β€” art: 35.035.0BASIC β€” art: 38.038.0artText Γ— ImagePic2WordSEARLEFreeDomMagicLensBASICmAP (%)

Textual categories

mAP (%) averaged over the queries of each primary modification type.

0102030405060Text Γ— Image β€” projection: 16.316.3Pic2Word β€” projection: 20.020.0SEARLE β€” projection: 27.027.0FreeDom β€” projection: 21.521.5MagicLens β€” projection: 31.131.1BASIC β€” projection: 53.153.1projectionText Γ— Image β€” domain: 23.123.1Pic2Word β€” domain: 24.424.4SEARLE β€” domain: 16.216.2FreeDom β€” domain: 22.322.3MagicLens β€” domain: 31.131.1BASIC β€” domain: 39.339.3domainText Γ— Image β€” attribute: 11.511.5Pic2Word β€” attribute: 11.911.9SEARLE β€” attribute: 17.217.2FreeDom β€” attribute: 16.416.4MagicLens β€” attribute: 24.124.1BASIC β€” attribute: 26.326.3attributeText Γ— Image β€” appearance: 26.226.2Pic2Word β€” appearance: 19.019.0SEARLE β€” appearance: 36.836.8FreeDom β€” appearance: 32.432.4MagicLens β€” appearance: 25.825.8BASIC β€” appearance: 48.848.8appearanceText Γ— Image β€” addition: 15.615.6Pic2Word β€” addition: 16.816.8SEARLE β€” addition: 15.715.7FreeDom β€” addition: 19.019.0MagicLens β€” addition: 28.228.2BASIC β€” addition: 24.024.0additionText Γ— Image β€” context: 26.226.2Pic2Word β€” context: 34.634.6SEARLE β€” context: 32.832.8FreeDom β€” context: 17.017.0MagicLens β€” context: 36.436.4BASIC β€” context: 35.635.6contextText Γ— Image β€” viewpoint: 30.930.9Pic2Word β€” viewpoint: 32.532.5SEARLE β€” viewpoint: 36.736.7FreeDom β€” viewpoint: 17.717.7MagicLens β€” viewpoint: 40.140.1BASIC β€” viewpoint: 47.847.8viewpointText Γ— ImagePic2WordSEARLEFreeDomMagicLensBASICmAP (%)

BASIC ranks first in six of eight visual and five of seven textual categories. MagicLens is stronger on fashion, household, addition and context.

SigLIP backbone

MethodImageNet-RNICO++MiniDNLTLL
Text0.831.120.744.43
Image5.026.255.6119.20
Text + Image8.608.959.7418.44
Text Γ— Image5.943.033.284.18
FreeDom41.8231.8153.6337.22
BASIC46.9229.6851.9442.05

Average mAP (%) with a SigLIP backbone. Pic2Word, SEARLE, MagicLens and similar methods train heads on CLIP features and cannot be moved to SigLIP, so the comparison is against the training-free FreeDom.

i-CIR with SigLIP, per visual category
Methodfict.land.mobi.hous.techfash.prod.art
Text Γ— Image25.5621.4212.9332.8628.5318.3012.1822.67
FreeDom27.7527.1019.3243.0231.3635.3120.2733.11
BASIC50.6553.4539.5648.8750.8452.8345.1144.43
i-CIR with SigLIP, per textual category
Methodproj.doma.attr.appe.view.addi.cont.
Text Γ— Image21.6117.8920.1926.4711.1528.2425.13
FreeDom22.9425.2520.0331.5238.3439.1725.59
BASIC53.4151.1742.1950.6263.7347.4250.90

mAP (%). With SigLIP, BASIC leads FreeDom in every i-CIR category.

CIRR, CIRCO and FashionIQ

These datasets have a different objective: image pairs are picked automatically and their difference is described afterwards, so the text alone often solves the query (the text-only baseline beats the image-only baseline on all three). Methods that shine there, such as MagicLens or CompoDiff, do not on i-CIR, and vice versa. No single approach is best everywhere; we argue that i-CIR is closer to real use.

CIRR
MethodR@1R@5R@10R@50
Text20.9644.8956.8079.16
Image7.4223.6134.0757.40
Text + Image12.4136.1549.1878.27
Text Γ— Image22.5550.3662.8486.02
Pic2Word23.9051.7065.3087.80
SEARLE24.2052.5066.3088.80
CompoDiff18.2053.1070.8090.30
FreeDom21.0048.7061.9088.10
CIReVL24.6052.3064.9086.30
MagicLens30.1061.7074.4092.60
BASIC15.8340.8953.9082.27
BASIC β˜…17.9844.9258.8086.51
CIRCO
MethodmAP@5mAP@10mAP@25mAP@50
Text3.093.253.764.01
Image1.602.022.763.13
Text + Image4.065.206.296.85
Text Γ— Image11.6412.2913.6414.28
Pic2Word8.709.5010.7011.30
SEARLE11.7012.7014.3015.10
CompoDiff12.6013.4015.8016.40
FreeDom14.0014.8016.4017.20
CIReVL18.6019.0020.9021.80
MagicLens29.6030.8033.4034.40
BASIC15.9516.7718.1918.94
BASIC β˜…15.9516.7718.2119.00
FashionIQ (average over dress, shirt, toptee)
MethodR@10R@50
Text18.9535.73
Image7.7116.39
Text + Image20.8537.08
Text Γ— Image25.9543.42
Pic2Word24.7043.70
SEARLE25.6046.20
CompoDiff36.0048.60
FreeDom21.6039.50
CIReVL28.6048.60
MagicLens30.7052.50
BASIC22.9441.14
BASIC β˜…25.3643.83

β˜… BASIC without the Harris penalty, which helps when one modality dominates. CLIP ViT-L/14 throughout.

01020304030 ms40 ms50 ms60 ms70 ms80 msBASIC: 32.13 mAP, 70.96 ms per queryBASICBASIC, no query expansion: 27.54 mAP, 70.48 ms per queryBASIC, no query expansionBASIC, no contextualization: 26.18 mAP, 33.23 ms per queryBASIC, no contextualizationFreeDom: 29.91 mAP, 34.66 ms per queryFreeDomSEARLE-XL: 14.04 mAP, 36.09 ms per querySEARLE-XLPic2Word: 7.88 mAP, 33.5 ms per queryPic2WordWeiCom: 10.47 mAP, 32.91 ms per queryWeiComtime per query on ImageNet-R (ms, no FAISS or other index acceleration)mAP (%)

Lightweight by construction

Everything is a dot product against stored CLIP features plus a few query-side vector operations, so retrieval scales like plain nearest-neighbour search and any FAISS index applies. The one real overhead is contextualization, which embeds 100 caption-like phrases per text query: with it BASIC takes 71 ms per query on ImageNet-R at 32.1 mAP, without it 33 ms at 26.2 mAP, on par with FreeDom (35 ms), SEARLE-XL (36 ms) and Pic2Word (34 ms).

CompoDiff (a diffusion model at inference) and CIReVL (a captioner plus an LLM) are orders of magnitude heavier and are not shown.

Want your method listed here? Evaluate it with the code on GitHub, report macro-mAP on i-CIR, and get in touch.

License & responsible use

CC BY-NC-SA 4.0

i-CIR is released as an evaluation-only benchmark under the Creative Commons Attribution-NonCommercial-ShareAlike 4.0 license. Please cite the paper when you use it. Every image keeps the license of its original source; the dataset does not alter upstream terms.

Privacy first

Annotators steered clear of faces, licence plates and private premises whenever the task allowed, and visible faces were pixelated where people are part of an instance. The license explicitly prohibits using i-CIR, or models evaluated with it, to identify, profile or track people, directly or through their belongings, and any surveillance, biometric or otherwise privacy-invasive application.

Report misuse or inappropriate content

Found an image that reveals personal information, that you own and want removed, or that is otherwise inappropriate? Suspect the dataset is being misused? Tell us. Reports are acknowledged promptly and can lead to content removal or, for license violations, revocation of access.

Report an issue

Get the full benchmark package: images, queries, positives and curated hard negatives for all 202 instances.

Evaluation-only release under CC BY-NC-SA 4.0. Please read the terms of responsible use before downloading.

Download i-CIR πŸ€— Hugging Face Dataset browser

Get in touch

Citation

If you find our project useful, please consider citing us:

@inproceedings{icir2025,
title={Instance-Level Composed Image Retrieval},
author={Psomas, Bill and Retsinas, George and Efthymiadis, Nikos and Filntisis, Panagiotis and Avrithis, Yannis and Maragos, Petros and Chum, Ondrej and Tolias, Giorgos},
booktitle={Advances in Neural Information Processing Systems (NeurIPS)},
year={2025},
}

Questions about the dataset, the code or the benchmark? Reach out to Bill Psomas at vasileios.psomas@fel.cvut.cz.

To report misuse of the dataset or an image that should not be there, use the report channel. Reports are acknowledged promptly and can lead to content removal.