Can AI paint a smell?
A team at Tsinghua University asked 30 people to smell real things and put words to them. They fed those words to an AI image tool, made 560 pictures, then asked 28 more people how well each picture matched the smell. We make pictures of smells too. This page sets out what they found, what it confirmed about our work, what it rules out, and the one result that argues against the way we do it.
On this page
What they did
Paint by Odor: An Exploration of Odor Visualization through Large Language Model and Generative AI is a 2026 preprint by Gang Yu, Yuchi Sun, Weining Yan, Xinyu Wang and Qi Lu, at Tsinghua University and Beijing University of Posts and Telecommunications. There were two experiments.
In the first, 30 people smelled 20 everyday things. Lemon, green tea, coffee beans, vinegar, glue, shampoo, mouthwash, shoe polish. For each one they picked words for the smell and words for how it made them feel. The team then asked GPT-4 to do the same job from the name alone, and compared the two.
In the second, they turned those words into pictures with Midjourney. They used seven different recipes, from a plain photo of the source right through to an abstract painting in the style of Mark Rothko. That gave 560 pictures. Then 28 people rated every one on four things: does it show the source, does it match the smell, do you like it, and is it interesting.
Colour does more of the work than anything else
This is the clearest result in the paper, and the most useful one to us. The team sorted 1,076 coded remarks from the people rating the pictures. Colour drew nearly twice as many as the next channel, and nothing else came close. The team call it the most intuitive visual element related to odours, and it drew the largest share of their coded feedback.
That lands directly on something we already do. Most notes we publish carry a colour, and most of those were set by hand rather than by a model. Their results are worth holding our palette up against, though it is a spot comparison rather than a test: they named broad colour groups for whole odour categories, and the notes below are one we picked to stand for each. Here are their groups against the colours we store.
| Colour their participants named | Smells they linked it to | Ours |
|---|---|---|
| Orange and yellow | Orange, fruit | Lemon #EBE047 |
| Green | Fresh, sour, mint, pungent | Mint #30C08B |
| Brown | Roasted, burnt | Coffee #874C22 |
| Blue | Fresh, mint | Sea Salt #5CADD6 |
| Black | Roasted, soil, ink | Oud #3A2417 |
| Red | Spicy, soil | Black Pepper #974226, the one we read differently |
Scroll the table sideways to see every column
Five of the six sit in the group their participants named, which is encouraging on six hand-picked notes rather than conclusive about the palette. The sixth is a real difference and worth naming: their participants reached for red when a smell was spicy, and we use a red-brown. We have left ours alone. Their red came from chilli and soil. Ours is Black Pepper's own colour, and a brown- shifted red keeps it legible next to the woods it usually sits beside. It is a considered disagreement rather than an oversight.
What it confirmed, and what it rules out
| What they found | What it means here |
|---|---|
| Colour carries more of a smell than any other part of a picture | Confirms the note colour on every note, chip, smellmap and ScentArt canvas we draw. It also tells us where to spend effort: a wrong colour costs more than a wrong shape |
| Curved, flowing lines read as juicy and fluid. Sharp lines read as pungent and spicy | Both of our renderers split them the same way, and the first did it before we read this. The classic engine drew sweet notes as loops, floral ones as curves and gave spicy notes a zigzag - taste, with no evidence behind it at the time. The material engine that now draws every perfume page keeps the mapping in physical form: spicy notes sharpen into crystalline shards with an angularity control, florals curve through convex petal cells, sweet matter sags into slow viscous lobes. Now the taste has some evidence |
| Texture works when it matches the physical stuff the smell comes from, like oil looking oily | This finding is now load-bearing rather than a limit. The material renderer resolves each of the 34 olfactory facets to a scanned physical material - tanned hide for leather, lichen for oakmoss, charred wood for smoke - so the texture on the canvas is the physical stuff the smell comes from, not an invented pattern. Where a smell has no photograph (sulphur is the character of a chemical bond) we say so and translate deliberately |
| An abstract picture needs a written description alongside it. Without one it scored worst of all seven recipes on matching the smell | The strongest argument for never showing a ScentArt canvas on its own. Every one sits beside its named notes, and the smellmap prints note names straight onto the picture |
| Pictures built from AI descriptions rather than human ones scored lower on matching the smell, and got called stereotyped | Why ScentArt draws only from note data we already hold, and is never asked what a smell looks like. Where a model does help assemble a note list, its output is a list of named notes a reader can check, not an image nobody can argue with |
Scroll the table sideways to see every column
The shape result is not theirs alone. Ryan Ward, Sophie Wuerger and Alan Marshall at the University of Liverpool tested 68 people across ten smells and found consistent matches for shape angularity too, along with colour, pitch and emotion. They also tested something different from the picture texture in the table above: not how a picture looks, but how smooth or rough the smell itself feels. That is a channel we do not encode at all, on the smellmap or anywhere else, and we are noting it here rather than quietly leaving it out.
The result that argues against us
A page like this is only worth reading if it reports the findings that go against us, so here is the main one. Plain, realistic pictures beat abstract ones at matching the smell, and the gap was statistically solid. ScentArt is abstract. On their evidence, a photograph of a lemon would tell you more about a lemon than our canvas does.
We have not changed it, for two reasons we can defend. The first is in their own numbers. Abstract pictures lost on matching the smell, but not on being liked or being interesting, and they drew far more discussion: roughly 69 coded remarks per recipe against 41.5 for the realistic ones. The second is that a perfume is not a lemon. It is fifteen ingredients moving over eight hours, and there is no photograph of that. The realistic option is not available to us.
Their finding that abstract pictures need words next to them is the one we act on. We never ship a ScentArt canvas as the whole answer. It sits with the note list, the notes are named, and the perfume page carries the plain-English verdict. The picture is the way in, not the evidence.
Why we do not let AI draw ours
This is a paper about making pictures with generative AI, and it is the reason we do not. Their own results show three places where a language model gets a smell wrong, and all three would hurt us badly.
- Odd things. GPT called an unlit cigarette earthy, woody and leathery. The people with it under their nose said fruity and sour, and 15 of them mentioned it. One said plum sweets.
- Strength. GPT called lily floral, fresh and sweet. The people in the room called it pungent, and five said it was overpowering. A model reads the idea of a flower, not the amount of it.
- Time. The grapes fermented during the study. Eight people smelled the alcohol. GPT kept saying fruity, juicy and fresh, because a model has no way to know a smell has moved on.
Time is the one that settles it for us. A perfume is a smell that changes for hours, and that is the exact thing the model could not see. So our evolution timeline is built from a stored how-long-it-lasts figure per ingredient, applied through one fixed curve, rather than from asking a model what the perfume smells like as it wears. We should be straight about the limit of that: those hours are themselves estimates made from the material class, not laboratory measurements. The difference that matters here is narrower than it sounds, and it is still the difference between a rule we can show you and an opinion you cannot check.
The image tool had its own problems. Midjourney drew cigarettes lit from the wrong end, could not handle see-through things like glue and mouthwash, and produced four noticeably different styles from one prompt. Our ScentArt is not generated. It is drawn by code from the perfume's stored profile, with a seed taken from that same profile. So one perfume always produces one picture, and that picture only changes when the data underneath it does.
One honest qualification, because "we do not let AI draw ours" is too broad as it stands. We do use an image model in one place: the pictures of raw ingredients on our note pages are model-made. The line we hold is the one this paper draws for us. A model may be asked to depict a named object, because a lemon is a thing it and you have both seen plenty of. It is never asked what a smell looks like, because that is the question the people in this study watched it get wrong.
Where it stops being about us
Two limits, both of which the authors raise themselves. Every participant was Chinese, and they show it mattering: the AI drew vinegar red and green, from Western wine vinegars, when Chinese vinegar runs from pale yellow to dark brown. Colour and smell do not pair the same way everywhere, and our readers are mostly in the UK. So we treat their colour table as strong evidence, not as a rule.
The bigger gap is simpler. Their participants had the real smell under their nose while they looked at the picture. Nobody reading a perfume page has that. We are asking a picture to do a harder job than the one it was tested on, which is a good reason to keep our claims for it modest.
See it
Six perfumes spanning the colour rows in their table, each drawn live from its own data.
- Acqua di Parma Colonia - a citrus cologne, and mostly yellow
- Acqua di Gio Profondo - an aquatic, and mostly blue
- Gucci Flora Gorgeous Gardenia - a fruity white floral, pink and orange
- Lattafa Khamrah - a spiced gourmand, and mostly brown
- Lattafa Asad - a woody spice, and the reddest of the six
- Louis Vuitton Ombre Nomade - a dense oud, nearly black
Every perfume on ScentVerdict has one - append /scentart to any perfume
URL. There is more on how it is built in the
ScentArt guide, and every renderer we have written
since March 2026 - eleven of them, including four studies that were built and never
shown - is in the renderers, each one
openable on any perfume.