Two long brush arcs running almost parallel with a narrow channel of bare ground held between them.

Can you measure whether two perfumes smell the same?

Research 7 min read

Every dupe on this site is a claim that two things smell alike, and we publish thousands of them. Two research groups have built proper ways to measure that, one starting from molecules and one from a neural network. We do it a third way, and ours is cruder than either. All three are set out below, because if you are going to trust a similarity score you should know what is behind it.

On this page
  1. A hundred-year-old question
  2. The first measure, and the perfect dupe
  3. The second measure, built by a neural network
  4. What we do
  5. Why we cannot use theirs
  6. What we do not measure at all
  7. See it
  8. Keep going

A hundred-year-old question

Alexander Graham Bell, 1914

"Can you measure the difference between one kind of smell and another? It is very obvious that we have very many different kinds of smells, all the way from the odor of violets and roses up to asafetida. But until you can measure their likenesses and differences you can have no science of odor."

The Weizmann team open their paper with that. Colour got its measure long ago: wavelength is physical, and we understand how it maps to what you see well enough to build screens out of it. Smell had no equivalent. Both groups below set out to close that gap.

The first measure, and the perfect dupe

Kobi Snitz, Noam Sobel and colleagues at the Weizmann Institute published Predicting Odor Perceptual Similarity from Odor Structure in 2013. They asked 139 people to rate how alike pairs of odour mixtures smelled, using mixtures of anything from 4 to 43 different molecules, then tested ways of predicting those ratings from chemistry alone.

The version that worked treated each mixture as a single vector rather than a bag of separate ingredients, and got to a correlation of 0.85 with what people said. The authors point out that this fits how the brain seems to handle smell: you get one impression, not a list of components. Anyone who has tried to pick out the individual notes in a perfume will recognise that.

In 2020 the same group published A measure of smell enables the creation of olfactory metamers in Nature, with the perfumer Christophe Laudamiel among the authors. They gathered 49,788 pairwise judgements from 199 people smelling 242 mixtures, and boiled the result down to a single number, measured in radians, built from 21 physical and chemical properties of the molecules involved.

Then they did the thing that should interest anyone who cares about dupes. They found a threshold, 0.05 radians, within which people found two mixtures very difficult to tell apart. They then used it to build metamers: pairs of mixtures with no molecules in common at all that smell the same.

That is the scientific version of a perfect dupe, sharing not one molecule with the original. It exists, and it is extremely hard: it took a formula, a lab and a 199-person study to confirm. We know of no dupe that has been tested against that standard.

A metamer. Two mixtures with not one molecule in common, producing a smell people find very difficult to tell apart. The dots are ingredients. No dot appears on both sides.

The second measure, built by a neural network

In 2023 a team at Google Research and the Monell Chemical Senses Center, led by Brian Lee, Emily Mayhew, Joel Mainland and Alexander Wiltschko with collaborators including the University of Reading, published A principal odor map unifies diverse tasks in olfactory perception in Science. They trained a graph neural network on the structures of around 5,000 molecules that had been described by a trained smelling panel, and produced a map they call the Principal Odor Map.

Two results stand out. On 400 molecules the model had never seen, its description of the smell matched the panel's average better than the median human panel member did, on 53% of them. And they used the map to place around 500,000 molecules nobody has ever smelled. For scale, the authors reckon smelling that set with their trained panel would take something like 70 person-years, which is the point of having a map at all.

The part that matters here is what the map's geometry means. One map, read three ways.

  1. Direction which descriptors apply
  2. Length whether it can be smelled at all
  3. Distance apart how alike two smells are
Three questions answered by one geometry. The third is the one this page is about, and on it a model built on the newer map beat a standard chemistry-based baseline on the Weizmann group's own 2013 data.

What we do

Ours is the same idea done with far less. We do not have molecules. What we have is notes, and for each note a set of 34 scores describing what it smells of. A perfume becomes the combination of its notes, and two perfumes are compared by the angle between those two sets of numbers.

The scoring is deliberately not name-matching. Lemon and bergamot read as close because they share the same underlying citrus scores, not because anyone wrote down that they are similar. We tried adding a name-based accord term and it made things measurably worse, so we removed it.

The honest comparison:

What is being compared Weizmann, 2020 Principal Odor Map, 2023 Us
Starts from Molecules Molecules Notes
Describes a smell as 21 physical properties A learned map position 34 facet scores
Similarity is An angle, in radians A distance on the map An angle between vectors
Checked against 199 people, 49,788 pairs A trained panel, 400 molecules held back from training 20 pairs of researched community opinion, the same 20 the method was chosen on
How well it did Good enough to build metamers Closer to the panel average than the median panellist, on 53% of them Rank correlation of 0.755, on the fitting set. No held-out test

Scroll the table sideways to see every column

That last column is three orders of magnitude less data than the Weizmann column, and it is weaker still than that makes it sound. Those 20 pairs are not a test set. We used them to pick the method in the first place, choosing cosine over the alternative on 0.755 against 0.747, and we fitted the displayed scale to them as well. The number describes how well the method fits the pairs it was chosen on. No perfume outside those 20 has ever been used to check it.

What we would say in our defence is that the kind of evidence is the same kind: real people saying whether two things smell alike. Ours are the fragrance community rather than a lab panel, gathered from research into what people say about specific pairs and cross-checked against two independent perfume databases. That makes it a decent sanity check on the design. It does not make it validation, and the two get confused often enough that we would rather spell it out.

It does mean our score is tuned to agree with people who have smelled both. When our engine says Armaf Club de Nuit sits very close to Creed Aventus, that number was calibrated against a community that has argued about exactly that pair for years. It also means the pairs it was calibrated on are the pairs it should be expected to do best on.

Why we cannot use theirs

The obvious question is why we do not use the better method. Two reasons, and the first is fatal.

We do not have the molecules. Both measures need the actual chemical make-up of what you are comparing. Perfume formulas are trade secrets, and not one of the perfumes in our catalogue has a published formula. We know a perfume contains bergamot. We do not know which molecules, in what proportion. Their methods are not too advanced for us. They run on something that does not exist outside a manufacturer's safe.

And there may be a licence to negotiate. The 2013 paper records that the authors disclosed the work to the institute's technology licensing company and began a patent application on the method and algorithm. We have not traced what became of it in each jurisdiction, so treat that as a question to answer rather than a settled obstacle. It is secondary in any case: without the formulas we cannot feed the method its input.

The Weizmann group do run a public tool, Odorspace, where you can compare formulas with their measure. It is worth a look. It also neatly demonstrates the problem, because to use it properly you need to type in the formula, which is the thing nobody has.

What our score cannot measure

The input is a published note list, not a formula. So the score does not know the proportions of anything, the concentration it is worn at, how the materials behave together, how the perfume changes over the hours, or how it behaves on a given person's skin. The 34 facets also flatten a whole perfume into one average vector, which is why two perfumes can score close while smelling noticeably different in the first ten minutes. Read a high score as a lead worth following, not as evidence that two bottles are interchangeable.

Reading these two papers made two further gaps obvious. The Principal Odor Map gets three things out of one geometry: which descriptors apply, how alike two smells are, and whether a molecule can be smelled at all, its detection threshold. We use the first two. We derive nothing about detectability or strength from our facets - our intensity and projection numbers come from elsewhere entirely - and the paper itself says the map cannot predict how strong a smell is above its threshold.

And a study by Ryan Ward, Sophie Wuerger and Alan Marshall at the University of Liverpool tested what 68 people match to smells across a whole range of channels. They found consistent results for colour and shape, and for one channel nobody else on these pages tested: texture, how smooth or rough a smell itself feels. We encode colour and form. Texture we ignore.

None of this is a promise to fix any of it. These are things we know we are not doing, and we would rather write them down than let a page like this read as a victory lap.

See it

Five perfumes, each showing what our engine thinks sits closest to them.

Every perfume on ScentVerdict has one - append /similar to any perfume URL. The dupe pairs built on top of this live in the dupe directory.

Keep going