What we have measured

A coordinate system that cannot be checked is an opinion with decimal places. These are the numbers so far, with their limits attached rather than in a footnote.

The figures

Each figure carries its own caveat, on the same card rather than in a footnote at the bottom of the page.

0.94 to 0.90 Agreement between two independently trained models scoring the same notes 0.94 on common notes, 0.90 on rare ones. Chance floor: 0.17
92% to 78% Of notes where both models picked the same leading facet, common to rare The disputed ones are genuinely disputable: saffron, vetiver, rubber
32 / 32 Facets that resolved to terms in an external, perfumery-native reference vocabulary, at v8 Guarded by a test, not by a promise. Facets added since are covered by the same test
4,626 Materials imported with their published odour terms, keyed by CAS number Carrying 20,448 term assignments between them
32,386 Perfumes positioned in the space, the set the model must survive contact with Built from 292,082 note placements

Where the models disagree, and why that is fine

The disagreements are the interesting part, because they are not random. On common notes the models split on saffron and vetiver, which working perfumers would also split on. On rare notes they split on rubber, read as smoky by one and leather by the other, and on mignonette, read as floral by one and green by the other. Every one of those readings is defensible.

We went looking at the rare end deliberately, because the first measurement only covered well-documented notes and could have been flattering itself. Agreement does fall there: 0.94 to 0.90, and leading-facet agreement from 92% to 78%. The median barely moves, which tells us this is a handful of genuinely ambiguous materials rather than a general loss of quality.

The most useful failure

The worst score in that test was jojoba, at 0.52. Jojoba is a carrier oil with almost no smell of its own, so there was nothing for the two models to agree about and both were fitting noise. That is not a modelling problem, it is a data problem: a material with no meaningful odour should not carry a confident position at all. Finding it is the reason to run the test, and separating those materials out is now part of the build.

Are families even the right shape?

The release checks that too, on the hardest case. Fougere is not really a region of a space - it is a structure, the co-presence of lavender, coumarin and oakmoss. If a position-only model can recover it, families are regions. If it cannot, they are not.

On 202 fougeres defined by that structure, a position-only classifier scored 0.79, well above chance, but 27% of them had a non-fougere as their nearest neighbour in the space. So position carries real signal and is not sufficient on its own. The model reads a family as a region and a rule, which is why the note list matters as much as the coordinates.

What none of this proves

Reproducibility is not truth

Two models agreeing means the space is stable, not that it is right. Both learned from overlapping public text, so they could share a mistake rather than a fact. Showing the numbers are correct rather than merely stable means anchoring them to external measured data. That work is under way, and this page will report the result whichever way it goes.

What we will not claim

Published note lists are marketing documents, not formulas. Brands substitute appealing names for unglamorous chemicals, invent notes that correspond to no material, and revise their pyramids in step with fashion rather than with the liquid in the bottle. Any system built on them inherits that.

So the honest description of ScentPrism is that it is one classification, applied identically to every fragrance we hold and published with its own error bars. We have not benchmarked it against every alternative and will not claim it beats them. It is not a chemical analysis either, and we will not describe it as one.

Keep going