What we have measured
A coordinate system that cannot be checked is an opinion with decimal places. These are the numbers so far, with their limits attached rather than in a footnote.
The figures
Each figure carries its own caveat, on the same card rather than in a footnote at the bottom of the page.
Where the models disagree, and why that is fine
The disagreements are the interesting part, because they are not random. On common notes the models split on saffron and vetiver, which working perfumers would also split on. On rare notes they split on rubber, read as smoky by one and leather by the other, and on mignonette, read as floral by one and green by the other. Every one of those readings is defensible.
We went looking at the rare end deliberately, because the first measurement only covered well-documented notes and could have been flattering itself. Agreement does fall there: 0.94 to 0.90, and leading-facet agreement from 92% to 78%. The median barely moves, which tells us this is a handful of genuinely ambiguous materials rather than a general loss of quality.
The worst score in that test was jojoba, at 0.52. Jojoba is a carrier oil with almost no smell of its own, so there was nothing for the two models to agree about and both were fitting noise. That is not a modelling problem, it is a data problem: a material with no meaningful odour should not carry a confident position at all. Finding it is the reason to run the test, and separating those materials out is now part of the build.
Are families even the right shape?
The release checks that too, on the hardest case. Fougere is not really a region of a space - it is a structure, the co-presence of lavender, coumarin and oakmoss. If a position-only model can recover it, families are regions. If it cannot, they are not.
On 202 fougeres defined by that structure, a position-only classifier scored 0.79, well above chance, but 27% of them had a non-fougere as their nearest neighbour in the space. So position carries real signal and is not sufficient on its own. The model reads a family as a region and a rule, which is why the note list matters as much as the coordinates.
What none of this proves
Two models agreeing means the space is stable, not that it is right. Both learned from overlapping public text, so they could share a mistake rather than a fact. Showing the numbers are correct rather than merely stable means anchoring them to external measured data. That work is under way, and this page will report the result whichever way it goes.
What we will not claim
Published note lists are marketing documents, not formulas. Brands substitute appealing names for unglamorous chemicals, invent notes that correspond to no material, and revise their pyramids in step with fashion rather than with the liquid in the bottle. Any system built on them inherits that.
So the honest description of ScentPrism is that it is one classification, applied identically to every fragrance we hold and published with its own error bars. We have not benchmarked it against every alternative and will not claim it beats them. It is not a chemical analysis either, and we will not describe it as one.