How Our Sensory Panel Works: Not "Liking" It, but Seeing the Difference

6 min read

The panel's job is to run the human nose and palate in a way that yields repeatable results. Its output is not preference but data that can be translated into a formulation.

The most common misunderstanding we meet in flavour development is this: people assume the sensory panel convenes to tell us whether we like a sample. Yet what the panel produces is not preference but data. The sentence "this is nice" cannot be translated into a formulation. The sentence "the top note is short, it drops after the first ten seconds; the trace left in the mouth is weak" can be translated directly — it tells us which component's volatility to adjust and which carrier to change.

That is why, when we built our panel, the question we asked was not "who has a good palate?" but "who can see a difference consistently and describe it?". The two are not the same thing.

Selecting panellists: what we are looking for is repeatability

We screen candidate panellists with threshold tests. We present dilute solutions of the basic tastes (sweet, sour, salty, bitter, umami) as a coded, ascending series starting from the lowest concentration; we measure at which concentration the candidate can detect the taste and at which point they can discriminate it. We apply the same method for odour thresholds. Then we repeat it on different days. Because hitting the mark once may be luck; what we are really after is the same candidate producing a similar result weeks apart.

An exceptionally sensitive nose does not on its own make a good panellist. Inconsistent sensitivity is noise. A moderately sensitive but extremely stable panellist is far more valuable in data terms.

A shared language: a difference that cannot be described cannot be used

Selected panellists are then trained on a descriptive vocabulary. The aim is for the word "fresh" to mean the same thing in everyone's head. We work with reference samples: this is the cooling intensity of a menthol reference, this is the peel-like sharpness that limonene gives, this is the creamy softness of vanillin. Unless terms are anchored to references, panel notes turn into a text in which everyone tells their own story.

  • Note position: top, middle, base — which dominates, which is missing.
  • Time profile: speed of opening, peak point, decay curve, the trace left in the mouth.
  • Intensity: on a zero-to-ten scale, anchored to a reference.
  • Character: defined terms such as green, roasted, fruity, soapy, metallic.
  • Defect: burnt, cardboard, oxidised, plastic-like — if present, it is written down by name.
Coded flavour samples and evaluation forms in the laboratory
Samples go to the panel with three-digit codes, with the order of presentation rotated.

Blinding and the triangle test

No sample reaches the panel under its own name. Every sample is given a random three-digit code and the order of presentation is varied from panellist to panellist. The panellist does not know which sample is the new formulation and which is the current product. Knowing that would corrupt the result — the human brain is all too willing to taste something differently when it knows it is new.

The tool we use most often is the triangle test. Three samples are placed in front of the panellist: two identical, one different. There is only one question — which one is different? Preference is not asked. This test gives a statistical answer to the question "is there a difference or not?". When a customer wants to change a raw material, or when we substitute a component for cost optimisation, this is the critical question: will the consumer notice the change? If a trained panel cannot catch the difference in a triangle test, the odds of that difference registering with the consumer are low. But this is not a guarantee: failing to find a difference does not prove that no difference exists — it only shows that it could not be measured under those conditions. That is why we base a substitution decision not on a single test but on evaluations repeated in the final matrix and across the shelf life.

Once a difference is found, we move to the second stage: descriptive analysis. The question is no longer "is there a difference" but "where is the difference". The panel scores each sample against defined terms and the results are plotted as a profile chart. The profile of the target sample and that of the formulation under development are overlaid; where it falls short and on which axis it overshoots becomes visible. What R&D will change in the next trial comes out of the deviation on that chart; this is also the data that feeds the revision rounds on the road from brief to series production.

“Before we trust the panel we test the panel: the same sample, on different days, has to give the same answer.”

What the instrument says, what the panel says

Chromatographic analysis gives composition: which volatile component, and in what quantity. That is indisputable data, but it is incomplete. A component that appears as a tiny peak on the chromatogram can define the character of the product single-handedly because its perception threshold is low. Conversely, a component that dominates by mass may do almost no work in perception. Analysis measures quantity; it does not measure perception.

That is why we read the two together. When the panel says "the roasted note is too much, there is a slight burnt trace", we go back to the analytical data and look at the group of components that carry that character. The reverse also happens: we add a component we judged to be missing relative to the reference sample, then test with a triangle test whether the panel actually notices it. Analysis tells us where to look; the panel tells us whether it was worth looking there.

Evaluation in the matrix: testing the flavour where it is going to live

The panel's final task is to test the flavour in the matrix it is destined for. A flavour that looks perfect in solution may not survive oven temperature in a biscuit dough; it may give an entirely different profile when heated in shisha tobacco; it may lose its day-one sharpness within a month in a fatty seasoning powder. So we repeat the evaluation as far as possible in the final product, on samples close to production conditions, and across the shelf life.

Blind reference comparison comes in here: the product the customer is targeting and our formulation go to the panel in the same matrix, under the same coding discipline. As long as the panel can tell them apart, we have a described deviation and a next trial to run. At the point where it cannot tell them apart, we have closed in on the target — but we confirm that not with the result of a single session, but with evaluations repeated across matrix and shelf life.

Let us define your flavour requirement together

Share the profile you are looking for, your application area and your process conditions; our team will come back to you with a sample and a quotation.