Beyond the First Sip: What Jack Daniel's RTDs Taught Us About Measuring Taste Down to the Last Drop
The quant said near-parity. The open-ends said something was off. Here's what we found when we stopped testing one moment and started testing the whole drink.
There's a particular kind of uneasy feeling in product development. Your test results come back strong, liking is at parity, purchase intent is healthy. Everything on the scorecard says go.
And then you read the open-ends…
That's roughly where Brown-Forman found themselves with a low-calorie Jack Daniel's RTD. Which is why, at Quirk's New York, I shared the stage with Amber Banet, Manager of Product Insights at Brown-Forman, to talk through what happened next - and what it taught us both about a gap sitting quietly in the middle of most product testing.
The gap is this: taste isn't a snapshot. But almost all of us test it like one.
RTDs are booming. They're also the most reformulated category in BevAlc
A bit of context on why this matters now, and not in five years.
The global RTD market is on track to pass $40 billion by 2028, growing around three times faster than full-strength spirits (IWSR, 2024). That growth alone would be a story. But the more interesting number is what's happening inside the liquid.
RTDs are now the number one category for reformulation spend in BevAlc (Mintel Global New Products), and more than 60% of new RTD launches include a low-calorie variant (Kantar Worldpanel, 2024).
So: brands are reformulating fast, at scale, in a category where there is essentially no margin for error on taste. Which puts a lot of weight on how well your testing actually reflects the way people drink.
Most product tests take one snapshot
Think about how someone actually drinks a can.
There's the first sip. There's mid-drink. There's three-quarters through. And there's the last drop - the bit that decides whether they reach for another one.
Most testing measures the first of those and treats the rest as an afterthought. Sip-and-spit protocols, first-sip evaluations, single-moment liking scores. All perfectly rigorous. All looking at one frame of a much longer film.
Here's the part worth sitting with: timing of measurement isn't just a research decision, it's a product decision. What you measure, and when, determines what your R&D team ends up optimizing for. If your data only exists at sip one, sip one is what gets fixed.
Low-calorie variants amplify every timing risk
This is where the snapshot problem stops being theoretical, because alternative sweeteners simply don't behave like sugar across a full serving.
Aftertaste builds. Stevia, monk fruit and acesulfame-K tend to reveal themselves in the finish rather than the opening. A first-sip test routinely misses this entirely.
Bitterness creeps in. Some low-calorie sweetener blends develop a bitterness mid-drink that genuinely wasn't there at the start - invisible in a single-moment evaluation.
Mouthfeel fades. Reduced-sugar formulations often change the textural arc of a drink. What feels crisp at sip one can feel thin or watery by the finish.
And liking diverges. Full-calorie and low-calorie variants frequently score almost identically at first sip, then pull apart significantly at mid-drink and finish. A snapshot test tells you that you have parity. The full arc tells you that you have a gap.
Which is, more or less, exactly what Amber's team was starting to suspect.
The brief: "Tell us how they feel at the bottom of the can"
Jack Daniel's RTDs are one of Brown-Forman's fastest-growing platforms, and low-calorie was a clear white space in the portfolio. The question was never should we go after it? It was how do we make sure it's genuinely great before we do?
As Amber put it on stage, the ambition wasn't to make a lighter version of the original. It was to make something that still delivered like the original - that felt like Jack Daniel's and tasted like Jack Daniel's, from the first sip to the bottom of the can.
So the team ran the standard toolkit. Central location tests. First-sip evaluation. Head-to-head against the full-calorie benchmark. And on those measures, the low-calorie variant performed solidly: headline liking near-parity, purchase intent strong, all the things you want to see.
But the open-ends kept saying something else. Comments about the finish. About aftertaste. About something being slightly off that respondents couldn't quite name.
What they knew: first-sip liking was at near-parity, the flavor direction was right, and consumers said they'd buy it.
What they didn't know: whether that parity held at mid-drink, which specific attribute was driving the finish dissatisfaction, whether the gap was real or anecdotal - and whether they were optimizing the wrong moment entirely.
The brief they brought us was refreshingly plain: don't just tell us what people think at the first sip. Tell us how they feel at the bottom of the can.
Simple to say. It required a fundamentally different research design.
How we captured the full arc
Four things made it work, and none of them were built as a one-off - this is how multi-moment measurement works in Product Hub as standard.
Multi-moment measurement. Evaluation at first sip, mid-drink and finish, within a single serving, with the same sensory and liking KPIs at every moment. That consistency is what makes within-occasion trend analysis possible at all.
Standardized sensory KPI tracking. Sweetness, bitterness, carbonation, mouthfeel, aftertaste intensity and aftertaste pleasantness, tracked at each moment using Product Hub's built-in sensory framework.
Liking and intent at every timepoint. Not as a blended average. This is the piece most testing misses: ask "how much do you like this overall" and you get a number smoothed across the whole experience. Ask at each moment and you can see the shape of it - where liking holds, where it slips, and how fast.
One connected platform. All timepoint data unified, so cross-moment comparisons, variant benchmarking and norm generation happen in one place rather than across a dozen spreadsheets.
Amber flagged three design decisions from the process that I think are worth calling out, because they weren't obvious:
- What does a "moment" mean? There's no universal standard for when to interrupt someone mid-drink. We landed on first sip, mid-can and last drop because those are real behavioral moments - the decision to keep drinking after trial, the sustained experience, and the impression that lingers. They were anchored to the commercial questions, not chosen arbitrarily.
- Which attributes matter? Sweetness and aftertaste were obvious. We added carbonation, mouthfeel and bitterness because those are the attributes most likely to shift across a full serving in a carbonated RTD. That selection was a real conversation, not a default list.
The framing changed. In a standard test you evaluate, get your data, and you're done. This felt closer to tracking: the same respondent, the same product, across the same occasion. It stopped being "what do consumers think of this drink" and became "what is their experience of this drink over time."
What the full arc revealed
The headline finding, as Amber delivered it:
First-sip liking was near-identical for both variants. By the finish, there was a meaningful gap - driven by aftertaste intensity in the low-calorie variant. The concept that looked like a winner at sip one looked materially different by the last drop.
Three things sat underneath that.
Aftertaste was the culprit. It built progressively across the serving. Barely detectable at first sip - which is precisely why the standard protocol had missed it - and the dominant sensory signal by the bottom of the can. That confirmed what consumers had been trying to articulate, and turned it into something actionable.
Sweetness diverged mid-drink. The full-calorie variant held a stable sweetness perception. The low-calorie version showed sweetness building - not dramatically, but consistently. By mid-can, consumers were describing the low-calorie variant as sweeter than the full-calorie original. That's the opposite of the intended positioning, and it is completely invisible at first sip.
And the finding that genuinely surprised us: mid-can liking for the low-calorie product began to decline before aftertaste had peaked. The experience was degrading in real time, not just leaving a poor final impression. The arc of the problem mattered as much as the endpoint.
That last one reframed the reformulation brief entirely. This wasn't a finish problem to patch. It was a sustained experience problem to solve.
The pattern we now look for across low-calorie RTDs
We see a consistent shape across multiple category studies - this isn't specific to any one brand or product:
Carbonation sitting at parity was reassuring - but it was also useful, because it told the team where not to spend reformulation effort. That's an underrated output.
The point of the table isn't that every low-calorie RTD has all of these problems. It's that if you aren't measuring these attributes across the full occasion, you won't find out whether you have a problem until consumers quietly stop rebuying. And by then it's an in-market problem, not a development one.
Three things this changes about formulation goals
Optimize for the finish, not just the open. If your low-calorie variant scores a seven at first sip and a four at the finish, you're working on the wrong moment. The finish is where loyalty is won or lost.
Measure what matters, at the moment it matters. Aftertaste and bitterness build. Mouthfeel fades. If you only measure at sip one, those attributes don't exist in your data. That isn't a consumer truth - it's a measurement gap.
Connected data makes iteration faster. When multi-moment data lives in one platform, you can compare across moments, variants and reformulation rounds, and your norm base compounds with every study. You stop running isolated tests and start building a product intelligence program.
Beyond the first sip
Product Hub by MMR is built specifically for end-to-end consumer product testing, with 35 years of MMR's sensory science built into the platform: multi-moment measurement as standard, standardized KPI tracking so data stays consistent across studies and markets, AI-powered analysis through Verbatim Explorer and Penalty Analysis, global reach with local execution, and automated reporting the moment fieldwork wraps - so your team reads insight instead of building charts.
If you're reformulating anything where the experience unfolds over time - and in food and drink, that's nearly everything - the question worth asking isn't “do consumers like this?”.
It's “do they still like it at the end?”
Book a demo and we'll show you what multi-moment testing looks like in your own category.
Blog.
The latest insights from our product experts