What a sealed study looks like, start to finish.
The cold brew concept screen, from the questionnaire to the scorecard, with the two questions we got wrong.
Neon Blue researchFigures from the scorecard Northwind Beverages received.
In short
- The seal comes first. Every prediction was hashed and time-stamped before one real person answered. Record 7c3e91a2, 9 September 2026, 14:02 UTC.
- 91 of 100, beside three other numbers. Chance scores 50, a published table 84, and the same people asked twice agree 93. The score means nothing without the three.
- Two questions missed. Where people would buy it (11 apart) and which name they remembered (9 apart). Both are printed first, and both changed the next study.
- Stated intent overshot by 26 points; the simulation landed within 3. Eight weeks of sales settled it.
The study arrives
Northwind sent us a questionnaire they had already fielded once: fourteen questions about a cold brew concept, a 12 oz can at $4.49, aimed at adults 21 to 54 in metros of 250,000 or more who had bought ready-to-drink coffee in the last thirty days. Nothing was rewritten. The point of a sealed study is that the customer's own questions go through unchanged, so the scorecard is about their decision and not ours.
The simulation was built the same afternoon: 1,200 people, one per census record inside that audience, given the habits of their ZIP at the rate they occur there, and, for the 340 who matched Northwind's loyalty file, a twin built from their own orders with seven of each person's real orders held back as the test.
Figure 1. The audience as it was drawn: real people as solid marks sized by the census weight they carry, simulated people clustered around the person whose cell they match.
The seal
Each simulated person answered each question alone, with its reasoning, and the answers were written to a record whose hash went to Northwind and to a public timestamp. From that moment nothing in the run could change. This is the part of the method that costs nothing and settles the most arguments: when the real answers arrive there is no version of the prediction that was made after them.
The scorecard ended a two-year argument with finance in one meeting.
The real people answer
A thousand people took the same fourteen questions over the following six days, recruited to the same audience definition. Their answers attached to the sealed record only after the field closed. We then asked 180 of them the same questions again a week later, which is where the ceiling comes from: the same people agree with themselves 93 times in 100, so no simulation should be judged against 100.
The scorecard, misses first
Agreement across the fourteen questions was 91: one hundred minus the average gap between simulated and real answer shares, across every option. Chance on these questions scores 50 and the best published table 84. Two questions sat far from the ceiling and the card opens with them.
Where would you expect to buy it? Real people said convenience store (31%); the simulation said grocery (36%). Eleven points apart. The simulated people carried the ZIP's shopping rates but not the commute, and cold brew is bought on the way to somewhere. Which name do you remember after one look? Real people 44%, simulated 53% for Slow Pour. Nine apart: the simulation over-remembers, which we now know and correct for.
The four closest questions, the claim, purchase intent, the pack and who the drink is for, were within three points. Those are the ones the launch was built on.
What changed
Two things. Commute and channel now enter every audience built for a ready-to-drink category, drawn from the same public sources as the rest. And recall questions carry a confidence label of Medium until a category has earned better, which shows on the card and in the report. Northwind's next study, the Q4 price ladder, scored 93.
- 17 Sep 2026
One accuracy number is not a number.
Why every score we publish sits beside chance, a published table and what the same people say when asked twice.
- 10 Sep 2026
What one person's data can and cannot tell you.
A twin answers as a person is likely to answer. It never says what they will do, and we refuse the questions that ask.
- 3 Sep 2026
A quarter with a hundred answered ideas.
What changes in a team when the research line stops rationing questions.