Testing packaging or a marketing message by comparing variants
Two or three versions of a packaging or a tagline, a synthetic panel, a measured comparison of purchase-intent distributions instead of a single opinion in a meeting.
A new packaging design is ready, or a new tagline just got written, and the question lands in the meeting: which of the two or three versions should ship? The answer is often decided by a show of hands, shaped by whoever in the room speaks loudest. That is fast, but it measures the opinion of the room, not the reaction of the market.
The problem with gut-feel choices
Asking internally "which one do you prefer" almost always produces a polite consensus that reflects the taste of whoever is deciding, not the buyers the product is aimed at. Asking the same question to a general-purpose chat assistant does not change much either: you get an opinion, a single point of view, with no variance, no distribution of responses, no measure of uncertainty. That is fine for rewording a sentence or catching an off tone, not for deciding between two options that will reach thousands of different people.
Compare distributions, not preferences
The right question is not "which one looks better" but "which one drives more purchase intent, and with whom." Showing each variant to a panel of distinct respondents and comparing the resulting distributions reveals things a show of hands never will: one variant can win a majority while sharply polarizing part of the audience, another can seem less popular in the room while generating a stronger and more consistent purchase intent overall.
Isolate what matters
For the comparison to mean anything, isolate one variable at a time. Three common cases:
- Visual packaging: two designs for the same product, same copy, same displayed price.
- The core message: the same image, two different taglines, to see which one actually carries the purchase intent.
- The positioning angle: the same product framed around two different promises (convenience versus enjoyment, for instance).
Testing several variables at once in a single run makes the result hard to interpret, since it is no longer clear whether the gap comes from the visual or the copy.
What Panelia adds in practice
Panelia simulates hundreds of synthetic respondents per test, calibrated against real human data, and returns for each variant a purchase-intent distribution with confidence intervals, along with verbatims that explain the reasoning behind the reactions. Comparing two or three versions of a packaging or a message takes about ten minutes and costs roughly one euro per test, far less than a classic creative test that requires recruiting a sample and waiting weeks. The method follows a protocol published on arXiv (2510.08338), not a model's hunch.
This is not a tool that replaces creative judgment; it is a decision-support tool that settles a team debate with a measurement instead of the loudest voice in the room. For a launch with very high stakes, a complementary human test is still worth running before the final commitment.
In practice
The whole loop fits into one working session: describe the base concept, prepare two or three variants that isolate the variable being tested, run the test, compare the distributions and read the associated verbatims. The variant that wins is not always the one that got the most nods in the meeting, and that is exactly the point: it corrects a bias nobody in the room sees coming. Running the same comparison again a few weeks later, once the wording or the visual has been refined, costs little enough to make it a habit rather than a one-off exercise.
Frequently asked questions
- How many variants can be compared at once?
- Two or three variants is usually enough; beyond that, the gaps between distributions become harder to read and to act on clearly.
- Is asking ChatGPT which one it prefers the same thing?
- No, a general-purpose assistant gives a single opinion with no distribution and no confidence interval, while a synthetic panel simulates hundreds of distinct reactions.
- Should packaging and message be tested together?
- It is better to isolate one variable at a time (the visual, or the copy, or the price shown) so you know which one actually explains the gap you observe.