Do AI product recommendations work if you can’t measure them against a holdout?

You add recommendations to a product page. A customer clicks one. They add it to their cart and buy.
That looks like a win.
But it leaves one hard question: would they have bought that item anyway?
Maybe they came to the store looking for it. Maybe they found it through search. Maybe they would have seen it in your collection pages. A recommendation click tells you the customer interacted with the widget. It does not, on its own, show that the widget caused the sale.
That difference matters when you are deciding whether a recommendation tool earns its place on your store.
Affina is being built around that question. It will provide AI product recommendations on product and cart pages, learned for each store. It will measure them against a real 10% holdout, rather than reporting only activity from shoppers who saw recommendations.
Affina is pre-launch. It is not available to install today.
Sales after a click are not the same as incremental sales
Most recommendation reporting starts with an easy chain of events:
- A shopper saw a recommendation.
- They clicked it.
- They bought something.
Those events are useful. You should know whether shoppers notice a recommendation module. You should know which products get clicks. You should know whether a recommendation appears before an order.
But the chain does not prove cause.
Say your store sells running shoes. A shopper opens a shoe product page, sees recommended socks, clicks the socks, and adds them to their order.
The recommendation may have introduced the socks.
Or the shopper may already have planned to buy socks. They may have searched for them next. The recommendation got credit because it was in front of them at the right moment.
This is the attribution problem. A system can count orders connected to recommendations while overstating the sales it actually created.
Shopify’s guidance on causal inference makes the same distinction. Descriptive data can describe what happened, but it provides no direct evidence that one thing caused another. Shopify ranks randomized A/B tests as the strongest form of evidence for this kind of question.
A holdout is the practical version of that test.
What a 10% holdout does
A holdout means a small, randomly selected group does not receive the treatment being measured.
For Affina, the treatment is the AI recommendation experience.
The plan is straightforward:
- Most eligible shoppers see Affina recommendations.
- A real 10% holdout does not see those recommendations.
- Both groups are measured over the same period.
- The difference between the groups is the evidence to examine.
If shoppers who see recommendations place more orders, add more items, or spend more than the holdout group, that gap is more meaningful than a raw recommendation-attributed sales total.
It is not perfect knowledge of every shopper’s intent. No measurement is. But random assignment gives you a comparison group that did not receive the recommendation experience.
Without that group, you are comparing recommendation activity to nothing.
With it, you can ask a better question: did showing recommendations change what happened?
Why a small holdout is worth keeping
It can feel uncomfortable to hide recommendations from some visitors.
After all, if you believe recommendations help, why not show them to everyone?
Because turning them on for everyone removes your best way to tell whether they help.
A holdout is the cost of learning. In this case, 10% of eligible traffic gives Affina a group for comparison while the remaining 90% sees recommendations.
The alternative is often worse: show a widget to everyone, see some clicks and sales, then make a long-term decision from numbers that cannot separate cause from coincidence.
This is especially important for product and cart pages.
Customers on those pages already have strong buying intent. A shopper who reaches the cart may be close to checkout before any recommendation appears. If they add another product after seeing a recommendation, the recommendation may have helped. It may also have been present when a high-intent shopper was already going to add more.
The holdout gives that question somewhere to land.
Offline scores are not your store’s answer
Recommendation models can be tested before they reach a shopper. You can check whether a model predicts products a customer later viewed or bought. You can compare one model’s ranking against another.
Those checks matter during development. They are not the final answer for a merchant.
Your store has its own catalogue, margins, stock position, product relationships, traffic sources, and customer behaviour. A model can look promising in a test dataset and still fail to improve the actual shopping experience.
Shopify evaluates its own recommendation systems online, not only with offline model scores. In one online A/B test, Shopify reported changes in orders, high-quality click-through rate, conversion rate, and product recall. The important part is not that every merchant should expect those figures. You should not assume that.
The important part is the method: test with real shoppers in a live environment.
Affina’s measurement is intended to follow that same principle. The question is not whether an AI model can produce a plausible list of related products. The question is whether its recommendations create a measurable difference for your store.
What to check before you judge any recommendation tool
You can check one thing in under a minute: open the reporting for your current recommendation app, theme module, or analytics setup.
Look for the comparison group.
Ask these questions:
- Does the report show results for shoppers who did not see recommendations?
- Were shoppers assigned to groups randomly?
- Is the reported sales number just orders after a recommendation click?
- Can you see the period being compared?
- Does the tool explain what happens when a shopper returns later on another device or session?
If the report only says “recommendation revenue” or “sales influenced,” treat that as activity reporting, not proof of incremental revenue.
That does not make the report useless. It can help you spot products that are often viewed together. It can show whether the placement gets attention. It can help you find weak recommendation sets.
It just cannot answer the causal question by itself.
What Affina will need to earn
Affina should not be judged by whether its recommendations look clever.
It should be judged by whether the measured group that sees them performs differently from the group that does not.
That means looking past clicks alone. A click can be curiosity. A recommendation can also shift a purchase from one product to another, rather than add value to the order. A useful measurement setup needs to show the business outcome you care about, with the holdout beside it.
For one store, that may be orders. For another, it may be items per order. For a catalogue with repeat purchases, it may be different again.
The useful promise is not that AI recommendations always work.
It is that you should be able to tell whether they work on your store.
Affina is pre-launch, and its purpose is to make that test possible on product and cart pages: store-specific recommendations, measured against a real 10% holdout.
If you are reviewing recommendation tools now, ask each one to show you its control group. If it has none, you can still learn what shoppers clicked. You cannot reliably claim what the recommendations caused.