Quick answer: A catalog holdout test splits your products into two matched groups, enriches one, leaves the other untouched, and compares them over the same period. It is the only way to separate the effect of your data work from seasonality, promotions, creative changes, and attribution drift. It is also easier to get wrong in fashion than in almost any other category, because a single flash sale landing on one side can manufacture a result that was never there.
Why "fix everything and compare to last month" does not work
The instinct after any catalog project is to look at last month against this month. In apparel that comparison is close to meaningless, and this year it is worse than usual.
Four things move underneath you at the same time. Demand is seasonal, and the seasonality in fashion is severe rather than gentle. Your promotion calendar is not flat. Creative refreshes on its own schedule. And reported attribution changed twice in 2026, once in January and once in March, which made conversion numbers smaller without a single order changing hands. We covered that in why Meta ROAS is vanishing.
There is a fifth problem specific to apparel. Between two periods, the catalog itself is different. New arrivals landed, a season sold through, some styles went out of stock and some were retired. You are not comparing a catalog to its earlier self. You are comparing two different catalogs and attributing the gap to your work.
A holdout removes all of it in one move. Both groups live through the same season, the same promotions, the same creative, and the same attribution rules. Whatever differs between them is much more likely to be the thing you changed.
The unit of randomization is the product, not the customer
Most ecommerce testing splits people. Geo holdouts, audience splits, lift studies. Catalog work cannot be measured that way, because product data is global. The moment you enrich a garment, every shopper sees the enriched version, on Google, on Meta, in onsite search, on the product page. There is no way to show one visitor the thin version and another visitor the rich one.
So you split the catalog instead. Half the products get the work, half stay as they are, and every shopper sees a mix of both.
That has one consequence worth sitting with before you start. Your sample size is the number of products, not the number of visitors. A store with 40 styles is running a 20 against 20 test. That is a small n, and it is exactly why how you build the groups matters more than anything else you will do in this project.
What breaks a fashion holdout
Most failed catalog tests fail for one of these reasons, and none of them are statistical exotica. They are ordinary operational events.
A promotion lands on one side. This is the big one. Flash sales are not random. They are triggered by excess inventory, which correlates with poor sell through, which is related to the thing you are measuring. If a sale hits three products in the enriched group, you have measured the sale.
Category imbalance. Dresses, outerwear, and tops behave differently on conversion rate, return rate, and price point. If one group ends up dress heavy, you measured category and called it enrichment.
Bestseller concentration. Fashion catalogs are top heavy. A small number of styles usually carry most of the revenue. If two of those land on one side, that side's numbers are that side's bestsellers, and everything else is noise around them.
Season transition. A test that runs across a season change is not one test. Demand moves between categories, and your two groups almost certainly do not hold the same mix of categories.
Stockouts. A style selling out mid test stops earning for the rest of the period. In apparel this happens by size first, which is quieter and easier to miss.
Email and organic social. A newsletter featuring three products can move more revenue in a week than the entire effect you are trying to detect. If those three products sit in one group, the test is contaminated.
New arrivals. Anything launched during the test window has no baseline and a launch curve of its own.
How to build balanced groups
Do not randomize the whole catalog in one pass. With small product counts, plain randomization produces unbalanced groups often enough that you should assume it will.
Stratify first. Sort every product into blocks on the dimensions that actually predict performance:
- Category
- Price band
- Trailing 90 day revenue decile
- Conversion rate
- Stock depth
- Product age, so that new and mature styles spread evenly
Then randomly assign within each block, so each block contributes half its products to each side. Handle your top revenue styles explicitly rather than letting chance decide. If five products carry 40 percent of revenue, pair them off deliberately.
Then check your work before you start. Pull the last 90 days for both groups and compare revenue per product, sessions per product, conversion rate, and return rate. If those do not line up closely, reshuffle and check again. Two groups that already differed before the test cannot tell you anything after it.
Finally, write the operating rules down before day one, because the pressure to bend them arrives later:
- Any promotion applies to both groups, or the affected products are excluded from the analysis.
- Products launched during the test are excluded, or assigned evenly and analyzed separately.
- Products that go out of stock are flagged with the date, not silently dropped.
- Nobody changes titles, images, or descriptions on the control group, including for reasons unrelated to the test.
That last rule is broken more often than any of the others, usually by someone doing their job well.
Step zero: baseline both groups before you touch anything
Here is the step almost everyone skips. Before any enrichment happens, record how complete the product data already is on both sides.
Without that baseline, you cannot answer the only question that matters at the end. If the enriched group outperformed, was it because you enriched it, or because it started richer? And if nothing moved, was the work ineffective, or did the enrichment barely change anything because those products were already well described?
You need two numbers per group, and they need to be separate. Required field coverage, which decides whether an item is eligible to show at all. And attribute depth, which decides how well the delivery and ranking systems can place an item that is already eligible. A catalog scoring high on the first and low on the second is common, and it is the pattern where enrichment has the most room to move.
SixFit Catalog Checkup is a free Shopify app that produces exactly these two scores against your live catalog, plus a per field breakdown of how many products each gap affects. Run it once before the test and once after, on both groups. It reads products and collections only, it does not write to your store, and it does not touch orders or customer data.
If you would rather do the baseline by hand, our channel guides cover what each side reads and in what order: Google Merchant Center Audit: A Step by Step Guide for Shopify and How to Audit Your Shopify Catalog for Meta Ads.
Save both reports. A test without a recorded starting point is an opinion with numbers attached.
How long to run it
Two separate constraints set the length, and the longer one wins.
The statistical constraint is that you need enough orders per product for the comparison to mean anything. A store doing a handful of orders per style per month needs a longer window than one doing hundreds.
The algorithmic constraint is that delivery systems have to relearn. Enriched data changes what Meta and Google can do with an item, and neither redistributes budget instantly. Reading the test in week one tells you about the learning period, not about the change.
Four weeks is a reasonable floor for most stores, and longer if your order volume per style is thin. Avoid windows that span a major sale event or a season launch. If you cannot avoid it, plan to exclude those weeks rather than explaining them away afterwards.
Agree on the read date before you start, and do not read the result early. Interim peeks change what you decide to do next, which is how a clean test becomes a story you already believed.
What to measure
Reported ROAS is the wrong primary metric this year, for reasons that have nothing to do with your catalog. Use per product measures, so that unequal group sizes and unequal starting revenue do not distort the comparison.
| Metric | Why it earns its place |
|---|---|
| Impressions or sessions per product | Moves first. If enrichment worked, discoverability changes before revenue does. |
| Conversion rate per product | Separates "more people saw it" from "the right people saw it". |
| Revenue per product | The headline number, but only trustworthy alongside the two above. |
| Return rate per product | Apparel specific. Better garment data can raise returns if it drives volume, or lower them if it sets expectations correctly. You want to know which. |
| Net revenue after returns | The number the business actually runs on. |
Track them per group and per product, not just as group totals. Product level data is what lets you check whether a result is broad or is one style carrying everything.
Reading the result without flattering yourself
Look for direction and consistency across several metrics rather than significance on one. A real effect usually shows up as impressions up, then conversion rate steady or up, then revenue up. That chain is more convincing than a single revenue number.
Some patterns and what they usually mean:
Impressions moved, revenue did not. The data made products findable and the constraint sits elsewhere, on the product page, the price, or the images.
Revenue moved, impressions did not. Be suspicious. Check for a promotion, a newsletter, or a stockout on the control side before you accept it.
Nothing moved. This is a real result and worth taking seriously. First check whether the enrichment actually changed the scores. If completeness and attribute depth barely moved, you tested a change that did not happen.
The lift is spread evenly across every product, including ones that were already complete. Suspect a confounder. The products with the thinnest starting data should move the most, and that gradient is one of the better internal checks available to you.
Segment the result by starting score. It is the closest thing to a control on your control.
When a holdout is the wrong tool
Two cases, and being honest about them is part of running a credible test.
Very small catalogs. Below roughly 30 products, the split is too noisy to read. Use a long before and after window instead, hold the promotion calendar as steady as you can, and accept that the evidence is weaker.
Required fields that block listing. Do not hold these out. If products are being rejected by Google or Meta for a missing field, fixing half of them to preserve a clean experiment means deliberately leaving revenue on the table for a month to prove something you can already verify in the platform's own diagnostics. Fix all of them, then run the holdout on attribute depth, which is the part where the effect size is genuinely unknown.
The test is there to answer open questions. It is not there to slow down work whose value is not in question.
The point of doing this at all
Every operator running paid acquisition in 2026 has a list of things they tried and cannot tell you whether any of them worked. Creative programs, landing page rebuilds, bidding changes, new photography. The reason is not that the work was bad. It is that it all shipped at once, into a moving market, and was measured against last month.
Catalog work is one of the few interventions in this category that can be split cleanly, because the unit is the product and you own every product. That makes it testable in a way that most of the funnel is not. It is worth using that.
Sources
- Meta Ads Insights API attribution changes, January 12, 2026, and click redefinition, March 3, 2026. On why reported conversion metrics fell in 2026 independently of delivery.
- Meta Andromeda delivery system, phased rollout through 2026. On machine selected delivery and the relearning period following changes to product data.
- Coresight Research and Loop Returns benchmarks. On online apparel return rates and the case for measuring net revenue after returns in fashion.
- PRIME AI, analysis of one million fashion returns. Finding that 67 percent of returns were driven by fit and sizing issues.
Baseline your catalog before you test anything
You cannot measure a change to your product data without knowing what your product data looked like first.
SixFit Catalog Checkup scores your Shopify catalog for Google and Meta readiness in a few minutes, free. Two scores out of 100, completeness and enrichment measured separately, every missing field ranked by how many products it affects, and a shareable report you can save before you start and compare against after.
Install it here: https://apps.shopify.com/catalog-checkup
Frequently asked questions
- What is a catalog holdout test?
- It is an experiment where you split your products into two matched groups, enrich the product data on one group, leave the other unchanged, and compare performance over the same period. Because both groups experience the same season, promotions, creative, and attribution rules, the difference between them isolates the effect of the data work.
- Why split products instead of customers?
- Product data is global. Once a garment is enriched, every shopper and every channel sees the enriched version, so there is no way to serve a thin version to one audience and a rich version to another. The only thing you can split is the catalog itself.
- How many products do I need for a holdout test?
- There is no hard threshold, but below roughly 30 products the groups become too small and too easily distorted by a single bestseller or stockout. Below that, a long before and after comparison with a stable promotion calendar is more practical, with the understanding that it is weaker evidence.
- How long should a catalog holdout test run?
- At least four weeks for most stores, and longer if you sell only a few units per style per month. Delivery systems need time to redistribute budget after product data changes, so early readings mostly reflect the learning period rather than the result.
- What should I measure instead of ROAS?
- Per product impressions or sessions, conversion rate, revenue, return rate, and net revenue after returns. Reported ROAS is a poor primary metric in 2026 because attribution windows changed twice during the year, which moved the number without moving the business.
- A promotion hit one of my groups mid test. Is the test ruined?
- Not necessarily. Exclude the affected products and the affected days from the analysis, and report that you did. A test with a documented exclusion is more credible than one that quietly includes a sale.
- Do I need to hold back products that are being rejected by Google or Meta?
- No. If an item is missing a required field and is not eligible to show, fix it now. Hold out on attribute depth instead, which is where the size of the effect is actually uncertain.
