How Data Purl validates its data
Data Purl checks its data at every step. Over the last two years, 95.9% of scheduled retailer-weeks were captured without a gap, and every gap is filled from the last good week and flagged. Every product passes outlier screens before it is published, and 91% of Hardlines & Consumables products have a high-confidence category.
Figures as of 26 September 2026, read from the production pipeline and the classification audits.
1. Collection
Retailer and brand websites are collected several times a week and published weekly. Over the last 104 weeks, across 365 retailer feeds and 34,933 scheduled retailer-weeks:
- Captured without a gap: 95.9% of retailer-weeks.
- Gap-filled: 4.1%. A week counts as a gap when a retailer returns no data or fewer than a quarter of its usual number of SKUs. The gap is filled with that retailer's last good week and flagged (IS_FILLED), so it can be excluded from any analysis.
- Capture frequency: each product is captured on about 4 days a week on average; 54% of retailer-weeks are captured on five or more days and 26% every day.
- Data load jobs retry automatically and alert the operations team when they fail.
2. Quality screens before publication
Every product-week passes automated screens before it reaches clients. A product-week is held back when:
- It was seen on fewer than four days in the week.
- Its discount is above 80%, or its selling price is above its full price.
- Its price is an outlier for its category and market (more than three interquartile ranges from typical prices, on a log scale, where the category has at least 20 products).
- It has no category or no price.
Thin retailer and category combinations are also held back, so a handful of listings cannot move a category figure. Automated tests check that product keys are unique, that weekly records are not duplicated, and that brand ownership dates never overlap.
3. Classification
Every product is placed in one standardized tree of about 4,000 categories. 91% of Hardlines & Consumables products in scope have a high-confidence category: one that has been individually reviewed and confirmed, or assigned by audited classification rules. The rest are classified by a model, used only where its calibrated confidence is at least 0.75. Fashion products are classified by an analyst-maintained rule model.
Brands are resolved to their owners, public or private. Listed owners are mapped to tickers with dated start and end dates, including licensing arrangements, so each observation is attributed to the owner at the time it was made.
4. Freshness
- Weeks run Monday to Sunday.
- The weekly build starts early on Wednesday, and the new week is published on Wednesday evening US Eastern time, before Thursday's US open.
- A price change or sell-out on a retailer's site is usually captured within a few days, and reaches clients between about three days and two weeks after it happens, depending on when in the week it happens.
- When history is restated, the publish timestamp on each row changes, so clients can see which records moved.
5. Backtests
Data Purl tests whether its shelf data tracks the numbers companies report. For each covered company, a model turns year-on-year changes in Data Purl measures (average listed price, markdown breadth, active SKUs and core mix) into a signal for the company's reported gross margin, and each model is re-scored every quarter against the reported result. Four examples:
Brands are attributed to the company that owned them at the time of each observation, so acquisitions and disposals do not leak future information into the history. Models for other companies, and current-quarter signals, are available to clients. Not investment advice.
What the data does not measure
Data Purl observes what retailers list online: prices, discounts, availability and assortment. It does not observe transactions, unit sales, checkout-only or loyalty discounts, or physical stores. See the methodology for full definitions.
The evidence on this page (history collected the same way every week, flagged gaps, audited classification and backtests) is what a one-off AI scrape cannot produce. See Data Purl vs AI web scraping.
Need more detail for due diligence?
We answer vendor due diligence questionnaires and can walk your team through collection, classification and quality controls.