Curated, researched data for your project, not a raw scrape
This is not a raw scrape. Every Data Purl extract is curated and researched: already categorized into one standardized tree, with brands resolved to their owners across public and private companies, listed owners mapped to tickers, products tagged for markdown, availability and pack size, and the output audited before release. Choose the sectors, markets, dates, grain and fields your project needs across 400+ retailers, and get the exact row count, the price and a free sample of your own cut before anything is pulled.
Already done before it reaches you
Raw scraped data leaves the hardest work to you: matching categories across retailers, resolving brand names, linking brands to companies and checking the result. In a Data Purl extract that work is done, researched and audited.
Categories
Raw scrape: Each retailer's own site menu and product names, all different.
Every product placed in one standardized tree of about 4,000 categories, the same at every retailer.
Brands and owners
Raw scrape: Brand names as typed on each site, with spelling variants and sub-brands scattered.
Brands resolved to one master brand and its owner, with licensed lines kept separate.
Public and private
Raw scrape: No link to the companies behind the brands.
Brands resolved to owners, public and private; listed owners mapped point-in-time to tickers, with acquisition and licensing dates.
Research
Raw scrape: None: the data is only as good as the page it came from.
Brand ownership, licensing deals and acquisitions researched and dated by our research team; the category tree built and maintained by analysts.
Tags
Raw scrape: Prices and text as displayed.
Markdown, in-stock, first-seen and last-seen flags, and standardized pack size and unit price.
Quality
Raw scrape: Whatever the collector captured, including misfiled and duplicate listings.
Classification and brand mapping audited before release; outliers and duplicates handled.
History
Raw scrape: Starts the day collection starts.
Continuous history back to 2013.
| Raw scraped data | Data Purl extract | |
|---|---|---|
| Categories | Each retailer's own site menu and product names, all different. | Every product placed in one standardized tree of about 4,000 categories, the same at every retailer. |
| Brands and owners | Brand names as typed on each site, with spelling variants and sub-brands scattered. | Brands resolved to one master brand and its owner, with licensed lines kept separate. |
| Public and private | No link to the companies behind the brands. | Brands resolved to owners, public and private; listed owners mapped point-in-time to tickers, with acquisition and licensing dates. |
| Research | None: the data is only as good as the page it came from. | Brand ownership, licensing deals and acquisitions researched and dated by our research team; the category tree built and maintained by analysts. |
| Tags | Prices and text as displayed. | Markdown, in-stock, first-seen and last-seen flags, and standardized pack size and unit price. |
| Quality | Whatever the collector captured, including misfiled and duplicate listings. | Classification and brand mapping audited before release; outliers and duplicates handled. |
| History | Starts the day collection starts. | Continuous history back to 2013. |
How it works
- 1Scope your cut. Choose sectors, markets, dates, grain and fields, and name any brands or retailers. The data is already categorized, mapped and audited, so you only choose what to include.
- 2Get the count, price and a free sample. Within one business day we send the exact row count, the price and a free sample of your own cut.
- 3Buy and receive the data. Pay for the extract as a one-off project purchase or with download credits, and receive the files.
Scope your cut
Which retailers are available
Extracts can draw on any of the retailers we cover. See the leading retailers by market segment, and name the retailers you need in the scope builder: we confirm coverage for each with your quote.
What the data looks like
A preview of core fields. The full data dictionary and a sample file of your own cut come with your quote.
| Field | Group | Description |
|---|---|---|
| WEEK_START_DATE | Time | Start of the observation week. |
| RETAILER_NAME | Identity | Retailer or brand website the product was observed on. |
| COUNTRY | Identity | Market of the retailer website. |
| MASTER_BRAND | Identity | Consolidated brand name, resolved across retailers. |
| BRAND_TICKER | Identity | Point-in-time listed identifier, where the brand's owner is a public company. |
| CATEGORY_GROUP / CATEGORY | Taxonomy | Position in the standardized category tree. |
| PRODUCT_KEY / SKU_KEY | Keys | Stable identifiers for the product and its size and colour variants. |
| AVGPRICE_TOTAL_INSTOCK | Price | Average selling price of in-stock items in the week. |
| AVGPRICE_LIST_INSTOCK | Price | Average list (pre-markdown) price of in-stock items. |
| AVGPRICE_FIRST_TOTAL | Price | Price the product first appeared at. |
| PCT_BREADTH_INSTOCK | Promotion | Share of in-stock items on markdown. |
| PCT_DEPTH_INSTOCK | Promotion | Average discount on items that are marked down. |
| AVGDAILYSKUS_TOTAL_INSTOCK | Assortment | Average daily count of in-stock SKUs. |
| REVIEWS_PER_DAY | Demand proxy | Average reviews accumulated per day, a proxy for sales velocity. |
| NET_STANDARDIZED_SIZE / _UOM | Units | Standardized pack size and unit of measure.from October 2026 |
| PRICE_PER_UNIT | Units | Price for one standard unit (for example per ounce or per litre).from October 2026 |
What drives the price
- Grain: weekly aggregates cost less than product-level, and product-level less than SKU-level rows.
- Breadth: the number of retailers, brands, categories and markets in the cut.
- History: the length of the date range requested.
- Fields: only the field groups you need are priced.
Need every record, refreshed every week? The full dataset is delivered to Snowflake or S3 and arranged with a founder.
Common questions
Can I buy a custom cut of Data Purl data without a subscription?
Yes. Project extracts are one-off purchases scoped to the brands, retailers, categories, markets and dates you need. You see the row count and price, and a free sample, before anything is pulled.
What does a custom extract include?
Weekly observations at the grain you choose (weekly aggregates, product level or SKU level) with the field groups you select: prices, promotions, availability, assortment, demand proxies, pack size and unit price, and owner and ticker mapping.
How is a Data Purl extract different from buying scraped data?
Scraped data gives you listings as each retailer displays them. A Data Purl extract arrives already categorized into one standardized tree, with brands resolved to their owners, public or private, listed owners mapped to tickers, products tagged for markdown, availability and pack size, and the classification and mapping audited before release. It is curated and researched for your use case, not a raw scrape you have to clean.
How is a custom extract priced?
By scope: grain, breadth (retailers, brands, categories and markets), length of history and field groups. We quote the exact price with the row count before you commit.
How is this different from the full dataset?
An extract is a one-off cut licensed for a single project. The full dataset is the complete, continuously refreshed history delivered to your Snowflake account or Amazon S3, arranged with a founder.
Can I use an extract in client work?
Licensing terms for your use case, including client deliverables, are confirmed with your quote. Redistribution of the underlying data is not permitted.