Should we build our own retailer scraper or buy eCommerce pricing data?
Building a scraper gets you raw pages; the cost is in everything after. To use retail pricing data for investment or strategy work you also need to keep collection running through site changes, place every retailer's products in one category tree, resolve brands to their owners and map listed owners to tickers, tag markdowns, stock and pack sizes, audit the output and carry history back years. Build when you need a narrow, short-lived cut of a few sites; buy when you need many retailers, entity mapping or history.
Last reviewed
What building involves
- Collection upkeep. Retailer sites change layouts and product pages often, and each change can silently break a scraper or the history behind it.
- Categorization. Every retailer uses its own menus and product names, so comparing a category across retailers means building and maintaining your own category tree.
- Brands and owners. Brand names appear with variants and sub-brands, and linking them to owners, public and private, takes ongoing research into ownership, licensing and acquisitions.
- Tickers. Mapping listed owners to tickers point-in-time is needed before any backtest.
- Tagging. Markdown, in-stock, first-seen and last-seen flags and standardized pack sizes all have to be derived.
- Quality. Misfiled and duplicate listings have to be found and handled before results can be trusted.
- History. A new scraper starts on the day it is switched on, so there is no history to compare against.
When building makes sense
- A one-off question about a handful of websites, where categories and owners do not matter.
- Data a provider does not cover, such as a niche retailer or a single brand's site.
When buying makes sense
- Many retailers or markets, or comparisons across retailers.
- Company-level views that need brands resolved to owners and tickers.
- Trend work that needs years of consistent history.
- A short project where engineering time is the constraint.
A middle route
Data Purl's custom extracts give you only the cut a project needs (chosen sectors, markets, dates, grain and fields) already categorized, resolved to owners, tagged and audited, with a free sample before you commit.
Metrics used
- Alternative data (shelf-side): Non-traditional data used in investment research; shelf-side data measures what is listed, at what price and availability, rather than what was bought.
- Brand-to-ticker mapping: Linking each consumer brand to the company that owns it and to that company's listed security, with ownership changes dated.
- Assortment count: The number of distinct products a brand or retailer has live for sale at a point in time.
How each measure is built: methodology.
See the data for your coverage
Start free in the analytics portal, or talk to the founders about the full dataset delivered to Snowflake or S3.