How to scrape product data from a website into Excel
You need a supplier's catalog or a competitor's prices in a spreadsheet, and copying them by hand would take days. There are several ways to get there, from simply asking for a file to running a scraper on a schedule. Here is how to pick one, what to collect and what usually gets in the way.
First, check whether you need to scrape at all
The cheapest data is data someone has already exported. Before building anything, try these:
- Ask the supplier. Many suppliers have a price list, a product feed or a dealer export in CSV or Excel. It is often more complete than the website and already has SKUs and wholesale prices.
- Look for an official API or feed. Some platforms and larger suppliers give partners product data through an API. It is the most stable source, because it is built for machines.
- Check the site itself. Some stores have a price list download or a catalog PDF tucked away in the footer or the dealer section.
If none of these exist or the data is incomplete, you are choosing between copying by hand, a ready-made tool and a custom scraper.
Manual copying, tools or a custom scraper
| Option | Works well when | Limits |
|---|---|---|
| Manual copying | A few dozen products, one time | Slow, typos creep in, no updates |
| Excel Power Query, From Web | Data sits in plain HTML tables on a page | Struggles with catalogs that load with JavaScript or span many pages |
| Browser extensions and no-code scraping tools | A simple site and someone willing to set it up | A monthly fee per the provider's plan, setup and fixes are on you |
| Custom script | Hundreds or thousands of products, several sites, regular updates | Someone has to write it and keep it working |
Excel's Power Query has a From Web option on the Data tab that can pull a table from a web page. It is worth a try on simple pages. For a real product catalog with categories, several pages of results and product variants, a script that walks through the pages and writes a clean spreadsheet is usually the reliable option.
Which fields to collect
Decide on the columns before anything is collected. Sketch one row of the final spreadsheet by hand: it settles more questions than a long description.
- Product name and brand
- SKU, manufacturer part number or barcode, if the site shows them
- Price, sale price and currency
- Stock status or quantity
- Category and subcategory
- Product page URL, so any row can be checked by hand
- Image links, or images saved into folders
- Specifications such as size, weight, material or compatibility, each in its own column
- Variants: colors and sizes as separate rows or listed in one cell
- Date of collection
The SKU or the page URL is what lets you match the same product between runs and between sites. Without a stable identifier, comparing prices over time turns into guesswork.
Keeping prices and stock up to date
A one-time export starts going out of date the day it is made. If you use the data for pricing or to fill your own store, you need regular updates. A typical setup works like this:
- The scraper runs on a schedule, daily or weekly, depending on how fast prices move in your niche.
- Each run is compared with the previous one by SKU or URL.
- New products, removed products, price changes and stock changes are flagged in separate columns or on a separate tab.
- The fresh file goes to your email, a shared folder or a Google Sheet that your other spreadsheets read from.
Run it only as often as you actually act on the data. Hourly runs on a site whose prices change once a month put load on that site and give you nothing new.
What usually gets in the way
- Content that appears as you scroll or click. Many modern stores load products with JavaScript, so the plain page is nearly empty. The scraper has to behave more like a browser, which takes longer to build and to run.
- Variants hidden behind selectors. Prices and stock for each size or color often appear only after a choice is made on the page.
- Prices visible only after login. Wholesale prices behind a login are not public data. The right route there is asking the supplier for an export.
- Protection against automated visitors. Some sites block bots on purpose. That is a signal to look for an official export or API, not something to get around.
- Messy source data. The same spec written three ways, mixed units, prices with and without tax. Bringing this into one format is often half the work.
- Redesigns. When the site changes its layout, a scraper may stop finding the data. A good script reports this instead of quietly sending you an empty file.
Legal and ethical limits
This is general practical guidance, not legal advice. Rules differ between countries and between websites, so if the data matters to your business, check with a lawyer.
- Collect only public data that anyone can see without logging in.
- Do not collect personal data, such as names, emails or phone numbers of private individuals.
- Read the site's terms of use and its robots.txt file, and respect what they say about automated access.
- Keep the load low: pauses between requests, sensible intervals between runs, no flood of parallel requests.
- Be careful with reuse. Comparing prices is one thing. Copying product descriptions and photos into your own store is another: those are usually protected and need the owner's permission.
What it costs
If you would rather not build this yourself, I collect public product data from websites into Excel or Google Sheets with filters, duplicates removed and prices in one format. The price grows with the number of sites and products, with how the site loads its pages, and with extras such as images or a schedule.
| Task | Price |
|---|---|
| One-time export of a single website's catalog | from $40 |
| Data from several websites in one spreadsheet | from $65 |
| Scheduled scraping with a change report | from $80 |
A one-time export of one site usually takes 1–3 days. Before you pay, I check the site and tell you if the data is not public or the site blocks automated access, and you see a sample of the columns before the full run. Send the links and the list of columns you need to get an exact quote.
I can do this for you
Web scraping service: website data to Excel. A website's product catalog, prices and specifications in an organized Excel file.
from $40Timeline: 1–3 days
FAQ
Is it legal to scrape product data from a website?
It depends on your country, the site's terms and what exactly you collect. Collecting public product data gently, without personal data and within the site's terms, is the safest ground. This is not legal advice, so ask a lawyer if the data is important to your business.
Can I scrape Amazon or other large marketplaces?
Large marketplaces usually restrict automated collection in their terms and actively block it. For your own listings, use the seller reports and tools the marketplace provides. For market research, look at official APIs or licensed data services first.
Will Excel cope with a large catalog?
Tens of thousands of rows are fine in Excel. Google Sheets slows down sooner on very large files, so for big catalogs an Excel file or a CSV is the more practical delivery format.
How often should I update competitor prices?
Match it to how often you change your own prices. Weekly is enough in many niches. Daily makes sense only if competitors really change prices daily and you are ready to react.