A universal web scraper for e-commerce sites with automatic platform detection and intelligent data extraction.
- Auto E-commerce Scraper
- Features
- Architecture
- Installation
- Usage
- How It Works
- API Reference
- Examples
- Tested Sites
- Configuration
- Links
- Disclaimer
- Troubleshooting
- Intelligent Container Detection - Finds repeating product cards without configuration
- HasData API Integration - Optional API support for JavaScript rendering and advanced features
- AI Extraction - Extract structured product data using AI (via HasData)
- Multiple Export Formats - Export to CSV, JSON, and Excel
- User-Friendly Interface - Built with Streamlit for easy interaction
One entry point, one package, no framework beyond Streamlit.
.
├── app.py # Main application entry point
├── scraper/
│ ├── core.py # AutoScraper class and main logic
│ ├── platforms.py # Platform-specific scrapers (WooCommerce, Shopify)
│ └── utils.py # Helper functions for container detection
└── ui/
├── layout.py # Sidebar and results rendering
└── export.py # Data export functionality
The scraper package works standalone, app.py only adds the UI.
Clone the repository and install the dependencies.
# Clone the repository
git clone https://github.com/hasdata/auto-ecommerce-scraper.git
cd auto-ecommerce-scraper
# Install dependencies
pip install -r requirements.txtNo build step follows, the app runs from source.
The list in requirements.txt is the whole footprint.
streamlit>=1.30.0
pandas>=2.0.0
beautifulsoup4>=4.12.0
requests>=2.31.0
openpyxl>=3.1.0Python 3.11 or newer.
The tool runs as a local Streamlit app.
Start the app and open the printed local URL.
streamlit run app.pyThen:
- Enter a product listing page URL
- Click "Start Scraping"
- Review extracted data
- Export to your preferred format
For JavaScript-heavy sites or AI extraction:
- Check "Use HasData API"
- Enter your API key
- Configure scraping options:
- JS Rendering
- Proxy settings
- AI product extraction
- Screenshot capture
- Email/link extraction
Each stage hands over to a fallback when it comes up empty.
The scraper identifies the platform from the page source. WooCommerce shows itself through woocommerce and product_type_ class prefixes, Shopify through its asset host and product-card markup, and everything else goes down the generic route.
For Shopify the scraper asks the storefront's own /products.json before touching the HTML, and for any platform it tries schema.org Product microdata before falling back to container detection.
Scores potential product containers based on:
- Presence of images and links
- Text content length
- HTML structure complexity
- CSS class keywords
- Price indicators
Extracts:
- Product titles
- Prices and original prices
- Product URLs
- Images
- Stock status
- Categories and tags
- SKUs
- Ratings and reviews
- Groups similar fields
- Removes low-frequency selectors
- Creates human-readable field names
- Handles links and images properly
The package is importable outside the UI.
The class carries the whole pipeline.
from scraper.core import AutoScraper
# Initialize scraper
scraper = AutoScraper(
url="https://example.com/products",
api_key="your_hasdata_key", # Optional
scrape_config={
"jsRendering": True,
"proxyType": "datacenter",
"blockAds": True
}
)
# Scrape data
result = scraper.scrape(
container_index=0, # Select container group
force_generic=False # Force generic method
)
# Access results
print(result['platform']) # 'woocommerce', 'shopify', or 'generic'
print(result['data']) # List of extracted itemsThe result dict names the platform it detected.
Each route is callable on its own.
from scraper.platforms import (
scrape_woocommerce,
scrape_shopify,
scrape_shopify_products_json,
scrape_microdata,
)
# WooCommerce card parsing
result = scrape_woocommerce(scraper)
# Shopify: the storefront's own /products.json first, HTML cards second
result = scrape_shopify_products_json(scraper)
result = result or scrape_shopify(scraper)
# schema.org Product microdata, works on any platform that marks its cards
result = scrape_microdata(scraper)On a Shopify store the JSON route replaces the whole card parse. One paginated request returns title, handle, price, SKU, vendor and image per product, verified on two live stores at 50 of 50 products each. Microdata runs before the generic container fallback, so a marked-up theme parses with zero site-specific selectors.
The two snippets below match the two ways people run the tool.
The minimal run needs only a URL.
scraper = AutoScraper("https://scrapeme.live/shop/")
result = scraper.scrape()
# result will contain:
# {
# 'platform': 'woocommerce',
# 'count': 48,
# 'data': [
# {
# 'title': 'Product Name',
# 'price': '$19.99',
# 'url': 'https://...',
# 'image': 'https://...',
# 'stock_status': 'In Stock'
# },
# ...
# ]
# }The platform parser picks the fields, no schema needed.
AI rules trade credits for structure on messy layouts.
from scraper.core import UNIVERSAL_PRODUCT_RULES
scraper = AutoScraper(
url="https://example.com/products",
api_key="your_key",
scrape_config={
"jsRendering": True,
"aiExtractRules": UNIVERSAL_PRODUCT_RULES
}
)
result = scraper.scrape()
ai_data = scraper.get_hasdata_extras()['aiResponse']The AI rules return one structured object per detected product.
- ✅ WooCommerce stores
- ✅ Shopify stores
- ✅ Custom e-commerce platforms
- ✅ Book stores (books.toscrape.com)
- ✅ Fashion marketplaces (Vinted)
- ✅ And many more...
Everything below passes through to the HasData request.
The table lists what the UI exposes.
| Option | Description | Default |
|---|---|---|
jsRendering |
Enable JavaScript rendering | true |
screenshot |
Capture page screenshot | false |
extractEmails |
Extract email addresses | false |
extractLinks |
Extract all links | false |
blockResources |
Block images/fonts/media | true |
blockAds |
Block advertisements | true |
proxyType |
Proxy type (datacenter/residential/mobile) | datacenter |
proxyCountry |
Proxy country code | US |
Unset options keep the API defaults.
Define custom extraction schemas:
custom_rules = {
"products": {
"type": "list",
"description": "all products",
"output": {
"name": {"type": "string"},
"price": {"type": "string"},
"rating": {"type": "number"}
}
}
}Any JSON schema in this shape works as a rule set.
This tool is for educational purposes only. Learn more about the legality of web scraping.
Three failure shapes account for nearly every report.
"Could not find repeating structures"
- Try checking "Use generic method"
- The page needs several similar items for detection to lock on
- Try with HasData API for JS-rendered content
"API Error"
- Verify your HasData API key
- Check your API quota
- Check the URL opens in a normal browser first
Empty or incomplete data
- Select a different container group
- Enable JS rendering
- Check if the site requires authentication
Made with ❤️ for the web scraping community

