Python site crawler (Scrapy)
A crawler for large sites, built on the Scrapy framework. Can't run client-side JavaScript.
my_actor/main.py
my_actor/items.py
my_actor/middlewares.py
my_actor/pipelines.py
my_actor/settings.py
my_actor/spiders/title.py
my_actor/__main__.py
1"""Module defines the main entry point for the Apify Actor.2
3Module defines the main coroutine for the Apify Scrapy Actor, executed from the __main__.py file. The coroutine4processes the Actor's input and executes the Scrapy spider. Additionally, it updates Scrapy project settings by5applying Apify-related settings. Which includes adding a custom scheduler, retry middleware, and an item pipeline6for pushing data to the Apify dataset.7
8Customization:9--------------10
11Feel free to customize this file to add specific functionality to the Actor, such as incorporating your own Scrapy12components like spiders and handling Actor input. However, make sure you have a clear understanding of your13modifications. For instance, removing `apply_apify_settings` break the integration between Scrapy and Apify.14
15Documentation:16--------------17
18For an in-depth description of the Apify-Scrapy integration process, our Scrapy components, known limitations and19other stuff, please refer to the following documentation page: https://docs.apify.com/cli/docs/integrating-scrapy.20"""21
22from __future__ import annotations23
24from apify import Actor25from apify.scrapy import apply_apify_settings26from scrapy.crawler import AsyncCrawlerRunner27
28# Import your Scrapy spider here.29from .spiders import TitleSpider as Spider30
31
32async def main() -> None:33 """Apify Actor main coroutine for executing the Scrapy spider."""34 async with Actor:35 # Retrieve and process Actor input.36 actor_input = await Actor.get_input() or {}37 start_urls = [url['url'] for url in actor_input.get('startUrls', [])]38 allowed_domains = actor_input.get('allowedDomains')39 proxy_config = actor_input.get('proxyConfiguration')40
41 # Apply Apify settings, which will override the Scrapy project settings.42 settings = apply_apify_settings(proxy_config=proxy_config)43
44 # Create AsyncCrawlerRunner and execute the Scrapy spider.45 crawler_runner = AsyncCrawlerRunner(settings)46 await crawler_runner.crawl(47 Spider,48 start_urls=start_urls,49 allowed_domains=allowed_domains,50 )A template example built with Scrapy to scrape page titles from URLs defined in the input parameter. It shows how to use Apify SDK for Python and Scrapy pipelines to save results.
- Apify SDK for Python - a toolkit for building Apify Actors and scrapers in Python
- Input schema - define and easily validate a schema for your Actor's input
- Request queue - queues into which you can put the URLs you want to scrape
- Dataset - store structured data where each object stored has the same attributes
- Scrapy - a fast high-level web scraping framework
This code is a Python script that uses Scrapy to scrape web pages and extract data from them. Here's a brief overview of how it works:
- The script reads the input data from the Actor instance, which is expected to contain a
start_urlskey with a list of URLs to scrape. - The script then creates a Scrapy spider that will scrape the URLs. This Spider (class
TitleSpider) is storing URLs and titles. - Scrapy pipeline is used to save the results to the default dataset associated with the Actor run using the
push_datamethod of the Actor instance. - The script catches any exceptions that occur during the web scraping process and logs an error message using the
Actor.log.exceptionmethod.
- Web scraping with Scrapy
- Python tutorials in Academy
- Alternatives to Scrapy for web scraping in 2023
- Beautiful Soup vs. Scrapy for web scraping
- Integration with Zapier , Make, Google Drive, and others
- Video guide on getting scraped data using Apify API
- A short guide on how to build web scrapers using code templates:
Python site crawler (Crawlee + BeautifulSoup)
A crawler that follows links and gets data from static pages, with Crawlee handling retries and request queues. Uses BeautifulSoup, Python's most popular HTML parser. Can't run client-side JavaScript.
Empty Python Actor
An Actor with the Apify SDK set up, so you can build any tool you need.
Python one-page scraper (BeautifulSoup)
A scraper that gets data from one web page with BeautifulSoup. The simplest way to start scraping.
Empty Python Actor (uv)
An Actor with the Apify SDK set up and dependencies managed by the uv package manager, so you can build any tool you need.
Python site crawler (BeautifulSoup)
A lightweight crawler that follows links and gets data from static pages with BeautifulSoup. Good for blogs, news, or product listings, but it can't run client-side JavaScript.
Python browser crawler (Playwright)
A crawler that uses a real browser through Playwright, so it gets data HTTP crawlers miss. Good for social feeds, dashboards, or single-page apps.