ruia

1,754 3 8 Apache-2.0

0.8.5 (6 Sep 2022) Oct 17 2018 911 (month)

Ruia is an async web scraping micro-framework, written with asyncio and aiohttp, aims to make crawling url as convenient as possible.

Ruia is inspired by scrapy however instead of Twisted it's based entirely on asyncio and aiohttp.

It also supports various features like cookies, headers, and proxy, which makes it very useful in dealing with complex web scraping tasks.

Example Use

#!/usr/bin/env python
"""
 Target: https://news.ycombinator.com/
 pip install aiofiles
"""
import aiofiles

from ruia import AttrField, Item, Spider, TextField


class HackerNewsItem(Item):
    target_item = TextField(css_select="tr.athing")
    title = TextField(css_select="a.storylink")
    url = AttrField(css_select="a.storylink", attr="href")

    async def clean_title(self, value):
        return value.strip()


class HackerNewsSpider(Spider):
    start_urls = [
        "https://news.ycombinator.com/news?p=1",
        "https://news.ycombinator.com/news?p=2",
    ]
    concurrency = 10
    # aiohttp_kwargs = {"proxy": "http://0.0.0.0:1087"}

    async def parse(self, response):
        async for item in HackerNewsItem.get_items(html=await response.text()):
            yield item

    async def process_item(self, item: HackerNewsItem):
        async with aiofiles.open("./hacker_news.txt", "a") as f:
            self.logger.info(item)
            await f.write(str(item.title) + "\n")


if __name__ == "__main__":
    HackerNewsSpider.start(middleware=None)

Alternatives / Similar

scrapy

54,211 2.12.0 (9 months ago) Jul 26 2019 compare

scrapyd

2,980 1.5.0 (10 months ago) Sep 04 2013 compare

autoscraper

6,638 1.1.14 (3 years ago) Jul 26 2019 compare

gracy

247 1.34.0 (8 months ago) Feb 05 2023 compare

scrapydweb

3,218 1.6.0 (6 months ago) Sep 30 2018 compare

gerapy

3,365 0.9.13 (2 years ago) Jul 04 2017 compare

photon

11,149 1.1.9 (6 years ago) Aug 24 2018 compare

dude

428 0.1.3 (2 years ago) Feb 20 2022 compare

Other Languages

colly

23,747 v2.1.0 (5 years ago) May 14 2018 compare

pholcus

7,580 v1.3.4 (5 years ago) Feb 15 2020 compare

geziyor

2,667 2025-02-18 (6 months ago) Jun 06 2019 compare

dataflowkit

676 2025-02-16 (6 months ago) Feb 09 2017 compare

rvest

1,498 1.0.4 (3 years ago) Nov 22 2014 compare

gocrawl

2,039 (4 years ago) Nov 20 2016 compare

ferret

5,716 v0.18.0 (2 years ago) Aug 06 2019 compare

node-crawler

6,733 2.0.2 (1 year, 1 month ago) Sep 10 2012 compare

panther

2,977 v2.2.0 (6 months ago) Jul 17 2018 compare

spidr

813 0.7.2 (6 months ago) Jul 25 2009 compare

wombat

1,316 3.0.0 (3 years ago) Dec 27 2011 compare

ralger

156 2.2.4 (4 years ago) Dec 22 2019 compare

roach

1,384 v3.2.0 (1 year, 4 months ago) Dec 27 2021 compare

ayakashi

213 1.0.0-beta8.4 (2 years ago) Apr 18 2019 compare

phpscraper

554 3.0.0 (1 year, 4 months ago) May 04 2020 compare

php-spider

1,335 v0.7.2 (1 year, 8 months ago) Mar 16 2013 compare

crwlr-crawler

356 v3.2.3 (6 months ago) Apr 18 2022 compare