ruiavsralger

python html-parser framework http-client

Apache-2.0 8 3 1,754

911 (month) Oct 17 2018 0.8.5(2 years ago)

156 1 3 MIT

Dec 22 2019 349 (month) 2.2.4(4 years ago)

Ruia is an async web scraping micro-framework, written with asyncio and aiohttp, aims to make crawling url as convenient as possible.

Ruia is inspired by scrapy however instead of Twisted it's based entirely on asyncio and aiohttp.

It also supports various features like cookies, headers, and proxy, which makes it very useful in dealing with complex web scraping tasks.

ralger is a small web scraping framework for R based on rvest and xml2.

It's goal to simplify basic web scraping and it provides a convenient and easy to use API.

It offers functions for retrieving pages, parsing HTML using CSS selectors, automatic table parsing and auto link, title, image and paragraph extraction.

Example Use

#!/usr/bin/env python
"""
 Target: https://news.ycombinator.com/
 pip install aiofiles
"""
import aiofiles

from ruia import AttrField, Item, Spider, TextField


class HackerNewsItem(Item):
    target_item = TextField(css_select="tr.athing")
    title = TextField(css_select="a.storylink")
    url = AttrField(css_select="a.storylink", attr="href")

    async def clean_title(self, value):
        return value.strip()


class HackerNewsSpider(Spider):
    start_urls = [
        "https://news.ycombinator.com/news?p=1",
        "https://news.ycombinator.com/news?p=2",
    ]
    concurrency = 10
    # aiohttp_kwargs = {"proxy": "http://0.0.0.0:1087"}

    async def parse(self, response):
        async for item in HackerNewsItem.get_items(html=await response.text()):
            yield item

    async def process_item(self, item: HackerNewsItem):
        async with aiofiles.open("./hacker_news.txt", "a") as f:
            self.logger.info(item)
            await f.write(str(item.title) + "\n")


if __name__ == "__main__":
    HackerNewsSpider.start(middleware=None)

library("ralger")

url <- "http://www.shanghairanking.com/rankings/arwu/2021"

# retrieve HTML and select elements using CSS selectors:
best_uni <- scrap(link = url, node = "a span", clean = TRUE)
head(best_uni, 5)
#>  [1] "Harvard University"
#>  [2] "Stanford University"
#>  [3] "University of Cambridge"
#>  [4] "Massachusetts Institute of Technology (MIT)"
#>  [5] "University of California, Berkeley"

# ralger can also parse HTML attributes
attributes <- attribute_scrap(
  link = "https://ropensci.org/",
  node = "a", # the a tag
  attr = "class" # getting the class attribute
)

head(attributes, 10) # NA values are a tags without a class attribute
#>  [1] "navbar-brand logo" "nav-link"          NA
#>  [4] NA                  NA                  "nav-link"
#>  [7] NA                  "nav-link"          NA
#> [10] NA
#

# ralger can automatically scrape tables:
data <- table_scrap(link ="https://www.boxofficemojo.com/chart/top_lifetime_gross/?area=XWW")

head(data)
#> # A tibble: 6 × 4
#>    Rank Title                                      `Lifetime Gross`  Year
#>   <int> <chr>                                      <chr>            <int>
#> 1     1 Avatar                                     $2,847,397,339    2009
#> 2     2 Avengers: Endgame                          $2,797,501,328    2019
#> 3     3 Titanic                                    $2,201,647,264    1997
#> 4     4 Star Wars: Episode VII - The Force Awakens $2,069,521,700    2015
#> 5     5 Avengers: Infinity War                     $2,048,359,754    2018
#> 6     6 Spider-Man: No Way Home                    $1,901,216,740    2021

Alternatives / Similar

colly

23,747 compare

pholcus

7,580 compare

geziyor

2,667 compare

dataflowkit

676 compare

scrapy

54,211 compare

rvest

1,498 compare

gocrawl

2,039 compare

ferret

5,716 compare

scrapyd

2,980 compare

node-crawler

6,733 compare

panther

2,977 compare

autoscraper

6,638 compare

gracy

247 compare

spidr

813 compare

scrapydweb

3,218 compare

gerapy

3,365 compare

wombat

1,316 compare

photon

11,149 compare

ralger

156 compare

roach

1,384 compare

dude

428 compare

ayakashi

213 compare

phpscraper

554 compare

php-spider

1,335 compare

crwlr-crawler

356 compare

rvest

1,498 compare

requests

52,519 compare

node-fetch

8,825 compare

colly

23,747 compare

axios

106,345 compare

aiohttp

15,425 compare

sax-js

1,101 compare

parse5

3,698 compare

htmlparser2

4,529 compare

httpx

13,703 compare

beautifulsoup

- compare

lxml

2,737 compare

got

14,454 compare

pholcus

7,580 compare

xmltodict

5,577 compare

curl-impersonate

4,221 compare

superagent

16,610 compare

needle

1,637 compare

cheerio

28,873 compare

geziyor

2,667 compare

html5lib

1,153 compare

cssselect

293 compare

dataflowkit

676 compare

faraday

5,785 compare

nokogiri

6,173 compare

httpclient

703 compare

feedparser

2,048 compare

excon

1,163 compare

httparty

5,837 compare

pyquery

2,312 compare

pycurl

1,094 compare

parsel

1,187 compare

scrapy

54,211 compare

typhoeus

4,084 compare

requests-html

13,780 compare

httr

988 compare

xml2

221 compare

curl-cffi

1,751 compare

guzzle

23,055 compare

em-http-request

1,217 compare

selectolax

1,186 compare

symfony-http

1,976 compare

html5-php

1,638 compare

untangle

619 compare

domcrawler

3,985 compare

wreck

381 compare

http-2

898 compare

treq

590 compare

goquery

14,273 compare

xpath

699 compare

resty

10,341 compare

cascadia

717 compare

soup

2,191 compare

req

4,374 compare

gocrawl

2,039 compare

htmlquery

760 compare

ferret

5,716 compare

nestful

505 compare

html5-parser

683 compare

crul

107 compare

scrapyd

2,980 compare

chompjs

202 compare

requests

3,576 compare

node-crawler

6,733 compare

hrequests

780 compare

panther

2,977 compare

gazpacho

764 compare

buzz

1,913 compare

embed

2,103 compare

httpful

1,741 compare