dudevsralger

python html-parser framework http-client

AGPL-3.0 29 2 428

157 (month) Feb 20 2022 0.1.3(2 years ago)

156 1 3 MIT

Dec 22 2019 349 (month) 2.2.4(4 years ago)

Dude (dude uncomplicated data extraction) is a very simple framework for writing web scrapers using Python decorators. The design, inspired by Flask, was to easily build a web scraper in just a few lines of code. Dude has an easy-to-learn syntax.

The simplest web scraper will look like this:

from dude import select


@select(css="a")
def get_link(element):
    return {"url": element.get_attribute("href")}

dude supports multiple parser backends: - playwright
- lxml
- parsel - beautifulsoup - pyppeteer - selenium

ralger is a small web scraping framework for R based on rvest and xml2.

It's goal to simplify basic web scraping and it provides a convenient and easy to use API.

It offers functions for retrieving pages, parsing HTML using CSS selectors, automatic table parsing and auto link, title, image and paragraph extraction.

Example Use

from dude import select

"""
This example demonstrates how to use Parsel + async HTTPX
To access an attribute, use:
    selector.attrib["href"]
You can also access an attribute using the ::attr(name) pseudo-element, for example "a::attr(href)", then:
    selector.get()
To get the text, use ::text pseudo-element, then:
    selector.get()
"""


@select(css="a.url", priority=2)
async def result_url(selector):
    return {"url": selector.attrib["href"]}


# Option to get url using ::attr(name) pseudo-element
@select(css="a.url::attr(href)", priority=2)
async def result_url2(selector):
    return {"url2": selector.get()}


@select(css=".title::text", priority=1)
async def result_title(selector):
    return {"title": selector.get()}


@select(css=".description::text", priority=0)
async def result_description(selector):
    return {"description": selector.get()}


if __name__ == "__main__":
    import dude

    dude.run(urls=["https://dude.ron.sh"], parser="parsel")

library("ralger")

url <- "http://www.shanghairanking.com/rankings/arwu/2021"

# retrieve HTML and select elements using CSS selectors:
best_uni <- scrap(link = url, node = "a span", clean = TRUE)
head(best_uni, 5)
#>  [1] "Harvard University"
#>  [2] "Stanford University"
#>  [3] "University of Cambridge"
#>  [4] "Massachusetts Institute of Technology (MIT)"
#>  [5] "University of California, Berkeley"

# ralger can also parse HTML attributes
attributes <- attribute_scrap(
  link = "https://ropensci.org/",
  node = "a", # the a tag
  attr = "class" # getting the class attribute
)

head(attributes, 10) # NA values are a tags without a class attribute
#>  [1] "navbar-brand logo" "nav-link"          NA
#>  [4] NA                  NA                  "nav-link"
#>  [7] NA                  "nav-link"          NA
#> [10] NA
#

# ralger can automatically scrape tables:
data <- table_scrap(link ="https://www.boxofficemojo.com/chart/top_lifetime_gross/?area=XWW")

head(data)
#> # A tibble: 6 × 4
#>    Rank Title                                      `Lifetime Gross`  Year
#>   <int> <chr>                                      <chr>            <int>
#> 1     1 Avatar                                     $2,847,397,339    2009
#> 2     2 Avengers: Endgame                          $2,797,501,328    2019
#> 3     3 Titanic                                    $2,201,647,264    1997
#> 4     4 Star Wars: Episode VII - The Force Awakens $2,069,521,700    2015
#> 5     5 Avengers: Infinity War                     $2,048,359,754    2018
#> 6     6 Spider-Man: No Way Home                    $1,901,216,740    2021

Alternatives / Similar

colly

23,747 compare

pholcus

7,580 compare

geziyor

2,667 compare

dataflowkit

676 compare

scrapy

54,211 compare

rvest

1,498 compare

gocrawl

2,039 compare

ferret

5,716 compare

scrapyd

2,980 compare

node-crawler

6,733 compare

panther

2,977 compare

autoscraper

6,638 compare

gracy

247 compare

spidr

813 compare

scrapydweb

3,218 compare

gerapy

3,365 compare

wombat

1,316 compare

ruia

1,754 compare

photon

11,149 compare

ralger

156 compare

roach

1,384 compare

ayakashi

213 compare

phpscraper

554 compare

php-spider

1,335 compare

crwlr-crawler

356 compare

rvest

1,498 compare

requests

52,519 compare

node-fetch

8,825 compare

colly

23,747 compare

axios

106,345 compare

aiohttp

15,425 compare

sax-js

1,101 compare

parse5

3,698 compare

htmlparser2

4,529 compare

httpx

13,703 compare

beautifulsoup

- compare

lxml

2,737 compare

got

14,454 compare

pholcus

7,580 compare

xmltodict

5,577 compare

curl-impersonate

4,221 compare

superagent

16,610 compare

needle

1,637 compare

cheerio

28,873 compare

geziyor

2,667 compare

html5lib

1,153 compare

cssselect

293 compare

dataflowkit

676 compare

faraday

5,785 compare

nokogiri

6,173 compare

httpclient

703 compare

feedparser

2,048 compare

excon

1,163 compare

httparty

5,837 compare

pyquery

2,312 compare

pycurl

1,094 compare

parsel

1,187 compare

scrapy

54,211 compare

typhoeus

4,084 compare

requests-html

13,780 compare

httr

988 compare

xml2

221 compare

curl-cffi

1,751 compare

guzzle

23,055 compare

em-http-request

1,217 compare

selectolax

1,186 compare

symfony-http

1,976 compare

html5-php

1,638 compare

untangle

619 compare

domcrawler

3,985 compare

wreck

381 compare

http-2

898 compare

treq

590 compare

goquery

14,273 compare

xpath

699 compare

resty

10,341 compare

cascadia

717 compare

soup

2,191 compare

req

4,374 compare

gocrawl

2,039 compare

htmlquery

760 compare

ferret

5,716 compare

nestful

505 compare

html5-parser

683 compare

crul

107 compare

scrapyd

2,980 compare

chompjs

202 compare

requests

3,576 compare

node-crawler

6,733 compare

hrequests

780 compare

panther

2,977 compare

gazpacho

764 compare

buzz

1,913 compare

embed

2,103 compare

httpful

1,741 compare