curl-cffivsralger

python html-parser framework http-client

MIT 34 2 1,751

594.9 thousand (month) Feb 23 2022 0.7.1(1 year, 1 month ago)

156 1 3 MIT

Dec 22 2019 349 (month) 2.2.4(4 years ago)

Curl-cffi is a Python library for implementing curl-impersonate which is a HTTP client that appears as one of popular web browsers like: - Google Chrome - Microsoft Edge - Safari - Firefox Unlike requests and httpx which are native Python libraries, curl-cffi uses cURL and inherits it's powerful features like extensive HTTP protocol support and detection patches for TLS and HTTP fingerprinting.

Using curl-cffi web scrapers can bypass TLS and HTTP fingerprinting.

ralger is a small web scraping framework for R based on rvest and xml2.

It's goal to simplify basic web scraping and it provides a convenient and easy to use API.

It offers functions for retrieving pages, parsing HTML using CSS selectors, automatic table parsing and auto link, title, image and paragraph extraction.

Highlights

bypasshttp2tls-fingerprinthttp-fingerprintsyncasync

Example Use

curl-cffi can be accessed as low-level curl client as well as an easy high-level HTTP client:

from curl_cffi import requests

response = requests.get('https://httpbin.org/json')
print(response.json())

# or using sessions
session = requests.Session()
response = session.get('https://httpbin.org/json')

# also supports async requests using asyncio
import asyncio
from curl_cffi.requests import AsyncSession

urls = [
  "http://httpbin.org/html",
  "http://httpbin.org/html",
  "http://httpbin.org/html",
]

async with AsyncSession() as s:
    tasks = []
    for url in urls:
        task = s.get(url)
        tasks.append(task)
    # scrape concurrently:
    responses = await asyncio.gather(*tasks)

# also supports websocket connections
from curl_cffi.requests import Session, WebSocket

def on_message(ws: WebSocket, message):
    print(message)

with Session() as s:
    ws = s.ws_connect(
        "wss://api.gemini.com/v1/marketdata/BTCUSD",
        on_message=on_message,
    )
    ws.run_forever()

library("ralger")

url <- "http://www.shanghairanking.com/rankings/arwu/2021"

# retrieve HTML and select elements using CSS selectors:
best_uni <- scrap(link = url, node = "a span", clean = TRUE)
head(best_uni, 5)
#>  [1] "Harvard University"
#>  [2] "Stanford University"
#>  [3] "University of Cambridge"
#>  [4] "Massachusetts Institute of Technology (MIT)"
#>  [5] "University of California, Berkeley"

# ralger can also parse HTML attributes
attributes <- attribute_scrap(
  link = "https://ropensci.org/",
  node = "a", # the a tag
  attr = "class" # getting the class attribute
)

head(attributes, 10) # NA values are a tags without a class attribute
#>  [1] "navbar-brand logo" "nav-link"          NA
#>  [4] NA                  NA                  "nav-link"
#>  [7] NA                  "nav-link"          NA
#> [10] NA
#

# ralger can automatically scrape tables:
data <- table_scrap(link ="https://www.boxofficemojo.com/chart/top_lifetime_gross/?area=XWW")

head(data)
#> # A tibble: 6 × 4
#>    Rank Title                                      `Lifetime Gross`  Year
#>   <int> <chr>                                      <chr>            <int>
#> 1     1 Avatar                                     $2,847,397,339    2009
#> 2     2 Avengers: Endgame                          $2,797,501,328    2019
#> 3     3 Titanic                                    $2,201,647,264    1997
#> 4     4 Star Wars: Episode VII - The Force Awakens $2,069,521,700    2015
#> 5     5 Avengers: Infinity War                     $2,048,359,754    2018
#> 6     6 Spider-Man: No Way Home                    $1,901,216,740    2021

Alternatives / Similar

curl-impersonate

4,221 compare

hrequests

780 compare

requests

52,519 compare

node-fetch

8,825 compare

axios

106,345 compare

aiohttp

15,425 compare

httpx

13,703 compare

got

14,454 compare

superagent

16,610 compare

needle

1,637 compare

faraday

5,785 compare

httpclient

703 compare

undetected-chromedriver

10,683 compare

excon

1,163 compare

httparty

5,837 compare

pycurl

1,094 compare

typhoeus

4,084 compare

puppeteer-stealth

89,751 compare

httr

988 compare

rvest

1,498 compare

guzzle

23,055 compare

em-http-request

1,217 compare

symfony-http

1,976 compare

wreck

381 compare

http-2

898 compare

treq

590 compare

resty

10,341 compare

req

4,374 compare

nestful

505 compare

crul

107 compare

requests

3,576 compare

selenium-driverless

718 compare

buzz

1,913 compare

httpful

1,741 compare

ralger

156 compare

http.rb

3,013 compare

rvest

1,498 compare

requests

52,519 compare

node-fetch

8,825 compare

colly

23,747 compare

axios

106,345 compare

aiohttp

15,425 compare

sax-js

1,101 compare

parse5

3,698 compare

htmlparser2

4,529 compare

httpx

13,703 compare

beautifulsoup

- compare

lxml

2,737 compare

got

14,454 compare

pholcus

7,580 compare

xmltodict

5,577 compare

curl-impersonate

4,221 compare

superagent

16,610 compare

needle

1,637 compare

cheerio

28,873 compare

geziyor

2,667 compare

html5lib

1,153 compare

cssselect

293 compare

dataflowkit

676 compare

faraday

5,785 compare

nokogiri

6,173 compare

httpclient

703 compare

feedparser

2,048 compare

excon

1,163 compare

httparty

5,837 compare

pyquery

2,312 compare

pycurl

1,094 compare

parsel

1,187 compare

scrapy

54,211 compare

typhoeus

4,084 compare

requests-html

13,780 compare

httr

988 compare

xml2

221 compare

curl-cffi

1,751 compare

guzzle

23,055 compare

em-http-request

1,217 compare

selectolax

1,186 compare

symfony-http

1,976 compare

html5-php

1,638 compare

untangle

619 compare

domcrawler

3,985 compare

wreck

381 compare

http-2

898 compare

treq

590 compare

goquery

14,273 compare

xpath

699 compare

resty

10,341 compare

cascadia

717 compare

soup

2,191 compare

req

4,374 compare

gocrawl

2,039 compare

htmlquery

760 compare

ferret

5,716 compare

nestful

505 compare

html5-parser

683 compare

crul

107 compare

scrapyd

2,980 compare

chompjs

202 compare

requests

3,576 compare

node-crawler

6,733 compare

hrequests

780 compare

panther

2,977 compare

gazpacho

764 compare

buzz

1,913 compare

embed

2,103 compare

httpful

1,741 compare

autoscraper

6,638 compare

gracy

247 compare

spidr

813 compare

scrapydweb

3,218 compare

gerapy

3,365 compare

wombat

1,316 compare

ruia

1,754 compare

chopper

22 compare

photon

11,149 compare

roach

1,384 compare

dude

428 compare

http.rb

3,013 compare

ayakashi

213 compare

phpscraper

554 compare

php-spider

1,335 compare

crwlr-crawler

356 compare