scrapydwebvssplash

GPL-3.0 61 1 3,111

1.4 thousand (month) Sep 30 2018 1.5.0(8 months ago)

4,078 15 404 BSD-3-Clause

Apr 25 2014 727 (month) 3.5(4 years ago)

ScrapydWeb is a web-based management tool for the Scrapyd service. It is built using the Python Flask framework and allows you to easily manage and monitor your Scrapy spider projects through a web interface.

ScrapydWeb allows you to view the status of your running spiders, view the logs of completed spiders, schedule new spider runs, and manage spider settings and configurations.

ScrapydWeb provides a simple way to manage your scraping tasks and allows you to schedule and run multiple spiders simultaneously. It also provides a user-friendly web interface that makes it easy to view the status of your spiders and monitor their progress.

You can install the package via pip by running pip install scrapydweb and then you can run the package by running scrapydweb command in your command prompt.

It will start a web server that you can access through your web browser at http://localhost:6800/ You will need to have Scrapyd running in order to use ScrapydWeb, Scrapyd is a service for running Scrapy spiders, it allows you to schedule spiders to run at regular intervals and also allows you to run spiders on remote machines.

Splash is a javascript rendering service with an HTTP API. It's a lightweight browser with an HTTP API, implemented in Python 3 using Twisted and QT5.

It is built on top of the QtWebkit library and allows developers to interact with web pages in a headless mode, which means that the web pages are rendered in the background, without displaying them on the screen.

splash is particularly useful for web scraping and web testing tasks, as it allows developers to interact with web pages in a way that is very similar to how a human user would interact with the browser.

It also allows you to execute javascript and interact with web pages even if they use heavy javascript.

Unlike Selenium or Playwright, splash is powered by webkit embedded browser instead of a real browser like Chrome or Firefox. As a down-side splash requests are easy to detect and block when scraping websites with anti-scraping features.

One benefit of splash is that it seemlesly integrates with Scrapy.

Example Use

# once splash server is started it can be requested to render pages through 
# HTTP requests:
import requests

url = "http://localhost:8050/render.html"
payload = {
    'url': 'https://www.example.com',
    'timeout': 30,
    'wait': 2
}

response = requests.get(url, params=payload)

# Get the page HTML
print(response.text)

Alternatives / Similar

gerapy

3,324 compare

scrapy

52,353 compare

scrapyd

2,942 compare

colly

23,105 compare

pholcus

7,566 compare

geziyor

2,602 compare

dataflowkit

658 compare

rvest

1,489 compare

ferret

5,716 compare

gocrawl

2,037 compare

node-crawler

6,693 compare

autoscraper

6,193 compare

panther

2,931 compare

spidr

802 compare

wombat

1,310 compare

gracy

246 compare

ruia

1,748 compare

splash

4,078 compare

photon

10,939 compare

roach

1,353 compare

ralger

156 compare

phpscraper

526 compare

ayakashi

209 compare

dude

420 compare

php-spider

1,330 compare

crwlr-crawler

322 compare