wombatvsspidr

MIT 24 2 1,360

1.4 thousand (month) Dec 27 2011 3.3.0(2026-04-07 16:31:34 ago)

835 2 16 MIT

Jul 25 2009 2.4 thousand (month) 0.7.2(2025-02-03 07:58:27 ago)

Wombat is a Ruby gem that makes it easy to scrape websites and extract structured data from HTML pages. It is built on top of Nokogiri, a popular Ruby gem for parsing and searching HTML and XML documents, and it provides a simple and intuitive API for defining and running web scraping operations.

One of the main features of Wombat is its ability to extract structured data from HTML pages using a simple, CSS-like syntax. It allows you to define a set of rules for extracting data from a page, and then automatically applies those rules to the page's HTML to extract the desired data. This makes it easy to extract data from even complex and dynamic pages, without having to write a lot of custom code.

In addition to its data extraction capabilities, Wombat also provides a variety of other features that can simplify the web scraping process. It can automatically follow links and scrape multiple pages, it can handle pagination and AJAX requests, and it can handle cookies and authentication. It also provides a built-in support for parallelism and queueing to speed up the scraping process.

Spidr is a Ruby gem that provides a simple and flexible way to spider and scrape websites. It allows you to easily navigate through a website, following links and scraping data as you go. It is built on top of Nokogiri, a popular Ruby gem for parsing and searching HTML and XML documents, and it provides a simple and intuitive API for defining and running web scraping operations.

One of the main features of Spidr is its ability to spider a website, following all the links on a page and visiting all the pages of a website. This allows you to easily and quickly scrape large amounts of data from a website, without having to manually specify which pages to visit.

In addition to its spidering capabilities, Spidr also provides a variety of other features that can simplify the web scraping process. It can automatically filter which links to follow and which pages to visit, it can handle cookies and authentication, and it can automatically store the scraped data in a database or file. It also provides a built-in support for parallelism and queueing to speed up the scraping process.

Example Use

```ruby require 'wombat' Wombat.crawl do base_url "https://www.github.com" path "/" headline xpath: "//h1" subheading css: "p.alt-lead" what_is({ css: ".one-fourth h4" }, :list) links do explore xpath: '/html/body/header/div/div/nav[1]/a[4]' do |e| e.gsub(/Explore/, "Love") end features css: '.nav-item-opensource' business css: '.nav-item-business' end end ``` will result in: ```json { "headline"=>"How people build software", "subheading"=>"Millions of developers use GitHub to build personal projects, support their businesses, and work together on open source technologies.", "what_is"=>[ "For everything you build", "A better way to work", "Millions of projects", "One platform, from start to finish" ], "links"=>{ "explore"=>"Love", "features"=>"Open source", "business"=>"Business" } } ```

```ruby require 'spidr' Spidr.start_at("http://example.com") do |spider| spider.every_page do |page| puts "Visiting: #{page.url}" # Extract data from the page using Nokogiri doc = Nokogiri::HTML(page.body) title = doc.css("title").text puts "Title: #{title}" end end ```

Alternatives / Similar

colly

25,231 compare

katana new

16,499 compare

pholcus

7,594 compare

geziyor

2,772 compare

dataflowkit

711 compare

scrapy

61,276 compare

crawl4ai new

63,373 compare

rvest

1,517 compare

scrapling new

36,206 compare

crawlee new

22,720 compare

mechanize new

4,440 compare

scrapegraphai new

23,278 compare

ferret

5,964 compare

gocrawl

2,053 compare

scrapyd

3,087 compare

botasaurus new

4,321 compare

node-crawler

6,790 compare

panther

3,062 compare

goutte new

9,215 compare

gracy

248 compare

spidr

835 compare

kimurai new

1,098 compare

scrapydweb

3,400 compare

photon

12,807 compare

autoscraper

7,136 compare

roach

1,454 compare

gerapy

3,495 compare

ruia

1,743 compare

ralger

165 compare

ayakashi

217 compare

phpscraper

583 compare

dude

425 compare

php-spider

1,341 compare

crwlr-crawler

369 compare

firecrawl new

- compare