htmlparser2vsnokogiri

javascript html-parser css-selectors xpath

MIT 12 4 4,789

277.1 million (month) Aug 28 2011 12.0.0(2026-03-20 23:08:40 ago)

6,248 23 108 MIT

Jul 25 2009 5.8 million (month) 1.19.2(2026-03-19 21:12:43 ago)

htmlparser2 is a Node.js library for parsing HTML and XML documents. It works by building a tree of elements, similar to the Document Object Model (DOM) in web browsers. This allows you to easily traverse and manipulate the structure of the document.

htmlparser2 is a low-level html tree parser but it can still be useful in web scraping as it's a powerful tool for HTML restructuring and serialization.

Nokogiri is a Ruby gem that provides a simple and powerful way to parse and search XML and HTML documents. It is built on top of the underlying C library libxml2, which is known for its speed and reliability.

Nokogiri provides a simple and intuitive API for parsing and searching XML and HTML documents, and it is widely used in the Ruby ecosystem for web scraping and data extraction.

One of the main features of Nokogiri is its ability to search and navigate through XML and HTML documents using a CSS or XPath selectors.

Nokogiri also provides a variety of other features that can simplify the process of working with XML and HTML documents. It can automatically handle character encodings and normalize documents, it can parse and search large documents with low memory usage, and it can validate documents against a DTD or schema.

Highlights

css-selectorsxpathpopular

Example Use

```javascript const htmlparser = require("htmlparser2"); const parser = new htmlparser.Parser({ onopentag: (name, attribs) => { console.log(`Opening tag: ${name}`); }, ontext: (text) => { console.log(`Text: ${text}`); }, onclosetag: (name) => { console.log(`Closing tag: ${name}`); } }, {decodeEntities: true}); const html = "

Hello, world!

"; parser.write(html); parser.end(); ```

```ruby require 'nokogiri' html_string = 'Page Title

Hello World!

This is a sample webpage.

' # Parse the HTML string doc = Nokogiri::HTML(html_string) # Extract the class attribute of h1 tag using CSS selector h1_class = doc.css("h1")[0]['class'] # or XPath h1_class = doc.xpath("//h1")[0]['class'] puts "H1 class: #{h1_class}" ```

Alternatives / Similar

parse5

3,886 compare

sax-js

1,153 compare

lxml

3,010 compare

beautifulsoup

- compare

jsdom new

21,552 compare

xmltodict

5,734 compare

cheerio

30,265 compare

html5lib

1,220 compare

cssselect

309 compare

feedparser

2,351 compare

nokogiri

6,248 compare

parsel

1,324 compare

selectolax

1,607 compare

pyquery

2,381 compare

xml2

223 compare

requests-html

13,863 compare

rvest

1,517 compare

untangle

632 compare

scrapling new

36,206 compare

html5-php

1,772 compare

domcrawler

4,038 compare

goquery

14,926 compare

cascadia

754 compare

htmlquery

781 compare

xpath

739 compare

soup

2,227 compare

chompjs

218 compare

html5-parser

700 compare

gazpacho

768 compare

embed

2,103 compare

chopper

23 compare

simple-html-dom new

- compare

ralger

165 compare

domcrawler

4,038 compare

parse5

3,886 compare

sax-js

1,153 compare

htmlparser2

4,789 compare

lxml

3,010 compare

beautifulsoup

- compare

jsdom new

21,552 compare

xmltodict

5,734 compare

cheerio

30,265 compare

html5lib

1,220 compare

cssselect

309 compare

feedparser

2,351 compare

parsel

1,324 compare

selectolax

1,607 compare

pyquery

2,381 compare

xml2

223 compare

requests-html

13,863 compare

rvest

1,517 compare

untangle

632 compare

scrapling new

36,206 compare

html5-php

1,772 compare

goquery

14,926 compare

cascadia

754 compare

htmlquery

781 compare

xpath

739 compare

soup

2,227 compare

chompjs

218 compare

html5-parser

700 compare

gazpacho

768 compare

embed

2,103 compare

chopper

23 compare

simple-html-dom new

- compare

ralger

165 compare