htmlparser2vspyquery

MIT 12 4 4,789

277.1 million (month) Aug 28 2011 12.0.0(2026-03-20 23:08:40 ago)

2,381 5 55 NOASSERTION

Dec 05 2008 2.0 million (month) 2.0.1(2024-08-30 08:12:22 ago)

htmlparser2 is a Node.js library for parsing HTML and XML documents. It works by building a tree of elements, similar to the Document Object Model (DOM) in web browsers. This allows you to easily traverse and manipulate the structure of the document.

htmlparser2 is a low-level html tree parser but it can still be useful in web scraping as it's a powerful tool for HTML restructuring and serialization.

PyQuery is a Python library for working with XML and HTML documents. It is similar to BeautifulSoup and is often used as a drop-in replacement for it.

PyQuery is inspired by javascript's jQuery and uses similar API allowing selecting of HTML nodes through CSS selectors. This makes it easy for developers who are already familiar with jQuery to use PyQuery in Python.

Unlike jQuery, PyQuery doesn't support XPath selectors and relies entirely on CSS selectors though offers similar HTML parsing features like selection of HTML elements, their attributes and text as well as html tree modification.

PyQuery also comes with a http client (through requests) so it can load and parse web URLs by itself.

Highlights

css-selectors

Example Use

```javascript const htmlparser = require("htmlparser2"); const parser = new htmlparser.Parser({ onopentag: (name, attribs) => { console.log(`Opening tag: ${name}`); }, ontext: (text) => { console.log(`Text: ${text}`); }, onclosetag: (name) => { console.log(`Closing tag: ${name}`); } }, {decodeEntities: true}); const html = "

Hello, world!

"; parser.write(html); parser.end(); ```

```python from pyquery import PyQuery as pq # this is our HTML page: html = """ Hello World!

Product Title

paragraph 1

paragraph2

$10

""" doc = pq(html) # we can use CSS selectors: print(doc('#product .price').text()) "$10" # it's also possible to modify HTML tree in various ways: # insert text into selected element: print(doc('h1').append('discounted')) "

Product Titlediscounted

" # or remove elements doc('p').remove() print(doc('#product').html()) """

Product Titlediscounted

$10 """ # pyquery can also retrieve web documents using requests: doc = pq(url='http://httpbin.org/html', headers={"User-Agent": "webscraping.fyi"}) print(doc('h1').html()) ```

Alternatives / Similar

parse5

3,886 compare

sax-js

1,153 compare

lxml

3,010 compare

beautifulsoup

- compare

jsdom new

21,552 compare

xmltodict

5,734 compare

cheerio

30,265 compare

html5lib

1,220 compare

cssselect

309 compare

feedparser

2,351 compare

nokogiri

6,248 compare

parsel

1,324 compare

selectolax

1,607 compare

pyquery

2,381 compare

xml2

223 compare

requests-html

13,863 compare

rvest

1,517 compare

untangle

632 compare

scrapling new

36,206 compare

html5-php

1,772 compare

domcrawler

4,038 compare

goquery

14,926 compare

cascadia

754 compare

htmlquery

781 compare

xpath

739 compare

soup

2,227 compare

chompjs

218 compare

html5-parser

700 compare

gazpacho

768 compare

embed

2,103 compare

chopper

23 compare

simple-html-dom new

- compare

ralger

165 compare

parse5

3,886 compare

sax-js

1,153 compare

htmlparser2

4,789 compare

lxml

3,010 compare

beautifulsoup

- compare

jsdom new

21,552 compare

xmltodict

5,734 compare

cheerio

30,265 compare

html5lib

1,220 compare

cssselect

309 compare

feedparser

2,351 compare

nokogiri

6,248 compare

parsel

1,324 compare

selectolax

1,607 compare

xml2

223 compare

requests-html

13,863 compare

rvest

1,517 compare

untangle

632 compare

scrapling new

36,206 compare

html5-php

1,772 compare

domcrawler

4,038 compare

goquery

14,926 compare

cascadia

754 compare

htmlquery

781 compare

xpath

739 compare

soup

2,227 compare

chompjs

218 compare

html5-parser

700 compare

gazpacho

768 compare

embed

2,103 compare

chopper

23 compare

simple-html-dom new

- compare

ralger

165 compare