readability
python-readability is a python package that allows developers to extract the main content of a web page, removing any unnecessary or unwanted elements, such as ads, navigation, and sidebars.
It is based on the algorithm used by the popular web-based service, Readability, and it uses the beautifulsoup4 package to parse the HTML and extract the main content.
Readability is similar to Newspaper in terms that it's extracting HTML data
Example Use
```python import requests from readability import document
response = requests.get('http://example.com') doc = document(response.content) doc.title() 'example domain'
doc.summary() """
example domain
\nthis domain is established to be used for illustrative examples in documents. you may use this\n domain in examples without prior coordination or asking for permission.
\n
\n\n\n
""" ```