Skip to content

newspapervsextractnet

MIT 513 6 15,018
1.0 million (month) Dec 28 2012 0.2.8(2018-09-28 04:58:18 ago)
297 9 9 MIT
Dec 11 2020 131 (month) 2.0.7(2022-11-06 07:33:14 ago)

newspaper is a Python package that allows developers to easily extract text, images, and videos from articles on the web.

It is designed to be fast, easy to use, and compatible with a wide variety of websites. It uses advanced algorithms to extract relevant information and metadata from articles, and it also supports several languages.

newspaper includes a http client or can ingest pre-scraped HTML documents.

ExtractNet is an automated web data extraction tool using machine learning to parse HTML and text data.

This tool can be used in web scraping to automatically extract details from scraped HTML documents. While it's not as accurate as structured extraction using HTML parsing tools like CSS selectors or XPath it can still parse a lot of details.

Example Use


```python from newspaper import Article # Create a new article object article = Article('https://www.example.com/article') # Download the article article.download() # Parse the article article.parse() # Print the article text print(article.text) # Print the article title print(article.title) # Print the article authors print(article.authors) # Print the article publication date print(article.publish_date) ```
```python import requests from extractnet import Extractor raw_html = requests.get('https://currentsapi.services/en/blog/2019/03/27/python-microframework-benchmark/.html').text results = Extractor().extract(raw_html) {'phone_number': '555-555-5555', 'email': 'example@example.com'} ```

Alternatives / Similar


Was this page helpful?