sumyvstrafilatura

Apache-2.0 28 4 3,670

152.5 thousand (month) Oct 20 2013 0.12.0(2026-02-14 21:00:12 ago)

5,650 4 107 Apache-2.0

Jul 17 2019 5.2 million (month) 2.0.0(2024-12-03 15:23:21 ago)

sumy is a Python library for automatic summarization of text documents. It can be used to extract summaries from various input formats such as plaintext, HTML, and URLs. It supports multiple languages and multiple summarization algorithms, including Latent Semantic Analysis (LSA), Luhn, Edmundson, TextRank, and SumBasic.

Trafilatura is a Python package and command-line tool designed to gather text on the Web. It includes discovery, extraction and text processing components. Its main applications are web crawling, downloads, scraping, and extraction of main texts, metadata and comments. It aims at staying handy and modular: no database is required, the output can be converted to various commonly used formats.

Going from raw HTML to essential parts can alleviate many problems related to text quality, first by avoiding the noise caused by recurring elements (headers, footers, links/blogroll etc.) and second by including information such as author and date in order to make sense of the data. The extractor tries to strike a balance between limiting noise (precision) and including all valid parts (recall). It also has to be robust and reasonably fast, it runs in production on millions of documents.

This tool can be useful for quantitative research in corpus linguistics, natural language processing, computational social science and beyond: it is relevant to anyone interested in data science, information extraction, text mining, and scraping-intensive use cases like search engine optimization, business analytics or information security.

Example Use

```python # -*- coding: utf-8 -*- from __future__ import absolute_import from __future__ import division, print_function, unicode_literals from sumy.parsers.html import HtmlParser from sumy.parsers.plaintext import PlaintextParser from sumy.nlp.tokenizers import Tokenizer from sumy.summarizers.lsa import LsaSummarizer as Summarizer from sumy.nlp.stemmers import Stemmer from sumy.utils import get_stop_words LANGUAGE = "english" SENTENCES_COUNT = 10 if __name__ == "__main__": url = "https://en.wikipedia.org/wiki/Automatic_summarization" parser = HtmlParser.from_url(url, Tokenizer(LANGUAGE)) # or for plain text files # parser = PlaintextParser.from_file("document.txt", Tokenizer(LANGUAGE)) # parser = PlaintextParser.from_string("Check this out.", Tokenizer(LANGUAGE)) stemmer = Stemmer(LANGUAGE) summarizer = Summarizer(stemmer) summarizer.stop_words = get_stop_words(LANGUAGE) for sentence in summarizer(parser.document, SENTENCES_COUNT): print(sentence) ```

```python # it can be used to clean HTML files from trafilatura import clean_html html = 'My Title

This is some bold text.

' cleaned_html = clean_html(html) print(cleaned_html) # can strip away tags: clean_html(html, tags_to_remove=["title"]) # or attributes clean_html(html, attributes_to_remove=["title"]) ```

Alternatives / Similar

html2text

2,140 compare

readability

2,894 compare

newspaper

15,018 compare

extruct

961 compare

youtube-dl

140,026 compare

sumy

3,670 compare

gofeed

2,824 compare

you-get

56,813 compare

embed

2,103 compare

embera

353 compare

photon

12,807 compare

essence

769 compare

extractnet

297 compare