使用Python从HTML文件中提取文本

我想使用Python从HTML文件中提取文本。我想从本质上得到相同的输出，如果我从浏览器复制文本，并将其粘贴到记事本。

我想要一些更健壮的东西，而不是使用正则表达式，正则表达式可能会在格式不佳的HTML上失败。我见过很多人推荐Beautiful Soup，但我在使用它时遇到了一些问题。首先，它会抓取不需要的文本，比如JavaScript源代码。此外，它也不解释HTML实体。例如，我会期望'在HTML源代码中转换为文本中的撇号，就像我将浏览器内容粘贴到记事本一样。

更新html2text看起来很有希望。它正确地处理HTML实体，而忽略JavaScript。然而，它并不完全生成纯文本;它产生的降价，然后必须转换成纯文本。它没有示例或文档，但代码看起来很干净。

相关问题:

在python中过滤HTML标签并解析实体在Python中将XML/HTML实体转换为Unicode字符串

当前回答

Perl方式(对不起妈妈，我永远不会在生产中这样做)。

import re

def html2text(html):
    res = re.sub('<.*?>', ' ', html, flags=re.DOTALL | re.MULTILINE)
    res = re.sub('\n+', '\n', res)
    res = re.sub('\r+', '', res)
    res = re.sub('[\t ]+', ' ', res)
    res = re.sub('\t+', '\t', res)
    res = re.sub('(\n )+', '\n ', res)
    return res

2018-07-06 11:36:06

其他回答

如果你想从网页中自动提取文本段落，有一些可用的python包，如Trafilatura。作为基准测试的一部分，比较了几个python包:

https://github.com/adbar/trafilatura#evaluation-and-alternatives

html_text https://github.com/TeamHG-Memex/html-text inscriptis https://github.com/weblyzard/inscriptis newspaper3k justext boilerpy3 https://github.com/jmriebold/BoilerPy3 基线 goose3 https://github.com/goose3/goose3 readability-lxml https://github.com/predatell/python-readability-lxml news-please https://github.com/fhamborg/news-please readabilipy https://github.com/alan-turing-institute/ReadabiliPy trafilatura

2022-09-18 20:12:04

有用于数据挖掘的模式库。

http://www.clips.ua.ac.be/pages/pattern-web

你甚至可以决定保留什么标签:

s = URL('http://www.clips.ua.ac.be').download()
s = plaintext(s, keep={'h1':[], 'h2':[], 'strong':[], 'a':['href']})
print s

2012-11-29 19:28:38

这不是一个完全的Python解决方案，但它会将Javascript生成的文本转换为文本，我认为这是重要的(例如google.com)。浏览器Links(不是Lynx)有一个Javascript引擎，可以通过-dump选项将源代码转换为文本。

所以你可以这样做:

fname = os.tmpnam()
fname.write(html_source)
proc = subprocess.Popen(['links', '-dump', fname], 
                        stdout=subprocess.PIPE,
                        stderr=open('/dev/null','w'))
text = proc.stdout.read()

2012-05-18 10:02:26

另一个非python解决方案:Libre Office:

soffice --headless --invisible --convert-to txt input1.html

我更喜欢这种方法的原因是，每个HTML段落都转换为单个文本行(没有换行符)，这正是我所寻找的。其他方法需要后处理。Lynx的输出确实不错，但并不是我想要的。此外，Libre Office可以用来从各种格式转换…

2015-12-11 04:11:45

安装html2text using

PIP安装html2text

然后,

>>> import html2text
>>>
>>> h = html2text.HTML2Text()
>>> # Ignore converting links from HTML
>>> h.ignore_links = True
>>> print h.handle("<p>Hello, <a href='http://earth.google.com/'>world</a>!")
Hello, world!

2017-04-05 07:16:30

使用Python从HTML文件中提取文本

推荐文章

最新文章

标签