在Python中从字符串中剥离HTML

from mechanize import Browser
br = Browser()
br.open('http://somewebpage')
html = br.response().readlines()
for line in html:
  print line

当在HTML文件中打印一行时，我试图找到一种方法，只显示每个HTML元素的内容，而不是格式本身。如果它发现'<a href="等等。例如">some text</a>'，它只会打印'some text'， 'hello'打印'hello'，等等。该怎么做呢?

当前回答

如果您需要剥离HTML标记来进行文本处理，那么一个简单的正则表达式就可以了。如果您希望清除用户生成的HTML以防止XSS攻击，请不要使用此方法。删除所有<script>标签或跟踪<img>s不是一个安全的方法。下面的正则表达式将相当可靠地剥离大多数HTML标记:

import re

re.sub('<[^<]+?>', '', text)

对于那些不理解regex的人来说，这将搜索字符串<…>，其中内部内容由一个或多个不是<的(+)字符组成。的吗?意味着它将匹配它能找到的最小字符串。例如，给定Hello，它将分别用?匹配<'p>和。没有它，它将匹配整个字符串<..Hello..>。

如果非标签<出现在html(例如。2 < 3)，它应该被写成转义序列&…总之，^<可能是不必要的。

2011-02-02 01:09:16

其他回答

如果你需要保留HTML实体(即&)，我在Eloff的答案中添加了“handle_entityref”方法。

from HTMLParser import HTMLParser

class MLStripper(HTMLParser):
    def __init__(self):
        self.reset()
        self.fed = []
    def handle_data(self, d):
        self.fed.append(d)
    def handle_entityref(self, name):
        self.fed.append('&%s;' % name)
    def get_data(self):
        return ''.join(self.fed)

def html_to_text(html):
    s = MLStripper()
    s.feed(html)
    return s.get_data()

2012-12-04 13:25:42

你可以使用BeautifulSoup get_text()特性。

from bs4 import BeautifulSoup

html_str = '''
<td><a href="http://www.fakewebsite.example">Please can you strip me?</a>
<br/><a href="http://www.fakewebsite.example">I am waiting....</a>
</td>
'''
soup = BeautifulSoup(html_str)

print(soup.get_text())
#or via attribute of Soup Object: print(soup.text)

建议显式地指定解析器，例如BeautifulSoup(html_str, features="html.parser")，以便输出可重现。

2015-12-30 15:31:13

如果你想去掉所有HTML标签，我发现最简单的方法是使用BeautifulSoup:

from bs4 import BeautifulSoup  # Or from BeautifulSoup import BeautifulSoup

def stripHtmlTags(htmlTxt):
    if htmlTxt is None:
            return None
        else:
            return ''.join(BeautifulSoup(htmlTxt).findAll(text=True))

我尝试了接受的答案的代码，但我得到了“RuntimeError:最大递归深度超出”，这没有发生在上面的代码块。

2013-01-30 11:47:49

我正在解析Github自述，我发现下面的工作真的很好:

import re
import lxml.html

def strip_markdown(x):
    links_sub = re.sub(r'\[(.+)\]\([^\)]+\)', r'\1', x)
    bold_sub = re.sub(r'\*\*([^*]+)\*\*', r'\1', links_sub)
    emph_sub = re.sub(r'\*([^*]+)\*', r'\1', bold_sub)
    return emph_sub

def strip_html(x):
    return lxml.html.fromstring(x).text_content() if x else ''

然后

readme = """<img src="https://raw.githubusercontent.com/kootenpv/sky/master/resources/skylogo.png" />

            sky is a web scraping framework, implemented with the latest python versions in mind (3.4+). 
            It uses the asynchronous `asyncio` framework, as well as many popular modules 
            and extensions.

            Most importantly, it aims for **next generation** web crawling where machine intelligence 
            is used to speed up the development/maintainance/reliability of crawling.

            It mainly does this by considering the user to be interested in content 
            from *domains*, not just a collection of *single pages*
            ([templating approach](#templating-approach))."""

strip_markdown(strip_html(readme))

正确移除所有markdown和html。

2016-01-02 10:10:52

您可以使用不同的HTML解析器(如lxml或Beautiful Soup)——它提供只提取文本的函数。或者，您可以在行字符串上运行一个regex来删除标记。请参阅Python文档了解更多信息。

2009-04-15 18:31:01

在Python中从字符串中剥离HTML

推荐文章

最新文章

标签