提取正则表达式匹配的一部分

我想用正则表达式从HTML页面中提取标题。目前我有这个:

title = re.search('<title>.*</title>', html, re.IGNORECASE).group()
if title:
    title = title.replace('<title>', '').replace('</title>', '')

是否有正则表达式来提取<title>的内容，这样我就不必删除标签了?

当前回答

Krzysztof krasosk目前投票最多的答案是<title>a</title><title>b</title>。此外，它忽略了跨越行边界的标题标签，例如，由于行长原因。最后，它失败于<title >a</title>(这是有效的XML/HTML标记中的空白)。

因此，我提出以下改进建议:

import re

def search_title(html):
    m = re.search(r"<title\s*>(.*?)</title\s*>", html, re.IGNORECASE | re.DOTALL)
    return m.group(1) if m else None

测试用例:

print(search_title("<title   >with spaces in tags</title >"))
print(search_title("<title\n>with newline in tags</title\n>"))
print(search_title("<title>first of two titles</title><title>second title</title>"))
print(search_title("<title>with newline\n in title</title\n>"))

输出:

with spaces in tags
with newline in tags
first of two titles
with newline
  in title

最后，我和其他人一起推荐一个HTML解析器——不仅要处理HTML标记的非标准使用。

2020-10-30 08:17:27

其他回答

re.search('<title>(.*)</title>', s, re.IGNORECASE).group(1)

2009-08-25 10:28:53

请注意，从Python 3.8开始，并引入了赋值表达式(PEP 572)(:=操作符)，可以通过直接在if条件中捕获匹配结果作为变量并在条件体中重用它来改进Krzysztof krasosco的解决方案:

# pattern = '<title>(.*)</title>'
# text = '<title>hello</title>'
if match := re.search(pattern, text, re.IGNORECASE):
  title = match.group(1)
# hello

2019-04-27 15:06:30

我可以向您推荐美丽汤吗?Soup是一个解析所有html文档的很好的库。

soup = BeatifulSoup(html_doc)
titleName = soup.title.name

2013-03-01 19:22:25

因此，我提出以下改进建议:

import re

def search_title(html):
    m = re.search(r"<title\s*>(.*?)</title\s*>", html, re.IGNORECASE | re.DOTALL)
    return m.group(1) if m else None

测试用例:

print(search_title("<title   >with spaces in tags</title >"))
print(search_title("<title\n>with newline in tags</title\n>"))
print(search_title("<title>first of two titles</title><title>second title</title>"))
print(search_title("<title>with newline\n in title</title\n>"))

输出:

with spaces in tags
with newline in tags
first of two titles
with newline
  in title

最后，我和其他人一起推荐一个HTML解析器——不仅要处理HTML标记的非标准使用。

2020-10-30 08:17:27

我需要一些东西来匹配package-0.0.1(名称，版本)，但想拒绝一个无效的版本，如0.0.010。

参见regex101示例。

import re

RE_IDENTIFIER = re.compile(r'^([a-z]+)-((?:(?:0|[1-9](?:[0-9]+)?)\.){2}(?:0|[1-9](?:[0-9]+)?))$')

example = 'hello-0.0.1'

if match := RE_IDENTIFIER.search(example):
    name, version = match.groups()
    print(f'Name:     {name}')
    print(f'Version:  {version}')
else:
    raise ValueError(f'Invalid identifier {example}')

输出:

Name:     hello
Version:  0.0.1

2021-05-20 14:00:38

提取正则表达式匹配的一部分

推荐文章

最新文章

标签