从字符串中删除标点符号的最佳方法

似乎应该有一种比以下更简单的方法：

import string
s = "string. With. Punctuation?" # Sample string 
out = s.translate(string.maketrans("",""), string.punctuation)

有？

当前回答

对于严肃的自然语言处理（NLP），您应该让像SpaCy这样的库通过标记化处理标点符号，然后您可以根据需要手动调整。

例如，您希望如何处理单词中的连字符？例外情况，如缩写？开始和结束引号？URL？在NLP中，将“let’s”这样的收缩分隔为“let”和“s”以进行进一步处理通常很有用。

2022-03-31 01:53:41

其他回答

myString.translate(None, string.punctuation)

2010-03-08 15:19:09

试试那个：）

regex.sub(r'\p{P}','', s)

2020-09-02 07:51:45

字符串标点符号漏掉了现实世界中常用的大量标点符号。一个适用于非ASCII标点符号的解决方案怎么样？

import regex
s = u"string. With. Some・Really Weird、Non？ASCII。 「（Punctuation）」?"
remove = regex.compile(ur'[\p{C}|\p{M}|\p{P}|\p{S}|\p{Z}]+', regex.UNICODE)
remove.sub(u" ", s).strip()

我个人认为，这是在Python中删除字符串标点符号的最佳方法，因为：

它删除所有Unicode标点符号它很容易修改，例如，如果您想删除标点符号，可以删除\｛s｝，但保留像$这样的符号。您可以非常具体地了解要保留的内容和要删除的内容，例如，Pd只会删除破折号。此正则表达式还规范了空白。它将制表符、回车符和其他奇怪的字符映射到漂亮的单个空格。

这使用了Unicode字符财产，您可以在Wikipedia上阅读更多有关该属性的信息。

2016-10-06 16:46:01

从效率的角度来看，你不会击败

s.translate(None, string.punctuation)

对于更高版本的Python，请使用以下代码：

s.translate(str.maketrans('', '', string.punctuation))

它使用查找表在C语言中执行原始字符串操作——除了编写自己的C代码之外，没有什么能比这更好的了。

如果速度不令人担忧，另一个选择是：

exclude = set(string.punctuation)
s = ''.join(ch for ch in s if ch not in exclude)

这比用每个字符替换s.replace更快，但不会像正则表达式或字符串转换等非纯python方法那样执行得好，正如您从下面的计时中看到的那样。对于这种类型的问题，在尽可能低的水平上解决是有回报的。

计时代码：

import re, string, timeit

s = "string. With. Punctuation"
exclude = set(string.punctuation)
table = string.maketrans("","")
regex = re.compile('[%s]' % re.escape(string.punctuation))

def test_set(s):
    return ''.join(ch for ch in s if ch not in exclude)

def test_re(s):  # From Vinko's solution, with fix.
    return regex.sub('', s)

def test_trans(s):
    return s.translate(table, string.punctuation)

def test_repl(s):  # From S.Lott's solution
    for c in string.punctuation:
        s=s.replace(c,"")
    return s

print "sets      :",timeit.Timer('f(s)', 'from __main__ import s,test_set as f').timeit(1000000)
print "regex     :",timeit.Timer('f(s)', 'from __main__ import s,test_re as f').timeit(1000000)
print "translate :",timeit.Timer('f(s)', 'from __main__ import s,test_trans as f').timeit(1000000)
print "replace   :",timeit.Timer('f(s)', 'from __main__ import s,test_repl as f').timeit(1000000)

结果如下：

sets      : 19.8566138744
regex     : 6.86155414581
translate : 2.12455511093
replace   : 28.4436721802

2008-11-05 18:36:11

我还没有看到这个答案。只需使用正则表达式；它删除了除单词字符（\w）和数字字符（\d）之外的所有字符，后跟一个空白字符（\s）：

import re
s = "string. With. Punctuation?" # Sample string 
out = re.sub(ur'[^\w\d\s]+', '', s)

2016-06-18 06:38:57

从字符串中删除标点符号的最佳方法

推荐文章

最新文章

标签