使用多个单词边界分隔符将字符串拆分为单词

我想我想做的是一项相当常见的任务，但我在网上找不到任何参考资料。我有带标点符号的文本，我想要一个单词列表。

"Hey, you - what are you doing here!?"

应该是

['hey', 'you', 'what', 'are', 'you', 'doing', 'here']

但Python的str.split（）只对一个参数有效，所以在用空格拆分后，所有单词都带有标点符号。有什么想法吗？

当前回答

使用替换两次：

a = '11223FROM33344INTO33222FROM3344'
a.replace('FROM', ',,,').replace('INTO', ',,,').split(',,,')

结果是：

['11223', '33344', '33222', '3344']

2012-03-30 13:27:30

其他回答

re.split（）

re.split（模式，字符串[，maxsplit=0]）按模式的出现次数拆分字符串。如果模式中使用了捕获括号，那么模式中所有组的文本也会作为结果列表的一部分返回。如果maxsplit为非零，则最多发生maxsplit拆分，字符串的剩余部分将作为列表的最后一个元素返回。（不兼容注意：在最初的Python1.5版本中，maxsplit被忽略。这在以后的版本中得到了修复。）

>>> re.split('\W+', 'Words, words, words.')
['Words', 'words', 'words', '']
>>> re.split('(\W+)', 'Words, words, words.')
['Words', ', ', 'words', ', ', 'words', '.', '']
>>> re.split('\W+', 'Words, words, words.', 1)
['Words', 'words, words.']

2009-06-29 17:57:49

我正在重新熟悉Python，需要同样的东西。findall解决方案可能更好，但我想到了这个：

tokens = [x.strip() for x in data.split(',')]

2012-04-20 16:53:46

使用panda的series.str.split方法可以获得相同的结果，而不是使用re-module函数re.split。

首先，使用上述字符串创建一个系列，然后将该方法应用于该系列。

thestring=pd.Series（“嘿，你-你在这里干什么！？”）thestring.str.split（pat='，|-'）

参数pat接受分隔符并将拆分字符串作为数组返回。这里，使用|（或运算符）传递两个分隔符。输出如下：

[嘿，你，你在这里干什么！？]

2018-09-10 15:32:40

我认为以下是满足您需求的最佳答案：

\W+可能适合这种情况，但可能不适合其他情况。

filter(None, re.compile('[ |,|\-|!|?]').split( "Hey, you - what are you doing here!?")

2012-03-09 08:30:11

这是一个有一些解释的答案。

st = "Hey, you - what are you doing here!?"

# replace all the non alpha-numeric with space and then join.
new_string = ''.join([x.replace(x, ' ') if not x.isalnum() else x for x in st])
# output of new_string
'Hey  you  what are you doing here  '

# str.split() will remove all the empty string if separator is not provided
new_list = new_string.split()

# output of new_list
['Hey', 'you', 'what', 'are', 'you', 'doing', 'here']

# we can join it to get a complete string without any non alpha-numeric character
' '.join(new_list)
# output
'Hey you what are you doing'

或者在一行中，我们可以这样做：

(''.join([x.replace(x, ' ') if not x.isalnum() else x for x in st])).split()

# output
['Hey', 'you', 'what', 'are', 'you', 'doing', 'here']

更新的答案

2016-06-04 19:35:58

使用多个单词边界分隔符将字符串拆分为单词

推荐文章

最新文章

标签