如何提取两个标记之间的子字符串?

假设我有一个字符串'gfgfdAAA1234ZZZuijjk'，我想提取'1234'部分。

我只知道在AAA之前的几个字符，以及在ZZZ之后的我感兴趣的部分1234。

使用sed，可以对字符串执行如下操作:

echo "$STRING" | sed -e "s|.*AAA\(.*\)ZZZ.*|\1|"

结果是1234。

如何在Python中做同样的事情?

当前回答

import re
print re.search('AAA(.*?)ZZZ', 'gfgfdAAA1234ZZZuijjk').group(1)

2011-01-12 09:18:00

其他回答

text = 'I want to find a string between two substrings'
left = 'find a '
right = 'between two'

print(text[text.index(left)+len(left):text.index(right)])

给了

string

2019-03-04 01:31:31

此外，您可以在波纹函数中找到所有的组合

s = 'Part 1. Part 2. Part 3 then more text'
def find_all_places(text,word):
    word_places = []
    i=0
    while True:
        word_place = text.find(word,i)
        i+=len(word)+word_place
        if i>=len(text):
            break
        if word_place<0:
            break
        word_places.append(word_place)
    return word_places
def find_all_combination(text,start,end):
    start_places = find_all_places(text,start)
    end_places = find_all_places(text,end)
    combination_list = []
    for start_place in start_places:
        for end_place in end_places:
            print(start_place)
            print(end_place)
            if start_place>=end_place:
                continue
            combination_list.append(text[start_place:end_place])
    return combination_list
find_all_combination(s,"Part","Part")

结果:

['Part 1. ', 'Part 1. Part 2. ', 'Part 2. ']

2021-10-05 19:02:30

在python中，可以使用正则表达式(re)模块中的findall方法从字符串中提取子字符串。

>>> import re
>>> s = 'gfgfdAAA1234ZZZuijjk'
>>> ss = re.findall('AAA(.+)ZZZ', s)
>>> print ss
['1234']

2018-03-14 09:11:23

使用sed，可以对字符串执行如下操作:

echo "$STRING" | sed -e "s|.*AAA$.*$ZZZ.*|\1|"

结果是1234。

你可以使用相同的正则表达式对re.sub函数做同样的事情。

>>> re.sub(r'.*AAA(.*)ZZZ.*', r'\1', 'gfgfdAAA1234ZZZuijjk')
'1234'

在基本sed中，捕获组由$..$表示，但在python中由(..)表示。

2015-01-31 08:29:21

以防有人要做和我一样的事。我必须在一行中提取圆括号内的所有内容。例如，如果我有这样一句话，‘美国总统(巴拉克·奥巴马)会见了……，我只想得到“巴拉克·奥巴马”，这是解决方案:

regex = '.*\((.*?)\).*'
matches = re.search(regex, line)
line = matches.group(1) + '\n'

也就是说，你需要用斜杠\符号来阻止括号。尽管这是一个关于更多正则表达式的问题。

此外，在某些情况下，你可能会在正则表达式定义之前看到'r'符号。如果没有r前缀，你需要像在c中那样使用转义字符。这里有更多关于这个的讨论。

2014-01-19 19:29:00

如何提取两个标记之间的子字符串?

推荐文章

最新文章

标签