非贪婪(不情愿)正则匹配sed?

我试图使用sed来清理url行来提取域。

所以从:

http://www.suepearson.co.uk/product/174/71/3816/

我想要:

http://www.suepearson.co.uk/

(不管后面有没有斜杠，都没有关系)

我试过:

 sed 's|\(http:\/\/.*?\/\).*|\1|'

And(转义非贪婪量词)

sed 's|\(http:\/\/.*\?\/\).*|\1|'

但我似乎不能让非贪婪量词(?)工作，所以它总是匹配整个字符串。

当前回答

非贪婪的解决方案超过一个字符

这个帖子真的很老了，但我认为人们仍然需要它。让我们假设你想杀死所有直到HELLO第一次出现的东西。你不能说[^HELLO]…

因此，一个很好的解决方案包括两个步骤，假设您可以在输入中留出一个您不期望的惟一单词，例如top_secit。

在这种情况下，我们可以:

s/HELLO/top_sekrit/     #will only replace the very first occurrence
s/.*top_sekrit//        #kill everything till end of the first HELLO

当然，对于一个简单的输入，你可以使用一个更小的单词，甚至可能是一个字符。

HTH!

2013-10-30 13:05:53

其他回答

在sed中模拟惰性(非贪婪)量词

以及所有其他正则表达式口味!

Finding first occurrence of an expression: POSIX ERE (using -r option) Regex: (EXPRESSION).*|. Sed: sed -r ‍'s/(EXPRESSION).*|./\1/g' # Global `g` modifier should be on Example (finding first sequence of digits) Live demo: $ sed -r 's/([0-9]+).*|./\1/g' <<< 'foo 12 bar 34' 12 How does it work? This regex benefits from an alternation |. At each position engine tries to pick the longest match (this is a POSIX standard which is followed by couple of other engines as well) which means it goes with . until a match is found for ([0-9]+).*. But order is important too. Since global flag is set, engine tries to continue matching character by character up to the end of input string or our target. As soon as the first and only capturing group of left side of alternation is matched (EXPRESSION) rest of line is consumed immediately as well .*. We now hold our value in the first capturing group. POSIX BRE Regex: $\(\(EXPRESSION$.*\)*.\)* Sed: sed 's/$\(\(EXPRESSION$.*\)*.\)*/\3/' Example (finding first sequence of digits): $ sed 's/$\(\([0-9]\{1,\}$.*\)*.\)*/\3/' <<< 'foo 12 bar 34' 12 This one is like ERE version but with no alternation involved. That's all. At each single position engine tries to match a digit. If it is found, other following digits are consumed and captured and the rest of line is matched immediately otherwise since * means more or zero it skips over second capturing group $\([0-9]\{1,\}$.*\)* and arrives at a dot . to match a single character and this process continues. Finding first occurrence of a delimited expression: This approach will match the very first occurrence of a string that is delimited. We can call it a block of string. sed 's/$END-DELIMITER-EXPRESSION$.*/\1/; \ s/$\(START-DELIMITER-EXPRESSION.*$*.\)*/\1/g' Input string: foobar start block #1 end barfoo start block #2 end -EDE: end -SDE: start $ sed 's/$end$.*/\1/; s/$\(start.*$*.\)*/\1/g' Output: start block #1 end First regex $end$.* matches and captures first end delimiter end and substitues all match with recent captured characters which is the end delimiter. At this stage our output is: foobar start block #1 end. Then the result is passed to second regex $\(start.*$*.\)* that is same as POSIX BRE version above. It matches a single character if start delimiter start is not matched otherwise it matches and captures the start delimiter and matches the rest of characters.

直接回答你的问题

使用方法#2(带分隔符的表达式)，你应该选择两个合适的表达式:

艾德:[^]\ / SDE: http:

用法:

$ sed 's/\([^:/]\/\).*/\1/g; s/\(\(http:.*\)*.\)*/\1/' <<< 'http://www.suepearson.co.uk/product/174/71/3816/'

输出:

http://www.suepearson.co.uk/

注意:对于相同的分隔符，这将不起作用。

2016-09-28 16:26:21

基本的和扩展的Posix/GNU regex都不能识别非贪婪的量词;你需要稍后的正则表达式。幸运的是，这个上下文的Perl regex非常容易获得:

perl -pe 's|(http://.*?/).*|\1|'

2009-07-09 10:58:23

不幸的是，如前所述，sed不支持这一点。为了克服这个问题，我建议使用次优方法(实际上甚至更好)，使用类似vim sed的功能。

在.bash-profile中定义

vimdo() { vim $2 --not-a-term -c "$1"  -es +"w >> /dev/stdout" -cq!  ; }

这将创建无头vim来执行命令。

现在你可以这样做:

回声路径美元| vimdo“% s_ \ c: [a-zA-Z0-9 \ \ /] python (a-zA-Z0-9 \ {-} \\/]\{-}:__ g”,

过滤掉$PATH中的python。

使用-在vimdo中从管道中输入。

而大多数语法是相同的。Vim具有更高级的特性，并且使用\{-}是非贪婪匹配的标准。参见帮助regexp。

2022-01-03 00:09:28

另一种方法，不使用正则表达式，是使用字段/分隔符方法，如

string="http://www.suepearson.co.uk/product/174/71/3816/"
echo $string | awk -F"/" '{print $1,$2,$3}' OFS="/"

2009-07-09 10:59:12

@Daniel H(关于你对andcoz的回答的评论，虽然是很久以前的事了):删除后面的零

s,([[:digit:]]\.[[:digit:]]*[1-9])[0]*$,\1,g

这是关于清楚地定义匹配条件……

2020-07-27 13:34:43

非贪婪(不情愿)正则匹配sed?

推荐文章

最新文章

标签