如何使用grep跨多行找到模式?

我想找到有“abc”和“efg”的文件，这两个字符串在该文件中的不同行。一个包含以下内容的文件:

blah blah..
blah blah..
blah abc blah
blah blah..
blah blah..
blah blah..
blah efg blah blah
blah blah..
blah blah..

应该匹配。

当前回答

Grep是这种操作的笨拙工具。

在大多数现代Linux系统中都可以找到pcregrep，可以用作

pcregrep -M  'abc.*(\n|.)*efg' test.txt

where -M，——multiline允许模式匹配多行

还有一个更新的pcre2grep。两者都是由PCRE项目提供的。

pcre2grep可以通过Mac Ports作为pcre2端口的一部分用于Mac OS X:

% sudo port install pcre2

并通过Homebrew为:

% brew install pcre

或者pcre2

% brew install pcre2

pcre2grep在Linux (Ubuntu 18.04+)上也可用

$ sudo apt install pcre2-utils # PCRE2
$ sudo apt install pcregrep    # Older PCRE

2010-04-21 21:29:25

其他回答

如果可以使用Perl，就可以很容易地做到这一点。

perl -ne 'if (/abc/) { $abc = 1; next }; print "Found in $ARGV\n" if ($abc && /efg/); }' yourfilename.txt

您也可以使用单个正则表达式来实现这一点，但这涉及到将文件的整个内容放入单个字符串中，对于大型文件，这可能会占用太多内存。为了完整起见，下面是该方法:

perl -e '@lines = <>; $content = join("", @lines); print "Found in $ARGV\n" if ($content =~ /abc.*efg/s);' yourfilename.txt

2010-04-21 20:36:10

Grep是这种操作的笨拙工具。

在大多数现代Linux系统中都可以找到pcregrep，可以用作

pcregrep -M  'abc.*(\n|.)*efg' test.txt

where -M，——multiline允许模式匹配多行

还有一个更新的pcre2grep。两者都是由PCRE项目提供的。

pcre2grep可以通过Mac Ports作为pcre2端口的一部分用于Mac OS X:

% sudo port install pcre2

并通过Homebrew为:

% brew install pcre

或者pcre2

% brew install pcre2

pcre2grep在Linux (Ubuntu 18.04+)上也可用

$ sudo apt install pcre2-utils # PCRE2
$ sudo apt install pcregrep    # Older PCRE

2010-04-21 21:29:25

我非常依赖于pcregrep，但是对于更新的grep，您不需要安装它的许多特性。只需使用grep -P。

在OP的问题的例子中，我认为以下选项很好地发挥了作用，第二好的选项符合我对问题的理解:

grep -Pzo "abc(.|\n)*efg" /tmp/tes*
grep -Pzl "abc(.|\n)*efg" /tmp/tes*

我将文本复制为/tmp/test1，删除'g'并保存为/tmp/test2。下面的输出显示，第一个显示匹配的字符串，第二个只显示文件名(典型的-o显示匹配，典型的-l只显示文件名)。请注意，'z'对于多行是必要的，'(.|\n)'意味着匹配'换行符以外的任何内容'或'换行符' -即任何内容:

user@host:~$ grep -Pzo "abc(.|\n)*efg" /tmp/tes*
/tmp/test1:abc blah
blah blah..
blah blah..
blah blah..
blah efg
user@host:~$ grep -Pzl "abc(.|\n)*efg" /tmp/tes*
/tmp/test1

要确定你的版本是否足够新，运行man grep，看看顶部是否出现类似的内容:

   -P, --perl-regexp
          Interpret  PATTERN  as a Perl regular expression (PCRE, see
          below).  This is highly experimental and grep -P may warn of
          unimplemented features.

它来自GNU grep 2.10。

2015-10-29 15:27:51

你至少有几个选择

DOTALL方法

用(?s) DOTALL the。包含\n的字符你也可以使用一个超前(?=\n)——不会在匹配中被捕获

example-text:

true
match me

false
match me one

false
match me two

true
match me three
third line!!
{BLANK_LINE}

命令:

grep -Pozi '(?s)true.+?\n(?=\n)' example-text

-p用于perl正则表达式

-o只匹配模式，而不是整行

-z允许换行

-i不区分大小写

输出:

true                                                  
match me                                              
true                                                  
match me three                                        
third line!!

注:

- +? makes modifier non-greedy so matches shortest string instead of largest (prevents from returning one match containing entire text)

你可以使用老式的O.G.手动方法，使用\n

命令:

grep -Pozi 'true(.|\n)+?\n(?=\n)'

输出:

true                                                  
match me                                              
true                                                  
match me three                                        
third line!!

2021-08-19 19:22:48

用银搜索器:

ag 'abc.*(\n|.)*efg' your_filename

与戒指持有者的答案相似，但用ag代替。银色搜索者的速度优势可能在这里大放异彩。

2015-01-13 21:04:10

如何使用grep跨多行找到模式?

推荐文章

最新文章

标签