从文件中获取第n行的Bash工具

有没有一种“规范”的方法？我一直在使用head-n|tail-1来完成这个任务，但我一直在想是否有一个Bash工具专门从文件中提取一行（或一系列行）。

我所说的“规范”是指一个主要功能就是这样做的程序。

当前回答

我将上面的一些答案放入了一个简短的bash脚本中，您可以将其放入名为get.sh的文件中，并链接到/usr/local/bin/get（或您喜欢的任何其他名称）。

#!/bin/bash
if [ "${1}" == "" ]; then
    echo "error: blank line number";
    exit 1
fi
re='^[0-9]+$'
if ! [[ $1 =~ $re ]] ; then
    echo "error: line number arg not a number";
    exit 1
fi
if [ "${2}" == "" ]; then
    echo "error: blank file name";
    exit 1
fi
sed "${1}q;d" $2;
exit 0

确保它可通过

$ chmod +x get

将其链接以使其在PATH上可用

$ ln -s get.sh /usr/local/bin/get

2019-01-28 15:22:05

其他回答

您也可以使用sed打印并退出：

sed -n '10{p;q;}' file   # print line 10

2011-05-17 11:49:24

已经有很多好答案了。我个人喜欢awk。为了方便起见，如果您使用bash，只需将以下内容添加到~/.bash_profile中即可。下次登录时（或者如果您在本次更新后获取.bash_profile的源代码），您将有一个新的漂亮的“第n”函数可用于管道传输文件。

执行此命令或将其放入~/.bash_profile（如果使用bash）并重新打开bash（或执行源~/.bach_profile）

# print just the nth piped in line
nth () { awk -vlnum=${1} 'NR==lnum {print; exit}'; }

然后，要使用它，只需通过管道。例如：

$ yes line | cat -n | nth 5
     5  line

2017-11-17 15:42:57

有了awk，速度相当快：

awk 'NR == num_line' file

如果为true，则执行awk的默认行为：｛print$0｝。

替代版本

如果您的文件恰好很大，最好在读取所需的行后退出。这样可以节省CPU时间请参见答案末尾的时间比较。

awk 'NR == num_line {print; exit}' file

如果要从bash变量中给出行号，可以使用：

awk 'NR == n' n=$num file
awk -v n=$num 'NR == n' file   # equivalent

查看使用exit节省了多少时间，特别是如果该行恰好位于文件的第一部分：

# Let's create a 10M lines file
for ((i=0; i<100000; i++)); do echo "bla bla"; done > 100Klines
for ((i=0; i<100; i++)); do cat 100Klines; done > 10Mlines

$ time awk 'NR == 1234567 {print}' 10Mlines
bla bla

real    0m1.303s
user    0m1.246s
sys 0m0.042s
$ time awk 'NR == 1234567 {print; exit}' 10Mlines
bla bla

real    0m0.198s
user    0m0.178s
sys 0m0.013s

因此，两者的差异是0.198秒对1.303秒，大约快了6倍。

2014-01-22 09:49:02

您也可以使用Perl实现这一点：

perl -wnl -e '$.== NUM && print && exit;' some.file

2011-05-16 19:43:41

根据我的测试，就性能和可读性而言，我的建议是：

尾部-n+n|头部-1

N是您想要的行号。例如，tail-n+7 input.txt | head-1将打印文件的第7行。

tail-n+n将打印从第n行开始的所有内容，head-1将使其在一行之后停止。

可选的head-N|tail-1可能更可读。例如，这将打印第7行：

head-7 input.txt | tail-1

当谈到性能时，较小的文件大小没有太大的差异，但当文件变大时，尾部|头部（从上方）的性能会优于尾部|头部。

排名靠前的是“NUMq；d’很有意思，但我认为，与头/尾解决方案相比，开箱即用的人更少，而且它也比尾/头慢。

在我的测试中，两个尾部/头部版本都优于sed的NUMq；d’一致。这与发布的其他基准一致。很难找到尾巴/脑袋真的很坏的案例。这也不奇怪，因为这些操作在现代Unix系统中会被大量优化。

为了了解性能差异，以下是我从一个巨大文件（9.3G）中得到的数字：

tail-n+n | head-1:3.7秒头-N|尾-1:4.6秒sed Nq；d： 18.8秒

结果可能有所不同，但总体而言，性能头部|尾部和尾部|头部对于较小的输入来说是可比的，sed总是慢了一个重要因素（大约5倍左右）。

要复制我的基准测试，您可以尝试以下操作，但请注意，它将在当前工作目录中创建一个9.3G文件：

#!/bin/bash
readonly file=tmp-input.txt
readonly size=1000000000
readonly pos=500000000
readonly retries=3

seq 1 $size > $file
echo "*** head -N | tail -1 ***"
for i in $(seq 1 $retries) ; do
    time head "-$pos" $file | tail -1
done
echo "-------------------------"
echo
echo "*** tail -n+N | head -1 ***"
echo

seq 1 $size > $file
ls -alhg $file
for i in $(seq 1 $retries) ; do
    time tail -n+$pos $file | head -1
done
echo "-------------------------"
echo
echo "*** sed Nq;d ***"
echo

seq 1 $size > $file
ls -alhg $file
for i in $(seq 1 $retries) ; do
    time sed $pos'q;d' $file
done
/bin/rm $file

这是在我的机器上运行的输出（ThinkPad X1 Carbon，带有SSD和16G内存）。我假设在最后一次运行中，所有内容都将来自缓存，而不是磁盘：

*** head -N | tail -1 ***
500000000

real    0m9,800s
user    0m7,328s
sys     0m4,081s
500000000

real    0m4,231s
user    0m5,415s
sys     0m2,789s
500000000

real    0m4,636s
user    0m5,935s
sys     0m2,684s
-------------------------

*** tail -n+N | head -1 ***

-rw-r--r-- 1 phil 9,3G Jan 19 19:49 tmp-input.txt
500000000

real    0m6,452s
user    0m3,367s
sys     0m1,498s
500000000

real    0m3,890s
user    0m2,921s
sys     0m0,952s
500000000

real    0m3,763s
user    0m3,004s
sys     0m0,760s
-------------------------

*** sed Nq;d ***

-rw-r--r-- 1 phil 9,3G Jan 19 19:50 tmp-input.txt
500000000

real    0m23,675s
user    0m21,557s
sys     0m1,523s
500000000

real    0m20,328s
user    0m18,971s
sys     0m1,308s
500000000

real    0m19,835s
user    0m18,830s
sys     0m1,004s

2017-07-31 13:10:02

从文件中获取第n行的Bash工具

推荐文章

最新文章

标签