如何将一个大的文本文件分割成具有相等行数的小文件?

我有一个大的(按行数)纯文本文件，我想把它分成更小的文件，也按行数。因此，如果我的文件有大约2M行，我想把它分成10个包含200k行的文件，或100个包含20k行的文件(加上一个文件;是否能被均匀整除并不重要)。

我可以在Python中相当容易地做到这一点，但我想知道是否有任何一种忍者方法来使用Bash和Unix实用程序(而不是手动循环和计数/分区行)。

当前回答

将一个大的文本文件分成1000行的小文件:

分裂<文件> -l 1000

使用实例将一个较大的二进制文件分割为多个10M大小的小文件。

split <file> -b

将多个文件合并为一个文件:

Cat x* > <文件>

拆分一个文件，每个拆分有10行(除了最后一个拆分):

Split -l 10文件名

将一个文件拆分为5个文件。文件被分割，使得每个分割都有相同的大小(除了最后一个分割):

Split -n 5文件名

每次分割一个512字节的文件(除了最后一次分割;千字节使用512k，兆字节使用512m):

Split -b 512 filename

拆分文件，每次拆分最多512字节，不断行:

split -C 512 filename

2022-01-24 10:54:49

其他回答

看看split命令:

$ split --help
Usage: split [OPTION] [INPUT [PREFIX]]
Output fixed-size pieces of INPUT to PREFIXaa, PREFIXab, ...; default
size is 1000 lines, and default PREFIX is `x'.  With no INPUT, or when INPUT
is -, read standard input.

Mandatory arguments to long options are mandatory for short options too.
  -a, --suffix-length=N   use suffixes of length N (default 2)
  -b, --bytes=SIZE        put SIZE bytes per output file
  -C, --line-bytes=SIZE   put at most SIZE bytes of lines per output file
  -d, --numeric-suffixes  use numeric suffixes instead of alphabetic
  -l, --lines=NUMBER      put NUMBER lines per output file
      --verbose           print a diagnostic to standard error just
                            before each output file is opened
      --help     display this help and exit
      --version  output version information and exit

你可以这样做:

split -l 200000 filename

它将创建文件，每个文件有200000行，命名为xaa xab xac…

另一个选项，按输出文件的大小分割(仍然在换行符上分割):

 split -C 20m --numeric-suffixes input_filename output_prefix

创建类似output_prefix01 output_prefix02 output_prefix03…每个最大大小为20兆字节。

2010-01-06 22:44:37

将一个大的文本文件分成1000行的小文件:

分裂<文件> -l 1000

使用实例将一个较大的二进制文件分割为多个10M大小的小文件。

split <file> -b

将多个文件合并为一个文件:

Cat x* > <文件>

拆分一个文件，每个拆分有10行(除了最后一个拆分):

Split -l 10文件名

将一个文件拆分为5个文件。文件被分割，使得每个分割都有相同的大小(除了最后一个分割):

Split -n 5文件名

每次分割一个512字节的文件(除了最后一次分割;千字节使用512k，兆字节使用512m):

Split -b 512 filename

拆分文件，每次拆分最多512字节，不断行:

split -C 512 filename

2022-01-24 10:54:49

使用split命令:

split -l 200000 mybigfile.txt

2010-01-06 22:45:05

Use:

sed -n '1,100p' filename > output.txt

这里，1和100是您将在output.txt中捕获的行号。

2016-04-21 09:27:07

是的，有一个拆分命令。它将按行或字节分割文件。

$ split --help
Usage: split [OPTION]... [INPUT [PREFIX]]
Output fixed-size pieces of INPUT to PREFIXaa, PREFIXab, ...; default
size is 1000 lines, and default PREFIX is `x'.  With no INPUT, or when INPUT
is -, read standard input.

Mandatory arguments to long options are mandatory for short options too.
  -a, --suffix-length=N   use suffixes of length N (default 2)
  -b, --bytes=SIZE        put SIZE bytes per output file
  -C, --line-bytes=SIZE   put at most SIZE bytes of lines per output file
  -d, --numeric-suffixes  use numeric suffixes instead of alphabetic
  -l, --lines=NUMBER      put NUMBER lines per output file
      --verbose           print a diagnostic just before each
                            output file is opened
      --help     display this help and exit
      --version  output version information and exit

SIZE may have a multiplier suffix:
b 512, kB 1000, K 1024, MB 1000*1000, M 1024*1024,
GB 1000*1000*1000, G 1024*1024*1024, and so on for T, P, E, Z, Y.

2010-01-06 22:46:46

如何将一个大的文本文件分割成具有相等行数的小文件?

推荐文章

最新文章

标签