如何将一个大的文本文件分割成具有相等行数的小文件?

我有一个大的(按行数)纯文本文件，我想把它分成更小的文件，也按行数。因此，如果我的文件有大约2M行，我想把它分成10个包含200k行的文件，或100个包含20k行的文件(加上一个文件;是否能被均匀整除并不重要)。

我可以在Python中相当容易地做到这一点，但我想知道是否有任何一种忍者方法来使用Bash和Unix实用程序(而不是手动循环和计数/分区行)。

当前回答

split(来自GNU coreutils，从2010-12-22版本8.8开始)包含以下参数:

-n, --number=CHUNKS     generate CHUNKS output files; see explanation below

CHUNKS may be:
  N       split into N files based on size of input
  K/N     output Kth of N to stdout
  l/N     split into N files without splitting lines/records
  l/K/N   output Kth of N to stdout without splitting lines/records
  r/N     like 'l' but use round robin distribution
  r/K/N   likewise but only output Kth of N to stdout

因此，split -n 4输入输出。将生成四个具有相同字节数的文件(output.a{a,b,c,d})，但行可能在中间被打断。

如果我们想保留完整的行(即按行分割)，那么这应该是有效的:

split -n l/4 input output.

相关答案:https://stackoverflow.com/a/19031247

2018-05-30 13:49:16

其他回答

HDFS getmerge小文件并分割成合适的大小。

这个方法会导致换行:

split -b 125m compact.file -d -a 3 compact_prefix

我尝试为每个文件getmerge和分割成大约128 MB。

# Split into 128 MB, and judge sizeunit is M or G. Please test before use.

begainsize=`hdfs dfs -du -s -h /externaldata/$table_name/$date/ | awk '{ print $1}' `
sizeunit=`hdfs dfs -du -s -h /externaldata/$table_name/$date/ | awk '{ print $2}' `
if [ $sizeunit = "G" ];then
    res=$(printf "%.f" `echo "scale=5;$begainsize*8 "|bc`)
else
    res=$(printf "%.f" `echo "scale=5;$begainsize/128 "|bc`)  # Celling ref http://blog.csdn.net/naiveloafer/article/details/8783518
fi
echo $res
# Split into $res files with a number suffix. Ref:  http://blog.csdn.net/microzone/article/details/52839598
compact_file_name=$compact_file"_"
echo "compact_file_name: "$compact_file_name
split -n l/$res $basedir/$compact_file -d -a 3 $basedir/${compact_file_name}

2017-11-08 07:43:09

看看split命令:

$ split --help
Usage: split [OPTION] [INPUT [PREFIX]]
Output fixed-size pieces of INPUT to PREFIXaa, PREFIXab, ...; default
size is 1000 lines, and default PREFIX is `x'.  With no INPUT, or when INPUT
is -, read standard input.

Mandatory arguments to long options are mandatory for short options too.
  -a, --suffix-length=N   use suffixes of length N (default 2)
  -b, --bytes=SIZE        put SIZE bytes per output file
  -C, --line-bytes=SIZE   put at most SIZE bytes of lines per output file
  -d, --numeric-suffixes  use numeric suffixes instead of alphabetic
  -l, --lines=NUMBER      put NUMBER lines per output file
      --verbose           print a diagnostic to standard error just
                            before each output file is opened
      --help     display this help and exit
      --version  output version information and exit

你可以这样做:

split -l 200000 filename

它将创建文件，每个文件有200000行，命名为xaa xab xac…

另一个选项，按输出文件的大小分割(仍然在换行符上分割):

 split -C 20m --numeric-suffixes input_filename output_prefix

创建类似output_prefix01 output_prefix02 output_prefix03…每个最大大小为20兆字节。

2010-01-06 22:44:37

这里有一个例子，把文件“toSplit.txt”分成200行小文件“splited00.txt”，splited01.txt，…"splited25.txt"…

split -l 200——numeric-suffix——additional-suffix=".txt

2022-12-17 07:21:55

使用分割:

将文件分割为固定大小的片段，创建包含连续INPUT部分的输出文件(如果没有给定或INPUT为' -'则为标准输入)

语法split [options] [INPUT [PREFIX]]

2010-01-06 22:44:55