在字符集之间转换文本文件的最佳方法?

在字符集之间转换文本文件的最快、最简单的工具或方法是什么?

具体来说，我需要从UTF-8转换为ISO-8859-15，反之亦然。

一切都可以:你最喜欢的脚本语言的一行程序，命令行工具或其他用于操作系统的实用程序，网站等等。

目前为止的最佳解决方案:

在 Linux/UNIX/OS X/cygwin 上：

Gnu iconv suggested by Troels Arvin is best used as a filter. It seems to be universally available. Example: $ iconv -f UTF-8 -t ISO-8859-15 in.txt > out.txt As pointed out by Ben, there is an online converter using iconv. recode (manual) suggested by Cheekysoft will convert one or several files in-place. Example: $ recode UTF8..ISO-8859-15 in.txt This one uses shorter aliases: $ recode utf8..l9 in.txt Recode also supports surfaces which can be used to convert between different line ending types and encodings: Convert newlines from LF (Unix) to CR-LF (DOS): $ recode ../CR-LF in.txt Base64 encode file: $ recode ../Base64 in.txt You can also combine them. Convert a Base64 encoded UTF8 file with Unix line endings to Base64 encoded Latin 1 file with Dos line endings: $ recode utf8/Base64..l1/CR-LF/Base64 file.txt

在Windows Powershell (Jay Bazuzi)上:

PS C:\> gc - zh utf8 in.txt | out - zh ascii out.txt

(但是没有ISO-8859-15支持;它说支持的字符集是unicode, utf7, utf8, utf32, ascii, bigendianunicode, default和oem。)

Edit

你是指iso-8859-1支持吗?使用"String"可以做到这一点，反之亦然

gc -en string in.txt | Out-File -en utf8 out.txt

注意:可能的枚举值是“Unknown, String, Unicode, Byte, BigEndianUnicode, UTF8, UTF7, Ascii”。

CsCvt - Kalytta的字符集转换器是另一个伟大的基于命令行的Windows转换工具。

当前回答

联机使用find，具有自动字符集检测功能

自动检测所有匹配文本文件的字符编码，并将所有匹配文本文件转换为utf-8编码:

$ find . -type f -iname *.txt -exec sh -c 'iconv -f $(file -bi "$1" |sed -e "s/.*[ ]charset=//") -t utf-8 -o converted "$1" && mv converted "$1"' -- {} \;

要执行这些步骤，sub shell sh和-exec一起使用，运行带有-c标志的一行程序，并使用——{}将文件名作为位置参数“$1”传递。在这两者之间，utf-8输出文件临时命名为convert。

file -bi表示:

- b,短暂的不要在输出行前加上文件名(简单模式)。我,mime 导致文件命令输出mime类型字符串，而不是更传统的人类可读字符串。例如，它可以说text/plain;charset=us-ascii而不是ASCII文本。sed命令按照iconv的要求将其仅切割为us-ascii。

find命令对于这样的文件管理自动化非常有用。点击这里获取更多信息。

2016-08-28 19:46:57

其他回答

还有一个转换文件编码的网络工具:https://webtool.cloud/change-file-encoding

它支持广泛的编码，包括一些罕见的编码，如IBM代码页37。

2020-08-18 09:34:35

使用这个Python脚本:https://github.com/goerz/convert_encoding.py 适用于任何平台。需要Python 2.7。

2018-07-01 10:17:32

独立实用程序方法

iconv -f ISO-8859-1 -t UTF-8 in.txt > out.txt

-f ENCODING  the encoding of the input
-t ENCODING  the encoding of the output

您不必指定这两个参数中的任何一个。它们将默认使用您的当前语言环境，通常是UTF-8。

2008-09-15 17:24:23

如“如何纠正文件的字符编码?”Synalyze它!可以让你在OS X上轻松转换ICU库支持的所有编码。

此外，您还可以显示从所有编码转换为Unicode的文件的一些字节，以便快速查看哪个字节适合您的文件。

2013-06-26 19:42:37

尝试VIM

如果你有vim，你可以使用这个:

没有对每种编码进行测试。

最酷的部分是你不需要知道源编码

vim +"set nobomb | set fenc=utf8 | x" filename.txt

注意，这个命令直接修改文件

解释部分!

+: vim打开文件时直接输入命令。通常用于在特定行打开文件:vim +14 file.txt |:多个命令的分隔符(如;在bash中) set nobomb:没有utf-8 BOM set fenc=utf8:设置新的编码为utf-8 doc link x:保存并关闭文件 Filename.txt:文件的路径 :这里的报价是因为管道。(否则bash将使用它们作为bash管道)

2015-09-30 08:41:28

在字符集之间转换文本文件的最佳方法?

推荐文章

最新文章

标签