将Unicode字符串转换为Python中的字符串(包含额外符号)

如何将Unicode字符串(包含额外的字符，如£$等)转换为Python字符串?

当前回答

在我的例子中，没有答案，因为我有一个包含unicode字符的字符串变量，这里解释的编码-解码都不起作用。

如果我在终点站做

echo "no me llama mucho la atenci\u00f3n"

python3
>>> print("no me llama mucho la atenci\u00f3n")

输出是正确的:

output: no me llama mucho la atención

但是使用脚本加载这个字符串变量不起作用。

我的案子就是这么办的，说不定能帮到谁

string_to_convert = "no me llama mucho la atenci\u00f3n"
print(json.dumps(json.loads(r'"%s"' % string_to_convert), ensure_ascii=False))
output: no me llama mucho la atención

2019-11-05 20:40:38

其他回答

看到unicodedata.normalize

title = u"Klüft skräms inför på fédéral électoral große"
import unicodedata
unicodedata.normalize('NFKD', title).encode('ascii', 'ignore')
'Kluft skrams infor pa federal electoral groe'

2009-07-30 15:44:32

>>> text=u'abcd'
>>> str(text)
'abcd'

如果字符串只包含ascii字符。

2012-10-25 16:27:20

在我的例子中，没有答案，因为我有一个包含unicode字符的字符串变量，这里解释的编码-解码都不起作用。

如果我在终点站做

echo "no me llama mucho la atenci\u00f3n"

python3
>>> print("no me llama mucho la atenci\u00f3n")

输出是正确的:

output: no me llama mucho la atención

但是使用脚本加载这个字符串变量不起作用。

我的案子就是这么办的，说不定能帮到谁

string_to_convert = "no me llama mucho la atenci\u00f3n"
print(json.dumps(json.loads(r'"%s"' % string_to_convert), ensure_ascii=False))
output: no me llama mucho la atención

2019-11-05 20:40:38

如果您有一个Unicode字符串，并且希望将其写入文件或其他序列化形式，则必须首先将其编码为可存储的特定表示形式。有几种常见的Unicode编码，例如UTF-16(大多数Unicode字符使用两个字节)或UTF-8(1-4字节/码点取决于字符)，等等。要将该字符串转换为特定的编码，您可以使用:

>>> s= u'£10'
>>> s.encode('utf8')
'\xc2\x9c10'
>>> s.encode('utf16')
'\xff\xfe\x9c\x001\x000\x00'

可以将这个原始字节字符串写入文件。但是，请注意，当读取它时，您必须知道它是什么编码，并使用相同的编码进行解码。

当写入文件时，您可以使用codecs模块来摆脱这个手动编码/解码过程。因此，要打开一个将所有Unicode字符串编码为UTF-8的文件，请使用:

import codecs
f = codecs.open('path/to/file.txt','w','utf8')
f.write(my_unicode_string)  # Stored on disk as UTF-8

请注意，使用这些文件的任何其他程序如果想读取这些文件，就必须了解文件的编码。如果你是唯一一个读/写的人，这不是问题，否则请确保你写的是一种其他使用文件的人都能理解的形式。

在Python 3中，这种形式的文件访问是默认的，内置的open函数将接受编码参数，并始终将以文本模式打开的文件转换为Unicode字符串(Python 3中的默认字符串对象)。

2009-07-30 16:44:54

下面是一个示例代码

import unicodedata    
raw_text = u"here $%6757 dfgdfg"
convert_text = unicodedata.normalize('NFKD', raw_text).encode('ascii','ignore')

2016-12-19 07:59:44

将Unicode字符串转换为Python中的字符串(包含额外符号)

推荐文章

最新文章

标签