将非ascii字符替换为单个空格

我需要用空格替换所有非ascii (\x00-\x7F)字符。我很惊讶，这在Python中不是非常容易的，除非我遗漏了什么。下面的函数简单地删除所有非ascii字符:

def remove_non_ascii_1(text):

    return ''.join(i for i in text if ord(i)<128)

这一个替换非ascii字符与空格的数量在字符编码点的字节数(即-字符替换为3个空格):

def remove_non_ascii_2(text):

    return re.sub(r'[^\x00-\x7F]',' ', text)

如何用一个空格替换所有非ascii字符?

在无数类似的SO问题中，没有一个是针对字符替换而不是剥离的，另外是针对所有非ascii字符而不是特定字符。

当前回答

作为一种原生且高效的方法，您不需要使用ord或任何字符循环。用ascii码编码，忽略错误。

下面只会删除非ascii字符:

new_string = old_string.encode('ascii',errors='ignore')

现在，如果你想替换被删除的字符，只需执行以下操作:

final_string = new_string + b' ' * (len(old_string) - len(new_string))

2018-01-23 14:39:32

其他回答

作为一种原生且高效的方法，您不需要使用ord或任何字符循环。用ascii码编码，忽略错误。

下面只会删除非ascii字符:

new_string = old_string.encode('ascii',errors='ignore')

现在，如果你想替换被删除的字符，只需执行以下操作:

final_string = new_string + b' ' * (len(old_string) - len(new_string))

2018-01-23 14:39:32

如果替换字符可以是'?'而不是空格，那么我建议result = text。编码(“ascii”、“替换”).decode ():

"""Test the performance of different non-ASCII replacement methods."""


import re
from timeit import timeit


# 10_000 is typical in the project that I'm working on and most of the text
# is going to be non-ASCII.
text = 'Æ' * 10_000


print(timeit(
    """
result = ''.join([c if ord(c) < 128 else '?' for c in text])
    """,
    number=1000,
    globals=globals(),
))

print(timeit(
    """
result = text.encode('ascii', 'replace').decode()
    """,
    number=1000,
    globals=globals(),
))

结果:

0.7208260721400134
0.009975979187503592

2017-01-03 06:31:18

def filterSpecialChars(strInput):
    result = []
    for character in strInput:
        ordVal = ord(character)
        if ordVal < 0 or ordVal > 127:
            result.append(' ')
        else:
            result.append(character)
    return ''.join(result)

像这样叫它:

result = filterSpecialChars('Ceñía mañana')
print(result)

2022-09-23 14:35:11

你的" .join()表达式正在过滤，删除任何非ascii的东西;你可以使用条件表达式:

return ''.join([i if ord(i) < 128 else ' ' for i in text])

这将逐个处理字符，并且仍然会使用一个空格替换每个字符。

你的正则表达式应该用空格替换连续的非ascii字符:

re.sub(r'[^\x00-\x7F]+',' ', text)

注意这里的+。

2013-11-19 18:11:35

我的问题是我的字符串包含了像BelgiÃ的België和&#x20AC的€符号。我不想把它们换成空格。但要有正确的符号。

我的解决方案是string.encode('Latin1').decode('utf-8')

2021-06-10 10:21:24

将非ascii字符替换为单个空格

推荐文章

最新文章

标签