将非ascii字符替换为单个空格

我需要用空格替换所有非ascii (\x00-\x7F)字符。我很惊讶，这在Python中不是非常容易的，除非我遗漏了什么。下面的函数简单地删除所有非ascii字符:

def remove_non_ascii_1(text):

    return ''.join(i for i in text if ord(i)<128)

这一个替换非ascii字符与空格的数量在字符编码点的字节数(即-字符替换为3个空格):

def remove_non_ascii_2(text):

    return re.sub(r'[^\x00-\x7F]',' ', text)

如何用一个空格替换所有非ascii字符?

在无数类似的SO问题中，没有一个是针对字符替换而不是剥离的，另外是针对所有非ascii字符而不是特定字符。

当前回答

可能是为了一个不同的问题，但我提供了@Alvero的答案(使用unidecode)的我的版本。我想在我的字符串上做一个“常规”剥离，即空白字符的字符串的开始和结束，然后用“常规”空间替换仅其他空白字符，即。

"Ceñíaㅤmañanaㅤㅤㅤㅤ"

"Ceñía mañana"

def safely_stripped(s: str):
    return ' '.join(
        stripped for stripped in
        (bit.strip() for bit in
         ''.join((c if unidecode(c) else ' ') for c in s).strip().split())
        if stripped)

我们首先将所有非unicode空格替换为常规空格(并再次将其连接)，

''.join((c if unidecode(c) else ' ') for c in s)

然后我们再次用python的正常拆分，剥离每个“位”，

(bit.strip() for bit in s.split())

最后将它们重新连接起来，但前提是字符串通过了if测试，

' '.join(stripped for stripped in s if stripped)

correctly returns ' Cenia manana '。

2019-04-08 15:03:03

其他回答

当我们使用ascii()时，它会转义非ascii字符，并且不会正确地更改ascii字符。所以我的主要想法是，它不会改变ASCII字符，所以我在字符串中迭代并检查字符是否被改变。如果它改变了，那就用替代品替换它。例如:' '(单个空格)或'?(带一个问号)。

def remove(x, replacer):

     for i in x:
        if f"'{i}'" == ascii(i):
            pass
        else:
            x=x.replace(i,replacer)
     return x
remove('hái',' ')

结果:“h i”(中间有一个空格)。

语法:remove(str,non_ascii_replacer) 这里你将给出你想要使用的字符串。 non_ascii_replacer =在这里，您将给出要替换所有非ASCII字符的替换器。

2020-12-22 08:48:46

你的" .join()表达式正在过滤，删除任何非ascii的东西;你可以使用条件表达式:

return ''.join([i if ord(i) < 128 else ' ' for i in text])

这将逐个处理字符，并且仍然会使用一个空格替换每个字符。

你的正则表达式应该用空格替换连续的非ascii字符:

re.sub(r'[^\x00-\x7F]+',' ', text)

注意这里的+。

2013-11-19 18:11:35

将所有非ascii (\x00-\x7F)字符替换为空格:

''.join(map(lambda x: x if ord(x) in range(0, 128) else ' ', text))

要替换所有可见字符，请尝试以下操作:

import string

''.join(map(lambda x: x if x in string.printable and x not in string.whitespace else ' ', text))

这将给出相同的结果:

''.join(map(lambda x: x if ord(x) in range(32, 128) else ' ', text))

2021-12-06 21:01:54

这个怎么样?

def replace_trash(unicode_string):
     for i in range(0, len(unicode_string)):
         try:
             unicode_string[i].encode("ascii")
         except:
              #means it's non-ASCII
              unicode_string=unicode_string[i].replace(" ") #replacing it with a single space
     return unicode_string

2016-08-20 22:35:18

作为一种原生且高效的方法，您不需要使用ord或任何字符循环。用ascii码编码，忽略错误。

下面只会删除非ascii字符:

new_string = old_string.encode('ascii',errors='ignore')

现在，如果你想替换被删除的字符，只需执行以下操作:

final_string = new_string + b' ' * (len(old_string) - len(new_string))

2018-01-23 14:39:32

将非ascii字符替换为单个空格

推荐文章

最新文章

标签