字符串格式：%vs..format vs.f-String文字

有多种字符串格式设置方法：

Python<2.6:“您好%s”%namePython 2.6+：“Hello｛｝”.format（name）（使用str.format）Python 3.6+：f“｛name｝”（使用f-string）

哪种情况更好？在什么情况下？

以下方法具有相同的结果，那么有什么区别？name=“爱丽丝”“你好%s”%name“您好｛0｝”.format（名称）f“您好｛name｝”#使用命名参数：“您好%（kwarg）s”%｛'kwarg'：name｝“你好｛kwarg｝”.format（kwarg=name）f“您好｛name｝”字符串格式化何时运行，如何避免运行时性能损失？

如果您试图结束一个重复的问题，该问题只是在寻找一种格式化字符串的方法，请使用How do I put a variable value in a string？。

当前回答

但是请注意，刚才我在尝试用现有代码中的.format替换所有%时发现了一个问题：“｛｝”.format（unicode_string）将尝试对unicode_string进行编码，并且可能会失败。

看看这个Python交互式会话日志：

Python 2.7.2 (default, Aug 27 2012, 19:52:55) 
[GCC 4.1.2 20080704 (Red Hat 4.1.2-48)] on linux2
; s='й'
; u=u'й'
; s
'\xd0\xb9'
; u
u'\u0439'

s只是一个字符串（在Python3中称为“byte array”），u是一个Unicode字符串（在Python 3中称“string”）：

; '%s' % s
'\xd0\xb9'
; '%s' % u
u'\u0439'

当您将Unicode对象作为参数提供给%operator时，即使原始字符串不是Unicode，它也会生成Unicode字符串：

; '{}'.format(s)
'\xd0\xb9'
; '{}'.format(u)
Traceback (most recent call last):
  File "<stdin>", line 1, in <module>
UnicodeEncodeError: 'latin-1' codec can't encode character u'\u0439' in position 0: ordinal not in range(256)

但.format函数将引发“UnicodeEncodeError”：

; u'{}'.format(s)
u'\xd0\xb9'
; u'{}'.format(u)
u'\u0439'

并且只有当原始字符串是Unicode时，它才能使用Unicode参数。

; '{}'.format(u'i')
'i'

或者如果参数字符串可以转换为字符串（称为“字节数组”）

2012-09-03 18:15:42

其他回答

回答第一个问题。格式在许多方面似乎更为复杂。关于%的一个令人讨厌的问题是，它可以接受变量或元组。你会认为以下方法总是有效的：

"Hello %s" % name

然而，如果name恰好是（1，2，3），它将抛出一个TypeError。为了保证它总是打印出来，你需要

"Hello %s" % (name,)   # supply the single argument as a single-item tuple

这太难看了。格式没有这些问题。同样在您给出的第二个示例中，.format示例看起来更简洁。

仅用于向后兼容Python 2.5。

为了回答第二个问题，字符串格式化与任何其他操作同时发生-当计算字符串格式化表达式时。Python不是一种惰性语言，它在调用函数之前会对表达式求值，因此表达式log.debug（“somedebuginfo:%s”%some_info）将首先将字符串求值为，例如“somedebug-info:roflcopters is active”，然后将该字符串传递给log.debug（）。

2011-02-22 18:49:21

模运算符（%）无法执行的操作，afaik：

tu = (12,45,22222,103,6)
print '{0} {2} {1} {2} {3} {2} {4} {2}'.format(*tu)

后果

12 22222 45 22222 103 22222 6 22222

非常有用。

另一点：format（）是一个函数，可以用作其他函数中的参数：

li = [12,45,78,784,2,69,1254,4785,984]
print map('the number is {}'.format,li)   

print

from datetime import datetime,timedelta

once_upon_a_time = datetime(2010, 7, 1, 12, 0, 0)
delta = timedelta(days=13, hours=8,  minutes=20)

gen =(once_upon_a_time +x*delta for x in xrange(20))

print '\n'.join(map('{:%Y-%m-%d %H:%M:%S}'.format, gen))

结果如下：

['the number is 12', 'the number is 45', 'the number is 78', 'the number is 784', 'the number is 2', 'the number is 69', 'the number is 1254', 'the number is 4785', 'the number is 984']

2010-07-01 12:00:00
2010-07-14 20:20:00
2010-07-28 04:40:00
2010-08-10 13:00:00
2010-08-23 21:20:00
2010-09-06 05:40:00
2010-09-19 14:00:00
2010-10-02 22:20:00
2010-10-16 06:40:00
2010-10-29 15:00:00
2010-11-11 23:20:00
2010-11-25 07:40:00
2010-12-08 16:00:00
2010-12-22 00:20:00
2011-01-04 08:40:00
2011-01-17 17:00:00
2011-01-31 01:20:00
2011-02-13 09:40:00
2011-02-26 18:00:00
2011-03-12 02:20:00

2011-06-13 20:20:32

顺便说一句，在日志记录中使用新样式的格式并不一定会影响性能。您可以将任何对象传递给实现__str__魔术方法的logging.debug、logging.info等。当日志模块决定必须发出消息对象（无论是什么）时，它会在发出消息之前调用str（message_object）

import logging


class NewStyleLogMessage(object):
    def __init__(self, message, *args, **kwargs):
        self.message = message
        self.args = args
        self.kwargs = kwargs

    def __str__(self):
        args = (i() if callable(i) else i for i in self.args)
        kwargs = dict((k, v() if callable(v) else v) for k, v in self.kwargs.items())

        return self.message.format(*args, **kwargs)

N = NewStyleLogMessage

# Neither one of these messages are formatted (or calculated) until they're
# needed

# Emits "Lazily formatted log entry: 123 foo" in log
logging.debug(N('Lazily formatted log entry: {0} {keyword}', 123, keyword='foo'))


def expensive_func():
    # Do something that takes a long time...
    return 'foo'

# Emits "Expensive log entry: foo" in log
logging.debug(N('Expensive log entry: {keyword}', keyword=expensive_func))

Python 3文档中对此进行了描述(https://docs.python.org/3/howto/logging-cookbook.html#formatting-样式）。但是，它也可以与Python 2.6一起使用(https://docs.python.org/2.6/library/logging.html#using-作为消息的任意对象）。

使用此技术的一个优点是，它允许延迟值，例如上面的函数expensive_func，而不是格式化样式不可知。这为Python文档中给出的建议提供了一个更优雅的替代方案：https://docs.python.org/2.6/library/logging.html#optimization.

2014-08-21 18:00:48

format还有另一个优点（我在答案中没有看到）：它可以获取对象财产。

In [12]: class A(object):
   ....:     def __init__(self, x, y):
   ....:         self.x = x
   ....:         self.y = y
   ....:         

In [13]: a = A(2,3)

In [14]: 'x is {0.x}, y is {0.y}'.format(a)
Out[14]: 'x is 2, y is 3'

或者，作为关键字参数：

In [15]: 'x is {a.x}, y is {a.y}'.format(a=a)
Out[15]: 'x is 2, y is 3'

据我所知，%是不可能的。

2014-12-04 18:33:46

在格式化正则表达式时，%可能会有所帮助。例如

'{type_names} [a-z]{2}'.format(type_names='triangle|square')

引发IndexError。在这种情况下，您可以使用：

'%(type_names)s [a-z]{2}' % {'type_names': 'triangle|square'}

这避免了将正则表达式写成“｛type_names｝[a-z]｛｛2｝｝”。当您有两个正则表达式时，这可能很有用，其中一个正则表达式单独使用而不使用格式，但两者的连接是格式化的。

2015-04-09 20:41:40

字符串格式：%vs..format vs.f-String文字

推荐文章

最新文章

标签