什么时候汇编比C快?

了解汇编程序的原因之一是，有时可以使用汇编程序来编写比用高级语言(特别是C语言)编写的代码性能更好的代码。然而，我也听人说过很多次，尽管这并非完全错误，但实际上可以使用汇编程序来生成性能更好的代码的情况极其罕见，并且需要汇编方面的专业知识和经验。

这个问题甚至没有涉及到这样一个事实，即汇编程序指令将是特定于机器的、不可移植的，或者汇编程序的任何其他方面。当然，除了这一点之外，了解汇编还有很多很好的理由，但这是一个需要示例和数据的具体问题，而不是关于汇编程序与高级语言的扩展论述。

谁能提供一些具体的例子，说明使用现代编译器汇编代码比编写良好的C代码更快，并且您能否用分析证据支持这一说法?我相信这些案例确实存在，但我真的很想知道这些案例到底有多深奥，因为这似乎是一个有争议的问题。

当前回答

只有在使用编译器不支持的特殊用途指令集时。

为了最大限度地利用具有多个管道和预测分支的现代CPU的计算能力，您需要以这样一种方式来构造汇编程序:a)人类几乎不可能编写b)甚至更不可能维护。

此外，更好的算法、数据结构和内存管理将为您提供至少一个数量级的性能，而不是在汇编中进行的微观优化。

2009-02-23 13:11:37

其他回答

GCC已经成为广泛使用的编译器。它的优化通常不是很好。比编写汇编程序的普通程序员好得多，但就实际性能而言，并没有那么好。有些编译器产生的代码简直令人难以置信。所以一般来说，有很多地方你可以进入编译器的输出，调整汇编器的性能，和/或简单地从头重写例程。

2009-05-24 15:14:32

http://cr.yp.to/qhasm.html有很多例子。

2009-02-23 16:27:11

尽管C语言“接近”于对8位、16位、32位和64位数据的低级操作，但仍有一些C语言不支持的数学操作通常可以在某些汇编指令集中优雅地执行:

Fixed-point multiplication: The product of two 16-bit numbers is a 32-bit number. But the rules in C says that the product of two 16-bit numbers is a 16-bit number, and the product of two 32-bit numbers is a 32-bit number -- the bottom half in both cases. If you want the top half of a 16x16 multiply or a 32x32 multiply, you have to play games with the compiler. The general method is to cast to a larger-than-necessary bit width, multiply, shift down, and cast back: int16_t x, y; // int16_t is a typedef for "short" // set x and y to something int16_t prod = (int16_t)(((int32_t)x*y)>>16);` In this case the compiler may be smart enough to know that you're really just trying to get the top half of a 16x16 multiply and do the right thing with the machine's native 16x16multiply. Or it may be stupid and require a library call to do the 32x32 multiply that's way overkill because you only need 16 bits of the product -- but the C standard doesn't give you any way to express yourself. Certain bitshifting operations (rotation/carries): // 256-bit array shifted right in its entirety: uint8_t x[32]; for (int i = 32; --i > 0; ) { x[i] = (x[i] >> 1) | (x[i-1] << 7); } x[0] >>= 1; This is not too inelegant in C, but again, unless the compiler is smart enough to realize what you are doing, it's going to do a lot of "unnecessary" work. Many assembly instruction sets allow you to rotate or shift left/right with the result in the carry register, so you could accomplish the above in 34 instructions: load a pointer to the beginning of the array, clear the carry, and perform 32 8-bit right-shifts, using auto-increment on the pointer. For another example, there are linear feedback shift registers (LFSR) that are elegantly performed in assembly: Take a chunk of N bits (8, 16, 32, 64, 128, etc), shift the whole thing right by 1 (see above algorithm), then if the resulting carry is 1 then you XOR in a bit pattern that represents the polynomial.

尽管如此，除非有严重的性能限制，否则我不会求助于这些技术。正如其他人所说，汇编代码比C代码更难记录/调试/测试/维护:性能的提高伴随着一些严重的代价。

编辑:3。溢出检测在汇编中是可能的(在C中不能真正做到)，这使得一些算法更容易。

2009-02-23 14:34:56

使用SIMD指令的矩阵操作可能比编译器生成的代码更快。

2009-02-23 13:06:09

第一点不是答案。即使你从来没有用它编程，我发现至少知道一个汇编指令集是有用的。这是程序员永无止境的追求的一部分，他们想知道得更多，从而变得更好。当你进入一个没有源代码的框架时，它也很有用，至少对正在发生的事情有一个粗略的了解。它还可以帮助您理解JavaByteCode和. net IL，因为它们都类似于汇编程序。

To answer the question when you have a small amount of code or a large amount of time. Most useful for use in embedded chips, where low chip complexity and poor competition in compilers targeting these chips can tip the balance in favour of humans. Also for restricted devices you are often trading off code size/memory size/performance in a way that would be hard to instruct a compiler to do. e.g. I know this user action is not called often so I will have small code size and poor performance, but this other function that look similar is used every second so I will have a larger code size and faster performance. That is the sort of trade off a skilled assembly programmer can use.

我还想补充一点，这里有很多中间地带，您可以用C编译代码并检查生成的程序集，然后更改C代码或调整并作为程序集进行维护。

我的朋友从事微控制器的工作，目前是用于控制小型电动机的芯片。他在低级c和汇编的组合中工作。他曾经告诉我，有一天他在工作中把主循环从48条指令减少到43条。他还面临着各种选择，比如代码已经增长到填满256k芯片，业务需要一个新功能，你呢

删除现有功能减少部分或全部现有特性的大小，可能会以性能为代价。提倡改用成本更高、功耗更高、外形更大的更大芯片。

我想补充一点，作为一个商业开发人员，我有很多的投资组合或语言、平台、应用程序类型，我从来没有觉得有必要深入编写程序集。我一直都很感激我所学到的知识。有时会被调试进去。

我知道我已经回答了“为什么我要学习汇编器”这个问题，但我觉得这是一个更重要的问题，而不是什么时候更快。

所以让我们再试一次你应该考虑组装

致力于底层操作系统功能在编译器上工作。工作在一个极其有限的芯片，嵌入式系统等

记住比较你的程序集和生成的编译器，看看哪个更快/更小/更好。

大卫。

2009-02-23 13:44:14

什么时候汇编比C快?

推荐文章

最新文章

标签