C和c++中联合的目的

我以前很轻松地使用过工会;今天当我读到这篇文章并知道这个代码时，我很震惊

union ARGB
{
    uint32_t colour;

    struct componentsTag
    {
        uint8_t b;
        uint8_t g;
        uint8_t r;
        uint8_t a;
    } components;

} pixel;

pixel.colour = 0xff040201;  // ARGB::colour is the active member from now on

// somewhere down the line, without any edit to pixel

if(pixel.components.a)      // accessing the non-active member ARGB::components

实际上是未定义的行为，即从工会成员中读取除最近写的人以外的内容会导致未定义的行为。如果这不是联合的预期用途，那么什么才是?谁能详细解释一下吗?

更新:

我想事后澄清一些事情。

The answer to the question isn't the same for C and C++; my ignorant younger self tagged it as both C and C++. After scouring through C++11's standard I couldn't conclusively say that it calls out accessing/inspecting a non-active union member is undefined/unspecified/implementation-defined. All I could find was §9.5/1: If a standard-layout union contains several standard-layout structs that share a common initial sequence, and if an object of this standard-layout union type contains one of the standard-layout structs, it is permitted to inspect the common initial sequence of any of standard-layout struct members. §9.2/19: Two standard-layout structs share a common initial sequence if corresponding members have layout-compatible types and either neither member is a bit-field or both are bit-fields with the same width for a sequence of one or more initial members. While in C, (C99 TC3 - DR 283 onwards) it's legal to do so (thanks to Pascal Cuoq for bringing this up). However, attempting to do it can still lead to undefined behavior, if the value read happens to be invalid (so called "trap representation") for the type it is read through. Otherwise, the value read is implementation defined. C89/90 called this out under unspecified behavior (Annex J) and K&R's book says it's implementation defined. Quote from K&R: This is the purpose of a union - a single variable that can legitimately hold any of one of several types. [...] so long as the usage is consistent: the type retrieved must be the type most recently stored. It is the programmer's responsibility to keep track of which type is currently stored in a union; the results are implementation-dependent if something is stored as one type and extracted as another. Extract from Stroustrup's TC++PL (emphasis mine) Use of unions can be essential for compatness of data [...] sometimes misused for "type conversion".

最重要的是，这个问题(它的标题从我的提问开始就没有改变)是为了理解联合的目的而提出的，而不是关于标准允许什么。例如，使用继承来实现代码重用当然是c++标准允许的，但这并不是将继承引入c++语言特性的目的或初衷。这就是为什么安德烈的回答仍然被人们所接受的原因。

当前回答

我经常遇到的联合最常见的用法是别名。

考虑以下几点:

union Vector3f
{
  struct{ float x,y,z ; } ;
  float elts[3];
}

这有什么用?它允许通过任意名称干净利落地访问Vector3f的vec;成员:

vec.x=vec.y=vec.z=1.f ;

或者通过整数访问数组

for( int i = 0 ; i < 3 ; i++ )
  vec.elts[i]=1.f;

在某些情况下，通过名称访问是最清晰的方法。在其他情况下，特别是当以编程方式选择轴时，更简单的方法是通过数值索引访问轴- x为0,y为1,z为2。

2013-08-11 22:43:16

其他回答

行为可能没有定义，但这只是意味着没有一个“标准”。所有优秀的编译器都提供#pragmas来控制打包和对齐，但可能有不同的默认值。默认值也会根据所使用的优化设置而改变。

此外，工会不仅仅是为了节省空间。它们可以帮助现代编译器使用类型双关语。如果你reinterpret_cast<>所有的东西，编译器就不能假设你正在做什么。它可能不得不放弃它所知道的类型并重新开始(强制写回内存，与CPU时钟速度相比，这是非常低效的)。

2012-01-18 11:21:31

结合的目的是相当明显的，但由于某种原因，人们经常忽略它。

联合的目的是通过使用相同的内存区域在不同的时间存储不同的对象来节省内存。就是这样。

它就像旅馆里的一个房间。不同的人住在里面的时间不重叠。这些人从来没有见过面，而且通常对彼此一无所知。通过合理管理房间的分时(即确保不同的人不会同时被分配到一个房间)，一个相对较小的酒店可以为相对大量的人提供住宿，这就是酒店的目的。

这正是工会所做的。如果您知道程序中的几个对象所保存的值具有不重叠的值生存期，那么您可以将这些对象“合并”为一个联合，从而节省内存。就像酒店房间在每个时刻最多有一个“活跃”租户一样，工会在每个节目时间最多有一个“活跃”成员。只能读取“活动”成员。通过写入其他成员，您将“活动”状态切换到其他成员。

出于某种原因，联合的最初目的被完全不同的东西“覆盖”了:写联合的一个成员，然后通过另一个成员检查它。这种内存重解释(又名“类型双关语”)不是联合的有效使用。它通常会导致未定义的行为，在C89/90中被描述为产生实现定义的行为。

编辑:在C99标准的一个技术勘误表(参见DR#257和DR#283)中，为了类型双关的目的而使用联合(即写入一个成员，然后读取另一个成员)给出了更详细的定义。但是，请记住，在形式上，这并不能防止您在试图读取陷阱表示时遇到未定义的行为。

2010-02-22 19:52:02

在C语言中，这是实现像变体这样的东西的好方法。

enum possibleTypes{
  eInt,
  eDouble,
  eChar
}


struct Value{

    union Value {
      int iVal_;
      double dval;
      char cVal;
    } value_;
    possibleTypes discriminator_;
} 

switch(val.discriminator_)
{
  case eInt: val.value_.iVal_; break;

在内存较少的情况下，此结构体比具有所有成员的结构体使用更少的内存。

顺便说一下，C提供了

    typedef struct {
      unsigned int mantissa_low:32;      //mantissa
      unsigned int mantissa_high:20;
      unsigned int exponent:11;         //exponent
      unsigned int sign:1;
    } realVal;

访问位值。

2010-02-22 12:10:35

你可以使用a并集，主要有两个原因:

一种以不同方式访问相同数据的方便方法，就像在您的示例中一样当存在不同的数据成员，其中只有一个可以“活动”时，一种节省空间的方法

1实际上更像是一个c风格的黑客，在你知道目标系统的内存架构是如何工作的基础上，以捷径编写代码。就像之前所说的，如果你没有瞄准许多不同的平台，你便能够避开这一问题。我相信一些编译器可能也会让你使用打包指令(我知道他们在结构上这样做)?

2的一个很好的例子。可以在COM中广泛使用的VARIANT类型中找到。

2010-02-22 11:46:30

其他人提到了架构上的差异(小端到大端)。

我读到的问题是，由于变量的内存是共享的，那么写入一个变量，其他变量就会改变，根据它们的类型，值可能是没有意义的。

如。联盟{ 浮动f; int我; });

如果你从x.f读取数据，那么写入x.i是没有意义的——除非你想要查看浮点数的符号、指数或尾数分量。

我认为还有一个对齐的问题:如果一些变量必须字对齐，那么你可能得不到预期的结果。