在XML中什么是无效字符

我正在使用一些包含字符串的XML:

<node>This is a string</node>

我传递给节点的一些字符串将有&，#，$等字符:

<node>This is a string & so is this</node>

由于&，这是无效的。

我不能在CDATA中包装这些字符串，因为它们需要这样。我尝试寻找一个字符列表，这些字符不能放在XML节点中，而不能放在CDATA中。

有人能给我指个方向或者给我一份非法字符的列表吗?

当前回答

综上所述，文本中的有效字符为:

制表符，换行和换行。除&和<外，所有非控制字符都有效。如果使用]]，则>无效。

XML规范的2.2节和2.4节详细给出了答案:

字符

合法字符包括制表符、回车符、换行符以及Unicode和ISO/IEC 10646的合法字符

字符数据

The ampersand character (&) and the left angle bracket (<) must not appear in their literal form, except when used as markup delimiters, or within a comment, a processing instruction, or a CDATA section. If they are needed elsewhere, they must be escaped using either numeric character references or the strings " & " and " < " respectively. The right angle bracket (>) may be represented using the string " > ", and must, for compatibility, be escaped using either " > " or a character reference when it appears in the string " ]]> " in content, when that string is not marking the end of a CDATA section.

2018-10-24 14:41:42

其他回答

预先声明的字符是:

& < > " '

有关更多信息，请参阅“XML中的特殊字符是什么?”

2009-04-08 13:59:20

在Woodstox XML处理器中，无效字符由以下代码分类:

if (c == 0) {
    throw new IOException("Invalid null character in text to output");
}
if (c < ' ' || (c >= 0x7F && c <= 0x9F)) {
    String msg = "Invalid white space character (0x" + Integer.toHexString(c) + ") in text to output";
    if (mXml11) {
        msg += " (can only be output using character entity)";
    }
    throw new IOException(msg);
}
if (c > 0x10FFFF) {
    throw new IOException("Illegal unicode character point (0x" + Integer.toHexString(c) + ") to output; max is 0x10FFFF as per RFC");
}
/*
 * Surrogate pair in non-quotable (not text or attribute value) content, and non-unicode encoding (ISO-8859-x,
 * Ascii)?
 */
if (c >= SURR1_FIRST && c <= SURR2_LAST) {
    throw new IOException("Illegal surrogate pair -- can only be output via character entities, which are not allowed in this content");
}
throw new IOException("Invalid XML character (0x"+Integer.toHexString(c)+") in text to output");

来自这里

2014-12-03 10:27:05

唯一的非法字符是&，<和>(以及属性中的"或'，这取决于使用哪个字符来分隔属性值:attr="必须使用"这里，' is allowed '和attr='必须使用'在这里，“is allowed”)。

它们是用XML实体转义的，这里你需要&&。

实际上，您应该使用一个工具或库来为您编写XML，并为您抽象这类东西，这样您就不必担心了。

2009-04-08 13:59:48

另一个简单的方法是在c#中转义可能不需要的XML / XHTML字符:

WebUtility.HtmlEncode(stringWithStrangeChars)

2014-02-19 10:01:02

有效字符的列表在XML规范中:

Char       ::=      #x9 | #xA | #xD | [#x20-#xD7FF] | [#xE000-#xFFFD] | [#x10000-#x10FFFF]  /* any Unicode character, excluding the surrogate blocks, FFFE, and FFFF. */

2011-02-24 20:34:52

在XML中什么是无效字符

推荐文章

最新文章

标签