按子字符串条件筛选panda DataFrame

我有一个熊猫DataFrame，其中包含一列字符串值。我需要根据部分字符串匹配来选择行。

类似于这个成语：

re.search(pattern, cell_in_question)

返回布尔值。我熟悉df[df['A']==“helloworld”]的语法，但似乎找不到一种方法来处理部分字符串匹配，比如“hello”。

当前回答

如果有人想知道如何执行相关问题：“按部分字符串选择列”

Use:

df.filter(like='hello')  # select columns which contain the word hello

若要通过部分字符串匹配来选择行，请将axis=0传递到筛选器：

# selects rows which contain the word hello in their index label
df.filter(like='hello', axis=0)

2016-10-12 21:04:32

其他回答

一个更一般的例子-如果在字符串中查找单词或特定单词的部分：

df = pd.DataFrame([('cat andhat', 1000.0), ('hat', 2000000.0), ('the small dog', 1000.0), ('fog', 330000.0),('pet', 330000.0)], columns=['col1', 'col2'])

句子或单词的特定部分：

searchfor = '.*cat.*hat.*|.*the.*dog.*'

创建显示受影响行的列（可以根据需要过滤掉）

df["TrueFalse"]=df['col1'].str.contains(searchfor, regex=True)

    col1             col2           TrueFalse
0   cat andhat       1000.0         True
1   hat              2000000.0      False
2   the small dog    1000.0         True
3   fog              330000.0       False
4   pet 3            30000.0        False

2021-02-16 09:41:59

假设我们在数据帧df中有一个名为“ENTITY”的列。我们可以过滤df，以获得整个数据帧df，其中“实体”列的行不包含“DM”，方法如下：

mask = df['ENTITY'].str.contains('DM')

df = df.loc[~(mask)].copy(deep=True)

2021-03-30 12:06:24

这是我最后为部分字符串匹配所做的。如果有人有更有效的方法，请告诉我。

def stringSearchColumn_DataFrame(df, colName, regex):
    newdf = DataFrame()
    for idx, record in df[colName].iteritems():

        if re.search(regex, record):
            newdf = concat([df[df[colName] == record], newdf], ignore_index=True)

    return newdf

2012-07-06 17:08:46

df[df['A'].str.contains("hello", case=False)]

2022-10-04 11:41:57

我在ipython笔记本电脑的macos上使用熊猫0.14.1。我尝试了上面的建议行：

df[df["A"].str.contains("Hello|Britain")]

并得到一个错误：

无法使用包含NA/NaN值的矢量进行索引

但当添加了“==True”条件时，效果非常好，如下所示：

df[df['A'].str.contains("Hello|Britain")==True]

2014-11-10 17:05:17

按子字符串条件筛选panda DataFrame

推荐文章

最新文章

标签