按子字符串条件筛选panda DataFrame

我有一个熊猫DataFrame，其中包含一列字符串值。我需要根据部分字符串匹配来选择行。

类似于这个成语：

re.search(pattern, cell_in_question)

返回布尔值。我熟悉df[df['A']==“helloworld”]的语法，但似乎找不到一种方法来处理部分字符串匹配，比如“hello”。

当前回答

矢量化字符串方法（即Series.str）允许您执行以下操作：

df[df['A'].str.contains("hello")]

这在熊猫0.8.1及以上版本中可用。

2012-07-17 21:52:18

其他回答

这是我最后为部分字符串匹配所做的。如果有人有更有效的方法，请告诉我。

def stringSearchColumn_DataFrame(df, colName, regex):
    newdf = DataFrame()
    for idx, record in df[colName].iteritems():

        if re.search(regex, record):
            newdf = concat([df[df[colName] == record], newdf], ignore_index=True)

    return newdf

2012-07-06 17:08:46

有点类似于@cs95的答案，但这里不需要指定引擎：

df.query('A.str.contains("hello").values')

2022-02-22 20:23:22

我的2c价值：

我执行了以下操作：

sale_method = pd.DataFrame(model_data['Sale Method'].str.upper())
sale_method['sale_classification'] = \
    np.where(sale_method['Sale Method'].isin(['PRIVATE']),
             'private',
             np.where(sale_method['Sale Method']
                      .str.contains('AUCTION'),
                      'auction',
                      'other'
             )
    )

2021-06-24 01:40:21

您可以尝试将它们视为字符串：

df[df['A'].astype(str).str.contains("Hello|Britain")]

2021-05-29 08:16:45

我在ipython笔记本电脑的macos上使用熊猫0.14.1。我尝试了上面的建议行：

df[df["A"].str.contains("Hello|Britain")]

并得到一个错误：

无法使用包含NA/NaN值的矢量进行索引

但当添加了“==True”条件时，效果非常好，如下所示：

df[df['A'].str.contains("Hello|Britain")==True]

2014-11-10 17:05:17

按子字符串条件筛选panda DataFrame

推荐文章

最新文章

标签