random.choice的加权版本

我需要写一个加权版的random。选择(列表中的每个元素有不同的被选择的概率)。这是我想到的:

def weightedChoice(choices):
    """Like random.choice, but each element can have a different chance of
    being selected.

    choices can be any iterable containing iterables with two items each.
    Technically, they can have more than two items, the rest will just be
    ignored.  The first item is the thing being chosen, the second item is
    its weight.  The weights can be any numeric values, what matters is the
    relative differences between them.
    """
    space = {}
    current = 0
    for choice, weight in choices:
        if weight > 0:
            space[current] = choice
            current += weight
    rand = random.uniform(0, current)
    for key in sorted(space.keys() + [current]):
        if rand < key:
            return choice
        choice = space[key]
    return None

这个函数对我来说太复杂了，而且很丑。我希望这里的每个人都能提供一些改进的建议或其他方法。对我来说，效率没有代码的整洁和可读性重要。

当前回答

从Python 3.6开始，随机模块中有一个方法选择。

In [1]: import random

In [2]: random.choices(
...:     population=[['a','b'], ['b','a'], ['c','b']],
...:     weights=[0.2, 0.2, 0.6],
...:     k=10
...: )

Out[2]:
[['c', 'b'],
 ['c', 'b'],
 ['b', 'a'],
 ['c', 'b'],
 ['c', 'b'],
 ['b', 'a'],
 ['c', 'b'],
 ['b', 'a'],
 ['c', 'b'],
 ['c', 'b']]

注意随机。选择将与替换样本，每个文档:

返回一个k大小的元素列表，这些元素是从替换的填充中选择的。

为确保回答的完整性，请注意:

当从一个有限的总体中抽取一个抽样单位并返回时对于该种群，在其特征被记录下来之后，在绘制下一个单元之前，采样被称为“与” 更换”。它基本上意味着每个元素可以被选择多于一次。

如果您需要在不替换的情况下进行采样，那么就像@ronan-paixão的精彩回答所说的那样，您可以使用numpy。Choice，其replace参数控制这种行为。

2016-10-11 12:11:30

其他回答

另一种方法是，假设我们的权重与元素数组中的元素的下标相同。

import numpy as np
weights = [0.1, 0.3, 0.5] #weights for the item at index 0,1,2
# sum of weights should be <=1, you can also divide each weight by sum of all weights to standardise it to <=1 constraint.
trials = 1 #number of trials
num_item = 1 #number of items that can be picked in each trial
selected_item_arr = np.random.multinomial(num_item, weights, trials)
# gives number of times an item was selected at a particular index
# this assumes selection with replacement
# one possible output
# selected_item_arr
# array([[0, 0, 1]])
# say if trials = 5, the the possible output could be 
# selected_item_arr
# array([[1, 0, 0],
#   [0, 0, 1],
#   [0, 0, 1],
#   [0, 1, 0],
#   [0, 0, 1]])

现在我们假设，我们要在一次试验中抽取3个项目。你可以假设有三个球R、G、B大量存在，它们的权重由权重数组给定，可能的结果如下:

num_item = 3
trials = 1
selected_item_arr = np.random.multinomial(num_item, weights, trials)
# selected_item_arr can give output like :
# array([[1, 0, 2]])

您还可以将要选择的项目数量视为一组中二项/多项试验的数量。所以，上面的例子仍然可以作为工作

num_binomial_trial = 5
weights = [0.1,0.9] #say an unfair coin weights for H/T
num_experiment_set = 1
selected_item_arr = np.random.multinomial(num_binomial_trial, weights, num_experiment_set)
# possible output
# selected_item_arr
# array([[1, 4]])
# i.e H came 1 time and T came 4 times in 5 binomial trials. And one set contains 5 binomial trails.

2019-10-24 12:42:57

在Udacity免费课程AI for Robotics中，Sebastien Thurn对此进行了演讲。基本上，他用mod运算符%做了一个权重索引的圆形数组，将变量beta设为0，随机选择一个索引， for循环遍历N，其中N是指标的数量，在for循环中，首先按公式增加beta:

Beta = Beta +来自{0…2 * Weight_max}

然后在for循环中嵌套一个while循环per:

while w[index] < beta:
    beta = beta - w[index]
    index = index + 1

select p[index]

然后到下一个索引，根据概率(或课程中介绍的情况下的归一化概率)重新采样。

在Udacity上找到第8课，机器人人工智能的第21期视频，他正在讲粒子滤波器。

2019-12-22 22:39:50

将权重排列成a 累积分布。使用random.random()来选择一个随机的浮点0.0 <= x < total。搜索用等分法进行分布。二等分的如http://docs.python.org/dev/library/bisect.html#other-examples中的示例所示。

from random import random
from bisect import bisect

def weighted_choice(choices):
    values, weights = zip(*choices)
    total = 0
    cum_weights = []
    for w in weights:
        total += w
        cum_weights.append(total)
    x = random() * total
    i = bisect(cum_weights, x)
    return values[i]

>>> weighted_choice([("WHITE",90), ("RED",8), ("GREEN",2)])
'WHITE'

如果需要做出多个选择，可以将其分成两个函数，一个用于构建累积权重，另一个用于对随机点进行等分。

2010-12-01 09:37:07

我需要做这样的事情非常快速非常简单，从搜索的想法，我终于建立了这个模板。其思想是以json的形式从api接收加权值，这里是由dict模拟的。

然后将其转换为一个列表，其中每个值都与它的权重成比例地重复，只需使用random。选择从列表中选择一个值。

我尝试了10次、100次和1000次迭代。分布似乎很稳定。

def weighted_choice(weighted_dict):
    """Input example: dict(apples=60, oranges=30, pineapples=10)"""
    weight_list = []
    for key in weighted_dict.keys():
        weight_list += [key] * weighted_dict[key]
    return random.choice(weight_list)

2018-10-23 12:30:39

从Python v3.6开始，是随机的。选项可用于从给定的填充中返回具有可选权重的指定大小的元素列表。

随机的。select (population, weights=None， *， cum_weights=None, k=1)

总体:包含独特观测值的列表。(如果为空，则引发IndexError) 权重:进行选择所需的更精确的相对权重。 Cum_weights:进行选择所需的累积权重。 K:要输出列表的大小(len)。(默认len () = 1)

一些注意事项:

1)利用加权抽样与替换，使绘制的项目以后可以被替换。权重序列中的值本身并不重要，但它们的相对比例却很重要。

np.random.choice只能将概率作为权重，也必须确保个人概率的总和达到1个标准，但这里没有这样的规定。只要它们属于数值类型(int/float/fraction, Decimal类型除外)，就仍然可以执行。

>>> import random
# weights being integers
>>> random.choices(["white", "green", "red"], [12, 12, 4], k=10)
['green', 'red', 'green', 'white', 'white', 'white', 'green', 'white', 'red', 'white']
# weights being floats
>>> random.choices(["white", "green", "red"], [.12, .12, .04], k=10)
['white', 'white', 'green', 'green', 'red', 'red', 'white', 'green', 'white', 'green']
# weights being fractions
>>> random.choices(["white", "green", "red"], [12/100, 12/100, 4/100], k=10)
['green', 'green', 'white', 'red', 'green', 'red', 'white', 'green', 'green', 'green']

2)如果既没有指定weights，也没有指定cum_weights，则以等概率进行选择。如果提供了权重序列，则它必须与填充序列的长度相同。

同时指定weights和cum_weights将引发TypeError。

>>> random.choices(["white", "green", "red"], k=10)
['white', 'white', 'green', 'red', 'red', 'red', 'white', 'white', 'white', 'green']

3) cum_weights通常是itertools的结果。累加函数在这种情况下非常方便。

从文档链接: 在内部，相对权重被转换为累积权重在进行选择之前，提供累计权重可以节省工作。

因此，无论是提供weights=[12,12,4]还是cum_weights=[12,24,28]，对于我们所设计的情况都会产生相同的结果，并且后者似乎更快/更有效。

2017-01-10 09:06:25

random.choice的加权版本

推荐文章

最新文章

标签