random.choice的加权版本

我需要写一个加权版的random。选择(列表中的每个元素有不同的被选择的概率)。这是我想到的:

def weightedChoice(choices):
    """Like random.choice, but each element can have a different chance of
    being selected.

    choices can be any iterable containing iterables with two items each.
    Technically, they can have more than two items, the rest will just be
    ignored.  The first item is the thing being chosen, the second item is
    its weight.  The weights can be any numeric values, what matters is the
    relative differences between them.
    """
    space = {}
    current = 0
    for choice, weight in choices:
        if weight > 0:
            space[current] = choice
            current += weight
    rand = random.uniform(0, current)
    for key in sorted(space.keys() + [current]):
        if rand < key:
            return choice
        choice = space[key]
    return None

这个函数对我来说太复杂了，而且很丑。我希望这里的每个人都能提供一些改进的建议或其他方法。对我来说，效率没有代码的整洁和可读性重要。

当前回答

这取决于你想对分布进行多少次抽样。

假设要对分布进行K次抽样。当n是分布中的项数时，每次使用np.random.choice()的时间复杂度为O(K(n + log(n)))。

在我的例子中，我需要对相同的分布进行多次采样，阶数为10^3其中n阶数为10^6。我使用了下面的代码，它预先计算了累积分布，并在O(log(n))中对其进行采样。总体时间复杂度为O(n+K*log(n))。

import numpy as np

n,k = 10**6,10**3

# Create dummy distribution
a = np.array([i+1 for i in range(n)])
p = np.array([1.0/n]*n)

cfd = p.cumsum()
for _ in range(k):
    x = np.random.uniform()
    idx = cfd.searchsorted(x, side='right')
    sampled_element = a[idx]

2017-11-06 10:29:03

其他回答

import numpy as np
w=np.array([ 0.4,  0.8,  1.6,  0.8,  0.4])
np.random.choice(w, p=w/sum(w))

2013-12-11 16:38:41

从Python v3.6开始，是随机的。选项可用于从给定的填充中返回具有可选权重的指定大小的元素列表。

随机的。select (population, weights=None， *， cum_weights=None, k=1)

总体:包含独特观测值的列表。(如果为空，则引发IndexError) 权重:进行选择所需的更精确的相对权重。 Cum_weights:进行选择所需的累积权重。 K:要输出列表的大小(len)。(默认len () = 1)

一些注意事项:

1)利用加权抽样与替换，使绘制的项目以后可以被替换。权重序列中的值本身并不重要，但它们的相对比例却很重要。

np.random.choice只能将概率作为权重，也必须确保个人概率的总和达到1个标准，但这里没有这样的规定。只要它们属于数值类型(int/float/fraction, Decimal类型除外)，就仍然可以执行。

>>> import random
# weights being integers
>>> random.choices(["white", "green", "red"], [12, 12, 4], k=10)
['green', 'red', 'green', 'white', 'white', 'white', 'green', 'white', 'red', 'white']
# weights being floats
>>> random.choices(["white", "green", "red"], [.12, .12, .04], k=10)
['white', 'white', 'green', 'green', 'red', 'red', 'white', 'green', 'white', 'green']
# weights being fractions
>>> random.choices(["white", "green", "red"], [12/100, 12/100, 4/100], k=10)
['green', 'green', 'white', 'red', 'green', 'red', 'white', 'green', 'green', 'green']

2)如果既没有指定weights，也没有指定cum_weights，则以等概率进行选择。如果提供了权重序列，则它必须与填充序列的长度相同。

同时指定weights和cum_weights将引发TypeError。

>>> random.choices(["white", "green", "red"], k=10)
['white', 'white', 'green', 'red', 'red', 'red', 'white', 'white', 'white', 'green']

3) cum_weights通常是itertools的结果。累加函数在这种情况下非常方便。

从文档链接: 在内部，相对权重被转换为累积权重在进行选择之前，提供累计权重可以节省工作。

因此，无论是提供weights=[12,12,4]还是cum_weights=[12,24,28]，对于我们所设计的情况都会产生相同的结果，并且后者似乎更快/更有效。

2017-01-10 09:06:25

如果不介意使用numpy，可以使用numpy.random.choice。

例如:

import numpy

items  = [["item1", 0.2], ["item2", 0.3], ["item3", 0.45], ["item4", 0.05]
elems = [i[0] for i in items]
probs = [i[1] for i in items]

trials = 1000
results = [0] * len(items)
for i in range(trials):
    res = numpy.random.choice(items, p=probs)  #This is where the item is selected!
    results[items.index(res)] += 1
results = [r / float(trials) for r in results]
print "item\texpected\tactual"
for i in range(len(probs)):
    print "%s\t%0.4f\t%0.4f" % (items[i], probs[i], results[i])

如果你知道你需要提前做多少选择，你可以不像这样循环:

numpy.random.choice(items, trials, p=probs)

2013-03-21 15:14:38

我可能已经来不及提供任何有用的东西了，但这里有一个简单，简短，非常有效的片段:

def choose_index(probabilies):
    cmf = probabilies[0]
    choice = random.random()
    for k in xrange(len(probabilies)):
        if choice <= cmf:
            return k
        else:
            cmf += probabilies[k+1]

不需要排序你的概率或用你的cmf创建一个向量，它一旦找到它的选择就会终止。内存:O(1)，时间:O(N)，平均运行时间~ N/2。

如果你有权重，只需添加一行:

def choose_index(weights):
    probabilities = weights / sum(weights)
    cmf = probabilies[0]
    choice = random.random()
    for k in xrange(len(probabilies)):
        if choice <= cmf:
            return k
        else:
            cmf += probabilies[k+1]

2015-01-27 21:55:52

假设你有

items = [11, 23, 43, 91] 
probability = [0.2, 0.3, 0.4, 0.1]

你有一个函数，它生成一个介于[0,1)之间的随机数(我们可以在这里使用random.random())。现在求概率的前缀和

prefix_probability=[0.2,0.5,0.9,1]

现在，我们只需取一个0-1之间的随机数，然后使用二分搜索来查找该数字在prefix_probability中的位置。这个索引就是你的答案

代码是这样的

return items[bisect.bisect(prefix_probability,random.random())]

2022-11-29 07:35:14

random.choice的加权版本

推荐文章

最新文章

标签