如何找到统计模式?

在R中，mean()和median()是标准函数，它们执行您所期望的功能。Mode()告诉您对象的内部存储模式，而不是参数中出现次数最多的值。但是是否存在一个标准库函数来实现向量(或列表)的统计模式?

当前回答

我浏览了所有这些选项，开始想知道它们的相对特性和性能，所以我做了一些测试。如果其他人也好奇，我在这里分享我的结果。

我不想为这里发布的所有函数而烦恼，我选择了一个基于一些标准的示例:函数应该对字符、因子、逻辑和数字向量都有效，它应该适当地处理na和其他有问题的值，输出应该是“合理的”，即没有数字作为字符或其他类似的愚蠢行为。

我还添加了一个我自己的函数，它是基于与chrispy相同的想法，除了适应更一般的用途:

library(magrittr)

Aksel <- function(x, freq=FALSE) {
    z <- 2
    if (freq) z <- 1:2
    run <- x %>% as.vector %>% sort %>% rle %>% unclass %>% data.frame
    colnames(run) <- c("freq", "value")
    run[which(run$freq==max(run$freq)), z] %>% as.vector   
}

set.seed(2)

F <- sample(c("yes", "no", "maybe", NA), 10, replace=TRUE) %>% factor
Aksel(F)

# [1] maybe yes  

C <- sample(c("Steve", "Jane", "Jonas", "Petra"), 20, replace=TRUE)
Aksel(C, freq=TRUE)

# freq value
#    7 Steve

最后，我通过微基准测试在两组测试数据上运行了五个函数。函数名指的是它们各自的作者:

Chris的函数被设置为method="modes"和na。rm=TRUE默认值，以使其更具可比性，但除此之外，这里使用的函数是由它们的作者提供的。

In matter of speed alone Kens version wins handily, but it is also the only one of these that will only report one mode, no matter how many there really are. As is often the case, there's a trade-off between speed and versatility. In method="mode", Chris' version will return a value iff there is one mode, else NA. I think that's a nice touch. I also think it's interesting how some of the functions are affected by an increased number of unique values, while others aren't nearly as much. I haven't studied the code in detail to figure out why that is, apart from eliminating logical/numeric as a the cause.

2016-05-27 02:49:33

其他回答

在r邮件列表中发现了这个，希望对你有帮助。我也是这么想的。您将希望table()数据，排序，然后选择第一个名称。这有点粗俗，但应该有用。

names(sort(-table(x)))[1]

2010-03-30 18:19:29

下面是可以用来找到R中矢量变量的模式的代码。

a <- table([vector])

names(a[a==max(a)])

2017-02-21 10:58:20

计算包含离散值的向量“v”的MODE的一个简单方法是:

names(sort(table(v)))[length(sort(table(v)))]

2016-08-27 07:54:50

虽然我喜欢肯威廉姆斯简单的功能，我想检索多种模式，如果他们存在。考虑到这一点，我使用下面的函数，它返回多个模式或单个模式的列表。

rmode <- function(x) {
  x <- sort(x)  
  u <- unique(x)
  y <- lapply(u, function(y) length(x[x==y]))
  u[which( unlist(y) == max(unlist(y)) )]
}

2014-12-24 16:08:02

这里有另一个解决方案:

freq <- tapply(mySamples,mySamples,length)
#or freq <- table(mySamples)
as.numeric(names(freq)[which.max(freq)])

2010-03-30 20:21:29

如何找到统计模式?

推荐文章

最新文章

标签