什么时候不使用Cassandra?

最近有很多关于卡桑德拉的话题。

Twitter, Digg, Facebook等都在使用它。

什么时候有意义:

使用卡桑德拉, 不用卡桑德拉，还有使用RDMS而不是Cassandra。

当前回答

没有什么是银弹，任何东西都是为了解决特定的问题而构建的，有自己的优点和缺点。这取决于你，你有什么问题陈述，什么是该问题的最佳解决方案。

我会按照你问的顺序一个一个地回答你的问题。因为Cassandra是基于NoSQL数据库家族的，所以在我回答你的问题之前，理解为什么使用NoSQL数据库是很重要的。

为什么使用NoSQL

In the case of RDBMS, making a choice is quite easy because all the databases like MySQL, Oracle, MS SQL, PostgreSQL in this category offer almost the same kind of solutions oriented toward ACID properties. When it comes to NoSQL, the decision becomes difficult because every NoSQL database offers different solutions and you have to understand which one is best suited for your app/system requirements. For example, MongoDB is fit for use cases where your system demands a schema-less document store. HBase might be fit for search engines, analyzing log data, or any place where scanning huge, two-dimensional join-less tables is a requirement. Redis is built to provide In-Memory search for varieties of data structures like trees, queues, linked lists, etc and can be a good fit for making real-time leaderboards, pub-sub kind of system. Similarly there are other databases in this category (Including Cassandra) which are fit for different problem statements. Now lets move to the original questions, and answer them one by one.

何时使用卡桑德拉

Being a part of the NoSQL family, Cassandra offers a solution for problems where one of your requirements is to have a very heavy write system and you want to have a quite responsive reporting system on top of that stored data. Consider the use case of Web analytics where log data is stored for each request and you want to built an analytical platform around it to count hits per hour, by browser, by IP, etc in a real time manner. You can refer to this blog post to understand more about the use cases where Cassandra fits in.

什么时候使用RDMS而不是Cassandra

Cassandra基于NoSQL数据库，不提供ACID和关系数据属性。如果您对ACID属性有强烈的需求(例如财务数据)，Cassandra将不适合这种情况。显然，您可以为此制定一个变通方案，但是您最终将编写大量的应用程序代码来模拟ACID属性，并将严重延误上市时间。同时，使用Cassandra管理这种系统对您来说也是复杂而乏味的。

什么时候不用卡桑德拉

我认为上面的解释是否有意义不需要回答。

2015-06-21 11:33:24

其他回答

Cassandra是一个特定问题的答案:当您有太多数据，以至于无法在一台服务器上存储时，您该怎么办?如何将所有数据存储在多个服务器上，同时不破坏银行账户，不让开发人员抓狂?Facebook每天都会收到4tb的压缩数据。这个数字很可能在一年内增长两倍以上。

如果您没有这么多数据，或者您有数百万美元来支付企业Oracle/DB2集群安装费用，以及安装和维护它所需的专家，那么您可以使用SQL数据库。

然而，Facebook不再使用cassandra，现在几乎只使用MySQL，在应用程序堆栈中移动分区，以获得更快的性能和更好的控制。

2010-04-24 19:30:22

如果你需要一个SQL语义完全一致的数据库，Cassandra不是你的解决方案。Cassandra支持键值查找。它不支持SQL查询。Cassandra中的数据“最终是一致的”。数据的并发查找可能不一致，但最终查找是一致的。

如果你需要严格的语义，需要对SQL查询的支持，可以选择其他的解决方案，比如MySQL, PostGres，或者结合使用Cassandra和Solr。

2017-03-09 04:23:29

除了上面给出的关于何时使用和何时不使用Cassandra的答案外，如果你决定使用Cassandra，你可能会考虑不使用Cassandra本身，而是使用它的众多表亲之一。

上面的一些答案已经指出了各种“NoSQL”系统，它们与Cassandra有许多相同的属性，有一些或大或小的差异，并且可能比Cassandra本身更适合您的特定需求。

Additionally, recently (several years after this question was originally asked), a Cassandra clone called Scylla (see https://en.wikipedia.org/wiki/Scylla_(database)) was released. Scylla is an open-source re-implementation of Cassandra in C++, which claims to have significantly higher throughput and lower latencies than the original Java Cassandra, while being mostly compatible with it (in features, APIs, and file formats). So if you're already considering Cassandra, you may want to consider Scylla as well.

2017-11-07 09:51:11

在评估分布式数据系统时，您必须考虑CAP定理——您可以选择以下两个:一致性、可用性和分区容差。

Cassandra是一个可用的、支持最终一致性的分区容忍系统。要了解更多信息，请参阅我写的这篇博客文章:NoSQL系统的可视化指南。

2010-04-20 19:01:38

除了这里的其他答案之外，沉重的单个查询与无数的轻查询负载是另一个需要考虑的问题。在nosql风格的DB中自动优化单个查询本身就比较困难。我使用过MongoDB，在尝试计算复杂查询时遇到了性能问题。我没有使用Cassandra，但我预计它会有同样的问题。

另一方面，如果您的负载预期是许多小型查询的负载，并且您希望能够轻松地向外扩展，那么您可以利用大多数NoSql数据库提供的最终一致性。注意，最终一致性实际上不是非关系数据模型的特性，但是在基于nosql的系统中实现和设置一致性要容易得多。

For a single, very heavy query, any modern RDBMS engine can do a decent job parallelizing parts of the query and take advantage of as much CPU and memory you throw at it (on a single machine). NoSql databases don't have enough information about the structure of the data to be able to make assumptions that will allow truly intelligent parallelization of a big query. They do allow you to easily scale out more servers (or cores) but once the query hits a complexity level you are basically forced to split it apart manually to parts that the NoSql engine knows how to deal with intelligently.

根据我使用MongoDB的经验，由于查询的复杂性，MongoDB最终无法对其进行优化，也无法在多个数据上运行部分查询。Mongo可以并行多个查询，但不太擅长优化单个查询。

2013-04-09 14:36:09

什么时候不使用Cassandra?

推荐文章

最新文章

标签