Go to file
Sylvain Lebresne e2ecdf268a Remove broken "defragment-on-read" optimization
The read path for names queries has had a "defragment-on-read"
optimization for a while whereby if too many sstables are hit by the
read, the result is written back into memtable, in the hope that later
reads will only read that newly written data in a single sstable (or at
least fewer).

The principle of that optimisation does not work however as data is
written back with the same timestamp as it originally has and that means
future reads cannot know to skip older sstables (at least with the
metadata we currently store).

As such, this optimisation never saved anything and in fact added load.

The patch removes that broken code.

Patch by Sylvain Lebresne, reviewed by Aleksey Yeschenko for
CASSANDRA-15432
2020-08-17 11:31:03 +02:00
.circleci When jvm dtest apis differ Circle CI's dtest_jars_build can fail to detect this and will use the jars from the older version 2020-08-03 15:15:32 -07:00
.jenkins Run in-jvm upgrade dtests in ci-cassandra 2020-08-08 10:07:24 +02:00
bin Backport CASSANDRA-12189, formatting fixes 2020-07-15 14:15:53 -05:00
conf Operational improvements and hardening for replica filtering protection 2020-07-30 12:53:52 +01:00
debian Prepare debian changelog for 3.0.21 2020-07-14 22:15:26 +02:00
doc Prevent client requests from blocking on executor task queue 2019-07-15 13:13:40 +01:00
examples Updated trigger example 2015-11-10 15:35:13 +00:00
ide Merge branch 'cassandra-2.2' into cassandra-3.0 2020-01-21 19:08:52 +01:00
interface Swap all references to 3.0 with 2.2 2015-05-14 17:16:54 +03:00
lib Allow dropping COMPACT STORAGE flag 2017-11-06 15:44:51 +01:00
pylib Backport CASSANDRA-12189, formatting fixes 2020-07-15 14:15:53 -05:00
redhat Merge branch 'cassandra-2.2' into cassandra-3.0 2020-03-03 22:45:53 +01:00
src Remove broken "defragment-on-read" optimization 2020-08-17 11:31:03 +02:00
test Merge branch 'cassandra-2.2' into cassandra-3.0 2020-08-04 10:45:03 +02:00
tools support frozen collections: list, set in cassandra-stress 2019-04-10 21:12:53 +10:00
.gitignore Remove generated files from apache-cassandra-*-src.tar.gz artifacts 2020-06-08 22:31:41 +02:00
.rat-excludes Remove Pig support 2015-10-16 13:23:20 +01:00
CASSANDRA-14092.txt Merge branch 'cassandra-2.2' into cassandra-3.0 2018-02-10 14:57:53 -02:00
CHANGES.txt Remove broken "defragment-on-read" optimization 2020-08-17 11:31:03 +02:00
CONTRIBUTING.md Add CONTRIBUTING.md 2015-01-05 22:45:53 +03:00
LICENSE.txt merge with 0.6 branch (post-850) 2010-03-26 16:31:52 +00:00
NEWS.txt Bump version to 3.0.21 2020-02-14 18:44:09 -06:00
NOTICE.txt Eliminate the dependency on jgrapht for UDT resolution 2015-12-18 15:39:33 +00:00
README.asc Merge branch 'cassandra-2.2' into cassandra-3.0 2020-07-29 14:17:28 +02:00
build.properties.default Switch http to https URLs in build.xml 2019-05-22 15:03:43 -04:00
build.xml Operational improvements and hardening for replica filtering protection 2020-07-30 12:53:52 +01:00
eclipse_compiler.properties Add Static Analysis to warn on unsafe use of Autocloseable instances 2015-05-27 17:53:26 -04:00

README.asc

Executive summary
-----------------

Cassandra is a partitioned row store.  Rows are organized into tables with a required primary key.

http://wiki.apache.org/cassandra/Partitioners[Partitioning] means that Cassandra can distribute your data across multiple machines in an application-transparent matter.  Cassandra will automatically repartition as machines are added and removed from the cluster.

http://wiki.apache.org/cassandra/DataModel[Row store] means that like relational databases, Cassandra organizes data by rows and columns.  The Cassandra Query Language (CQL) is a close relative of SQL.

For more information, see http://cassandra.apache.org/[the Apache Cassandra web site].

Requirements
------------
. Java >= 1.8 (OpenJDK and Oracle JVMS have been tested)
. Python 2.7 (for cqlsh)

Getting started
---------------

This short guide will walk you through getting a basic one node cluster up
and running, and demonstrate some simple reads and writes.

First, we'll unpack our archive:

  $ tar -zxvf apache-cassandra-$VERSION.tar.gz
  $ cd apache-cassandra-$VERSION

After that we start the server.  Running the startup script with the -f argument will cause
Cassandra to remain in the foreground and log to standard out; it can be stopped with ctrl-C.

  $ bin/cassandra -f

****
Note for Windows users: to install Cassandra as a service, download
http://commons.apache.org/daemon/procrun.html[Procrun], set the
PRUNSRV environment variable to the full path of prunsrv (e.g.,
C:\procrun\prunsrv.exe), and run "bin\cassandra.bat install".
Similarly, "uninstall" will remove the service.
****

Now let's try to read and write some data using the Cassandra Query Language:

  $ bin/cqlsh

The command line client is interactive so if everything worked you should
be sitting in front of a prompt:

----
Connected to Test Cluster at localhost:9160.
[cqlsh 2.2.0 | Cassandra 1.2.0 | CQL spec 3.0.0 | Thrift protocol 19.35.0]
Use HELP for help.
cqlsh>
----

As the banner says, you can use 'help;' or '?' to see what CQL has to
offer, and 'quit;' or 'exit;' when you've had enough fun. But lets try
something slightly more interesting:

----
cqlsh> CREATE KEYSPACE schema1
       WITH replication = { 'class' : 'SimpleStrategy', 'replication_factor' : 1 };
cqlsh> USE schema1;
cqlsh:Schema1> CREATE TABLE users (
                 user_id varchar PRIMARY KEY,
                 first varchar,
                 last varchar,
                 age int
               );
cqlsh:Schema1> INSERT INTO users (user_id, first, last, age)
               VALUES ('jsmith', 'John', 'Smith', 42);
cqlsh:Schema1> SELECT * FROM users;
 user_id | age | first | last
---------+-----+-------+-------
  jsmith |  42 |  john | smith
cqlsh:Schema1>
----

If your session looks similar to what's above, congrats, your single node
cluster is operational!

For more on what commands are supported by CQL, see
https://github.com/apache/cassandra/blob/trunk/doc/cql3/CQL.textile[the CQL reference].  A
reasonable way to think of it is as, "SQL minus joins and subqueries, plus collections."

Wondering where to go from here?

  * Getting started: http://wiki.apache.org/cassandra/GettingStarted
  * Join us in #cassandra on irc.freenode.net and ask questions
  * Subscribe to the Users mailing list by sending a mail to
    user-subscribe@cassandra.apache.org
  * Planet Cassandra aggregates Cassandra articles and news:
    http://planetcassandra.org/