mirror of https://github.com/apache/cassandra
The SimpleProgressLog had a number of problems: 1. It polled for progress with no attempt to determine whether progress could realistically be made, so: - as the number of pending transactions grew, the proportion of useful work dropped (as many would be unable to make progress without earlier transactions completing) - each transaction in the chain could recover only on average 1/2 poll interval behind the last transaction to complete 2. It requested full transaction state from every replica on each attempt 3. It maintained a lot of in-memory state 4. Polling happened en-masse, allowing for little per-transaction control We also separately maintained fairly expensive per-command listener state that negatively affected our command loading and caching. The new DefaultProgressLog makes use of several new features: LocalListeners, RemoteListeners, Timers and Await messages. - LocalListeners provide a memory-efficient collection for managing each CommandStore<E2><80><99>s transaction listeners, with dedicated record keeping for inter-transaction relationships. - RemoteListeners provide a mechanism for request/response pairs that may be separated by longer than the normal Cassandra message timeout, and require minimal state on sender and recipient. This permits replicas to cheaply update their local state machine as soon as distributed information becomes available. The DefaultProgressLog tracks each transaction with separate timers to handle per-transaction scheduling, backoff etc, and a succinct state machine. To reduce overhead correspondence is preferentially limited to a handful of replicas, and limited to the home shard where appropriate. patch by Benedict; reviewed by Ariel Weisberg for CASSANDRA-19870 |
||
|---|---|---|
| .build | ||
| .circleci | ||
| .github | ||
| .jenkins | ||
| bin | ||
| ci | ||
| conf | ||
| debian | ||
| doc | ||
| examples | ||
| ide | ||
| lib | ||
| modules | ||
| pylib | ||
| redhat | ||
| src | ||
| test | ||
| tools | ||
| .asf.yaml | ||
| .gitignore | ||
| .gitmodules | ||
| .snyk | ||
| CASSANDRA-14092.txt | ||
| CHANGES.txt | ||
| CONTRIBUTING.md | ||
| LICENSE.txt | ||
| NEWS.txt | ||
| NOTICE.txt | ||
| README.asc | ||
| TESTING.md | ||
| accord_demo.txt | ||
| build-shaded-dtest-jar.sh | ||
| build.properties.default | ||
| build.xml | ||
| relocate-dependencies.pom | ||
| simulator.sh | ||
README.asc
Apache Cassandra
-----------------
Apache Cassandra is a highly-scalable partitioned row store. Rows are organized into tables with a required primary key.
https://cwiki.apache.org/confluence/display/CASSANDRA2/Partitioners[Partitioning] means that Cassandra can distribute your data across multiple machines in an application-transparent matter. Cassandra will automatically repartition as machines are added and removed from the cluster.
https://cwiki.apache.org/confluence/display/CASSANDRA2/DataModel[Row store] means that like relational databases, Cassandra organizes data by rows and columns. The Cassandra Query Language (CQL) is a close relative of SQL.
For more information, see http://cassandra.apache.org/[the Apache Cassandra web site].
Issues should be reported on https://issues.apache.org/jira/projects/CASSANDRA/issues/[The Cassandra Jira].
Requirements
------------
- Java: see supported versions in build.xml (search for property "java.supported").
- Python: for `cqlsh`, see `bin/cqlsh` (search for function "is_supported_version").
Getting started
---------------
This short guide will walk you through getting a basic one node cluster up
and running, and demonstrate some simple reads and writes. For a more-complete guide, please see the Apache Cassandra website's https://cassandra.apache.org/doc/latest/cassandra/getting_started/index.html[Getting Started Guide].
First, we'll unpack our archive:
$ tar -zxvf apache-cassandra-$VERSION.tar.gz
$ cd apache-cassandra-$VERSION
After that we start the server. Running the startup script with the -f argument will cause
Cassandra to remain in the foreground and log to standard out; it can be stopped with ctrl-C.
$ bin/cassandra -f
Now let's try to read and write some data using the Cassandra Query Language:
$ bin/cqlsh
The command line client is interactive so if everything worked you should
be sitting in front of a prompt:
----
Connected to Test Cluster at localhost:9160.
[cqlsh 6.3.0 | Cassandra 5.0-SNAPSHOT | CQL spec 3.4.8 | Native protocol v5]
Use HELP for help.
cqlsh>
----
As the banner says, you can use 'help;' or '?' to see what CQL has to
offer, and 'quit;' or 'exit;' when you've had enough fun. But lets try
something slightly more interesting:
----
cqlsh> CREATE KEYSPACE schema1
WITH replication = { 'class' : 'SimpleStrategy', 'replication_factor' : 1 };
cqlsh> USE schema1;
cqlsh:Schema1> CREATE TABLE users (
user_id varchar PRIMARY KEY,
first varchar,
last varchar,
age int
);
cqlsh:Schema1> INSERT INTO users (user_id, first, last, age)
VALUES ('jsmith', 'John', 'Smith', 42);
cqlsh:Schema1> SELECT * FROM users;
user_id | age | first | last
---------+-----+-------+-------
jsmith | 42 | john | smith
cqlsh:Schema1>
----
If your session looks similar to what's above, congrats, your single node
cluster is operational!
For more on what commands are supported by CQL, see
http://cassandra.apache.org/doc/latest/cql/[the CQL reference]. A
reasonable way to think of it is as, "SQL minus joins and subqueries, plus collections."
Wondering where to go from here?
* Join us in #cassandra on the https://s.apache.org/slack-invite[ASF Slack] and ask questions.
* Subscribe to the Users mailing list by sending a mail to
user-subscribe@cassandra.apache.org.
* Subscribe to the Developer mailing list by sending a mail to
dev-subscribe@cassandra.apache.org.
* Visit the http://cassandra.apache.org/community/[community section] of the Cassandra website for more information on getting involved.
* Visit the http://cassandra.apache.org/doc/latest/development/index.html[development section] of the Cassandra website for more information on how to contribute.