Go to file
Sam Tunnicliffe 04533e6cda Avoid blocking AntiEntropyStage when submitting validation requests
Patch by Sam Tunnicliffe; reviewed by Benjamin Lerer for CASSANDRA-15812

Switches ValidationExecutor's work queue to LinkedBlockingQueue to
avoid blocking AntiEntropyStage when the executor is saturated. This
requires VE.corePoolSize to be set to concurrent_validations as now
it will always prefer to queue requests rather than start new threads.

This commit also adds a hard limit on concurrent_validations, as allowing
an unbounded number of validations to run concurrently is never safe.
This was always true, but setting a high value here is now more
dangerous as it controls the number of core, not max, threads.
This hard limit is linked to concurrent_compactors, so operators may
set concurrent_validations between 1 and concurrent_compactors.
The meaning of setting it < 1 has changed from "unbounded" to
"whatever concurrent_compactors is set to".

This safety valve can be overridden with a system property at startup
and/or a JMX property.

CASSANDRA-9292 removed the 1hr timeout on prepare messages, but this
was inadvertently undone when CASSANDRA-13397 was committed. As nothing
long running is done in the repair phase anymore, this timeout can
safely be reduced.

If using RepairCommandPoolFullStrategy.queue, the core pool size
for repairCommandExecutor must be increased from the default
value of 1 or else all concurrent tasks will be queued and no
more threads created.
2020-06-09 13:41:05 +01:00
.circleci Use Docker image for dtests in CircleCI w/ JAVA8_HOME environment variable & Allow different pip-source-install repos in requirements.txt 2020-06-04 10:34:43 +02:00
.jenkins Merge branch 'cassandra-3.11' into trunk 2020-05-07 13:35:36 +02:00
bin Fix CQLSH UTF-8 encoding issue for Python 2/3 compatibility 2020-04-22 12:44:47 -07:00
conf Avoid blocking AntiEntropyStage when submitting validation requests 2020-06-09 13:41:05 +01:00
debian Add fqltool and auditlogviewer to rpm and deb packages 2020-06-03 18:31:26 +02:00
doc Improving Cassandra configuration docs 2020-06-03 11:24:38 -07:00
examples/triggers Fix trigger example on 4.0 2017-08-24 08:34:34 -07:00
ide Add compaction allocation measurement test 2020-03-11 12:50:50 -07:00
lib Update Python driver for cqlsh to 3.23 2020-05-07 09:35:10 -07:00
pylib Merge branch 'cassandra-3.11' into trunk 2020-04-29 21:10:51 +02:00
redhat Add fqltool and auditlogviewer to rpm and deb packages 2020-06-03 18:31:26 +02:00
src Avoid blocking AntiEntropyStage when submitting validation requests 2020-06-09 13:41:05 +01:00
test Avoid blocking AntiEntropyStage when submitting validation requests 2020-06-09 13:41:05 +01:00
tools Fix tools/bin/fqltool for all shells 2020-05-19 16:36:09 +02:00
.gitignore Merge branch 'cassandra-3.11' into trunk 2020-06-09 09:48:14 +02:00
.rat-excludes Remove Pig support 2015-10-16 13:23:20 +01:00
CASSANDRA-14092.txt Merge branch 'cassandra-2.2' into cassandra-3.0 2018-02-10 14:57:53 -02:00
CHANGES.txt Avoid blocking AntiEntropyStage when submitting validation requests 2020-06-09 13:41:05 +01:00
CONTRIBUTING.md Remove comment from CONTRIBUTING.md regarding closing GitHub PRs. 2019-01-09 14:17:19 +00:00
LICENSE.txt merge with 0.6 branch (post-850) 2010-03-26 16:31:52 +00:00
NEWS.txt Fix CQLSH UTF-8 encoding issue for Python 2/3 compatibility 2020-04-22 12:44:47 -07:00
NOTICE.txt Remove Java Driver dependency for UDFs and UDAs / Limit the dependencies used by UDFs/UDAs 2020-02-12 09:58:43 +01:00
README.asc Update README.asc 2020-03-30 14:16:00 -07:00
TESTING.md Ninja trivial typo in TESTING 2018-06-20 21:28:09 -07:00
build.properties.default Switch http to https URLs in build.xml 2019-05-22 15:03:43 -04:00
build.xml Merge branch 'cassandra-3.11' into trunk 2020-06-09 09:48:14 +02:00
eclipse_compiler.properties Add Static Analysis to warn on unsafe use of Autocloseable instances 2015-05-27 17:53:26 -04:00

README.asc

Apache Cassandra
-----------------

Apache Cassandra is a highly-scalable partitioned row store. Rows are organized into tables with a required primary key.

http://wiki.apache.org/cassandra/Partitioners[Partitioning] means that Cassandra can distribute your data across multiple machines in an application-transparent matter. Cassandra will automatically repartition as machines are added and removed from the cluster.

http://wiki.apache.org/cassandra/DataModel[Row store] means that like relational databases, Cassandra organizes data by rows and columns. The Cassandra Query Language (CQL) is a close relative of SQL.

For more information, see http://cassandra.apache.org/[the Apache Cassandra web site].

Requirements
------------
. Java >= 1.8 (OpenJDK and Oracle JVMS have been tested)
. Python 2.7 (for cqlsh)

Getting started
---------------

This short guide will walk you through getting a basic one node cluster up
and running, and demonstrate some simple reads and writes. For a more-complete guide, please see the Apache Cassandra website's http://cassandra.apache.org/doc/latest/getting_started/[Getting Started Guide].

First, we'll unpack our archive:

  $ tar -zxvf apache-cassandra-$VERSION.tar.gz
  $ cd apache-cassandra-$VERSION

After that we start the server. Running the startup script with the -f argument will cause
Cassandra to remain in the foreground and log to standard out; it can be stopped with ctrl-C.

  $ bin/cassandra -f

****
Note for Windows users: to install Cassandra as a service, download
http://commons.apache.org/daemon/procrun.html[Procrun], set the
PRUNSRV environment variable to the full path of prunsrv (e.g.,
C:\procrun\prunsrv.exe), and run "bin\cassandra.bat install".
Similarly, "uninstall" will remove the service.
****

Now let's try to read and write some data using the Cassandra Query Language:

  $ bin/cqlsh

The command line client is interactive so if everything worked you should
be sitting in front of a prompt:

----
Connected to Test Cluster at localhost:9160.
[cqlsh 2.2.0 | Cassandra 1.2.0 | CQL spec 3.0.0 | Thrift protocol 19.35.0]
Use HELP for help.
cqlsh> 
----

As the banner says, you can use 'help;' or '?' to see what CQL has to
offer, and 'quit;' or 'exit;' when you've had enough fun. But lets try
something slightly more interesting:

----
cqlsh> CREATE SCHEMA schema1 
       WITH replication = { 'class' : 'SimpleStrategy', 'replication_factor' : 1 };
cqlsh> USE schema1;
cqlsh:Schema1> CREATE TABLE users (
                 user_id varchar PRIMARY KEY,
                 first varchar,
                 last varchar,
                 age int
               );
cqlsh:Schema1> INSERT INTO users (user_id, first, last, age) 
               VALUES ('jsmith', 'John', 'Smith', 42);
cqlsh:Schema1> SELECT * FROM users;
 user_id | age | first | last
---------+-----+-------+-------
  jsmith |  42 |  john | smith
 cqlsh:Schema1> 
----

If your session looks similar to what's above, congrats, your single node
cluster is operational! 

For more on what commands are supported by CQL, see
http://cassandra.apache.org/doc/latest/cql/[the CQL reference]. A
reasonable way to think of it is as, "SQL minus joins and subqueries, plus collections."

Wondering where to go from here?

  * Join us in #cassandra on the https://s.apache.org/slack-invite[ASF Slack] and ask questions
  * Subscribe to the Users mailing list by sending a mail to
    user-subscribe@cassandra.apache.org
  * Visit the http://cassandra.apache.org/community/[community section] of the Cassandra website for more information on getting involved.
  * Visit the http://cassandra.apache.org/doc/latest/development/index.html[development section] of the Cassandra website for more information on how to contribute.