Go to file
Michael Marshall 5172a0df3b Dynamically skip sharding L0 when SAI Vector index present
This is a partial solution to IllegalStateException thrown by VectorPostings. It works by using a single shard at L0 when a vector index is present. As noted in the jira ticket, there are edge cases that may still produce errors, notably the case where there are multiple data directories.
The key trade offs here are related to the time complexity for search. Since graph search is log(n), and searching m graphs is m * log(n), we see better search performance by building bigger graphs which is essentially log(m * n). We could pre-shard, which comes at a cost of increased search time complexity.

patch by Michael Marshall,Dmitry Konstantinov; reviewed by Caleb Rackliffe,Dmitry Konstantinov,Michael Semb Wever for CASSANDRA-19661

Co-authored-by: Michael Marshall <mmarshall@apache.org>
Co-authored-by: Dmitry Konstantinov <netudima@gmail.com>
2026-02-26 12:45:05 +00:00
.build Cleanup of dependency-check-suppressions.xml, suppressing CVE-2025-67735 2026-02-17 11:40:01 +11:00
.circleci Minor bugs in generate.sh -d 2024-03-20 08:06:29 +01:00
.github Add a github action that runs .build/docker/check-code.sh 2025-10-13 12:34:51 +02:00
.jenkins Improve diagnostics for JUnit tests which crashed or timed out on Jenkins 2026-02-18 22:53:39 +00:00
bin Do not source cassandra-env.sh unnecessarily in nodetool and other tooling 2025-07-31 09:15:37 +02:00
conf Upgrade logback version to 1.5.18 and slf4j dependencies to 2.0.17 2026-02-06 13:28:03 +11:00
debian Prepare debian changelog for 5.0.6 2025-10-21 13:42:35 +02:00
doc Upgrade logback version to 1.5.18 and slf4j dependencies to 2.0.17 2026-02-06 13:28:03 +11:00
examples Merge branch 'cassandra-4.1' into cassandra-5.0 2023-09-25 14:26:14 -06:00
ide Merge branch 'cassandra-4.1' into cassandra-5.0 2025-10-24 14:26:39 -04:00
lib Upgrade Python driver to 3.29.0 2024-01-19 17:14:57 +00:00
pylib Merge branch 'cassandra-4.1' into cassandra-5.0 2025-12-19 23:44:01 +01:00
redhat Merge branch 'cassandra-4.1' into cassandra-5.0 2025-03-30 09:32:10 +02:00
src Dynamically skip sharding L0 when SAI Vector index present 2026-02-26 12:45:05 +00:00
test Dynamically skip sharding L0 when SAI Vector index present 2026-02-26 12:45:05 +00:00
tools Update jackson-dataformat-yaml to 2.19.2 and snakeyaml to 2.1 2025-09-19 13:29:05 +02:00
.asf.yaml Notify the corresponding JIRA issue as soon as the PR is raised 2023-03-29 06:24:34 -05:00
.gitignore Merge branch 'cassandra-4.1' into cassandra-5.0 2025-09-08 00:38:46 +01:00
.snyk Cleanup of dependency-check-suppressions.xml, suppressing CVE-2025-67735 2026-02-17 11:40:01 +11:00
CASSANDRA-14092.txt Default to nb instead of nc for sstable formats 2023-11-13 09:26:11 +01:00
CHANGES.txt Dynamically skip sharding L0 when SAI Vector index present 2026-02-26 12:45:05 +00:00
CONTRIBUTING.md Merge branch 'cassandra-3.11' into trunk 2021-04-22 08:32:58 -05:00
LICENSE.txt Merge branch 'cassandra-3.11' into cassandra-4.0 2023-08-31 22:39:56 +02:00
NEWS.txt Merge branch 'cassandra-4.1' into cassandra-5.0 2025-12-18 10:59:09 -08:00
NOTICE.txt Merge branch 'cassandra-3.11' into cassandra-4.0 2023-02-22 10:25:08 -06:00
README.asc Merge branch 'cassandra-4.1' into cassandra-5.0 2026-01-20 15:27:10 +01:00
TESTING.md Improve and clean up documentation and fix typos 2023-01-26 14:42:47 +01:00
build-shaded-dtest-jar.sh Merge branch 'cassandra-3.11' into trunk 2021-04-19 17:39:10 +02:00
build.properties.default Add snapshot remote repo to build resolution and build.properties.default 2024-09-16 15:49:14 -04:00
build.xml Implement microbench test target type 2026-01-27 11:31:17 +01:00
relocate-dependencies.pom Update maven-shade-plugin to version 3.6.1 2026-02-25 09:13:04 +01:00

README.asc

Apache Cassandra
-----------------

Apache Cassandra is a highly-scalable partitioned row store. Rows are organized into tables with a required primary key.

https://cwiki.apache.org/confluence/display/CASSANDRA2/Partitioners[Partitioning] means that Cassandra can distribute your data across multiple machines in an application-transparent matter. Cassandra will automatically repartition as machines are added and removed from the cluster.

https://cwiki.apache.org/confluence/display/CASSANDRA2/DataModel[Row store] means that like relational databases, Cassandra organizes data by rows and columns. The Cassandra Query Language (CQL) is a close relative of SQL.

For more information, see https://cassandra.apache.org/[the Apache Cassandra web site].

Issues should be reported on https://issues.apache.org/jira/projects/CASSANDRA/issues/[The Cassandra Jira].

Requirements
------------
- Java: see supported versions in build.xml (search for property "java.supported").
- Python: for `cqlsh`, see `bin/cqlsh` (search for function "is_supported_version").


Getting started
---------------

This short guide will walk you through getting a basic one node cluster up
and running, and demonstrate some simple reads and writes. For a more-complete guide, please see the Apache Cassandra website's https://cassandra.apache.org/doc/5.0/cassandra/getting-started/index.html[Getting Started Guide].

First, we'll unpack our archive:

  $ tar -zxvf apache-cassandra-$VERSION.tar.gz
  $ cd apache-cassandra-$VERSION

After that we start the server. Running the startup script with the -f argument will cause
Cassandra to remain in the foreground and log to standard out; it can be stopped with ctrl-C.

  $ bin/cassandra -f

Now let's try to read and write some data using the Cassandra Query Language:

  $ bin/cqlsh

The command line client is interactive so if everything worked you should
be sitting in front of a prompt:

----
Connected to Test Cluster at localhost:9160.
[cqlsh 6.2.0 | Cassandra 5.0-SNAPSHOT | CQL spec 3.4.7 | Native protocol v5]
Use HELP for help.
cqlsh>
----

As the banner says, you can use 'help;' or '?' to see what CQL has to
offer, and 'quit;' or 'exit;' when you've had enough fun. But lets try
something slightly more interesting:

----
cqlsh> CREATE KEYSPACE schema1
       WITH replication = { 'class' : 'SimpleStrategy', 'replication_factor' : 1 };
cqlsh> USE schema1;
cqlsh:Schema1> CREATE TABLE users (
                 user_id varchar PRIMARY KEY,
                 first varchar,
                 last varchar,
                 age int
               );
cqlsh:Schema1> INSERT INTO users (user_id, first, last, age)
               VALUES ('jsmith', 'John', 'Smith', 42);
cqlsh:Schema1> SELECT * FROM users;
 user_id | age | first | last
---------+-----+-------+-------
  jsmith |  42 |  john | smith
cqlsh:Schema1>
----

If your session looks similar to what's above, congrats, your single node
cluster is operational!

For more on what commands are supported by CQL, see
https://cassandra.apache.org/doc/5.0/cassandra/developing/cql/index.html[the CQL reference]. A
reasonable way to think of it is as, "SQL minus joins and subqueries, plus collections."

Wondering where to go from here?

  * Join us in #cassandra on the https://s.apache.org/slack-invite[ASF Slack] and ask questions.
  * Subscribe to the Users mailing list by sending a mail to
    user-subscribe@cassandra.apache.org.
  * Subscribe to the Developer mailing list by sending a mail to
    dev-subscribe@cassandra.apache.org.
  * Visit the https://cassandra.apache.org/community/[community section] of the Cassandra website for more information on getting involved.
  * Visit the https://cassandra.apache.org/doc/latest/development/index.html[development section] of the Cassandra website for more information on how to contribute.