- Don't try to update metrics when no TxnId (e.g. GetMaxConflict)
- maybeExecuteImmediately did not guarantee mutual exclusivity
- Cancellation of a running 'plain' task could corrupt AccordExecutor state
- GetLatestDeps message serializer
- txn_blocked_by StackOverflowError
- Not updating CommandsForKey in all cases on restart
Also Improve:
- TableId.from/toString
- Route toString methods
- Tracing coverage of FetchRoute
- Replay command store parallelism
patch by Benedict; reviewed by Alex Petrov for CASSANDRA-20763
- Use AsymmetricComparator for seeking
- Reimplement AbstractTransformer for more efficienct operation when only transforming a subset (i.e. Subtraction)
- Reuse source node IntervalMaxIndex calculations for first branch level (to save iterating leaves)
patch by Benedict; reviewed by Alex Petrov for CASSANDRA-20756
- WatermarkCollector should not report same closed/retired epoch N times
- AccordSyncPropagator can merge pending requests and back-off retries
- AccordCommandLoader should notify listeners
- AccordSegmentCompactor should estimate number of keys to ensure bloom filters work
- txn_blocked_by table should report what it can, not throw IllegalStateException
- Permit uncompressed system tables
patch by Benedict; reviewed by Alex Petrov for CASSANDRA-20754
- Journal debugging vtable support
- Background task tracing support
Fix:
- HLC_BOUND only valid for strictly lower HLC
- HAS_UNIQUE_HLC can only be safely computed if READY_TO_EXECUTE
- Break recursion in CommandStore.ensureReadyToCoordinate
- Fix find intersecting shard scheduler
- Separate adhoc ShardScheduler from normal to avoid overwriting
- Should still use execution listeners to detect invalid reads for certain transactions even with dataStoreDetectsFutureReads
patch by Benedict; reviewed by Alex Petrov for CASSANDRA-20746
- Lock inversion when restarting durability for erased sync point
- Initialise durability cycle start time
Improve:
- Improve durability reporting
- Timeout durability requests
- DurabilityQueue concurrency should limit only until quorum is achieved
- Some exceptions that may be thrown when starting coordination may not be propagated
patch by Benedict; reviewed by Alex Petrov for CASSANDRA-20328
- Cannot shrink CommandsForKey if loading pruned
- NoSpamLogger to suppress AccordSyncPropagator repeats
- Minor performance improvement to BTree.Subtraction
- EphemeralReads should retry in a later epoch if replication factor changes
- Fix no local epoch for NotAccept
- Invalidate a command with no route but for which we know we own some key that has applied locally
- Don't pre-merge existing DurableBefore, to reduce duplicate/redundant persistence
- Misc purging bugs and improve testing of purging
patch by Benedict; reviewed by Alex Petrov for CASSANDRA-20739
- Fix shouldCleanup handling of erase/expunge
- Fix CommandChange handling of minUniqueHlc being cleared
- Don't clear minUniqueHlc when fast applying; instead simply validate !isWaiting
patch by Benedict; reviewed by Alex Petrov for CASSANDRA-20726
- Attempt to fix CommandsForKey StackOverflowError (presumed to be reachable via SaveStatus->InternalStatus map returning null)
- Bound recursion with SafeCommandStore.tryRecurse()
- IndexOutOfBoundsException in CINTIA checkpoint list encoding bounds logic
- Maintain MaxDecidedRX to save majority of work when deciding RX that are newer than those we have previously agreed
- Introduce IntervalBTree and use in CommandsForRanges to limit time spent in critical section when there are millions of RX
patch by Benedict; reviewed by Alex Petrov for CASSANDRA-20727
- Introduce SequentialAgentExecutor
- ExecuteEphemeralRead.LocalExecute (matching ExecuteTxn.LocalExecute)
- remove AgentExecutor (agent() only used as implementation detail)
- Skip CommandStore/CommandsForKey on read or apply, when safe to do so
- Implement maybeExecuteImmediately
- Faster maxTimestamp
Fix:
- Marking vestigial without knowing all covering keys
- Incorrect size in WaitingOn.none
- Cleanup RX that are not known but are no longer needed, due to being defunct
- Home should not abort because of local truncation; may still be incomplete on another shard
- Invalidate infinite loop due to interpreting Vestigial as a possible decision (as though it were another Truncated state)
- Known.isTruncated reporting incorrect answer in some cases, leading to infinite RecoverWithRoute loops as not slicing to correct subset and some shards are truncated and reject the deps collection
- BurnTest slowCoordinatorDelay should grow with retryCount
- NotAccept should check both ballot and saveStatus; PreCommitted or later can be propagated by Apply that has an earlier ballot
- Handle awaiting locally retired epoch
- Fix propagate low epoch of sync points
- Fix GetEphemeralReadDeps.reduce NPE
- Don't use tombstones to ERASE journal entries, as want to produce an Erased SaveStatus until EXPUNGE
patch by Benedict; reviewed by Alex Petrov for CASSANDRA-20726
* Fix typos in AccordDebugKeyspace
* Make durability service listen to topology changes
* Account for truncated epoch when computing the low bound
* Add a test to make sure GC issues do not slip past us in the future
* Fix compilation
Patch by Alex Petrov; reviewed by Benedict Elliott Smith for CASSANDRA-20718
* Further relax conditions of 'unsafe' startup
* Allow skipping unfinished allocations
* Relax assertion: allow same key accross multiple different stores
* Fix SegmentTest
Patch by Alex Petrov; reviewed by Benedict Elliott Smith for CASSANDRA-20718
* Ignore storage compat mode errors
* Fix restart test (Cannot invoke "org.slf4j.Logger.debug(String, Object, Object)" because "this.inInstancelogger" is null)
* Always speculate in accord replacement test to avoid read timeouts
* Make sure not to cleanup fields before persisting them
Background: we would clean up redundantBefore and some other fields.
We are calling persistFieldUpdates, during which we "flush" field updates to command store, but then set it back to null:
fieldUpdates = null;
so when we get to persistFieldUpdatesInternal in AccordTask, we don't have any field updates.
Patch by Alex Petrov; reviewed by Benedict Elliott Smith for CASSANDRA-20718
Also Fix:
- slowCoordinatorDelay should bound its start point, and log more detail if its expectations are breached
- Erased SaveStatus can be reported for all queried owned participants on a replica (the saved participants have been erased)
patch by Benedict; reviewed by Alex Petrov for CASSANDRA-20722
SAI predicate search currently has a bug that could result in missing rows due to a
concurrent flush during a query. The new test created in this PR shows the point of failure.
The problem is that we get the SSTable index references before getting the Memtable index
references. Note that we do it in the correct order in the ANN OF query path, but not
in the WHERE query path.
This commit updates the QueryView object to hold references to the appropriate Memtable indexes.
It also removes the problematic search methods from MemtableIndexManager to prevent future misuse.
patch by Michael Marshall; reviewed by Caleb Rackliffe and Ekaterina Dimitrova for CASSANDRA-20709
- AccordJournal should not attempt to construct expunged commands
- ready/notReady logic can break if we receive a partial Route that does not cover the local store
- Invariant.require -> TopologyRetiredException
- validate min/max in TopologyManager.preciseEpochs
Improve:
- ExecuteSyncPoint should send Apply
patch by Benedict; reviewed by Alex Petrov for CASSANDRA-20713
- Do not query local topology when deciding what keys to fetch to avoid TopologyRetiredException
- Ensure we propagate information back to the requesting CommandStore by using the store id to ensure it is included
- BurnTest topology fetching was broken by earlier patch
- Topology callbacks were not being invoked as we were not calling .begin()
- Topology mismatch failure during notAccept phase was not being reported due to CoordinatePreAccept already having isDone==true
patch by Benedict; reviewed by David Capwell for CASSANDRA-20711
- Don't assume stillExecutes applies to remote request - must mark unavailable any keys we don't have the txn definition for
- Don't exit notifyManagedPreBootstrap notify loop early, as could have later transactions with lower txnId
- CoordinateEphemeralRead must retry if insufficient responses from replicas that still own the range
- Burn test terminates while non-recurring tasks pending on command store queues
- Infinite recovery due to not sending InformDurable when partially truncated
- Incorrect participants when invoking removeRedundantDependencies
- Infinite bootstrap loop on retired topology
Improve
- Do not perform linear filters of CommandStores or Topologies
- Some default implementations of forEach/iterator
- TopologyRetiredException messages
patch by Benedict; reviewed by Alex Petrov for CASSANDRA-20707
- cfk pruning+prebootstrap=invalid future dependency
- exclude retired ranges when filtering RX stillTouches
- propagate uses incorrect lowEpoch when fetch finds additional owned/touched ranges
- node.withEpoch should callback with TopologyRetiredException, not throw
- Recovery can race with durable-applied pruning; must not send durable unless latest ballot on apply
- removeRedundantDependencies was not slicing pre-bootstrap range calculation to participating ranges
- NPE in TopologyManager.atLeast caused by referencing an epoch that has been GC'd
- use journal durableBeforePersister in burn test, not NOOP_PERSISTER
- ServerUtils.cleanupDirectory use tryDeleteRecursive
- FsyncRunnable shutdown
- fix NPE in AccordJournalBurnTest
patch by Benedict; reviewed by Alex Petrov for CASSANDRA-20688
Updates SystemKeyspace.writePreparedStatement to accept a timestamp
associated with the Prepared creation time. Using this timestamp
will ensure that an INSERT into system.prepared_statements will
always precede the timestamp for the same Prepared in
SystemKeyspace.removePreparedStatement.
This is needed because Caffeine 2.9.2 may evict an entry as soon
as it is inserted if the maximum weight of the cache is exceeded
causing the DELETE to be executed before the INSERT.
Additionally, any clusters currently experiencing a leaky
system.prepared_statements table from this bug may struggle to
bounce into a version with this fix as
SystemKeyspace.loadPreparedPreparedStatements currently does
not paginate the query to system.prepared_statements, causing heap
OOMs. To fix this this patch adds pagination at 5000 rows and
aborts loading once the cache size is loaded. This should allow
nodes to come up and delete older prepared statements that may no
longer be used as the cache fills up (which should happen immediately).
This patch does not address the issue of Caffeine immediately evicting
a prepared statement, however it will prevent the
system.prepared_statements table from growing unbounded. For most users
this should be adequate, as the cache should only be filled when there
are erroneously many unique prepared statements. In such a case we can
expect that clients will constantly prepare statements regardless
of whether or not the cache is evicting statements.
patch by Andy Tolbert; reviewed by Berenguer Blasi and Caleb Rackliffe for CASSANDRA-19703
As part of a broader effort to decouple java driver code from the
server code, this moves sstableloader to its own tools directory.
As sstableloader is also used as a library (CASSANDRA-10637), added
a new artifact 'cassandra-sstableloader' that will get deployed to
maven along with 'cassandra-all'.
While I expect this is likely a niche use case, this will allow users
to continue using BulkExport as a library.
Moves sstableloader-specific targets to its own build.xml in
tools/sstableloader/build.xml.
Also updates IDE project files and circleci to utilize new
sstableloader-specific targets.
patch by Andy Tolbert; reviewed by Stefan Miklosovic and Mick Semb Wever for CASSANDRA-20328
patch by Stefan Miklosovic; reviewed by Bernardo Botella Corbi for CASSANDRA-20563
Co-authored-by: Bernardo Botella Corbi <contacto@bernardobotella.com>
Memory-mapping is done in buffers of size less than 2GiB.
When these buffers aren't aligned to 4KiB and the trie-index file
spans many buffers then reading it results in going out of buffer
bounds.
This patch fixes it by making sure that the buffers are correctly
aligned.
patch by Szymon Miezal; reviewed by blambov and brandonwilliams for CASSANDRA-20351
Memory-mapping is done in buffers of size less than 2GiB.
When these buffers aren't aligned to 4KiB and the trie-index file
spans many buffers then reading it results in going out of buffer
bounds.
This patch fixes it by making sure that the buffers are correctly
aligned.
patch by Szymon Miezal; reviewed by blambov and brandonwilliams for CASSANDRA-20351
* cassandra-5.0:
Add LittleEndianMemoryUtil and NativeEndianMemoryUtil, switch memtable-related off-heap objects and Memory to use them and have Little Endian now. Add BE offsets detection on Summary loading. Add test SSTables in an old format with BE offsets in Summary component to LegacySSTableTest.
Add BE offsets detection on Summary loading.
Add test SSTables in an old format with BE offsets in Summary component to LegacySSTableTest.
Patch by Dmitry Konstantinov; reviewed by Branimir Lambov, Michael Semb Wever for CASSANDRA-20190