* Fix typos in AccordDebugKeyspace
* Make durability service listen to topology changes
* Account for truncated epoch when computing the low bound
* Add a test to make sure GC issues do not slip past us in the future
* Fix compilation
Patch by Alex Petrov; reviewed by Benedict Elliott Smith for CASSANDRA-20718
* Further relax conditions of 'unsafe' startup
* Allow skipping unfinished allocations
* Relax assertion: allow same key accross multiple different stores
* Fix SegmentTest
Patch by Alex Petrov; reviewed by Benedict Elliott Smith for CASSANDRA-20718
* Make sure to wait for Mutation stage quiesence during replay
* Fix visiting same key multiple times during replay
Patch by Alex Petrov; reviewed by Benedict Elliott Smith for CASSANDRA-20675
Background: we want to have an escape hatch to pull / copy journal files to investigate locally, but at that stage, we would rather allow the cluster to skip a log entry rather than wipe an entire cluster.
Patch by Alex Petrov; reviewed by Benedict Elliott Smith for CASSANDRA-20702
* Ignore storage compat mode errors
* Fix restart test (Cannot invoke "org.slf4j.Logger.debug(String, Object, Object)" because "this.inInstancelogger" is null)
* Always speculate in accord replacement test to avoid read timeouts
* Make sure not to cleanup fields before persisting them
Background: we would clean up redundantBefore and some other fields.
We are calling persistFieldUpdates, during which we "flush" field updates to command store, but then set it back to null:
fieldUpdates = null;
so when we get to persistFieldUpdatesInternal in AccordTask, we don't have any field updates.
Patch by Alex Petrov; reviewed by Benedict Elliott Smith for CASSANDRA-20718
Also Fix:
- slowCoordinatorDelay should bound its start point, and log more detail if its expectations are breached
- Erased SaveStatus can be reported for all queried owned participants on a replica (the saved participants have been erased)
patch by Benedict; reviewed by Alex Petrov for CASSANDRA-20722
SAI predicate search currently has a bug that could result in missing rows due to a
concurrent flush during a query. The new test created in this PR shows the point of failure.
The problem is that we get the SSTable index references before getting the Memtable index
references. Note that we do it in the correct order in the ANN OF query path, but not
in the WHERE query path.
This commit updates the QueryView object to hold references to the appropriate Memtable indexes.
It also removes the problematic search methods from MemtableIndexManager to prevent future misuse.
patch by Michael Marshall; reviewed by Caleb Rackliffe and Ekaterina Dimitrova for CASSANDRA-20709
- AccordJournal should not attempt to construct expunged commands
- ready/notReady logic can break if we receive a partial Route that does not cover the local store
- Invariant.require -> TopologyRetiredException
- validate min/max in TopologyManager.preciseEpochs
Improve:
- ExecuteSyncPoint should send Apply
patch by Benedict; reviewed by Alex Petrov for CASSANDRA-20713
- Do not query local topology when deciding what keys to fetch to avoid TopologyRetiredException
- Ensure we propagate information back to the requesting CommandStore by using the store id to ensure it is included
- BurnTest topology fetching was broken by earlier patch
- Topology callbacks were not being invoked as we were not calling .begin()
- Topology mismatch failure during notAccept phase was not being reported due to CoordinatePreAccept already having isDone==true
patch by Benedict; reviewed by David Capwell for CASSANDRA-20711
- Don't assume stillExecutes applies to remote request - must mark unavailable any keys we don't have the txn definition for
- Don't exit notifyManagedPreBootstrap notify loop early, as could have later transactions with lower txnId
- CoordinateEphemeralRead must retry if insufficient responses from replicas that still own the range
- Burn test terminates while non-recurring tasks pending on command store queues
- Infinite recovery due to not sending InformDurable when partially truncated
- Incorrect participants when invoking removeRedundantDependencies
- Infinite bootstrap loop on retired topology
Improve
- Do not perform linear filters of CommandStores or Topologies
- Some default implementations of forEach/iterator
- TopologyRetiredException messages
patch by Benedict; reviewed by Alex Petrov for CASSANDRA-20707
- cfk pruning+prebootstrap=invalid future dependency
- exclude retired ranges when filtering RX stillTouches
- propagate uses incorrect lowEpoch when fetch finds additional owned/touched ranges
- node.withEpoch should callback with TopologyRetiredException, not throw
- Recovery can race with durable-applied pruning; must not send durable unless latest ballot on apply
- removeRedundantDependencies was not slicing pre-bootstrap range calculation to participating ranges
- NPE in TopologyManager.atLeast caused by referencing an epoch that has been GC'd
- use journal durableBeforePersister in burn test, not NOOP_PERSISTER
- ServerUtils.cleanupDirectory use tryDeleteRecursive
- FsyncRunnable shutdown
- fix NPE in AccordJournalBurnTest
patch by Benedict; reviewed by Alex Petrov for CASSANDRA-20688