Commit Graph

8861 Commits

Author SHA1 Message Date
David Capwell 3fb23e0247 Fix flakey test org.apache.cassandra.simulator.test.ShortPaxosSimulationTest#casOnAccordSimulationTest
patch by David Capwell; reviewed by Ariel Weisberg for CASSANDRA-20322
2025-04-17 11:59:55 -07:00
David Capwell b81d13f994 Expose epoch ready state as a vtable for operators to inspect things
patch by David Capwell; reviewed by Benedict Elliott Smith for CASSANDRA-20302
2025-04-17 11:59:55 -07:00
Alex Petrov 1aaf0bb336 Fix null field accounting
Patch by Alex Petrov; reviewed by Benedict Elliott Smith for CASSANDRA-20297
2025-04-17 11:59:55 -07:00
Alex Petrov 3324e23f64 Fix topology loading
Use initializeTopologyUnsafe to for-load the last topology and create shards rather than replaying all topologies.

Patch by Alex Petrov; reviewed by Benedict Elliott Smith for CASSANDRA-20294
2025-04-17 11:59:55 -07:00
David Capwell 5e9e2b2b04 Accord: TableParamsTest is flakey due to bad generators and production validation logic missed the argument to String.format causing confusing errors
patch by David Capwell; reviewed by Blake Eggleston for CASSANDRA-20286
2025-04-17 11:59:55 -07:00
Benedict Elliott Smith b017ae8175 Add @Ignore to failing tests while we triage 2025-04-17 11:59:55 -07:00
Benedict Elliott Smith d7ed1ad391 Fix:
- Even if we can decide autonomously that we took the fast path, we must still wait for earlier transactions to decide themselves
 - We must update a command that is in CFK.loadingPruned whether or not it is outOfRange
 - We must visit pruned commands that are transitive dependencies of RX to ensure dependencies are propagated
 - flagsWithoutDomainAndKind() -> flagsWithoutDomainOrKindOrCardinality() [to fix integration regression]
 - Don't invoke uniqueNow() twice when allocate nextTxnId
 - BTreeReducingRangeMap edge case that corrupts map when merge function does not return one of its inputs

patch by Benedict; reviewed by Alex Petrov for CASSANDRA-20292
2025-04-17 11:59:55 -07:00
Alex Petrov 080d6c633d Avoid double loading in Cassandra side of journal; make sure to include records appended to journal.
Patch by Alex Petrov; reviewed by Benedict Elliott Smith for CASSANDRA-20244
2025-04-17 11:59:55 -07:00
David Capwell 4f2e306c26 Accord: jvm-dtest changed semantics of uncaughtExceptions handling which broke tests that depended on the prior semantics
patch by David Capwell; reviewed by Benedict Elliott Smith for CASSANDRA-20248
2025-04-17 11:59:54 -07:00
David Capwell f345011822 When generating AbstractTypeGenerators.safeTypeGen() for tests composite types limit depth but collections didnt
patch by David Capwell; reviewed by Caleb Rackliffe for CASSANDRA-20283
2025-04-17 11:59:54 -07:00
Benedict Elliott Smith af7dfcc327 Refactor RedundantStatus to encode vector of states that can be merged independently
Also fix:
 - Truncate command on first access, without participants
 - Use Ballot.ZERO when invoking CFK.insertOutOfRange where appropriate
 - Don't supply a command's own route to ProgressLog.waiting to ensure new keys are incorporated
 - Ensure progress in CommandsForKey by setting vestigial commands to ERASED
 - Add any missing owned keys to StoreParticipants.route to ensure fetch can make progress
 - Recovery must wait for earlier not-accepted transactions if either has the privileged coordinator optimisation
 - Inclusive SyncPoint used incorrect topologies for propose phase
 - Barrier must not register local listener without up-to-date topology information
 - Stop home shard truncating a TxnId to vestigial rather than Invalidated so other shards can make progress
Also improve:
 - Validate commands are constructed with non-empty participants
 - Remove some unnecessary synchronized keywords
 - Clear ok messages on PreAccept and Accept to free up memory
 - Introduce TxnId.Cardinality flag so we can optimise single key queries
 - Update CommandsForKey serialization to better handle larger flag space
 - Configurable which Txn.Kind can result in a CommandStore being marked stale
 - Process DefaultProgressLog queue synchronously when relevant state is resident in memory
 - Remove defunct CollectMaxApplied version of ListStore bootstrap
 - Standardise linearizability violation reporting
 - Improve CommandStore.execute method naming to reduce chance of misuse
 - Prune and address some comments

patch by Benedict; reviewed by Alex Petrov for CASSANDRA-20282
2025-04-17 11:59:54 -07:00
Alex Petrov c4a5971af3 Improve topology fetching
- preclude TCM of being behind Accord if newer epoch is reported via withEpoch/fetchTopologyInternal
  - improve topology discovery during first boot and replay
  - fix races between config service TCM listener reporting topologies, and fetched topologies during

Patch by Alex Petrov, reviewed by Ariel Weisberg and David Capwell for CASSANDRA-20245
2025-04-17 11:59:54 -07:00
David Capwell d543cc9107 To improve accord interoperability test coverage, need to extend the harry model domain to handle more possible CQL states
patch by David Capwell; reviewed by Ariel Weisberg for CASSANDRA-20156
2025-04-17 11:59:54 -07:00
Benedict Elliott Smith 113d83417e Follow-up to CASSANDRA-20222:
- Implement missing parts of protocol optimisations, refine some particulars and remove MEDIUM_PATH_WAIT_ON_RECOVERY
Also fix:
 - Deps.txnIds -> Deps.txnIdsWithFlags to make clear unsafety, and validate Command isn't created with flags
 - Save a lossy low/high epoch we're waiting on in WaitingState so that we can reconstruct the same Route on callback
 - Load any potentially invalidated commands we had in ProgressLog to ensure they are marked invalidated locally and notify any waiters
 - Refine Invariant to accept case where a truncated dependency should apply first but both transactions are redundant and only the waiter could not be cleaned up
 - only update CFK.maxUniqueHlc using commands we execute
 - Don't update uniqueHlc when zero to avoid incompatible transaction invariant check

patch by Benedict; reviewed by Alex Petrov for CASSANDRA-20222
2025-04-17 11:59:54 -07:00
Ariel Weisberg f55653ecb0 Fix paging with Accord range reads
With paging it's possible for the range Accord needs to read from to have the start and end token be the same in which case the LHS token needs to be converted to a MinTokenKey

Patch by Ariel Weisberg; Reviewed by David Capwell for CASSANDRA-20262
2025-04-17 11:59:54 -07:00
Ariel Weisberg d88596b6a4 Fix ShortAccordSimulationTest
Patch by Ariel Weisberg; Reviewed by Alex Petrov for CASSANDRA-20264
2025-04-17 11:59:54 -07:00
Ariel Weisberg 6522d52b0b Live migration for non-serial reads
Patch by Ariel Weisberg; Reviewed by David Capwell for CASSANDRA-19439
2025-04-17 11:59:54 -07:00
Alex Petrov cf9e34df45 Fix empty LatestDepsSerializer
Patch by Alex Petrov; reviewed by Benedict Elliott Smith for CASSANDRA-20252
2025-04-17 11:59:54 -07:00
Benedict Elliott Smith 7f0e5f8e8b Follow-up to CASSANDRA-20228:
- Fix waiting state callback computes different route to initiator
 - Invariants.checkX -> Invariants.requireX (to allow complementary Invariants.expectX as appropriate)

patch by Benedict; reviewed by Alex Petrov for CASSANDRA-20228

Follow-up to CASSANDRA-20228:
 - Fix AccordUpdateParameters to correctly supply List cell paths derived from applyAt
 - Fix ExecuteAtSerialize.serialiseNullable
 - Fix topologies flush regression caused by markRetired forcing a flush on every call

patch by Benedict; reviewed by Ariel Weisberg for CASSANDRA-20228
2025-04-17 11:59:54 -07:00
Benedict Elliott Smith 5be214a417 Agree a distributed uniqueHlc to use for Apply
- detect and save whether RX can be used as a GC bound for HLCs
 - have RX durably record if they witnessed any superseding epoch and use this for HLC GC
 - rework GC to retain applyAt for writes

Also improve:
 - PreCommitted overwrites fast path vote, then loses tie-break with NotAccept, so we don't recover
 - track executeAtLeast and uniqueHlc separately from WaitingOn in CommandChange
 - Always propose executeAt.hlc() >= txnId.hlc()
 - Short-circuit recovery if ApplyAt already known
Also fix:
 - Dangling recovery when Commit rejected
 - Failed epoch fetch handling in DefaultProgressLog
 - CommandsForKey must track Accept executeAt for computing earlierWait collection on recovery
 - Report DefaultProgressLog exceptions before retrying
 - Fix inverted condition in RecoverWithRoute
 - Remove laterNoWait from laterWait collection in recovery
 - Invalidate incorrectly selected maxNotTruncated if max had a not truncated component
 - WaitingState.fetch not request/propagating prior epochs
 - CFK visit active pruning was comparing txn against maxCommittedWriteBefore rather than maxCommittedWriteForEpoch

patch by Benedict; reviewed by Alex Petrov for CASSANDRA-20228
2025-04-17 11:59:54 -07:00
Alex Petrov b50536f1a4 Fix Simulator Tests
Patch by Alex Petrov and Benedict Elliott Smith; reviewed by Alex Petrov and Benedict Elliott Smith for CASSANDRA-20240.
2025-04-17 11:59:54 -07:00
Aleksey Yeschenko b8a3706ef9 Remove Journal record ownership tags functionality
patch by Aleksey Yeschenko; reviewed by Benedict Elliott Smith for
CASSANDRA-20229
2025-04-17 11:59:54 -07:00
David Capwell 3c7cb1e3b3 Migrate route index from commands table to journal, and drop the commands table
patch by David Capwell; reviewed by Alex Petrov for CASSANDRA-20144
2025-04-17 11:59:54 -07:00
Benedict Elliott Smith 3e16efa490 Protocol optimisations:
- Privileged coordinator. If the coordinator is a replica we can reduce our quorum sizes by including the coordinator's vote.
   - with deps: if we include coordinator's preaccept deps we can reliably reduce quorum size by 1, at the expense of recovery sometimes requiring additional phases and waiting for future txns
   - with only vote: if we only include the vote we can avoid any additional recovery phases or waiting for future txns, but can reduce our quorum size for only some configurations
 - Medium path. If t=t0 at a simple majority we can take just two rounds.
   - with additional phase: on recovery, earlier txns must wait for medium path to be disabled (or committed) so we do not accidentally recover a transaction that won't be witnessed
   - with unstable deps: on slow path commit, deps not found in accept as committed as unstable and do not affect recovery decisions for earlier transactions
Also improve:
 - Recovery of await conditions simply recalculates the rejects flag rather than restarting to ensure faster forward progress
 - refactor Command hierarchy
 - tweak: don't save Writes for non-write transactions
Also fix:
 - isOutOfRange when invoked from callback for some txn < Committed
 - handle loadingPruned that is pre-bootstrap on update to bootstrappedAt
 - journal replay can overwrite in memory state loaded for a command not yet replayed
 - subtle ordering bug in LatestDeps.mergeProposal
 - fix occasional class initialisation order bug
 - handle race condition on learning of topology

patch by Benedict; reviewed by Aleksey Yeschenko for CASSANDRA-20222
2025-04-17 11:59:54 -07:00
Alex Petrov 13625288c6 Fix ForceSnapshotTest
Patch by Alex Petrov; reviewed by David Capwell for CASSANDRA-20223.
2025-04-17 11:59:54 -07:00
Alex Petrov 72f3828c0b Fix CoordinatorReadLatencyMetricTest
Patch by Alex Petrov; reviewed by David Capwell for CASSANDRA-20216
2025-04-17 11:59:54 -07:00
Alex Petrov 21f9f248cd Fix serialization order for topology updates
Patch by Alex Petrov; reviewed by David Capwell for CASSANDRA-20219.
2025-04-17 11:59:54 -07:00
Benedict Elliott Smith ed2ca6e2fb FetchRequest should report as unavailable any slice that executes in a later epoch that is not owned by the replicas
Improve:
 - Remove GetMaxConflict; use local MaxConflict collection

Fix:
 - Invalidated should retain StoreParticipants so we can update CFK on journal replay
 - loading pruned uninitialised commands via CFK: make sure hasTouched contains key so that if invalidated we are notified
 - updateExecuteAtLeast should always be higher than TxnId
 - Async CFK callbacks treated pruned transactions incorrectly
 - node.withEpoch when ExecuteEphemeralRead in futureEpoch
 - Deps.without should be key/range aware
 - Durably mark bootstrapBeganAt and safeToReadAt in MaxConflicts
 - Don't attempt to calculate local deps when DepsErased in GetLatestDeps (command has durably applied to this shard)
 - filter StoreParticipants before invoking shouldCleanup

patch by Benedict; reviewed by Alex Petrov for CASSANDRA-20183
2025-04-17 11:59:54 -07:00
Alex Petrov dfcf3aff4e Fix topology replay during bootstrap and startup, decouple Accord from TCM
Includes multiple changes, primary ones:
  * Make HarrySimulatorTest work with Accord
  * Removed nodes now live in TCM, no need to discover historic epochs in order to find removed nodes
  * CommandStore <-> RangesForEpochs mappings required for startup are now stored in journal, and CS can be set up _without_ topology replay
  * Topology replay is fully done via journal (where we store topologies themselves), and topology metadata table (where we store redundant/closed information)
  * Fixed various bugs related to propagation and staleness
     * TCM was previously relied on for "fetching" epoch: we can not rely on it as there's no guarantee we will see a consecutive epoch when grabbing Metadata#current
     * Redundant / closed during replay was set with incorrect ranges in 1 of the code paths
     * TCM was contacted multiple times for historical epochs, which made startup much longer under some circumstances

Patch by Alex Petrov; reviewed by  Benedict Elliott Smith for CASSANDRA-20142
2025-04-17 11:59:54 -07:00
David Capwell 17dbba770b Refactor the ast package to enable harry model based testing
patch by David Capwell; reviewed by Abe Ratnofsky, Caleb Rackliffe for CASSANDRA-20155
2025-04-17 11:59:54 -07:00
Alex Petrov b664d3e616 Migrate in memory journal to CommandChange logic shared with AccordJournal
Patch by Alex Petrov; reviewed by Benedict Elliott Smith for CASSANDRA-20115
2025-04-17 11:59:54 -07:00
Benedict Elliott Smith 52242f23a8 Remove TimestampsForKey
patch by Benedict; reviewed by Alex Petrov for CASSANDRA-20174
2025-04-17 11:59:53 -07:00
Benedict Elliott Smith 44cd181438 Fixes
- Fix notifying unmanaged after update redundant before/bootstrap
     - Do not infer invalid if we have a single round of replies with minKnown not decided and maxKnown erased - in this case store the knowledge for next request.
     - Fix SyncPoint topology selection
     - Fix CheckStatusOkFull.with(InvalidIf)
     - Fix NotifyWaitingOn
     - ExecuteTxn should only contact latest topology for follow-up requests
     - DurableBefore.min should not go backwards on new epoch topology, journal replay was not correctly handling PreApplied, partialTxn can be null if not owned
     - Fix notify pre-bootstrap that arrives post-bootstrap
     - Avoid GC race condition on Propagate where we can incorrectly infer a shard is stale
     - Ensure redundantBefore on previously-owned range does not imply redundant before for overlapping queries on still-owned range
     - Ensure we don't mark stale unless all of the quorum we contacted had erased, else we may have raced with the agreement and erase
     - Fix Invalidate when no route found for FetchData does not report to all requested local epochs
     - Fix WAS_OWNED_RETIRED without durableBefore at Universal can lead to assertions with RX that we permit to execute but that have not yet
     - Fix initialiseWaitingOn can in some cases transitively notify the command we're updating via maybeCleanup of dependencies, but the command isn't yet updated so isn't ready
     - Fix encountering a command that is pre-bootstrap, and for which we have locally 'applied' a supserseding RX, so that we do not know its outcome locally (so we do not cleanup the command), but also it must have been decided - and we should not respond with future dependencies.
     - Epoch failures on CoordinatePreAccept should trigger the CoordinatePreAccept failure handler
     - Use the shard bound rather than GC bound for fallback dependency
     - LatestDeps should be sliced to actual route, so as not to use both PreAccepted AND Stable deps as though Stable
     - Fix various callback issues with node.withEpoch and Recover/Propose.isDone
     - RecoverWithRoute can encounter a partially truncated transaction where the Deps for one shard are not committed. Must fetch LatestDeps.
     - Tighten LatestDeps semantics for Recover
     - CommandsForKey: do not restore pruned as APPLIED
     - Ensure prune points execute in the epoch in which they are declared
     - must merge all fast path votes including those from earlier epochs that may have witnessed a later transaction
     - Recoveries that know the transaction is committed a priori should skip the Accept phase
     - Maintain GC behaviour for redundant commands that are pre-bootstrap
     - don't apply ERASE to CommandsForKey to avoid breaking pruning
     - Introduce clearBefore to ProgressLog to more consistently handle cleaning up redundant transactions (and avoid triggering burn test invariants)
     - don't replay journal of a bootstrapping node in burn test
     - Recover, Accept or Commit reply from epoch that has been retired should be treated as Success rather than Redundant
     - Distinguish completely REDUNDANT+PRE_BOOTSTRAP from partially GC_BEFORE and REDUNDANT+PRE_BOOTSTRAP - latter can make stronger inferences based on the GC_BEFORE intersection (could perhaps be treated as simply GC_BEFORE)
     - RX must register historical transactions with CFK
     - CommandStore.bootstrapper must wait for coordinate sync via same mechanism as sync()
     - Don't start topology change for shard where all replicas are already bootstrapping
     - Reify executes et al in StoreParticipants
     - LocalListeners txn listener reentry may erase the entry entirely
     - use registerAt in AbstractRequest for expirations, use correct time for expiresAt in ListAgent
     - use txnId.epoch() for pruning, as must be before both txnId and executeAt of prune point for coordinating dependencies
     - compute accurate KnownMap when affected by bootstrap or staleness
     - upgradeTruncated should calculate Definition and Deps separately
     - Invalidate should not sort before Erased when calculating max reply or max knowledge reply
     - avoid another infinite loop at end of burn test
     - avoid another epoch loading edge case
     - pass through low/high epochs to ensure we propagate information to all waiting command stores
     - RX must adopt a non-pruned dependency that has a higher TxnId (if is itself behind prune point)
     - rejects should also be calculated on COMMITTED started before
     - remove Apply Factory wrapper for RX, redundant now we have CoordinationAdapters (and has faulty epoch logic)
     - for RX ensure we return maximum  writes for each epoch we intersect (same effectively as pruning logic)
     - rework updateUnmanaged to improve clarity
     - BeginRecovery constructor of LatestDeps should use touches() not owns() for compute localDeps
     - BeginRecovery superseding calculation was incorrectly treating startedBefore Committed and Accepted the same, when the point at which a dep should be known differs
     - Refactor Command visiting, porting C* integration to accord-core
     - RelationMultiMap Builder should resize keys and keyLimits independently
     - CommandsForKey Serialization moved to accord-core
     - losing ownership of range should trigger re-registration of unmanaged waiting on commit of a no-longer owned txn

    patch by Benedict; reviewed by Alex Petrov for CASSANDRA-20172
2025-04-17 11:59:53 -07:00
Ariel Weisberg 21383d2789 Fix Accord SAI tests and Accord double apply
Patch by Ariel Weisberg and David Capwell; Reviewed by David Capwell for CASSANDRA-20089
2025-04-17 11:59:53 -07:00
Alex Petrov 211249b8a9 Implement field saving/loading in AccordJournal
Patch by Alex Petrov; reviewed by Benedict Elliott Smith for CASSANDRA-20114
2025-04-17 11:59:53 -07:00
David Capwell c114fc46d2 Accord simulation test failing with "ClassNotFoundException: WARN [AccordExecutor[1,0]"
patch by David Capwell; reviewed by Benedict Elliott Smith for CASSANDRA-20127
2025-04-17 11:59:53 -07:00
Alex Petrov 74817bc83e Semi integrate burn test:
* Move serializers that previously were in static field into corresponding swappable accessors
  * Add a (no-op in Cassandra impl) Result serializer
  * Allow more than one Accord Journal table
  * Move some previously-Cassandra implementation details (such as FieldUpdates) to Accord
  * Fix a bug with sorting of Journal Key (incorrect non-identity flags)
  * Accord side: Make sure not to re-sort collection of events scheduled for application
  * Implement serializers for Accord/BurnTest List* items

Patch by Alex Petrov; reviewed by Benedict Elliott Smith for CASSANDRA-20112.
2025-04-17 11:59:53 -07:00
Alex Petrov bad9805b7a Make it easier to reuse generators, make Harry more extensible and accord-proof, refactor Harry's major subsystems
* SELECT is now a regular Harry operation, not a separate "validate" step
  * Flushes (or any custom operation) are now a part of a model, Harry will run them for you
  * Simplified validation of subset/wildcard queries
  * All Operations and are now SchemaSpec-agnostic, and can use any generators (like QT or whatever)
  * You read it right, Harry now can use any generator and still be able to not to keep generated values in memory
  * There are no concepts of “descriptor selectors”, or “bijection” exposed anymore. You just provide a Generator, Harry figures out the rest.
  * All Cassandra data types are now supported
  * No upper limit on the number of columns (also, clustering columns)
  * Basic (for now) support for Transactions
  * It is now possible to visit multiple partition in the same Visit (essentially, for transactions)
  * Much much simpler to generate synthetic workloads now (basically, you just write a Generator
  * Add support for all single-cell types.

Most of it was just a refactoring of Harry, to allow the following (still pending):

  * Easier to add subset SELECTION (important for SAI testing)
  * Easier add support for functions
  * Easier to add support for collections and UDTs
  * Easier to add support for LET bindings

Patch by Alex Petrov; reviewed by David Capwell, Caleb Rackliffe, and Abe Ratnofsky for CASSANDRA-20080.
2025-04-17 11:59:53 -07:00
Benedict Elliott Smith ec7c7c9fce Accord: Fix unit tests and improve burn test stability
patch by Benedict; reviewed by Alex Petrov for CASSANDRA-20113
2025-04-17 11:59:53 -07:00
Benedict Elliott Smith f2ea674140 Key transaction recovery should witness range transactions
patch by Benedict; reviewed by David for CASSANDRA-20105
2025-04-17 11:59:53 -07:00
Aleksey Yeschenko 304749ebda Set Accord debug tables partitioner to LocalPartitioner
patch by Aleksey Yeschenko; reviewed by Benedict Elliott Smith for
CASSANDRA-20062
2025-04-17 11:59:53 -07:00
Ariel Weisberg 16b5e191f9 Non-serial reads and range reads without live migration support
Patch by Ariel Weisberg; Reviewed by Benedict Elliott Smith for CASSANDRA-19437
2025-04-17 11:59:53 -07:00
Aleksey Yeschenko bcf1abbd16 Implement missing virtual tables for Accord debugging
patch by Aleksey Yeschenko; reviewed by Benedict Elliott Smith and David
Capwell for CASSANDRA-20062
2025-04-17 11:59:53 -07:00
Ariel Weisberg b29eac16fe Split accord migration into two phases
Patch by Ariel Weisberg; Reviewed by Benedict Elliott Smith for CASSANDRA-19436
2025-04-17 11:59:53 -07:00
David Capwell 513509ee2c Accord's ConfigService lock is held over large areas which cause deadlocks and performance issues
patch by David Capwell; reviewed by Benedict Elliott Smith for CASSANDRA-20065
2025-04-17 11:59:53 -07:00
Ariel Weisberg 6c0ad476ed Miscellaneous migration test fixes
Patch by Ariel Weisberg; Reviewed by David Capwell for CASSANDRA-20060
2025-04-17 11:59:53 -07:00
David Capwell 9fe1a977b5 Get Harry working on top of Accord and fix various issues found by TopologyMixupTestBase
patch by David Capwell; reviewed by Alex Petrov, David Capwell for CASSANDRA-20054
2025-04-17 11:59:53 -07:00
Benedict Elliott Smith 5a12875b3a Use ExclusiveSyncPoints to join a new topology
For correctness, the dependencies we adopt on joining a new topology must exclude the possibility of respondents accepting additional transactions with a lower TxnId, so proxying on the existing `ExclusiveSyncPoint` mechanisms is logical for the time-being. This patch removes the `FetchMajorityDeps` logic in favour of simply waiting for a suitable `ExclusiveSyncPoint` to be proposed.

patch by Benedict, reviewed by Alex Petrov for CASSANDRA-20056
2025-04-17 11:59:53 -07:00
Alex Petrov 4e9110e881 Check for splittable ranges
Patch by Alex Petrov; reviewed by Ariel Weisberg for CASSANDRA-20032

Accord Deps tests have incorrect range semantics

patch by David Capwell; reviewed by Ariel Weisberg for CASSANDRA-20029
2025-04-17 11:59:53 -07:00
David Capwell e94d86b1e5 TopologyMixupTestBase does not fix replication factor for Keyspaces after reaching rf=3
patch by David Capwell; reviewed by Alex Petrov for CASSANDRA-19975
2025-04-17 11:59:53 -07:00
David Capwell 1f80f99b71 Accord should block currently unsafe operations
patch by David Capwell; reviewed by Ariel Weisberg for CASSANDRA-20020
2025-04-17 11:59:53 -07:00
David Capwell 5afe99b7ca Accord metrics are isolated which cause existing coordination metrics to be empty, should also populate there as well
patch by David Capwell; reviewed by Benedict Elliott Smith for CASSANDRA-20017
2025-04-17 11:59:53 -07:00
Alex Petrov d1af562705 Shut down scheduler with "now"
Fix NPE in MockJournal on null onFlush
Fix SavedCommandTest.
After the serialization change that serializes "changed" before "is null", null flag can no be written.
2025-04-17 11:59:53 -07:00
Ariel Weisberg a51bed1c77 Accord should not block partition restricted index queries
Patch by Ariel Weisberg; Reviewed by David Capwell for CASSANDRA-19955

Non-serial single partition reads on Accord

Patch by Ariel Weisberg; Reviewed by Benedict Elliott Smith for CASSANDRA-19951
2025-04-17 11:59:53 -07:00
Alex Petrov 042a6e9769 Add bounce to load test 2025-04-17 11:59:52 -07:00
Benedict Elliott Smith 22df42c7be Store historical transactions per epoch
update durability scheduling and majority deps fetching
do not deserialize deps in CommandsForRangesLoader unless required

AccordJournalPurger should use shouldCleanupPartial

load historical transactions when loading topology
2025-04-17 11:59:52 -07:00
Benedict Elliott Smith 1734011032 Fix truncatedApply deserialization 2025-04-17 11:59:52 -07:00
Benedict Elliott Smith b8f3b74587 Journal diff serialization: validateFlags and WaitingOn size 2025-04-17 11:59:52 -07:00
Benedict Elliott Smith 25b1bd74b8 Enable and test purging 2025-04-17 11:59:52 -07:00
Benedict Elliott Smith aa6ebc140f Halve cache memory consumption by not retaining 'original' to diff; dedup RoutingKey tableId; avoid calculating rejectsFastPath in more cases; delay retry of fetchMajorityDeps; fix SetShardDurable marking shards durable 2025-04-17 11:59:52 -07:00
Benedict Elliott Smith 82ae3adcb2 ExclusiveSyncPoints should always wait for a simple quorum
split JournalKey in journal table so we can index it
reorder journal fields so we can easily index on route (when present)
use Message.expiresAtNanos for callback expiration
do not notify slow for range barriers

Accord: Do not contact faulty replicas, and promptly report slow replies for preaccept/read. Do not wait for stale or left nodes for durability.
2025-04-17 11:59:52 -07:00
Benedict Elliott Smith d8ccc36e4d visit journal backwards to save time parsing
don't load range commands that are redundant, and load least possible
use MISC verb handler for maintenance tasks
2025-04-17 11:59:52 -07:00
Benedict Elliott Smith 0e1819bbb3 Accord: Share DurableBefore between CommandStores 2025-04-17 11:59:52 -07:00
Alex Petrov 8f7b1c2351 Follow-up to CASSANDRA-19967 and CASSANDRA-19869
patch by Alex Petrov, Ariel Weisberg, Benedict Elliott Smith, Blake Eggleston and David Capwell
2025-04-17 11:59:52 -07:00
Benedict Elliott Smith 282bacb84b ninja: fix CFK serializer 2025-04-17 11:59:52 -07:00
David Capwell d034661ae2 Support Restart node in Accord
patch by David Capwell; reviewed by Alex Petrov for CASSANDRA-19969
2025-04-17 11:59:52 -07:00
Alex Petrov d0a3586bd9 Add purging to Accord Journal table
Patch by Alex Petrov; reviewed by Aleksey Yeshchenko and Benedict Elliott Smith for CASSANDRA-19877
2025-04-17 11:59:52 -07:00
Benedict Elliott Smith 562894dcb1 improve AccordLoadTest to support more keys 2025-04-17 11:59:52 -07:00
Alex Petrov 1c269348a4 Implement Journal replay on startup:
* reconstruct CFK, TFK, progressLog
  * migrate CommandStore collection state from Accord table to the log
  * make memtable writes non-durable; reconstruct memtable state from Writes

Patch by Alex Petrov and Benedict Elliott Smith; reviewed by Benedict Elliott Smith and Alex Petrov for CASSANDRA-19869
2025-04-17 11:59:52 -07:00
David Capwell 1ed52038ce This commits contains the following two patches in order to reduce the amount of conflicts resolution necessary for future rebasing:
(Accord): C* stores table in Range which will cause ranges to be removed from Accord when DROP TABLE is performed
patch by David Capwell, Sam Tunnicliffe; reviewed by Sam Tunnicliffe for CASSANDRA-18675

CEP-15: (Accord) sequence EpochReady.coordinating to allow syncComplete to be learned from newer epochs
patch by David Capwell; reviewed by Alex Petrov, Blake Eggleston for CASSANDRA-19769
2025-04-17 11:59:52 -07:00
Blake Eggleston fbfb633bda CEP-15 (C*) - misc accord perf improvements
Patch by Blake Eggleston; Reviewed by David Capwell for CASSANDRA-19940

Changes:
Increase accord repair range splitting
Streamline table metadata fetching - removes some unnecessary abstraction from the table metadata lookup path
Remote unnecessary set building when building lists of overlapping keys
Add separate recover delay for repair and increase default recover delay
2025-04-17 11:59:51 -07:00
Blake Eggleston dd1230f2a2 CEP-15 (C*): Read accord repair cfk keys from sstable index.
Patch by Blake Eggleston; Reviewed by David Capwell for CASSANDRA-19920
2025-04-17 11:59:51 -07:00
David Capwell 8f7532bc10 Rebase fixup: Accord should follow the pattern and use requestTime.computeDeadline like the rest of the code, and Accord timeout MUST be less than user timeout
Rebase fixup: when a local keyspace is being open but it isnt present return null so error msg can be provided
Rebase fixup: improved metrics error msg when the exception doesnt match what is expected
Rebase improvement: when we see a timeout or preempt use the new vtable to show the status cross the cluster
Rebase improvement: Cluster.checkForThreadLeaks now groups similar stack traces to make the output less dense
2025-04-17 11:59:51 -07:00
David Capwell f87c22cbbd Rebase fixup: SerializationsTest needed to recreate the service.SyncComplete.bin file 2025-04-17 11:59:51 -07:00
Alex Petrov ef5f793dab Journal segment compaction
Patch by Alex Petrov and Aleksey Yeschenko, reviewed by Aleksey Yeschenko and Alex Petrov for CASSANDRA-19876
2025-04-17 11:59:51 -07:00
Benedict Elliott Smith 9ebce5d6df Redesign progress mechanisms to be memory efficient, use fewer messages and to resolve dependency chains promptly.
The SimpleProgressLog had a number of problems:

1. It polled for progress with no attempt to determine whether progress could realistically be made, so:
 - as the number of pending transactions grew, the proportion of useful work dropped (as many would be unable to make progress without earlier transactions completing)
 - each transaction in the chain could recover only on average 1/2 poll interval behind the last transaction to complete
2. It requested full transaction state from every replica on each attempt
3. It maintained a lot of in-memory state
4. Polling happened en-masse, allowing for little per-transaction control

We also separately maintained fairly expensive per-command listener state that negatively affected our command loading and caching.

The new DefaultProgressLog makes use of several new features: LocalListeners, RemoteListeners, Timers and Await messages.
 - LocalListeners provide a memory-efficient collection for managing each CommandStore<E2><80><99>s transaction listeners, with dedicated record keeping for inter-transaction relationships.
 - RemoteListeners provide a mechanism for request/response pairs that may be separated by longer than the normal Cassandra message timeout, and require minimal state on sender and recipient. This permits replicas to cheaply update their local state machine as soon as distributed information becomes available.

The DefaultProgressLog tracks each transaction with separate timers to handle per-transaction scheduling, backoff etc, and a succinct state machine. To reduce overhead correspondence is preferentially limited to a handful of replicas, and limited to the home shard where appropriate.

patch by Benedict; reviewed by Ariel Weisberg for CASSANDRA-19870
2025-04-17 11:59:51 -07:00
David Capwell f1db115e73 Create a fuzz test that randomizes topology changes, cluster actions, and CQL operations
patch by David Capwell; reviewed by Alex Petrov for CASSANDRA-19847
2025-04-17 11:59:51 -07:00
David Capwell a21c0a75aa Add a table to inspect the current state of a txn
patch by David Capwell; reviewed by Benedict Elliott Smith for CASSANDRA-19838
2025-04-17 11:59:51 -07:00
Caleb Rackliffe 2b01b5fa79 Command to Exclude Replicas from Durability Status Coordination
patch by Caleb Rackliffe; reviewed by David Capwell, Sam Tunnicliffe, and Benedict Elliott Smith for CASSANDRA-19321
2025-04-17 11:59:51 -07:00
Alex Petrov 325e48ac39 Switch to streaming serialization of SavedCommand
Patch by Alex Petrov; reviewed by David Capwell for CASSANDRA-19865

Co-authored-by: dcapwell <dcapwell@gmail.com>
2025-04-17 11:59:51 -07:00
David Capwell 26d517f4ff txns that update a static row when the desired row doesn't exist leads to an error
patch by David Capwell; reviewed by Caleb Rackliffe for CASSANDRA-19855
2025-04-17 11:59:51 -07:00
Alex Petrov b757cb8730 Add size to the segment index for safer journal reads
Patch by Alex Petrov; reviewed by Marcus Eriksson for CASSANDRA-19871
2025-04-17 11:59:51 -07:00
Alex Petrov c16297f862 Add an ability to reconstruct arbitrary epoch state from the log to TCM
Patch by Alex Petrov; reviewed by Marcus Eriksson for CASSANDRA-19790.
2025-04-17 11:59:51 -07:00
David Capwell a7c2bcafcd CommandsForRanges does not support slice which cause over returned data being sent
patch by David Capwell; reviewed by Alex Petrov for CASSANDRA-19857
2025-04-17 11:59:51 -07:00
Ariel Weisberg 8ab5003118 Accord migration and interop correctness
Patch by Ariel Weisberg; Reviewed by Blake Eggleston for CASSANDRA-19744
2025-04-17 11:59:51 -07:00
Benedict Elliott Smith c377a066af CASSANDRA-19825: Fix various bugs and abstraction deficiencies, including:
- Remove concept of non-participating home keys; home keys are required to be a participant in the transaction
- Remove covering/covers concept
- Various invalidation/truncation/erase behaviours

patch by Benedict; reviewed by Blake for CASSANDRA-19825
2025-04-17 11:59:51 -07:00
Benedict Elliott Smith 6905157b0e CommandsForKey Improvements incl Pruning
CommandsForKey periodically self-prunes, so as to continue functioning well in-between garbage collections. Once we prune we are left with potentially incomplete information, and have to sometimes load per-command information from disk. But the payoff is ensuring CommandsForKey objects - which drive the majority of the state machine - are kept to a reasonable size.

patch by Benedict; reviewed by Blake Eggleston and David Capwell
2025-04-17 11:59:51 -07:00
Youki Shiraishi 6d01bc2535 CEP-15 (Accord): When starting a transaction in a table where Accord is not enabled, should fail fast rather than fail with lack of ranges
patch by Youki Shiraishi; reviewed by Caleb Rackliffe, David Capwell for CASSANDRA-19759
2025-04-17 11:59:51 -07:00
Alex Petrov 1439fe8d31 Bring back Journal simulator (w/o Accord at least for now); add semaphore interceptor.
Patch by Alex Petrov; reviewed by David Capwell for CASSANDRA-19695.
2025-04-17 11:59:51 -07:00
David Capwell 0d4c0961be CEP-15: (Accord) When nodes are removed from a cluster, need to update topology tracking to avoid being blocked
patch by David Capwell; reviewed by Blake Eggleston for CASSANDRA-19719
2025-04-17 11:59:51 -07:00
Benedict Elliott Smith ac58a62c51 Introduce Periodic mode to Accord Journal
patch by Benedict; reviewed by Aleksey Yeschenko, Alex Petrov and David Capwell for CASSANDRA-19720
2025-04-17 11:59:51 -07:00
David Capwell d83de7fbb4 CEP-15: (Accord) SyncPoint timeouts become a Exhausted rather than a Timeout and doesn’t get retried
patch by David Capwell; reviewed by Ariel Weisberg for CASSANDRA-19718
2025-04-17 11:59:50 -07:00
David Capwell 39efba9c6a CEP-15: (Accord) Bootstraps LocalOnly txn can not be recreated from SerializerSupport
patch by David Capwell; reviewed by Benedict Elliott Smith for CASSANDRA-19674
2025-04-17 11:59:50 -07:00
Ariel Weisberg feb8d2e2b2 Accord barrier/inclusive sync point fixes
Patch by Ariel Weisberg, Benedict Elliott Smith; reviewed by Benedict Elliott Smith for CASSANDRA-19641
2025-04-17 11:59:50 -07:00
Caleb Rackliffe 1c87bb96ed Baseline Diagnostic vtables for Accord
patch by Caleb Rackliffe; reviewed by David Capwell and Ariel Weisberg for CASSANDRA-18732
2025-04-17 11:59:50 -07:00
David Capwell 11ed5309b5 IndexOutOfBoundsException while serializing CommandsForKey
patch by David Capwell; reviewed by Blake Eggleston for CASSANDRA-19642
2025-04-17 11:59:50 -07:00
Aleksey Yeschenko 3127c558d3 Move preaccept expiration logic away from Agent
patch by Aleksey Yeschenko; reviewed by Alex Petrov and Benedict Elliott Smith for CASSANDRA-18888
2025-04-17 11:59:50 -07:00
Caleb Rackliffe 1b3c3f32a4 post-rebase fixes, mostly around CASSANDRA-19341 and CASSANDRA-19567 2025-04-17 11:59:50 -07:00
Caleb Rackliffe 418f8bf3f0 Prohibit counter column access in Accord transactions
patch by Caleb Rackliffe; reviewed by David Capwell for CASSANDRA-18987
2025-04-17 11:59:50 -07:00
Blake Eggleston 543210ae12 CEP-15 (C*) Integrate accord with repair
Patch by Blake Eggleston; Reviewed by Ariel Weisberg and David Capwell for CASSANDRA-19472
2025-04-17 11:59:50 -07:00