- cfk pruning+prebootstrap=invalid future dependency
- exclude retired ranges when filtering RX stillTouches
- propagate uses incorrect lowEpoch when fetch finds additional owned/touched ranges
- node.withEpoch should callback with TopologyRetiredException, not throw
- Recovery can race with durable-applied pruning; must not send durable unless latest ballot on apply
- removeRedundantDependencies was not slicing pre-bootstrap range calculation to participating ranges
- NPE in TopologyManager.atLeast caused by referencing an epoch that has been GC'd
- use journal durableBeforePersister in burn test, not NOOP_PERSISTER
- ServerUtils.cleanupDirectory use tryDeleteRecursive
- FsyncRunnable shutdown
- fix NPE in AccordJournalBurnTest
patch by Benedict; reviewed by Alex Petrov for CASSANDRA-20688
- Share TableMetadatas and PartitionKey for PartialTxn serialization
Also:
- TxnReference et al should reference uniqueId, and avoid serializing ksName/cfName
- Don't double count shared keys when estimating size on heap
patch by Benedict; reviewed by David Capwell for CASSANDRA-20578
Also improve:
- TxnId serialization
- StoreParticipants serialization
- compareUnsigned Node.Id for consistency with serialized TxnId
patch by Benedict; reviewed by Alex Petrov for CASSANDRA-20546
Improve:
- InMemoryJournal compaction should simulate files to ensure we compact the same cohorts of records (so shadowing applied correctly)
- Don't update CFK deps if already known
- avoid unnecessary heapification of LogGroupTimers
- begin removal of Guava (mostly unused)
- consider medium path on fast path delayed
- add Seekables.indexOf to support faster serialization
- unboxed Invariant.requires variant
- AsyncChains.addCallback -> invoke
Fix:
- Recovery must wait for earlier transactions on shards that have not yet been accepted, even if we recover a fast Stable record on another shard
- only skip Dep calculation on PreAccept if vote is required by coordinator
- ReadTracker must slice Minimal the first unavailable collection
- If PreLoadContext cannot acquire shared caches, still consider existing contents
- CFK: make sure to insert UNSTABLE into missing array
- Fix failed task stops DelayedCommandStore task queue processing further tasks
- short circuit AbstractRanges.equals()
- Handle edge case where we can take fast/medium path but one of the Deps replies contains a future TxnId
- update executeAtEpoch when retryInFutureEpoch
- Fix incorrect validation in validateMissing (should only validate newInfo is not in missing when in witnessedBy; all other additions should be included)
patch by Benedict; reviewed by David Capwell for CASSANDRA-20522
- Accord Journal purging was disabled
- remove unique_id from schema keyspace
- avoid String.format in Compactor hot path
- avoid string concatenation on hot path; improve segment compactor partition build efficiency
- Partial compaction should update records in place to ensure truncation of discontiguous compactions do not lead to an incorrect field version being used
- StoreParticipants.touches behaviour for RX was erroneously modified; should touch all non-redundant ranges including those no longer owned
- SetShardDurable should correctly set DurableBefore Majority/Universal based on the Durability parameter
- fix erroneous prunedBefore invariant
- Journal compaction should not rewrite fields shadowed by a newer record
- Don't save updates to ERASED commands
- Simplify CommandChange.getFlags
- fix handling of Durability for Invalidated
- Don't use ApplyAt for GC_BEFORE with partial input, as might be a saveStatus >= ApplyAtKnown but with executeAt < ApplyAtKnown
patch by Benedict; reviewed by Alex Petrov for CASSANDRA-20441
- Decouple command serialization from TableMetadata version; introduce ColumnMetadata ids; gracefully handle missing TableId
- DataInputPlus.readLeastSignificantBytes must truncate high bits
- Fix RandomPartitioner accord serialization
- Fast path stable commits must not override recovery propose/commit decisions regarding visibility of a transaction
- RejectBefore must mergeMax, not merge, to ensure we maintain epoch and hlc increasing independently
- Bad commitInvalidate decision
- consistent filtering for touches and stillTouches
- ensure TRUNCATE_BEFORE implies SHARD_APPLIED
- TopologyManager.unsyncedOnly off-by-one error
- DurabilityQueue should not retry SyncPointErased
- handle rare case of no deps but none needed
- not updating CFK synchronously on recovery, which can lead to erroneous recovery decisions for other transactions
- Don't return partial read response when one commandStore rejects the commit
- Filter touches/stillTouches consistently
- WaitingState computeLowEpoch must use hasTouched to handle historic key with no route
Improve:
- Use format parameters to defer building Invariants.requireArgument string
- streamline RedundantStatus/RedundantBefore
Also improve:
- Introduce DurabilityService
- Retire SyncPoint, replace Barrier with Write and RX
- MessageType -> enum, restore GetMaxConflict
- Standardise backoff logic with WaitStrategy
- improve TimeoutStrategy/RetryStrategy specification strings
- Forbid KX, remove directKeyDeps
- Introduce UniqueTimeService, permitting hlc reservations for sync points avoid delay when min TxnId is sufficiently in the past
- Remove ListStore custom purge logic
Also fix:
- RejectBefore should reject on both epoch and hlc
- Do not record sync success for removed nodes
- Support GlobalDurability detecting no command store to run on
- Incorrect ballot constructor
- Serializing 15-bit ballot flags incorrectly
- TopologyManager.hasEpoch deadlock
- Computing withOpenEpochs incorrectly, sometimes stopping one epoch short
- PartitionKey serializer should not depend on schema information that can be erased
patch by Benedict; reviewed by Alex Petrov for CASSANDRA-20395
- Bad ArrayBuffers recycling logic
- RX must ensure dependencies TRANSITIVE_VISIBLE
- Permit constructing "antiRange" that spans multiple prefixes
- Not computing range CommandSummary IsDep correctly
- Truncated commands that aren't shard durable could not repopulate CFK on replay, permitting recovery of another command to make an incorrect decision
- NPE on async persist of RX (i.e. supplying no callback)
- NPE in Builder.shouldCleanup when durability is null
patch by Benedict; reviewed by Alex Petrov for CASSANDRA-20370
- Only use persisted RedundantBefore for compaction
- RouteIndex should index only touches, not Route
- Flush RangesForEpoch updates to journal immediately, so we do not rely on the command we are processing succeeding
- DurableBefore updates must wait for the epochs to be known locally
- Shard.mustWitnessEpoch to support guaranteeing to witness relevant non-topology schema changes
- We must propagate RedundantBefore RX shard bounds along with epoch syncs
- Prevent a truncated transaction FetchData infinite loop
- GC_BEFORE status being overwritten by bootstrappedAt, permitting old transaction state to be resurrected
- Avoid CFK.maxUniqueHlc read race on bootstrap
- TopologyManager.awaitEpoch could wait for wrong epoch
- Journal fsync thread could miss notifications
Also improve:
- CommandStores uses SearchableRangeList for finding matching stores
- Refactor RedundantBefore to use a sorted array of TxnId/RedundantStatus pairs (to better fix GC_BEFORE issue)
- Accord debug keyspace operates on keyspace/table, and sorts correctly by token
patch by Benedict; reviewed by Alex Petrov for CASSANDRA-20361
Improve:
- Introduce pre/accept fast execution flags
- Introduce searchable Deps serialization
- Flatten AccordRoutingKey(s) into single type, using sentinel bits
- Introduce new fast byte-comparable serialization methods for Token and TableId to support above
Fix:
- Fix journal re-serialization logic
- Enable RandomPartitioner for Accord by supporting fixed-width serialization for RouteIndex
patch by Benedict; reviewed by Alex Petrov for CASSANDRA-20349
Also fix:
- Topology slicing must declare whether we share/slice node ownership (to assist above)
- CFK.visit removes transitive dependencies too eagerly across epoch change
- apply cleanup to builder consistently, and construct the same value we would produce by purge (so that replay is idempotent)
- Invoke ExecuteTxn.LocalExecute callbacks on originating CommandStore
- misc other minor issues
patch by Benedict; reviewed by Alex Petrov for CASSANDRA-20325
Use initializeTopologyUnsafe to for-load the last topology and create shards rather than replaying all topologies.
Patch by Alex Petrov; reviewed by Benedict Elliott Smith for CASSANDRA-20294
- Even if we can decide autonomously that we took the fast path, we must still wait for earlier transactions to decide themselves
- We must update a command that is in CFK.loadingPruned whether or not it is outOfRange
- We must visit pruned commands that are transitive dependencies of RX to ensure dependencies are propagated
- flagsWithoutDomainAndKind() -> flagsWithoutDomainOrKindOrCardinality() [to fix integration regression]
- Don't invoke uniqueNow() twice when allocate nextTxnId
- BTreeReducingRangeMap edge case that corrupts map when merge function does not return one of its inputs
patch by Benedict; reviewed by Alex Petrov for CASSANDRA-20292
Also fix:
- Truncate command on first access, without participants
- Use Ballot.ZERO when invoking CFK.insertOutOfRange where appropriate
- Don't supply a command's own route to ProgressLog.waiting to ensure new keys are incorporated
- Ensure progress in CommandsForKey by setting vestigial commands to ERASED
- Add any missing owned keys to StoreParticipants.route to ensure fetch can make progress
- Recovery must wait for earlier not-accepted transactions if either has the privileged coordinator optimisation
- Inclusive SyncPoint used incorrect topologies for propose phase
- Barrier must not register local listener without up-to-date topology information
- Stop home shard truncating a TxnId to vestigial rather than Invalidated so other shards can make progress
Also improve:
- Validate commands are constructed with non-empty participants
- Remove some unnecessary synchronized keywords
- Clear ok messages on PreAccept and Accept to free up memory
- Introduce TxnId.Cardinality flag so we can optimise single key queries
- Update CommandsForKey serialization to better handle larger flag space
- Configurable which Txn.Kind can result in a CommandStore being marked stale
- Process DefaultProgressLog queue synchronously when relevant state is resident in memory
- Remove defunct CollectMaxApplied version of ListStore bootstrap
- Standardise linearizability violation reporting
- Improve CommandStore.execute method naming to reduce chance of misuse
- Prune and address some comments
patch by Benedict; reviewed by Alex Petrov for CASSANDRA-20282
- preclude TCM of being behind Accord if newer epoch is reported via withEpoch/fetchTopologyInternal
- improve topology discovery during first boot and replay
- fix races between config service TCM listener reporting topologies, and fetched topologies during
Patch by Alex Petrov, reviewed by Ariel Weisberg and David Capwell for CASSANDRA-20245
- Implement missing parts of protocol optimisations, refine some particulars and remove MEDIUM_PATH_WAIT_ON_RECOVERY
Also fix:
- Deps.txnIds -> Deps.txnIdsWithFlags to make clear unsafety, and validate Command isn't created with flags
- Save a lossy low/high epoch we're waiting on in WaitingState so that we can reconstruct the same Route on callback
- Load any potentially invalidated commands we had in ProgressLog to ensure they are marked invalidated locally and notify any waiters
- Refine Invariant to accept case where a truncated dependency should apply first but both transactions are redundant and only the waiter could not be cleaned up
- only update CFK.maxUniqueHlc using commands we execute
- Don't update uniqueHlc when zero to avoid incompatible transaction invariant check
patch by Benedict; reviewed by Alex Petrov for CASSANDRA-20222
- Fix waiting state callback computes different route to initiator
- Invariants.checkX -> Invariants.requireX (to allow complementary Invariants.expectX as appropriate)
patch by Benedict; reviewed by Alex Petrov for CASSANDRA-20228
Follow-up to CASSANDRA-20228:
- Fix AccordUpdateParameters to correctly supply List cell paths derived from applyAt
- Fix ExecuteAtSerialize.serialiseNullable
- Fix topologies flush regression caused by markRetired forcing a flush on every call
patch by Benedict; reviewed by Ariel Weisberg for CASSANDRA-20228
- detect and save whether RX can be used as a GC bound for HLCs
- have RX durably record if they witnessed any superseding epoch and use this for HLC GC
- rework GC to retain applyAt for writes
Also improve:
- PreCommitted overwrites fast path vote, then loses tie-break with NotAccept, so we don't recover
- track executeAtLeast and uniqueHlc separately from WaitingOn in CommandChange
- Always propose executeAt.hlc() >= txnId.hlc()
- Short-circuit recovery if ApplyAt already known
Also fix:
- Dangling recovery when Commit rejected
- Failed epoch fetch handling in DefaultProgressLog
- CommandsForKey must track Accept executeAt for computing earlierWait collection on recovery
- Report DefaultProgressLog exceptions before retrying
- Fix inverted condition in RecoverWithRoute
- Remove laterNoWait from laterWait collection in recovery
- Invalidate incorrectly selected maxNotTruncated if max had a not truncated component
- WaitingState.fetch not request/propagating prior epochs
- CFK visit active pruning was comparing txn against maxCommittedWriteBefore rather than maxCommittedWriteForEpoch
patch by Benedict; reviewed by Alex Petrov for CASSANDRA-20228
- Privileged coordinator. If the coordinator is a replica we can reduce our quorum sizes by including the coordinator's vote.
- with deps: if we include coordinator's preaccept deps we can reliably reduce quorum size by 1, at the expense of recovery sometimes requiring additional phases and waiting for future txns
- with only vote: if we only include the vote we can avoid any additional recovery phases or waiting for future txns, but can reduce our quorum size for only some configurations
- Medium path. If t=t0 at a simple majority we can take just two rounds.
- with additional phase: on recovery, earlier txns must wait for medium path to be disabled (or committed) so we do not accidentally recover a transaction that won't be witnessed
- with unstable deps: on slow path commit, deps not found in accept as committed as unstable and do not affect recovery decisions for earlier transactions
Also improve:
- Recovery of await conditions simply recalculates the rejects flag rather than restarting to ensure faster forward progress
- refactor Command hierarchy
- tweak: don't save Writes for non-write transactions
Also fix:
- isOutOfRange when invoked from callback for some txn < Committed
- handle loadingPruned that is pre-bootstrap on update to bootstrappedAt
- journal replay can overwrite in memory state loaded for a command not yet replayed
- subtle ordering bug in LatestDeps.mergeProposal
- fix occasional class initialisation order bug
- handle race condition on learning of topology
patch by Benedict; reviewed by Aleksey Yeschenko for CASSANDRA-20222
Improve:
- Remove GetMaxConflict; use local MaxConflict collection
Fix:
- Invalidated should retain StoreParticipants so we can update CFK on journal replay
- loading pruned uninitialised commands via CFK: make sure hasTouched contains key so that if invalidated we are notified
- updateExecuteAtLeast should always be higher than TxnId
- Async CFK callbacks treated pruned transactions incorrectly
- node.withEpoch when ExecuteEphemeralRead in futureEpoch
- Deps.without should be key/range aware
- Durably mark bootstrapBeganAt and safeToReadAt in MaxConflicts
- Don't attempt to calculate local deps when DepsErased in GetLatestDeps (command has durably applied to this shard)
- filter StoreParticipants before invoking shouldCleanup
patch by Benedict; reviewed by Alex Petrov for CASSANDRA-20183
Includes multiple changes, primary ones:
* Make HarrySimulatorTest work with Accord
* Removed nodes now live in TCM, no need to discover historic epochs in order to find removed nodes
* CommandStore <-> RangesForEpochs mappings required for startup are now stored in journal, and CS can be set up _without_ topology replay
* Topology replay is fully done via journal (where we store topologies themselves), and topology metadata table (where we store redundant/closed information)
* Fixed various bugs related to propagation and staleness
* TCM was previously relied on for "fetching" epoch: we can not rely on it as there's no guarantee we will see a consecutive epoch when grabbing Metadata#current
* Redundant / closed during replay was set with incorrect ranges in 1 of the code paths
* TCM was contacted multiple times for historical epochs, which made startup much longer under some circumstances
Patch by Alex Petrov; reviewed by Benedict Elliott Smith for CASSANDRA-20142
- Fix notifying unmanaged after update redundant before/bootstrap
- Do not infer invalid if we have a single round of replies with minKnown not decided and maxKnown erased - in this case store the knowledge for next request.
- Fix SyncPoint topology selection
- Fix CheckStatusOkFull.with(InvalidIf)
- Fix NotifyWaitingOn
- ExecuteTxn should only contact latest topology for follow-up requests
- DurableBefore.min should not go backwards on new epoch topology, journal replay was not correctly handling PreApplied, partialTxn can be null if not owned
- Fix notify pre-bootstrap that arrives post-bootstrap
- Avoid GC race condition on Propagate where we can incorrectly infer a shard is stale
- Ensure redundantBefore on previously-owned range does not imply redundant before for overlapping queries on still-owned range
- Ensure we don't mark stale unless all of the quorum we contacted had erased, else we may have raced with the agreement and erase
- Fix Invalidate when no route found for FetchData does not report to all requested local epochs
- Fix WAS_OWNED_RETIRED without durableBefore at Universal can lead to assertions with RX that we permit to execute but that have not yet
- Fix initialiseWaitingOn can in some cases transitively notify the command we're updating via maybeCleanup of dependencies, but the command isn't yet updated so isn't ready
- Fix encountering a command that is pre-bootstrap, and for which we have locally 'applied' a supserseding RX, so that we do not know its outcome locally (so we do not cleanup the command), but also it must have been decided - and we should not respond with future dependencies.
- Epoch failures on CoordinatePreAccept should trigger the CoordinatePreAccept failure handler
- Use the shard bound rather than GC bound for fallback dependency
- LatestDeps should be sliced to actual route, so as not to use both PreAccepted AND Stable deps as though Stable
- Fix various callback issues with node.withEpoch and Recover/Propose.isDone
- RecoverWithRoute can encounter a partially truncated transaction where the Deps for one shard are not committed. Must fetch LatestDeps.
- Tighten LatestDeps semantics for Recover
- CommandsForKey: do not restore pruned as APPLIED
- Ensure prune points execute in the epoch in which they are declared
- must merge all fast path votes including those from earlier epochs that may have witnessed a later transaction
- Recoveries that know the transaction is committed a priori should skip the Accept phase
- Maintain GC behaviour for redundant commands that are pre-bootstrap
- don't apply ERASE to CommandsForKey to avoid breaking pruning
- Introduce clearBefore to ProgressLog to more consistently handle cleaning up redundant transactions (and avoid triggering burn test invariants)
- don't replay journal of a bootstrapping node in burn test
- Recover, Accept or Commit reply from epoch that has been retired should be treated as Success rather than Redundant
- Distinguish completely REDUNDANT+PRE_BOOTSTRAP from partially GC_BEFORE and REDUNDANT+PRE_BOOTSTRAP - latter can make stronger inferences based on the GC_BEFORE intersection (could perhaps be treated as simply GC_BEFORE)
- RX must register historical transactions with CFK
- CommandStore.bootstrapper must wait for coordinate sync via same mechanism as sync()
- Don't start topology change for shard where all replicas are already bootstrapping
- Reify executes et al in StoreParticipants
- LocalListeners txn listener reentry may erase the entry entirely
- use registerAt in AbstractRequest for expirations, use correct time for expiresAt in ListAgent
- use txnId.epoch() for pruning, as must be before both txnId and executeAt of prune point for coordinating dependencies
- compute accurate KnownMap when affected by bootstrap or staleness
- upgradeTruncated should calculate Definition and Deps separately
- Invalidate should not sort before Erased when calculating max reply or max knowledge reply
- avoid another infinite loop at end of burn test
- avoid another epoch loading edge case
- pass through low/high epochs to ensure we propagate information to all waiting command stores
- RX must adopt a non-pruned dependency that has a higher TxnId (if is itself behind prune point)
- rejects should also be calculated on COMMITTED started before
- remove Apply Factory wrapper for RX, redundant now we have CoordinationAdapters (and has faulty epoch logic)
- for RX ensure we return maximum writes for each epoch we intersect (same effectively as pruning logic)
- rework updateUnmanaged to improve clarity
- BeginRecovery constructor of LatestDeps should use touches() not owns() for compute localDeps
- BeginRecovery superseding calculation was incorrectly treating startedBefore Committed and Accepted the same, when the point at which a dep should be known differs
- Refactor Command visiting, porting C* integration to accord-core
- RelationMultiMap Builder should resize keys and keyLimits independently
- CommandsForKey Serialization moved to accord-core
- losing ownership of range should trigger re-registration of unmanaged waiting on commit of a no-longer owned txn
patch by Benedict; reviewed by Alex Petrov for CASSANDRA-20172
* Move serializers that previously were in static field into corresponding swappable accessors
* Add a (no-op in Cassandra impl) Result serializer
* Allow more than one Accord Journal table
* Move some previously-Cassandra implementation details (such as FieldUpdates) to Accord
* Fix a bug with sorting of Journal Key (incorrect non-identity flags)
* Accord side: Make sure not to re-sort collection of events scheduled for application
* Implement serializers for Accord/BurnTest List* items
Patch by Alex Petrov; reviewed by Benedict Elliott Smith for CASSANDRA-20112.
For correctness, the dependencies we adopt on joining a new topology must exclude the possibility of respondents accepting additional transactions with a lower TxnId, so proxying on the existing `ExclusiveSyncPoint` mechanisms is logical for the time-being. This patch removes the `FetchMajorityDeps` logic in favour of simply waiting for a suitable `ExclusiveSyncPoint` to be proposed.
patch by Benedict, reviewed by Alex Petrov for CASSANDRA-20056
Patch by Alex Petrov; reviewed by Ariel Weisberg for CASSANDRA-20032
Accord Deps tests have incorrect range semantics
patch by David Capwell; reviewed by Ariel Weisberg for CASSANDRA-20029
update durability scheduling and majority deps fetching
do not deserialize deps in CommandsForRangesLoader unless required
AccordJournalPurger should use shouldCleanupPartial
load historical transactions when loading topology