Add START_MOVE_KEYS_MAX_RETRIES knob (default 50, BUGGIFY=10) and throw
start_move_keys_too_many_retries when exceeded. Wire the new error into
DDRelocationQueue (same light-weight re-queue handler as
finish_move_keys_too_many_retries and move_to_removed_server) and
normalDDQueueErrors so DD does not restart. Also propagate actor_cancelled
before the retry limit check in startMoveKeys.
* DD 7.3 finish movekeys backoff (#12991)
* Add jitter to finishMoveKeys backoff
Jitter the exponential backoff delay to [0.75x, 1.25x] of the base
value. This prevents all 15 FlowLock slots from retrying in lockstep
when they all hit transaction_too_old at the same time.
* Cap finishMoveKeys backoff at 5s after jitter
Apply the 5.0s cap after jitter so the final delay never exceeds
the documented maximum.
Problem: some of these transactions have been observed to involve cascades of many shards trying to execute them simultaneously owing to somewhat unpredictable DD dynamics. This then results in DD being pegged on CPU with many failing transactions.
Solution: the counters here will make it clear if this is happening and if so, where, so that further mitigations/solutions can be devised. Some of these code paths have TraceEvents, but some don't and of the existing TraceEvents, many are sampled. This PR avoids all of that and counts every time through, making us less reliant on guesswork.
An alternative approach (not taken here for many reasons) would be to instrument transaction code to let callers pass a tag. Then have the transaction client and/or server side emit a few metrics (started, committed, aborted) parameterized by tag.
One may reasonably object that the boilerplate here is kind of fugly. I guess my view is that at this late date the time for more subtle approaches has come and gone. We need this code to tell us what it is doing and if this looks a little intrusive, so be it.
Testing:
20260422-234326-gglass-2f514436068d99e6 compressed=True data_size=35519571 duration=5298275 ended=100000 fail=1 fail_fast=10 max_runs=100000 pass=99999 priority=100 remaining=0 runtime=0:55:44 sanity=False started=100000 stopped=20260423-003910 submitted=20260422-234326 timeout=5400 username=gglass
At commit: fff5439e with clang, seed -f ./tests/slow/SharedDefaultBackupCorrectness.toml -s 2189316179 -b on
We found a corruption where the destination storage server can get the incorrect
serverKeys mutations. Note this only happens when shard_encode_location_metadata is enabled.
The reason is that one of the actors in the previous iteration encountered
transaction_too_old error, and the transaction restarted. However, because the
actors are not cancelled, these can still modify the next transaction that
retried.
* bulkload support general engine and fix bugs
* add comments
* improve test coverage and fix bug
* nits and address comments
* nit
* nits
* fix data inconsistency bug due to bulkload metadata
* fix ss bulkload task metadata bugs
* nit and fix CI issue
* fix bugs of restore ss bulkload metadata
* use ssBulkLoadMetadata for fetchKey and general kv engine
* cleanup bulkload file for fetchkey
* fix CI issue
* fix simulation stuck due to repeated re-recruitment of unfit dd
* randomly do available space check when finding the dest team for bulkload in simulation
* address conflict
* code clean up
* update BulkDumping.toml same to BulkLoading.toml
* consolidate ss fetchkey and fetchshard failed to read bulkload task metadata
* fix DD bulkload job busy loop bug which causes segfault and test terminate unexpectedly in joshua test
* nit
* fix ss busy loop for bulkload in fetchkey
* use sqlite for bulkload ctest
* fix bulkload ctest stuck issue due to merge and change storage engine to ssd
* fix comments for CC recruit DD
* address comments
* address comments
* add comments
* fix ci format issue
* address comments
* add comments
* Improve BulkLoad/Dump implementation
* make bulkload test data folder inside simfdb folder
* simplify code
* use manifest in bulkdump metadata
* use manifest in bulkload
* apply bulkload fileset to bulkload and fix bugs of bytesampling value generation
* remove BulkDumpFileFullPathSet
* address comments
* address comments
* address comments
* Add usable region check per shard for encode shard location metadata
* nits
* nit
* address comments
* fix SS assertion failed for a wrong data move type generated by an old binary which does not encode the data move type in the data move id
* fix ClientTransactionProfilingCorrectness 7.3 upgrade test considering physical shard move compatibility
* code clean
* split CycleTestRestart in upgrading test from release-7.3
* address comments
* nits
To retrieve storage metadata for every status json request is very expensive
for clusters with a large number of storage servers. So I change the logic so
that ClusterController actively monitors changes to storage metadata, and only
retrieves them when there is a change.
prepare transactions
add DD Restore Preparing state; actor accept blob migrator requests
Refactor DDEnabledState and PrepareBlobRestoreReply
Improve the initialization, remove unused DDSharedContext method
move prepareBlobRestore to moveKeys
Make context->lock and DataDistributor.lock share the reference; change checkMoveKeysLock to checkPersistenMoveKeysLock; Add more debug trace
fix requestId assignment bug
Add DDEnabledState::sameId method
Throw movekeys_conflict after hybrid restore preparation to force reload
server list and shard mapping; format code; remove unused
methods,definition and comments
fix rebase conflicts
Make DD only load initial Data Distribution after enabled
use dd_config_changed error and throw it in serveBlobMigratorRequests
move empty range check before blob restore to the transaction lock database
Rename BlobRestore*