* [Store] feat: CXL storage full features.
* fix(ci): resolve cxl test failure and code format
* fix(ci): add CXL protocol support and fix code format issues
* [Store] feat: CXL storage full features, reset and rm extern/pybind
* [TE] Support TCP fallback in EFA build and improve EFA documentation (#1523)
When building with USE_EFA=ON, auto_discover is disabled to prevent
RDMA transport installation (QP creation fails on EFA devices). This
means TCP transport is also not installed automatically. Add explicit
TCP transport installation for non-EFA protocols in the EFA build path.
Documentation changes:
- build.md: Add USE_EFA option and clarify USE_CUDA default/purpose
- supported-protocols.md: Add EFA as a supported protocol
- efa_transport.md: Add USE_CUDA=ON to build command, document GPU
memory requirement
Co-authored-by: whn09 <whn09@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Teng Ma <teng-ma@linux.alibaba.com>
* [Store] Optimize BucketStorageBackend for reduced lock contention and add delete safety (#1456)
* fix(ci): resolve cxl test failure and code format
* [Store] Add Local Cache Mechanism for Mooncake Store Client (#1226)
* feat(Store): add local hot cache for client
* feat(Store): add client local hot cache log to show performance
* fix: local hot cache initialize bug
* fix(Store): Mooncake put slice is max 16MB, so make local hot cache block 16MB
* feat(Store): move local hot cache initialization to Client::Create
* feat(Store): local hot cache remove unused small block implementation
* feat(Store): add client local hot cache unit test
* fix(Store): modify client local hot cache suit with v0.3.7
* feat(Store): change local hot cache unit tes
* fix: initialize local hot cache with negative value
* feat: use in process master and metadata fro local hot cache unit test.
* feat: update local hot cache to one replica one slice version
* fix: local hot cache unit test use in process master service
* fix: code style fix
* fix: fix dirty read when client wants to read a previously hitted hot block but the hot block is modified by incoming put actions
* fix: local hot cache unit test use in process master service
* fix: code format fix
* fix: fix comment problems for
* feat: add local hot asynchronous queue size limit
* fix: local hot cache task involves the block so that there is no memcpy operation when inserting local hot cache
* fix: code check fix
* fix: update block in_use prop to reference count
---------
Co-authored-by: shichangzhang064 <zhangshichang@h-partners.com>
* fix(ci): add CXL protocol support and fix code format issues
* fix: address comments from code review
* fix(ci): resolve cxl test failure
* fix(ci): resolve ci error
---------
Co-authored-by: 王鹤男 <wanghenan09@gmail.com>
Co-authored-by: whn09 <whn09@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Teng Ma <teng-ma@linux.alibaba.com>
Co-authored-by: Mahesh Bapatu <153306023+maheshrbapatu@users.noreply.github.com>
Co-authored-by: Shichang Zhang <77728761+Shichang-Zhang@users.noreply.github.com>
Co-authored-by: shichangzhang064 <zhangshichang@h-partners.com>
* feat(ci): add sglang epd test case
Signed-off-by: shicanwei <shicanwei.scw@alibaba-inc.com>
* refactor: update dual-machine test cases to adapt to new common functions in common.sh
Signed-off-by: shicanwei <shicanwei.scw@alibaba-inc.com>
---------
Signed-off-by: shicanwei <shicanwei.scw@alibaba-inc.com>
* [Store]: add task executor feature with unit and executor test
Signed-off-by: Vincent Gao <vincentbo@linux.alibaba.com>
* [Store]: add some optimizations to task executor
Refactor the task_executor into the client_service and add a
task structure on the client side. Additionally, remove the
existence check in the execute function;
only retrying should be performed if replica allocation fails.
Signed-off-by: Vincent Gao <vincentbo@linux.alibaba.com>
* [Store]: get source replica from copyStart or moveStart api
* [Store]: call move or copy end if the target replica already exist to complete the replication task
* [Store]: use the max_retry_attempts in master side
* [Store]: add client integration test and set default max_retry_attempts to 10
* [Doc] update task api introduction
Signed-off-by: Vincent Gao <vincentbo@linux.alibaba.com>
* [Store] Set copy and move as private methods
Signed-off-by: Vincent Gao <vincentbo@linux.alibaba.com>
* [Doc]: change the default task max_retry_attempts to 10
* [Doc]: fix some description error
* [Store]: allocate the buffer size to be a multiple of 16MB
* [Store]: validate the replica is in local and directly construct slices from replica buffer address instead of copy the data to local buffer
* [Store]: change the validate logic to directly use transfer engine endpoint or local_hostname_.
* [Store] add e2e ci test for copy and move api
* [Store] refactor the client move and copy function
* [Store] fix the e2e test
* [Store] remove unused code
* [Store] add source field when build replica copy payload in the task_manager_test
* [Store] rename back to snake case for split_into_slices function and also remove hard code for client poll count
* [Store] revert mis deleted field when resolve conflicts
* [Store] change test to validate the real behaviour
* [Store] change the default task fetch size to 16
* [Doc]: change the replica copy/move sequence diagram
* [Store] add new split_into_slice method
* [Store] change the real client to use split_to_slice with buffer handle parameters
---------
Signed-off-by: Vincent Gao <vincentbo@linux.alibaba.com>
Co-authored-by: Vincent Gao <vincentbo@linux.alibaba.com>
* [Store] Retrieve the actual glibc version during the build process
Signed-off-by: Cruz Zhao <CruzZhao@linux.alibaba.com>
* [build] avoid use hardcode libc.so.6 and avoid use python in shell script
Signed-off-by: Cruz Zhao <CruzZhao@linux.alibaba.com>
---------
Signed-off-by: Cruz Zhao <CruzZhao@linux.alibaba.com>
* feat(tone_tests): add E2E test cases and bilingual documentation
Signed-off-by: lukotong-7 <shicanwei.scw@alibaba-inc.com>
* update T-One job link to show real-time logs
Signed-off-by: lukotong-7 <shicanwei.scw@alibaba-inc.com>
* Add description for test_hicache_storage_mooncake_backend.py
Signed-off-by: lukotong-7 <shicanwei.scw@alibaba-inc.com>
* Add parse and cleanup command options to test_1p1d_erdma.sh
Signed-off-by: lukotong-7 <shicanwei.scw@alibaba-inc.com>
* Ensure that cleanup and parse are always executed, and replace exit with return.
Signed-off-by: lukotong-7 <shicanwei.scw@alibaba-inc.com>
---------
Signed-off-by: lukotong-7 <shicanwei.scw@alibaba-inc.com>
* [Store]: Add independent deployment implementation for Client
Signed-off-by: Xingrui Yi <yixingrui@linux.alibaba.com>
* [Doc]: mooncake-store: add introduce for client standalone mode
Signed-off-by: Xingrui Yi <yixingrui@linux.alibaba.com>
* [CI]: Add dummy client test
Signed-off-by: Xingrui Yi <yixingrui@linux.alibaba.com>
* [Store]: Change dummy client setup into a new func
Signed-off-by: Xingrui Yi <yixingrui@linux.alibaba.com>
* [Store]: Add more log for dummy client
Signed-off-by: Xingrui Yi <yixingrui@linux.alibaba.com>
---------
Signed-off-by: Xingrui Yi <yixingrui@linux.alibaba.com>
* ci: add non-CUDA release workflow and update documentation
- Add release-non-cuda.yaml workflow for building non-CUDA version
- Modify build_wheel.sh to support dynamic package name modification
- Update README.md and docs with installation instructions for both versions
- CUDA version includes Mooncake-EP and GPU topology detection
- Non-CUDA version for environments without CUDA dependencies
* Update README.md
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
---------
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
* Always build with EP in CI
* Add missing LIBRARY_PATH
* Do not build tests and examples for the final steps to save up disk space
* Remove the requirement of USE_CUDA=ON for transfer-engine (without EP)
* Initialize a mooncake backend
* Add pybind
* Fix incorrect backend registration
* Fix wheel building of mooncake_ep
* Add a fake allreduce implementation
* Introduce transfer_engine to mooncake_backend
* Add a basic CPU proxy execution framework
* Implement a seemingly working allgather
* Remove mooncake_ep's dependency on etcd
* Implement `_allgather_base`
* Implement `allreduce`
* Implement `alltoall`
* Use an even-odd pattern for data transfer
* Add a `set_host_ip` method
* Switch to an extended-API implementation of the Mooncake backend
* Implement `broadcast`
* Implement `barrier`
* Extend Mooncake backend to CPU
* Support more operations for reduction
* Fix the backend-worker coordination logic
* Optimize CPU worker with a callback pattern
* Add a timeout-based broken-ranks detection
* Merge EP module into Mooncake's build system
* Share transfer buffer across all worker instances
* Switch to a more robust approach to detect broken ranks
* Specify CUDA device for test_mooncake_backend.py
* Explicitly stop mooncake worker
* Use transfer engine's notifications to implement collective signals
* Remove the unused `all_reduce_without` API
* Switch to mooncake backend for test_mooncake_ep.py
* Support both IB and RoCE
* Fix EP unit test
* Pass the auto-detected nic_id to EP Buffer
* Fix CMake conditional branches when `PYTORCH_CMAKE_PATH` is not set
* Fix ibgda syncing for RoCE
* Revert "Share transfer buffer across all worker instances"
This reverts commit 964e0a96
* Implement `_reduce_scatter_base`
* Make CPU backends aware of broken ranks
* Fix .typos.toml
* Add a perf test for mooncake backend
* Support more dtypes for reduction
* Revert "Use transfer engine's notifications to implement collective signals"
This reverts commit f20ffb21
* Share worker thread among all process groups
* Share transfer engine among all process groups
* Fix unit tests
* Add a warmup phase for transfer engine
* Fix transfer engine buffer locations
* Fix incorrect calculation of mooncake ep buffer
* Do not use timeout detection in mooncake_ep tests
* Update mooncake backend perf test
* Demangle per-group buffer offset from the shared taskId
* Stop allocating the useless `cuda_counter_buffer` and `cuda_data_buffer`
* Split the task list into a CPU region and a CUDA region
* Add a warmup for test_mooncake_backend_perf.py
* Switch from raw cudaEvent to `torch::Event`
* Fix MooncakeWorkCuda::wait() to make it compatible with cuda graphs
* Add doc
* Fix perf test
* Implement all-gather for perf test
* Move impl of `MooncakeEpBuffer`'s member functions to .cpp
* Change `gathered_experts` to `broken_nodes` to make the API more consistent
* `broken_nodes` should be `broken_ranks`
* API rename
* Fix format
* Enable WITH_EP option in CI
* Try installing torch in advance in CI
* Set `TORCH_CUDA_ARCH_LIST` in CMakeLists.txt
* Install required dependencies in the CI CUDA environment
* [CI] Add the matching PyTorch
* [CI] Add a workaround for missing `CUDA::nvToolsExt`
* Remove unused pybind base class declaration of `MooncakeBackendOptions`
* Support `set_device_filter`
* Remove unused headers for ep_py.cpp
* Build the EP-wheel with setuptools on CI
* [CI] Add the build-with-ep process to release.yaml
* Minor format fix
* Update build guide
* Fix docs
* Only build EP wheel with torch==2.8.0
* Add a torch version assertion for Mooncake Backend
* Fix some python typing
* Use the correct group for EP's initial data sharing
* API: invert `broken_ranks` and change into `active_ranks`
* Followup fix for inverting the API
* Fix format
* Bug-fix in mooncake_ep_kernel.cu
* Mooncake EP has to be built with USE_CUDA on
* Fixed some issues according to the review
* Fix bug
* Prepare test action env.
* [Trial] fix wildcard.
* [Trial] fix libcuda.so.1 has been modified after cuda-12.8 installed in alternative install.
* Revert to previous release.yml.
* Try to add repo for prevent publish to pypi failed in fork repo.
* [Store] config: start http server from master
Signed-off-by: Teng Ma <sima.mt@alibaba-inc.com>
* [Refactor] Use coro http server for metadata
* fix cmake
* fix merge
* add doc and test
* clear
---------
Signed-off-by: Teng Ma <sima.mt@alibaba-inc.com>