Commit Graph

57 Commits

Author SHA1 Message Date
Feng Ren e8db9e54cf
[TE] Add TENT codebase to main (Phase 1: structural import) (#1213)
* [TE] Add TENT to main branch: Phase 1

* Fix CI issues

* Apply suggestions from code review

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

* Fix build bug

* Update code format

* Fix code format issue

* disable building Mooncake TENT by default

* Code reformat

* Revise benchmark code and add docs

* format bench code

* Retrigger

* Update Python APIs

* Fix bugs

* Register memory in parallel

* Fix wheel packing

* Format fix

* reformat

---------

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2025-12-22 19:22:13 +08:00
Cruz Zhao 1a5983e756
[store] zero copy for get_tensor() and batch_get_tensor() (#1192) 2025-12-19 08:38:17 +08:00
Xuchun Shang 31c45e47f5
refactor tensor api and add tests (#1217)
* refactor tensor api and and tests

Signed-off-by: Xuchun Shang <xuchun.shang@gmail.com>
Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
2025-12-17 16:45:49 +08:00
shicanwei.scw cb657895a9
[CI] Add sglang e2e tests (#1181)
* feat(tone_tests): add E2E test cases and bilingual documentation

Signed-off-by: lukotong-7 <shicanwei.scw@alibaba-inc.com>

* update T-One job link to show real-time logs

Signed-off-by: lukotong-7 <shicanwei.scw@alibaba-inc.com>

* Add description for test_hicache_storage_mooncake_backend.py

Signed-off-by: lukotong-7 <shicanwei.scw@alibaba-inc.com>

* Add parse and cleanup command options to test_1p1d_erdma.sh

Signed-off-by: lukotong-7 <shicanwei.scw@alibaba-inc.com>

* Ensure that cleanup and parse are always executed, and replace exit with return.

Signed-off-by: lukotong-7 <shicanwei.scw@alibaba-inc.com>

---------

Signed-off-by: lukotong-7 <shicanwei.scw@alibaba-inc.com>
2025-12-16 00:52:27 +08:00
Xuchun Shang fa8dc059c2
[Store] add tp awareness for get_tensor (#1127)
* add tp awareness for get_tensor
2025-12-11 11:12:44 +08:00
Cruz Zhao 50442be8cc
[Store] pub_tensor for multiple replica (#1148) 2025-12-08 22:38:36 +08:00
ascend-direct-dev 6cca832c2d
change cmake for ascend (#1114)
Co-authored-by: youxiao <youxiao@huawei.com>
2025-12-02 14:30:46 +08:00
EkiRui 556fe771d1
feat(Store): Support Connecting Multiple Dummy Clients to One Real Client (#1122) 2025-11-28 17:59:09 +08:00
Xun Sun 63f395e165
[EP] Support multiple torch versions (#1098)
* Debug

* Debug

* Debug

* Debug

* Debug

* Debug

* Debug

* Debug

* Debug

* Debug

* Debug

* Debug

* Debug

* Debug

* Debug

* Debug

* Debug

* Debug

* Debug

* Debug

* Debug

* Debug

* Debug

* Debug

* Debug

* Debug
2025-11-28 14:25:45 +08:00
ZechaoZhang-beta a9313dabe3
[TE] feat: add barex_transport by build with USE_BAREX (#1045)
* feat[accl-barex]: add barex_transport by build with USE_BAREX

* feat[accl-barex]: spell fix

* feat[accl-barex]: clang-format

* feat[accl-barex]: fix clang format

* feat[barex]: add log

* Update mooncake-common/common.cmake

* Update mooncake-common/common.cmake

* fix all issues

---------

Co-authored-by: Teng Ma <teng-ma@linux.alibaba.com>
Co-authored-by: Teng Ma <sima.mt@alibaba-inc.com>
2025-11-25 23:59:58 +08:00
EkiRui 7517db4585
[Store] feat: Add standalone deployment implementation for Client (#1084)
* [Store]: Add independent deployment implementation for Client

Signed-off-by: Xingrui Yi <yixingrui@linux.alibaba.com>

* [Doc]: mooncake-store: add introduce for client standalone mode

Signed-off-by: Xingrui Yi <yixingrui@linux.alibaba.com>

* [CI]: Add dummy client test

Signed-off-by: Xingrui Yi <yixingrui@linux.alibaba.com>

* [Store]: Change dummy client setup into a new func

Signed-off-by: Xingrui Yi <yixingrui@linux.alibaba.com>

* [Store]: Add more log for dummy client

Signed-off-by: Xingrui Yi <yixingrui@linux.alibaba.com>

---------

Signed-off-by: Xingrui Yi <yixingrui@linux.alibaba.com>
2025-11-25 23:57:30 +08:00
Xuchun Shang 3a6cd2903e
[Store] add batch tensor (#1044)
* add batch [put/get] tensor

Signed-off-by: Xuchun Shang <xuchun.shang@gmail.com>

* fix

Signed-off-by: Xuchun Shang <xuchun.shang@gmail.com>

* add ci

Signed-off-by: Xuchun Shang <xuchun.shang@gmail.com>

* fix

Signed-off-by: Xuchun Shang <xuchun.shang@gmail.com>

* fix

Signed-off-by: Xuchun Shang <xuchun.shang@gmail.com>

---------

Signed-off-by: Xuchun Shang <xuchun.shang@gmail.com>
2025-11-13 11:25:37 +08:00
JinYan Su 223d74933b
ci: add non-CUDA release workflow and update documentation (#969)
* ci: add non-CUDA release workflow and update documentation

- Add release-non-cuda.yaml workflow for building non-CUDA version
- Modify build_wheel.sh to support dynamic package name modification
- Update README.md and docs with installation instructions for both versions
- CUDA version includes Mooncake-EP and GPU topology detection
- Non-CUDA version for environments without CUDA dependencies

* Update README.md

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

---------

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
2025-10-27 13:41:54 +08:00
ykwd 26948206e2
[CI] Fix CI Error Due to RDMA Fail (#930) 2025-10-20 19:34:23 +08:00
Xun Sun f89c207cfa
[CI/Build] Always build with EP in CI (#922)
* Always build with EP in CI

* Add missing LIBRARY_PATH

* Do not build tests and examples for the final steps to save up disk space

* Remove the requirement of USE_CUDA=ON for transfer-engine (without EP)
2025-10-17 14:17:13 +08:00
Xun Sun c5829aad1b
[Misc] Mooncake EP & Mooncake Backend (#805)
* Initialize a mooncake backend

* Add pybind

* Fix incorrect backend registration

* Fix wheel building of mooncake_ep

* Add a fake allreduce implementation

* Introduce transfer_engine to mooncake_backend

* Add a basic CPU proxy execution framework

* Implement a seemingly working allgather

* Remove mooncake_ep's dependency on etcd

* Implement `_allgather_base`

* Implement `allreduce`

* Implement `alltoall`

* Use an even-odd pattern for data transfer

* Add a `set_host_ip` method

* Switch to an extended-API implementation of the Mooncake backend

* Implement `broadcast`

* Implement `barrier`

* Extend Mooncake backend to CPU

* Support more operations for reduction

* Fix the backend-worker coordination logic

* Optimize CPU worker with a callback pattern

* Add a timeout-based broken-ranks detection

* Merge EP module into Mooncake's build system

* Share transfer buffer across all worker instances

* Switch to a more robust approach to detect broken ranks

* Specify CUDA device for test_mooncake_backend.py

* Explicitly stop mooncake worker

* Use transfer engine's notifications to implement collective signals

* Remove the unused `all_reduce_without` API

* Switch to mooncake backend for test_mooncake_ep.py

* Support both IB and RoCE

* Fix EP unit test

* Pass the auto-detected nic_id to EP Buffer

* Fix CMake conditional branches when `PYTORCH_CMAKE_PATH` is not set

* Fix ibgda syncing for RoCE

* Revert "Share transfer buffer across all worker instances"

This reverts commit 964e0a96

* Implement `_reduce_scatter_base`

* Make CPU backends aware of broken ranks

* Fix .typos.toml

* Add a perf test for mooncake backend

* Support more dtypes for reduction

* Revert "Use transfer engine's notifications to implement collective signals"

This reverts commit f20ffb21

* Share worker thread among all process groups

* Share transfer engine among all process groups

* Fix unit tests

* Add a warmup phase for transfer engine

* Fix transfer engine buffer locations

* Fix incorrect calculation of mooncake ep buffer

* Do not use timeout detection in mooncake_ep tests

* Update mooncake backend perf test

* Demangle per-group buffer offset from the shared taskId

* Stop allocating the useless `cuda_counter_buffer` and `cuda_data_buffer`

* Split the task list into a CPU region and a CUDA region

* Add a warmup for test_mooncake_backend_perf.py

* Switch from raw cudaEvent to `torch::Event`

* Fix MooncakeWorkCuda::wait() to make it compatible with cuda graphs

* Add doc

* Fix perf test

* Implement all-gather for perf test

* Move impl of `MooncakeEpBuffer`'s member functions to .cpp

* Change `gathered_experts` to `broken_nodes` to make the API more consistent

* `broken_nodes` should be `broken_ranks`

* API rename

* Fix format

* Enable WITH_EP option in CI

* Try installing torch in advance in CI

* Set `TORCH_CUDA_ARCH_LIST` in CMakeLists.txt

* Install required dependencies in the CI CUDA environment

* [CI] Add the matching PyTorch

* [CI] Add a workaround for missing `CUDA::nvToolsExt`

* Remove unused pybind base class declaration of `MooncakeBackendOptions`

* Support `set_device_filter`

* Remove unused headers for ep_py.cpp

* Build the EP-wheel with setuptools on CI

* [CI] Add the build-with-ep process to release.yaml

* Minor format fix

* Update build guide

* Fix docs

* Only build EP wheel with torch==2.8.0

* Add a torch version assertion for Mooncake Backend

* Fix some python typing

* Use the correct group for EP's initial data sharing

* API: invert `broken_ranks` and change into `active_ranks`

* Followup fix for inverting the API

* Fix format

* Bug-fix in mooncake_ep_kernel.cu

* Mooncake EP has to be built with USE_CUDA on

* Fixed some issues according to the review

* Fix bug
2025-09-26 10:02:17 +08:00
ascend-direct-dev c0f83b1ab3
add ascend direct transport to mooncake store (#835)
Co-authored-by: youxiao <youxiao@huawei.com>
2025-09-17 09:25:19 +08:00
Mumupika a7e1f1c360
[CI] Fix Release build_wheel.sh to make python 3.8 auditwheel happy (#801)
* Prepare test action env.

* [Trial] fix wildcard.

* [Trial] fix libcuda.so.1 has been modified after cuda-12.8 installed in alternative install.

* Revert to previous release.yml.

* Try to add repo for prevent publish to pypi failed in fork repo.
2025-09-05 10:37:28 +08:00
zuochunwei 0cc51a60ac
[TransferEngine] heterogeneous_ascend support kv-cache transfer between npu and gpu (#759)
* heterogeneous_ascend
Co-authored-by: AscendTransport<ascend_transport@yeah.net>

* update desc for USE_ASCEND_HETEROGENEOUS option

* format

* fix bug

* format

* lock_guard

* a new fix

---------

Co-authored-by: zuochunwei <zuochunwei@meituan.com>
Co-authored-by: ascend_transport <ascend_transport@yeah.net>
Co-authored-by: AscendTransport <ascendtransport@yeah.net>
2025-09-03 09:41:48 +08:00
qicosmos 185c5d229b
[coro_rpc] use client pool and enable rdma (#789) 2025-09-02 00:40:21 +08:00
Vladislav Nosivskoy aa9d4471a1
[Store] Add replication guarantees (#744)
* add replication guarantees

Signed-off-by: Vladislav Nosivskoy <vladnosiv@gmail.com>

* refactor implementation

Signed-off-by: Vladislav Nosivskoy <vladnosiv@gmail.com>

* fix merge aftermath

Signed-off-by: Vladislav Nosivskoy <vladnosiv@gmail.com>

* less copies

Signed-off-by: Vladislav Nosivskoy <vladnosiv@gmail.com>

* try fix wheel tests

Signed-off-by: Vladislav Nosivskoy <vladnosiv@gmail.com>

* fix wheel tests

Signed-off-by: Vladislav Nosivskoy <vladnosiv@gmail.com>

* remove old comment

Signed-off-by: Vladislav Nosivskoy <vladnosiv@gmail.com>

* fix wheel tests and add replication fault tolerance wheel test

Signed-off-by: Vladislav Nosivskoy <vladnosiv@gmail.com>

* split wheel tests

Signed-off-by: Vladislav Nosivskoy <vladnosiv@gmail.com>

* simplify code

Signed-off-by: Vladislav Nosivskoy <vladnosiv@gmail.com>

* replica allocation is best effort operation

Signed-off-by: Vladislav Nosivskoy <vladnosiv@gmail.com>

* fix tests for new best-effort behaviour

Signed-off-by: Vladislav Nosivskoy <vladnosiv@gmail.com>

* fix

Signed-off-by: Vladislav Nosivskoy <vladnosiv@gmail.com>

* update docs

Signed-off-by: Vladislav Nosivskoy <vladnosiv@gmail.com>

* update another docs

Signed-off-by: Vladislav Nosivskoy <vladnosiv@gmail.com>

---------

Signed-off-by: Vladislav Nosivskoy <vladnosiv@gmail.com>
2025-08-25 20:24:41 +08:00
Teng Ma f76c92295b
[Store] add c++ http metadata server in mooncake master (#766)
* [Store] config: start http server from master

Signed-off-by: Teng Ma <sima.mt@alibaba-inc.com>

* [Refactor] Use coro http server for metadata

* fix cmake

* fix merge

* add doc and test

* clear

---------

Signed-off-by: Teng Ma <sima.mt@alibaba-inc.com>
2025-08-25 13:56:37 +08:00
ascend-direct-dev 4f94433525
add ascend direct transport (#740)
* add ascend direct transport

* fix style

---------

Co-authored-by: youxiao <youxiao@huawei.com>
2025-08-18 22:34:17 +08:00
Sgt.Pepper 81f492c033
[Store]feat: Migrate Persistence Metadata from Client to Master Service (#690)
* initial commit

* fix client::query return fault

* fix isexist return fault

* fix test bug

* fix clearinvalidhandles problem

* add file description for 3fs

* change ssd function start from client to master

* fix naming error

* edit doc description

* edit doc

* clang format

* fix as the review comment

* fix formmat

* add master service test for ssd

* fix format

* add log and cli

* fix putend test
2025-08-14 14:34:51 +08:00
Houjiang Chen 4e49040172
[TransferEngine] exclude packaging ascend precompiled libraries (#737) 2025-08-13 09:39:06 +08:00
AscendTransport fc1eb07da9
[TransferEngine] Update to support CANN 8.2.RC1 (#714) 2025-08-07 13:51:52 +08:00
doujiang24 4e03dbe013
code format & enable code format checking in ci (#677)
Signed-off-by: doujiang24 <doujiang24@gmail.com>
2025-08-02 15:50:25 +08:00
AscendTransport d10893f1a7
open ascend (#658) 2025-07-24 10:36:42 +08:00
ykwd 1f71517d72
[Store] Enlarge the Default KV TTL to 5 Seconds (#660)
* change default kv lease ttl to 5 sec
2025-07-24 10:34:35 +08:00
AscendTransport 7e1ef1adc6
[TransferEngine] Ascend Transport: add batch_transfer_sync, Debian support & bug fixes (#619)
* ascend batch sumbit

* init fix

* fix cmake

* review

* review
2025-07-18 14:28:24 +08:00
Teng Ma 7a6b3d3acd
[Store] feat: put/get tensor API for store (#579)
* [Store] feat: add put/get tensor API

* add new paramter for get

* add random test

* resolve conflicts

* add tensor api test to ci

* fix ci

* modify position

* use put_from

* opt

* install torch for ci

* reduce the scope of pybind11 torch

* pass ci

* add dereg

* fix get tensor

---------

Co-authored-by: qicosmos <qicosmos@linux.alibaba.com>
2025-07-11 11:47:16 +08:00
ykwd 60ccc133a9
[Store] Soft Pin for Important Object (#587)
This PR adds a soft pin mechanism for important and frequently used objects, such as system prompts and hot objects.
2025-07-08 19:37:41 +08:00
Sgt.Pepper 6515b50d8c
[Store] Add ungister_buffer python binding for Mooncake Store (#596)
* test

* fix removefile bug

* add unregister_buffer python api and pytest

* add test fix
2025-07-07 13:01:29 +08:00
JinYan Su bde2fcaa13
refactor: introduce expected pattern for error handling in master service (#562)
* refactor: introduce expected pattern for error handling in master service

- Replace ErrorCode return types with tl::expected<T, ErrorCode> pattern
- Improve error handling clarity by separating success values from error codes
- Update MasterService methods to return expected<void, ErrorCode> or expected<T, ErrorCode>
- Modify RPC service interfaces to support expected pattern
- Update all related tests to handle new expected return types
- Add necessary includes for ylt/util/expected.hpp

This change makes error handling more explicit and type-safe:
- Success cases can be accessed via .value()
- Error cases can be accessed via .error()
- Eliminates ambiguity between success and error states

Future work:
- Extend expected pattern to RPC response types
- Enhance error code system for more comprehensive error handling

* Refactor error handling to use expected<T,E> pattern instead of ErrorCode

- Updated MasterService methods to replace ylt::expected with tl::expected for better error handling.
- Modified BatchGetReplicaList, BatchPutStart, BatchPutEnd, and other methods to return tl::expected types.
- Enhanced ClientIntegrationTest to handle expected results from Put, Get, Remove, and other operations using tl::expected.
- Adjusted error handling in clientctl and master_metrics_test to utilize the new expected type.
- Improved overall error reporting in tests to provide clearer feedback on operation failures.

* fix test compile

* refactor: update client implementation and remove master.proto

- Enhanced client.h and client.cpp with new functionality
- Removed obsolete master.proto file
- Updated master_client.cpp and transfer_task.cpp
- Improved integration and stress tests

* fix: update Python integration to work with new batch API

- Replace BatchObjectInfo with vector<vector<Replica::Descriptor>>
- Update BatchPut to handle new return type vector<tl::expected<void, ErrorCode>>
- Fix BatchQuery API usage to work with new expected pattern
- Convert unordered_map to vector format for BatchPut parameter compatibility

* refine master log and metric

* fix(ci): should alloc first

* merge main

* fix tests
2025-07-03 14:54:06 +08:00
Sgt.Pepper 77f5a7b9a1
[Store] Enable Client SSD Offload And Storage Persistence (#437)
* enable client ssd offload and storage persistence

* add storage_root_path in all tests setup() initialization and add related description in doc

* clean up headers and improve code readability - Added consistent Doxygen-style comments to all header files - Removed redundant code and outdated comments - Optimized function execution logic in LocalFile

* Revert "add storage_root_path in all tests setup() initialization and add related description in doc"

This reverts commit 159442d44e.

revert old high-level api test and doc modification

* Restore the high-level API to its original state and modify it to introduce the storage path through environment variables.

* add local_file_test and thread_pool_test

* feat(client_ssd_offload): implement async writes and fix locking bugs

- Refactor write operations to use thread pool for async file I/O
- Fix potential double-unlock bug by adding atomic is_locked_ flag
- Add corrupted file cleanup on write failure:
  - Auto-delete files with failed writes in destructor
  - Prevent subsequent reads of corrupted data

* add support for remove , remove_all , isexist interface etc.

* feat(kvcache): implement cluster isolation with session IDs

    * Remove precompilation parameters to simplify build configuration
    * Add session ID mechanism for cluster isolation:
      - Master node now generates unique session IDs on initialization
      - All persistent operations are scoped under session-specific subdirectories

* edit two parameters client get, add persisitence path in client rather than store_py.cpp

* add support for batch api conflict , refactor replica.descriptor to support file and memory type

* add test branch

* add ci ssd

* change python test

* add log for fail

* change querykey return value type

* fix bug

* fix bug

* add sleep for removefile

* fix sleep

* edit ci.yml and fix delete before write problem

* add comment for storage_backend

* spell check

* fix name problem and decrease errorcode for file

* add pytest for ssd offload

* edit test

* edit test

* fix test

* fix bug

* fix test

* Modify the thread pool value capture to reference capture to fix the issue of significant performance degradation when writing files with put.

* add async getfrom file in batchget transfertask. delete file_storage_backend

* add support for HA in cluster_id subdirectory, change session_id to fsdir

* add persistence in batchput

* add disk allocate for get_into py interface

* fix bug

* temp

* fix bug in submit fileread task for std:move(slices)

* edit querykey to return optional<descriptor>, add interface batchquerykey for storagebackend

* fix confict in batchget, add batchget/batchput test

* fix bug

* fix bug

* comment batch test

* fix conflict and add batch_get_into file test

* fix test bug

* fix test bug

* fix conflict
2025-07-02 19:15:10 +08:00
AscendTransport 6ff0f35211
[TransferEngine] Enable Huawei Ascend Transport for TransferEngine (#502)
* Ascend Transport

* fix ci

* reviewed changes
2025-07-01 15:47:34 +08:00
JinYan Su 8e2e0adcb4
feat(store): add zero-copy operations for python binding (#532)
* feat(store): add zero-copy operations for python binding

* test: rename dict fuzz e2e test to run last

* test: remove obsolete test_multicards.py from repository

* chore(tests): remove multicards test execution from script
2025-06-20 17:46:24 +08:00
Shangming Cai a79f770bcc
[Build] Optimize store build control for wheel and local build (#531)
* [Build] Optimize store build control for wheel and local build

Signed-off-by: Shangming Cai <caishangming@linux.alibaba.com>

* fix typo

Signed-off-by: Shangming Cai <caishangming@linux.alibaba.com>

* fix rm

Signed-off-by: Shangming Cai <caishangming@linux.alibaba.com>

---------

Signed-off-by: Shangming Cai <caishangming@linux.alibaba.com>
2025-06-20 15:46:35 +08:00
Shangming Cai f0cb5618c9
[Build] Deprecate stale adaptor usage to reduce whl package size (#529)
Signed-off-by: Shangming Cai <caishangming@linux.alibaba.com>
2025-06-20 14:19:41 +08:00
Shangming Cai 4e78b6e676
[Build] Add allocator class to support nvlink for more use-cases (#524)
Signed-off-by: Shangming Cai <caishangming@linux.alibaba.com>
2025-06-19 20:00:09 +08:00
Shangming Cai 0a8e4c3074
[Build] Optimize nvlink allocator build logic and fix name issue (#523)
Signed-off-by: Shangming Cai <caishangming@linux.alibaba.com>
2025-06-19 19:51:52 +08:00
Teng Ma e6bbc3dae8
[Build] add TE bench into wheel package (#514) 2025-06-18 19:04:09 +08:00
JinYan Su 31b814a4af
feat(release): add support for arm64 architecture (#486)
* feat(release): add support for arm64 architecture

* Update scripts/build_wheel.sh

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>

---------

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
2025-06-14 12:08:59 +08:00
Teng Ma 4d11e25f28
[Build] fix build wheel if nvlink is disabled (#480)
* [Build] fix build wheel if nvlink is disabled

* fix
2025-06-13 00:08:08 +08:00
JinYan Su 9c95392d07
chore: automate build output directory and update scripts (#460)
* chore: automate build output directory and update scripts

* feat(ci): enable CUDA support in workflows

Signed-off-by: Jinyang Su <751080330@qq.com>

* chore(build): switch to nvcc for nvlink hook compilation

Signed-off-by: Jinyang Su <751080330@qq.com>

---------

Signed-off-by: Jinyang Su <751080330@qq.com>
2025-06-10 10:45:36 +08:00
jiafu zhang 3c27b0bd91
leave endpoint status unchanged when delete endpoint reference to avoid endpoint deconstruction before CQ being generated (#384)
* leave endpoint status unchanged when delete endpoint reference to avoid endpoint deconstruction before CQ being generated

Signed-off-by: jiafu.zhang <jiafu.zhang@intel.com>

* leave endpoint status unchanged when delete endpoint reference to avoid endpoint deconstruction before CQ being generated

Signed-off-by: jiafu.zhang <jiafu.zhang@intel.com>

* leave endpoint status unchanged when delete endpoint reference to avoid endpoint deconstruction before CQ being generated

Signed-off-by: jiafu.zhang <jiafu.zhang@intel.com>

---------

Signed-off-by: jiafu.zhang <jiafu.zhang@intel.com>
2025-05-27 11:23:50 +08:00
doujiang24 04b29e4b5a
[Build]: exclude cuda so files in auditwheel. (#379)
Signed-off-by: doujiang24 <doujiang24@gmail.com>
2025-05-20 15:18:47 +08:00
maobaolong 076022eda8
[Integration] feat(Mooncake Integration): Supply a MooncakeConfig into whl file (#338)
* supply a MooncakeConfig into whl

* Add example comment and require fields check

* Add test to ci

* Revert the rename to keep consistence
2025-05-12 17:52:19 +08:00
Johnny da24425450
[Build] ARM build_wheel.sh (#344) 2025-05-11 14:15:09 +08:00
JinYan Su a6a491dcc7
[Mooncake Integration] feat: add pybind support for Python 3.12 (#249)
* feat: use pyproject.toml and simplify setup.py

* feat: add apache license and update pyproject.toml

* chore: use string quotes for python version in CI matrix

* feat: build wheels for python 3.10 and 3.12

* chore: update wheel build process in CI/CD

* ci: refactor ci and release workflows for matrix builds

* chore: remove unused workflow and fix license path

* feat: initialize git submodules in dependencies.sh
2025-04-15 15:46:11 +08:00