Mooncake/mooncake-integration
zbtrs2 fd7bbf720a ollama-kvcache: add deployment, observability and docs
deploy packages the topology as docker-compose (master, store proxy,
sidecar, patched Ollama, Prometheus, Grafana) with the RDMA/GPUDirect
host passthrough it needs, a Prometheus scrape config and a Grafana
dashboard whose primary panel integrates the live count of avoided
prefill tokens.

docs/REPORT.md is the full write-up: design, what is real versus
simulated, the measured store bandwidth, the multi-agent results, the
adaptive-arbiter behaviour, the Stage-2 microbenchmark and the honest
limitations. The README is the entry point and quickstart, and the
figures and result JSONs are the measured data behind the report.
2026-06-29 02:13:48 +08:00
..
ollama ollama-kvcache: add deployment, observability and docs 2026-06-29 02:13:48 +08:00
store [Store] Expose batch_replica_clear in Python binding (#1848) 2026-04-17 17:36:32 +08:00
transfer_engine [TransferEngine][ROCm] Add ROCm HIP support to the Mooncake Python package (#1742) 2026-04-13 11:13:10 +08:00
CMakeLists.txt [Docker] fix(docker): respect PYTHON_VERSION build-arg when building wheel (#1745) 2026-04-10 09:56:19 +08:00
allocator.py [TE] Add early mem backend detection method in NVLINK_allocator (#1296) 2026-01-04 15:20:52 +08:00
allocator_ascend_npu.py [TE] Ubshmem transport support ipc memory and build allocator when set USE_UBSHMEM=ON (#1519) 2026-02-14 19:30:06 +08:00
integration_utils.h [Store] validate metadata when put_tensor (#1396) 2026-01-26 16:31:48 +08:00