forked from ccf-ai-infra/Intro-ops
Co-authored-by: wawahejun <hejunlbbc@gmail.com> |
||
|---|---|---|
| cmake | ||
| docs | ||
| examples | ||
| include/operator_runtime | ||
| ops | ||
| python | ||
| tests | ||
| .cnb.yml | ||
| .gitignore | ||
| CMakeLists.txt | ||
| README.md | ||
| README.zh.md | ||
| requirements.txt | ||
README.md
Operator Runtime Training Camp
This repository is a training-oriented GPU operator runtime. It is intentionally small, but its workflow mirrors production operator libraries:
- Create a directory under
ops/<op>/with backend implementations. - The build system auto-discovers sources by directory convention.
- Implement a backend-specific descriptor lifecycle (create, workspace, execute, destroy).
- Expose a Python API with out-of-place, out-variant, and prepared execution.
- Validate correctness against PyTorch and benchmark steady-state execution.
Python Layout
python/
operator_runtime/
backend.py
ops/
_internal/
operator_runtime_testing/
operator_runtime.opscontains public operator bindings.operator_runtime._internalcontains private FFI/runtime plumbing.operator_runtime_testingcontains test-only helpers such as assertions and benchmark utilities.
Operators
| Operator | NVIDIA C++ | TileLang | MetaX |
|---|---|---|---|
copy |
runnable | runnable when TileLang is installed | stub |
vector_add |
runnable | runnable when TileLang is installed | stub |
reduce_sum |
runnable, row-wise fp32 | runnable when TileLang is installed | stub |
softmax |
runnable, row-wise fp32 | runnable when TileLang is installed | stub |
Setup
pip install -r requirements.txt
Build
mkdir -p build
cd build
cmake .. -DCAMP_ENABLE_NVIDIA=ON -DCAMP_ENABLE_METAX=OFF
cmake --build . -j$(nproc)
Validate
python tests/run_ops.py --op copy --backend nvidia --mode all
CAMP_BUILD_DIR=build pytest tests/ -v --backend nvidia
pytest tests/ -v --backend tilelang
python tests/run_ops.py --op all --backend nvidia --mode bench
The TileLang backend requires the tilelang Python package.
Adding a New Operator
- Create
ops/<name>/nvidia/<name>_cuda.hwith the C API (4 functions: create, workspace, execute, destroy). - Create
ops/<name>/nvidia/<name>_cuda.cuwith the CUDA implementation. - Create
python/operator_runtime/ops/<name>.pyusingoperator_runtime._internal.bind_*functions. - Create
tests/cases/<name>.pywithcorrectness_cases(),api_error_cases(), andbenchmark_cases(). - Create
tests/ops/test_<name>.pyandtests/bench/<name>.py. - Re-run
cmake ..in the build directory (the glob will pick up the new.cufile). - Register the public API in
python/operator_runtime/__init__.py.
No YAML, no code generation, no registration step.
Production Mapping
| Training concept | Production equivalent |
|---|---|
directory convention ops/<op>/nvidia/*.cu |
build system auto-discovery / operator registry |
C header ops/<op>/nvidia/<op>_cuda.h |
reviewed operator API contract |
| descriptor lifecycle | create, workspace, execute, destroy |
tests/cases/<op>.py |
correctness, layout, and API contract coverage |
PerformanceResult |
profiler report row with latency, bytes, flops, bandwidth |
| eager TileLang kernel | puzzle-stage kernel using T.empty(...) return values |
lazy TileLang out_idx template |
TileOPs-style kernel factory and output-position contract |