Group modules by responsibility:
- _runtime/: internal C ABI layer (loader, ctypes_bindings,
prepared, tensor_view); underscore prefix signals non-public API
- testing/: test utilities split by concern (assertions, benchmark,
profiler) instead of mixed in top-level files
- ops/__init__.py: per-operator registration; adding a new operator
only requires one line here, not touching the top-level __init__
- operator_runtime/__init__.py: thin gateway that re-exports from ops/
External import paths for tests are unchanged:
operator_runtime.testing still exports require_cuda, assert_close,
cuda_time_ms, PerformanceResult
Co-authored-by: wawahejun <hejunlbbc@gmail.com>
Replace the YAML manifest + Python code-generation pipeline with
CMake file(GLOB_RECURSE) auto-discovery. Operators are now registered
by directory convention (ops/<name>/nvidia/*.cu) rather than by
declaring sources and symbols in operator.yaml.
Deleted: ops/*/operator.yaml, cmake/GenerateOperators.cmake,
tools/generate_operator_artifacts.py, tools/operator_schema.yaml,
tools/validate_operator_manifest.py, operator_registry.py,
tests/test_manifest_contract.py
Modified: ops/CMakeLists.txt uses GLOB_RECURSE, requirements.txt
drops pyyaml
Co-authored-by: wawahejun <hejunlbbc@gmail.com>