Closes #1741
## Summary
- Add a shared `isolated_dynamo` fixture in `tests/conftest.py` that
runs `torch._dynamo.reset()` before and after every `torch.compile`
test.
- Request the fixture from all compile tests (`test_compile.py`,
`test_elementwise_compile.py`, `test_norm_ops.py`, `test_pool.py`),
replacing the module-local reset fixture in
`test_elementwise_compile.py`.
- Fixes `FailOnRecompileLimitHit` on all three `test_mha_kernel_compile`
cases in the push-tier gpu-smoke run (cross-test dynamo cache pollution,
not intrinsic guard churn).
## Test plan
- [x] Modified files pass unit tests (297/297 nodes across the 4 touched
test files at 8de7df1e; pre-commit green)
- [x] All three `test_mha_kernel_compile` parametrizations pass under
the push-tier selection `-m "smoke or full" -k compile` with push-tier
ordering on H200 (11/11; 3 failed pre-fix). CI push-tier confirmation
lands on the post-merge run — PR-event gpu-smoke runs `-m smoke` only.
- [x] This PR documents the flip commit/mechanism and the layer choice
(see Regression)
- [ ] Subsequent push-to-main GPU Smoke run is green (verifiable only
after merge)
## Regression
- **Flip commit**:
|
||
|---|---|---|
| .claude | ||
| .foundry | ||
| .github | ||
| assets | ||
| benchmarks | ||
| docs | ||
| scripts | ||
| tests | ||
| tileops | ||
| workloads | ||
| .gitignore | ||
| .pre-commit-config.yaml | ||
| CLAUDE.md | ||
| LICENSE | ||
| Makefile | ||
| README.md | ||
| THIRD_PARTY_NOTICES.md | ||
| constraints.txt | ||
| pyproject.toml | ||
README.md
TileOPs
Spec-driven GPU operator library for LLMs — designed for AI agents to build, evaluate, and optimize
Built on TileLang
Status: TileOPs is under active development. APIs may change.
Overview
TileOPs is a GPU operator library for LLM training and inference, built on TileLang. Beyond providing a growing collection of production-quality operators, TileOPs explores a spec-driven development model where AI agents can read declarative operator specifications, generate kernel implementations, and evaluate them against hardware-theoretical performance bounds — with minimal human scaffolding.
Architecture
Every operator is split into two layers with a strict boundary:
- Op (L2) — stateless Python entry point. Handles validation, dtype casting, and memory layout. Compatible with CUDA-Graph and
torch.compile. - Kernel (L1) — TileLang GPU implementation with hardware-specific optimizations (Ampere, Hopper).
This separation keeps user-facing behavior independent of GPU strategy, allowing agents and developers to modify either layer without side effects on the other.
Key Properties
- Spec-driven — each operator is declared in a machine-readable manifest (
tileops/manifest/) that specifies signatures, workloads, and roofline formulas, serving as the entry point for both agent code generation and automated validation - Roofline-evaluated — kernel performance is measured against Speed-of-Light hardware bounds, not relative baselines
- Auto-tuning — built-in search over tile sizes, pipelines, and scheduling parameters
- Lightweight — depends only on TileLang, PyTorch, and einops
Installation
TileOPs can be installed from PyPI or built from source. A CUDA-capable GPU is required.
Prerequisites
- Python >= 3.10
- PyTorch >= 2.1
- CUDA Toolkit
- NVIDIA GPU: Hopper (SM_90)
- TileLang == 0.1.9
From PyPI
pip install tileops
From source
git clone https://github.com/tile-ai/TileOPs
cd TileOPs
make install # dev dependencies + pre-commit hooks
[!NOTE] If CUDA and TileLang are already installed system-wide and you encounter build issues:
PIP_NO_BUILD_ISOLATION=1 pip install -e '.[dev]' -v && pre-commit install
Verify:
python -m pytest tests/ -q # requires a CUDA GPU
Quick Start
import torch
from tileops.ops import GemmOp
M, N, K = 1024, 1024, 512
dtype = torch.float16
gemm = GemmOp(M, N, K, dtype=dtype)
A = torch.randn(M, K, device="cuda", dtype=dtype)
B = torch.randn(K, N, device="cuda", dtype=dtype)
C = gemm(A, B)
Documentation
Design docs and development guides are in docs/. The full API reference and performance tables are published at TileOPs.github.io.
Contributing
See docs/ for design docs. Branch and commit conventions are in .claude/conventions/types.sh.
License
TileOPs is released under the MIT License.