CuTeDSL
Agent skill
Scope markers delimit regions of a CuTeDSL kernel. G-Watch records when each region runs on every thread.
Example
The runnable example is
examples/cuda/trace/trace_cute_matmul.py.
pip install nvidia-cutlass-dsl
# for humanized visualization
python3 examples/cuda/trace/trace_cute_matmul.py --report trace.html
# for agentic visualization
python3 examples/cuda/trace/trace_cute_matmul.py --report trace.json
gwatch show trace.json
Step 1: Mark scopes in the kernel
Wrap each region with scope_start and scope_end from the CuTe trace helper.
Markers go inside the @cute.kernel body.
Each region needs a unique integer id.
import gwatch.cuda.xtrace.cute as gw_trace
import cutlass.cute as cute
@cute.kernel
def my_kernel(x: cute.Tensor, out: cute.Tensor, n: Int32):
tidx, _, _ = cute.arch.thread_idx()
bidx, _, _ = cute.arch.block_idx()
i = bidx * BLOCK_SIZE + tidx
if i < n:
gw_trace.scope_start(1)
v = x[i]
gw_trace.scope_end(1)
out[i] = v * 2.0
Step 2: Build and trace
CuTe's JIT cubin skips CUPTI's module capture. So tracing reads CuTe's dumped PTX instead. The order of these steps matters:
- Set
CUTE_DSL_KEEP_PTX=1andCUTE_DSL_DUMP_DIR. - Import
gwatch.cuda.xtrace.cute, which creates the capsule. - Call
init_cupti_hooks(). - Import
cutlasslast.
This way CUPTI is already listening when @cute.jit loads the module.
import os
os.environ.setdefault("CUTE_DSL_KEEP_PTX", "1")
os.environ.setdefault("CUTE_DSL_DUMP_DIR", "/tmp/gw_cute_ptx")
import gwatch.libpygwatch as pygwatch
import gwatch.cuda.xtrace.cute as gw_trace # creates the capsule
from gwatch.cuda.xtrace import do_trace
pygwatch.init_cupti_hooks() # install CUPTI hooks
import cutlass # import cutlass only now
# ... define the @cute.kernel above (with scope markers) and its @cute.jit launcher ...
result = do_trace(
fn=lambda: run_my_kernel(...), # first execution of the kernel
kernel_name_pattern=r".*my_kernel.*", # regex on the mangled prototype
dsl="cute", # CuTeDSL: trace uses CuTe's dumped PTX
scope_name_map={1: "load"}, # optional: id -> label
instrumentation_tier="ptx",
)
A few things to note:
instrumentation_tier="ptx"traces the scopes marked in source. To trace a compiled cubin without markers, see SASS.- PTX dump. The env vars and the import order above let tracing find the kernel's PTX.
dsl="cute"tells G-Watch to recover PTX from CuTe's dump directory.scope_name_map(optional) turns the integer ids into the labels shown in the report.
Step 3: Render the report
from gwatch.common.format import File
from gwatch.cuda.xtrace.format import Section_IntraKernelTrace
section = Section_IntraKernelTrace()
section.add_run(result)
report = File(title="Xtrace")
report.add_section(section)
report.render("trace.html") # interactive panel; use .json for the agent archive