G-Watch.玑衡
Open Traces文档博客GitHub
English·中文
  • Getting Started

    • Introduction
    • Installation
  • Intra-kernel Tracing
    • Supported DSLs
      • CUDA
      • TileLang
      • CuTeDSL
      • Triton
    • Visualize Xtrace for Agent
  • Intra-kernel Profiling
    • Counter-based Profiling
    • PC Sampling
  • Binary Analysis
    • Control Flow
      • Control Flow Graph
      • Control Divergence Graph
      • Warp Specialization
      • Software Pipeline
    • Data Flow
      • Data Flow Graph
      • SMEM Usage
      • DMEM/SMEM Access Pattern
    • Hardware Execution
      • Instruction Scheduling
      • RMEM Liveness
      • RMEM Cache

Documentation

文档

每篇文档都提供 humanize(人类友好)与 agent(机器友好)两种视图。点击进入即为 humanize 渲染,页内可切换到 agent 视图。

Getting Started

Introduction

Installation

Pick your platform and CUDA version to install G-Watch, or build it from source.

Intra-kernel Tracing

CUDA

Trace the internal phase timeline of a hand-written CUDA C++ kernel by adding device-side scope markers.

TileLang

Trace the internal phase timeline of a TileLang kernel by adding device-side scope markers.

CuTeDSL

Trace the internal phase timeline of a CuTeDSL kernel by adding device-side scope markers.

Triton

Trace the internal phase timeline of a Triton kernel by adding device-side scope markers.

Visualize Xtrace for Agent

Render an intra-kernel trace as compact text that an LLM agent can read, with gwatch show.

Intra-kernel Profiling

Counter-based Profiling

PC Sampling

Binary Analysis

Control Flow Graph

Control Divergence Graph

Warp Specialization

Software Pipeline

Data Flow Graph

SMEM Usage

DMEM/SMEM Access Pattern

Instruction Scheduling

RMEM Liveness

RMEM Cache

G-Watch.

为面向智能体的 GPU/TPU kernel 优化提供性能剖析与程序分析。

G-Watch 是 Mars Compute 旗下的开源项目。

Open Traces文档博客GitHub

G-Watch 2024 – 2026