<< All versions

Skill v1.0.2

currentAutomated scan100/100
bbuf/ai-infra-auto-driven-skills/llm-pipeline-analysis

~5 modified

──Details
PublishedSeptember 29, 2026 at 11:01 PM
Content Hashsha256:b44c9c7d425fe8cd...
Git SHA475af5a803db
Bump Typepatch
Compare with v1.0.1
──Files
Files (1 file, 12.4 KB)
SKILL.md12.4 KBactive
SKILL.md · 307 lines · 12.4 KB

version: "1.0.2" name: llm-pipeline-analysis description: "Inspect LLM torch profiler traces at forward-pass, layer, and kernel level. Use when you need layer timings, anchor-kernel boundaries, representative kernel flows, or Perfetto time ranges."


LLM Pipeline Analysis

Overview

Use this when a whole-trace profiler summary is too coarse. The scripts read a Chrome-trace JSON file, find layer-boundary anchor kernels, group kernels into forward passes and layers, and print timing tables you can use for Perfetto navigation or detailed timing analysis.

When To Use It

  • when you need to know which layers contribute most
  • when the model has alternating layer types (e.g. models with

compress_ratios like DeepSeek-V4 NSA, or hybrid GDN/GQA stacks such as Qwen3.8-27B)

  • when you need to compare cold-start vs steady-state forward passes
  • when you need to navigate to a specific layer in Perfetto UI
  • when you need to select representative layers for deep-dive analysis

Verify inputs from artifacts

Resolve model config, phase, rank/device and TP/EP/DP from the run manifest, server arguments and trace. Ask only for missing information that changes the analysis. TP-0 identifies a rank, not TP world size; do not default an unknown run to TP8/EP8 or decode BS1. Speculative request BS and target-verify M differ. If no config is available, use an explicit profile and verified layer count.

Model Profiles

Scripts use ModelProfile to determine layer boundary detection and kernel classification. Profiles are auto-inferred from config.json or selected via --profile:

ProfileAnchor kernelBlocks/layerLayer structureAuto-infer condition
dsv41explicit verified --anchor-kernel1target or draft phase, verified separatelymodel_type=deepseek_v41
dsv4_csa_hcamhc_post_tilelang2attn + ffn halvescompress_ratios non-empty
dsv3_mlaflash_fwd_mla_combine1full layerkv_lora_rank > 0
genericrepeated RMSNorm/AllReduce or --anchor-kernel1full layerfallback

Use --profile generic --anchor-kernel YOUR_KERNEL for models not covered by built-in profiles. Generic TP=1 traces can auto-detect a repeated RMSNorm anchor without requiring an NCCL AllReduce marker.

Prerequisites

  • A torch.profiler trace in Chrome-trace JSON format (.json or .json.gz)
  • The model's config.json (for profile inference, compress_ratios, etc.)
  • The trace must contain a recognizable layer-boundary anchor kernel

(auto-detected from the profile, or specified via --anchor-kernel)

Layer Boundary Detection

The scripts use an anchor kernel as a layer-boundary marker. The anchor and layer structure are determined by the active ModelProfile.

The legacy dsv4_csa_hca profile assumes an unfused trace with two matching mHC post calls per layer. This does not hold for every current V4/V4.1 path. The dsv41 profile refuses automatic anchor selection: verify a once-per-layer anchor in one phase and use the matching config. Read DSV4.1 lessons before interpreting mHC/AR fusion, PDL or shared-expert overlap. Anchor intervals are navigation guides; kernels on another stream can cross their boundaries.

For an unfused trace compatible with dsv4_csa_hca, each transformer layer produces 2 consecutive mhc_post_tilelang calls:

mhc_post_tilelang ← end of attn half (attention + O-proj + AllReduce)
... ffn computation ...
mhc_post_tilelang ← end of ffn half (MoE experts + AllReduce)
... next layer attn ...
mhc_post_tilelang ← next layer's attn boundary

So for N layers with the dsv4_csa_hca profile, one forward pass has 2N anchor blocks. With dsv3_mla or generic, each layer has 1 block.

Forward pass P starts at block index P * (N * blocks_per_layer).

Scripts

1. layer_timeline_analyzer.py — Per-layer timeline and cluster stats

bash
# Show all forward passes summary (cold-start vs steady-state)
python3 scripts/layer_timeline_analyzer.py \
--trace /path/to/TP-0.trace.json.gz \
--config /path/to/config.json \
--show-all-passes
# Detailed per-layer breakdown for a specific forward pass
python3 scripts/layer_timeline_analyzer.py \
--trace /path/to/TP-0.trace.json.gz \
--config /path/to/config.json \
--fwd-pass 5
# Auto-select the first relatively stable pass window
python3 scripts/layer_timeline_analyzer.py \
--trace /path/to/TP-0.trace.json.gz \
--config /path/to/config.json

The script prints:

  • Per-layer wall-clock time, sum-duration, and category breakdown (MLA, MoE, GEMM, NCCL, MHC, Hadamard)
  • Layer cluster statistics grouped by type (C4_LIGHT, C128_HEAVY, HASH, etc.)
  • All-passes summary showing cold-start → steady-state growth

Automatic steady-state selection requires two consecutive layer-0 timing changes within 5%. It does not use an absolute latency threshold, so the same rule works across model sizes and accelerators. If no stable window exists, choose --fwd-pass explicitly.

2. layer_kernel_breakdown.py — Per-layer kernel detail and compute flow

bash
# Single layer kernel dump
python3 scripts/layer_kernel_breakdown.py \
--trace /path/to/TP-0.trace.json.gz \
--config /path/to/config.json \
--fwd-pass 5 --layer 3
# Compute flow format (with model architecture summary and category column)
python3 scripts/layer_kernel_breakdown.py \
--trace /path/to/TP-0.trace.json.gz \
--config /path/to/config.json \
--fwd-pass 5 --layer 3 --format compute-flow
# JSON export
python3 scripts/layer_kernel_breakdown.py \
--trace /path/to/TP-0.trace.json.gz \
--config /path/to/config.json \
--fwd-pass 5 --layer 3 --format json
# Compare two layers side-by-side
python3 scripts/layer_kernel_breakdown.py \
--trace /path/to/TP-0.trace.json.gz \
--config /path/to/config.json \
--fwd-pass 5 --layer 2 --compare-layer 3

Output formats:

  • --format text (default): grouped summary + top hot kernels ranked by duration, with simplified names and percentages
  • --format compute-flow: model architecture summary + per-kernel hotness table with Category, %, and ts_rel(ms) columns
  • --format json: one machine-readable JSON document ranked by duration; with

--compare-layer, it contains primary, comparison, and kernel_diff

  • Kernel diff when comparing two layers (unique kernels in each)

3. perfetto_time_mapper.py — Perfetto UI time navigation

bash
# Show all forward pass time ranges in Perfetto
python3 scripts/perfetto_time_mapper.py \
--trace /path/to/TP-0.trace.json.gz \
--config /path/to/config.json
# Layer-level time ranges for a specific forward pass
python3 scripts/perfetto_time_mapper.py \
--trace /path/to/TP-0.trace.json.gz \
--config /path/to/config.json \
--fwd-pass 5 --layers 2,3,38,42

The script prints:

  • Forward pass time ranges in Perfetto-relative seconds
  • Per-layer start/end times with compress_ratio labels

Workflow

Step 1: Identify steady-state forward pass

bash
python3 scripts/layer_timeline_analyzer.py \
--trace $TRACE --config $CONFIG --show-all-passes

Read the "all-passes" table. The first pass is cold-start (few tokens). Use the first relative-stability window selected by the script, or choose a pass explicitly when timings continue to change.

Step 2: Per-layer breakdown on steady-state pass

bash
python3 scripts/layer_timeline_analyzer.py \
--trace $TRACE --config $CONFIG --fwd-pass 5

Identify:

  • Which layer type dominates (C4_LIGHT vs C128_HEAVY vs HASH)
  • The MLA / MoE / GEMM / NCCL proportion per layer type
  • Which layer type is the best next target

Step 3: Compute flow for representative layer(s)

Select 1-2 representative layers (one per bottleneck type), then:

bash
# Human-readable compute flow table
python3 scripts/layer_kernel_breakdown.py \
--trace $TRACE --config $CONFIG \
--fwd-pass 5 --layer 3 --format compute-flow
# JSON export
python3 scripts/layer_kernel_breakdown.py \
--trace $TRACE --config $CONFIG \
--fwd-pass 5 --layer 3 --format json > /tmp/layer3_detail.json

The --format compute-flow output includes:

  • Model architecture summary at the top
  • Per-kernel hotness table with # | Half | Category | Simplified Name | dur(us) | % | ts_rel(ms) | Input Dims
  • Rows are ranked by dur(us) descending by default; use ts_rel(ms) to jump back to the kernel's trace location.

Step 4: Compare layer types (optional)

bash
python3 scripts/layer_kernel_breakdown.py \
--trace $TRACE --config $CONFIG \
--fwd-pass 5 --layer 2 --compare-layer 3

This shows the exact kernel difference between the two layer types.

Step 5: Navigate in Perfetto UI (optional)

bash
python3 scripts/perfetto_time_mapper.py \
--trace $TRACE --config $CONFIG \
--fwd-pass 5 --layers 2,3,38,42

Use the printed time ranges to navigate directly in Perfetto.

Layer Type Classification

The scripts classify layers based on config.json fields:

Config fieldValueLayer TypeDescription
compress_ratios[i]0FULL_ATTNNo NSA compression (layers 0-1)
compress_ratios[i]4C4_LIGHTC128 sparse attention, fastest
compress_ratios[i]128C128_HEAVYC4 attention + Hadamard + Indexer, bottleneck
i >= N - num_hash_layers—HASHHash-table routing with paged MQA
i == 0—FIRSTFirst layer (empty KV cache)
i == N - 1—FINALFinal layer (lm_head output)

Kernel Categories

Kernels are classified by the active ModelProfile's rules. Categories marked with (DSv4) are specific to the dsv4_csa_hca profile; all profiles include the universal categories.

CategoryMatch PatternProfileTypical Share (DSv4)
★ MLA Attentionflash_fwd_splitkv_mlaDSv4, DSv321-33%
★ MoE Fusedfused_moe_kernelDSv4, DSv311-17%
● NCCL AllReduceAllReduceuniversal5-8%
GEMM fp8deep_gemmuniversal12-25%
GEMM bf16nvjetuniversal11-13%
Hadamard XformhadamardDSv40-2.4%
Indexer CacheindexerDSv40-0.1%
Paged MQApaged_mqa_logitsDSv40-1.8%
MHCmhc_pre_gemm_sqrsum, mhc_pre_big_fuse, mhc_post_tilelangDSv410-15%
C4/C128 Prefillc4_prefill, c128_prefillDSv40-0.3%
RMSNormRMSNorm, rms_normalizeuniversal1-2%
FP8 Quantquant, Quantuniversal1-2%
TopKtopkuniversal0-0.7%
RoPEdeepseek_rope, fused_norm_ropeDSv4, DSv31-2%
Activationsilu_mul_clamp, act_and_muluniversal0-0.5%
Other—universal2-5%

Reporting Checklist

Include:

  1. Trace metadata: trace path, model config path, GPU type, TP/EP
  2. Model Architecture Summary (from config.json):
  • model name, num_layers, hidden_size, num_attention_heads, num_key_value_heads, head_dim
  • Attention type (e.g. csa_hca), Q/O LoRA ranks
  • MoE config: num_experts, topk, num_shared_experts, intermediate_size
  • MHC config (if applicable)
  • NSA config (if applicable): index_n_heads, index_head_dim, index_topk, qk_rope_head_dim, sliding_window
  • compress_ratios distribution (how many C4_LIGHT / C128_HEAVY / FULL_ATTN / HASH layers)
  1. Per-batch forward passes summary table (from layer_timeline_analyzer.py --show-all-passes):
  • Columns: Fwd#, Start(s), End(s), Duration(ms), Avg Layer(ms), First Layer(ms), Notes
  • Identifies cold-start vs steady-state passes
  1. Chosen forward pass: index and rationale (cold-start vs steady-state)
  2. Per-layer wall-clock and sum-duration table (from layer_timeline_analyzer.py --fwd-pass N):
  • Columns: L, c_r, Type, Wall(ms), SumDur(ms), MLA, MoE, GEMM, NCCL, MHC, Hadam, AR#, K#
  • Each row is one layer, with layer type label
  1. Layer cluster statistics table grouped by type:
  • Columns: Cluster, #, Avg Wall(ms), Avg Sum(ms), MLA%, MoE%, GEMM%, NCCL%, MHC%, Hadam%
  • Identifies bottleneck layer type and likely next target
  1. Compute Flow Table for selected representative layer(s):
  • Produced by layer_kernel_breakdown.py --format compute-flow
  • Columns: # | Half | Category | Simplified Name | dur(us) | % | ts_rel(ms) | Input Dims
  • Rows are sorted by top hot kernels (dur(us) descending) by default
  • Optional JSON export (--format json)
  1. Perfetto UI time ranges when requested
  2. One-line summary: bottleneck layer type and likely next target
← v1.0.1All versions