This proposal outlines a two-three week effort to measure real GPU power and energy during single-turn 8k-input/1k-output inference on MI355X, B200, GB200, and GB300. The final deliverable will be a reproducible cross-platform dataset and an article draft explaining the results.

Model

I recommend using Qwen/Qwen3.5-397B-A17B-FP8 as the shared anchor model. Matching 8k/1k recipes already exist on all four platforms:

This should be described as a deployed-platform comparison, not a pure silicon comparison, because the GB systems use a different serving topology from MI355X and B200. The practical question is: at a given level of interactivity or throughput, how much GPU energy does each deployed system use? Concurrency controls the load but is not the final comparison metric.

The core experiment covers concurrency {4, 16, 64} on four platforms, with three independent repeats per point, for 36 accepted results. If the core data is clean and the article needs smoother curves, {8, 32, 128} can be added afterward. Prefill and decode energy will be reported separately only for GB200 and GB300, where those worker roles are directly observable. MI355X and B200 will remain end-to-end measurements in v1.

What I need from reviewers

Please confirm or adjust these defaults:

  1. Use one anchor model, Qwen3.5 FP8, for v1 — recommended: yes.
  2. Use core load points {4, 16, 64} with three independent repeats — recommended: yes.
  3. Compare platforms by achieved interactivity and throughput; treat concurrency as the controlled load — recommended: yes.
  4. Use 1P1D only on GB200 and GB300 — recommended: yes. Treat 4P1D and 8P1D as a separate scale-out study.
  5. Treat multiple-model coverage as follow-up work unless it remains a required v1 article claim.
  6. Reuse or coordinate with Aryaman’s PR #1635 and selectively port useful ideas from PR #1574.
  7. Name one methodology reviewer and one cluster escalation owner.

1. Objective and scope