GPU Infrastructure
NVIDIA H100 / H200 — Certification & Benchmarking Validation
4 nodes running
Nodes in queue
12
3 added today
Running now
4
avg 2h 14m elapsed
Passed (30d)
89
92% pass rate
Failed (30d)
8
5 NVLink · 2 HBM · 1 PCIe
Avg cert time
3h 42m
per 8-GPU node
Phase progress
    Overall progress62%
    Live metrics
    GPU thermal map — node-001
    <75°C — nominal 75–82°C — warm >82°C — flag
    Event log
    18
    tests passed
    6
    tests running
    8
    tests pending
    1
    Stage 1 — Firmware & Config Baseline
    Environment must be fully configured before any hardware test runs
    4 pass~12 min
    System preparation
    Drivers · firmware · runtime stack · cluster config
    4 pass
    ~12 min
    2
    Stage 2 — Hardware Health Diagnostics
    Validate individual GPU health, storage, networking baseline, and observability
    7 pass11 pending~50 min
    GPU health & diagnostics
    Enumeration · DCGM L1–L4 · ECC · XID · PCIe · inforom
    7 pass
    ~20 min
    Ethernet & DNS validation
    iperf3 frontend throughput · NIC link state · DNS health checks
    3 pending
    ~10 min
    Storage benchmarks
    fio · seq read/write BW · random IOPS · P99 latency · NVMe + block
    5 pending
    ~20 min
    Observability & continuous monitoring
    Telegraf · cluster-level dashboards · host-level GPU health · 24×7 alerting
    3 pending
    ongoing
    3
    Stage 3 — Interconnect & Fabric Performance
    Validate intra-node NVLink and inter-node InfiniBand / RoCE at full throughput
    3 running5 pending~65 min
    NVLink & NVSwitch validation
    nvbandwidth GPU↔GPU matrix · NCCL intra-node · NVSwitch fault isolation
    3 running
    ~25 min
    InfiniBand / RoCE fabric
    ibping · ib_read/write_bw · GPUDirect RDMA · NCCL multi-node · switch isolation
    5 pending
    ~40 min
    4
    Stage 4 — AI Workload Benchmark
    Prove production readiness under sustained AI training load — the final gate before certification
    3 pass3 pending~95 min
    GPU stress & burn-in
    gpu-burn sustained · DCGM targeted stress · pulsed power · thermal soak
    3 pass
    ~35 min
    End-to-end workload validation — MFU
    PyTorch FSDP · Llama-3 8B · TPS · MFU % · all_reduce latency · 16 nodes
    3 pending
    ~60 min
    GPU-level diagnostics
    GPU ECCPCIeHBM BWPowerTempXID errorsDCGM L4
    Power draw (W)
    HBM bandwidth (TB/s)
    InfiniBand / GPUDirect RDMA
    Fabric speed400 Gb/s HDR
    NCCL all_reduce BW372 GB/s (93%)
    ib_write_bw390 GB/s
    ib_read_bw388 GB/s
    ibping latency1.2 µs
    GPUDirect RDMAenabled
    IB error rate0.00%
    Multi-node NCCL (4 nodes)pass
    Ethernet (iperf3) & DNS
    Frontend NIC25 GbE
    iperf3 TX24.7 Gbps
    iperf3 RX24.6 Gbps
    DNS — S3 / R20 errors
    DNS — wandb.ai0 errors
    All NIC link statesup
    NVMe storage (fio)
    Sequential read BW21.8 GB/s
    Sequential write BW18.3 GB/s
    Random read IOPS1.04M
    Random write IOPS820K
    P99 write latency2.8 ms
    Checkpoint stalls0
    MFU reference run — Llama-3 8B / FSDP
    Tokens / secpending
    Model flops utilizationpending
    GPU utilizationpending
    all_reduce latencypending
    Phase 9 starts after NVLink validation completes
    node-009 · 8× H100 SXM5 80GB — CERTIFIED
    All 10 phases passed · Serial: HGX-001-A74F · Cert ID: CERT-2026-0628-009
    Jun 27, 2026
    3h 38m total
    node-010 · 8× H200 SXM5 141GB — CERTIFIED
    All 10 phases passed · Serial: HGX-002-B91C · Cert ID: CERT-2026-0628-010
    Jun 27, 2026
    3h 51m total
    node-007 · 8× H100 SXM5 — FAILED — NVLink fault
    GPU 3↔5 NVLink BW measured 241 GB/s (threshold 385 GB/s). NVSwitch port fault suspected.
    Phase 5 · Jun 26, 2026 14:22
    DCGM reported elevated NVLink error counters on GPU 3 in phase 3 — soft signal caught early.
    Phase 3 · Jun 26, 2026 13:45