Nodes in queue
12
3 added today
Running now
4
avg 2h 14m elapsed
Passed (30d)
89
92% pass rate
Failed (30d)
8
5 NVLink · 2 HBM · 1 PCIe
Avg cert time
3h 42m
per 8-GPU node
Active nodes — select to inspect
Phase progress
Overall progress62%
Live metrics
GPU thermal map — node-001
<75°C — nominal
75–82°C — warm
>82°C — flag
Event log
9 test categories · 32 individual tests · Together AI + Crusoe methodology
18
tests passed
6
tests running
8
tests pending
1
Stage 1 — Firmware & Config Baseline
Environment must be fully configured before any hardware test runs
4 pass~12 min
System preparation
Drivers · firmware · runtime stack · cluster config
4 pass
~12 min
2
Stage 2 — Hardware Health Diagnostics
Validate individual GPU health, storage, networking baseline, and observability
7 pass11 pending~50 min
GPU health & diagnostics
Enumeration · DCGM L1–L4 · ECC · XID · PCIe · inforom
7 pass
~20 min
Ethernet & DNS validation
iperf3 frontend throughput · NIC link state · DNS health checks
3 pending
~10 min
Storage benchmarks
fio · seq read/write BW · random IOPS · P99 latency · NVMe + block
5 pending
~20 min
Observability & continuous monitoring
Telegraf · cluster-level dashboards · host-level GPU health · 24×7 alerting
3 pending
ongoing
3
Stage 3 — Interconnect & Fabric Performance
Validate intra-node NVLink and inter-node InfiniBand / RoCE at full throughput
3 running5 pending~65 min
NVLink & NVSwitch validation
nvbandwidth GPU↔GPU matrix · NCCL intra-node · NVSwitch fault isolation
3 running
~25 min
InfiniBand / RoCE fabric
ibping · ib_read/write_bw · GPUDirect RDMA · NCCL multi-node · switch isolation
5 pending
~40 min
4
Stage 4 — AI Workload Benchmark
Prove production readiness under sustained AI training load — the final gate before certification
3 pass3 pending~95 min
GPU stress & burn-in
gpu-burn sustained · DCGM targeted stress · pulsed power · thermal soak
3 pass
~35 min
End-to-end workload validation — MFU
PyTorch FSDP · Llama-3 8B · TPS · MFU % · all_reduce latency · 16 nodes
3 pending
~60 min
Per-GPU DCGM results — node-001 · 8× H100 SXM5 80GB
GPU-level diagnostics
| GPU | ECC | PCIe | HBM BW | Power | Temp | XID errors | DCGM L4 |
|---|
Power draw (W)
HBM bandwidth (TB/s)
NVLink bandwidth matrix — nvbandwidth device_to_device_memcpy_write_ce (GB/s)
≥385 GB/s — pass
370–384 — marginal
<370 — fail
NCCL all_reduce_perf — intra-node results
Fabric, network & storage validation
InfiniBand / GPUDirect RDMA
Fabric speed400 Gb/s HDR
NCCL all_reduce BW372 GB/s (93%)
ib_write_bw390 GB/s
ib_read_bw388 GB/s
ibping latency1.2 µs
GPUDirect RDMAenabled
IB error rate0.00%
Multi-node NCCL (4 nodes)pass
Ethernet (iperf3) & DNS
Frontend NIC25 GbE
iperf3 TX24.7 Gbps
iperf3 RX24.6 Gbps
DNS — S3 / R20 errors
DNS — wandb.ai0 errors
All NIC link statesup
NVMe storage (fio)
Sequential read BW21.8 GB/s
Sequential write BW18.3 GB/s
Random read IOPS1.04M
Random write IOPS820K
P99 write latency2.8 ms
Checkpoint stalls0
MFU reference run — Llama-3 8B / FSDP
Tokens / secpending
Model flops utilizationpending
GPU utilizationpending
all_reduce latencypending
Phase 9 starts after NVLink validation completes
Completed certifications — cleared for production deployment