This is a frozen, versioned document. Contracts cite this exact version and it will not change. Newer requirements are published at a new version URL; this one remains available indefinitely.
These requirements define the minimum a provider GPU cluster must meet to engage with SF Compute.
Every requirement is tagged with a Tier:
- Core: a hard requirement the cluster must meet
- Nice-to-have: a miss may be accepted as a tradeoff
Unmet requirements are reviewed manually before acceptance. SF Compute's validation suite (§10) verifies acceptance. Contracts reference this document by version.
For H100 / H200 and B200 requirements, refer to Cluster Requirements v1 (effective July 7, 2026). This revision covers B300 and GB300.
| Component | Requirement | Tier |
|---|---|---|
| Chassis | NVIDIA HGX reference platform (AMD or Intel baseboard) or NVIDIA DGX. Validated OEM chassis: Dell PowerEdge XE9780(L), NVIDIA DGX, Supermicro HGX, Pegatron HGX, ASRock HGX (other OEMs / SKUs upon review) | Core |
| GPU | 8× NVIDIA B300 288GB SXM6 | Core |
| CPU | 2× AMD EPYC 9005 series or newer, or 2× Intel 6th gen Xeon (Granite Rapids) or newer | Core |
| System RAM | ≥ 3 TB ECC | Core |
The BOM must state the chassis manufacturer as well as the GPU SKU. Every chassis must support bulk (non-GUI, non-USB) BIOS/firmware updates (§9). Off-list chassis, including the "upon review" entries, are accepted only after manual validation (§10) and a firmware-manageability review (§9).
We support GB300 NVL72 from all OEMs that adhere to NVIDIA's reference design
| Component | Requirement | Tier |
|---|---|---|
| Unit | 72× NVIDIA B300 288GB in one NVLink domain. A compute tray ("node") is 4× B300 GPUs + 2× NVIDIA Grace CPUs | Core |
| Platform | Any OEM build that adheres to NVIDIA's reference design: 18× compute trays + 9× NVLink switch trays per rack | Core |
| Cooling | Direct liquid cooling. Facility requirements (density, CDUs, loop, leak detection) are in Colocation §1–2. | Core |
| GPU interconnect | 2× ConnectX-8 mezzanine boards per compute tray. Multi-rack builds need the east-west fabric in §2. | Core |
| In-band / OOB | Dual-port BlueField-3 in-band and 1Gbps BMC per tray | Core |
| Item | Requirement | Tier |
|---|---|---|
| Cooling | Air or liquid; no immersion or other out-of-spec cooling for DGX/HGX systems. Air vs. liquid is the supplier's choice. Facility cooling capacity and redundancy requirements are in Colocation §2 and must match the build. | Core |
| Rack density | Liquid cooling preferred where higher rack density is required; per-rack density by cooling type is in Colocation §1. | Nice-to-have |
| In-band NIC | Single dual-port 400 Gbps BlueField-3, with the two ports cabled to separate switches for switch-failure resilience | Core |
| GPU interconnect | 8× ConnectX-8 per node for B300; 2× ConnectX-8 mezzanine boards for GB300 | Core |
| OOB Mgmt NIC | 1 Gbps | Core |
| Boot storage | Mirrored boot drives: RAID-1 pair, ≥ 960 GB usable | Core |
| Ephemeral storage | Local NVMe scratch: 30 TB target, 20 TB minimum. BOM must indicate drive count, size, and RAID layout per node | Core |
| Operating system | Ubuntu 24.04 or 26.04 pre-installed; re-imageable by SF Compute (§9) | Core |
All NICs / HCAs must be SR-IOV-capable for VM passthrough (§9).
| Network | Requirement | Tier |
|---|---|---|
| In-band / north-south Ethernet | Converged network carries tenant access, data, monitoring, and network-attached storage; non-blocking; 400 Gbps per GPU server; 200 Gbps per CPU or storage server. Dual-switch redundancy; NVIDIA SN5610 switch preferred | Core |
| GPU east-west interconnect | ≥ 800 Gbps InfiniBand or RoCEv2, non-blocking. For InfiniBand: rail-optimized spine-leaf topology with redundant UFM servers | Core |
| OOB management | 1 Gbps; isolated from the in-band network | Core |
| Fabric topology visibility | Leaf / spine / core fabric topology is discoverable programmatically (e.g. via API), so SF Compute can map a tenant's GPUs to their fabric placement | Nice-to-have |
| Public IPs | Each machine must have a public IP or sufficient access to set up BGP to advertise routes | Core |
Network-attached storage runs over the in-band / north-south fabric; no separate storage network is required. Storage sizing is in §4.
ECN/PFC tuning for RoCE is agreed per deployment; it is not a listing requirement.
GB300 NVL72: the NVLink domain spans the rack. A single-rack deployment needs no external GPU interconnect fabric. Multi-rack deployments interconnect over the east-west fabric at 800 Gbps XDR
| Requirement | Detail | Tier |
|---|---|---|
| Management nodes | 6× CPU-only (2× AMD EPYC or Xeon, 128 GB RAM, 100 GB disk, public IP); 9× for clusters over 128 GPU nodes | Core |
These nodes run SF Compute's on-site control plane (k8s control plane + workers, UFM proxy, storage services); quorum services run on an odd subset. State the node count and per-node RAM.
State your NAS solution and its usable (not raw) capacity; high-performance filesystems (VAST/WEKA) fill to only ~80–90% before performance degrades; throughput is recorded against per-GPU reference tiers, not as an SLA.
| Requirement | Detail | Tier |
|---|---|---|
| Supported vendors | VAST, WEKA, Dell PowerScale/ObjectScale/Lightning | Core |
| Usable capacity per GPU | ≥ 4 TB usable per GPU (≈ 2 PB for 512 GPUs; ≈ 8 PB for 2,048 GPUs). Primary sizing metric | Core (information) |
| Capacity floor | ≥ 100 TB usable minimum for small clusters, NVMe-backed (shared-file protocols served on top) | Nice-to-have |
| Throughput per GPU | Reference tiers: read 0.5 / 1.0 / 5.0 GB/s, write 0.25 / 0.5 / 2.5 GB/s per GPU. Target the middle tier or better (read ≥ 1.0 GB/s, write ≥ 0.5 GB/s per GPU). Recorded, not an SLA | Nice-to-have |
| Resilience & growth | ≥ 20% of physical disks dedicated to parity; capacity expandable to 2× within 60 days without service interruption | Nice-to-have |
| Requirement | Detail | Tier |
|---|---|---|
| Bare-metal delivery | No provider-imposed hypervisor or control-panel layer. SF Compute runs its own virtualization on top of bare metal (§9) | Core |
| BMC access | On GPU, management, and DPU nodes | Core |
SF Compute provides first-line support for nodes sold on the Marketplace, which requires BMC access at scale. It is used to:
- View console output during boot, network outages, device-enumeration failures, or host OS boot failures
- Reboot nodes via IPMI without a support ticket
- Re-image the host OS to keep nodes current and reset between tenants
- Verify BIOS configuration is secure and firmware is current
- Tune BIOS parameters for performance
We recommend the provider place SF Compute on a separate management VLAN, issue separate credentials, put the BMC behind a VPN, and allowlist IPs on the management network. In return, SF Compute notifies before persistent BIOS changes, requests approval before flashing firmware or BIOS, and provides an export of the original BIOS settings for restore.
| Requirement | Detail | Tier |
|---|---|---|
| DIA / ISP connections | Two or more circuits with < 50 ms ping to AWS S3 or Cloudflare R2. Must have redundant routing across physically diverse paths with automatic failover and no single point of failure. Failover mechanism is the provider's choice; diverse paths and dual connections are required | Core |
| Internet uplink capacity | Aggregate external uplink ≥ the greater of (a) 2× 100 Gbps and (b) 125 Mbps per GPU (≈ 1 Gbit/s per 8-GPU node). This is oversubscribed internet-facing capacity, not a per-tenant guarantee. For clusters > 256 nodes the minimum can be negotiated | Core |
| Firewall | HA pair with automatic failover, supporting multiple 100 Gbps connections | Core |
| Ingress / egress | No restrictions or metering | Core |
Facility details (power density, cooling, certifications, physical security) are in Colocation Requirements.
| Requirement | Detail | Tier |
|---|---|---|
| Tier & power redundancy | Minimum Tier III; ≥ N+1 generator redundancy; on-site fuel under a continuous-replenishment / resupply contract (see Colocation §1 for generator and fuel detail) | Core |
| 2N redundancy | 2N power / cooling distribution where available | Nice-to-have |
| Power | Adequate provisioning confirmed at 100% load | Core |
| Cooling | Adequate capacity and ≥ N+1 redundancy confirmed at maximum thermal output, with airflow clearance where applicable (see Colocation §2) | Core |
| rPDUs | Redundant power to each rack | Core |
| Physical security | Measures preventing unauthorized entry into the DC hall and access to hardware | Core |
| Inspection | Physical or virtual inspection on request: neat cabling, hot/cold separation (blanking panels + baffles), clean space, reverse-airflow switches not co-located with forward-airflow switches | Core |
| Requirement | Detail | Tier |
|---|---|---|
| Remote hands | 24/7/365 | Core |
| Initial response time | 30 minutes or better | Core |
| OEM support contract | Active OEM hardware support / warranty, minimum 3-year next-business-day (NBD). No cold-spare chassis required; chassis go to OEM RMA or on-site OEM repair | Core |
| Spare-parts strategy | On-hand spares locker of 1–3% of high-failure-rate FRUs (PSUs, RAM, GPUs): ~1% at scale, up to ~3% for small clusters, scaled to cluster size and OEM self-service provisions | Core |
| SLA | Per contract | Core |
These settings let SF Compute provision, virtualize, and operate the cluster on bare metal. A node that misses them is unusable even if it passes every §10 performance test.
Exact BIOS values, version floors, and provisioning steps are validated internally and are not part of this spec.
| Requirement | Detail | Tier |
|---|---|---|
| BMC + power-cycle | BMC with power-cycle and console, via API or portal; remote reboot/console/BIOS/reimaging preferred | Core |
| BIOS virtualization | VT-x / AMD-SVM; x2APIC + interrupt remapping; IOMMU (VT-d); SR-IOV enabled globally and per InfiniBand adapter; "Above 4G decoding" + Resizable BAR; a kexec-capable BIOS. | Core |
| BIOS standardization | Validated against a known-good "golden BIOS"; full per-node export (Redfish or on-node IPMI/Redfish), standardized across machines | Core |
| Mellanox firmware | ATC enabled on each card | Core |
| Firmware management | Out-of-band flashing of BMC, baseboard, and NVIDIA-card firmware; vendor-published firmware bundles; a bulk BIOS/firmware update tool (no per-node GUI or USB-stick flashing) | Core |
| Bare-metal handoff | Vanilla Ubuntu 24.04 or 26.04 on bare metal with BMC access; SF Compute re-images nodes with its own provisioning system and OS/driver image | Core |
| Networking | Consistent, stable per-host IPs (static routed /31 or DHCP); per-VM IPv4 reachability | Core |
| Storage auth | API credentials where shared storage is offered | Core (when storage offered) |
| Tenant isolation (RDMA) | Hard isolation: no tenant can reach another tenant's GPU memory or traffic. SF Compute administers fabric partitioning; the supplier grants administrative control of the subnet manager / UFM and enables subnet-manager virtualization (ib sm virt enable) where applicable. | Core |
Suppliers must provide proof of a successful burn-in, or allow time for SF Compute to run one. SF Compute validates before delivery rather than relying on vendor attestations. Burn-in establishes a performance baseline and detects hardware faults (bad GPUs, marginal optics, weak DIMMs, miscabled rails) while remediation is still the supplier's responsibility.
Sev-1 criteria are Core: the node or cluster must pass them. Sev-2 criteria are scored (Nice-to-have); a miss requires a documented disposition.
Burn-in runs in two tiers. Continuous monitoring runs across both and can fail either tier: XID events, PCIe AER events, IB counter deltas, thermal/power throttling, node liveness, and, on liquid-cooled builds, coolant-loop integrity (manifold quick-connector joints checked under sustained load). Cluster-wide burn-in begins only after per-host burn-in is clean, and is skipped for single-node / VM instances (no shared fabric to validate).
| Stage | Scope | Duration |
|---|---|---|
| Per-host | CPU, RAM, local NVMe, individual GPUs | 16–24 h / node (24 h minimum for GPU soak) |
| Cluster-wide | NVLink, NCCL collectives, InfiniBand fabric, sustained soak | 24 h – 5 days (default 48 h; scaled to fleet confidence and the customer SLA) |
Tooling: DCGM, gpu-burn, stressapptest, stress-ng, memtester, fio.
| Criterion | Threshold | Severity |
|---|---|---|
dcgmi diag -r 4 on every GPU, pre- and post-soak | 0 failures or warnings | Sev-1 |
| Uncorrectable ECC errors (volatile + aggregate) | 0 | Sev-1 |
| Correctable ECC errors | Within threshold | Sev-2 |
GPU row-remap (nvidia-smi -q) | Clean, no pending/failed remaps | Sev-1 |
| Gating XID events (hardware-implicating) | 0 | Sev-1 |
| Thermal / power throttling under sustained GPU load | Not asserted; clocks/power sustained at TDP | Sev-1 |
| Host memory stress (stressapptest / memtester) | 0 errors | Sev-1 |
| CPU stress (stress-ng) | 0 errors; no MCEs; no thermal throttle | Sev-1 |
| Logical CPU count vs. spec (SMT/core config) | Matches PO/BOM (e.g. 128, not 64) | Sev-1 |
| Local NVMe (fio) | ≥ 100k read IOPS / ≥ 100k write IOPS per node (symmetric); throughput within SKU spec | Sev-2 |
Tooling: nccl-tests, ClusterKit / perftest (ib_write_bw, ib_read_bw). Soak: NCCL all-reduce loop (default) or Megatron-LM / HPL LINPACK, 24 h – 5 days (default 48 h).
| Criterion | Threshold | Severity |
|---|---|---|
| NVLink / NVSwitch via intra-node NCCL all-reduce | Within 5% of SKU reference | Sev-1 |
| Cluster-scale NCCL all-reduce busbw (ClusterKit) | ≥ 0.93 bandwidth tolerance (btol); latency within 2.1× (ltol) | Sev-1 |
| Run-to-run variance on cluster collectives | ≤ 3% | Sev-2 |
| All IB ports up at expected speed and width | 100% | Sev-1 |
| Per-port IB bandwidth (ClusterKit / perftest) | ≥ 95% of line rate | Sev-1 |
| Bisection bandwidth across cluster | ≥ 90% of theoretical | Sev-1 |
| IB link recoveries / flaps during burn-in | 0 | Sev-1 |
| NIC negotiated link speed (per node) | Asserted at expected line rate | Sev-1 |
| Soak: node uptime / GPU enumeration / IB stability | 100% / 0 dropouts / 0 flaps | Sev-1 |
| Soak: performance drift end vs. start | < 2% | Sev-2 |
GPU count and SKU, host RAM, NVMe, NIC count/speed/model, and firmware must match the PO and BOM (mismatches are Sev-1). The firmware baseline (BIOS, BMC, GPU VBIOS) is pinned and recorded per node as an acceptance artifact.
A node that passes all of the following is considered working. The NCCL and DCGM checks are identical across SKUs; NVLink and InfiniBand bandwidth targets vary by SKU (reference table below).
- Single-node NCCL, on every node:
NCCL_DEBUG=INFO python3 -m torch.distributed.run --standalone --nproc_per_node=8 nccl-test.py - Multi-node NCCL, concurrently across nodes (rendezvous on node 0):
NCCL_DEBUG=INFO python3 -m torch.distributed.run --nproc_per_node 8 --nnodes NUM_NODES \ --rdzv-backend c10d --rdzv-endpoint HOST-0:8888 nccl-test.py - NVLink P2P bandwidth: per-SKU target below, via
p2pBandwidthLatencyTest(P2P enabled). - InfiniBand inter-node bandwidth: per-SKU target below, via
ib_write_bw -d mlx5_0 --use_cuda 0 -a --report_gbits. - DCGM:
sudo dcgmi diag -r 4passes.
| SKU | NVLink P2P (uni / bidi) | InfiniBand per port (line rate / real) |
|---|---|---|
| H100 / H200 | ~375 / ~740 GB/s | 400 Gbit/s NDR / ~390 Gbit/s |
| B200 | ~750 / ~1,480 GB/s | 400 Gbit/s NDR / ~390 Gbit/s |
| B300 | ~750 / ~1,480 GB/s | 800 Gbit/s XDR / ~750 Gbit/s |
| GB300 NVL72 | TBD: NVL72-domain targets published after first-article validation | 800 Gbit/s XDR / ~750 Gbit/s |
Line rates are firm. Measured NVLink P2P and real IB throughput are fleet-confirmed references, not SLAs.
On completion the customer receives a per-node / per-GPU hardware inventory matched against the PO, a tooling manifest (image digests, driver/CUDA/OFED versions, pinned firmware baseline), and a signed acceptance certificate referencing this document by version. Acceptance requires that all Sev-1 (Core) criteria pass on all nodes, every Sev-2 finding has a documented disposition (remediated, waived with reason, or scheduled), and the customer countersigns. A Sev-1 failure triggers remediation and a re-run of the affected tier; remediation during burn-in is the supplier's responsibility.
| Requirement | Tier |
|---|---|
| Full bill of materials (BOM) | Core |
| Rack elevation diagrams | Core |
| Network architecture diagram | Core |
