This is a frozen, versioned document. Contracts cite this exact version and it will not change. Newer requirements are published at a new version URL; this one remains available indefinitely.
This document describes the administrative access SF Compute asks for in order to provision, operate, monitor, and support a cluster running on a provider's infrastructure. It expands on Cluster Requirements §5 and is scoped entirely to the contracted cluster.
This is a starting point, not a checklist. Every deployment differs. Our platform has the flexibility to work in many scenarios. Please reach out to the SF Compute team to discuss integration options.
Each item is tagged with a Tier:
- Core: automation depends on it; without it the cluster can't be operated the way it's sold
- Flexible: we ask for it, but alternatives are common and usually workable
The reason we ask for as much as we do: SF Compute resells a node in minutes rather than months, and that depends on provisioning, isolation, and recovery running without a human in the loop.
| Domain | Access requested | Primary purpose |
|---|---|---|
| Out-of-band / BMC | Admin (read/write) on server and DPU BMCs — Redfish, IPMI, SSH, web, Serial-over-LAN | Power control, health monitoring, remote recovery |
| Power / PDU | Programmatic per-outlet control | Deep power-cycling for troubleshooting and firmware ops |
| DPU / BlueField OS | Full admin on the running DPU OS | Overlay networking, bare-metal provisioning, tenant isolation |
| Provisioning | DHCP / PXE / TFTP on cluster networks, or control of the existing services | Re-image nodes between tenants without a provider ticket |
| Firmware & BIOS | Out-of-band flashing, bulk update tooling, per-node BIOS export, version audit | Standardization, performance tuning, CVE posture |
| InfiniBand fabric | Admin on UFM and its hosts; subnet-manager virtualization | Automated pkey partitioning; RDMA tenant isolation |
| Ethernet / IP fabric | Admin on switches and firewall; VLAN / port / segmentation control | Re-segmentation as nodes move between tenants |
| Shared storage | Admin on the storage system and its hosts | Dynamic quota and partition assignment between tenants |
| Management nodes | Full admin and out-of-band access on the CPU management nodes | Runs the SF Compute control plane |
| Remote hands | 24/7 on-site support tickets, 30-minute initial response | Hardware troubleshooting, repair, replacement |
| Requirement | Detail | Tier |
|---|---|---|
| BMC administrator accounts | Administrator-level accounts on all server and DPU BMCs, across management interfaces in use: Redfish, IPMI, SSH, web UI, and Serial-over-LAN where supported | Flexible |
| Dedicated out-of-band network | All management interfaces (BMC, PDU, switches, firewall) reachable on a dedicated OOB network, isolated from production and customer traffic | Core |
| Disk erase | The ability to erase every disk present on a managed system | Core |
| Account management | The ability to create and manage administration and monitoring accounts on managed devices | Flexible |
BMC access lets us control power, manage system configuration, and apply firmware updates. It also gives us visibility into GPU errors, thermal events, hardware health, and failed boots across both bare-metal and VM tenants.
This access also gives us the ability to act as first-line support for the end user. We triage and close most customer reports without involving the provider at all. Without direct console and power access, an unreachable host or an unresponsive BMC requires human triage. With it, we diagnose, drain, and recover the node within the platform.
If provider policy prevents issuing us BMC accounts directly, a jump host with equivalent Redfish/IPMI reach is usually workable.
| Requirement | Detail | Tier |
|---|---|---|
| Per-outlet power control | Administrative access to the rack PDUs, or equivalent programmatic control at the outlet level | Flexible |
| Power topology map | A cable map describing each outlet-to-PSU connection | Flexible |
Some troubleshooting and firmware operations need a true power cut at the source. Outlet-level control also gives us per-cluster power consumption telemetry.
Both items are marked Flexible. If the PDUs are shared infrastructure or the vendor has no usable API, power requests can route through a ticket system instead. It's slower, and it puts the work on the provider's team rather than ours, but it isn't a blocker.
| Requirement | Detail | Tier |
|---|---|---|
| DPU OS administration | Full administrative access to the BlueField DPU operating system on all DPUs, where BlueField DPUs are part of the build | Core |
| Privileged mode | DPUs configured in privileged mode at handoff, with the ability for SF Compute to reconfigure them in restricted mode | Core |
| Firmware control | The ability to update DPU firmware | Core |
| On-board storage | The ability to erase or reflash all on-board DPU storage devices | Core |
| DPU-level Redfish | Redfish access at the DPU level | Flexible |
Access to BlueField DPUs allows us to configure networking external to the running host and its operating system. This allows configuration of multi-tenant isolation. It underpins bare-metal provisioning, letting us provision and attest a node without trusting host state the tenant controls. Having full control of the DPU also allows us to manage firmware and custom operating systems on the DPU at scale.
| Requirement | Detail | Tier |
|---|---|---|
| DHCP / PXE / TFTP | Either the ability to run these services on the cluster's in-band and out-of-band networks, or administrative control of the existing services, or DHCP forwarding to ours | Core |
| Per-host PXE boot control | The ability to re-image any node without data-center involvement | Core |
| SF Compute PXE server | Running our own PXE server on the cluster networks | Flexible |
Our cluster management system uses an embedded DHCP and PXE service to provision nodes. We need the ability to run our platform's DHCP/PXE stack in parallel with any existing tooling. We will work with the provider to ensure proper isolation between our software and any existing automation.
| Requirement | Detail | Tier |
|---|---|---|
| BIOS configuration | Read/write BIOS configuration on all GPU and management nodes via the out-of-band Redfish API | Core |
| Out-of-band flashing | OOB flashing of BMC, baseboard, and NVIDIA-card firmware | Core |
| Bulk update tooling | Vendor-published firmware bundles and a bulk BIOS/firmware update tool. Per-node GUI flashing and USB-stick flashing do not satisfy this | Core |
| Per-node BIOS export | Full per-node BIOS export (Redfish, or on-node IPMI/Redfish), so configuration can be standardized across the cluster and restored to its original state | Core |
| Version audit | The ability to audit current firmware and BIOS versions across the fleet | Core |
We standardize BIOS and firmware versions and settings across a cluster against known-good baselines, which is how hard-to-trace performance bugs and single-node anomalies get prevented rather than chased. Version auditing lets us check the fleet against known CVEs and confirm BIOS security posture, which we attest to customers as part of our compliance program.
In return: we notify before persistent BIOS changes, request approval before flashing firmware or BIOS, and hand back an export of the original BIOS settings so the cluster can be restored to its delivered state.
| Requirement | Detail | Tier |
|---|---|---|
| UFM administration | Administrative access to UFM and the host(s) it runs on, for partition-key (pkey) and fabric configuration management | Core |
| Subnet-manager virtualization | ib sm virt enable where applicable, per Cluster Requirements §9 | Core |
| Redundant SM hosts | Two subnet-manager hosts, each with root and out-of-band console access | Core |
| SM host L2 adjacency | Both SM hosts' management NICs on the same L2 segment, with bidirectional TCP and UDP permitted between them | Core |
| SM host diversity | Each SM host's HCA cabled to a different leaf switch; where the build allows, the two hosts in separate racks | Flexible |
Fabric partitioning is a hard security boundary. Pkey management is what guarantees that no tenant can reach another tenant's GPU memory or traffic. Because tenant assignments change with market activity, this partitioning has to be automated.
| Requirement | Detail | Tier |
|---|---|---|
| VLAN / port / segmentation | The ability to configure VLANs, ports, and segmentation on the management and data networks | Core |
| Switch administration | Administrative access to the cluster's Ethernet switches, via out-of-band and/or management ports | Core |
| Firewall administration | Administrative access to the cluster firewall | Flexible |
Network configuration changes as nodes are assigned, released, and moved between tenants, and it changes on the same timescale the market moves. Direct control keeps tenant segmentation correct through those transitions, supports fabric-health monitoring, and is what makes a node resellable rather than statically assigned to one customer for the life of the contract.
Firewall access is marked Flexible: where the firewall is shared with other provider tenants, a documented change process with a committed turnaround is workable.
| Requirement | Detail | Tier |
|---|---|---|
| Shared storage platform | A shared filesystem (VAST, WEKA, or equivalent — see Cluster Requirements §4) with fast network access from every node | Core |
| Administrative access | Admin access to the storage system and the server(s) running it | Core |
| Dedicated storage network | Preferred, though fast in-band Ethernet is acceptable | Flexible |
We need the ability to manage the shared storage infrastructure the provider has in place. We need to automate storage provisioning and teardown between tenant transitions. It also lets us tune for performance and debug storage issues directly rather than relaying symptoms.
| Requirement | Detail | Tier |
|---|---|---|
| Management nodes | Administrative and out-of-band access to the CPU management nodes. Count and specification per Cluster Requirements §3: 6 nodes, or 9 for clusters over 128 GPU nodes | Core |
| Out-of-band parity | OOB access on the same management network, at the same level of access as the compute nodes | Core |
| In-band reachability | In-band connectivity permitted between management nodes and compute nodes during provisioning | Core |
These are the nodes the SF Compute control plane runs on — provisioning, deprovisioning, monitoring, and administration of every compute node in the cluster, plus any additional managed services. They are operated exactly like the compute nodes they manage, which is why the access level needs to match rather than being a lesser tier.
| Requirement | Detail | Tier |
|---|---|---|
| Remote hands | 24/7/365 on-site support, per Cluster Requirements §8 | Core |
| Initial response time | 30 minutes or better | Core |
| Request channel | An agreed method for opening a request, ideally programmatic | Flexible |
We need a way to get a technician in front of the hardware quickly, and a response time we can commit to with our customers. Our customer SLAs are aligned with the provider to ensure the least amount of downtime and the best user experience possible.
The access above is paired with commitments running the other direction, consistent with Cluster Requirements §5.
| Commitment | Detail |
|---|---|
| Dedicated network and credentials | All management access rides a dedicated out-of-band network, separate from production and customer traffic. We recommend placing SF Compute on a separate management VLAN with dedicated credentials, and we restrict our own access to an agreed set of source IPs. How we reach the management network is decided per provider |
| Change control | We notify the provider and request approval before flashing new firmware or BIOS, and we provide an export of the original BIOS settings for restore |
| Scoped use | All access is limited to the cluster under contract, and used solely for provisioning, monitoring, support, security verification, and performance management |
| Documented deltas | Where a requirement above is adjusted, waived, or satisfied a different way, it is written down in the deployment's access schedule rather than left as a verbal understanding |
Full terms are in Cluster Requirements §5 (Access & permissions) and §9 (Machine state & provisioning configuration).
