SF Compute
v1Effective September 23, 2026

This is a frozen, versioned document. Contracts cite this exact version and it will not change. Newer requirements are published at a new version URL; this one remains available indefinitely.

This document describes the administrative access SF Compute asks for in order to provision, operate, monitor, and support a cluster running on a provider's infrastructure. It expands on Cluster Requirements §5 and is scoped entirely to the contracted cluster.

This is a starting point, not a checklist. Every deployment differs. Our platform has the flexibility to work in many scenarios. Please reach out to the SF Compute team to discuss integration options.

Each item is tagged with a Tier:

  • Core: automation depends on it; without it the cluster can't be operated the way it's sold
  • Flexible: we ask for it, but alternatives are common and usually workable

The reason we ask for as much as we do: SF Compute resells a node in minutes rather than months, and that depends on provisioning, isolation, and recovery running without a human in the loop.

DomainAccess requestedPrimary purpose
Out-of-band / BMCAdmin (read/write) on server and DPU BMCs — Redfish, IPMI, SSH, web, Serial-over-LANPower control, health monitoring, remote recovery
Power / PDUProgrammatic per-outlet controlDeep power-cycling for troubleshooting and firmware ops
DPU / BlueField OSFull admin on the running DPU OSOverlay networking, bare-metal provisioning, tenant isolation
ProvisioningDHCP / PXE / TFTP on cluster networks, or control of the existing servicesRe-image nodes between tenants without a provider ticket
Firmware & BIOSOut-of-band flashing, bulk update tooling, per-node BIOS export, version auditStandardization, performance tuning, CVE posture
InfiniBand fabricAdmin on UFM and its hosts; subnet-manager virtualizationAutomated pkey partitioning; RDMA tenant isolation
Ethernet / IP fabricAdmin on switches and firewall; VLAN / port / segmentation controlRe-segmentation as nodes move between tenants
Shared storageAdmin on the storage system and its hostsDynamic quota and partition assignment between tenants
Management nodesFull admin and out-of-band access on the CPU management nodesRuns the SF Compute control plane
Remote hands24/7 on-site support tickets, 30-minute initial responseHardware troubleshooting, repair, replacement

RequirementDetailTier
BMC administrator accountsAdministrator-level accounts on all server and DPU BMCs, across management interfaces in use: Redfish, IPMI, SSH, web UI, and Serial-over-LAN where supportedFlexible
Dedicated out-of-band networkAll management interfaces (BMC, PDU, switches, firewall) reachable on a dedicated OOB network, isolated from production and customer trafficCore
Disk eraseThe ability to erase every disk present on a managed systemCore
Account managementThe ability to create and manage administration and monitoring accounts on managed devicesFlexible

BMC access lets us control power, manage system configuration, and apply firmware updates. It also gives us visibility into GPU errors, thermal events, hardware health, and failed boots across both bare-metal and VM tenants.

This access also gives us the ability to act as first-line support for the end user. We triage and close most customer reports without involving the provider at all. Without direct console and power access, an unreachable host or an unresponsive BMC requires human triage. With it, we diagnose, drain, and recover the node within the platform.

If provider policy prevents issuing us BMC accounts directly, a jump host with equivalent Redfish/IPMI reach is usually workable.

RequirementDetailTier
Per-outlet power controlAdministrative access to the rack PDUs, or equivalent programmatic control at the outlet levelFlexible
Power topology mapA cable map describing each outlet-to-PSU connectionFlexible

Some troubleshooting and firmware operations need a true power cut at the source. Outlet-level control also gives us per-cluster power consumption telemetry.

Both items are marked Flexible. If the PDUs are shared infrastructure or the vendor has no usable API, power requests can route through a ticket system instead. It's slower, and it puts the work on the provider's team rather than ours, but it isn't a blocker.

RequirementDetailTier
DPU OS administrationFull administrative access to the BlueField DPU operating system on all DPUs, where BlueField DPUs are part of the buildCore
Privileged modeDPUs configured in privileged mode at handoff, with the ability for SF Compute to reconfigure them in restricted modeCore
Firmware controlThe ability to update DPU firmwareCore
On-board storageThe ability to erase or reflash all on-board DPU storage devicesCore
DPU-level RedfishRedfish access at the DPU levelFlexible

Access to BlueField DPUs allows us to configure networking external to the running host and its operating system. This allows configuration of multi-tenant isolation. It underpins bare-metal provisioning, letting us provision and attest a node without trusting host state the tenant controls. Having full control of the DPU also allows us to manage firmware and custom operating systems on the DPU at scale.

RequirementDetailTier
DHCP / PXE / TFTPEither the ability to run these services on the cluster's in-band and out-of-band networks, or administrative control of the existing services, or DHCP forwarding to oursCore
Per-host PXE boot controlThe ability to re-image any node without data-center involvementCore
SF Compute PXE serverRunning our own PXE server on the cluster networksFlexible

Our cluster management system uses an embedded DHCP and PXE service to provision nodes. We need the ability to run our platform's DHCP/PXE stack in parallel with any existing tooling. We will work with the provider to ensure proper isolation between our software and any existing automation.

RequirementDetailTier
BIOS configurationRead/write BIOS configuration on all GPU and management nodes via the out-of-band Redfish APICore
Out-of-band flashingOOB flashing of BMC, baseboard, and NVIDIA-card firmwareCore
Bulk update toolingVendor-published firmware bundles and a bulk BIOS/firmware update tool. Per-node GUI flashing and USB-stick flashing do not satisfy thisCore
Per-node BIOS exportFull per-node BIOS export (Redfish, or on-node IPMI/Redfish), so configuration can be standardized across the cluster and restored to its original stateCore
Version auditThe ability to audit current firmware and BIOS versions across the fleetCore

We standardize BIOS and firmware versions and settings across a cluster against known-good baselines, which is how hard-to-trace performance bugs and single-node anomalies get prevented rather than chased. Version auditing lets us check the fleet against known CVEs and confirm BIOS security posture, which we attest to customers as part of our compliance program.

In return: we notify before persistent BIOS changes, request approval before flashing firmware or BIOS, and hand back an export of the original BIOS settings so the cluster can be restored to its delivered state.

RequirementDetailTier
UFM administrationAdministrative access to UFM and the host(s) it runs on, for partition-key (pkey) and fabric configuration managementCore
Subnet-manager virtualizationib sm virt enable where applicable, per Cluster Requirements §9Core
Redundant SM hostsTwo subnet-manager hosts, each with root and out-of-band console accessCore
SM host L2 adjacencyBoth SM hosts' management NICs on the same L2 segment, with bidirectional TCP and UDP permitted between themCore
SM host diversityEach SM host's HCA cabled to a different leaf switch; where the build allows, the two hosts in separate racksFlexible

Fabric partitioning is a hard security boundary. Pkey management is what guarantees that no tenant can reach another tenant's GPU memory or traffic. Because tenant assignments change with market activity, this partitioning has to be automated.

RequirementDetailTier
VLAN / port / segmentationThe ability to configure VLANs, ports, and segmentation on the management and data networksCore
Switch administrationAdministrative access to the cluster's Ethernet switches, via out-of-band and/or management portsCore
Firewall administrationAdministrative access to the cluster firewallFlexible

Network configuration changes as nodes are assigned, released, and moved between tenants, and it changes on the same timescale the market moves. Direct control keeps tenant segmentation correct through those transitions, supports fabric-health monitoring, and is what makes a node resellable rather than statically assigned to one customer for the life of the contract.

Firewall access is marked Flexible: where the firewall is shared with other provider tenants, a documented change process with a committed turnaround is workable.

RequirementDetailTier
Shared storage platformA shared filesystem (VAST, WEKA, or equivalent — see Cluster Requirements §4) with fast network access from every nodeCore
Administrative accessAdmin access to the storage system and the server(s) running itCore
Dedicated storage networkPreferred, though fast in-band Ethernet is acceptableFlexible

We need the ability to manage the shared storage infrastructure the provider has in place. We need to automate storage provisioning and teardown between tenant transitions. It also lets us tune for performance and debug storage issues directly rather than relaying symptoms.

RequirementDetailTier
Management nodesAdministrative and out-of-band access to the CPU management nodes. Count and specification per Cluster Requirements §3: 6 nodes, or 9 for clusters over 128 GPU nodesCore
Out-of-band parityOOB access on the same management network, at the same level of access as the compute nodesCore
In-band reachabilityIn-band connectivity permitted between management nodes and compute nodes during provisioningCore

These are the nodes the SF Compute control plane runs on — provisioning, deprovisioning, monitoring, and administration of every compute node in the cluster, plus any additional managed services. They are operated exactly like the compute nodes they manage, which is why the access level needs to match rather than being a lesser tier.

RequirementDetailTier
Remote hands24/7/365 on-site support, per Cluster Requirements §8Core
Initial response time30 minutes or betterCore
Request channelAn agreed method for opening a request, ideally programmaticFlexible

We need a way to get a technician in front of the hardware quickly, and a response time we can commit to with our customers. Our customer SLAs are aligned with the provider to ensure the least amount of downtime and the best user experience possible.

The access above is paired with commitments running the other direction, consistent with Cluster Requirements §5.

CommitmentDetail
Dedicated network and credentialsAll management access rides a dedicated out-of-band network, separate from production and customer traffic. We recommend placing SF Compute on a separate management VLAN with dedicated credentials, and we restrict our own access to an agreed set of source IPs. How we reach the management network is decided per provider
Change controlWe notify the provider and request approval before flashing new firmware or BIOS, and we provide an export of the original BIOS settings for restore
Scoped useAll access is limited to the cluster under contract, and used solely for provisioning, monitoring, support, security verification, and performance management
Documented deltasWhere a requirement above is adjusted, waived, or satisfied a different way, it is written down in the deployment's access schedule rather than left as a verbal understanding

Full terms are in Cluster Requirements §5 (Access & permissions) and §9 (Machine state & provisioning configuration).

Point Mugu, CA