BlogEvaluation guide
How to Evaluate Firmware and Patch Ownership on an On-Prem GPU Pod
Before production data lands, know who can change BIOS, BMC, GPU firmware, and drivers—and who signs the risk when a patch breaks a locked stack.
The ISSM asked who authorizes a Sev-1 firmware flash after hours. Operations said the vendor would do it remotely. The vendor said the customer change board owned the window. Neither party owned the freeze. The locked training stack sat on a BIOS revision nobody could name.
That is a firmware- and patch-ownership failure, not a remote-support-boundary review and not a key-management review. Those evaluations exist elsewhere. This one is narrower: who may change BIOS, BMC, GPU firmware, NIC firmware, switch firmware, host drivers, and the container runtime—and who signs the risk when a required security patch conflicts with a locked workload.
Inventory every mutable plane
Do not begin with a vendor slide that says “we keep firmware current.” Begin with a plane inventory. For each mutable surface, record the current revision, who can change it, from where, with which identity, whether the customer must be present, how the change is logged, and how it is rolled back.
- System BIOS and related platform firmware.
- BMC and other out-of-band management firmware.
- GPU firmware and VBIOS on each module.
- NIC firmware and option ROM.
- Switch firmware on the fabric the pod depends on.
- Host drivers, including GPU, NIC, and storage.
- Container runtime and the node agent that launches jobs.
A change to any of those planes can take a locked ML stack offline. Treating “firmware” as one checkbox hides the plane that will actually break the job. GPU firmware is not BMC firmware. A host-driver bump is not a container-runtime bump. If one vendor ticket can touch three planes, write that as three approvals or as one expressly combined change.
Name who can change each plane in production: customer operations, customer security, the integrator, the OEM, or an automated pipeline. If two organizations believe they are the only writer, you do not have ownership. You have a race.
Separate the freeze from the emergency patch path
A freeze is a written rule that production revisions do not move except through a named process. An emergency patch path is a break-glass exception for a Sev-1 that cannot wait for the next change board. They are not the same document, and one does not silently cancel the other.
The change board should own planned moves: security bulletins that can wait, driver qualifications, BIOS updates that need a maintenance window, and any change that requires re-running the locked-stack suite. The emergency path should name the severity, the customer approver who can be reached after hours, the planes that may be touched, the time limit, the individual identity used, and the rollback that must be proven before the session closes.
Walk a Sev-1 firmware flash before production data lands. If operations can only describe “the vendor will flash it,” and the vendor can only describe “your change board must approve it,” nobody can restore the pod at 02:00. Decide in advance what happens if rollback images are missing, if the BMC becomes unreachable mid-flash, and if session recording or change-ticket systems are down. Those decisions belong in the runbook, not in the first outage.
Rollback evidence is part of ownership. A flash without a known-good image, a recorded pre-change revision, and a verification command is not a controlled change. Prefer customer-held golden images for each frozen plane. A vendor “we can re-image” promise is not the same artifact.
Qualify patches against the locked ML stack
A locked workload is a commercial and technical baseline: model family, framework versions, compiler or CUDA stack, driver range, and the acceptance suite that proved the reserved pod. A security bulletin that requires a driver or GPU-firmware move is a conflict until that suite passes again—not an automatic override of the freeze.
Build a qualification matrix. Rows are the mutable planes. Columns are the locked-stack constraints, the bulletin or CVE that wants to move the plane, the test that must pass, the owner of the test, and the decision: apply, defer with a written risk acceptance, or apply on a canary node first. A verbal “it should be fine” is not a cell in that matrix.
Evaluate the matrix alongside on-prem GPU infrastructure for CUI and ITAR and the mothership comparisons index. Pacific Intelligent Technologies, Inc. does not make a customer CMMC certified. A firmware-ownership map is evidence of a change boundary, not a certification claim.
If the program cannot accept an untested flash, write the deferral: who accepts residual risk, for how long, and which compensating controls apply. If the program cannot accept an unpatched bulletin, write the emergency qualification: a shorter suite that is still enough to declare the locked stack usable. Silence between those two positions is how a Friday CVE becomes an unauthorized weekend flash.
Pack the evidence an auditor can actually read
Auditors and ISSMs will not reconstruct ownership from Slack. Pack the artifacts before the first production job: the plane inventory with revisions, the freeze revision number, the change-board charter, the after-hours break-glass workflow, the qualification matrix, rollback images and their checksums, and a sample change record from a tabletop flash.
The sample record should show who requested the change, who approved it, which identity performed it, which planes moved, which tests passed, and how rollback would have been executed. If you cannot produce that pack from an exercise, you will not produce it from an incident.
Teams evaluating reserved GPU capacity should treat firmware ownership as part of the capacity boundary. Compute that can be scheduled is not the same as a stack that can be silently rewritten by a vendor bulletin.
If you are evaluating an on-prem GPU pod and want to walk the mutable-plane map with Pacific, start from Pacific Intelligent Technologies, Inc. and book 30 minutes with Harper. Bring the current revision list, the freeze rule, and one locked workload. Do not put controlled details in a website form.
FAQ
Does the vendor automatically own firmware after handoff?
No. Ownership is a written assignment per plane. The integrator may hold BMC and BIOS, the customer may hold host drivers, and the OEM may hold GPU firmware—or the customer may hold all of them. If the contract is silent, assume conflict and write the map before production.
Should every security bulletin break the freeze?
No. A bulletin is an input to the qualification matrix, not an override. Apply, defer with signed risk, or canary. An untested flash that restores a CVE score and breaks the locked stack is still an outage.
Is remote vendor flashing the same as patch ownership?
No. Remote access is a path. Ownership is who may authorize the change and who accepts the residual risk. A vendor who can reach the BMC without a customer approver has access, not a freeze. Evaluate the path separately from this ownership review.
Does this evaluation make an organisation CMMC certified?
No. A documented firmware and patch boundary can help you operate inside a declared environment, but certification depends on the customer’s systems, policies, people, scope, evidence, and assessment. Pacific Intelligent Technologies, Inc. provides infrastructure, including on-prem GPU infrastructure for CUI and ITAR, not certification itself.
Where does this sit relative to capacity and architecture choice?
Patch ownership is a change-boundary question. Capacity availability is a commercial question on Pacific's GPU capacity page. Architecture comparisons live on the comparisons index. Start from Pacific Intelligent Technologies, Inc. if you need both in one conversation, then book 30 minutes with Harper.
The evaluation rule is straightforward: inventory every mutable plane, separate freeze from break-glass, qualify patches against the locked stack, and pack evidence an auditor can read. A vendor flash that nobody can authorize is an open change boundary.
Continue on the mothership
This satellite stops at the playbook. Transactions, specs, and comparisons live on pacific.space. If the next step is a human, book 30 minutes with Harper.