BlogEvaluation guide

How to Evaluate Break-Glass Procedures Without Breaking CUI Boundaries

An ISSM/CISO procurement checklist for evaluating emergency GPU-pod access without quietly expanding the CUI boundary.

Consider an evaluation scenario: an on-prem GPU pod stops accepting jobs, the identity service is unavailable, and a vendor engineer offers to open a remote shell. Someone has the emergency credential. Nobody can say whether that shell exposes CUI, who must approve it, or how access ends.

The outage is operational. The access decision is a boundary decision. Evaluate break-glass procedures before procurement approval, not while a support bridge is asking for a password.

Ask the supplier to disclose the emergency path as precisely as the normal admin path. Use Pacific's CMMC-focused infrastructure planning to frame the discussion around scope, responsibilities, and evidence—not a certification claim.

1. Define the emergency before defining the credential

Break-glass should name a bounded exception to normal access controls. Require a written trigger, an authorized objective, and a stopping condition for each emergency path.

  • Separate identity-provider failure, orchestration failure, management-network failure, and suspected compromise. They need not use the same access procedure.
  • Record affected assets: GPU hosts, scheduler, storage, management controllers, credential vault, and any support gateway.
  • Identify permitted actions and prohibited actions. Restoring a service is not blanket permission to inspect datasets or export diagnostics.
  • Name the incident owner and the authority that can stop the session if CUI exposure or attacker activity is suspected.

Procurement gate: reject a procedure whose entire trigger is “when support needs access.” Require a scenario-to-access-path matrix with owners.

2. Name who invokes, approves, and operates

Ask for a role matrix, not a shared emergency account name. Invocation, approval, credential release, operation, and review should each have an accountable owner.

  • Define eligible operators, eligibility checks, and alternates. Keep emergency privileges limited to people authorized for the systems and information involved.
  • Where dual control is required by the organization’s procedure, require a separate authorized approver before release or activation—not a second person copied afterward.
  • Test approval when the primary approver, identity provider, or normal ticket system is unavailable. Define an authenticated alternate channel.
  • Disclose any single-person exception, its narrow trigger, and its compensating safeguards. Post-event review is not equivalent to prior dual approval.

Retain attributable evidence of the operator, approver, reason, and authorized scope. A generic account must not erase the identity of the human using it.

3. Make elevation scoped and time-bound

A checked-out password with a calendar reminder is not demonstrated expiration. Ask which component enforces the time limit and what happens if that component fails.

  • Specify the target systems, privilege level, allowed connection source, and approved duration. Avoid pod-wide root access when a narrower recovery path works.
  • Demonstrate credential or token expiration and active-session termination separately. An expired credential may leave an established session running.
  • Require fresh approval for extensions, with the continuing reason and revised end time recorded.
  • For unavoidable standing recovery credentials, disclose storage, custody, retrieval evidence, post-use rotation, and controls that limit where they work.

Compare this exception with the baseline in the GPU-pod administrative access evaluation checklist. Break-glass should not become the easier route for routine maintenance.

4. Trace CUI exposure through the recovery task

Physical location does not settle information access. An on-prem pod can expose CUI through a console, mounted volume, application output, diagnostic bundle, or remote screen.

  • Map what the emergency role can read, not just what the operator intends to read. Include job directories, checkpoints, swap, dumps, and mounted storage.
  • Identify potentially sensitive GPU diagnostics, memory dumps, training artifacts, and model outputs. Determine their handling from content and applicable requirements.
  • Define approved destinations for screenshots, command output, recordings, and support bundles. Do not assume a vendor ticket portal is authorized for CUI.
  • Prefer synthetic-data demonstrations and minimal diagnostics. Treat redaction as a reviewed process, not a guarantee that a bundle contains no CUI.

If the procedure depends on the operator voluntarily avoiding files that remain readable, disclose that limitation. Administrative intent is not technical isolation.

5. Evaluate the vendor remote path separately

Vendor participation is a separate access decision, even during an approved emergency. Disclose every hop from the support engineer’s endpoint to the pod.

  • Document the gateway, relay services, endpoint controls, authentication, session owner, and destination systems. Include outbound support tunnels.
  • Distinguish verbal guidance, screen viewing, remote control, and direct shell access. Each can create a different exposure path.
  • Evaluate clipboard sharing, file transfer, local recording, and diagnostic upload. Disable unnecessary capabilities and verify the restriction.
  • Require customer-controlled activation and a demonstrated disconnect mechanism. Disclose whether the vendor can reconnect without new approval.

For the detailed path review, use the GPU-pod remote support boundary evaluation. If no authorized remote path exists, use an approved local recovery procedure or defer the action; urgency does not authorize a new destination for CUI.

6. Preserve evidence without creating another exposure

Logging must survive the failure being recovered from. Ask for evidence of authorization, credential release, connection, privileged actions, termination, and cleanup.

  • Correlate records to one incident identifier and attributable operator. Document clock synchronization and how reviewers handle gaps or drift.
  • Protect records from alteration by the emergency operator where feasible. Disclose local buffering and reconciliation when central logging is unavailable.
  • Choose metadata, command auditing, and session capture deliberately. Recordings can capture CUI or secrets and therefore need appropriate protection.
  • Specify record access, retention, storage location, and review ownership under applicable policy and obligations. Avoid arbitrary universal retention claims.

Ask what happens if both normal logging and fallback evidence collection fail. The procedure needs a preapproved decision rule, not an improvised silent session.

7. Prove restoration closes the exception

Restoration is more than a green scheduler. Require evidence that emergency privileges, temporary connectivity, and diagnostic artifacts have been addressed before closure.

  • Terminate sessions, revoke temporary access, rotate exposed or used recovery secrets as required, and verify that old credentials no longer work.
  • Remove temporary accounts, keys, tunnels, firewall exceptions, and support agents. Compare relevant configuration against the approved baseline.
  • Restore normal identity controls, monitoring, and log forwarding. Verify that the next emergency path remains usable after credential rotation.
  • Account for collected artifacts and approved copies. Investigate suspected exposure under the incident process and evaluate applicable reporting obligations.

Assign a reviewer other than the operator to reconcile approvals, actions, changes, and remaining exceptions. Keep unresolved items owned and dated.

8. Turn the checklist into an acceptance test

Run a controlled demonstration with synthetic data before accepting the procedure. Do not create a production outage or expose live CUI merely to prove recoverability.

  • Exercise unavailable normal identity, an absent primary approver, interrupted central logging, and a vendor request to export diagnostics.
  • Observe approval, scoped activation, denied out-of-scope access, expiration, disconnection, credential invalidation, and baseline restoration.
  • Retain the approved runbook, role matrix, access-path diagram, test evidence, disclosed limitations, and remediation owners.
  • Write procurement findings as demonstrated, not demonstrated, or accepted exception. Do not convert a successful drill into a CMMC certification assertion.

FAQ

Does break-glass automatically violate a CUI boundary?

No. Evaluate whether the emergency access, people, systems, and information flows remain within the authorized scope and applicable requirements. A documented exception is not permission to bypass those requirements.

Can vendor support join if no files are transferred?

Possibly, but no file transfer does not mean no CUI exposure. Screen viewing, shell output, and recordings may reveal information. Approve the actual access path and participant authorization before enabling the session.

Bring the runbook, access-path diagram, and unresolved exceptions to a 30-minute GPU-pod boundary evaluation discussion. Start with what the emergency operator can actually reach, then decide what procurement still needs demonstrated.

Continue on the mothership

This satellite stops at the playbook. Transactions, specs, and comparisons live on pacificmachines.com. If the next step is a human, book 30 minutes with Harper.

Book 30 min