Skip to content

When the Platform Can Say No Without Saying Why

No externally governed inference or action provider should constitute an unmitigated single point of failure. Where provider behavior cannot be sufficiently characterized, treat the provider as an untrusted-availability dependency.

Something very ordinary happened to me this week.

I was working with an AI platform that had access to a repository. We were writing architecture notes. Nothing exotic. Nothing that looked remotely like a dangerous operation. The system prepared a Markdown file, called the repository tool to create it, and the provider blocked the call.

The only useful explanation was essentially: safety checks prevented the operation.

  • No policy identifier.
  • No reason code that told me what class of control had fired.
  • No statement that the repository operation itself was prohibited.
  • No way to tell whether the trigger had been the text, the account state, a heuristic, some transient classifier result, accumulated context, a platform-wide condition, or something else entirely.

Later, another ordinary repository operation worked.

Then another another ordinary repository operation failed.

That is annoying when you are writing documents.

It becomes something else entirely when you put on your standards glasses.

The problem is not that a platform has safety controls.

Platforms should have controls.

The problem is that the control has become part of the execution path while remaining uncharacterized to the system designer who depends on that path.

If I cannot know what conditions will cause an operation to be refused, cannot determine why a refusal occurred, cannot reproduce the refusal reliably, cannot distinguish a deliberate policy denial from an availability failure, and cannot establish what guarantees exist around the behavior, then from an engineering perspective I do not have a dependable execution primitive.

I have an external dependency with an opaque failure mode.

That distinction matters.

Opaque provider interlocks may be legitimate controls. In a critical workflow, however, unexplained refusals become an uncharacterized failure mode, and the standards world already knows what to do with that.

Refusal Is Not the Failure

Security people are comfortable with refusal.

  • A firewall refuses traffic.
  • An access-control system refuses credentials.
  • A safety controller refuses an unsafe state transition.
  • A payment system refuses a transaction.

Those systems can still be reliable because the conditions around the refusal are part of the system contract.

You may not know every proprietary detection technique inside a fraud engine, but you can still know what the service promises, what outcomes it returns, which errors mean what, what gets logged, who owns the decision, how it is appealed, how it is monitored, and what happens when the service itself is unavailable.

The refusal is observable enough to engineer around.

That is not the same as requiring a vendor to publish every defensive rule. Security controls often need some secrecy. A detection provider does not have to explain its entire threat model to every caller.

But a system architect still needs a usable boundary around the service.

  • What are the possible outcomes?
  • What constitutes success?
  • What constitutes denial?
  • What constitutes transient failure?
  • What evidence survives?
  • What state may have changed before the denial?
  • Can the operation be retried safely?
  • What service-level expectations apply?
  • What happens to the larger process if this component refuses service?

If those questions cannot be answered, the safety control has also become an availability dependency.

Reliability Is Something You Can Characterize

Suppose a component fails one request in every thousand.

That may be perfectly acceptable.

If I know the failure rate, know the failure semantics, can detect the failure, can retry safely, can fail over, and can demonstrate recovery, I can build a system around it.

We have been doing this for decades.

  • Disks fail.
  • Networks fail.
  • Processes crash.
  • Databases deadlock.
  • Certificates expire.
  • Cloud regions go dark.

Engineering does not require components to be perfect. It requires failure to be understood well enough that the larger system can remain within its required operating envelope.

An opaque provider interlock is more difficult because the failure distribution itself may be hidden.

The same apparent class of operation may succeed and fail at different times. The caller may receive too little information to distinguish a policy decision from a service fault. A previously successful operation does not establish that the next similar operation will be permitted. A failed operation may provide no stable rule that a designer can encode into test cases.

At that point the only safe assumption is brutally simple:

This path may become unavailable at any time.

That is not a moral judgment about the provider.

It is a reliability classification.

The Standards Are Already Looking at This Problem

NIST's AI Risk Management Framework does not say, “commercial AI platforms must never have opaque interlocks.” Standards rarely speak in sentences that convenient.

What it does do is define trustworthy AI in terms that include being valid and reliable, safe, secure and resilient, accountable and transparent, and explainable and interpretable. The Framework treats reliability as foundational and emphasizes that these characteristics have to be considered in context of use. NIST Publications

That last phrase matters.

A system that is perfectly useful for drafting an article can be completely unsuitable as the sole execution path for a safety-critical control.

NIST's 2026 concept note for a forthcoming AI RMF profile for critical infrastructure goes even closer to the bone. It explicitly identifies stringent critical-infrastructure needs including deterministic behavior, explainability, graceful degradation, and fail-safe operation. NIST

That is the language of a world in which the consequences of hidden behavior matter.

Standards lens: AI trustworthiness

  • NIST AI Risk Management Framework 1.0 - frames trustworthy AI around valid and reliable behavior, safety, security and resilience, accountability and transparency, and explainability and interpretability. NIST Publications
  • NIST 2026 AI RMF Critical Infrastructure Profile concept note - explicitly calls out deterministic behavior, explainability, graceful degradation and fail-safe operation as critical-infrastructure needs. NIST
  • ISO/IEC 42001:2023 - establishes an AI management-system framework for managing AI risks and opportunities and identifies traceability, transparency and reliability among its benefits and objectives. ISO
  • ISO/IEC 23894:2023 - provides guidance for integrating AI-specific risk management into organizational activities. ISO

None of those documents says that a general-purpose AI platform is automatically prohibited from serious systems.

They do say something more useful: risk has to be managed as a property of the actual system in its actual context.

If an opaque provider decision can interrupt a required function, that interruption belongs in the risk model.

External Providers Do Not Make Your Responsibility Disappear

This is where the argument stops being specifically about AI.

NIST SP 800-53 has dealt with external system services for years.

Control SA-9 requires organizations to define expectations for external service providers, document roles and responsibilities, and monitor provider compliance. Its discussion is unusually relevant here: when an external provider implements part of your system, you do not directly control the provider's controls, but the responsibility for managing the resulting risk remains with the organization using the service.

It also points directly at service-level agreements: expectations, measurable outcomes, remedies and response requirements. NIST Publications

NIST SP 800-171 Rev. 3 carries the same principle into protection of Controlled Unclassified Information. Requirement 03.16.03 tells organizations to define external-service security requirements, document shared responsibilities, and monitor provider compliance on an ongoing basis. Its discussion explicitly says that service-level agreements define performance expectations, measurable outcomes, and remedies or mitigations for noncompliance. NIST Publications

That is a fairly devastating standards question for any opaque dependency:

What exactly are you measuring?

If a provider can interrupt a function through an internal decision process whose behavior you cannot sufficiently characterize, what availability objective do you assign to that function?

What failure rate do you test?

What error class do you alert on?

What recovery-time objective do you promise?

What is the provider actually obligated to restore?

What evidence can your auditor inspect?

Standards lens: external services and supplier assurance

  • NIST SP 800-53 Rev. 5, SA-9 - External System Services - requires documented oversight, provider requirements and ongoing monitoring; its discussion ties external services to service-level expectations and measurable outcomes. NIST Publications
  • NIST SP 800-171 Rev. 3, 03.16.03 - External System Services - requires defined provider requirements, documented shared responsibilities and ongoing monitoring of provider compliance. NIST Publications
  • ISO/IEC 27036-2:2022 - Cybersecurity - Supplier relationships - Part 2: Requirements - specifies requirements for defining, implementing, operating, monitoring, reviewing, maintaining and improving supplier/acquirer relationships, including cloud services. ISO
  • ISO/IEC 20000-1:2018 - Service management system requirements - covers planning, delivery, monitoring, measurement, review and improvement of managed services so that service requirements can be met. ISO

The common thread is not “vendors must tell you all their secrets.”

It is that a service relationship used in a controlled system must be governable.

The consuming organization needs enough visibility, evidence, commitments and compensating controls to make a defensible claim about the service it is delivering.

The Common-Cause Failure Nobody Sees

There is another trap here that matters enormously as AI platforms become operating surfaces instead of chat windows.

Imagine that the platform gives me five different tools.

  • GitHub.
  • Gmail.
  • Slack.
  • Google Drive.
  • A remote execution connector.

From the application layer, that looks like diversification.

Five tools. Five services. Five different downstream systems.

But if every one of those actions is mediated through the same provider-level policy and safety gate, the architecture actually looks like this:

GitHub --------\
Gmail ----------\
Slack ------------> provider control plane ---> action
Drive ----------/
Remote host ----/

Those are not five independent action paths.

They are five adapters hanging from one upstream decision point.

If that common decision point becomes unavailable, all five can disappear together. If its policy state changes, all five can change behavior together. If its account-level posture changes, all five can be affected together.

Moving from GitHub to Drive is not failover if the same opaque gate decides whether either call leaves the platform.

This is classic common-cause thinking.

Functional-safety engineering has spent decades warning us not to confuse apparent redundancy with independent redundancy. IEC 61508 explicitly treats avoidance and control of faults and failures as a core safety concern, and Part 6 includes methodology for quantifying hardware-related common-cause failures. IEC Webstore

The exact mathematics of IEC 61508 are not directly transferable to a cloud AI policy engine.

The architectural lesson absolutely is.

Two things are not redundant simply because they have different names.

They have to fail differently.

Standards lens: failure, recovery and independence

  • NIST SP 800-53 Rev. 5, CP-10 - System Recovery and Reconstitution - calls for recovery and reconstitution to a known state after disruption, compromise or failure. NIST Publications
  • NIST SP 800-53 Rev. 5, SI-13 - Predictable Failure Prevention - addresses failure characteristics, standby components and transfer of responsibilities without compromising safety, readiness or security. NIST Publications
  • IEC 61508 series - provides the general functional-safety lifecycle for electrical/electronic/programmable electronic safety-related systems; Parts 2 and 3 include measures for avoiding and controlling faults and failures. IEC Webstore
  • 0IEC 61508-6:2010 - includes methodology for assessing hardware-related common-cause failures. IEC Webstore
  • IEC 62443-3-3:2013 - provides system security requirements and security levels for industrial automation and control systems. IEC Webstore
  • ISO/IEC 25010:2023 - provides an ICT product-quality model intended to support requirements definition, testing, measurement, evaluation and acceptance criteria. ISO

This is why “we have several connectors” is not enough.

The question is: where is the shared failure domain?

Continuity Standards Make the Answer Boring

The good news is that none of this requires a revolutionary architecture.

Business-continuity standards give us a wonderfully boring answer.

ISO 22301 requires organizations to prepare for, respond to and recover from disruptions while continuing required products and services at an acceptable capacity. The current published edition is ISO 22301:2019, with a 2024 amendment; a successor edition is under development in 2026. ISO

ISO/TS 22318 applies those continuity principles directly to supplier relationships and supply chains. Its objective is to help organizations remain prepared for disruption in the resources and services they depend on. The 2021 edition was reviewed and confirmed in 2025. ISO

That means the correct response to an opaque provider gate is not outrage as an architecture.

It is continuity engineering.

Decide what function must survive.
Decide how much interruption is acceptable.
Decide what state has to persist outside the provider.
Decide whether there is an alternate path.
Decide whether a human can take over.
Decide how you know the alternate path still works.
Exercise it.
Record the result.

Standards lens: continuity

  • ISO 22301:2019, with Amendment 1:2024 - the current published business-continuity management standard; it provides a framework to plan, operate, monitor, maintain and improve resilience and recovery from disruptive incidents. A successor edition is under development in 2026. ISO
  • ISO/TS 22318:2021 - extends business-continuity principles into supplier and supply-chain relationships and remains current following confirmation in 2025. ISO
  • NIST SP 800-53 Rev. 5 contingency controls - include recovery and reconstitution requirements and related contingency mechanisms. NIST Publications

Once you describe the problem that way, the design rule becomes obvious:

an externally governed AI provider should not be an unmitigated single point of failure for a critical function.

This Does Not Mean “Route Around the Safety System”

There is an important trap on the other side.

Suppose Provider A refuses an operation. A terrible design would say:

  • Try Provider B.
  • If B refuses, try C.
  • Keep shopping until somebody says yes.

That is not resilience.

That is policy laundering.

The local system still has to know whether the requested action is authorized. The provider's refusal means only one thing with certainty:

that provider path did not perform the operation.

It may have refused for an excellent safety reason. It may have refused because of a transient classifier result. It may have refused because the service was degraded. It may have refused because the account was in a different internal state. It may have refused simply because the provider has terrible policies which you cannot even see.

If the platform does not expose enough information to distinguish those cases, then the consuming architecture must not pretend that it knows.

The right sequence is local governance first.

  • Is this action authorized under the organization's own policy?
  • If not, stop.
  • If it is authorized, attempt the provider path.
  • If the provider refuses, record the refusal.
  • Then continue through an alternate path only if that alternate path is independently authorized under the same governing policy.
  • Otherwise, stop and hand the operation to the responsible human.

That is a very different thing from bypassing a safety control.

It preserves the provider's right to refuse service without silently granting the provider authority over the entire external system.

Put the AI Around the Critical Path, Not in Place of It

This leads to a design pattern I expect we will see more often:

Use an external AI platform heavily around critical systems.

  • Let it analyze.
  • Let it correlate.
  • Let it summarize.
  • Let it detect anomalies.
  • Let it draft.
  • Let it recommend.
  • Let it explain.
  • Let it simulate.

But keep the authoritative state, safety envelope, recovery logic and critical actuation somewhere you can actually characterize.

External AI provider
    |
    +--> analysis
    +--> synthesis
    +--> recommendation
    +--> diagnosis
    +--> proposed action
             |
             v
      locally governed system
             |
             +--> authorization
             +--> canonical state
             +--> safety envelope
             +--> execution
             +--> verification
             +--> recovery

In other words, you can put a brilliant consultant in the control room.

Do not make the consultant the emergency-stop circuit.

For many ordinary enterprise workflows, the provider can still be the primary execution path. The consequences are modest enough that a failed call followed by a human retry is acceptable.

But as the consequence of failure rises, the assurance requirement rises with it.

At some point, “the platform usually does it” is no longer a control.

The Real Standard Is Evidence

There is a tendency in AI discussions to make this philosophical.

Can we trust the model? Is the platform aligned? Is the provider safe?

Those questions are too large to be operationally useful.

Standards work is more pedestrian.

  • What is the context of use?
  • What does the component promise?
  • What can fail?
  • How do we detect failure?
  • What evidence survives?
  • Who owns the risk?
  • What happens next?
  • Can we recover?
  • Can we prove we recovered?
  • Can the process survive losing this supplier?

That is why the unexplained repository refusal bothered me more after I stopped being irritated by it. As a user experience problem, it was trivial. As a standards specimen, it was nearly perfect.

  • A provider-controlled mechanism changed whether an operation could occur.
  • The mechanism was outside my control.
  • Its decision basis was not available to me.
  • Its behavior was not sufficiently predictable for me to turn it into a dependable rule.
  • And the only responsible architectural conclusion was to classify it as a failure mode.

Not because the safety mechanism should not exist.

Because a control you cannot characterize is still part of the system you have to engineer.

The system designer does not get to ignore it merely because the uncertainty lives inside somebody else's cloud.

Relevant Standards and Frameworks

NIST AI Risk Management Framework 1.0
Trustworthiness characteristics include valid and reliable, safe, secure and resilient, accountable and transparent, and explainable and interpretable. NIST Publications

NIST - Concept Note: Development of the AI RMF Trustworthy Use of AI in Critical Infrastructure Profile, April 2026
Identifies critical-infrastructure needs including deterministic behavior, explainability, graceful degradation and fail-safe operation. This is a concept note for a profile under development, not a completed standard. NIST

NIST SP 800-53 Rev. 5 - Security and Privacy Controls for Information Systems and Organizations
Relevant controls include SA-9 External System Services, CP-10 System Recovery and Reconstitution, and SI-13 Predictable Failure Prevention. NIST Publications

NIST SP 800-171 Rev. 3 - Protecting Controlled Unclassified Information in Nonfederal Systems and Organizations
Requirement 03.16.03 addresses external system services, provider requirements, shared responsibility and ongoing monitoring. NIST Publications

ISO/IEC 42001:2023 - Information technology - Artificial intelligence - Management system
Provides requirements for establishing and continually improving an AI management system, including structured management of AI risks and opportunities. ISO

ISO/IEC 23894:2023 - Information technology - Artificial intelligence - Guidance on risk management
Provides guidance for integrating AI risk management into organizational activities and functions. ISO

ISO/IEC 27036-2:2022 - Cybersecurity - Supplier relationships - Part 2: Requirements
Specifies requirements for defining, operating, monitoring, reviewing and improving supplier/acquirer relationships, including cloud services. ISO

ISO 22301:2019 - Security and resilience - Business continuity management systems - Requirements
The current published business-continuity management standard, with Amendment 1:2024. A successor edition is under development in 2026. ISO

ISO/TS 22318:2021 - Security and resilience - Business continuity management systems - Guidelines for supply chain continuity management
Extends continuity principles to supplier relationships and supply-chain disruption. It was reviewed and confirmed in 2025. ISO

ISO/IEC 20000-1:2018 - Information technology - Service management - Part 1: Service management system requirements
Covers planning, delivery, monitoring, measurement, review and improvement of managed services and service requirements. ISO

ISO/IEC 25010:2023 - Systems and software engineering - SQuaRE - Product quality model
Provides a quality model for specifying, measuring, testing and evaluating ICT product quality and acceptance criteria. ISO

IEC 61508 series - Functional safety of electrical/electronic/programmable electronic safety-related systems
Provides the general functional-safety framework for E/E/PE safety-related systems. Parts 2 and 3 address techniques and measures for avoidance and control of faults and failures; Part 6 includes common-cause-failure methodology. IEC Webstore

IEC 62443-3-3:2013 - Industrial communication networks - Network and system security - Part 3-3: System security requirements and security levels
Provides technical system-security requirements and security levels for industrial automation and control systems. IEC Webstore

The Rule

The final rule is not complicated.

The system should be able to lose a provider call without losing the work.

For a critical function, I would go one step further:

No externally governed inference or action provider should constitute an unmitigated single point of failure. Where provider behavior cannot be sufficiently characterized, treat the provider as an untrusted-availability dependency. Keep canonical state, governing policy, safe-state behavior, resumption evidence and any required continuity path under independently controlled authority.

That is not anti-cloud.

It is not anti-AI.

It is not even anti-safety-interlock.

It is simply what engineering looks like when somebody else's black box is sitting in your control path.

Latest