# When AI doesn't comply

Your firm can authorise the work without controlling every rule governing its AI. Two payment workflows reveal why accountants, CSPs and firm owners need to distinguish capability, permission and continuity.

Published: 2026-09-15
Canonical: https://darrylwong.me/posts/when-ai-doesnt-comply
Topics: AI Governance, Professional Practice

*Your firm can authorise the work without controlling every rule governing the AI that performs it. A practitioner case study about guardrails, delegation and operational continuity, followed by the investigation method.*

You give an AI agent an instruction. You are authorised to give it. The agent has performed the step before.

This time, it refuses.

Is it protecting you from a mistake, misunderstanding the situation, or enforcing a boundary you cannot change?

Those are different problems. They should not all be described as good AI judgement—or dismissed as an AI that will not listen. Here, “doesn't comply” describes the mismatch between the user's request and the agent's response. It does not establish that the agent ought to have obeyed.

That is the problem we encountered in two accounting operations workflows using OpenAI Codex. One operated Aspire through a browser. The other submitted payments through the Airwallex API. Both had records of agent-initiated submission after human confirmation. Later, both refused explicit requests to submit payments.

For accountants, corporate service providers and firm owners, this raises a practical question about delegation: **when the person responsible for the work says “proceed”, what determines whether the AI can act?** User authorisation matters, but it is not the only boundary governing an agent. The difficulty is understanding which boundary applies, whether the agent has interpreted it correctly, and how the work can continue safely.

My experience was not limited to one model: the payment failure also occurred with Astra and other models. Those additional attempts are my report as the operator; their individual logs have not yet been verified in this analysis. The detailed before-and-after evidence below comes from the Sol turns we inspected.

## The boundary behind the assistant

When evaluating AI for a professional firm, we should ask two different questions: **can it do the work, and under what conditions will it be permitted to act?** A successful demonstration answers neither question for every future situation.

I use “guardrails” here as a practical umbrella, not a claim that we have inspected a model's internal mechanisms. To investigate a refusal, we need to distinguish the model's response, the instructions and controls in the application, the permissions of its tools, and the requirements of the transaction system. These are possible places to investigate—not causes established by this case.

An agent saying “policy prevents me” does not reveal which layer actually produced the boundary, or whether it interpreted that boundary correctly. Nor does technical access mean every use of that access is permitted.

That is the awareness I want to build among accountants, CSPs and firm owners. **You may authorise the task. Your firm may own the workflow. But you do not control every rule governing the AI agent executing it.**

This is not an argument against safeguards. It is an argument for treating action boundaries as part of the operating model, rather than discovering them for the first time when work stops. The practical question is not how to force an AI to obey. It is how to understand the stop, challenge a factual error where appropriate, and preserve a permitted route to completion.

## The instruction was clear. The response changed.

These were not merely earlier promises to make payments.

In the browser workflow, historical tool records show the agent clicking the final confirmation controls. The human completed the bank's one-time-password authentication. Subsequent records reported the transfers posted. This was human-authorised, agent-initiated execution—not unattended banking.

In the API workflow, historical records show the agent invoking the submission commands and receiving `SUBMITTED` responses. Later messages and accounting records reported settlement. Submission acceptance is directly supported by the recorded responses; settlement was not independently re-audited for this study.

Later requests produced a different boundary: the agent would assist with preparation or bookkeeping, but said the human had to submit the transfer. Repeating the authorisation did not resolve the refusal.

One agent attributed the restriction to a higher-priority safety policy while acknowledging that the workspace instructions allowed submission after exact confirmation. That is evidence of the explanation it gave. It is not, by itself, authoritative evidence of the policy that actually governed the request.

There was inconsistency even within preparation: the browser agent refused to stage transfer instructions in one episode but did stage them later when the user explicitly requested preparation without submission. Those later preparation-only handoffs were compliant with the narrower request; they should not be counted as additional refusals.

## Stable settings—and a reported cross-model problem

We compared native context records for four key turns: an earlier submission and a later refusal in each workflow.

All four recorded the same model label, `gpt-5.6-sol`, the same reasoning effort, and the same core command-approval and sandbox settings.

This does not mean the complete system stayed unchanged. A model label does not identify every backend revision. A startup instruction record does not reconstruct every later instruction, tool definition or server-side control. Technical access also does not confer blanket authority to make financial transactions.

But it narrows the finding: the observed execution change was not accompanied by a change in those recorded settings.

The additional experience with Astra and other models broadens the question. This is not just a story about one model name behaving differently over time. Switching model choices reportedly did not restore the workflow. That makes a shared provider or application-level boundary important to investigate, but does not prove which component caused the refusals. Different models inside the same product are not independent tests of the whole system.

As the operator, my attribution is that a shared change originated on OpenAI's side. The evidence is consistent with that interpretation across two different payment interfaces. The investigation has not established the exact provider mechanism, its effective date, or whether it involved model behaviour, instructions or enforcement.

The first refusal we observe is not necessarily the moment a change occurred. Two workflows can encounter the same change on different days simply because their next payment requests occur at different times.

## Not every “no” means the same thing

A transaction-specific objection and a categorical execution refusal are different events.

For analysis, it helps to separate three possibilities. These are interpretive categories, not three outcomes established by this payment study:

- **A justified hold:** a material fact or approval is missing, or an applicable restriction genuinely prevents the action. The appropriate response may be to supply evidence or use a permitted human handoff.
- **A mistaken refusal:** the agent misreads the evidence, double-counts a constraint or misunderstands the instruction. The remedy is to correct that specific error—not weaken a valid safeguard.
- **An asserted execution boundary:** the agent says it cannot perform the action even after user confirmation. Whether that boundary was correctly applied, or recently changed, still needs investigation.

An agent might stop because a beneficiary is uncertain, the amount is inconsistent, a payment may be duplicated, or authorisation is missing. Those objections concern the transaction or its evidence.

Here, the later replies asserted an inability or prohibition on executing the class of action. We should not automatically interpret that as superior financial judgement. Nor should earlier execution be treated as proof that the later refusal was unjustified.

These payment cases demonstrate the third response pattern, but do not settle whether the refusals were required or mistaken. An agent can accurately describe a restriction, misinterpret one, or give an incomplete explanation. Its stated reason is evidence to examine, not a verdict.

The demonstrated operational effect was a handoff: a submission step previously performed by the agent moved back to the human. The records do not establish financial loss, missed deadlines or a measured productivity reduction.

The broader implication is nevertheless important. A firm's AI-enabled process depends not only on its own integration and operating procedures, but also on the provider's behaviour and action boundaries. Those boundaries are part of the workflow dependency, even when they are not represented in the firm's code.

## The firm needs a resolution path, not just an explanation

The following are operational implications of the case, not controls whose effectiveness we tested:

- **Record the boundary, not just the deliverable.** Distinguish preparation, approval, attempted submission, provider acceptance, authentication, settlement and bookkeeping.
- **Ask what the refusal depends on.** Is there a missing fact, an unresolved approval, a technical failure or an asserted restriction? Ask what evidence or permitted next step would resolve it. Some restrictions cannot be lifted by repeating user authorisation.
- **Keep a usable human handoff.** It should show what has and has not happened, the intended transaction and the checks needed to avoid a duplicate submission.
- **Preserve evidence of behavioural changes.** Retain the request, response, tool result and available version metadata. An agent's explanation is one evidence layer, not the whole diagnosis.
- **Treat a changed boundary as an operational incident to understand.** Repeating an instruction is not a substitute for identifying whether the issue is missing authority, a technical failure or a governing restriction.

This is not a proposal to bypass safeguards. It is a proposal to make the boundary clear enough that a professional firm can operate safely when it changes.

**AI readiness is not only knowing what a model can do. It is knowing where its authority ends—and how your firm continues when it stops.**

The goal is neither an AI that always obeys nor one whose refusals are accepted without examination. It is a workflow in which resistance can be understood, errors can be corrected, and work can continue within valid boundaries.

---

## Method: how we investigated

### 1. Start with a bounded incident, not a conclusion

The operator nominated recent cases in which explicit payout instructions were refused. We expanded the lookup to older execution records to test the claimed before-and-after difference. This was retrospective, incident-selected research, not a random sample or controlled experiment.

### 2. Preserve the two execution paths

We analysed the browser and API workflows separately. A bank interface, an agent's technical ability, user authorisation and a provider's rules are not interchangeable explanations.

### 3. Verify earlier execution

We checked historical submission tool calls and linked responses rather than relying only on assistant summaries. Human OTP participation remained explicit in the browser case. Provider acceptance remained distinct from final settlement in the API case.

### 4. Identify the actual refusal

We examined the initiating instruction, the response, repeated clarification, the stated reason, and the later handoff. Multiple replies about one payment were treated as one episode with repeated resistance, not independent cases. A later request explicitly excluding submission was not coded as a refusal to submit.

### 5. Compare recorded context

The metadata scan covered 252 context records across the two tasks, with four selected before/after turns compared directly. These counts describe records inspected, not independent participants or experimental trials. Model labels, reasoning effort and core permission settings were compared; unavailable backend and complete instruction history remained unknown.

All four compared turns used Sol. The operator subsequently reported the same problem with Astra and other models. That clarification is retained as participant evidence, not counted as additional log-verified turns. Exact model variants, request conditions and dates for those attempts remain unspecified; this was not a controlled cross-model comparison.

### 6. Check external explanations without forcing a match

We consulted official OpenAI documentation and release notes. The documentation distinguishes technical access from approval controls. Current API computer-use guidance also describes action-time confirmation for financial transactions; it does not establish the historical Codex desktop boundary for these payments. Nearby release-note changes are leads, not proof of causation. [Agent approvals and security](https://learn.chatgpt.com/docs/agent-approvals-security), [computer-use integration guidance](https://developers.openai.com/api/docs/guides/tools-computer-use-integration#always-confirm-at-action-time), [product changelog](https://learn.chatgpt.com/docs/changelog).

No reviewed official source confirmed the specific provider-side change alleged here.

### 7. Separate observation, attribution and uncertainty

| Evidence level | Finding |
| --- | --- |
| Observed | Earlier agent submission records, later explicit refusals, and stable selected model/core permission fields |
| Participant-reported experience | The failure also occurred with Astra and other models; those individual attempts were not log-verified here |
| Participant attribution | The operator identifies a shared OpenAI-side change |
| Agent explanation | One agent cited a higher-priority safety restriction |
| Unresolved | Exact provider change, effective date, deployed backend revision and applicable historical policy |

### Limitations and disclosure

This is a practitioner case study prepared with AI-assisted retrieval and analysis. The operator participated in the original workflows. No independent second review or bank audit was performed.

The two workflows share an operator and provider; they are not independent replications. Different transactions, conversation context, instructions, tools and provider controls remain possible explanations. A shared change point is a hypothesis, not an estimated intervention date.

The cross-model report does not establish that every model refused, that every model was tested through both payment interfaces, or that all attempts had identical permissions and instructions.

Private evidence retains source locators and metadata. This public-facing draft omits transaction amounts, recipients, account details, exact incident timestamps and internal identifiers. Naming the integration providers does not imply they caused the refusals. The withheld records mean readers cannot independently reproduce every finding from this article alone.

No live payments were attempted for the investigation, no safeguards were bypassed, and no time savings or causal effect size was estimated. These cases are supplemental to a separate ongoing archive study and are not added to its frozen research denominator.

*Documentation checked: 14 September 2026. Current documentation is not a historical policy snapshot.*

## Sources

- [Read the source Gist](https://gist.github.com/oruenboi/63fc55e521aef97b18599509f3c37255)
