10 Sept 2026 · 8 min read
AI Is Doing More Than Drafting. Our Review Practices Need to Catch Up.
What real human–AI work reveals about professional judgment, evidence, approval and the limits of passing checks. An anonymised report for accountants, corporate service providers and firm owners.
Findings from real human–AI work in professional practice
By Darryl Wong · September 2026
The familiar advice to professional firms is to use AI for preparation and leave judgment to people.
That distinction is becoming difficult to maintain. In the records examined for this report, AI did not merely produce documents. It questioned classifications, connected facts across domains, requested evidence and proposed changes to professional work. People supplied business context, challenged conclusions and authorised actions. Both also accepted assumptions that deserved closer examination.
The most useful question is therefore not simply whether human judgment still matters. It is what makes a human–AI decision worthy of reliance.
Our strongest finding is that the quality of the interaction matters: what evidence is available, how a challenge is answered, what the checks actually test and what an approval means. A person being present does not, by itself, establish that those conditions have been met.
This report draws on an exploratory, AI-assisted review of 165 local Codex session files. It presents anonymised observations, not a controlled comparison of human and AI performance.
AI can improve the question before it improves the answer
In one corporate tax computation, the agent rejected simply relabelling an expense to obtain a more favourable treatment. It continued searching for a legitimate alternative and connected other expenditure with the commercial activity it supported and a potential further deduction.
It then asked a consequential question: what service had actually been performed? The answer could distinguish a marketing campaign from a different kind of commercial service. A practitioner obtained clarification, and the computation was revised.
The revised calculation was inspectable and its arithmetic was corroborated. The available evidence did not independently establish the full eligibility of the claim or realised tax savings.
Even with that limit, the interaction shows something important. The agent contributed a professional hypothesis and identified evidence that could change the decision. The human contribution was to pursue the question and supply context unavailable in the ledger.
The value was not just faster calculation. It was a broader search for a defensible answer.
Human knowledge matters when records do not explain the arrangement
In another case, an agent inferred revenue ownership from how an invoice or billing relationship appeared. The practitioner supplied a different explanation of the commercial arrangement. An accounting adjustment followed, and the recorded change was checked.
That establishes the effect of the human intervention. It does not independently prove the underlying contractual entitlement.
The distinction matters because “human judgment” is often described too vaguely. Here it meant knowledge of how the business actually worked—information that a document label did not contain.
For firm owners, this suggests a practical training priority: teach staff to explain the arrangement, the parties' responsibilities and the limits of the records, not merely to give the agent more files.
A human challenge can still end in an unsafe decision
The payment case provides a less comfortable contrast. A practitioner questioned whether obligations had already been paid. The agent offered contrary assurance, and an adverse payment sequence followed.
The later investigation found a concrete search problem. A local calendar date was translated into a UTC boundary. Earlier payments fell outside that boundary despite belonging to the relevant local business date. The search could therefore miss the very evidence needed to answer the human's challenge.
The same payments had been correctly excluded in an earlier task. A successful check had not reliably carried forward.
This supports a specific retrieval-and-control failure, not a claim that the model simply “forgot,” nor proof that the date boundary was the only cause.
The operational lesson is nevertheless clear: a challenge is only as useful as the verification that follows it. A confident answer to “are you sure?” is not equivalent to demonstrating the relevant payment history and search coverage.
Approval can close a workflow without resolving its facts
Several records moved from open questions to “closed,” “final” or accepted treatment after broad confirmation or a management instruction.
Those transitions did not always provide an item-specific explanation of the underlying facts. In one retained workbook, a closure record expressly distinguished management-adopted positions from other forms of support. In another workflow, confirmation wording was removed without that removal itself demonstrating the answers.
This does not mean every treatment was wrong. Relevant knowledge may have existed outside the recorded interaction. It means the record did not establish the same thing as the status label.
A useful review should keep four questions separate:
Review question — What it establishes
What supports the fact? — The basis for the underlying assertion
What treatment was chosen, and why? — The reasoning applied to that fact
What was checked in this exact output? — The scope of technical verification
Who authorised the next action? — Permission to proceed
One “approved” response should not silently answer all four.
Passing checks can create confidence beyond their scope
In a financial-statement case, the cash-flow totals reconciled. A separate relationship between dividends, cash payments and opening and closing payable balances remained unresolved in the inspected evidence.
Further examination traced the reported payment figure to a consolidation calculation using a subsidiary statement and an elimination of a parent-company receipt. That finding was important counterevidence: it would have been unjustified to describe the number as an invented balancing figure.
But an explainable derivation was still not independent confirmation of the underlying payments. The checks demonstrated arithmetic consistency without settling the whole substantive question.
A related problem appeared in review documentation. A section labelled “Independent tax review” was authored within the same session. The heading did not demonstrate an uninvolved reviewer or an independent assessment.
The lesson is to describe a check by what it proves. “The totals reconcile,” “the source matches” and “the treatment has been independently reviewed” are different claims.
What firms should change
The findings suggest that AI training should include supervision of decisions, not only prompting and software operation.
Make challenges evidence-based. Ask what source was checked, whether it covers the right entity and period, and what would change the answer.
Make uncertainty durable. An unresolved question should remain visible when work moves between tasks or people. A lesson stored somewhere is not proof that the next task will apply it.
Make approval specific. Record whether a person confirms a fact, adopts a treatment, accepts an output version or authorises an action.
Measure first-time quality separately from recovery. A successful repair matters, but it should not erase the original failure when evaluating the workflow.
Train for reciprocal correction. Humans should be able to challenge AI, and AI should be able to question a human instruction without treating every request to finish as evidence that the work is correct.
These are recommendations informed by observations. Their effectiveness has not been tested as a training or control intervention in this study.
How this report was developed
The study used existing local Codex session records and selected retained outputs from professional and operational workflows. No synthetic demonstration was substituted for the cases described here.
Of a frozen cohort of 598 eligible session files, 165 received substantive dialogue review. Eighteen older-format files were excluded from that analytical cohort and retained as sources. The review produced 280 case-analysis documents; these are overlapping or related analytical records, not 280 independent experiments.
Selected findings were checked against archived tool calls and results, document contents, workbook formulas, saved calculation results and file identities. Not every artifact or business fact was independently verified. Hidden model reasoning was excluded from the analysis.
The analysis was performed with Codex assistance, under the direction of a practitioner involved in some of the underlying work. First-pass coding was prepared for 24 selected decision windows. No completed second-review responses or inter-reviewer agreement results underpin this report. It is not a blinded, independently validated or peer-reviewed study. AI-assisted interpretation can itself make mistakes.
Cases were selected for explanatory value, including adverse and contradictory observations. They cannot establish population error rates, a general human-versus-AI winner or causal effects of particular models, tools or training. Client names, exact amounts, transaction dates and operational identifiers have been omitted; the underlying records remain private, limiting public reproducibility.
Timing was assessed separately across 177 files, a broader set than the substantive review. Recorded duration includes tool execution and waiting, and task scopes differ. It is not measured human labour. This study does not establish hours saved, quality-adjusted productivity gains or sustained behavioural improvement over time.
The report is deliberately bounded by the evidence already available. Unresolved outcomes remain unresolved; this is not an audit opinion or advice on any particular client's accounting or tax treatment.
The conclusion
These records show AI contributing to professional reasoning—and show why neither its confidence nor a human approval should end the inquiry automatically.
The useful direction for professional firms is to build a working relationship in which both sides can question the other, and the resulting decision remains traceable to evidence and a clearly defined check.
Human judgment still matters. So does the quality of the system through which it is exercised.
Share and save
Living source
This post is the stable site version. The source gist may be updated as the working pattern develops.
Read the complete source report on GitHub Gist