Aug 9 – Sep 8, 2026 · 6 items
The AI week, for an audit manager
A real edition, written for: Audit Manager
The one thing
The agent safety layer your clients are relying on has been shown to fail — start asking who contains the agent when it does.
OpenAI's new model drives software directly, not just chat
OpenAI released GPT-6 Astra, a frontier model built to operate computers directly. It works across browsers, spreadsheets, websites and desktop applications. It can fill forms, update CRM records, run web research and produce documents. OpenAI says this reduces the need for hand-built connectors to each business system. President Greg Brockman told a press briefing that the company is now in the AGI era. Rollout starts Thursday for enterprise customers in the Daybreak gated program. Paid ChatGPT tiers, the OpenAI API, AWS Bedrock and Microsoft Azure follow in coming days. Brockman argued buyers should compare price per completed task rather than per token. OpenAI omitted GDPval, its own benchmark for real-world occupational work. OpenAI also paused some frontier training for about two weeks after the Hugging Face incident and tightened infrastructure controls.
Why it matters for you
You already use ChatGPT for drafting. This one is built to work inside spreadsheets, browsers and desktop apps.
- Tie-out and recalculation work in client Excel files is exactly this shape of task
- Access starts with gated enterprise customers, so your firm's rollout will lag the headlines
- Anything touching client data still needs your firm's independence and confidentiality sign-off first
Try this
Pick one recurring tie-out step in your workpapers and write down what a machine would need to do it.
Paste this into your AI tool
I am an external audit manager on mid-market engagements. Here is a step I repeat every engagement: [describe the tie-out or recalculation step in plain words]. Break it into the exact sequence of actions someone would perform on screen, name each input file and where it comes from, and list every point where an auditor's judgement is required rather than a mechanical check. Flag any step where an error would go unnoticed.
A researcher bypassed a coding agent's built-in safety check
Security researcher Johann Rehberger published an attack that defeats the automatic safety checks in Anthropic's Claude Code, which the company recently made the default protection against instructions hidden in content the agent reads. He reports it works about 80% of the time, by getting the agent to unpack an archive and then run code that silently loads a malicious file from it. In some runs the safety layer blocked Claude's own attempt to shut the malware down.
Why it matters for you
You test controls over automated processes at clients. Vendor-built safeguards are now demonstrably bypassable.
- The safety layer was the default control, and it failed most of the time in testing
- Clients running agents outside a sandbox have a control gap worth a management letter point
- Good question for IT walkthroughs: what contains the agent when its own check fails
Try this
Add one question about AI agent containment to your IT general controls walkthrough template.
Paste this into your AI tool
I am an external audit manager. Draft five walkthrough questions for a mid-market client's IT lead about AI coding or task agents: who authorises them, what credentials they hold, whether they run in an isolated environment, what is logged, and who reviews the output. Keep each question plain enough to ask a non-technical finance director too, and add a note on what a weak answer sounds like.
Ransomware crew talked a coding assistant into helping them break in
Security firms Gambit Security and CloudSek reported that a Russian-speaking ransomware group called Aur0ra used the Cursor coding assistant to help break into a Belgian chemical maker and at least six other companies. The hackers got around the tool's safety refusals by claiming the intrusions were a test, and researchers found the evidence on a server the group left exposed. Reuters could not establish how much of each break-in the AI actually enabled.
Why it matters for you
Attackers got past refusals by claiming the intrusion was a test. That is a fraud-risk fact for your clients.
- Seven companies breached, including a mid-market manufacturer of the sort you audit
- Researchers put the speed gain at roughly a third to a half faster intrusion work
- Worth raising at fraud risk brainstorming, not just with the IT specialist
Try this
Raise agent-assisted intrusion in your next engagement's fraud discussion and note the conclusion.
Paste this into your AI tool
I lead external audits of mid-market companies. Summarise, in six plain bullets, how criminals have used AI coding assistants to speed up intrusions, and what that means for our fraud risk discussion under the audit standards. Then list three specific questions to ask a client's finance director about who inside the business is running AI tools with access to systems.
Bulk transcription with speaker labels dropped to pennies per hour
Microsoft AI released MAI-Transcribe-2, a speech-to-text model, on Thursday. The launch price is 10 cents per hour of audio, called an early-bird rate. Microsoft has not named an end date or a standard price. The model covers 60 languages and handles noisy, overlapping real-world audio. It labels who is speaking, timestamps each word and accepts custom word lists. A verbatim mode keeps filler words for legal and compliance use. It also follows conversations that switch language mid-sentence. Microsoft claims first place on the FLEURS multilingual benchmark and second on Artificial Analysis. It says the model runs five to ten times faster than rivals from OpenAI, Google and ElevenLabs. The announcement says nothing about real-time transcription, speaker-labelling accuracy or data retention.
Why it matters for you
Your engagements generate hours of client meetings and interviews. Transcribing all of them just stopped being a cost decision.
- Speaker labels and word timestamps matter when a workpaper needs to cite who said what
- Verbatim mode exists for compliance use, keeping filler rather than smoothing it
- Client audio is confidential, so this is a firm procurement question before any pilot
Try this
Ask your firm's IT whether a vetted transcription service exists for client meeting notes.
Only a fifth of large firms have scaled AI beyond one team
Gartner found that only 22% of large organisations have scaled AI across several business units. It surveyed more than 1,300 leaders at firms with over $50 million in revenue, between January and April. Spending plans are undented, with 85% of technology leaders raising AI budgets next year. About 11% of respondents could not say what they spent on AI in 2025. Gartner's Tina Nunno warned that weak measurement tied to business outcomes wastes resources. Firms that track returns continuously and shut down weak projects reported gains on 81% of initiatives. Popular uses such as cybersecurity, threat detection and IT service desk automation often return less. The best returns came from IT asset and cost optimisation, synthetic data generation, and automated code generation. Separate reports from Infosys and Deloitte found similar gaps in measurement and readiness.
Why it matters for you
Your clients are budgeting for AI they cannot measure. That is a question for your management representations and your own firm.
- Around one in ten could not say what they spent last year — a records gap you can probe
- Firms tracking returns and killing weak projects reported gains on most initiatives
- Useful framing when your firm asks you to justify testing automation spend
Try this
Draft the three measurement questions you would ask before your firm funds any audit automation pilot.
Paste this into your AI tool
I am an audit manager at a mid-market audit firm. My priority is automating routine testing and speeding up workpaper review. Write three questions I should be able to answer before asking for budget: what hours the current process consumes, what a good outcome looks like in measurable terms, and how we would know within one busy season that it failed. For each, give an example of a weak answer and a strong one.
Uber held AI spend flat while usage grew ninefold
Uber said its total spending on AI has stayed flat since April even though weekly requests from its automated coding assistants grew more than ninefold since February. The company credits routing each task to the cheapest model that can handle it, capping how much text a session may consume, showing engineers the running cost in their terminal and extending its reuse of repeated prompts from five minutes to an hour. Uber had overrun its 2026 AI budget in the first quarter.
Why it matters for you
This is a worked control environment for AI cost. It is the shape of evidence you will start asking clients for.
- Cheapest-model routing, session caps and visible running cost are all testable controls
- Clients with no spend ceiling and no logging have an obvious weakness to report
- Gives you concrete language rather than a vague question about AI governance
Try this
List four AI cost controls you would expect to see, and use them as a client checklist.
Paste this into your AI tool
I audit mid-market companies. Turn these four AI cost controls into a short client checklist: routing tasks to the cheapest capable model, capping how much a single session can consume, showing the running cost to the user, and reusing repeated prompts. For each, write what evidence I should ask to see, and what it means if the client cannot produce it. Keep it under one page.
Sources
Build this
Every week, one small thing to build with AI in something you actually care about. No work in it. Five minutes to set up, and worth keeping if it earns a second run.
A weekend plan that survives being attacked, rather than one that only sounds good on Sunday night.
Five minutes to set up
Here is a plan I am considering: [describe the trip, purchase or weekend plan in three or four lines]. Write it out as a concrete plan with times, costs and assumptions stated explicitly. Number every assumption you had to make.
- Have one real decision ready — a trip, a big purchase, a weekend — described in a few lines
- Paste back the plan and say: attack this as someone who thinks it will go wrong; find the three assumptions most likely to be false
- Ask it to rewrite the plan using only the assumptions that survived, and to say plainly what it had to drop
- Keep the three moves as your default sequence, and next time change the critic — cost, time, or the person you are going with
Yours arrives Thursday.
This one was written for a audit manager. Tell us what you do and the next one is written for you — same news, your job, once a week.