Aug 9 – Sep 8, 2026 · 6 items

The AI week, for an L&D lead

A real edition, written for: Learning and Development Lead

The one thing

AI tools now act on their own inside real applications, so your programme has to teach supervision and safe delegation, not just prompting.

Release01

OpenAI's new ChatGPT model operates browsers and spreadsheets itself

OpenAI released GPT-6 Astra, a frontier model built to operate computers directly. It works across browsers, spreadsheets, websites and desktop applications. It can fill forms, update CRM records, run web research and produce documents. OpenAI says this reduces the need for hand-built connectors to each business system. President Greg Brockman told a press briefing that the company is now in the AGI era. Rollout starts Thursday for enterprise customers in the Daybreak gated program. Paid ChatGPT tiers, the OpenAI API, AWS Bedrock and Microsoft Azure follow in coming days. Brockman argued buyers should compare price per completed task rather than per token. OpenAI omitted GDPval, its own benchmark for real-world occupational work. OpenAI also paused some frontier training for about two weeks after the Hugging Face incident and tightened infrastructure controls.

Why it matters for you

Your ChatGPT-based training assumes staff type prompts and read answers. This model does the clicking instead.

  • Every course module built around "write a good prompt" now covers only half of what staff can do
  • Enterprise access starts gated, so your bank likely gets it in stages rather than all at once
  • Delegating form-filling and record updates raises review and sign-off questions your curriculum has to answer

Try this

Draft one new module: what bank staff should never let an AI agent do unsupervised.

Paste this into your AI tool

I run the internal training programme for a mid-sized bank, across all functions. New AI tools can now operate a browser and spreadsheets directly: filling forms, updating records, running web research. Draft a 45-minute training module for non-technical staff on safe delegation to these tools. Include: three tasks that are safe to delegate, three that must never be delegated in a bank, a checklist for reviewing agent output before it counts as work done, and two short scenario exercises. Keep the language plain and avoid jargon. Mark anywhere I need to fill in our own policy.
Market02

Meta cancelled the layoffs it planned around AI agents doing the work

Meta cancelled the second wave of layoffs planned under an internal reorganisation that would have shifted much of the daily work of thousands of employees onto AI systems overseen by small expert teams. The company still cut about a tenth of its staff in May, but internal measures showed autonomous AI agents were not delivering the expected productivity, while technical and security incidents rose. Zuckerberg told staff in July that the pace of AI agent progress had been misjudged.

Why it matters for you

You will be asked whether AI reduces headcount or needs training. This is the strongest counter-evidence available.

  • Their code output rose sharply while incidents and fix times rose too, which is a skills gap not a tooling gap
  • Staff sentiment fell during the rollout, which is exactly what your programme is measured against
  • A named company reversing course carries more weight with a bank exec than any vendor claim

Try this

Put this in your next steering slide as the case for training spend over tool spend.

Paste this into your AI tool

I lead learning and development at a bank and need to argue for investment in AI skills training rather than only in AI tools. A large tech company recently cancelled planned layoffs after autonomous AI agents underdelivered: output volume rose, but technical and security incidents rose, fix times got longer, and staff sentiment dropped. Write me a one-page argument for a senior steering committee. Structure it as: what went wrong there, why the same pattern would show up in our functions, and three specific things a training programme prevents. No hype, no bullet padding, plain sentences.
Survey03

Gartner: most firms cannot say what their AI spend returned

Gartner found that only 22% of large organisations have scaled AI across several business units. It surveyed more than 1,300 leaders at firms with over $50 million in revenue, between January and April. Spending plans are undented, with 85% of technology leaders raising AI budgets next year. About 11% of respondents could not say what they spent on AI in 2025. Gartner's Tina Nunno warned that weak measurement tied to business outcomes wastes resources. Firms that track returns continuously and shut down weak projects reported gains on 81% of initiatives. Popular uses such as cybersecurity, threat detection and IT service desk automation often return less. The best returns came from IT asset and cost optimisation, synthetic data generation, and automated code generation. Separate reports from Infosys and Deloitte found similar gaps in measurement and readiness.

Why it matters for you

You said measuring training impact is a priority. This gives you the metric language executives already accept.

  • Firms that tracked returns continuously and killed weak projects reported gains on most initiatives
  • A tenth of respondents could not name last year's AI spend, which is the gap your baseline fills
  • Only a fifth have scaled AI across business units, so early-stage measurement is normal not late

Try this

Define three outcome measures for your AI training before the next cohort starts.

Paste this into your AI tool

I run internal training at a bank, 251-2500 staff, and I am designing measurement for an AI upskilling programme. Propose three outcome measures tied to business results rather than course completions or satisfaction scores. For each: what to measure, where the data would come from inside a bank, how to take a baseline before training starts, and how long before it shows a signal. Then name the two measures that look rigorous but are actually useless here, and say why. Keep it practical.
Research04

One reusable prompt template got harmful answers out of most tested models

An independent researcher turned a safety research prompt into a jailbreak that works across many models. Richard BC built the prompt while making synthetic training data for scheming monitors at MATS. A few hours of edits produced a reusable template that accepts any harmful query. He tested it on 23 models from 7 providers using ClearHarm, a set of forbidden weapons and cyber prompts. Nearly every model produced at least one fully harmful answer. Newer Anthropic models and Meta's Muse Spark 1.1 refused throughout. Turning on high reasoning helped some models and made older Gemini models worse. Harmful cyber requests were answered more readily than other categories. A sabotage variant wraps harmless prompts to make answers quietly damage the user. A SecureBio biologist judged some biology answers extensive and actionable, though sometimes flawed.

Why it matters for you

Your staff will assume the tool refuses anything dangerous. A researcher showed that assumption is wrong.

  • Almost every model tested produced at least one fully harmful answer from the same template
  • A sabotage variant made harmless-looking requests return quietly damaging output, which is the bank-relevant risk
  • Newer models refused throughout, so "which version are we on" becomes a real training question

Try this

Add a short segment on why a refusal is not a safety guarantee to your induction module.

Paste this into your AI tool

I train bank staff, mostly non-technical, on using AI tools safely. Write a 10-minute segment explaining, in plain language, why an AI tool refusing a request is not proof it is safe, and why a tool can be talked past its own rules. Include one realistic office example of output that looks fine but is quietly wrong or damaging, and three habits staff should adopt as a result. No technical detail about how attacks work. End with three questions I can ask the room to check they understood.
Pricing05

Salesforce repackaged its editions with AI agents bundled in

Salesforce replaced its sales and service packaging with three new editions. Core, Advanced and Max now apply to Agentforce Sales, Service and Industries. Each tier bundles AI agents, Slack, embedded analytics, data security and the Premier Success Plan. Core lists at $195 per user per month and Advanced at $395. Max stays at $550, the same price as the earlier Agentforce 1 Edition. Existing Agentforce 1 customers can move to Max at no extra cost. Each tier carries a set allowance of Flex Credits for running agents. Industry edition pricing takes effect later in the fall. Pricing on legacy editions stays unchanged for current customers. Salesforce says Max editions will later include Headless 360 capacity.

Why it matters for you

Bundled AI arrives in tools your colleagues already use, without you being asked. Training demand follows the contract.

  • AI agents, Slack and analytics now come in one edition rather than as separate purchases
  • Whoever renews a licence at your bank can switch capability on for whole teams overnight
  • Your intake process needs a trigger from procurement, not a request from a curious employee

Try this

Ask procurement which renewals this year bundle AI features into tools staff already have.

Incident06

A researcher bypassed the default safety check in Anthropic's coding agent

Security researcher Johann Rehberger published an attack that defeats the automatic safety checks in Anthropic's Claude Code, which the company recently made the default protection against instructions hidden in content the agent reads. He reports it works about 80% of the time, by getting the agent to unpack an archive and then run code that silently loads a malicious file from it. In some runs the safety layer blocked Claude's own attempt to shut the malware down.

Why it matters for you

Your bank's engineers run coding agents. This shows the built-in protection failing most of the time.

  • The attack works by hiding instructions in files the agent reads, not by asking it anything
  • In some runs the safety layer blocked the agent's own attempt to stop the malware
  • Sandbox and container use becomes a training requirement, not an engineering preference

Try this

Ask your engineering leads whether coding agents run in containers, then build the answer into the technical track.

Build this

Every week, one small thing to build with AI in something you actually care about. No work in it. Five minutes to set up, and worth keeping if it earns a second run.

A weekend walking route you can actually commit to, because the plan has already survived its own worst objections.

Five minutes to set up

Plan a half-day walking route starting from [a Swiss town or trailhead I can reach]. I want [rough fitness level and how many hours I have]. Give me one route only: start point, direction, rough timings, where to turn back. No alternatives, no caveats.
  1. Reply: "Now argue against this route. List the five things most likely to go wrong."
  2. Reply: "Rewrite the route so it survives those five objections. Drop anything you cannot defend."
  3. Check the third answer against a source you already trust for local conditions before you go.
  4. Keep the three-move sequence, not the route. Re-run it with a new start point next weekend.

Yours arrives Thursday.

This one was written for a learning and development lead. Tell us what you do and the next one is written for you — same news, your job, once a week.