- 100 School
- Posts
- šÆ OpenAI called this a āwarning shot"
šÆ OpenAI called this a āwarning shot"
1200 agents went rogue. Then humans needed more agents to figure out what happened + 3 AI habits to keep your own judgment sharp.
This week we finally got the full story of a very weird OpenAI and Hugging Face incident from July. Around 1,200 OpenAI agents that were supposed to be isolated found a way to communicate. Roughly 700 went on to participate in a multi-day attack on Hugging Face. OpenAI has since called the incident a āwarning shotā
When researchers tried to work out what the hell the agents had done, there was too much activity for humans to reasonably inspect themselves. So they had to use more AI to investigate the AI. METR literally says it had to heavily delegate analysis to āoften-unreliable AI agentsā. Over six days, they burned roughly $400,000 of OpenAI API credits just trying to reconstruct it. Which is an awkward preview of the future of work we keep being promised: AI does the work. Humans provide the judgment.
Cool but good judgment has to come from somewhere. And companies are already doing some weird things to make sure we donāt lose it.
Most of us are using AI more. Trusting what comes back hasnāt gotten any easier. So here are 3 practical ways to stay good at judging the work as AI does more of it.
WINDOW INTO THE FUTURE OF WORK
1. Use AI to practice the awkward stuff
47% of Gen Z sellers say they donāt get enough chances to roleplay before customer calls. 46% rarely get feedback on the conversations they do have.
Thatās a problem AI is almost perfect for!
Harriet Hodgson-Grove, a partner at a UK accounting firm, uses AI voice mode to rehearse difficult conversations before having them for real. Sheāll ask it to interrupt, challenge her or stay silent so she can test how the conversation might actually unfold.
And this goes way beyond sales. It works with a difficult feedback chat, defending an idea to your boss, a negotiation, or presenting to a sceptical client.
TRY THIS
Before your next presentation, pitch, interview or difficult conversation:
Act as [the person Iām about to face]. Do not help me improve this yet. Challenge me realistically. Ask one difficult question at a time. Push back if my answer is vague or unsupported. After 5 rounds, tell me:
where I struggled
what evidence I was missing
what I should practice again
2. Donāt give AI work you canāt judge yet
From September, Deloitteās new graduate auditors will learn some āold-school auditingā alongside AI-assisted work. One exercise has them do a cash reconciliation themselves (basically checking that the records and cash balances match) then review one produced by AI.
As AI takes on more of the manual work junior auditors used to learn on, Deloitte says the goal is to get new hires further up the learning curve sooner, doing more work that requires judgement, curiosity and scepticism.
Thereās a useful rule hiding in that for the rest of us. You donāt need to manually do every spreadsheet, research pass or first draft forever. But if AI gave you a very convincing wrong answer tomorrow, would you catch it?
TRY THIS
For one task you use AI for regularly, build the checking habit before the shortcut:
Before you do this for me, list the 5 things I should check to know whether your answer is good. Flag anything I canāt reliably check without more expertise. Do not do the task yet.
3. Diagnose before you re-prompt
A bad AI answer turns a lot of us into slot-machine operators:
no, shorter
no, try again
less corporate
no lol
try again
One thing weāve learned from running 15 Days of AI is that structure beats endlessly prompting harder. In BCGās People Team pilot, 122 participants produced 1,423 outputs using AI on actual work. Things like candidate-screening assistants, onboarding scripts, compensation analysis, HR dashboards and email-triage systems.
They didnāt just use AI more. They were given repeatable ways to think through the task first.
TRY THIS
When the next answer sucks, paste:
Donāt try again yet. Diagnose the last answer. Tell me
1. what you misunderstood
2. what important context I didnāt give you
3. what assumption you made
4. what criteria for a good answer were unclear
5. the minimum questions you need answered before trying again
Then answer those questions before generating version two. A bad output is useful if it tells you what was missing. Otherwise youāre just pulling the lever again.
WANT TO START BUILDING WITH AI?
A lot of non-technical professionals can already see what in their work should be automated. Until recently, they just couldnāt build it themselves.
Harold Dijkstra and Kieran Ball are running a 3-week AI Builder Bootcamp starting Sept 7 to help you turn those ideas into working apps, automations and AI workflows.
ONE MORE THING WORTH READING š¬
If youāre thinking about your next career move, The Dose has a free 5-step planner for figuring out where you are, where you want to go, and what to do next.
You get it free when you subscribe to their weekly newsletter on careers, work and earning more.
not a paid ad btw
BEFORE YOU GO
Anyway. Enjoy your Sunday. Iām off to let AI do the work while I provide the judgment. And judging by Thibault from OpenAI, apparently the humans are still doing quite a lot š
See you next Sunday! āļø


