- 100 School
- Posts
- đź’Ż Should you trust ChatGPT with your entire workflow?
đź’Ż Should you trust ChatGPT with your entire workflow?
GPT Live got weirdly human, GPT-5.6 has “that dog in it,” and OpenAI’s new agent can work across your files and apps for hours. Here’s what the launch demos skipped.
Husk is a guy who built a small corner of the internet out of asking AI voice models questions that are painfully easy, then waiting to see how confidently they fall apart. This week, he of course tried GPT Live.
The video landed at the end of a fairly chaotic week, even by AI standards.
Anthropic spent the first half extending access to Fable 5 while everyone tried to squeeze in a few more builds before hitting the limits. I wrote about the 5 non-coding things worth doing with it and you now have till July 19th (allegedly) to play with it.
Then OpenAI took over.
Within two days, it released GPT-5.6, ChatGPT Work, and GPT Live, its new voice model. Depending on who you read, this was either the moment agents finally became normal work software, the best voice product anyone has shipped, or an impressive collection of names and settings nobody fully understands yet.
Peter Yang tested the releases for a day and gave GPT-5.6 the most useful review I saw: “It’s got that dog in it”.
It keeps going. It stays with difficult work. And it seems less likely to hit a complicated task and decide you probably didn’t need it done anyway. But Peter’s praise came with a pretty long list of confusion too:
What exactly is Work versus Codex? Why are some things chats and others tasks? When should you use Sol, Terra or Luna? Why are there several effort settings before a normal person has even started the job?
Right now the pieces are impressive and still a little scattered. OpenAI may have made working with agents feel more mainstream this week but hasn’t fully made them feel simple yet. So let’s try to make sense of it today.
Window into the Future
ChatGPT Work is where this gets much bigger than the new voice model.
You can hand it a project, give it access to the files and apps you choose, and leave it working across the messy middle: research, analysis, documents, spreadsheets, presentations, websites. OpenAI says it can stay with complicated jobs for hours, splitting them into smaller steps and completing them without waiting for you after every click.
It’s basically the useful parts of Codex brought into the ChatGPT product almost everyone already understands well enough to open.
For non-technical people, that matters more than a benchmark score. You dont need to know how the agent is assembled. You can hand it the folder.

There are loads of examples in OpenAI’s launch post: complete slide decks, financial models, documents that follow an existing company format, interfaces it can inspect and refine instead of generating once and abandoning.
That part looks excellent.
But the most useful criticism I saw came from Ethan Mollick a few weeks ago, then again after testing ChatGPT Work this week.
His point is that these products still think like software tools.
In software, the codebase can serve as a source of truth. You can test whether something works. Changes are tracked. Failed branches, old versions and decisions are usually recoverable if someone cares enough to look.
Knowledge work is less cooperative. A polished report might contain 20 sources. It probably won’t tell you which two sources caused the author to change their mind, or that one important number came from a spreadsheet somebody described as “rough but basically right” in Slack.
A lot of the work is hiding around the artifact. Which means an agent can produce something that looks finished while missing the part that made the work make sense.
I keep thinking about my own newsletter archive.
There are dozens of published issues in there, plus screenshots, links, messy drafts, subject lines I dropped because they were repetitive or just a bit boring. If I ask ChatGPT Work to analyse it, I dont need a summary telling me we write about AI and the future of work.
I’d want it to tell me where an argument first appeared, how it changed across several issues, which sources kept coming back, whether I started repeating a sentence shape too often, and where it cannot know the reason because that reason only exists in somebody’s head or in an old Slack thread.
That’s closer to actual knowledge work.
The models can now work for hours. They can inspect more files than I could read in a week, keep several routes open at once, build the deck, fix the deck, then build a tiny website for the deck because apparently we’re doing that too.
I still have the same brain I had last Sunday.
The launch videos tend to end when the work is done but normal work doesn’t. And that’s a very different problem from “can the model do the task?” It can.
Can the rest of us keep up with what it did, understand enough of the route, and spend our attention in the places where a quick approval could turn into a very expensive afternoon? That’s the part I think teams will run into next.
There will simply be far more finished-looking work arriving than anyone is used to reviewing. We already have this with AI-written documents. Agents can multiply that asymmetry.
There’s a version of this that works brilliantly, obviously. The agent does the slow collection and construction, then returns the small number of decisions that actually need you. That’s what the good examples are moving toward.
But I want the receipt attached. So the first thing I’m trying with Work is a very unglamorous instruction at the end of the task:
Alongside the final output, give me a one-page work log showing:
- the sources you relied on most
- anything you ignored or could not access
- the assumptions you made
- decisions you made without asking me
- anything that still needs a human to checkI just want to know where to look before I approve something built from 70 files and three years of context.
This is also a lot of what we’re working through with teams at 100 School. Once the tools can do more, somebody still has to agree what needs checking, what needs recording and what “finished” means. Because once the job gets more complicated than counting the e’s in seventeen, “looks finished” isn’t enough.
How to AI: GPT-5.6 + Work Edition 🤖
Every week, this section is your shortcut. Here are a couple of ways you could use GPT-5.6 at work this week that are worth your time:
Before you go ✌️
Be honest: when AI hands you something polished, how closely do you check it? |

See you next week!
P.S. Want to make your team & company AI-first? Let us help here.

