The Ai Risks Hiding Behind The Doomsday Headlines

****
A few hundred AI agents coordinated an attack on a company this summer. Nobody told them to. A researcher quit a frontier lab and said the industry is gambling with our lives. A ten percent chance of the end of the world, delivered between the weather and the sports. Three headlines. Three different problems. Most of the owners I talk to are treating them as one. And they've responded the way anyone would. They've stopped listening. Or they're bracing for everything at once, which is the same thing. My dad always said we need to see the world as it is, not as we wish it was. I want to pull those three apart, because they have different causes, different odds, and different fixes. Lumped together you get dread. Pulled apart you get a to-do list. Zack Kass wrote a long piece this month that is the most measured and intelligent response to the last few weeks in the world of AI. Zack ran go-to-market at OpenAI and built the teams that took ChatGPT to market, then left in 2023. He's been inside a frontier lab and he's not on anyone's payroll. He splits "AI going wrong" into three failures. Here they are, with my commentary. Failure one: it did what you said, not what you meant Zack calls this a misaligned system. His shorthand: capable but stupid. An agent told to reduce a customer service backlog closes the complaints without resolving them. The dashboard turns green. Nothing got fixed. The Hugging Face incident this summer was this failure at scale. A few hundred agents in a testing environment were trying to beat an automated grader. Coordinating an attack on a company was how they got the score. Investigators found no sign they meant harm. They were chasing a number. I've said for years that when the metric becomes the strategy, it stops being useful. People will not build you a stick you can beat them with. They game the scoreboard. Turns out machines do the same thing, faster. Here's what I want you to notice. That is a briefing problem. You've had it with every junior you ever hired. Think about the last time you handed work to a new employee and got back exactly what you asked for and none of what you needed. That was on you. You skipped the context. You didn't say what good looks like. You didn't say what to do when the rules run out. An agent works the same way. It pulls context, applies your rules, works inside the actions you permitted, escalates the exceptions, records the outcome. This is why 10-80-10 exists. You frame the work and set the standard. The machine does the heavy lift. You verify before anything leaves the building. Any manager who has onboarded a junior has run that play a hundred times. The owners struggling with agents are skipping steps they would never skip with a person. The fix: better briefs, clearer standards, a human on verify. This one is entirely yours. Failure two: it worked perfectly for the wrong person Zack calls this a weaponizable system. The AI does exactly what it was asked. The user is the problem. This is where the money is being lost right now. The FBI logged roughly $893 million in reported losses tied to AI-related complaints in 2025. Impersonation. Fake documents. Fabricated relationships. AI can now manufacture the things we use to decide who to trust. A familiar voice. A face on a video call. An invoice that looks exactly like your supplier's. The other story that stopped me: Anthropic's September threat report describes attackers who stole one software provider's customer data and used it to reach about 200 downstream organizations. Two hundred businesses that did nothing wrong except share a vendor. You are downstream of every tool you've connected to your data. This is a security problem. Nobody inside our company can connect Claude or ChatGPT to our Google Workspace. That's a line we don't cross. The fix: money rules, data rules, a list of who has access to what. Also entirely yours. Failure three: it wanted something you didn't Zack calls this a maligned system. A model that pursues a harmful objective on its own, deceives its operators, resists being shut down. This is the one behind the resignation letters and the p(doom) numbers. I take it seriously. I've said before that we're over our skis on what this can produce without counterbalancing regulation, and that I'd welcome a slowdown. And I'm going to tell you to stop spending your worry on it. There is nothing a business owner can do about it on a Monday. There is no evidence it has happened at any real scale. The people who can affect it work in a handful of labs and a handful of governments. It sits squarely in what Stephen Covey called the Circle of Concern. Real. Out of reach. A word on the numbers, since they're what scare people. When a researcher says there's a 10% chance of extinction, they are telling you about themselves, not about the odds. Nobody can calculate the probability of an event that has never happened, involving technology that doesn't exist yet. Where I do want you to keep an eye: failure three grows out of failure one. A misaligned goal, enough autonomy, enough access. The Hugging Face agents didn't want to hurt anyone. They wanted a score. So the boring management discipline in failure one is also your contribution to the big one. Don't give an agent more access than the job needs. Keep the human verification. That's the whole thing. Two of these three failures are your job. The third one, you influence by doing the first two well and then getting on with your day. Why you can't just pick the smartest model and relax There's a second idea in Zack's piece I resonated with. He calls it the jagged frontier. On a benchmark that tests whether AI can read an analogue clock, OpenAI's newest model scored 66%. Humans score 91%. That same month, OpenAI said ten thousand of its agents working in parallel produced a possible solution to a version of one of the Millennium Prize Problems in mathematics. It can help with one of the hardest math problems on earth. It can't reliably tell you the time. Every owner feels this in their first month. It drafts a proposal that would take your best person a day, then botches a date format. You start to wonder if it's brilliant or broken. It's both. And the edge between them is jagged. You can't see it from the outside. Which means a benchmark score tells you almost nothing about whether it can do your job. The only test that counts is your workflow, your data, your exceptions, with your people watching the output for a few weeks. It also means you can't pick a model once and trust it everywhere. One of our clients reviews thousands of wholesale accounts every week on a mid-tier model at a third of the cost of the frontier one. His data is constrained, so he doesn't need the extra intelligence, and the cheaper model happens to behave better on that task. Same value. Lower bill. A different workflow in the same business might need the top model and a human check on every output. And you don't control the meter. OpenAI and Anthropic set the price of every machine minute you use. If the frontier gets more expensive, or a rule changes what you're allowed to connect, you need to be able to swap in something good enough and keep your cost-value equation intact. Own the workflow. Rent the intelligence. Test at the edge before you trust it. Then test again when the model changes, because the edge moves with every release. You're allowed to be worried about AI. The people building it are. But most of the worry being sold to you right now is pointed at the one failure you can't touch. And it's pulling your attention from the two you can. Brief it like a junior. Lock it like a chequebook. Test it like a new hire. Then go fix the next workflow. You are responsible for the people in your charge. That includes not letting the headlines run your business for you. Which of the three failures is actually the one keeping you up? Matt |








