A year ago we wrote that generative AI answered a question and agentic AI does a job, and we put numbers next to that claim so we could come back and check our own math. This is that check. What actually panned out, what surprised us, and where the line between a tool and a judgment call held over twelve months of real use.
The receipts, one year out
Last November, we published "AI agents are here," built around a benchmark that was honest about its own limits. Anthropic’s Claude 3.5 Sonnet had just gained the ability to operate a computer directly, and on the OSWorld benchmark, which tests whether an AI can complete real computer tasks the way a person does, it scored 14.9 percent. Best in the field at the time. Still wrong more than four times out of five. We said the shift was real and the capability was early, and that both of those things could be true at once. We didn’t know yet how fast the second half of that sentence would change.
It changed fast. On September 29, 2025, Anthropic released Claude Sonnet 4.5, and its OSWorld score came in at 61.4 percent, up from Claude Sonnet 4’s 42.2 percent earlier in the year (Anthropic). Run the full year: 14.9 percent to 42.2 percent to 61.4 percent. That’s not incremental improvement. That’s a system going from failing most of the time to succeeding most of the time inside twelve months, and Anthropic reports the newer model holding focus on complex, multi-step tasks for more than 30 hours at a stretch.
Adoption moved too, though not the way the loudest headlines predicted. McKinsey’s state of AI survey, published November 5, 2025, found 62 percent of organizations now experimenting with or scaling agentic AI systems, with 23 percent reporting they’re scaling an agent somewhere in the business (McKinsey). Read the next line before you get excited: no more than 10 percent of respondents report scaling agents within any single business function. Wide, shallow experimentation. Not the autonomous workforce the coverage promised. We called that gap a year ago. It’s still there, just measured now instead of guessed at.
What actually surprised us
We expected the capability curve to keep climbing. We didn’t expect this many organizations to walk projects back out the door. On June 25, 2025, Gartner predicted that more than 40 percent of agentic AI projects will be canceled by the end of 2027, and named the reasons plainly: escalating costs, unclear business value, and inadequate risk controls (Gartner). That’s not a skeptic’s guess. That’s the same analyst firm that told us a year ago agentic AI would drive 15 percent of daily work decisions by 2028, now telling us the market rushed the deployment part harder than the capability part could support.
“Most agentic AI projects right now are early stage experiments or proof of concepts that are mostly driven by hype and are often misapplied.”
Anushree Verma, Senior Director Analyst, Gartner
Gartner has a name for part of the problem: agent washing, vendors rebranding existing automation as an agent without the underlying capability to back it up. We saw versions of that pitch land in our own inbox this year, tools promising autonomous execution that turned out to be a chatbot with a new label. The lesson wasn’t that agentic AI failed to arrive. It arrived exactly on the schedule the benchmarks predicted. The lesson was that a company handing a client’s brand to a system just because the system now claims the word "agent" was always going to end up in that 40 percent.
What we actually handed to agents this year
Here’s the honest accounting, not the theoretical one we published in 2024. Over the past twelve months, agentic tools earned a real place inside a handful of our own procedural work. Research pulls that used to eat an afternoon, competitor scans, source-gathering for a piece like this one, now come back structured in minutes, with a person still verifying every fact before it goes anywhere near a client. First-pass content calendars against a brief get drafted by an agent and reshaped by a strategist before anyone sees them. Pre-launch QA, checking a site for broken links, missing alt text, or heading-hierarchy gaps, now runs as an automated sweep instead of a manual checklist. Project administrivia, meeting recaps, status summaries, restructuring a messy spreadsheet of campaign numbers into something a client can actually read, moved to agentic tools almost entirely.
That’s real time back. It’s also a narrow list on purpose. None of it touched what a client is paying us for.
Where the line never moved
No agent picked a concept this year. No agent decided whether a headline earned the reader’s attention or just filled the space. No agent sat across a table from a founder and heard the thing they didn’t put in the brief. That work stayed exactly where it lived in 2024, with people who have a stake in the outcome, and nothing this year gave us a reason to move it.
Regulation caught up to that instinct faster than we expected. On August 2, 2025, the governance rules and obligations for general-purpose AI models under the EU AI Act became applicable, the first real compliance deadline tied to how these systems get built and disclosed (European Commission). We don’t have EU clients bound by it yet, but watching a formal governance layer arrive this quickly, less than a year after the tools it governs went agentic, told us the caution in our original piece wasn’t overcorrection. It was pointed the right direction.
“The tools got faster this year. The judgment calls didn’t get any easier to hand off, and we’re not looking for them to.”
HAUS XXIV
One year in, the filter holds
We could have written a piece that just declared victory. The benchmark numbers would have let us. Instead we’re publishing the part that’s less comfortable: the market moved faster on capability than it did on sense, a lot of agentic projects are going to get quietly shut down over the next two years, and the discipline of checking what a system actually did before you claim credit for it matters more now than it did when we first wrote about this.
That’s the same filter we started with. Test what earns its place. Keep what does the work honestly. Cut what doesn’t, no matter how good the demo looked. A year of watching this shift up close didn’t change that filter. It sharpened it.
