For two years, generative AI meant one motion: you typed a prompt, it handed back an output, you judged whether the output was any good. That motion just changed. Late this year, the biggest names in AI stopped shipping better answer machines and started shipping systems that plan, act, and complete multi-step tasks on their own. Here's what that shift actually is, and where it does and doesn't belong in a creative process.
The single turn is over
Walk the timeline of this year and the pattern is hard to miss. In May, McKinsey’s global survey found 65 percent of organizations were regularly using generative AI, nearly double the share from ten months earlier, with the tools still living almost entirely inside that one familiar loop: prompt in, output out, human decides what happens next (McKinsey). That was the baseline. Then the baseline moved.
On September 12, Salesforce introduced Agentforce, a suite of AI agents built to "autonomously reason, make decisions, and complete business tasks" by pulling relevant data, building an execution plan, and acting on it inside a set of guardrails, without a human prompting every step (Salesforce). That same day, OpenAI released o1-preview, a model built to spend more time reasoning through a problem before answering, trying different approaches and catching its own mistakes along the way rather than producing a single first-pass response (OpenAI). Six weeks later, on October 21, Gartner named agentic AI a top strategic technology trend for 2025, defining it plainly: systems that "autonomously plan and take actions to meet user-defined goals," and predicting that by 2028 this kind of autonomous decision-making will account for at least 15 percent of daily work decisions, up from close to none this year (Gartner). And on October 22, Anthropic put the clearest, most concrete version of this shift in front of the public: Claude 3.5 Sonnet gained the ability to look at a screen, move a cursor, click, and type, carrying out tasks that early partners described as running "dozens, and sometimes even hundreds, of steps" without a human directing each one (Anthropic).
That’s four separate companies, in four separate announcements, over roughly six weeks. None of them coordinated. All of them pointed the same direction. Generative AI answered a question. Agentic AI does a job.
“Generative AI answered a question. Agentic AI does a job. That’s not a bigger version of the same tool. It’s a different kind of tool.”
What "agentic" actually means, without the hype
Strip the marketing language away and an AI agent is a system that can hold a goal, break it into steps, take an action, check whether that action worked, and decide what to do next, largely without a person re-prompting it at every turn. A generative tool waits for you. An agentic one moves toward an outcome. Anthropic’s own numbers are honest about where that capability actually stands right now: on the OSWorld benchmark, which tests how well an AI can operate a real computer the way a person does, Claude 3.5 Sonnet scored 14.9 percent using screenshots alone. That nearly doubled the next-best system’s 7.8 percent, and it’s still a system that gets the task right well under one time in five (Anthropic). Anthropic calls the capability "imperfect" and "at times cumbersome and error-prone," and recommends starting with low-risk tasks while it matures. That’s not a caveat we’re adding for comfort. It’s the company that built the thing telling you where it actually stands.
That gap between the headline and the benchmark matters more than either one alone. The shift is real. The capability is early. Both of those things are true at the same time, and a creative business that only hears the first one is going to make decisions it regrets.
What this could mean for a creative workflow, concretely
Here’s where it gets useful instead of theoretical. An agentic tool is well suited to the parts of creative work that are procedural: pulling competitor research into a working doc, drafting a first-pass content calendar against a brief, restructuring a spreadsheet of campaign performance into a summary someone can actually read, running a batch of image resizes or file renames across a project folder, checking a site for broken links or missing alt text before launch. Those are real hours. Handing them to a system that can execute a plan, not just answer one question, gives our team more of the day back for the work that actually needed a person in the first place.
What an agentic tool is not suited to, no matter how many steps it can chain together, is creative judgment. Whether a concept is actually good. Whether a headline earns the reader’s attention or just fills the space. Whether a brand’s tone lands as confident or as arrogant. Whether a client’s real problem is the one they described in the brief or something underneath it that took a conversation to surface. Those calls require taste, context, and a stake in the outcome that a system optimizing toward a goal doesn’t have and isn’t built to have. An agent can execute a plan. It cannot decide the plan was the wrong one to begin with, not the way a person who has sat across the table from a founder can.
The Picasso Standard, applied to a system that acts
We’ve held one line since AI first showed up in this practice: the tool was never the problem, misrepresenting how the work got made is the problem. Agentic AI doesn’t change that line. It raises the stakes on it. When a tool only answers a question, the human decides what to do with the answer at every turn. When a tool can execute a chain of steps on its own, the discipline of staying honest about what it did and what a person decided has to be built into the workflow itself, not left to memory after the fact.
That’s the same standard we apply to a single AI-drafted headline, scaled to a system that can now touch dozens of steps unsupervised. AI sits at our table. It doesn’t run the haus. That was true when the tools could only generate. It stays true now that some of them can act.
“AI sits at our table. It doesn’t run the haus. That was true when the tools could only generate. It stays true now that some of them can act.”
Where we go from here
We’re not adopting agentic tools because the word is trending. We’re adopting the parts of it that give our team back hours currently spent on work that never needed a human’s taste to begin with, and we’re leaving the judgment calls exactly where they’ve always lived: with the people in the room. The gap between a benchmark score in the teens and a headline that says "AI agents are here" is a gap worth sitting inside for a while before anyone hands a client’s brand over to a system that can act but can’t yet tell good from mediocre.
The defiant don’t chase every new tool because it’s loud. We test what earns its place, keep what does the work honestly, and cut what doesn’t. That filter doesn’t change because the tool got more capable. If anything, it matters more.
