Most organisations have adopted AI. They've purchased licences, run training sessions, and a myriad of people now float between Claude, Copilot, ChatGPT, and/or Gemini daily.

But almost none have genuinely redesigned the work those tools are supposed to change (let alone measure the impact of the roll out).

Microsoft's 2026 Work Trend Index found that 75% of knowledge workers use AI at work, but fewer than half say it has changed how their team actually operates. The tools are there, but the transformation isn't.

This same pattern has played out before.

  • When spreadsheets arrived, accounting firms didn't redesign how audits worked. They did the same audits faster, and then did more of them.

  • When email arrived, the promise was less paperwork. What organisations got was more communication, faster, with no change to the decision-making process underneath. The meetings stayed, the approval chains stayed, and the paperwork moved from desks to inboxes.

The technology changed but the work never did, and AI is following the same script. Organisations are treating it as a productivity tool, asking "how do we do this 20% faster?" when the real question is "why are we doing this at all, and what becomes possible when the cost of execution drops by 90%?".

The ops lead in one marketing function described the gap to me: "We see ourselves as project managers, not decision makers. We force the decision out by doing a whole lot of work, rather than make the decision first and do the work that falls out of that". Five internal systems with none talking to each other, and a piece of work scoped at 50 hours turning into 150 because nobody could see what happened last time. With AI tools everywhere, and not a single workflow redesigned.

That is not a technology problem, it is a work-design problem. And the void between "we have AI tools" and "we work differently" is now visible to the board, coming with it the gap in ROI.

This is a guide for how to think about closing that gap. Not by buying more tools, but by redesigning how work gets done, how information flows, how learnings get codified, and how the function gets smarter over time.

The shape of this process is two parallel workstreams.

Workstream 1: Prove it on one team, one workflow

Do not start with the tool the team already has (the tool might be the wrong one). Start with the work, and pick a metric. Something the leadership team already cares about and will be centrally focused on day in day out. For example, campaign cycle time, brief-to-execution days, hours spent on a compliance review, cost per piece of content.

Then ask:

  • What is deficient about that metric? Trace it backwards to where the bottlenecks are and the decisions that drive that bottleneck.

  • Where are the handoffs that add wait time but no value?

  • Where does rework happen because information was not available when the work started?

  • Where does knowledge get lost between one project and the next?

Take a campaign workflow as the example, where the team picks brief-to-execution days as the metric.

  1. Tracing backwards, the brief goes to an agency before the internal team has made its own decision on direction.

  2. The agency comes back with three options, and two rounds of revisions follow because the brief was underspecified.

  3. Meanwhile, compliance review happens after creative is finished, not during it, so a third of campaigns get sent back for rework at the end.

  4. Nobody can see what happened on the last campaign that looked like this, because that information lives in a project tracker nobody queries and a shared drive nobody searches.

  5. The metric is not slow because people are slow. It is slow because the workflow forces decisions to happen in the wrong order, at the wrong time, with the wrong information.

Those are the answers that the team needs to align on the baseline.

Why one team first?

Do not try to roll this out across the whole function. Cross-team transformation hits a wall fast because different teams work from different datasets, move at different speeds, and operate under different leaders with different priorities. They have different levels of AI readiness and different execution capabilities. The alignment cost alone can stall the whole thing before any work gets redesigned.

Start with one team where the people share similar context, data sources, and jobs. The measurement is cleaner, plus the route to experimentation and learning is shorter. It is easier to see what is working and what is not when the variables are fewer.

What "redesign" actually means

This is the part most teams get wrong. They point AI at the existing workflow and ask it to do the same steps faster. Workday surveyed 3,200 employees and found that nearly 40% of the time saved by AI is lost to rework, fixing hallucinations, rewriting robotic tone, fact-checking confident output. Only 14% of employees consistently report net-positive outcomes, which is not a tool problem. That is what happens when the workflow was never redesigned around the tool's actual capabilities and limitations.

Redesign means taking the workflow you traced and rethinking it across six dimensions:

  • What work gets done, in what order, triggered by what

  • Which steps are human, which are AI, which are both

  • Where AI changes what the human does, not just how fast they do the same task. In the campaign example: instead of the team writing a brief, sending it to an agency, waiting for three options, and revising twice, a decision-package agent generates three costed options with estimated timelines before anything leaves the building. The internal team decides first, and the agency executes a decision rather than a direction-finding exercise. That is not the same workflow done faster, it is a different workflow entirely.

  • Which data sources and tools feed the new workflow, assessed from first principles rather than defaulting to what exists today

  • How the team checks in, reviews, escalates, and learns. For example: a fortnightly review where the team looks at what the agents produced, what got overridden, and what the steering documents should absorb.

  • What changes for the roles, including new ones the function has never had before

The thing is, this is where most teams stop. They redesign the workflow, ship it, and move on, but the workflow is not the operating system. The operating system is what makes the workflow get smarter over time. And that requires three things most functions don't focus on building.

Aakash Gupta recently published a detailed guide to building a Team OS in Claude Code. He studied four implementations across DoorDash, Pendo, Google, and a solo builder, and every working system converged on the same three-layer architecture. His framing maps directly to what I see in non-technical functions: shared context, shared queries, and shared routines.

Shared context: the function's memory

When a project runs 3x over its original estimate, what happens to that information? In most functions, nothing happens, the project finishes, they might run a post-mortem, and then the team moves on. The next time something similar is scoped, nobody (including an agent) can see, or point to, what happened last time.

This is the difference between using AI as a productivity tool and building it into the operating system.

  • A tool sits there until someone picks it up.

  • An operating system runs in the background and gets smarter on its own.

In Stulberg's implementation at DoorDash, the architecture is deliberate. The root file is a routing table, not an encyclopedia. Every folder gets a navigation file that tells agents what is inside and when to read it. Summaries sit at the top level and raw data lives one layer deeper, loaded only when the summary cannot answer the question. That structure means an agent can resolve a query using a fraction of its context window instead of loading everything and hoping for the best.

The same principle applies to a marketing function. An AI-native function closes the loop, because every piece of work that runs through the redesigned workflow captures what actually happened: the effort, the decisions, and the blockers. That information flows back into the system, not into a document library nobody reads, but into structured context that agents can query.

Back to the campaign example, where the team runs a product launch campaign.

  • It takes 22 days from brief to execution instead of the historical 35.

  • The system captures that the compliance review ran in parallel with creative, cutting five days.

  • It captures that the internal decision was made before the agency brief, eliminating two revision rounds.

  • It also captures that the asset library was missing approved product images, which added three days of back-and-forth with the brand team.

  • All of that is now structured, queryable context, and when the next product launch campaign starts six weeks later, the intake agent surfaces: last time took 22 days, compliance-in-parallel saved five days, and the asset gap added three.

The team starts with that information rather than from zero, and the function becomes queryable. "What happened last time we ran a campaign like this? What was the actual cycle time, and where did the rework happen?". Those questions get answered before the brief is written, not after the budget is overrun.

Shared queries: agents surface patterns before decisions are made

Once the context exists, agents can act on it. How they reach the team depends on AI readiness, and most teams graduate through three models: one person running queries and sharing outputs (hub and spoke), everyone querying the context directly, or specialised agents handling different query types automatically. The agents themselves are what matter, regardless of which model delivers them.

This is about having agents wired to the function's own data, surfacing the function's own patterns, before the humans make their calls.

Shared routines: the governance that makes it compound

This is the hardest piece and the biggest behaviour change.

When someone on the team learns something that should change how work gets done, what happens? Say a team member identifies that the compliance review step should happen in parallel with creative development, not after it. That change, if adopted, would affect every project the team runs. It would change what the agents do and what the steering documents those agents work from.

The question is: does that team member make the change themselves? Does it immediately affect their colleagues' work, or does it go through a process where someone assesses the recommendation, checks whether the rest of the team agrees, and makes a deliberate decision about whether it becomes the new standard?

That process, the governance of how best practice evolves, is the operating system's immune system. Without it, one person's good idea can break six other people's workflows. With it, the function learns and improves with every cycle, deliberately.

The governance loop can be lightweight, something like a 30-minute fortnightly review. Three questions to guide on this step:

  1. What did the agents produce that the team overrode?

  2. What override should become the new default?

  3. Who owns making that change?

That is the entire loop, and it is fast, visible, and owned. Not bureaucracy but discipline, and the difference between a function that compounds and one that drifts back to the old way within three weeks.

Workstream 2: Build the function-wide blueprint from what Workstream 1 teaches

Most organisations that try a pilot do exactly one thing: prove it works on one team. Then the pilot sits there, and nobody extracts the implications for the rest of the function. The team that ran it moves on, and the proof never scales. That pattern is so common it is almost the default.

Workstream 2 exists specifically to prevent it, and while the pilot is running, the function's senior leadership needs to work through the questions that decide how the model scales. One question at a time, informed by what the build is revealing:

  • What is the purpose of this function in three to five years?

  • What is the structural principle that decides how it should be designed?

  • Who decides what, and at which level of risk?

  • What changes for the people, and what new roles need to exist?

  • What does the function run on, which platforms, which data, and what connects to what?

  • How does work get governed when AI is doing a material share of it?

  • How is the function funded and how is value measured?

The cadence can flex, and fortnightly is the ideal; monthly works if that is what the leadership team can commit to. But someone senior must be watching what the pilot teaches and asking what it means for the rest of the function. None of those questions can be skipped.

By the final session, the answers integrate into a blueprint the leadership team has co-authored across the whole period, not a strategy document someone wrote alone.

The pilot is not a one-off

The first pilot teaches the model, because it proves that the redesign works on real work, with real people, producing measurable improvement against the metric the leadership team chose at the start.

Then it repeats with the next workflow and the next team. Each cycle is faster because the three layers, shared context, shared queries, and shared routines, are already in place from the first build. The operating system grows with each workflow it absorbs, and a pilot without a blueprint stays isolated, an interesting experiment that never scales.

A blueprint without a pilot stays theoretical, a strategy deck that collects dust. The two workstreams reinforce each other because the pilot's evidence feeds every leadership decision, and the leadership decisions inform the next phase of the build.

What this looks like when it is working

The function has a memory, because work that runs through the system captures what actually happened, and that information is available before the next similar piece of work starts. Agents surface patterns, flag risks, and generate decision-ready options before briefs are written. The team makes the decision first, and the work falls out of that, not the other way around.

When someone identifies a better way to do something, there is a clear path for that recommendation to become the new standard, or to be deliberately set aside. The governance is visible, and the function gets smarter with every project it runs.

Not because it bought better tools, but because it redesigned the work, built the context, and committed to the routines that make it compound.

Passionate about all things AI, emerging tech and start-ups, Mike is the Founder of The AI Corner.

Subscribe to The AI Corner

The fastest way to keep up with AI in New Zealand, in just 5 minutes a week. Join thousands of readers who rely on us every Monday for the latest AI news.