Most people open ChatGPT. Type three paragraphs or upload files of context explaining the task, the audience, the tone, and the background, then wait for the output. Read it, decide it's 70% right, and manually rewrite the parts that missed, then close the tab.
Tomorrow, they do the same thing. Re-explain the same context, upload the same files, re-describe the same audience, get roughly the same 70% result, and manually fix it again.
That is how most people interact with AI right now, and most believe that's what "using AI" looks like.
Now apply an employee test to that same workflow.
Imagine a team member who arrives every morning with no memory of yesterday, no knowledge of the company's standards, no understanding of what was discussed last week. Every task starts from scratch. Every output needs manual review before it goes anywhere. The manager dictates the brief in full every single time, then edits the work personally before it can be used.
No leader would accept that from an employee. Nobody would call that delegation, and nobody would expect it to scale. But it describes the default AI experience in most organisations today.
That is not managing AI; that is clearly micromanaging it, and it's a significant barrier to unlocking value from the technology.
The whitepaper that quantified the gap
A whitepaper from Microsoft and Founders Forum Group surveyed 200+ EMEA tech founders and interviewed 11 leaders from companies including Mistral ($14 billion valuation), Lovable, Faculty, and Tessl. The data tells a clear story:
1 in 2 founders are developing AI tools in-house
But 1 in 3 remain in exploration mode
Data integration and tool selection are the biggest barriers
1 in 3 still lack trust in their own data infrastructure
The investment is happening. But, crucially, the structural shift mostly isn't.
Guy Podjarny, founder of Tessl (previously scaled Snyk to a multi-billion dollar security company), found that 3 kb of precise agent instructions outperformed 20 kb of broad ones. More information made the agents worse = not better (more on this article here). Precision beats volume, every time. The same lesson every experienced manager already knows: a clear brief with defined boundaries outperforms a 40-page document nobody reads.
The organisations breaking through aren't the ones with better models or bigger budgets, but the ones that learned to manage AI the way good managers manage people: with clear context, defined standards, and enough autonomy to let the work compound. And by creating an environment for AI to complete workflows before coming back to the human for feedback and approval. Not for employees to check back in with their manager on every email they wrote.
What management actually looks like (and what most teams are doing instead)
Nuha Hashem, founder of Cozmo AI, built what she calls an Agent Protocol Engine: "every AI employee runs on explicit policies, tools, and feedback loops instead of vibes and prompts".
Vibes and prompts vs policies and feedback loops. That distinction (and it is a massive one) is the gap between the ChatGPT-tab-open workflow described above and an AI operation that actually scales.
Podjarny frames it with a specific test: "When a problem comes up, do the people dive into the code and fix it themselves, or do they modify the agent's behaviour so it overcomes the problem on its own?". Apply that approach to any function (not just coding), and most teams are still doing the first, fixing the output manually. The AI-native organisations are doing the second, fixing the system so the problem doesn't recur.
The first is using a tool. The second is managing one.
In practice, that gap shows up everywhere. Someone opens ChatGPT, types three paragraphs of context, gets a 70% output, and manually fixes the rest. A manager reviews every AI output before it can be used. A developer asks Copilot to autocomplete a line of code. All of these feel productive, and all of them reset to zero tomorrow.
The management version looks completely different.
An agent with persistent context, voice guidelines, and quality checks produces 90%+ work without re-explaining anything. How this looks in practice:
Validation rules catch failures before the output is ever presented.
A structured spec gets delegated to an agent that delivers the entire feature.
The difference is not effort, it's compounding: every improvement carries forward instead of resetting overnight.
Most teams sit firmly in the first pattern (and honestly believe they're "doing AI"). The second is where the returns actually build and work feels different.
Why most organisations stay stuck
The whitepaper data explains the structural reason.
60% of leaders and founders are focused on upskilling current teams, compared with 30% hiring AI specialists.
1 in 4 are redesigning existing roles for hybrid human-AI collaboration.
67% are using AI in a leadership context to guide faster decision-making.
The intention is there, but only 17% are building cross-industry ecosystem partnerships to scale AI capability, and while 45% of founders say Responsible AI is a core strategic priority, 1 in 3 have no formal governance practice at all.
That governance gap is the reason most teams end up micromanaging. Without clear policies, quality standards, and defined boundaries for what AI can and cannot do autonomously, every single interaction requires manual oversight by default. The organisation doesn't choose to micromanage. It just never built the management infrastructure that would allow anything else.
Angie Ma from Faculty describes the pattern: most organisations end up with "a meadow of wildflowers: beautiful and colourful, but not connected and not transformative". The ones that succeed "intentionally create curated gardens, prune ruthlessly, and identify a handful of AI capabilities that genuinely move KPIs".
That pruning and curation is management. The wildflower meadow is what happens when everyone opens their own ChatGPT tab, prompts independently, and nobody orchestrates (which is most organisations right now).
Arthur Mensch, co-founder of Mistral, reinforces the discipline: measure "real-world outcomes over benchmark scores" and embed "observability, auditability, and versioning from day one". Clear metrics, visibility into how work gets done, and systems that improve with each cycle. The same standards any good manager applies to a high-performing team.
Three shifts worth making
1. Audit the management style honestly. How much AI interaction across the organisation is the open-ChatGPT-retype-everything-fix-it-manually loop? That number is the micromanagement score, and most leaders would be uncomfortable with the answer.
2. Build onboarding for AI the same way it exists for people. Context files, defined responsibilities, quality standards, clear boundaries on which decisions can be made autonomously and which need human sign-off. If none of those exist, the organisation is re-onboarding its most capable worker every single morning.
3. Change the problem-solving reflex. When an AI agent produces a bad output, the instinct in most teams is to fix the output manually. The management reflex is different: change the agent's operating context so the problem doesn't recur. One fixes a moment, the other fixes the system. Podjarny's 3KB finding proves the point: less information structured precisely beats more information delivered loosely.
Oana Jinga from Dexory put the urgency bluntly: "If an organisation treats AI like old-school procurement with two-year RFPs, it will be the biggest mistake". The same applies to management style. Treating AI agents the way nobody would treat a competent employee (dictating every task, reviewing every output, starting from scratch each morning) is the expensive path to staying exactly where most organisations already are.
The models aren't the bottleneck. The understanding of how AI works and the management layer around them is.

Passionate about all things AI, emerging tech and start-ups, Mike is the Founder of The AI Corner.
Subscribe to The AI Corner
