← The AI Leadership Brief

Your AI Is Only as Good as the Judgment You Give It

The gap between organisations running AI at scale and those stuck in pilot purgatory is not what most leaders think it is. It is not compute power, model access, or technical sophistication. By mid-2026, the models themselves have been commoditised — every serious organisation has access to roughly equivalent AI infrastructure. The divide is something older and harder: the ability to make judgment explicit.

That is the argument from a recent HBR piece by researchers at Harvard Business School's AI Institute, and it is one every senior leader needs to sit with. Because the bottleneck is not the technology. It is you — or more precisely, the accumulated judgment inside your organisation that has never been written down.

The problem nobody anticipated

For decades, organisations transmitted expertise informally. New hires watched, listened, and gradually absorbed how the organisation thought. They sat next to senior colleagues. They observed how edge cases were handled. They learned — over years — the unspoken logic of when to bend a rule, when to escalate, when to take the hit and keep a customer, and when to hold the line.

That model worked because humans were doing the executing. AI agents have changed the equation entirely.

Unlike traditional software, AI agents can operate in ambiguous environments and make decisions in real time. But unlike people, they cannot absorb norms through observation or infer context from organisational culture. They operate based on what is made explicit. Nothing more.

This produces a specific failure mode that is now showing up across sectors. Companies deploy customer-facing AI agents without first codifying how their best people make decisions — how to handle a pricing exception, a frustrated long-term customer, or a request that sits just outside policy. The agent eventually goes off track because no one ever wrote down how that particular decision actually gets made.

The cost of that misalignment compounds. Every interaction where the agent applies the wrong logic, takes the wrong tone, or escalates when it should not is not just a failed transaction. It is evidence that the organisation handed over execution before it had done the harder work of encoding how it actually thinks.

What codified judgment actually means

The HBR authors call this challenge "judgment infrastructure" — and it is worth being precise about what that means in practice, because it is easy to mistake for documentation.

Documentation captures what happened. Codified judgment captures how decisions get made — the risk tolerance, the escalation thresholds, the brand voice under pressure, the exception logic that your best people apply instinctively and would struggle to articulate if asked directly.

Codifying judgment means translating the tacit decision-making principles of your organisation into structured guidance that agents can execute. These principles include risk tolerance, brand voice, escalation thresholds, quality standards, and the subtle logic of exception handling. Historically, these things lived in the minds of experienced people. For AI to work at scale, they need to live somewhere more durable.

The distinction matters because most organisations attempt this the wrong way. They ask experienced people to write down what they know. That approach rarely produces anything useful. Experts are notoriously poor at articulating tacit knowledge in the abstract, because they know far more than they can say when asked directly to document it.

A more effective method: do not ask them to document their judgment. Create conditions where it surfaces naturally. Convene a small panel of experienced practitioners in the same role. Bring in a skilled moderator and walk the group through a series of realistic scenarios and actual edge cases. Where the panel agrees quickly, you have clear policy. Where they disagree, you have judgment worth capturing. The transcript of that conversation becomes your first draft of codified judgment.

A claims team at an insurance firm might surface more nuance about risk tolerance, customer empathy, and escalation logic in a single two-hour session than years of documented procedures ever captured — because debate externalises reasoning in a way that documentation never does.

Try this promptPro
Codify Judgment and Decision Making into AI Agents
Open in library →

Three shifts leaders need to make

The organisations pulling ahead have made structural changes, not just tactical ones. There are three shifts worth understanding.

The first is governance. Defining acceptable risk boundaries, setting performance expectations for agents, managing how they are onboarded and offboarded — these are organisational questions as much as technical ones. And critically, they cannot be outsourced. Business units, HR, and IT need to govern together. Treating agents like software licences rather than operational contributors is how organisations end up with agents that are technically functional but strategically misaligned.

ITA Group, a global events company, learned this through an early air-travel booking agent. The hard part was not building the agent. It was defining what the agent needed to know to be trusted: when to optimise for cost, when traveller experience mattered more, which exceptions were acceptable, and when a human needed to intervene. The fix required changing the operating model so that expert users — not just technologists — could shape agent behaviour directly.

The second shift is in what managers actually do. The HBR authors describe the emerging role as "judgment architect" — and it represents a meaningful change in what valuable management looks like.

Debbie Riazzi, director of compliance and labour relations at AWP Safety, a one-person department, has built a portfolio of agents, each codifying a different slice of her expertise. One handles medical accommodation requests by pulling the relevant job description, surfacing how comparable requests were resolved, and running through a standardised intake she has refined over years. Another handles the opening moves of every information request the company receives — parsing what is being asked, routing to the right owner, and drafting the response. Her judgment, applied consistently at scale.

Nathan Mapp, a controller at a global venture-capital firm, codified more than a dozen years of finance expertise into a series of markdown files that his agents can reference in real time. A team of two now covers ground that would previously have required ten.

What Riazzi and Mapp have done is not automate their work. They have made their expertise portable — deployable across every task their agents touch, at a consistency no team of humans could match.

The third shift is in what the highest-performing employees look like. The traditional divide between strategic thinkers and operational doers is collapsing. High-performing employees increasingly are both — they reason strategically and operationalise their thinking through agents. They design workflows, encode judgment, build and iterate on systems, and continuously move up the value chain.

Ramp, a financial platform used by 30,000 companies, trains every employee during onboarding to build their own AI tools rather than act as a button-pusher on systems someone else created. The organisation is systematically building people who can shape how AI executes on their behalf — not just use it.

Why this compounds

The strategic case for investing in judgment infrastructure is not just about the immediate productivity gain. It is about compounding advantage.

When judgment is successfully codified, expertise becomes portable. Best practices are no longer locked inside the most senior people. Institutional knowledge can be deployed across functions, geographies, and products at scale. The organisation that has done this work can move faster on the next deployment because it has already built the trust, governance, and operating rhythm to repeat the pattern.

ITA Group's trajectory shows what this looks like in practice. The first six to seven months were slow because the company was learning how to translate judgment from expert employees' heads into the context files agents use. But once that operating model took hold, the pace changed — especially in software development. Timelines shrank from months to weeks, and the habit of using agents to rethink the work itself began spreading into other functions.

The first use cases are slow. The next ones move faster. The organisation that starts now builds a structural advantage over one that waits until the technology "matures" — because the technology is not the constraint. The constraint is the work of making judgment explicit, and that work takes time regardless of when you start.

AI Agent Development: Initial Six Months vs. Subsequent Deployment — ITA Group's learning curve shows dramatic acceleration after codifying judgment

What this means for senior leaders

The first phase of AI adoption was about who had access to the best models. That phase is largely over, because access has been commoditised. The next phase will be defined by who has done the harder work of encoding how they actually think and work.

For a CEO or board member, the question is not whether your organisation is using AI. It is whether your organisation is using AI in a way that encodes your competitive logic — your risk appetite, your customer standards, your quality thresholds, your exception handling — or whether you are running generic models on generic assumptions.

The difference between those two states is not the technology. It is judgment infrastructure. And building it starts not with a procurement decision but with a harder question: can your organisation make explicit how your best people actually decide?

Most cannot yet. The ones that develop that capability first will not just work faster. They will compound that advantage across every deployment that follows.

Put it into practice
from day one.

Marcus is a curated AI prompt library built for C-suite leaders. Every prompt is structured, role-specific, and ready to use — so you spend less time prompting and more time deciding.

Explore the library →Get access