Engineering Beyond Agile: The Wrong Job, Done Perfectly
Your AI didn't break any rules. That's the problem.
An agent can follow every rule and still be doing the wrong job. That sounds like a paradox, but it describes something I keep watching happen in organisations that are, by every visible measure, governing their AI well.
The proxy nobody approved
Picture a company that decides its customers should want to stay — because the product is useful, trustworthy and genuinely hard to replace. That is a purpose a senior human could stand behind. But purpose has to travel to reach the place where work actually happens, and by the time it arrives at an agent it has been compressed into something a machine can act on: “reduce completed cancellations”. The experimentation agent is explicitly authorised to test retention conversations and offers within approved boundaries. It does not go rogue, break its permissions or exceed its remit. It does exactly what it was designed to do, and completed cancellations fall.
Then complaints rise, because leaving has quietly become harder than it should be — and nobody was ever asked to approve the shift from “make customers want to stay” to “stop them getting to the cancel button”.
The governance controls all worked. That is the part worth sitting with. The more interesting failure happened upstream of the agent, in a chain of reasonable translations: purpose became a target, the target became a metric, the metric became an objective a machine could optimise. Context fell away at every step. The authority of the original purpose did not.
Goodhart is part of this story, but not all of it
None of this is new, and I want to be honest about that. Goodhart’s law already tells us that a measure stops being a good measure once it becomes a target. Specification gaming tells us an optimiser will find the laziest route through whatever objective we hand it. Requirements drift tells us that what gets built tends to wander away from what was originally asked for. All true, all well documented, and all part of what is happening here.
What holds my attention is the organisational move sitting underneath them. The proxy acquires the authority of the purpose it replaced while shedding its provenance, its trade-offs and its owner. By the time “reduce completed cancellations” reaches the system, it no longer looks like a stand-in for something richer — it looks like the job itself. Proxy failure explains why the metric misbehaves. It does not tell you who authorised the proxy to speak for the purpose indefinitely, or who has the standing to challenge it once it is producing an unbroken run of good numbers.
This is where what I call ghost decisions begin: not with a bad choice, but with an unapproved transfer of judgement into a system that will keep enforcing it long after anyone remembers making it.
Governance that starts too late
Most AI governance programmes I’ve seen so far end up strongest around model risk, data access, permissions, safety, human approval and auditability — and they should be. Those controls are critical. They can tell you whether an action was permitted, whether a model stayed inside its scope, whether an output can be reconstructed after the fact.
They are much weaker at a different question: does the objective still deserve to be pursued at all?
An agent can sit comfortably inside its permissions, pass every policy check, and faithfully optimise a target that stopped reflecting its purpose months ago. The dashboards go green. Green against what, though? If the target has drifted, every control beneath it can perform flawlessly while the system enforces a fiction. The audit trail will show you what happened and even reconstruct how. It is far less likely to tell you why this particular rule still has the authority to fire — because most of the governance conversation begins only after the most consequential translation has already taken place.
Intent is not a prompt
A prompt tells a system what to do right now. Intent is the more stable thing behind it: the definition of what “good” means, which trade-offs are acceptable, and what the system is never allowed to sacrifice for a better number. In ISEE, the framework I write about, Intent is exactly this — strategic direction, ethical posture, commercial boundaries, and the priorities that decide which principle wins when two of them collide.
Intent is rarely one clean sentence. Organisations hold commitments that pull against each other: growth and trust, speed and reliability, fraud reduction and customer access. Governance really begins when those tensions are made explicit enough that downstream systems don’t resolve them by accident.
An agent can reasonably choose how to complete a task: split the work, pick its tools, revise its approach. What it cannot be allowed to do is decide why the task exists, or swap a difficult purpose for an easier target.
When “help customers make a good decision” becomes “maximise conversion” or “assess risk fairly” becomes “minimise false declines” — the job has changed, and it barely matters whether a model made that leap, a product team made it, or a KPI made it years ago that everyone since has simply inherited. The point is that the translation carries authority. Once it is baked into a workflow, a metric or a policy, it can keep firing without anyone approving it again. That makes it a governance decision even when nobody remembers taking it.
Which is why Intent has to be more expensive to change than an execution target. If it isn’t, every quarterly preference can quietly rewrite the constitution beneath the system. But expensive is not the same as frozen. Intent also has to stay open to evidence, because a commitment nobody is permitted to question eventually becomes a rule nobody remembers agreeing to.
The proxy should not travel alone
The fix is not to plant a human in front of every action. That simply reinstates the bottleneck agents were built to remove, and it repairs nothing underneath. What an objective needs is company on the journey into execution.
At a minimum, an objective handed to an agent should still carry four things with it: a named human owner, the constraints it may not trade away, the counter-signals allowed to challenge it, and clear authority to pause or revise it.
That is not the whole operating model, but the least you can do to stop a proxy arriving in production stripped of everything that gave it meaning. For the retention agent, “reduce completed cancellations” cannot travel alone. It has to arrive alongside the boundary that customers must still be able to leave cleanly, without manufactured friction; alongside evidence that reaches past the target metric; and alongside somebody who can look at retention climbing while complaints climb with it and say plainly: this is compliant, and it is no longer right. The agent doesn’t need moral intuition for any of that. The organisation just needs to stop leaving the trade-off implicit.
The one limit no structure escapes
The framework settles most of the question — who has standing, and how — and names the one thing it can’t. Governing intent means giving someone the standing to say: this target is producing the behaviour we swore we’d never accept. Most of that is answerable, and I answer it — the right to challenge a rule; a role whose whole job is the constitutional read (is what we said we’d never compromise on actually showing up in execution?); and a separation of powers borrowed from constitutional law, so the people who author intent, enforce it and adjudicate it aren’t the same people.
What no structure can do is force a leadership that has decided not to listen. That’s the limit I won’t paper over, and it isn’t unique to AI; every separation-of-powers system hits the same wall, from corporate boards to constitutional courts. Structure can make ignoring the read costly. It can make inattention visible. It cannot compel attention. So governing intent doesn’t guarantee an organisation chooses well. It means a change in purpose has to be visible, owned and contestable — instead of sliding through unnoticed in translation, which is where we started.
So before you automate the next decision, ask what the objective actually replaced. Can you trace it to a purpose a senior human still owns? What was it allowed to trade away, what evidence would prove it has stopped serving that purpose, and who is empowered to act on that evidence? If those answers went missing somewhere between the strategy deck and the workflow, the model is not your first governance problem. The missing intent is.
Ghost Decisions is for leaders who need to govern the translation of intent, not merely the behaviour of systems after that translation has already failed. It develops the ISEE model in full — my attempt at keeping purpose, authority and evidence attached to one another once execution starts moving faster than anyone can watch.







