← Back to Blog

Can AI Coding Agents Replace Developers?

My answer is no. Here's why, and why the market is already correcting.

The bottleneck isn't capability. It's context. And we've seen this movie before.

The context gap: four pillars labelled What, How, Why and Where hold up a beam marked production-ready code. The What and How pillars stand; the Why and Where pillars are crumbling and the beam is falling.

Two pillars stand. Two don't, and that's where production code falls.

The pattern in front of us

Working in agentic coding infrastructure, I've watched the same pattern enough times to call it.

An agent writes plausible code. The tests pass. The developer ships.

A week later, the code is gone: rewritten, refactored, or quietly deleted.

Big Tech is catching on. The companies that moved first on "AI replaces developers" are rehiring. Roughly a third of US hiring managers who cut roles after adopting AI have added them back (Robert Half). 55% of employers who cut headcount citing AI regret it (Forrester). Half of the companies that cut customer service staff for AI will rehire by 2027 (Gartner).

The reason shows up in production: 0% of the senior engineering leaders Lightrun surveyed are "very confident" AI code will behave correctly once deployed (Lightrun, 2026).

The narrative said AI was replacing developers. The data says:

We don't trust it to ship.

The question isn't whether AI can write code. It can.

The question is whether AI IDE coding agents can replace developers.

The answer, today, is no. And not for reasons that will be fixed by a better model.

Two staircases

Even the data above deserves a caveat.

Most studies of AI-generated code lump everything together: Copilot autocompletes, Cursor edits, full agent runs. That data covers the easy case.

The replacement narrative is about the hard case.

Assisted coding, meaning tab completions, inline edits, chat-based suggestions:

  • Single-shot prediction
  • Human verifies every output
  • Cost of error: roughly zero

Agentic coding, meaning Claude Code, Cursor, Codex, Antigravity:

  • Long-horizon decisions
  • Agent plans, executes, iterates autonomously
  • May run for 30 minutes and tens of thousands of tokens before anyone sees the result
  • Cost of error: compounds

These are parallel staircases, not the same one.

Assisted coding's bottleneck is model capability, and it's improving fast. Agentic coding's bottleneck is context, environment access, verification, planning, and intent, almost none of which is a model problem.

Better autocomplete does not compound into better agents, any more than a faster horse compounds into a car.

When we ask whether AI IDE coding agents can replace developers, we have to be specific about which AI we mean.

The agentic kind is the only kind that could plausibly do the job end-to-end. It's also the one that isn't ready.

The bottleneck isn't model capability. It's context.

The four pillars of context

We've seen this movie before.

Excel didn't change finance because the spreadsheet got better. It changed finance when Bloomberg and FactSet wired data feeds into it.

EC2 didn't change enterprise computing when AWS launched it. It changed enterprise computing when DirectConnect, VPC peering, and private link gave it a real integration story.

The capability existed for years before the utility arrived.

What closed the gap was the plumbing.

Capability is the demo. Plumbing is the production. The two phases never collapse into one.

Agentic coding is in the same phase. The agents work. Runtime access and test execution, the basic plumbing to the environment, are mostly there in the big four platforms (Claude Code, Cursor, Codex, Google's Antigravity).

What's still missing is the harder plumbing: context.

A developer holds four kinds of context every time they ship code. Most prompts contain a thin slice of one. Agents see less.

Call them the four pillars of context:

  • What. What needs to exist? The spec. The behavior. The success criteria.
  • Why. Why are we building it? The purpose. The user. The decisions already made. The alternatives already rejected.
  • How. How should it be built? The patterns. The conventions. The way this team does things.
  • Where. Where does it fit in production? The dependencies. The traffic patterns. The rate-limit pools. The blast radius.

They are not equally missing today.

What: partial. The prompt delivers a slice. Usually thin. Always incomplete.

How: partial. The codebase reveals patterns. The conventions and the reasoning behind them stay hidden.

Why: barely there. Decision history doesn't live in any file the agent reads. The PM's rejected alternatives, the user research, the strategy memo: none of it is in the repo.

Where: barely there. Production topology lives outside the codebase. Traffic patterns, shared rate limits, the brittle downstream service that breaks under sync calls: invisible to the agent.

The prediction worth holding onto: the four pillars won't fill in equally.

Agents will produce code that compiles, runs, and passes, and is still wrong, because it solved a problem nobody briefed the agent about, in a place the agent didn't know it lived.

The model won't have failed. The reviewer will.

That is the pattern teams describe: code that passes initial AI-assisted review and fails 30 to 90 days later in production.

That's not a quality problem. That's a context problem with a delay fuse.

Context comes with mileage

The four pillars aren't filled by the environment alone. Some of the context comes from the human at the keyboard.

And it shows in the prompt.

Engineers who have lived through the full arc of an application (inception, build, deploy, on-call, postmortem, decommission) carry that lived context into every prompt they write. They have seen what breaks at 3 AM. They have written the apology email. They know what the agent doesn't.

Their prompts have production ramifications baked in:

  • Idempotency, without thinking about it
  • Rate limits and timezone edge cases
  • The deprecated endpoint, the observability hook
  • The failure mode they want: graceful degradation, hard fail, retry with backoff

The agent reads that prompt and produces code that reflects it.

Engineers who haven't lived those stages prompt for the happy path. Not because they're worse, but because the context to prompt otherwise only comes from having seen what happens when the happy path isn't enough.

Both outputs are technically correct. Only one survives a code review.

The agent isn't broken. It's mirroring the quality of context it received.

Two prompts side by side. Happy-path prompt: add a function to send password reset emails. Production-aware prompt: idempotent per user, rate-limited by IP, token expires in one hour, audit log on send, test unhappy paths first. Caption: same agent, different output, different production outcome.

Context is king.

The implication is the part nobody likes saying out loud:

AI does not lower the experience bar for software engineering. It raises it.

You used to need to know how to write the code.

Now you need to know:

  • How to write the code
  • AND how to direct a system that will write it for you
  • Which means knowing what good output looks like
  • What the production constraints are
  • What to demand of the agent that the agent will not demand of itself

That's not less expertise. That's the same expertise, built from lived end-to-end context, applied at a different layer.

The engineers who extract the most from agents are the ones with the most lived context to bring to the prompt. AI doesn't replace expertise. It compounds whatever you walk in with.

The cost of guessing in production

All of this becomes an economics problem the moment you try to run agents in production.

From Lightrun's 2026 State of AI-Powered Engineering Report, a survey of 200 senior SRE and DevOps leaders running AI code at scale:

  • 43% of AI-generated changes need manual debugging in production after passing QA and staging
  • 0% of senior engineering leaders are "very confident" AI code will behave correctly once deployed
  • 88% of organizations need 2 to 3 redeploy cycles to verify a single AI-generated fix

That's the production reality. Now the math.

1. Rework is a budget line

A rework rate above 40% is not a quirk. It's a budget line.

People-hours, not just compute. Multiply across a team. Multiply across a year. The number gets serious.

2. Token cost compounds with iteration

A lazy prompt that takes six iterations isn't just slow. It's expensive in a way that doesn't show up until the bill arrives.

The more autonomous the agent, the longer the tail on each failed run. In my experience, an agent exploring three wrong paths before finding the right one burns 10,000 to 20,000 tokens of recovery per task.

Acceptable in development. Less acceptable in a CD pipeline running thousands of times a day.

3. Production wants determinism

Hardened code on the shortest path to the objective.

Not an agent:

  • Rediscovering the path on every run
  • Exploring whether a slightly different solution might work
  • Deciding mid-execution to refactor something it wasn't asked to

Non-determinism is a liability when uptime is the contract.

It's why production code today, the code actually running businesses, gets human-hardened before it ships.

Not because agents can't write code. Because they can't yet write code that survives production economics.

What closes the gap

AI IDE coding agents will replace developers when all four pillars stand:

  • When agents have persistent project memory: knowing why decisions were made, what was rejected, what the business actually needs
  • When agents have production awareness: knowing where code lives, what it touches, what depends on it
  • When the prompt isn't the limiting factor, because the agent already has the context the prompt was trying to supply

Until then, the future isn't humans-out. It's humans orchestrating agents that finally have enough context to be trusted with production.

The market has already answered

The "AI replaces developers" thesis isn't just being tested in benchmarks. It's being tested in P&L statements, and it's failing.

The numbers at the top of this article, a third rehiring, 55% regret, 0% confidence, aren't a blip. They're the four pillars showing up on a balance sheet.

The companies that moved first on "AI replaces developers" are now moving first on "actually, no."

Not because they changed their minds about AI.

Because they ran into the four pillars.

They built without the Why. They shipped without the Where. They learned, expensively, that the developer wasn't doing what they thought the developer was doing.

The developer doesn't disappear. The job changes shape: from writing every line to directing a partner that has, for the first time, the memory to be one.

The next question isn't whether agents will get smarter. They will.

The question is whether we'll give them what they actually need to do the job.

That's an infrastructure problem, not a model problem.

That's the work I spend my days on.

Imagine if AI didn't just generate code, but captured and carried the lived context that makes code production-ready in the first place. The decisions. The postmortems. The "we tried that and it broke." The institutional knowledge that usually leaves the building when an engineer does.

The mileage one engineer earns over a decade becomes context the whole team can prompt with. The 3 AM lessons stay in the codebase, not just in someone's head. The new engineer ramping up gets handed the context the previous one walked out with.

That's what Recallium does. A memory layer for coding agents: it captures lived context and passes it forward, across projects, sessions, tools, and the engineers who come next. See how it works.

The argument above is what I see from the inside. The product is what's left when you take the argument seriously.

Also published on Medium →

Related Resources: