I need to be honest with the community about something that’s been keeping me up at night.
Three months ago, our engineering team went all-in on AI coding assistants. GitHub Copilot, Cursor, the works. The promise was clear: write code faster, ship features faster, move faster.
And it worked. Sort of.
The Numbers Looked Great
Our engineering throughput shot up 59% in the first quarter. Individual developers were completing 21% more tasks. Pull requests increased by 98%. Every metric we tracked at the individual level was up and to the right.
I was ready to declare victory.
Then I Looked at the Velocity Data
Here’s what actually happened to our delivery velocity: nothing.
- Features that took 3 weeks to ship still take 3 weeks
- Cycle time from commit to production didn’t improve
- Sprint velocity stayed flat
- Release cadence unchanged
Worse, our main branch success rate dropped from 85% to 72%. More code, more failures, same delivery speed.
The Realization Hit Hard
We spent months optimizing code generation—the one part of software delivery that was never the bottleneck.
The real constraints?
PR Review: Our senior engineers now spend 60% of their time reviewing code instead of designing systems. Review time increased 91% because AI-generated code needs MORE scrutiny, not less. It’s “almost right but not quite” (66% of developers say this, according to Stack Overflow).
QA Saturation: Testing infrastructure couldn’t keep up with the volume. Our QA team went from handling 50 PRs/week to 100+ with the same headcount and tools.
Security Validation: Security scanning, compliance checks, deployment pipelines—all manual or semi-automated. AI didn’t touch these bottlenecks.
Amdahl’s Law Came for Us
Coding represents about 15% of the work involved in shipping software. We made that 15% faster and expected 100% of delivery to accelerate.
The math doesn’t work. The system moves as fast as its slowest link, and we just made the fast parts faster while ignoring the slow parts.
The Question I’m Wrestling With
What were we actually trying to fix?
Did we misidentify the constraint? Did leadership—myself included—get excited about AI’s promise without understanding our actual delivery pipeline?
The research backs this up: Waydev’s 2026 report shows this is industry-wide. Faros data confirms it: throughput up, velocity flat. InfoQ covered Agoda’s experience: “AI coding assistants haven’t sped up delivery because coding was never the bottleneck.”
I’m Looking for Wisdom
For those of you who’ve been through this:
- Did you see the same paradox? High individual productivity, flat delivery velocity?
- What actually moved the needle? Not on coding speed, but on delivery speed?
- How did you identify your real bottlenecks? What did you measure that mattered?
I’m not anti-AI. I’m pro-systems thinking. But I need to get honest about whether we’re optimizing the right parts of the system—or just making the easy parts easier while the hard parts get harder.
Anyone else living this paradox?