I’ve spent a good part of my career making the same argument in different conference rooms: Performance matters, and it deserves real investment.
And I’ve heard the same reply more times than I can count.
So the page takes a few extra seconds. What’s the worst that happens?
Honestly, for a long time, that pushback wasn’t entirely wrong. Unless you were a large retailer where a hundred milliseconds of latency showed up directly in the revenue numbers, or a mature engineering organization still carrying scars from a public outage, nothing dramatic happened when things got a little slower. So performance testing became the thing you did if there was budget left at the end — and there was never budget left at the end.
Here’s what always frustrated me about that. Every other discipline in the SDLC found its moat, its non-negotiable reason to exist. Developers build the product; you can’t ship without them. Functional testing became the quality gate nobody dared skip. Automation earned its keep by making manual testing scale. Security got its seat at the table the hard way — breaches, regulators, headlines. Nobody asks the security team to justify its existence anymore.
Performance? Performance was where organizations felt comfortable taking chances. Always the last priority, perpetually squeezed to the end of the release cycle — despite being, and I’ll say this with some bias but plenty of conviction, the most technically demanding discipline of the lot.
You need to understand the application, the infrastructure underneath it, the workload on top of it and the math connecting all three. Deep skills, thin mandate.
That was the deal for decades.
AI is tearing that deal up.
When Efficiency Becomes the Product
In the pre-LLM world, a slow application was mostly a UX problem. Annoying, sometimes costly, rarely existential.
In an LLM-powered product, the economics work differently. Every token consumes compute resources that someone is paying for, whether the output is brilliant or wasteful. An inefficient prompt or a bloated retrieval pipeline doesn’t just feel sluggish — it shows up as a line item, every single month, growing with your user base.
Now watch what happens when an engineer walks into a review and shows they can get the same output quality with 40% fewer tokens. Nobody asks what the worst case is. Nobody asks whether it can wait until next quarter. The benefit is immediate, visible and sits comfortably on a spreadsheet.
That’s the moment I realized performance finally has its moat — the thing I’d been trying to argue into existence for twenty years, AI economics built in about eighteen months.
GPUs are expensive, and the industry is still coming to terms with what that really means. Squeezing more inference out of the same cluster, trimming prompts without losing accuracy, batching and caching intelligently, and routing simple queries to smaller models instead of sending everything to the biggest one — this work used to be invisible. Nice engineering, sure, but nothing you’d mention to the board.
Today, it’s the difference between an AI product with viable unit economics and one that quietly loses money on every request.
I’d go further: Optimization skill is now among the highest-ROI capabilities an engineering organization can have. The same craft that once got labeled “premature optimization” and deferred indefinitely is suddenly a competitive advantage you can put in a pitch deck.
Please Don’t Fold This Into SRE
Whenever performance comes up in org-design conversations, someone suggests merging it into SRE. On paper, it looks tidy. Both care about production, both live in dashboards, both sit close to infrastructure.
I’ve watched how that plays out, and it’s predictable: Reliability always wins. A live incident will beat a latency drift every single time — and it should. Which means performance work, tucked inside SRE, gets deprioritized inside its new home just like it was deprioritized in the old one.
Different room, same last place.
The AI era makes the mismatch sharper. Keeping a system up is simply not the same craft as making a system efficient. The engineer who can untangle a failed Kubernetes rollout is rarely the same person who can profile GPU utilization, reason about cache behavior in an inference pipeline, or model what a RAG architecture costs at ten times the current scale.
What’s emerging looks less like traditional SRE and more like a blend of systems engineering, ML engineering and cost analysis. My bet is that performance engineering merges with this broader optimization discipline — and becomes a big, first-class function in its own right — rather than surviving as a sub-team under reliability.
It’s already happening.
Look at the job boards. Roles focused specifically on inference optimization, AI performance and model efficiency are increasingly prominent, with engineers hired to cut latency, improve throughput and reduce cost per token.
Look at the executive conversations, where “what does this cost per query?” has become a default question rather than one somebody has to raise awkwardly at the end of the meeting.
For those of us who spent years defending performance budgets line by line, this is a strange and satisfying moment.
We didn’t win the argument. The economics changed, and the argument won itself.
Performance is no longer the discipline you get to last. In the AI era, it might be the one you can least afford to skip.

