Four roles do not yet exist on most enterprise org charts, and the gap is no longer theoretical.
InRhythm founder and Chairman Gunjan Doshi recently wrote about a shift most enterprises are not budgeting for [link to be added once published]: the move from an operating model optimized for code correctness to one built for behavioral reliability. His argument is structural. Traditional software behaves predictably: write it, review it, test it, deploy it. AI-driven systems do not. Output depends on the underlying model, dynamic context, available tools, runtime permissions, and evaluation tests that shift over time. Managing that difference requires functions most standard org charts do not reflect.
InRhythm sees this gap directly, in nearly every enterprise engineering engagement currently underway. The technology conversation is largely solved. The organizational design conversation has barely started.
The Four Functions, and Why They Are Structural, Not Optional
Gunjan names four strategic functions that a behavioral-reliability operating model requires. Each is worth restating plainly, because the naming itself is doing real work.
Evaluation Engineering. In an agentic environment, the evaluation suite is the specification. This function does not ask whether code matches a fixed design. It defines what good looks like across stochastic outputs, builds gold-standard benchmark datasets, and runs continuous regression testing against models that keep changing underneath it.
Context Architecture. Base models are becoming a commodity. Dynamic context is not. This function controls what an AI layer can see, retrieve, call, and remember at any given moment. Most enterprise AI failures InRhythm has diagnosed trace back to fragmented context, not model limitations.
AI Product Intent. Standard product management prioritizes a backlog. This function designs boundaries for non-deterministic behavior: how the system handles ambiguity, when it escalates to a human, where its refusal boundaries sit, and how it fails gracefully when it fails at all.
Agentic Operations. Traditional site reliability engineering keeps infrastructure online. This function monitors autonomous reasoning paths, tool-invocation chains, token economics, drift, and blast radius. As systems move from answering questions to executing multi-step actions, observability has to move from uptime to action accuracy.
Why This Is a Budget Decision, Not a Hiring Decision
The organizations InRhythm sees making genuine progress are not simply adding these four functions as new job titles layered onto an unchanged structure. They are reallocating real capital away from traditional coordination roles and directly into these functions, on purpose, before operational risk forces the decision under worse conditions.
This is consistent with a pattern InRhythm has written about before in the context of software delivery cadence: when execution speed compresses from days to hours, coordination overhead that used to be invisible becomes an active tax on the organization. A three-day requirements process does not survive contact with a three-hour build cycle. The same compression that shrinks traditional engineering roles simultaneously expands the four functions above, and most enterprises are resourcing the shrinking side of that equation while leaving the expanding side unfunded.
What This Means Structurally
Three changes follow directly from this, and InRhythm works through each one with engineering leadership as a matter of course rather than a one-time exercise.
Career ladders built around deep, single-station specialization need to give way to ladders that reward systemic ownership, evaluation rigor, and leverage per person.
Teams organized around functional hand-off stages need to collapse into teams organized around end-to-end problem domains, removing the coordination tax at the source rather than trying to speed up the handoffs themselves.
Headcount budget needs to move deliberately toward Evaluation Engineering, Context Architecture, and Agentic Operations, ahead of the point where operational risk makes the reallocation mandatory rather than strategic.
Where This Leaves Engineering Leadership
The central question is no longer how many developers AI can replace. It is where human judgment needs to sit once AI handles execution, and whether the organization’s structure, career ladders, and budget reflect that answer today or are still built for a model of work that no longer exists. This is the operating model question InRhythm’s advisory work is built to answer, function by function, before the gap becomes the kind of operational risk that forces the decision on someone else’s timeline.