AI Ops Engineering: The Missing Engineering Discipline for the AI-Native Company
AI is moving from something employees use to something companies operate.
That distinction matters.
For the last few years, most companies have treated AI as a collection of tools. Give people ChatGPT. Buy Claude. Add Copilot. Connect a few workflows. Build some prompts.
That was the first phase.
The next phase is very different.
Companies are starting to have agents that monitor systems, analyse customers, prepare decisions, update databases, write reports, classify work, communicate with other systems and increasingly collaborate with other agents.
At that point, the problem is no longer:
“How do we use AI?”
It becomes:
“How do we engineer an organisation where humans and agents can reliably work together?”
I think there is a missing engineering discipline for this.
I call it AI Ops Engineering.
What is AI Ops Engineering?
AI Ops Engineering is the engineering required to enable an organisation to operate AI and agents at scale.
It sits somewhere between:
- AI Engineering
- Software Engineering
- IT
- Data Engineering
- Automation
- Internal Platform Engineering
But it isn’t exactly any of them.
An AI Engineer might build an AI-powered product feature.
A Software Engineer might build an application.
IT might manage identity, SaaS applications and access.
A Data Engineer might build the pipelines moving information between systems.
The AI Ops Engineer asks a different question:
What infrastructure, tools, context and operating systems need to exist for agents to do useful work across the company?
The output isn’t necessarily an application.
It might be an agent.
It might be a tool an agent can call.
It might be a data sync.
It might be a knowledge system.
It might be an evaluation pipeline.
It might be an MCP server.
It might be a small internal application sitting inside your mission-critical apps.
Or it might be the infrastructure allowing fifty different agents to safely use the same capabilities.
The unit of work is increasingly the organisation itself.
We have seen this happen before
Clay is a good example.
Clay didn’t just create another sales tool.
It helped make go-to-market programmable.
That eventually created a new type of builder: the GTM Engineer.
Clay describes GTM Engineering as building automated revenue systems using AI, data and workflow automation instead of running those processes manually. Clay says it started using the term in 2023.
The interesting part isn’t the title.
It is what caused the title to appear.
The abstraction level changed.
Previously, doing sophisticated revenue automation required engineering resources, data engineers, RevOps specialists and a collection of point solutions.
Clay compressed much of that complexity into a platform.
Suddenly one technically minded operator could build systems that previously required several functions.
I think something similar is now happening to the rest of the company.
The workspace is becoming a runtime
Look at where platforms like Notion are heading.
Notion started as a workspace for documents and databases.
Now its Custom Agents can react to events, run on schedules, access company knowledge and connected systems, take actions and delegate work to other agents. Notion also provides activity logs, permissions and version history around those agents.
More interestingly, Notion Workers introduces hosted TypeScript programs that can expose tools, syncs and webhooks to those agents.
That is a significant shift.
The workspace isn’t just where humans store information anymore.
It is becoming:
Knowledge + Data + Agents + Tools + Runtime
And Notion isn’t alone.
The broader direction is clear.
Our existing SaaS layer is becoming programmable by agents.
The company itself is becoming programmable.
Someone needs to engineer that environment.
This isn’t prompt engineering
The first generation of enterprise AI work focused heavily on prompts.
Prompts matter.
But prompts are probably one of the least difficult parts of operating AI inside an organisation.
The difficult problems look more like this:
How does the agent get the right context?
Where does customer information come from?
How fresh is it?
Who owns it?
What happens when HubSpot disagrees with the product database?
Can the agent access call transcripts?
Can it query millions of product events?
What can the agent actually do?
Can it update HubSpot?
Can it create an issue?
Can it run Python?
Can it query Postgres?
Can it call an internal API?
Can it trigger another agent?
What is the execution environment?
Where does long-running work happen?
How are failures handled?
How do you resume a run?
How do humans inspect what happened?
How do you version tools?
What is the permission model?
Which customers can this agent see?
Which actions require approval?
What happens if somebody puts malicious instructions inside a support ticket?
Which agent identities have access to which systems?
How do you know any of this actually works?
What does success mean?
How do you evaluate output?
How do you detect regressions?
How do you know whether an agent is saving money or quietly creating more work?
These are engineering problems.
The AI Ops stack
I think AI Ops Engineering will increasingly own five layers.
1. Context
Agents are only as useful as the context available to them.
That means connecting:
- CRM data
- product data
- support tickets
- call transcripts
- documents
- databases
- analytics
- internal APIs
- institutional knowledge
The goal isn’t simply moving data around.
It is making organisational context usable by machines.
2. Tools
Agents need capabilities.
Search for a customer.
Update an opportunity.
Analyse usage.
Generate a report.
Run a model.
Create a refund request.
Query the warehouse.
Send something for approval.
These capabilities need stable interfaces, permissions and documentation.
Increasingly, internal APIs will be designed not only for applications.
They will be designed for agents.
3. Agents and workflows
This is the orchestration layer.
Which agent handles what?
What triggers it?
Which model does it use?
Which tools can it access?
When should another agent take over?
When should a human enter the loop?
This is where prompts become systems.
4. Runtime and reliability
A demo that works once is easy.
A system that runs thousands of times across an organisation is different.
Now you need:
- retries
- queues
- state
- logging
- tracing
- observability
- evaluations
- versioning
- cost controls
- fallback behaviour
- human escalation
The same transition happened with software.
Eventually AI workflows need production engineering too.
5. Governance
Agents are effectively becoming another type of user inside the company.
They need identities.
Permissions.
Audit logs.
Ownership.
Policies.
Deployment processes.
Someone needs to know:
Which agents exist, what they can access, who owns them and what they are doing.
That becomes an engineering function very quickly.
What does an AI Ops Engineer actually build?
Imagine a SaaS company.
The company wants to identify customers at risk of churn.
The naive version is:
Send some HubSpot data to an LLM and ask whether the customer will churn.
The AI Ops Engineering version might combine:
- product usage from PostHog
- account and deal data from HubSpot
- support conversations
- call transcripts
- billing information
- customer configuration
- historical churn patterns
An agent could investigate accounts showing particular signals.
Another agent could pull the relevant evidence.
A model could classify the situation.
The result could appear directly inside the CRM.
The Customer Success Manager could ask follow-up questions.
The system could prepare a recommended intervention.
Actions with financial or customer impact could require human approval.
Every decision could retain its evidence.
Every run could be inspected.
The models could be evaluated against historical outcomes.
That’s not a prompt.
That’s an operational AI system.
Someone needs to engineer it.
AI Ops Engineering is not AI Platform Engineering
There will obviously be overlap.
But I think the distinction matters.
AI Platform Engineering asks:
How do we provide infrastructure for engineers building AI systems?
AI Ops Engineering asks:
How do we use that infrastructure to change how the company operates?
The first is infrastructure-oriented.
The second is operationally oriented.
An AI Ops Engineer might happily use Notion, Cloudflare, Claude, PostHog, HubSpot, MCP and a few hundred lines of TypeScript.
They aren’t trying to build a foundation model platform.
They are trying to make the organisation work differently.
That difference is important.
It isn’t traditional IT either
IT historically manages the technology humans use.
Devices.
Identity.
Applications.
Permissions.
Networks.
Security policies.
AI Ops increasingly has to manage something adjacent:
the technology doing work on behalf of humans.
The boundary between IT and engineering therefore starts getting fuzzy.
When your CRM contains agents…
Your workspace runs code…
Your knowledge base triggers workflows…
Your AI assistant can modify production systems…
…and your SaaS tools expose themselves through MCP…
Who owns that?
It can’t simply be left to individual employees experimenting with prompts.
But putting every change through a traditional software engineering team defeats much of the point.
AI Ops Engineering sits in that gap.
The AI Ops Engineer
The interesting thing about this role is that the best people probably won’t look like traditional specialists.
They need enough software engineering to build reliable systems.
Enough AI engineering to understand models, context, evaluation and agents.
Enough IT thinking to understand permissions, security and organisational systems.
Enough product thinking to identify the actual workflow.
And enough business understanding to know whether any of it is worth building.
That combination is unusual.
But the GTM Engineer would have looked unusual five years ago too.
From prototypes to organisational infrastructure
Most companies currently have a growing collection of:
- prompts
- GPTs
- Claude Projects
- automations
- notebooks
- scripts
- Zapier workflows
- n8n workflows
- internal tools
- agent experiments
Individually, many of them are useful.
Collectively, they become a mess.
Nobody knows who owns them.
Nobody knows which ones still work.
Data gets copied everywhere.
Permissions become unclear.
The same capabilities are rebuilt repeatedly.
Eventually companies will need to treat AI systems the same way they learned to treat software systems.
Not every experiment needs governance.
But successful experiments need a path to become infrastructure.
Something like:
Experiment → Workflow → Agent → Production System → Shared Capability
That transition is where AI Ops Engineering becomes important.
AI changes the shape of internal software
For the last twenty years, companies bought SaaS applications and humans operated them.
The next generation looks different.
Humans will still use applications.
But agents will increasingly operate applications too.
Agents will work across applications.
And some applications will mostly exist to provide context, state and tools to agents.
This changes what internal engineering looks like.
We won’t just build software for employees.
We will build an environment in which employees and agents operate together.
That environment needs architecture.
It needs interfaces.
It needs security.
It needs observability.
It needs ownership.
It needs engineering.
That is what I mean by AI Ops Engineering.
Not operating the models.
Not maintaining GPUs.
Not writing better prompts.
Engineering the organisation so that AI can operate inside it.