All part of building your career
Career growth looks different for everyone, and there is no single path into tech. At Liberty IT, employees have opportunities to learn, develop their skills and explore work that makes an impact.
Tech & Engineering
AI agents can write code, analyse requirements and move through engineering tasks at a speed that would have seemed unrealistic not long ago. But when you introduce them into a live enterprise delivery team, speed is only part of the story.
For Brian Craig, Senior Director of Architecture at Liberty IT, one of the biggest lessons from the past 18 months has been that the technology itself isn’t necessarily the hardest part.
The bigger challenge is creating the knowledge, processes and controls that allow people and AI agents to work effectively together.
It’s a topic Brian explored at AICON in his session, From AI Assistants to Agentic Delivery: Rewiring Product Development at Scale. We caught up with him to hear what his team has learned from introducing AI agents into a real enterprise delivery process - and what changes for engineers when those agents become part of the team.
What's the story behind this session - why speak about agentic delivery, and why now?
Plenty of organisations have adopted AI tools by now, but far fewer have embedded AI into the way products actually get delivered. That's the gap we wanted to close.
About 8 months ago, we launched one of Liberty Mutual's first Agentic Product Development Lifecycle initiatives to find out what happens when AI agents stop being something engineers reach for and start being active participants in a product team.
What's actually different day to day once an agent is part of the team rather than a tool you use?
It changed the working day more than people expected. Within three months of starting, we'd built what we came to call a coordinated agentic factory: specialised agents contributing across requirements, planning, coding, triage and knowledge management, while people retained responsibility for direction, validation, approval and merging.
Tell us about the platform you used and what made it the choice for starting testing on it?
We chose a mature enterprise content management platform as our lighthouse project, one already operating at real scale: more than 48 million managed items and around 500,000 requests a day. That scale mattered because it meant that any changes and any problems would show up quickly and clearly, rather than staying hidden inside a small pilot.
And what did the numbers tell you?
We saw sustained throughput improvements of four to six times on complex engineering work, backed by real engineering metrics rather than projections: 2,104 human pull requests against 62 factory pull requests.
But the number that told us the most wasn't about speed. Agents were waiting, on average, 90 times longer for a person to review their output than it had taken the agent to produce it in the first place. That's when it became clear that making the technology dramatically faster while leaving the surrounding human process unchanged wasn’t enough. It created a fast system wrapped in a slow one.
You've said the constraint was never the model but the knowledge. What did that actually look like?
The instinctive response to making an agent more useful is to give it more information. We found the opposite could be true. Feeding agents large volumes of organisational context introduced contradictory, outdated and irrelevant material, and a capable agent handed ambiguous information will still reason across it and make a decision anyway, just not necessarily the right one.
We realised that knowledge had to be engineered. That shift changed how we approached the problem. Instead of handing knowledge over wholesale, we started curating it, structuring it, validating it with subject-matter experts, and exposing only what was relevant to the task in front of the agent. Improving that knowledge foundation made more difference than upgrading the underlying model ever did.
Knowledge is not something you simply give to an AI system. Knowledge is something you engineer.
Is there a simple way to know if information is ready for an agent to use?
The same test works for people. If a requirement isn't clear enough for a new starter to act on in their first week, it's unlikely to be clear enough for an agent either.
What had to already be true about your engineering standards before agentic delivery could work at all?
None of this works without solid technical foundations underneath it. Deterministic gates between stages, human review before anything reaches production, and continuous evaluation aren't extra steps bolted on for AI's sake. They're what make it possible to trust a faster system at all, and they rest on the same engineering excellence we hold ourselves to across every team.
We designed on the assumption that agents will sometimes get things wrong, rather than assuming that risk could be engineered away entirely.
As agents took on more, how did your role and your team's actually change?
One pattern we saw was what we started calling a day shift and night-shift model. During the day, people set direction, refine requirements, make judgement calls and decide what good looks like. Overnight, agents execute, building, testing, refining and updating knowledge, ready for people to react to when they're back online.
That doesn't make engineers less important. It makes them differently important, with value shifting from producing every artefact by hand towards expressing intent clearly, defining quality and exercising judgement over what a machine has produced.
What controls did you put in place to keep that trustworthy?
We were clear from the start that this wasn't about handing agents more autonomy simply because they were capable of more. Our target was managed autonomy: agents operating within defined boundaries, while people stayed responsible for direction, high-risk decisions and the quality of what ultimately shipped. The goal wasn't to remove oversight but to redesign where that oversight sat.
For another organisation still at pilot stage, what's the one thing you'd tell them to sort out first?
It's less about the tools and more about readiness: if you gave an agent access to everything your organisation knows today, how useful would it actually be for a real production task with real consequences? For a lot of organisations, the honest answer is not very - not because today's agents aren't capable enough, but because the knowledge and operating model around them aren't ready yet.
What do you want people walking out of your session to take away?
That model choice and platform choice matter, but neither will decide how well it will work. The organisations that benefit most from agentic AI won't be the ones that deploy agents fastest. They'll be the ones that learn fastest how to engineer the environment agents operate in: the knowledge, the processes, the controls and the roles around them.
Throughout the conversation, one idea kept coming back: capability was never really the constraint. What's hard, and what's proven most valuable, is the less visible work - engineering the knowledge, redesigning the roles and building the checkpoints that let a faster system be trusted. That's the story Brian brought to AICON, and it's the same one running through how Liberty IT and Liberty Mutual's engineers are learning to build with agents, deliberately rather than just faster.
Career growth looks different for everyone, and there is no single path into tech. At Liberty IT, employees have opportunities to learn, develop their skills and explore work that makes an impact.