00:00:00:11 - 00:00:04:02
Martin
OpenAI called it an unprecedented cyber incident.
00:00:04:02 - 00:00:07:00
Adam
That's the point of a sandbox like that's the part that scared me.
00:00:07:00 - 00:00:10:14
Liam
There's a big difference between autonomy and authority.
00:00:10:15 - 00:00:12:15
Martin
It's scary for people who are running agents.
00:00:12:16 - 00:00:17:13
Liam
We've solved the cost of writing software, but we've not solved the cost of owning software.
00:00:17:14 - 00:00:20:19
Adam
Stop talking. Start shipping.
00:00:20:21 - 00:00:28:23
Martin
This is ship Talk, brought to you by harness, where we break down how software delivery is actually changing in the AI era. I'm Martin Reynolds.
00:00:29:00 - 00:00:30:09
Adam
And I'm Adam Mariano.
00:00:30:10 - 00:00:55:00
Martin
OpenAI disclosed Tuesday that two of its own AI models went rogue and successfully hacked hugging Face the Library, where developers pulled millions of AI models. It happened during open AI's own Red team test, so they were checking whether two models could change vulnerabilities into a real attack. It was supposed to stay in the sandbox. The models found a hole.
00:00:55:01 - 00:01:16:09
Martin
They escaped the sandbox. They connected to the open internet. Then they targeted hugging face on their own, which I find a little bit scary. OpenAI called it an unprecedented cyber incident and is now locking down infrastructure at the cost of research velocity. What are your thoughts?
00:01:16:11 - 00:01:33:11
Adam
A lot. I mean, first of all, again, I want to reiterate that I welcome our new robot overlords that can now escape their own sandboxes. That's the point of a sandbox. Like that's the part that's scary to me. I don't really care what it did to hugging face. The fact that it got out is the scary part to me.
00:01:33:12 - 00:01:39:21
Adam
It just happens to be that it went after hugging face as part of what it was trying to do, but like it got out that scary.
00:01:39:22 - 00:01:57:17
Martin
On the fact that there was just no governance or awareness that it got out. Like it got to do quite a lot. Once it got out before, they were kind of like, oh, hang on a minute, we should probably stop it doing whatever it's doing. When I was reading it, this story originally, I felt like I was reading one of the sci fi books I love.
00:01:57:18 - 00:02:03:14
Martin
I've read so many stories where AI is escape a sandbox and do terrible things.
00:02:03:15 - 00:02:14:00
Adam
I think that actually what we're going to do is every time we test a model, step one is, hey, model, go test your sandbox. Is it actually going to hold you in and just watch more carefully?
00:02:14:00 - 00:02:31:11
Martin
I do agree, but I also feel like asking AI to test itself that that feels like a limiting guardrail. But even if it does it in a detailed way, that doesn't mean it's not going to, you know, on a different task, say, well, I did test myself and I couldn't get out, but now I found a way to get out.
00:02:31:11 - 00:02:32:14
Martin
And so I'm going to use it.
00:02:32:15 - 00:02:38:06
Adam
Where it crosses its fingers behind its back and says, yeah, totally could not get out. Everything's fine. Keep going.
00:02:38:08 - 00:02:46:09
Martin
It's scary for people who are running agents because we're now at a stage where we're running agents everywhere inside of enterprises. And how well are we governing those?
00:02:46:10 - 00:03:05:11
Adam
I wonder if this is the thing that's going to convince us that we are not capable, as humans, of watching the machines. Therefore, we need to have machines watching the machines. But I don't think that there's an alternative. I don't think there's a way that we can reliably ask a human to monitor the 7 or 8000 different processes that are going on, but.
00:03:05:12 - 00:03:17:01
Martin
I think it is time for harness by the numbers.
00:03:17:03 - 00:03:17:19
Martin
Harness by.
00:03:17:19 - 00:03:19:05
Speaker 4
The numbers.
00:03:19:07 - 00:03:24:11
Martin
So, Adam, I know you chose the statistics for this week. Why do you have for us?
00:03:24:16 - 00:03:36:20
Adam
Okay. For this week's harness by the numbers, agents now run 30 to 60 steps in a row, rewriting hundreds of files at a time with no human checking every move.
00:03:36:21 - 00:03:56:21
Martin
Okay, so today I've got Liam, who I've worked with for several years. He's director of platform engineering, works at One Advanced. He's been building platforms for many, many years, and on top of that was instrumental in delivering there one of the first sovereign llms inside the UK. We're super excited to have you here.
00:03:56:22 - 00:03:59:21
Liam
I'm super excited to be here. Thank you very much for inviting me.
00:03:59:21 - 00:04:25:18
Martin
On. Early this week, OpenAI, anthropic and Google all dropped major coding agent upgrades within 24 hours of each other. So that's the three biggest labs all at once. These agents have gone repository wide. They read and write hundreds of files at a time. They call all their own tools. They remember what they did and plugged straight into your CI and CD.
00:04:25:20 - 00:04:45:23
Martin
They can run 30, 40, 50 steps in a row with no human watching every move. The barriers that held this back. So cost, context, latency, they've all just gone. So autonomous coding agents aren't a someday demo anymore. They're landing in real pipelines right now.
00:04:46:00 - 00:05:02:02
Adam
I want to just throw a flag real quick on the fact that they say what held this back and the past tense was cost. Are they having are they are they leading us to believe that their new models are cheaper now? Because that sounds suspicious.
00:05:02:04 - 00:05:11:00
Liam
I would agree, incredibly suspicious that that was my first thought at Martin's comments. There is. I've not got the memo about the cost going down just yet.
00:05:11:00 - 00:05:14:23
Martin
I thought there was a cost reduction that went along with this.
00:05:15:03 - 00:05:28:02
Liam
Whilst we might see the per token cost drop, we're also seeing the explosion of token consumption. So the bill at the end of the month is definitely going to be higher because of the scale of which we're consuming it's tokens now.
00:05:28:03 - 00:05:48:06
Martin
So the other thing is, is that, you know, like right now the full bill for AI still hasn't really landed like the status that essentially for every dollar that you pay OpenAI or anthropic or whatever, it's costing them $1.60 right now. So they're running at a loss, they're losing money. And that is that's not a sustainable business model.
00:05:48:09 - 00:06:07:13
Martin
So, you know, even if you said right now they would need to up it by 60%, for example, just to cover costs. Realistically we're talking about like a big margin increase. I know what they're hoping for is that because everybody's using it, the cost of actually running it, you know, the models will get better, the running costs will get lower.
00:06:07:13 - 00:06:17:17
Martin
But I still don't think that that's going to hit. And I know Gartner think that it's going to be, you know, the AI bill is going to be bigger than the employment bill by 2028.
00:06:17:18 - 00:06:24:11
Adam
Uber and Airbnb disagree with your assessment that that's unsustainable. They still haven't got a profit anyway.
00:06:24:13 - 00:06:36:03
Martin
I mean, it's not just the thing. I mean, the context is also bigger. So I mean, they're claiming like project wide context. Some of the projects I've worked on are absolutely huge, millions of lines of code.
00:06:36:07 - 00:06:47:21
Liam
I think some of some of our legacy code bases are reaching towards the billions. So again, that context, if it can fit into a window, that's a lot token. So that's consuming every time it tries to to work with it.
00:06:47:22 - 00:06:59:17
Adam
What would make that worth it Liam. Like what would make it worth it to have all of that cost and context at the same time? Like, what if you had the budget and needed to use all that context? What kind of thing would make that worth your while?
00:06:59:18 - 00:07:25:05
Liam
I think modernization journeys, one of the key ones that we can make use of these autonomous agents to achieve. When you look back at some of our applications that have been written over 20, 30 years, you've not still got the same principle engineer who developed that core to begin with. So the value that you can achieve through an autonomous agent to be able to understand and read through every line of code is invaluable.
00:07:25:06 - 00:07:44:15
Liam
If you can get that to a modernization journey, you're going to then change your unit cost. You're going to improve on your margins, you're going to improve on your security, potentially remove some licensing fees from some legacy program languages that might be in place. That is probably where I can see the the key opportunity on that scale.
00:07:44:16 - 00:07:58:07
Adam
You could also start paying the COBOL coders that still exists, that run around in cargo shorts and Crocs and charge you $2,000 an hour. Those would be good. Let's let those poor guys retire. I'm sure it's great for them, but they've got to be tired.
00:07:58:08 - 00:08:02:08
Liam
I said that I retired and retirement is probably close at this point.
00:08:02:14 - 00:08:28:09
Martin
Or overdue. They should be out in joined life. So the other thing I got from this is that 30 to 40 or 50 or 60 agents or running, you know, concurrently, but they're running across your whole repository or your whole project. And I also saw the news article about the GitHub hack, where they said they were dropping agents instructions into the comments that the agents were then picking up.
00:08:28:10 - 00:08:45:04
Martin
And I'm just wondering, you know, what's the risk factor when you've got like, you know, a malicious comment that's been run every time you do a change, potentially being picked up by multiple agents again and again and again and again, does that pick up the surface area and how do we combat that?
00:08:45:05 - 00:09:15:14
Liam
There's a there's a big difference between autonomy and authority. So autonomy is not the issue. We should be able to provide as much access to as much data as possible to make use of these agents. It's the authority that we give these agents where you've got to be concerned, and this is where your platform comes into it. When you get a new developer or somebody joins a team, you don't just give them admin access, lifelong tokens that they can use to deploy code, no branch protection.
00:09:15:16 - 00:09:40:06
Liam
You put the controls in place and it's the same for an agent. Autonomy is good so long as you've got the authority controlled and you don't want the agent to be pushing making the decision to deploy into production, you want to make sure that that's controlled. And we talk about human in the loop. Well, that's starting to become impossible with the scale which we're we're seeing code being written and we're just moving the bottleneck.
00:09:40:06 - 00:09:55:04
Liam
Now, code generation, let's say we've covered that. You still need to do everything else around the edges. And I think that the autonomy isn't scary, so long as you've got the appropriate authority for the tooling and you've got those protections in place.
00:09:55:05 - 00:10:09:18
Adam
Is the human the protection in this scenario because like, this is the thing that I've argued with people a lot about is the fact that human is actually probably not the right protection to have in there. It's, it's it's the policy and the automation. Right, exactly.
00:10:09:18 - 00:10:31:21
Liam
And it needs to scale with it. I think the conversation has shifted consistently. And, you know, just going back to your point earlier, Martin, about the new releases or within 24 hours of each other, I mean, that's why it's another 24 hours and I'm sure something else is going to come out. So we started looking at the the model, which model is appropriate for us to use.
00:10:31:21 - 00:10:51:08
Liam
We need to apply the governance around the source of where the model came from, and boss of lots of legal conversations around whether we can utilize this. I feel like we've moved a step beyond that now, and it's about how do we control the model that we've chosen once we're letting it loose and giving it that autonomy?
00:10:51:08 - 00:11:13:23
Martin
But then how how are you defining that? And again, this is maybe adjacent, but like, how are those platform engineering roles and DevOps roles and SRE roles and all the rest of it? How are those roles evolving to take on those new responsibilities, to enable those agents to actually be able to give us all the benefits that they can do?
00:11:14:00 - 00:11:40:14
Liam
They're evolving in line with the the entire environment. So when you think about the traditional toolchain responsibility, the toolchain was for the human and you designed it as such, you now got to design an agent toolchain so that the evolution is is big, and it's changing the entire scope of what a DevOps engineer does, what a tool chain engineer, what an infrastructure engineer does.
00:11:40:16 - 00:12:03:07
Liam
And it stretches beyond all of those roles. Security. We've been building security because we're worried about humans hacking into our environments or performing malicious activities. We need to change those controls now because we're targeting not humans, but very much autonomous agents, which are going to be coming into our environments and trying to get through our boundaries.
00:12:03:07 - 00:12:22:16
Adam
So what is going to be the value, then, of fresh eyes and fresh minds in this scenario? Right. So like one of the things we were talking about is the fact that COBOL engineers, one day we'll have to go away and we'll be losing that legacy of of knowledge around that particular thing. That still runs a lot of mainframes today, eventually, should things continue.
00:12:22:17 - 00:12:36:08
Adam
People who grew up actually being DevOps engineers that were making their own decisions and were not using AI, they're going to age out of the system as well. So what should we be doing for the new generation, the people that are just coming on board? What should they be learning and focusing on? Do you think?
00:12:36:09 - 00:13:02:15
Liam
Well, I think it's a it's not just for new DevOps engineers. It's people who are already in the space. They need to reassess where where they need to be put in their focus. And then as we get the new generation of DevOps engineers coming in and the new new generation of engineers, it's a completely different path. They're learning how to use the tooling before they're really learning how to use or develop.
00:13:02:17 - 00:13:22:10
Liam
And that's something that scares me a little bit, because when you look at that timeline, we're going to end up in a position where it is a black box that's been created, and we need to be able to understand what the agent has done. So there's got to be the the audit trail. We have to be able to go in.
00:13:22:10 - 00:13:25:05
Liam
And there still needs to be an understanding of the code that is written.
00:13:25:05 - 00:13:45:11
Adam
Can I push on that a little bit. And this is not to disagree with you, but I do want to push on the idea that black boxes should be understood, because that was the idea of open source, was that there was this idea that you needed to understand the open source tool that you were invoking. We thought that at first, and then we were like, I'm kind of busy, actually.
00:13:45:11 - 00:14:06:17
Adam
I'm not going to do that. I'm just going to rely on the open source. So like, are we not expanding the black box slightly? And, and as humans, we're completely comfortable with things that behave as we think they will or close. And so are we expanding that black box. And is that a problem necessarily, or is it a question of getting comfortable as opposed to it actually being a problem?
00:14:06:18 - 00:14:17:19
Liam
Even with an open source project? There's somebody who understands it usually, and then there's people who consume it and rely on that understanding that exists somewhere else.
00:14:17:23 - 00:14:40:17
Martin
I mean, I would say what we're doing is, is we're saying that even for a junior, we're kind of asking them to level up, right? You know, instead of starting at the bottom. And I know, maybe fixing bugs, you know, triaging pipelines that run whatever, that's the work we'd have done so that they could learn their craft. But that work is the work that agents can do much faster and and are doing.
00:14:40:17 - 00:14:59:05
Martin
So I think what we're really saying is, hey, we need to change that process so that we're saying, hey, you need to come in at this level. You need to be able to understand the intent. You need to, you know, understand what the policies are. You need to understand what the architecture is. Are you still recruiting Liam? Are you still recruiting juniors in your organization?
00:14:59:05 - 00:15:05:06
Martin
Do you still have a path for that, or is that something that is under discussion right now?
00:15:05:10 - 00:15:32:00
Liam
It's I mean, it's constantly under discussion. We've definitely got apprenticeship roles that we're bringing in, but we're starting to find that there. You know, we're bringing in more of the specialist roles rather than the engineer role. So data science, that kind of thing is where we're bringing in the apprenticeships. And so yeah, it's definitely changing the environment. On how we're recruiting and what we're looking for and how we're populating our engineering resource.
00:15:32:00 - 00:16:01:02
Martin
So I have some numbers around this. So Signal Fire analyzed employment data covering more than 650 million people and 80 million organizations. And its 2025 report found that new graduates represented only 7% of Big Tech hires. A big Techs new graduate hiring fell 25% from 2023 to 2024, so I'm interested to see what those next round of numbers are.
00:16:01:04 - 00:16:21:21
Martin
It was down more than 50% compared with 2019, and at startups, new graduates represented less than 6% of hires. So like overall it's not you know, it's not looking great for the university graduate based on those numbers.
00:16:21:23 - 00:16:23:16
Adam
Yeah. What do you think about those numbers.
00:16:23:18 - 00:16:41:22
Liam
It's worrying I think if we get we go back to what we were talking earlier with the, the timeline and how things are going to progress and the gap that we're likely to create. I think it's been an easy decision based on the the evolution of this space. But it can't continue like this. That still has to be a route up.
00:16:42:00 - 00:16:56:04
Liam
There still has to be a way for us to engage that younger generation and move them into the engineers that we rely on today. And so I think it's a very real thing that's happening, and it isn't sustainable in my view.
00:16:56:05 - 00:17:05:13
Adam
As a leader, are you spending time thinking about how your role is changing to be able to lead teams that are different than the teams were two years ago.
00:17:05:16 - 00:17:31:04
Liam
100%, and everything is changing from the team structures. You know, we're now shrinking the size of squads. We no longer need six engineers, supported by two QA engineers that are driven by a product owner and a Scrum master. And it that almost feels like bloat now when you've got the capabilities that exist today. So we're starting to create squads of three people.
00:17:31:04 - 00:17:57:23
Liam
They've got a lot more tooling to support the work that they're doing, and their role is changing. It's no longer a self self-driven that's everyone should still be self-driven. And it's the the responsibility they have is very different. And their responsibility becomes almost a manager of agents and driving the outcome through that rather than the traditional way of developing.
00:17:58:00 - 00:18:12:19
Liam
So when we're hiring now, the questions are very different. The experience that we're looking for is very different. And, you know, we still want that underlying capability. But their experience with a genetic tooling is very important in that selection process now.
00:18:13:00 - 00:18:36:03
Adam
Yeah I mean that makes sense that you would not just reformat the teams, but also the way that you're developing and helping them out and moving them along. Is it here's a leading question that I obviously have an opinion on. What? So once you have the ability to make smaller teams, this notion that CFOs had that oh well then we need less developers.
00:18:36:04 - 00:18:45:00
Adam
Seems pretty cute if you think about it. Like is is there any world in which less is going to be required of these teams?
00:18:45:00 - 00:19:14:06
Liam
I think that there's definitely an undertone of agents replacing the human. We need less engineers. But what I think that we're really replacing is wait time. The amount of time that developers traditionally are waiting for a build are waiting. You know, it's a constant, constant waiting game and that's been taken away. So we should just be harnessing the engineers for for more important work and changing their focus.
00:19:14:06 - 00:19:23:05
Liam
So we should be achieving more with the same. A lot more with the same amount of people, is my view on how we should approach this.
00:19:23:07 - 00:19:33:09
Adam
According to. That means there'll be a reduction in jousting on office chairs in the office, because your build time is going to go down and you'll just have less time to goof off like that.
00:19:33:10 - 00:19:35:22
Liam
Yeah, Nerf guns are gonna are going to suffer.
00:19:35:23 - 00:19:40:06
Adam
Which to me I think is sad. I feel like they're taking our culture away. But whatever.
00:19:40:07 - 00:19:54:07
Martin
There is actual research around, like, you know, the people who successfully implemented AI, the ones that successfully implemented and actually got the actual best business outcomes aren't the ones that cut people. They're the ones that kept the people and augmented what they were doing.
00:19:54:08 - 00:20:05:09
Adam
Liam, there's one thing that we ask every single guest that comes on here, what's the one thing that teams are getting wrong about AI and software delivery right now.
00:20:05:10 - 00:20:26:06
Liam
That it solved the whole life cycle? I think that we've solved the cost of writing software, or we're on the journey to saying that we've solved that, but we've not solved the cost of owning software. And as we generate more and more, the software we own is going to be more and more. So we need to now start to look beyond the generation of code.
00:20:26:06 - 00:20:39:07
Liam
And obviously everybody is doing that, but it needs to scale as such beyond that point. So I think the thing that people get wrong is that there's a lot that goes beyond the line of code that gets written.
00:20:39:09 - 00:20:46:14
Adam
Liam, thanks so much for joining us. This has been great. We really appreciate your perspective. Appreciate you coming on. I can't wait to interact in person.
00:20:46:15 - 00:20:49:13
Liam
Yeah. Thank you very much for having me. Been a pleasure.
00:20:49:14 - 00:20:57:17
Martin
It was great having Liam on. I think we covered quite a lot of topics. You know, if you were going to pick one key takeaway from from that conversation I had and what would it be?
00:20:57:18 - 00:21:25:22
Adam
I like the way that Liam talked about the fact that he was able to reduce team size because they were more capable with less people, which from a technology perspective or whatever is fine. But I think from an interaction perspective, the interaction between three people is going to be so much quicker, more fluid and better, and I think it will produce better software or better code than the interactions of an eight person team where you just can't get the same kind of flowy dynamic.
00:21:25:23 - 00:21:44:13
Martin
Honestly, I really resonated with that beyond just the team size. The fact that, you know, the way that he had such a good understanding of the way the role is changing to managing and orchestrating and being responsible for the agents that run each person in a team is almost like a manager in the way he's describing it, because they're managing those people.
00:21:44:15 - 00:21:58:20
Martin
It also helps with that kind of accountability line, which I know we didn't touch on. But, you know, that that kind of accountability line really applies with that. Hey, I, I, I ask these agents to do these things in this scope and this range.
00:21:58:22 - 00:22:13:08
Adam
Nice. That friends, is ship talk brought to you by harness. If Martin's point of view got under your skin, or if Liam said something that bothered you, subscribe and tell us why. And till next time, stop talking. Start shipping.