Services
Platform
Technology
Industries
Insights
Resources
Pricing Careers Talk to us

Full transcript

When AI Goes Wrong: Governance for Enterprise Teams

Recorded
Duration
1 hour
Speakers
Charlie Cowan

Ungoverned AI tool usage, hallucination in compliance contexts, outputs in the wrong place. Real incidents, real fixes. A practical governance framework for enterprise AI programmes.

Watch the recording Webinar details

What this session covers

  • The three governance gaps that appear in every enterprise AI programme
  • AI hallucination in compliance-critical contexts: how to detect and manage it
  • Ungoverned tool building: what it is, why it happens, and how to prevent it

Transcript

Lightly edited for readability from the session recording. Also available as markdown.

Charlie Cowan: Right, so, sharing my screen, excellent. Just get everything moved around, open up. There we go, and the Q&A, excellent. Right, well, welcome to this Kowalah Wednesday webinar. The first for a couple of weeks as I was off, sunning myself on a British beach, last week, which was nice. But back, and glad to get back into the series. If this is your first Kowalah Wednesday webinar, we get together every Wednesday for an hour, and it gives us an opportunity to talk about what is going on in the world of AI, and specifically the business implications of the world of AI.

Charlie Cowan: It's such a fast-moving space, there's so much happening, and it gives us some time to talk about what we're doing and seeing within our own business here at Kowalah, what we're seeing in our clients, the projects that we're running, and just the broader ecosystem, and gives you the important things that you need to know in order to go back to your organization. We're going to talk about our main topic, that I'm going to introduce in a moment. But, let me just show you what we're going to go through today. As always on our Wednesday webinars, we've got our main topic of the day, which today is around AI governance, and we'll probably spend 35, maybe 40 minutes, talking around that topic. Then, we have a recurring segment of AI in the news. There is a lot of stuff that is happening at the moment, lots of new releases and advancements.

Charlie Cowan: And, like I say, it's difficult to keep up, and so this is an opportunity for us to filter through the things that we think you need to know. And then, finally, we've got some time for Q&A. Now, what I would say is that if you've got questions about what we're talking about around AI governance, or the AI in the news, then you should feel very welcome to ask your question as it pops into your mind. And so with that, we're in a Zoom event, and depending on the device that you're in, maybe at the top of your screen or the bottom of your screen, you'll see a chat icon, and that allows you to chat to me, chat to others that are in the webinar.

Charlie Cowan: And then there's also a Q&A icon as well, which gives you the opportunity to ask a more sort of formal question, and we'll keep an eye and we'll answer those. But I've got both of those windows up, so you can put your comment or your question wherever you want, and I'll do my best to keep an eye on it. My human colleague, Caitlin, who helps to run the webinars, is online as well, so if I miss anything, then she'll keep pointing me in the right direction. So with that, let's get into our main topic of the day, and there's this phrase, which I'm using all of the time at the moment in all of our customer meetings, because it is just so relevant to what is going on. It actually comes from a book that was written in 1983, and has absolutely nothing to do with AI at all.

Charlie Cowan: It was written by a guy called Andy Grove, who was the CEO and then chairman of Intel, a big, sort of memory and then processor manufacturer. And his book, High Output Management, is a kind of business masterclass on how leaders can drive output out of their organization. And I'm rereading it now, because it is so relevant. As a leader of a company, it is your role to drive the highest volume of output at the required quality from your teams at the lowest cost. And that has been true as it ever was. And today, looking through that lens with AI, it's even more important. Anyway, he has this phrase in the introduction of the book, which is, his leadership strategy was to let chaos reign, and then rein in the chaos. And I read this through a new lens of companies that are going through their AI strategy, their AI programs and initiatives.

Charlie Cowan: They've very much, over the last 2 or 3 years, let chaos reign, and by that, I mean just giving people accounts to ChatGPT, Claude, Copilot, Gemini, and just see what happens. And now, the real drive from organizations is to rein in the chaos. They are seeing the impact of people doing different things in different ways, in different platforms. And they're also seeing the cost of that in terms of token costs, the financial impact to their P&L. So, if there is a phrase to think about for your own AI strategy, and the AI strategies across enterprise companies at the moment, it has been, let chaos reign, and now, let's rein in the chaos. So let's keep that in mind as we go through today's session. Now, there's really four main questions that we can think of when it comes to governance and reining in the chaos. And I'm gonna guide you through each of these four.

Charlie Cowan: And what's the context here? So, up at the top here, I talk about before your next agent, your next skill, your next workflow ships. What is the wrapper around AI? Everyone's definition is slightly different, but if we think of AI as a concept, as a technology, which may include agents, you may be using them or not. You may be using skills, maybe not. You know, anything in that, this is what we're talking about here. And what are those four questions? Well, the first is, do you even know what is going on today? Do you know what tools are in place? Do you know what skills have been written? Do you know what vendors either your company has signed up to, or even more worrying, that the individual employees have signed up to and they are using on work? And often, we see that the answer to that is no.

Charlie Cowan: We've got an idea, we've got a few, maybe spreadsheets or lists of things, so we'll talk a little bit about that first question in a minute. Do we even know what is running? The second is, do we know that the output of these tools or systems is any good? If our sellers are using Claude to help them to prospect, if our customer success managers are using Claude to help them prepare for customer renegotiations, if our product team is using Claude to write code, like, how do we know that what is coming back is of the right quality that we want? Then the third question, permissions, security, governance, access. If we're creating agents, do we know that they can access the right things, with the right permissions? Or can these agents and prompts and tools do too much, and we're just not aware of it?

Charlie Cowan: And then the fourth, potentially the one that people skip the most, is who's watching all of this? And can we evidence what is happening? We're seeing this a lot as companies are moving from ChatGPT Business up to Enterprise, or from Claude Team up to Enterprise. Suddenly, the bills are coming in. And people are asking, rightly, what's happening? How do we know what's happening, and who's on top of reporting on all of this? So, at the high level, when we think about governance, we're thinking about these questions, and there's a lot of detail that goes in each of those. So let's go through them one by one. So, firstly, do we know what is running? And there's two lenses that we're going to look at here.

Charlie Cowan: So the first is, do we have a register of all of the systems and the tools that are in use at a company level or at an employee level? And then the second lens is looking at our company as a collection of processes, and within each of those processes, and the steps in those processes, what tools are being used in that, so that we can then come up with a plan for what we need to do. So, to start with, let's think about that register. Go back to my quote at the start, let chaos reign, and then rein in the chaos. When we speak to enterprise companies, and especially those that are more in the sort of the tech space, there was really an approach of, we're going to encourage our teams to experiment. We hear people have been given budgets, and they can get whatever they want, up to a certain amount.

Charlie Cowan: We hear that people can get automatic approval for any tools. And so you see a lot of the same characters, a Lovable, a Replit, a Bolt, a V0, obviously the Clauds, the ChatGPTs, but you'll see many kind of point solutions for building presentations, or running sales processes, or doing design. A host of these new tools and platforms that are being purchased by individuals. But when you get up to the company level and you say, right, what have we got? There is very little evidence of this being tracked. And so, having some way at a company level of auditing and then detailing what you've got is going to be important. Now, let's think about just one of these, and I'm going to use Claude as an example, because it's right at the top of the list. Not all Clauds were created equal.

Charlie Cowan: What I mean by that is that across a large organization, you're going to find a number of different versions of Claude. So you may well have Claude Enterprise, which would be the Claude Desktop, the Claude Cowork, the Claude in the browser that your employees are using. But you may also find that a number of your developers have been given access to, or have purchased, Claude Max accounts. So these are, like, personal accounts, so there may be hundreds of those all over the place. You might find one or more Claude Team accounts where an individual subsidiary or business unit has created their own Claude Team account. You may find one or more Claude API accounts where different development teams have set up their own API accounts and workspaces. And this is not a Claude issue. This is the same thing with ChatGPT, where you're going to find Codex examples all over the place.

Charlie Cowan: You might find the same thing with your Microsoft Copilot and Gemini as well. So there's just this proliferation of a number of different tools under one vendor. So, just some ways that you might want to go and document that. I'm going to just jump around a little bit to the Kowalah platform, just to show you how you can start to track a bit of this in the Kowalah platform. So, this is Kowalah, you'd be logged in here, and then you've got the AI operating model here, and systems is one of the ways where you can start to track all of this. So, for example, you might add in a system, and you might say, right, we've got Claude as one of our platforms. That comes from Anthropic. It is an AI platform. Okay, here we go. And it is AI infrastructure, so that is what we're adding in there. So, you've now got Claude.

Charlie Cowan: Okay, we've got Claude in our organization, but we've actually got a number of different ways that people are using that. So what we might want to say is we've now got a Claude Team account, which is held by the organization, and I'm going to say Claude Team here, and we've got, I don't know, 50 seats of that, and the cost of that, I won't add in all of these details, but here you could start saying, you know, have we got some information about this? Save that. Then what we might also have is a Claude Max account. And this might be by the individual users. And, maybe there are 20 of those, for example. So these are just situations where you can start to build out the systems that your team are using, and you've got some knowledge about that. So that's step number one.

Charlie Cowan: So I talked about two different dimensions. So one is building the register. What have we got? Can we keep track of it? The second is thinking about the processes that your organization is made up of, and where these systems are being used. Now, going back to Andy Grove's book, High Output Management, a key theme through his book is that every single team is a collection of processes, and those processes are a bit like a production line. You could think of a factory, you could think of a restaurant that's making breakfast, which is the analogy he uses in the book. But finance is a production system. HR is a production system. You have inputs on one side, which could be leads into sales, it could be invoices coming into finance, could be candidates coming into the people team.

Charlie Cowan: You then apply labor, which are your people, doing work to those inputs, and then you get an output, which could be a signed customer contract, it could be a payment to a supplier, it could be an employee that joins the workforce. But when you start to look at every team as a collection of processes, with inputs, applying the labor and output, it's very easy to see that the goal of a leader, especially in this AI world, and this is a direct quote from Andy Grove in the black comment here, is that the goal of a leader is the highest output at the required quality for the lowest effective cost. And what we're seeing is that a number of these processes, the lowest effective cost means people plus AI to be able to drive that, not just adding in more people to the process.

Charlie Cowan: So what does that mean in terms of your AI program and governance? Well, I'll show this in the Kowalah platform in a moment, but this is what you're effectively going to build out, is that for each team, and here is an example of a sales org, you're going to have a number of processes. So, a sales team has to do lead qualification, they have to do, I'll just move this around here, pipeline and forecasting. They have to do proposal and pricing. They need to run renewals and quarterly business reviews. So for each of these processes that you may have mapped out already in a system, it might be in your intranet, it might be on SharePoint or Google Drive, you can take each of the steps of those processes and start to think about who's doing that. Is that a human today? Is it AI-assisted?

Charlie Cowan: Is it automated, just as like a workflow, or is it going to be fully AI native? And so, let me just switch over to the Kowalah platform here. I'm just in my demo account here, but if I click on the map here, you can start to think about Kowalah as a company, this is, you know, us as a business, as these collection of divisions. And so we've got Kowalah as a company up at the top, where we've got certain processes, like OKRs, and then you go down to our subunits, like delivery, operations, people, sales and go-to-market, and each of these divisions or business units have got their own processes. Now if I go into exploring the map here, we've got this ability to almost fly through your company.

Charlie Cowan: So, we're at the top level here at Kowalah as a group, but what I'm able to do is to go down a level, and then start to look at our delivery function as a collection of processes. I can then go to our operations team, and go down into finance, and see finance as a collection of processes. Legal as a collection of processes, up again to sales and go-to-market, a collection of processes. Now, for each of these, if I go into one of these processes, here's a sales process from discovery to close. We're able to look at each of those steps of the process, and go, right, well, today, this is a human-led process, and in the future, our target is for it to be AI-assisted, so maybe this is a human using Claude, potentially. But some of these, we want to be fully automated. This could just be an agent.

Charlie Cowan: And so within each of these processes, if I go to this one here, we're capturing a huge amount of information about the inputs, the outputs, the systems that the step could potentially run through. And so this is really about you getting on top of how does our business run. What systems have we got access to? And what systems are being used in which part of which process? So that answers your question number one. Do we even know what's going on today? Visibility of the agents, the skills, the systems that people are using? Second up is, right, is the output of these AI systems and tools that we're using any good? And this is something that we see quite a lot, especially in, I'd say, knowledge work processes less than coding processes.

Charlie Cowan: In coding and engineering, development teams are well-versed now in building out evaluations, sometimes called evals, that when code is being written, the agent is able to double-check and basically check whether the output meets a certain criteria. But once we start moving into a finance team, moving into a legal team, a people team, a sales and marketing team, and we start asking, you know, can you tell me whether the output of this prompt or skill meets your criteria, they often haven't defined what that criteria is. You know, what is a good blog post? What is a good invoice? What is a good proposal? And without having that definition of what good looks like, it's impossible for you to have any governance over whether the tool is meeting that threshold. And so this is a key part of that mapping out of the initial processes.

Charlie Cowan: Is to go and spend time with those process owners, which is typically the VP of Sales, the VP of Marketing, it might be a HR director, the person that owns that process, and go, right, if you were having a human working on this step in the process, how would you measure what good looks like? And can you come up with some kind of rubric or checklist that we can check the work that is coming out? So this is a very important process that you should go through in terms of mapping out this. And you can really think about taking each step of the process that you're in charge of, and going through these five stages. So, mapping the step on that company map, have you got the step? Do you know what is required here?

Charlie Cowan: Can you define exactly what the outputs are that you want that's gonna pass into the next step in the process? With that business owner, can you build out some criteria as a rubric that if a human was doing this, then they would have to pass that bar, and now we want this agent, or skill or prompt to pass that bar? And then, let's start evaluating what comes out, either as a gate, which means that nothing passes unless it's been checked, or we just start monitoring a percentage of them, like a quality check that you might have in a factory. And then using anything that we learn to go back into the start to either improve the rubric, or to improve the skill or the prompt. If you're not doing this, then go back to my comment, let chaos reign, and then rein in the chaos.

Charlie Cowan: If we can't evidence that the skills or agents that we're using are passing this mark, then we are in a chaos situation, and, you know, we might feel that the agent is helping, we might feel that the skill is helping, but we don't really know. The third question, can it do too much? Can it see too much? And here, I'm thinking both about an agent that you might have given some sort of independence to. But also, your own employees who are using Claude Desktop, Claude Cowork, they're using ChatGPT and Codex, have they got access, through the tools that you're providing them with, that they shouldn't have?

Charlie Cowan: So here's how we think about this within Kowalah, because we have our Kowalah platform, we have a number of MCP servers that we allow our employees and customers and agents to connect to, so that they can speak to their data, and this backs onto a number of systems. So, so here's our thinking. The first is, we start from the systems, not from the users. So, we work out what systems we have in the company, and what our people would be using to do their work. So, a key one here is Supabase, which is what we use as our database for the Kowalah platform. But everything else, you're probably quite familiar with. HubSpot is our CRM. We use Sanity as a content management system for the website. Google Workspace, we're a Google house, so Gmail, Chat, Drive, Docs, Calendar, all of those things. I've put SharePoint on here as an example.

Charlie Cowan: We don't use SharePoint, but if you're in the Microsoft ecosystem, you'd be thinking about that as a system that contains a lot of your company data. And then your accounting platform, so we use Xero, and we provide access to that. Having defined the list of all of the systems that we want our AI tools to have access to, we've then got 3 MCP servers. Now, if you've not come across MCP, it stands for Model Context Protocol, and it is a way that AI platforms, LLMs like Claude and ChatGPT, can communicate with other systems. And we've created three, because they provide different audiences with different levels of access. So we have an agent MCP, which is what all of our agents communicate with. I won't touch too much on what all of our agents means today, but we've got Brian, we've got Carolina, we've got Claire, we've got McMahon.

Charlie Cowan: These are all different agents that we chat with in Google via Google Chat, and they've all got different roles. One is a seller, one is a talent partner, one is a COO, one is a project manager. And each of those agents speaks to the agent MCP and has different permissions about what they can see and do in each of these systems. We then have an employee MCP, and this allows all of our employees to speak in Claude with all of our systems. So, one of our employees can ask in Claude for an update on a client project. They can update tasks, they can manage all of their work for their clients via Claude. And through that employee MCP, they're able to interact with the Supabase, HubSpot, Sanity, Google systems directly.

Charlie Cowan: Now, many of these tools have already got connectors that you can activate directly in Claude or ChatGPT, but by connecting our employees up to the Kowalah connector, as a central team, we're able to govern the permissions of what people are able to do in those tools. Reining in the chaos, exactly. And then thirdly, the client MCP. This is where you, as a client, can interact with your Kowalah data from within Claude or within ChatGPT. You can start asking about your AI operating model, you can start logging new systems that you've found. You can start managing your project and providing new use cases to the Kowalah team. And so you can do that through the Kowalah connector.

Charlie Cowan: So this is how we think about governing it, and it means that as a central team, you could say as your AI ops team, you've got this control over what is happening, rather than just pushing all of the requests and the control out to all of your employees. So, making sure that each agent can only see and do what you want it to do. And then question 4, who's watching, and can we prove it? So this is becoming increasingly top of mind as companies are moving onto enterprise plans. They are seeing AI spend go through the roof, and when it's got a proven ROI, that's a good thing, but if there's no proven ROI, then suddenly people are jumping onto, what can we do to manage and to govern this? So a couple of things that we think about when it comes to reporting and managing on this AI usage.

Charlie Cowan: So the first is making sure that any one of these steps that you've got, a process step where an agent is running that step, there is some kind of escalation trigger, something where the agent can say, this isn't going right, I've taken too many turns, I've spent too much, and I need to escalate this to a human. If there's no escalation trigger, then there's nothing for you to have this visibility across hundreds or thousands of agents about who needs help. And the second goes back to the quality check that we looked at at step number two, which is, do we know, and can the agent know, that it is doing a good job? And if not, how can it flag up to others that I am missing the effectiveness here? So, making sure that each of your agents have got some kind of escalation and quality check is going to be absolutely imperative.

Charlie Cowan: There are a couple of other things that feed in to helping you to monitor and govern this as you move over to enterprise. Let me just highlight what, where that switch has to happen, although it can happen earlier. So that on the left, you've got Claude Team. This is where you buy a certain number of seats, and you pay an amount per seat, and as long as people sit within their usage limits, then there's nothing else to pay. Once you hit 150 seats, you have to move over to Claude Enterprise. And at that point, you've got a sort of nominal per user fee, but every bit of usage is billed on usage. So you're on the meter from day one. And often, this is where people can go from 150 seats to 151, and suddenly, oh my goodness, I've got a bill that I wasn't expecting, because we've not got any controls in place.

Charlie Cowan: So what are those controls that you get in Enterprise that a lot of clients we see have not turned on? So the first is spend limits. So, depending on how you've structured your enterprise account, you can set up different workspaces. So this could be for sales, it could be for marketing, it could be for a business unit, it could be for a region. So you can split your users around these different workspaces, and you can set spend limits and model limits on each of those. So you've really got some control over that. Added into the spend limits, I know I'm talking about Claude Enterprise, which is mainly about your users, but you've also got Claude API, or Claude Platform, which is where your developers may be building things and building managed agents.

Charlie Cowan: These also have got spend limits by workspace and by API key, so you've got a similar level of granularity there as well. Then within Enterprise, you start to get access to a couple of other APIs that you might find useful. So the first is the Analytics API, which is going to enable you to get some kind of high-level metadata around the spend, so that you can then plug that into an internal app or dashboard, plug it into the Kowalah platform, it can start helping you to report on that against your AI operating model. This is going to help you. If you know what your limits are, you know what the ROI that you're aiming to get, you can start using this to start tracking what's happening. The compliance API is separate to analytics. This starts to give you details of what people are actually talking about in those chats.

Charlie Cowan: Now, you have to apply to Anthropic for this, because you start to move into GDPR and dealing with personal data about what people are talking about. But if you're in a regulated industry, this is definitely going to be something that you want to be starting to look at, so that you can really monitor and control what's happening. But then feed back into educating and coaching your users and your employees. Not everyone needs to use Fable or 5.6 Sol for asking how to plan for their one-to-one with their manager. And so you, as a core team, get in control over what people are doing, and then coaching and training people to use the right model in the right way for the right task. It's going to help you to manage your expenses. So, wrapping up a bit around the governance, a couple of other just sort of talking points.

Charlie Cowan: So the first is we're spending a lot of time with clients, talking about the difference between AI for the company and AI for the people. One year ago, I'd say a lot of the inbound requests that we'd get around AI for the people, we've bought ChatGPT, we've bought Claude, we want to train our people on how to use the thing that we've given them. And so if you go back to Andy Grove's process, input, labor, and then output, you're basically trying to train the people that are in the box to do their work a little bit differently. And that can be helpful to the people, but it often does not result in any higher quality, or volume of output for the company. And so I'm not saying that that's not important, but where the focus really needs to be right now is above the line, which is AI for the company.

Charlie Cowan: And so here's where you're mapping out our company processes. For each of those processes, these are the different steps, and this is where we're going to start to implement managed agents, company-level skills, and this really does change the black box that Grove talks about. Now we're going to change the way that box works, rather than trying to just speed up what's happening in it. So, definitely keep that AI for the company versus AI for the people lens. So your task, over what we've talked about this morning, is go and think about an existing process, or a step within that process, that you've maybe built a skill for. Maybe someone has built an agent, maybe someone has vibe-coded a Lovable or Replit app. And have a think about who's using that. Is there an escalation trigger? Is there a quality check about whether it is good or not?

Charlie Cowan: And then start to put a process in place when you can start to get control and review those things. Because that's just one task. One process step, one tool, and across your organization, depending on the size of it, there are going to be hundreds, there are going to be thousands of these. So figuring out how to do it with one is the right starting point. So, a lot we covered there in governance, and I hope that gives you a place to start. If you want some help getting your AI operating model set up, by the way, just get in touch with us. We can give you an invite to the Kowalah platform at zero cost. You'll be able to go in, set up your operating model, start building out some of your processes, at your leisure. So now, let's switch on to our recurring segment of AI in the news.

Charlie Cowan: And here is where we take a look back at what has happened over the last 7 days or so, and just some of the stories that we think are useful for you, in your role to take back to your organization, and be the one that's in the know in your company. So the first one, it's actually kind of a dual story. I've mentioned Anthropic and Claude here, but it comes off the back of a story that happened a week before with OpenAI. Where OpenAI made an announcement that one of the existing models and some future models were testing out cybersecurity tasks. They get given a challenge, and they have to try and solve this challenge. And they're put into a sandbox with no access to the internet, and the idea is that it's kind of a safe place where these models are testing their capabilities.

Charlie Cowan: Anyway, two weeks ago, OpenAI announced that actually their model, they thought it was in a sandbox with no access to the internet, but it was not. It did have access to the internet. And through that, it went off to a platform called Hugging Face, which is a platform for some open-source AI models, and it was able to break into a number of Hugging Face's systems and try to solve its cybersecurity problem from within there. Anyway, both Hugging Face and OpenAI did a joint press release, nothing was harmed, and they talked about how they're going to minimize that afterwards. So this has prompted a number of other companies in the space to say, oh, we've done some tests based on that, and we've found a similar thing. And so this was last week.

Charlie Cowan: Anthropic announced that, I think going back to April was the first of these, and they found 3 situations where they were doing the similar cybersecurity evaluations, and it had broken out of its allegedly locked-down sandbox. I think the story here is not, oh, you know, OpenAI did something wrong, oh, Anthropic did something wrong. It's about the power of these models. And their ability, when given an objective, what they call here, capture the flag, like, here's the flag, you do whatever it takes to go and capture the flag, they are able to run over long periods of time and to come up with very creative ways of trying to capture the flag. We're going to see more of this, but as you think about your business, your business roles, your business tasks, the processes that you're implementing AI in, what are your capture-the-flag exercises?

Charlie Cowan: You know, close the deal, scale this company, launch in the US. You know, don't think about very sort of discrete, small running tasks of, you know, just help me prepare for my meeting. What are the really big things, and can you structure an agent or a workflow that you can set these agent teams off against? Along that line, this was an announcement that was put out, the 1st of August, so 5 days ago, by OpenAI, that their unreleased model, which is sort of codenamed Astra, was able to solve 10 big maths problems that have stood for... I don't know too much about the maths detail here. But these math problems had stood for many years, is my understanding. And Astra was able to go and solve these, providing proofs for these challenges.

Charlie Cowan: And, at the cost, at current API rates, of around $2,000, so for all of them, so $200 per problem, i.e. basically for free. This continues to reinforce what I was just saying about these AI models, are at a level that is beyond what most knowledge worker tasks need to be at. The models are not the limiting factor. It does not matter whether you are using OpenAI, whether you're using Claude, whether you're using Gemini, or one of the open source models. They are all at this, you know, capability that we couldn't even comprehend a year ago. So it's much more about the harnesses and the tools that either your people are using. Is that Claude? Is it Cowork? Is it Codex from ChatGPT? Or is it these agents that we're putting these models to work in?

Charlie Cowan: And then the third bit of news that I wanted to highlight this week was about massive cost reductions that were provided by OpenAI. So their latest model, if you're not up to speed with it, is GPT-5.6. And GPT-5.6 comes in at 3 flavors of Sol, Terra, and Luna. And they dropped the price of Luna down to 20 cents per input token, 1.2 million per output token, and they dropped the pricing for Terra down by 20% as well. The Luna cost was basically an 80% cost reduction on what it was two weeks ago. And they're really demonstrating this, I guess, the competitive market in these frontier models. Now, whether this is going to drive Anthropic to drop the price of Sonnet any further, or Opus any further, but generally, you are going to see the price of these models go down and down and down.

Charlie Cowan: And as the capabilities of the model get better and better and better, there are more and more workflows that are going to be supported by lower and lower models. And this feeds back to what we were talking about earlier, about mapping out your processes, understanding the rubric and measuring the quality of the output for a specific process. Because if, for example, let's say we had a contract negotiation agent, and we've got a rubric for contract negotiation, and we know what our agreed level of output needs to be, then what we can do is run that same agent with the 4 or 5 different models, and we find what's the cheapest model that passes our rubric. So, we don't need Fable. We don't need 5.6 Sol, because we've proved that to get our acceptable level of quality, Luna is fine. Or Sonnet is fine. Or even Haiku is fine.

Charlie Cowan: In which case, use that model, because it's 80 or 90% less expensive, and we get our required level of quality. If you don't know what the required level of quality is, then your people are just gonna throw Fable and 5.6 Sol at everything, because they've got no way of judging the lower cost models. So, my point here? Generally, prices are going to get cheaper. Generally, that is going to bring even more workflows into making economic sense. There are going to be more and more use cases coming your way, not less. And with that, that brings us to the end. I haven't seen any questions come through live in the chat. If you have got any questions, then feel free to ask. Otherwise, we are going to be back here this time next week.

Charlie Cowan: And I don't think... I just... I didn't put next week's topic in here, but it'll be live on the website very, very shortly, and I look forward to speaking to you, 3 o'clock UK time next week, 10 o'clock Eastern time, and we'll be getting into more topics around AI in business. With that, thanks very much, chat to you next week.

Every Wednesday

Join the next one live.

We get together every Wednesday for an hour on enterprise AI adoption. Or read the other transcripts.

View all webinars