Podcasts

The Harness Problem: Governing AI That Starts Over Every Time

Written by Cheryl Brown | 1 Oct, 2026

Every time an AI agent runs, it starts fresh, with no memory of what it did last. Presidio CTO Rob Kim joins Adam Blue to explain what that means for keeping AI accountable, and why swapping one AI model for another can undo a process that was working fine. Their bigger point: Banks don't need smarter AI to win. Today's tools already hold decades of untapped value.

Watch

 

 

Subscribe

    

Related Links

AI for Everyone

"Tron," directed by Steven Lisberger

"Brain Rules," by John Medina

Transcript

Join Rob Kim and I on the next Cut to Context, where we talk about agent governance, how if agentic technology doesn't really improve dramatically at the frontier level, we've still got 20 or 30 years of capabilities and goodness to unpack, and how thinking about “Tron” might give you some insight into ways that non-LLM technologies are starting to arrive on the scene and have an impact on the way we're all using AI.

Adam Blue

Welcome to Cut to Context. I'm here with Rob Kim, who is CTO of Presidio, a really crucial and super effective partner for us at Q2. And today we're going to talk a little bit about agentic governance. So thanks for being on today, Rob.

Rob Kim

Thank you for having me, Adam.

Adam Blue

So let's dig right in. Talk about what it means from your perspective in the work you're doing to govern an agent and kind of put those words together in that order. It feels odd when I really introspect on what that might mean, but just dig in and tell us what that actually means.

Rob Kim

Well, what's interesting is we had this conversation with some of our technical thought leaders, like what is even an agent? Because technically speaking, these things don't persist. So are they really just agent templates given the fact that each run is so ephemeral and the concepts around governing it, at least in terms of how we provide restrictions or harness it, the guardrails, it's really just an application of instruction sets that we're applying at runtime. But I guess when we think of harnessing, that's precisely it. It's how do we provide as explicitly as possible the guardrails around or the restrictions around how the model should respond and what knowledge components should it consider when it's providing you the inference?

Adam Blue

Interesting. I think that breaking that into those two categories is really thoughtful because on the one hand we have some set of, we'll just call them output limiters, filters, whatever, that says let's end up with outputs that feel sensible. And that's a big, very loaded word that I'm just going to drop and then run away from. And then the other, and maybe the place for us to focus first is what does it mean to provide the agent with the correct context and what are some of the axes that go into assembling that context when you guys think about it?

Rob Kim

Yeah, so I guess the first thing is, all context for us is going to be how can I make sure that the response that the agent provides is something that is going to be meaningful and relevant to what it is that I want. So from that perspective, obviously people hear things about skills, so the idea that I can apply instructional formats to say, "Hey, think about things in this sort of way." Maybe you're applying guidelines around how you want the response to be formatted, or in some cases you could talk about the ethos or the branding of the response tailored to the way that your organization is. Other things that are probably much more around precision and accuracy are going to be, at this point, table stakes things like don't make stuff up, make sure that it is factually correct, cite your sources, don't create the sources.

And then the other thing would be in multi-agent and more comprehensive or complex workflows, make sure you understand what the previous agent run or if you're in a situation where you're providing functions like a set of subfunction agents that are rolling things up to an upstream coordinator agent, make sure that you all understand your part and what has already been done before.

So this idea of being able to take whatever context is being utilized and being able to pass that on to the various agents, which frontier model services do a tremendous job of, and that's why people love using it because it's so easy. But as we've found, especially with more and more of our clients talking about model routing, intelligent model routing, dynamic model routing, how that doesn't necessarily translate well when you start to look at open-weighted models unless you account for that somehow.

Adam Blue

Yeah, that makes sense. So let's move in a little different direction and talk about context from this perspective. I don't remember which French philosopher said it, but he famously said, "Hell is other people." And I feel that way most days. So when you talk about an agent and that agent lives on my desktop and I tell it what I want and I give it access to information that clearly only I have rights to because it lives here and it works on my behalf, that feels like one order of magnitude of problem to work out context. And I don't know if I claim that problem is solved. I'd claim that we're on a pretty reasonable path maybe to solving that problem.

And then I think about a company, and it doesn't even have to be a huge company. And you think about how much of the organization function and culture of a company exists based on some level of asymmetry of information between the people in the company. And sometimes that asymmetry of information is regulatory. Finance people should know whether or not you're going to make your quarter and maybe the rest of the business shouldn't know because that information is material non-public information if you're a publicly traded company. You have all manner of PII and data. And even if you're not in a heavily regulated space, there's still reasonable business expectations of privacy.

And so when you introduce agentic technology to the space where humans interact with each other, which feels like it's pretty important and people are doing every day, how do you think about the scenario within which the initial push is to say, let's just MCP up everything and then we'll just drop an agent in the middle and it will magically figure out our business. And then it turns out there's a lot of subtlety and nuance in the business that is not embedded in those underlying data sources.

So have you found that to be the case or do you think we just make the MCPs and we're off to the races?

Rob Kim

No, we definitely think this is an issue. In fact, we've gone as far as to look at user-initiated agent use very differently than enterprisewide or what's emerging as more autonomous complex agent workflows as something that we have to handle entirely differently, both from an identity access control and authority perspective, but also how the agent has to be able to delineate the materialized views that a certain persona should be able to see, to your point.

I mean, we do that with data sets all the time. I might be looking at the same data source, but my materialized view is going to be very different based on the persona that I am. I'm not going to be able to see everybody's accounts, but maybe I do get a roll-up, and if I'm the CRO, I get to see everything.

So how do you make sure that with any agent workflow that's going to be more autonomous in nature or being serviced by multiple personas is accurately managing that across all of them? So we haven't solved for that. And by the way, we don't believe anybody has solved for that. I mean, we're doing our own investigations obviously into things like agency and obviously ATA standards, what third-party access management identity access and governance platforms are trying to do. And there's bits and pieces that are all out there, but nobody has really stitched that together, at least from an authority and access control standpoint, let alone how are we delineating and making sure that the agents stay on top of that.

The other thing that we found that's a bit scary is, man, agents don't really … they're just like humans. They say they do stuff and may even tell you right to your face that they're doing stuff. But when you look at the trace and you actually see from the hotel data that while they may have retrieved the right answer or the answer that you've wanted, it didn't get there based on your intended workflow or in some cases some of the guardrails and guidelines that you put in place.

And so that's the other thing that's scary that we found that just because you receive the right response doesn't mean that that workflow worked the way that it intended. And that also made us start to think about, wow, we really have to look at how we access data quite differently when we release things that are more autonomous or not user-initiated in nature.

To your point, I think we've solved—high 90s% solved—user-initiated workspaces. And for us, as part of trying to accelerate user adoption and fluency, the core of that really was that if we can provide isolation of a user's workspace down to the runtime execution, where the files live, even if somebody makes a mistake, putting data in where they shouldn't, it's only isolated to them, to their workspace, as well as the execution of that. So if we can do that architecturally in the way that we're designing the harness, then even though we don't have all the governance things in place, because truthfully we don't know what they all are until we start to build some practitioners and see the patterns, we at least feel comfortable that no one user will affect everyone.

And some of this stuff is going to be learn as you go. And certainly for us in our conversations with clients, because we were willing to do that, a lot of lessons learned for us to go back and say, "Hey, don't do that because we already screwed that piece up before."

Adam Blue

Yeah. Yeah, it's interesting. And it's funny because the word “harness” didn't mean anything interesting 18 months ago, and now we use the word “harness” like we've been saying it since we fell out of the womb.

So when you think about the harness itself, and that seems to be where a lot of the competition is now in terms of taking market share or building reusable products or staking out where the value is going to be, that ability to create and manage the boundary between traditional repeatable deterministic software and then the nondeterminism of the LLMs that makes them interesting. And at the end of the day, the LLM is fundamentally a collection of weights and you give it an input stream of tokens and it produces an output stream of tokens. And we just don't have time in our day to say that repeatedly. And so we use these deeply anthropomorphized words.

But I think the way we talk about these things sometimes can create some of the misunderstandings and failures and fundamental ground truths that lead to these challenging outcomes. And I think the less familiar people are with the technology, the more likely they are to make those leaps of cognitive assumption about how stuff works.

And so do you have any interesting experiences you'd relate around places where it just really felt like the technology had it figured out and then you wake up on a Tuesday and it's like we were just wrong. We were just so wrong.

Rob Kim

That's funny that you say that. I think all of us come a year from now are going to look back and say there were times where we say, "Well, this is the way," and then find out we were completely wrong.

Yes, that happens more than you would think. I think the first thing is that I think the primary goal of any harness is not to make probabilistic, stochastic models deterministic. And I think that's one thing that we have to set aside, and it's really to make the authority, the inputs, and the actions much more governable. That's the way that I would look at it. Hallucination itself is not actually a bug or a bad thing. It's what these models are supposed to do. To your point, it's regressive in nature and it's really just about what's the probability of the next token.

And if we go back to what we said about agents being completely ephemeral, everything that we're putting on top is just guiding the way that it should provide the response and the inference. I do think one of the things that we do is we tend to maybe over rotate. I'm seeing it internally where yes, we're combining more deterministic services as part of a more complex workflow. And in some cases you almost hobble the agent from doing what it needs to do. And while we're getting the right output and response, there's a lot more care and feeding that we are building into this workflow that now we're going to be responsible for figuring out how to maintain versus as these models are getting smarter in some cases or just allowing the agents more flexibility so that we can better understand how to, I guess, minimize the amount of harnessing that we need actually in many cases produces a better result. That is something that we have definitely seen.

The other big one is that there's much more, like with model routing, which comes up quite a bit now. I know I mentioned it twice now, but it happens to be topic du jour, especially when you see what Jev has been doing. But one of the things that we saw was that relationship between how you harness and the model is much tighter and much more dependent than we originally thought.

And so this idea that I can create an agent workflow and then later on just substitute a different model and it doesn't seem to behave the same. Why is that? A lot of that has to do with some of the things that we talked about around context portability and how we're applying that context across multiple agents within a given workflow and solving for that.

One of our top engineers actually just had baby, and he was explaining to me how he was watching the nurses and doctors and different ones every single time, different nurse, different doctor, every single time would glance up, look at the board and know exactly what to do. When you were fed, what medicine to give you, never had to ask any questions. Everything was already outlined. And then he made that parallel to the, "Well, I had this complex agent workflow that I use for platform ops. When I used Sonnet, it worked flawlessly. I switch out and use MiniMax and I'm running into all these weird issues." And all he did was simply provide context portability for each of the agents by simply just creating a skill, applying it to every agent to say, "Before you execute, I want you to go to this area, see if it's done before, gather information that requires before you do the run. After you complete the run, I want you to drop in a file, tell me what you did, why you did it, the results of it in machine format for every single agent." And as soon as he provided that, the workflow worked without issue. And so it gets you to think that when we say model services, there's a lot of harnessing that occurs within the models themselves that we're not aware of.

And so, as we start to look at things like routing, we have to be much more conscious that that dependency in how we design these agent workflows is there. And that took a lot of learning. It took a lot of people using the AI studio that we built and users just complaining, "Hey, when I switched to MiniMax, it sucks." And well, it's not that it sucks, it's that there are harness components around context that are being handled by Claude that is not being handled by MiniMax. And so we have to figure out a way to replicate that in order to get the same level of efficacy for that workflow that's going to use that open-weighted model.

Adam Blue

Yeah. If you think about software engineering, when you have two components that are supposed to be interchangeable and they don't exhibit the same behavior from the same inputs, we say that one of them is broken or both of them are broken. And in AI, when we have two different models and they generate different outputs from the same inputs, I guess we're supposed to think that's amazing. And so I think it's just a totally different way to solve problems and to think about concepts like encapsulation and abstraction because they don't work the same way.

It's like when you're in high school and you learn a bunch of physics and you think, "Man, physics are pretty cool. They kind of describe the universe." And then you pick at some problem like, "What if all the collisions weren't perfectly inelastic between two bodies? Or what if air had friction?" And your physics teacher is like, "You don't really want to do that." And then eventually you get to university-level physics and you start to relax some of those constraints and you get into some very advanced theory and maybe you master that.

And then you say, "Well, what's underneath that?" And the instructor will say, "Well, what's underneath that is quantum physics. And you really don't … you just don't want to go in there.”

Rob Kim

Everything you learned previously is invalidated.

Adam Blue

Yeah. And it's like Newton was a genius, but quantum physics says everything he said is wrong. But he's not so wrong that we don't continue to use Newtonian physics every day, constantly. I mean, your brain is kind of hardwired to grasp Newtonian physics. You do it every time you pick up a glass or every time you shoot a pool shot or whatever.

And we now have a technology. I think the shift is as akin to the schism between traditional classical physics and maybe all the way to quantum physics where you have to learn a new way to think about the problem a little bit.

Rob Kim

Yeah. And I think to that point, it's a really good analogy because we seem to be so enamored with obviously each of these, the flagship models coming out and the fact that they can do so much more. And then when something doesn't work for a thing that we're building, we just keep saying, "Well, we just need the models to get better."

To your point, I think, so the whole idea of Newtonian versus quantum, I think we have plenty of Newtonian that we could be using. If intelligence stopped right now, nothing got any better, it's all going to be about how we configure, customize, provide the guardrails, provide the guidelines for specific use cases that's going to drive the enterprise use.

So to your point, I'm not thinking quantum physics when I'm hitting a cue ball. And most of the things we do in the enterprise do not require that level of intelligence. We're going to be implementing AI for 20, 30 years, even with the level of intelligence that we have in these models right now. And Jev was perfect to explain that because you start to see the practical deployment. It's almost like what Cohere did when they came out and said, "Hey, when it comes to automation and workflow, it is about enterprise use." So I completely agree.

Adam Blue

Yeah. So Jev is really interesting. So let's expand that a little bit here as we wrap up. So Jev, and you've probably looked at it more than I have, but you can keep me honest here. So it's a model that is, instead of producing an output in the stream of human readable structured language, produces an output in the form of a typed kind of set of capabilities. So you can have a list of probabilities, you can have kind of a Boolean sort of thing. They have something called a Noul or a Noul. I don't know how to pronounce it, but it's a different shape of data type.

But the idea being that the Jev model is not a large language model and is the model that concerns itself with classifying things, and really that's it, just classifying things. And it's fascinating because the massive constraint on the output space in theory would make Jev much less expensive and much more interesting for things like is this a good response or a bad response. Is this a complaint about the temperature of the room or is it a complaint about the temperature of the coffee? Those kinds of decisions, it feels like is what Jev's targeted at. Have I got that right from what I've read?

Rob Kim

Yeah, it's basically rank or scoring versus multiple choice versus is it right, true or false? And it's really great for that. And I think one of the places where we're looking to actually integrate it into what we're doing in the agent harnesses is around model routing and being able to help make those decisions with considerations into cache coherency, because obviously it has access to all of that hotel data to be able to make those sorts of choices, certainly better than we can or some sort of rules-based engine that is static and doesn't really work.

It's not a substitute, as you said, it's a daughter card to the CPU, but the way that we think it's really going to provide benefit is both just in terms of speed and ability for providing more concurrency, especially if you're looking to implement your own GPU environments versus curated services, as well as drop the overall token consumption for when we do have to send chunks up to upstream models.

Adam Blue

Yeah. And I think the evolution of Jev fits in really nicely with something you said, which is even if the frontier intelligence doesn't get much better, and I don't know if it will or it won't, I have my bets, but even if the frontier intelligence is close to or asymptotically approaching as good as we can get with large language model, we've got 10, 20, 30 years, I totally agree, of capability to unpack that we can use to make things more effective and more efficient and more productive. And then Jev a little bit is like saying, "What's the smallest possible problem I can solve with this machine intelligence and inference?"

I think it does qualify as AI for sure. And then how does the smallness of that actually give it tremendous advantages versus this other unwieldly. I think for me, this is an opinion. I will bracket this as an opinion. I think the obsession with AGI is actually limiting the usefulness of AI. I think that focusing on what it could do well and if everybody could just simmer down, we'd probably make more progress. But man, it's such a ring to reach for that is a distraction in my opinion.

Rob Kim

Yeah, it's interesting how we approach developing AI very much like the way humans are. I love John Medina. He's an affiliate professor over at University of Washington, but he wrote a book at this point decades ago called "Brain Rules." And I'm sort of rereading that because everything conceptually is exactly how we're approaching how we're building AI and test-time scaling components and how working memory works in our brain. All of those things have a lot of consistency.

And so it's funny to me that while we are building AI to pretty much emulate how we think and learn, we're expecting this AGI to be all-knowing when the point of getting smarter in an area is about specialization, it's about domain specifics, and that's what we see. And so why do we care about AGI when what I really want is intelligence around domain areas, around speciality that can provide me more insight than a person that might have trained for a lifetime in a particular area. What you're seeing in math as an example. That just seems unusual to me.

But yeah, I agree. We have plenty of intelligence in what we have now and 90%, 95% of the enterprise use cases really just boil down to how we handle more flexibility within complex workflows. And that's something that AI can do right now.

Adam Blue

Great. Well, I think that's a fantastic conclusion, so thank you for that well-timed and inevitable summary.

So one of the things we like to do at the end of Cut to Context is to pull a reference from art or literature, history, or pop culture. And today when you talk about Jev, I'm reminded, I'm sure you've seen the movie "Tron," right? Yeah. Yeah. So "Tron" is fascinating because it posits this interiority to a computer that in many ways matches human society, although with very fast motorcycles and lots of consequences for running into things. But my favorite part of "Tron" that no one ever talks about that is so great is, do you remember Bit? So Bit is the companion, and it's a Boolean, and it's represented by this kind of floating abstraction. But Bit is awesome because you can ask Bit a question and it just says yes or no, and it tells you true or false.

And it's interesting because in the world of "Tron," the Bit is like this underlying source of old Aristotelian-level truth. And I mean, there is no such thing in our world, and we know that from quantum mechanics that everything is a lie, but that's OK because we still have to live in the world and pretend like things can actually be true. And it's just interesting that even in 1981, they kind of posit this machine intelligence that literally can give you one answer, yes or no. And you think about how powerful that is, just the definitiveness of the answer.

And so it's interesting that people are working on non-LLM AI technologies. The way they might fit in with LLMs, I like your daughter card notion, I think makes a ton of sense. And I'm immediately reminded of this highly anthropomorphized Disney Fantasia Wonderland, and I wonder how many of the movies we watched as kids we’ll all end up reconstructing in our pursuit of this technology effectively.

Rob Kim

Well, as long as it doesn't go “Terminator”-wise, I think we're good, but who's to say?

Adam Blue

Yeah, fingers crossed. All right. Well, thanks for joining today, Rob. I think this was really great, and we really appreciate your enthusiasm and participation. So thanks everyone for watching Cut to Context. I'm Adam Blue, and you can find us whatever your top-quality podcasts are bought and sold. Thank you.

Rob Kim

Thanks.