TRANSCRIPT · CC BY 3.0

Cloud Native Live Fireside Chat: Building AI agents at scale on Kubernetes

How Intuit runs 100+ AI agents on GenOS atop its Kubernetes platform: what differs from a normal service, how agents are evaluated, and its SRE agents.

CNCF [Cloud Native Computing Foundation] · Published · 24 min · English · License: CC BY 3.0 · Source: watch on YouTube

Transcript source: automatic speech recognition on Vidleaf (unedited, may contain errors). Paragraph breaks and timestamps added by Vidleaf.

[0:00] Thank you. Alrighty, welcome everybody to our inaugural cloud native live fireside chat. My name is John O'Bacon. I'm the founder and CEO of State Shift. I'm really excited to do this today. We've got some amazing guests on this session. Make sure you are responding in the chat and engaging to make this as fun and as interesting as possible. Now, Today, we're going to be talking to some friends of ours at Intuit. Now, if you're not familiar with Intuit, Intuit has been consistent end-user members of the CNCF for a long period of time. Probably best known within the CNCF community as the creators of Argo, who donated Argo to the CNCF, and have been very involved in the AIOps and the MLOps communities for quite some time. What's really interesting here is that Intuit, has been on this journey to move all of their products, TurboTax, QuickBooks, etc., to become AI-native products. And what we're going to be digging into today is how Intuit has built its platform to build agents at scale, expanding from cloud-native into AI-native. OK, we've got two guests who are going to be speaking to you today. We've got Mukulika Kapas, who's director of product management at Intuit. How are you doing, Mukulika? Doing well?

[1:09] Awesome. Great. Thank you. Great to have you. And we've also got Marin Kurian, who's the distinguished engineer into it as well. Great to see you, Marin. How are you doing? Thank you, sir. Alrighty, so why don't we dig right into this? So let's get started. Like we're talking about AI agents, but how do you define an AI agent? What has been into its experience in this space so far? Yeah, I can take that. So before we go into agents, I want to quickly call out the difference between a workflow and an agent. So a workflow is when you have a predefined set of steps. These steps are executed in a sequence. Whereas AI agents can be used for solving open-ended problems where it's more difficult or even impossible to predict the sequence of steps or the required number of steps even. And you can't hard code a fixed path.

[1:54] So this is where we have seen in examples at Intuit where agents are usually successful. It's about working with unstructured data such as text, images, documents, speech. And going back to how do I define an agent, agent has key components. Primarily they have a GNI model. It can be a collection or variety of models, large to small. They have tools which are services or APIs which provide these agents both the data to load into the context and also the APIs and services to perform actions. The key component in the agent is then an orchestrator which manages states, leverages memory, past interactions, and resends with the tools available at the agent's disposal and the LLMs themselves to develop plans and execute so that they can achieve the goals. And finally, these AI agents need to be deployed on a safe and secure runtime and is activated either by a user request or through some other triggers in the environment. So at Intuit, we have been developing agents for the last two years or so. Our first generation agents were conversational assistants, so it will do data fetching, answer a question you have.

[3:06] what we are calling done for you experiences so these agents can act on your behalf execute tasks when permitted um what is different i think in industry is that we have enabled all of our engineers to develop agents so this is part of a transformation for us so along with the first set of agents we also launched our first version of the platform which we call genos short for generative ai operating system so teams can focus on solving customer problems rather than handling talks like responsibility, data governance, security, compliance, privacy, and observability. All of these are taken care of by the platform. Now, we didn't start from scratch. Huge thanks to Mukulika and team. All of these were developed on top of into its core development platform, which provided us the paved roads and the runtime guarantees. And that's how we were able to launch this for the velocity we need to scale.

[4:00] Amazing. Well, thank you. And speaking of Mukulika, who is a legend in this work, I'd like to kind of discuss a little bit about the development journey that you've been on at Intuit. Mukulika, could you share a bit about kind of the technologies you use, the scale? Like, why did you build it? Sure. So to take you through Intuit's cloud native journey and then AI native journey, in 2011, Intuit began its public cloud journey. They were one of the first companies to move to AWS, mainly doing lift and shift from data center. In 2018, we rolled out an Intuit-wide cloud native development platform for building and running backend microservices and frontend micro apps.

[4:47] This gave us the base infrastructure for building and deploying applications and services in a standard, repeatable and secured way using containers while sending all the operational data to a central operational data lake. for micro front-end app. The platform integrates with vendor solutions like, say, GitHub Artifactory, as well as key AWS cloud services. As you know, we open sourced Argo as a part of this journey and hence won the CNCF End User Award twice for our usage as well as contribution. It's no surprise as well as it enabled rolling out operational excellence standards and best practices that Merin talked about.

[5:56] leading to five nines of service availability across the company. Internally, we call this development platform Modern SaaS Platform. To increase our development productivity even further and to enable agents and AI at scale within our product experiences, starting 2022, we went from more cloud-native platform to AI-native development platform. Not only use AI for, say, code assistant in the inner loop leading to, say, 40 percent faster coding on average, but also use AI in the outer loop, say, for auto scaling of our services and agents. We use AI in operations like anomaly detection and observability and now SRE agents that we'll talk about soon.

[6:56] service API PaveRoad, as well as now agent development PaveRoad called GenOS, which provides standardized code templates, CICD, out-of-the-box observability, key integration, etc., etc. For example, for agent, it integrates with eval and LLMs. For regular service, say it integrates with just database. We recently open sourced another new project, NUMA Flow, to simplify event-driven or async agent and service development, which is also a pretty cool project.

[7:28] Now let's talk about scale. All backend microservices and apps are built and run at scale across 350 plus Kubernetes cluster with greater than million CPU cores. Today we have around 3000 plus services, hundreds of apps, tens and thousands of data pipelines and now 100 plus agents in production on this common platform. processing platform, batch processing platform, and now agent development platform on this common development platform. To summarize, the value the platform provides is powering more done-for-you product experiences using AI, accelerating development velocity, and finally enabling operational excellence across the company.

[8:22] Amazing. Amazing. What I love about this is a lot of people are talking about agents and how to build agents online, but you're doing it like you're living it, you're breathing it, you're doing it at scale. Maren, I'm curious in your position as a distinguished engineer, you know, how is building an agent any different to building another application or a service on your development platform? Yeah, good question. So fundamentally, it is an application or a service. It is the extra work that we need to do on top of building the service that makes the agent more effective and better for user experience. So then what is common is all the identity access management, runtime infrastructure, end-to-end tracing, observability, integrations with user experience, analytics, all of these.

[9:07] are required. They need to be secure, compliant, and reliable. All of that remains the same. So the extra work that I talked about are agents need large language models. So now we are getting into the realm of working with AI, where things are mostly non-deterministic. So depending on the number of elements you need to work with and the number of prompts, then you are getting into prompt engineering, optimization, evaluation. Agents need non-overlapping tools in order to be able to make decisions effectively. So these are natural language description of services optimized for the context of the agent. And you would probably be hearing context engineering a lot these days. So agents also need to optimize the context with which they need to work so they don't hallucinate, they don't go out of turn, out of sync. So all of these are individual work streams on their own. So this is a lot of the extra work do on top of developing a regular service and application. So that's just the coding part. Now, the hardest and most difficult in developing an agent is

[10:13] The evaluation step. So as agents have non deterministic outcomes and also you didn't write all the code. Right. Nobody actually wrote all the code. So it's harder to debug the test. Even the testing criteria of an agent is extremely subjective, because you are working with natural language like English at the time. Right. So previously you would have heard, oh, it works for in my machine. Now I'm hearing it works for my data. It works for my questions. So being. creating an objective test criteria and doing systematic and structured evaluation that's another step that you need to do in addition to you know making your agents work like a regular software and also ai solutions are about continuous iterations to solve the problem so we need the end-to-end data pipeline moving data from user integrations experiences traces logs back to these improve over time. So this specific aspect of evaluation and data management is new for a lot of software engineers. So I think

[11:21] That is a huge leap from what is a good software engineer to a good AI engineer. Right. And this is a lot of work. And we know that. Right. That is why we build GenOS. Right. We want to streamline and help people. uh move fast through this end-to-end life cycle of developing agent testing them deploying them continuously monitoring them and improving them so based on our own learning for the first two years This spring, we launched Agent Starter Kit, which further optimizes in the whole CI/CD or the lifecycle of managing the development of an agent. So we came up with the starter code, the CI/CD process and all the best practices and capabilities baked in.

[12:03] So you are not going to have to read hundreds of documents, talk to 10 different teams. It's all available for you in a way you can reference and then build your own agent. And we are also, again, given the importance of the space, we are also investing in tutorials, trainings and whatnot. providing standard pipelines. So again, our goal is to help teams move faster, through the life cycle so they actually go to production and not get stuck in the POC phase. Yeah, I love it. And by the way, everybody, just I want to take a moment before we go on.

[12:36] There's a lot of people talking about how to build agents on LinkedIn. A lot of people honestly don't know what they're talking about. Whereas Mookalika and Merin are doing it, they're living it, they're breathing it. Right. So get your questions in. This is a great opportunity to kind of pick the brains of people who are really doing this at scale in a significant company. So if you've got questions, get them in. It doesn't matter how what your question is. It could be a really simple question. It could be a really complex question. Just get them in. So we can tap into their experience. All right. So you mentioned Gen OS and you mentioned agent starter kit, but how has this enabled it into it? Like, you know, what would you say are some of the outcomes there?

[13:13] Yeah, I can start and then Mugli can add more. So remember, Induit has only about 8000 engineers and more than 1200 of them are now developing agents using genos with varying levels of maturity and so far we have been able to launch 3 500 plus experiments in production our customers have tested and given feedback and we get on a typical day for 50 000 requests on genoa's and in august alone the agents on genos used 4 trillion tokens so we have been on this journey um as i was saying we launched agent In that week alone, we got 900 downloads.

[13:54] and 100 plus teams were able to do a demo during the demo day, within five days. So we have significantly brought down the learning curve, the barrier to entry, so people can actually get started. And if you have an idea, then you are able to get going. So powered by the agent starter kit, QuickBooks launched a series of next generation agents done for experiences this summer. So we have a whole suite of agents working for customers in our QuickBooks platform. I just wanted to call out two of them. So our accounting agents saves about 12 hours per month for small business in helping them keep their books organized.

[14:40] screenshots which you've taken all of that and converts them into estimates, which the businesses can then send us invoices, which has helped them get paid five days faster with a 10% more likelihood to get paid in full. Just to add to what Marin said, Intuit is also redefining the mid-market ERP landscape with Intuit Enterprise Suite, a modern AI native ERP solution with built-in automation for, say, multi-entity management, instant AI-driven insights, and a virtual team of AI agents. What we are hearing already from our customers, 78% of the customers are saying they're getting more time to grow their business. They can focus more on the business than managing they have a better picture of their financial health. So we are already seeing value and ROI out of these agents.

[15:39] Amazing. Speaking as a QuickBooks user who wants to get paid quicker, I'm all in favor of this. So, you know, we've seen one final question before we get to some of the viewer questions. We've seen the rise of agents within the infrastructure space with the likes of, you know, SRE agents. But how are you thinking about incorporating agents in the platform engineering side of things? So like everybody else, we first rolled out multiple AI-powered code assistance for code generation, test generation in the inner loop of development. But now, agentic development is changing the whole product development lifecycle. Today, over 30% of developers' time is spent in outer loop, troubleshooting issues, managing infrastructure, managing incidents.

[16:31] Despite progress, change induced incidents still remain high, even within Intuit, and recovery times are inconsistent. So last few years, we focused on customer centric observability using golden signals, real user monitoring and distributed tracing on top of our central operational data lake. Because you have to remember, agents first need clean data. to work, just like any other AI. We started with implementing near real-time anomaly detection on top of that data for observability and progressive delivery first before going into the agentic space.

[17:15] Now that we have a matured operational data and data pipelines and end-to-end observability, we first rolled out Incident Recovery Agent, which during an incident traverses our asset dependencies because incidents are not just caused by one software asset. It can be caused by a downstream asset impacting multiple assets. are very important to identify what's the root cause of the incident step by step with developer feedback loop to enable faster incident recovery. Today, the agent is already being able to find the root cause of more than 50% of our incidents. As Merin said, we constantly do the eval of the agent with real incident data. Other areas where we are experimenting with agents in the platform are agents doing tasks asynchronously trying to remediate any security defects that our security team opens across the company. Or upgrade, say, a tech date, like upgrading Java version. Asynchronous agents going to your repo, finding dependencies, and doing the upgrade. We are doing on-call alert triaging. On-call engineers get lots of alerts, and to triage and figure out which one to take action and which one not to.

[18:44] Build and deployment failure. Our biggest developer question in support slacks is why my build failed and why my deployment is giving crash to back off. OK, I like triaging and insights and then developer support. We are trying to answer majority of our developers questions with conversational agents before sending them to the second tier of support. Yeah. Amazing. So we've got a question here from somebody who was a mystery. They did not add their name. So I'm going to assume it's obviously Brad Pitt.

[19:18] who said how can i add ai agents into kubernetes clusters to help me fix issues So we have platform engineers who maintain the Kubernetes clusters and obviously then the service developers who use the Kubernetes cluster to to build their own agents or their own services. So our platform engineering team are experimenting with agents first to triage issues because we need to be 100 percent accurate in triaging issues with human in the loop before making deterministic production changes.

[19:54] in the Kubernetes clusters. Amazing. And then we got a question here from Larry. I love this question. He says, I'd love to hear about the economics of the platform. One million plus cause is huge. literally and figuratively, the ROI must be there to make it viable for companies thinking about adopting agents. How do you balance the investment versus the results? Thanks. I'll take a first shot and then Merin can add. So just to clarify, one million plus cores was the overall platform on which we run services, APIs, data pipelines, web apps and agents. We are not yet running one million plus cores with agents because of every agent that we are running in production. The first thing we are measuring is the accuracy of the answers that they are giving to the customers.

[20:46] very careful to run them in production with especially giving answers to our customers. ROI, we already mentioned like how our customers are getting money back faster, how they're saying they're getting time back to invest in their business. Merin, do you want to add anything else? Yeah, just exactly what Mukulika said, right? We encourage a lot of experimentation. Doesn't mean always ruthlessly prioritizing what must get the investment, which provides a boost benefit for our customers.

[21:23] Love it. Progressive perfection. For example, our SRE agent, we are tracking that if it takes 20 minutes to find a root cause of the issue for a human, less than five minutes for an agent. And so far for 50% of the cases, the agent is finding within less than five minutes. Amazing. Amazing. And then I think this is probably going to be our final question because we're running out of time here. This is from Jan Harvey. I hope I'm pronouncing your name correctly. What are the first practical steps for. an individual developer can take to start building and deploying AI agents on Kubernetes, especially if they're new to cloud native tools and orchestration.

[22:00] Mary, you want to take it? Yeah, I'm assuming this is not about building infrastructure agents, but generally agents, but on top of Kubernetes. It is like how you would develop any agent. Don't worry about infrastructure and cloud first. Figure out how to build an agent effectively in your machine and then slowly build up on on whatever infrastructure you have available in your organization. Bye. And I think the fundamental difference, like I called out, is that this is a lot of software engineers' first experience working with AI. So that is the place where you need to upskill yourself and invest more. The infrastructure will follow. The tooling will follow. If last two years is any indication, things will be better. The tooling will get better. There will be a lot more streamlining. But the skill that you need to invest in yourself is how to leverage AI effectively. That you have to do on your own.

[22:55] Amazing. Well, Mukulika, Mehran, thank you so much for joining us today. This was the inaugural Cloud Native Live Fireside Chat. Thank you for joining us today. Everybody, go and check out the Intuit open source page on LinkedIn. You can see the link right there on the screen, and we'll hopefully have some people in the background maybe pop a link into the chat. YouTube and then stay tuned for future sessions as well. So thanks everyone. Have a great week and we'll see you on the next one.

Open in the Vidleaf workbench

Search the transcript, select lines, copy quotes with timestamps, translate.

Open in the workbench →

Attribution

"Cloud Native Live Fireside Chat: Building AI agents at scale on Kubernetes" by CNCF [Cloud Native Computing Foundation] (https://www.youtube.com/@cncf), licensed under CC BY 3.0 (https://creativecommons.org/licenses/by/3.0/). Source video: https://www.youtube.com/watch?v=iIzEfS5ThBA. This page is a text transcript of the video with paragraph breaks and timestamps added; the creator is not affiliated with and does not endorse Vidleaf.

Are you the creator or a rights holder? Request a correction or removal: copyright@vidleaf.app (see About these pages).

Last updated