The AI Kubernetes Show
The AI Kubernetes Show
0
The AI Kubernetes Show explores the practical challenges of running AI workloads on Kubernetes platforms. Episodes dig into real-world adoption stories, covering architecture decisions, tooling and operational pitfalls. Typical topics include GPU scheduling, model serving, autoscaling and resource management. The show is aimed at platform engineers, DevOps teams and machine learning practitioners working in cloud-native environments.
Epizódok
-
Is It Easier to Write a Vulnerability Than to Find One? 23.09.2026 38pStatic analysis tools used to be right about as often as a coin flip. Kathleen Goeschel, Principal Product Security Engineer at Red Hat, spent years using machine learning to fix that, and now she's watching large language models change both sides of the fight. How fast vulnerable code gets written, and how fast it gets found.Goeschel walks through how she trained machine learning models on aggregated scanner alerts, software metrics, and prior CVEs to cut false positives in static analysis, and the surprising finding that code churn alone predicted real vulnerabilities better than most of the signals built for that purpose. Read the blog post and follow us on LinkedIn! -
Durable Execution for AI Agents in Kubernetes 09.09.2026 44pIf your AI agent restarts and loses everything it was doing, your platform team is stuck retrofitting recovery logic Kubernetes was never built for. Mark Fussell, CEO of Diagrid and creator of the Dapr project, joins the AI Kubernetes Show to explain why that's the real infrastructure problem in the agentic AI era, not the model.Fussell built the platform running Azure SQL Server at scale during 20 years at Microsoft, then spent the last 9 years building Dapr (Distributed Application Runtime), now a CNCF-graduated project. He and host William Morgan cover durable execution and why he calls it "process reincarnation," why non-determinism and multi-minute latency break assumptions that held for a decade of microservices, and why he tells platform teams to translate what they already know instead of relearning it. TAKEAWAYS✓ Why durable execution, or "process reincarnation," matters more once language models are in the call path ✓ How to decide when to keep a workflow deterministic instead of routing it through a model ✓ Why non-determinism, not topology, is the real thing that changed since microservices ✓ How Dapr assigns identity at the process level using the SPIFFE standard, and why that matters for agents ✓ Why a log isn't the same thing as tamper-proof attestation ✓ Why Fussell says about 90% of existing platform engineering skills transfer directly to agentic systemsRead the blog post!FIND AND FOLLOW US ON LINKEDIN✦ LinkedIn: https://www.linkedin.com/company/the-ai-kubernetes-show/ -
Matt Barker (BoltMCP) on SPIFFE, AI Agent Identity, and Kubernetes 26.08.2026 46pAI-generated code is 40% more likely to leak a secret. Matt Barker, co-founder of the company that built cert-manager and later led secure workload access at CyberArk, explains why SPIFFE might finally fix workload identity for AI agents.Matt and William Morgan cover why SPIFFE replaces long-lived secrets with short-lived, cryptographic workload identities, what Matt is building at BoltMCP to control what an AI agent can see and do once it's trusted, and why CISOs are being told to enable AI adoption "at all costs."Topics covered:Why SPIFFE replaces secrets with short-lived workload identities (SVIDs)Where SPIFFE's job ends and human-delegated authorization beginsWhy "context window flooding" breaks MCP integrations at scaleWhy Kubernetes stays the default foundation for agentic AI infrastructureRead the blog post: SPIFFE Identity for AI Agents on KubernetesFind us on:YouTubeApple MusicLinkedIn -
Disaggregated Inference on Kubernetes 12.08.2026 56pEvery LLM inference request splits into two phases with completely different hardware profiles: pre-fill and decode. But most Kubernetes clusters still run them on the same pod. Morgan Foster, who works on the llm-d project and the AI Gateway Working Group out of Red Hat's Office of the CTO, joins William Morgan to explain what that means. Foster and Morgan get into why real-time inference workloads fit the primitives Kubernetes already has, while training often runs on Slurm instead. They walk through disaggregated inference: how pre-fill and decode get scheduled on separate pods, how KV cache blocks move between them over RDMA, and what happens today when a pre-fill or decode pod dies mid-request (nothing handles the failover cleanly, and the computed KV blocks just get thrown away). They also cover why faster token generation buys agentic systems more capability, not just quicker responses, and why Kubernetes proxies, built around header-based ingress routing, are struggling to govern AI traffic that depends on streamed JSON request bodies and inference happening outside the cluster. FIND AND FOLLOW US ON: ✦ Spotify: https://open.spotify.com/show/4LLTNjMbd7wNCYhSI9pkQU?nd=1&dlsi=3baba6bcc6b44cae ✦ Apple Music: https://podcasts.apple.com/us/podcast/the-ai-kubernetes-show/id1895650894 TAKEAWAYS✓ Why real-time inference workloads fit Kubernetes' existing primitives, while training often runs on Slurm instead ✓ How disaggregated inference splits the pre-fill and decode phases of an LLM request across separate pods ✓ Why KV cache blocks move between pods over RDMA, and what happens if the pre-fill or decode pod dies mid-transfer ✓ Why more tokens per second buys agentic systems more capability, not just faster responses ✓ Why Kubernetes proxies built for header-based routing struggle with AI traffic that depends on streamed request bodies ✓ What the AI Gateway Working Group is building to fix itRead the blog post at: https://www.buoyant.io/ai-kubernetes-episode/disaggregated-inference-on-kubernetes -
The State of Cloud Native AI and What’s Still Evolving 29.07.2026 48pThere is no agreed upon definition of what cloud native AI means yet, but if you're trying to figure out where Kubernetes AI standards are really headed, check out this episode with someone who's worked on the CNCF's first attempt to define cloud native AI. William Morgan talks with Ron Petty (RXM, CNCF founding contributor, and member of the CNCF's TCG AI) about what's actually settled in cloud native AI and what isn't. They cover the CNCF's shift from working groups to technical community groups (TCGs), Petty's own framing of cloud native AI as a two-way street between building models and using AI to run cloud native systems, and the 15-criteria self-reported conformance program (mostly GPU management) that gets you an "AI certified" stamp. TAKEAWAYS ✓ Why Ron Petty defines cloud native AI as a two-way street: building models on cloud native infrastructure, and using AI to run cloud native systems ✓ What the CNCF's self-reported AI conformance program actually checks, including GPU management and having an AI router on your cluster ✓ Why certificate management and rotation is still unsolved, and how that connects to why Linkerd defaults to mutual TLS ✓ What's actually wrong with Model Context Protocol's design, according to a proxy engineer's read on it ✓ Why long-running AI agents don't need fundamentally different guardrails than short-running ones ✓ How K8sGPT uses existing Kubernetes checks plus an LLM to help junior engineers diagnose cluster issues fasterRead the blog post at: www.buoyant.io/ai-kubernetes-episode/the-state-of-cloud-native-ai-and-whats-still-evolving -
Treat Testing as a Platform Service on Kubernetes 15.07.2026 48pMost testing tools bolt onto a CI pipeline and hope for the best. Ole Lensmar, CTO of Testkube, built a company that changes that script. Pull test execution out of CI/CD entirely and run tests as Kubernetes workloads instead.In this episode, Lensmar walks through why Testkube uses Kubernetes as its test execution engine, running Selenium, Playwright, K6, JMeter, and Postman tests as cluster workloads instead of asking teams to adopt new tools. He and podcast host William Morgan dig into who should own testing (developers own functional tests, while the platform team owns performance, security, and chaos testing). Lensmar's core recommendation is to treat testing and quality as a platform capability, the same way you already treat CI and CD. More AI-generated code means more tests are needed, intelligent test selection can keep pipelines fast as test counts grow, and AI is already useful for triaging failed test logs using Kubernetes and Grafana MCP servers. TAKEAWAYS: ✓ Where testing ownership splits between developers and the platform team, and where it doesn't ✓ Why "testing as a platform service" should get the same priority as CI and CD ✓ How intelligent test selection works, and why you should never rely on it alone ✓ Where AI already adds value today: triaging failed test logs with Kubernetes and Grafana MCP servers ✓ Why testing the AI components of your application (agents, evals) is becoming its own discipline.Read the blog post at: -
LLM Inference on Kubernetes: New Primitives, Real Challenges 01.07.2026 43pRunning LLM inference on Kubernetes requires new primitives for routing, autoscaling, and GPU scheduling. Here's what platform engineers need to know. -
Running Multi-agent AI on Kubernetes: Lessons from Imagine Learning 17.06.2026 48pIn this episode of The AI Kubernetes Show, Blake Romano, Staff Software Engineer at Imagine Learning, walks through what it actually looks like to build and run AI agents on Kubernetes at scale. He talks about the architecture choices, the failures, and why the organizational context you bring to the LLM matters more than which Software Development Kit (SDK) you use.Imagine Learning is a K-12 education company building digital platforms for students and educators, and Blake has been driving AI and platform engineering initiatives there. -
AI Agents Security: Guardrails & Production on Kubernetes 03.06.2026 41pLearn platform-level security patterns for AI agents on Kubernetes. Close the production gap with LLM guardrails, tool filtering, and short-lived tokens. -
Platform Engineering & Kubernetes: Guardrails For AI Code 20.05.2026 53pLearn how Schonfeld scaled their internal AI platform, SchonAI, using Kubernetes and established guardrails to manage AI agent code volume. Build your AI-native workflow now. -
One Dependency Away: Supply Chain Security in the Age of AI 06.05.2026 50pSecure your Kubernetes environment. Learn why zero trust cybersecurity is the only defense against AI agents and non-deterministic agentic software in your supply chain. -
Moving from Single Agents to AI Agent Fleets 14.01.2026 27pThe future of software development isn't about single agents—it's about building AI agent fleets! Dive into this conversation with Okteto CEO Ramiro Berrelleza to understand how this shift is fundamentally changing platform engineering and accelerating developer productivity. In this episode of The AI Kubernetes Show, we sat down with Ramiro to discuss AI adoption and the need for constant experimentation in the current "Cambrian explosion" of AI tooling. Berrelleza highlights the move from single-threaded AI tools to large, asynchronous AI agent fleets, which solves the bottleneck of waiting for a single AI response. This agentic model is a game-changer, with some early adopters seeing a massive increase in output. Organizations need to adapt for AI-native workflows, because the focus on traditional metrics like measuring code production (lines of code, number of PRs) for AI is flawed. Instead, organizations should identify and focus their AI projects on their real constraints, such as slow CI workflows. Ramiro also addresses the disproportionate challenge of open source maintainer overload caused by AI-generated contributions, proposing a policy of "human-proof code." Finally, AI agents are presented as a powerful technical context multiplier for everyone from sales engineers to the CEO, significantly speeding up the onboarding process and improving communication across the organization. Read the blog post: Takeaways✓ The future is moving from single-threaded AI tools to "AI agent fleets" to solve productivity bottlenecks.✓ Traditional metrics like lines of code or PR count are now ineffective for measuring AI-driven developer productivity.✓ The new focus for AI investment should be on organizational bottlenecks, such as optimizing slow CI workflows.✓ Open source projects should adopt policies like "human-proof code" to manage maintainer overload from AI contributions.✓ AI agents can serve as a technical context multiplier, speeding up onboarding and improving organization-wide understanding of complex code.Hit the like button, subscribe for more content on platform engineering and AI, and ring the notification bell. What is the biggest productivity bottleneck you've solved with AI agents? Let us know in the comments!#AIAgentFleets #PlatformEngineering #DeveloperProductivity #Kubernetes #KubeCon #Okteto #AgenticAI #OpenSource #SoftwareDevelopment #TechTrends -
Why Testing and Validation are the Unsolved AI Code Challenges 14.01.2026 27pIs your engineering org ready for the speed of AI? Grant Miller, CEO of Replicated, breaks down the intersection of AI and platform engineering, revealing why testing and validation are the biggest unsolved problems in the industry.In this episode of The AI Kubernetes Show, we sit down with Replicated CEO Grant Miller to discuss how the pace of AI is fundamentally reshaping software development. Miller argues that engineering velocity has become the core competitive differentiator and shares the concept of "leadership empathy," where leaders contribute to a pull request with AI to understand the new tools. This increased velocity, however, puts significant system pressure on platform engineering teams, leading to "Frankenstein-y" application footprints and a greater need for top-notch observability and optimized CI/CD pipelines to improve "iteration speed total."The unique distribution challenges of self-hosted AI applications and the difficulty of validating AI code generation, especially for templated infrastructure-as-code like Helm charts and Terraform. Unlike front-end code, the human validation loop for infrastructure-as-code is not intuitive, making the complexity of testing and validation the industry's most significant hurdle.Read the blog post: Takeaways✓ AI turns engineering velocity into the ultimate competitive advantage, requiring organizations to move incredibly fast.✓ Leaders must develop "leadership empathy" by using AI tools to understand the modern developer experience.✓ Rapid AI code generation can lead to complex, "Frankenstein-y" application architectures, increasing pressure on platform engineering for troubleshooting and observability.✓ The biggest challenge in AI-generated code is the lack of an intuitive validation loop for infrastructure-as-code like Helm charts.✓ Testing and validation are the key unsolved problems and future areas for discovery and job creation.Liked this podcast? Hit the like button, subscribe for more AI and platform engineering insights, and let us know in the comments: What is the biggest challenge your team faces with AI-generated code?#AI #PlatformEngineering #EngineeringVelocity #AIGeneratedCode #TestingAndValidation #Kubernetes #Replicated #TechPodcast #CloudNative -
AI: Bubble or Bug? A CTO’s Perspective on Engineering in the AI Era 14.01.2026 27pIs the AI boom a bubble, or is it a new technological wave? Dinesh Majrekar, CTO of Civo, breaks down the current state of software development, explains why data sovereignty is the paramount security concern, and details how AI's real value lies in increasing code auality, not just velocity.In this episode of The AI Kubernetes Show, Civo CTO Dinesh Majrekar tackles the AI bubble hype, suggesting it is a blend of market speculation and genuine, disruptive innovation, drawing a comparison to the historical hardware monopoly of IBM during the mainframe era. He dives into the challenge of data sovereignty in the age of large language models, explaining Civo's solution of using an "on-prem public cloud" to run an OpenAI-compatible endpoint on private GPUs. This approach ensures maximum security for sensitive data, like medical records, by guaranteeing the data "never leaving your building." We also discussed the flattening curve of open source LLM capabilities, noting that models like the Kimi K2 model are now matching and even beating proprietary benchmarks while using fewer resources.Majrekar challenges the prevailing focus on speed, arguing the true value for software development teams is in boosting code quality. He champions code generation as the best AI use case but stresses it must be a "partnership" where saved time is reinvested in tackling technical debt and strengthening the code base. This is important for managing deployment risk. Finally, he addresses the dilemma of non-deterministic outputs in deterministic processes, which engineers simply call "a bug," emphasizing that AI is not a universal solution.Read the blog post: www.buoyant.io/ai-kubernetes-episode/ai-bubble-or-bug-a-ctos-perspective-on-engineering-in-the-ai-eraKey Takeaways✓ Code Quality is the true benefit of integrating AI; the time saved on initial generation should be used to fix technical debt and strengthen code.✓ Achieving true Data Sovereignty requires running LLMs on private infrastructure (e.g., an on-prem public cloud) to keep data securely contained.✓ The non-deterministic outputs of LLMs can be considered a "bug" in core engineering processes that demand algorithmic certainty.✓ Code generation is the strongest AI use case, but developers must maintain ownership and set a high context standard for the LLM to follow.✓ Open source LLM capabilities are now "on par" with proprietary models.Hit the like button and subscribe to The AI Kubernetes Show for more AI content! What is your engineering team prioritizing with AI: velocity or quality? Let us know in the comments below!#AI #CodeQuality #DataSovereignty #SoftwareDevelopment #PlatformEngineering #Kubernetes #LLM -
Maintaining DevOps Integrity in the Age of AI Velocity 07.01.2026 27pIs AI velocity breaking your DevOps processes? We sat down with Principal Platform Engineer Ahmed Bebars to discuss the critical balance of using AI to ship code faster while maintaining DevOps integrity in your platform engineering team.Ahmed Bebars, a principal platform engineer and CNCF ambassador, breaks down the AI's impact on SDLC and the transformation of platform engineering. In this The AI Kubernetes Show episode, he argues that AI is a powerful productivity tool that increases velocity, but teams must uphold established DevOps processes, including rigorous integration and regression testing, to prevent disruption and ensure quality. We discuss how Large Language Model (LLM) output is directly tied to input context, leading to the highly favored concept of spec-driven development, where humans guide the AI with precise specifications.The discussion also explores the challenge of adopting a new mindset for the non-deterministic nature of LLMs. Ahmed explains the difference between deterministic vs. non-deterministic LLMs and how tooling can be used to make the outputs predictable. For leaders looking at preparing platform engineering teams for AI, he advises embracing the technology, starting small, and focusing on hosting local LLMs and building agentic workflows for use cases like incident response triage and data gathering. Read the blog post: www.buoyant.io/ai-kubernetes-episode/maintaining-devops-integrity-in-the-age-of-ai-velocityFollow us on LinkedIn: https://www.linkedin.com/company/the-ai-kubernetes-show/Takeaways✓ How to use AI to increase code velocity without compromising established DevOps processes.✓ The critical role of context and spec-driven development for high-quality LLM output.✓ Understanding the engineering mindset shift required to work with non-deterministic LLM outputs.✓ Practical advice for preparing platform engineering teams for AI through local knowledge and agentic workflows.✓ The broader role of AI in the development ecosystem, including documentation, testing, and observability.If you found this discussion valuable, please like this video, subscribe for more insights into The AI Kubernetes Show, and hit the notification bell! What's the biggest challenge your team faces with integrating AI into your SDLC? Let us know in the comments below! #AI #DevOps #Kubernetes #PlatformEngineering #SDLC #SpecDrivenDevelopment #CNCF #KubeCon #TechTalk #OpenSource -
AI's Double-Edged Sword: The Technical Imperative and the Path to Accessibility 07.01.2026 19pDiscover the dual nature of AI! Tech Lead Chris Khanoyan shares his view on the rapidly changing AI and data science landscape and the critical need for a technical foundation and the transformative power of AI accessibility for the deaf community. In this episode of The AI Kubernetes Show, we dive deep into the world of AI and data science with Chris Khanoyan, a tech lead and senior data scientist at Booz Allen. Chris highlights the rapidly changing data science landscape, noting the significant overlap between data scientists and data engineers. While auto-generated code has made coding more accessible to practically anyone, he stresses that a solid technical foundation remains critical for debugging and understanding the fundamental elements of a system.We covered the foundational challenge of data governance and the need for clean, trustworthy data. Chris explains the importance of establishing a data pipeline and provenance (where the data comes from and who owns the dataset) before training any Large Language Models (LLMs). He offers a core principle for starting any project: begin with the end in mind. We also explore the hurdles of overcoming data access and scarcity, which often require formal agreements with non-technical clients, especially in sectors like the federal government. Finally, as a deaf individual, Chris provides a unique perspective on AI accessibility. He discusses how AI assistance is easing the mental fatigue from constantly processing captions and the potential game-changer of AI-powered glasses for live captions, while also addressing the current security and data sensitivity barriers that prevent their widespread adoption.Read the blog post: www.buoyant.io/ai-kubernetes-episode/ais-technical-imperative-and-the-path-to-accessibilityFollow us on LinkedIn: https://www.linkedin.com/company/the-ai-kubernetes-show/Takeaways✓ A solid technical foundation is still vital for practitioners to manage bugs, even with the rise of AI code generation.✓ Data governance and establishing data provenance are primary challenges in successful AI implementation.✓ AI projects must always begin with the end in mind to effectively prepare and utilize data.✓ Workarounds for data scarcity involve combining and consolidating various datasets from different systems (on-prem and cloud).✓ AI accessibility tools, such as live captioning on glasses, offer a significant boost to productivity and ease mental fatigue, though data security remains a critical barrier.If you enjoyed this conversation on the technical imperative of AI, hit the Like button and subscribe for more expert interviews! Let us know in the comments: What is the single biggest data governance challenge your team is facing today? #AIandDataScience #DataGovernance #AIAccessibility #TechLeadInterview #DataProvenance #LLMData #GoogleCloud #DataEngineer #TechInterview #MachineLearning -
The AI Tug-of-War: Bridging the Divide Between Platform Engineering and Data Science 07.01.2026 25pKeith Maddox, co-lead of the Kubernetes AI Working Group, breaks down the architectural shifts and security challenges required to run enterprise AI agents at scale.In this The Kubernetes AI Show episode, we chat with Keith Maddox, senior principal software engineer lead at Microsoft and Istio maintainer, who shares his perspective on the convergence of data science, AI agents, and platform engineering on Kubernetes AI workflows. He details the organizational dissonance between traditional platform stacks and data science workflows and how the Kubernetes AI working group is working to create a seamless migration path. We cover advanced model specialization techniques like Low Rank Adaptation (LoRA) and Retrieval-Augmented Generation (RAG), which are crucial for enterprise use cases driven by data privacy and liability concerns.Maddox also provides advice for platform owners, including the technical and non-technical strategies for LLM token spend management—recommending an egress gateway to centralize policy—and the importance of customer empathy with application developers. A major focus is the AI agent identity security gap, which falls between traditional human and machine identities. He strongly advocates for a zero trust AI mindset and immediate mitigation through agent sandboxing (using technologies like gVisor, KVM, or Wazet) and short-lived, ephemeral machine identities to manage the non-deterministic nature of LLMs.Read the blog post: www.buoyant.io/ai-kubernetes-episode/the-ai-tug-of-war-bridging-the-divide-between-platform-engineering-and-data-science Follow us on LinkedIn: https://www.linkedin.com/company/the-ai-kubernetes-show/ Key Learnings✓ The core conflict is a "tug of war" over tech stacks between platform and data science teams.✓ Model specialization is necessary due to the high cost and lack of specificity of foundational models for enterprise applications.✓ Managing LLM costs requires centralizing policy through an egress gateway and open communication with development teams.✓ AI agents pose a new security challenge, requiring a move toward short-lived, ephemeral machine identities and agent sandboxing.✓ A "Zero Trust" mindset is the recommended security approach for non-deterministic AI agents and workflows.If you're building, deploying, or securing AI workflows, hit the Like button and subscribe for more deep-dive technical content! Let us know in the comments: What is the biggest challenge your team is facing with AI agent identity and security today? #PlatformEngineering #Kubernetes #AIAgents #LLMs #ZeroTrustAI #KubeCon #DataScience #TechSecurity #DevOps -
Platform Engineering, AI, and the Two Faces of Cloud Native Sustainability 30.12.2025 11pKubeCon 2025 proved that platform engineering is essential, but the next big challenge for the cloud native community is sustainability. Learn from KubeCon Co-Chair Faseela what this future looks like and what it means for your platform.In this episode of The AI Kubernetes Show, we break down the key KubeCon NA 2025 takeaways with co-chair Faseela, a cloud native developer at Ericsson. She discusses the immense success of the Atlanta event and highlights the convergence of AI and Kubernetes, noting that platform engineering and AI were the hottest tracks based on talk submissions. Looking ahead, the focus shifts to designing platforms that are secure, cost-efficient, and sustainable.Fazila explains the "two-sided coin" of sustainability in the cloud native community: environmental sustainability—meaning engineers must actively consider how to design platforms to be optimized and efficient to minimize cloud native environmental impact—and human sustainability, which is all about tackling CNCF maintainer burnout. She details how the CNCF's TOC performs health checks on open source projects and encourages companies that are making use of that project to put resources under that project. This deep dive into the future of cloud native is essential viewing for all software engineers interested in responsible collaboration.Read the blog post: URLFollow us on LinkedIn: https://www.linkedin.com/company/the-ai-kubernetes-show/Takeaways✓ Platform Engineering and AI were the most popular technical tracks at KubeCon North America 2025 based on submission volume.✓ The future of cloud native requires engineers to design platforms that are secure and cost efficient and sustainable.✓ Sustainability in cloud native is two-fold: minimizing environmental impact and ensuring the "human sustainability" of the open source community by preventing maintainer burnout.✓ The CNCF's TOC performs regular health checks and encourages companies to dedicate resources to under-maintained open source projects.If you're building in the cloud native space, hit the like button, subscribe for more interviews, and leave a comment below! What is the biggest challenge your team faces when trying to implement a sustainable platform?#Kubernetes #CloudNative #PlatformEngineering #KubeCon #CNCF #Sustainability #AIandKubernetes # -
Stop Bolting on Security: The Key to Reliable AI Agent Systems 30.12.2025 16pIs your AI infrastructure safe? Marina Moore, research scientist and co-chair of CNCF Tag Security, talks about her research on AI agent isolation and how to build robust platform engineering security. Build it securely from the start!In this must-watch episode of The AI Kubernetes Show, we sit down with security expert Marina Moore to discuss the paradigm shift in AI-driven systems security. Moore shares her latest research on Securing Autonomous AI Agents by applying a "decompose" approach, which breaks down complex tasks into smaller pieces of work and enforces a security boundary with gated pathways for data flow. This strategy is a pragmatic solution for AI agent isolation and surprisingly results in minimal security performance overhead because the LLM inference processing is the system's slowest part.The goal is to build security in from the start rather than trying to "bolt on" security later. Moore explains how the CNCF Tag Security assessment process leverages a core threat modeling question—listing system actors and data flow—to help projects improve their architecture early. This discussion is for anyone involved in cloud native security assessment and the future of secure AI development, including actionable Kubernetes security best practices advice for both platform engineers and software developers.Read the blog post: www.buoyant.io/ai-kubernetes-episode/stop-bolting-on-security-the-key-to-reliable-ai-agent-systemsFollow us on LinkedIn: https://www.linkedin.com/company/the-ai-kubernetes-show/Takeaways✓ Security can be practically applied by breaking down autonomous work into smaller, isolated agents with secured, gated data flow.✓ Adding security layers has a small impact on performance because the LLM tool calls and inference processing are the primary system bottlenecks.✓ The simple act of enumerating all system actors, data flows, and potential attack vectors is a critical self-assessment that illuminates hidden connections for better design.✓ Integrating security early in the development lifecycle is more efficient and enhances overall system reliability compared to bolting it on at the end.✓ Focus on designing a "secure by design" infrastructure and establishing a secure baseline to enable safe experimentation with non-deterministic AI systems.If you found this video valuable, hit that like button, subscribe for more AI security content, and hit the notification bell! Let us know in the comments: What is the biggest platform engineering security challenge you are facing with AI agents today?#AI #Security #Kubernetes #AIAgentIsolation #CloudNativeSecurity #ThreatModeling #PlatformEngineering #CNCF #DevSecOps #LLMSecurity -
Agentic AI Security on Kubernetes: Understanding Payload Routing and Defense-in-Depth 24.12.2025 1óIs your Kubernetes cluster ready for the next generation of AI workloads? Join Buoyant Tech Evangelist Flynn and Red Hat's Shane Utt from the AI Gateway Working Group as they reveal the critical shift from header-based to payload-based Kubernetes Networking for Agentic AI.In this episode of The Kubernetes AI Show, we dive deep into the challenges of running AI workloads on cloud native infrastructure with AI Gateway leaders, Flynn and Shane. They discuss how the AI Gateway Working Group evolved from the Gateway API inference extension (GIE) to focus on the average Kubernetes user who is often performing inference via egress outside the cluster. A key challenge is that AI models are fundamentally different from microservices—they are not interchangeable, making advanced routing a must.The conversation highlights a critical shift: networking infrastructure must now be retrofitted for payload processing for routing to ensure security, compliance (like GDPR), and proper model selection. We explore parallels between AI’s hardware demands and High Performance Computing (HPC), revealing the problem of low LLM hardware utilization and the talent gap between data scientists and network engineers. Finally, the co-leads detail the security implications of Agentic AI's autonomy, advocating for a defense-in-depth security strategy that includes multi-cluster routing and Model Context Protocol (MCP) servers with 'elicitation' to prevent accidental, destructive actions.Read the blog post: www.buoyant.io/ai-kubernetes-episode/agentic-ai-security-on-kubernetes-understanding-payload-routing-and-defense-in-depthFollow us on LinkedIn: https://www.linkedin.com/company/the-ai-kubernetes-show/Key Learnings ✓ Learn why AI models require advanced, non-interchangeable routing logic compared to standard microservices.✓ Discover how the shift to payload processing is essential for semantic routing, security, and GDPR compliance.✓ Understand why AI’s low hardware utilization (20-30%) is a known problem from High Performance Computing (HPC).✓ See how Agentic AI introduces a "deeply terrifying" security risk that demands a defense-in-depth security approach like MCP elicitation.If you're building or securing AI workloads on Kubernetes, you can't afford to miss this discussion! Subscribe to The Kubernetes AI Show for more foundational insights, like this one from the AI Gateway leaders. Hit the like button and let us know in the comments: What is the most critical challenge you are facing when building Kubernetes Networking for your Agentic AI applications?#KubernetesNetworking #AIGateway #AgenticAI #CloudNative #GatewayAPI #LLMs #K8s #Microservices #TechEvangelist
Népszerű itt:
Ez a podcast ezeknek az országoknak a podcast-listáin is szerepel.