KubeFM

KubeFM

KubeFM
Maa Yhdysvallat
Genret Teknologia
Kieli EN-US
Jaksot 100
Viimeisin 15.09.2026

KubeFM is a podcast dedicated to the Kubernetes ecosystem. It covers notable developments, expert opinions, and real-world experiences from teams running Kubernetes at scale. The show explores both successful strategies and common pitfalls, offering listeners practical insights and sometimes controversial viewpoints.

Jaksot

  • Building Platforms for AI Agents, with Mauricio (Salaboy) Salatino 15.09.2026 28min
    Non-deterministic agents pose specific challenges for platform teams in observability, state management, governance, and trust.Mauricio (Salaboy) Salatino explains why agentic applications behave like distributed multi-agent systems. His test assigned agents to take an order, cook the pizza, deliver it, and charge the customer. One order crossed 15 containers and produced 200 traces.In this interview:Why agent frameworks can recreate monolith scaling and resource contentionHow OpenTelemetry data can measure agent behavior and trust over timeWhy narrowly scoped agents are safer than agents that follow long sequencesHow the platform can become the learning layer that feeds better context back to LLMsSponsorKubernetes moves too fast to track everything. Learn Kubernetes Weekly filters out the noise to deliver one curated email with useful articles, tutorials, tools, jobs, events, and CFPs. Subscribe to Learn Kubernetes Weekly.More infoFind all the links and info for this episode here: https://ku.bz/TlVjXdnb6Interested in sponsoring an episode? Learn more.
  • Kubernetes Can Run Your Database. Your Team Can't., with Kat Cosgrove 08.09.2026 41min
    KubeSelect is a new show that tests one bold Kubernetes hypothesis with one expert. Hosts Salman Iqbal and Bart Farrell examine the evidence and ask whether the claim holds up.In episode two, Kat Cosgrove tackles a long-running question: should teams run databases on Kubernetes? StatefulSets, persistent volumes, CSI, and database operators changed the technical answer, but they did not remove the operational risk.In this episode:How Kubernetes storage has matured and why old assumptions persistWhere Kubernetes responsibility ends, and database responsibility beginsWhy operators do not replace database expertsWhen managed database services remain the better choiceSponsorKubernetes moves too fast to track everything. Learn Kubernetes Weekly filters out the noise to deliver one curated email with useful articles, tutorials, tools, jobs, events, and CFPs. Subscribe to Learn Kubernetes Weekly.More infoFind all the links and info for this episode here: https://ku.bz/7yDWlP8T5Interested in sponsoring an episode? Learn more.
  • Observability Won't Save You, with Henrik Rexed 01.09.2026 31min
    Kube Select takes one bold Kubernetes hypothesis and tests it with an expert.In the first episode, Salman Iqbal and Bart Farrell ask whether most organizations have an observability problem or a decision-making problem.Henrik Rexed, Senior Staff Engineer at Dynatrace, challenges the original claim. Teams can collect metrics, logs, and traces, but that telemetry needs system relationships, ownership, and deployment history to support a confident decision.In this episode:Why telemetry without context slows incident responseHow SLOs, ownership, and deployment events guide troubleshootingWhen more metrics increase cost without improving decisionsHow AI agents can investigate incidents without replacing human judgmentSponsorKubernetes moves too fast to track everything. Learn Kubernetes Weekly filters out the noise to deliver one curated email with useful articles, tutorials, tools, jobs, events, and CFPs. Subscribe to Learn Kubernetes Weekly.More infoFind all the links and info for this episode here: https://ku.bz/848pmBN5_Interested in sponsoring an episode? Learn more.
  • Why Kubernetes Needs to Learn GPUs, with Saiyam Pathak 25.08.2026 32min
    Kube Signals starts where the keynote ends: with the trends that platform teams will have to operationalize next.In this special episode, Brian Teller speaks with Saiyam Pathak about his KubeCon India keynote and the shift from developer platforms to AI factories. They examine what GPU scarcity, shared accelerators, and AI workloads mean after the conference slides meet real infrastructure.In this interview:Why GPU infrastructure is becoming a platform-engineering concernHow DRA, HAMI, MIG, and MPS change GPU allocation and utilizationWhere isolation, scheduling, and observability become harder for AI platformsWhich cloud-native AI trends and projects platform engineers need to watchSponsorThis episode is sponsored by LearnKube. Download the free book, The Technical Guide to Kubernetes Rightsizing, to understand what Prometheus and Grafana cannot tell you about safely reducing requests and limits.More infoFind all the links and info for this episode here: https://ku.bz/4QZDqrnf-Interested in sponsoring an episode? Learn more.
  • GitOps at Enterprise Scale, with Elad Cohen 19.08.2026 36min
    At enterprise scale, a deployment pipeline that runs Helm upgrades directly against Kubernetes hides drift, mixes configuration with CI logic, and makes the last pipeline run the source of truth.Elad Cohen explains how WSC Sports moved from Azure DevOps to GitHub Actions and redesigned delivery around Git and Argo CD. The resulting platform separates builds from deployments, keeps service configuration in values files, and continuously reconciles clusters.In this interview:Why CI should change Git instead of the clusterHow ApplicationSets create main and shadow deployments from one values fileHow AppProjects scope permissions and route alerts by teamWhy reusable Helm contracts make customization compound across servicesSponsorThis episode is sponsored by LearnKube. Download the free book, The Technical Guide to Kubernetes Rightsizing, to understand what Prometheus and Grafana cannot tell you about safely reducing requests and limits.More infoFind all the links and info for this episode here: https://ku.bz/wX5H5MjwvInterested in sponsoring an episode? Learn more.
  • Automating Pod Disruption Budgets with Kyverno, with Ahmad Asmar 10.08.2026 45min
    Karpenter can reduce Kubernetes infrastructure costs, but aggressive node consolidation can also expose workloads that lack disruption safeguards.Ahmad Asmar explains how Zencity uses Kyverno to automatically generate Pod Disruption Budgets, while accounting for existing PDBs, percentage-based availability targets, single-replica workloads, and environment-specific policies.In this interview:How Karpenter consolidation changes the availability risks of cluster operationsWhy Kyverno's generated policies can provide safer defaults than manual enforcementHow to handle duplicate PDBs, scaling workloads, and single-replica edge casesHow aggregated ClusterRoles keep custom permissions separate from Helm-managed resourcesSponsorThis episode is sponsored by LearnKube. Download the free book, The Technical Guide to Kubernetes Rightsizing, to understand what Prometheus and Grafana cannot tell you about safely reducing requests and limits.More infoFind all the links and info for this episode here: https://ku.bz/xrlPJg54DInterested in sponsoring an episode? Learn more.
  • From KIAM to EKS Pod Identities, with Fabián Sellés Rosa 04.08.2026 26min
    An unmaintained identity component can remain invisible until a routine Kubernetes upgrade turns it into an incident.Fabián Sellés Rosa, Platform Engineer and Runtime Tech Lead at Adevinta, explains how his team moved from KIAM to EKS Pod Identities without discarding the security boundaries and application interface that their internal platform depended on.In this interview:Why KIAM became urgent to replace after years of stable operationHow Crossplane, a custom controller, and KRO with ACK compared against the team's criteriaWhy managed EKS Capabilities reduced toil but introduced observability and rollout trade-offsHow Kyverno preserved namespace-level authorization for IAM rolesSponsorThis episode is sponsored by LearnKube. Download the free book, The Technical Guide to Kubernetes Rightsizing, to understand what Prometheus and Grafana cannot tell you about safely reducing requests and limits.More infoFind all the links and info for this episode here: https://ku.bz/R_06hwnCnInterested in sponsoring an episode? Learn more.
  • 1 Million Tokens Per Second on Kubernetes, with Federico Iezzi 28.07.2026 47min
    GPU inference throughput depends on more than accelerator generation or count.Memory bandwidth, model parallelism, cache configuration, and the load generator itself all influence measured throughput.Federico Iezzi, Customer Engineer at Google Cloud, explains how his team achieved 1 million output tokens per second using Qwen 3.5 27B, vLLM, GKE Autopilot, and NVIDIA B200 GPUs.The discussion covers:Why memory bandwidth limits decode performanceHow Federico chose between tensor and data parallelismWhat changed after enabling multi-token prediction and reducing the KV cache footprint with FP8 quantization.SponsorThis episode is sponsored by LearnKube. Download the free book, The Technical Guide to Kubernetes Rightsizing, to understand what Prometheus and Grafana cannot tell you about safely reducing requests and limits.More infoFind all the links and info for this episode here: https://ku.bz/1xD9Md0mbInterested in sponsoring an episode? Learn more.
  • The Hidden Cost of Slow Autoscaling, with John Ford 19.05.2026 21min
    Forced platform migrations are usually treated as something to survive. At Scout24, a mandatory OS migration became an opportunity to rethink Kubernetes autoscaling, node provisioning, and infrastructure efficiency.John Ford explains how Scout24 moved its EKS-based Infinity platform from a polling autoscaler and over-provisioned capacity to Karpenter and Bottlerocket. The result was faster node startup, a safer migration path, and about a 30% infrastructure reduction without major downtime.In this interview:Why two-minute node provisioning forced a 25% capacity bufferHow Karpenter made the Bottlerocket migration saferWhat broke around EC2 metadata, AWS SDKs, and cgroupsHow the new foundation enables Spot, ARM, and GPU workloadsSponsorThis episode is sponsored by LearnKube — get started on your Kubernetes journey through comprehensive online, in-person or remote training.More infoFind all the links and info for this episode here: https://ku.bz/DdmVC2_7vInterested in sponsoring an episode? Learn more.
  • The Namespaces Scaling Trap, with Brian Stack 12.05.2026 36min
    Most teams scale Kubernetes by thinking about pods and nodes. At Render, Brian Stack ran into a different dimension: hundreds of thousands of namespaces per cluster, multiplied across DaemonSets that list-watch every namespace.Brian explains how Render traced the issue through Calico and Vector, worked with upstream maintainers, and turned memory profiling into operational wins: lower node costs, lighter API-server load, and faster rollouts.In this interview:Why namespaces can become a hidden scaling bottleneckHow DaemonSets multiply memory and control-plane pressureHow profiling, staging clusters, and upstream collaboration freed 7 TiBWhy pushing from an 80% fix to a complete fix can make teams fasterSponsorThis episode is sponsored by LearnKube — get started on your Kubernetes journey through comprehensive online, in-person or remote training.More infoFind all the links and info for this episode here: https://ku.bz/0mrvCsXrVInterested in sponsoring an episode? Learn more.
  • AI Agents Running Kubernetes, with Mike Solomon 05.05.2026 38min
    What happens when an AI agent stops generating Kubernetes YAML and starts operating the cluster directly?Mike Solomon, software engineer at AIATELLA, explains how his team moved from a sprawling Helm setup to Markdown-driven infrastructure specs that Claude Code can execute, test, and refine.You will learnWhy Helm became hard to maintain for a fast-moving medical infrastructure repoHow Claude debugged Argo, TLS conflicts, kubectl patches, and private registry credentialsHow runbooks plus agent memory files capture failures so deployments become reproducible.It is a practical look at where Kubernetes automation may be heading: less hand-written YAML, more precise intent, and a sharper definition of when the human must stay in the loop.SponsorThis episode is sponsored by LearnKube — get started on your Kubernetes journey through comprehensive online, in-person or remote training.More infoFind all the links and info for this episode here: https://ku.bz/y70mLvWNsInterested in sponsoring an episode? Learn more.
  • SaaS with Kubernetes Operators and Garbage Collection, with Alexander Held 28.04.2026 35min
    A single Kubernetes CRD for every service request turns small changes into full-platform reconciliations.Alexander Held, former platform engineer at Mercedes-Benz Tech Innovation, describes a production refactor from a 2,000-line CRD to purpose-built resources and controllers. He shows how teams can model business workflows as Kubernetes APIs and then use owner references, finalizers, and events to keep platform operations predictable.You will learn:Why monolithic CRDs create performance and troubleshooting problemsHow controllers turn database provisioning and backups into reconciliation loopsHow finalizers clean up external resources such as S3 backupsWhy Kubernetes events make platform workflows easier to debugSponsorThis episode is sponsored by LearnKube — get started on your Kubernetes journey through comprehensive online, in-person or remote training.More infoFind all the links and info for this episode here: https://ku.bz/TGy4Qn7QsInterested in sponsoring an episode? Learn more.
  • What Hip-Hop Can Teach Us About Kubernetes, with Kelsey Hightower, Eric Abercrombie, and Julius Payne II 21.04.2026 1t 29min
    Kelsey Hightower, Eric Abercrombie, and Julius Payne II reflect on life after achievement, entering the Kubernetes world for the first time, and how music, creativity, and lived experience shape the way they think about technology.In this interview:Why fundamentals, patience, and repetition still matter more than shortcutsHow Kubernetes, community, and confidence intersect for people entering cloud-native workWhat hip-hop, production, and storytelling can teach us about ownership, authenticity, and finding your voiceSponsorThis episode is sponsored by LearnKube — get started on your Kubernetes journey through comprehensive online, in-person or remote training.More infoFind all the links and info for this episode here: https://ku.bz/czrCCXSLtInterested in sponsoring an episode? Learn more.
  • Intelligent Kubernetes Load Balancing, with Rohit Agrawal 07.04.2026 30min
    You're running gRPC services in Kubernetes, load balancing looks fine on the dashboard — but some pods are burning at 80% CPU while others sit idle, and adding more replicas only partially helps.Rohit Agrawal, a Staff Software Engineer on the traffic platform team at Databricks, explains why this happens and how his team replaced Kubernetes's default networking with a proxy-less, client-side load-balancing system built on the xDS protocol.In this episode:Why KubeProxy's Layer 4 routing breaks down under high-throughput gRPC: it picks a backend once per TCP connection, not per requestHow Databricks built an Endpoint Discovery Service (EDS) that watches Kubernetes directly and streams real-time pod metadata to every clientHow zone-aware spillover cut cross-availability-zone costs without sacrificing availabilityWhy CPU-based routing failed (monitoring lag creates oscillation) and what signals to use insteadThe system has been running in production for three years across hundreds of services, handling millions of requests.SponsorThis episode is sponsored by LearnKube — get started on your Kubernetes journey through comprehensive online, in-person or remote training.More infoFind all the links and info for this episode here: https://ku.bz/y803JMhBkInterested in sponsoring an episode? Learn more.
  • That Time I Found a Service Account Token in my Log Files, with Vincent von Büren 31.03.2026 28min
    You're integrating HashiCorp Vault into your Kubernetes cluster and adding a temporary debug log line to check whether the ServiceAccount token is being passed correctly. Three months later, that log line is still in production — and the token it prints has a 1-year expiry with no audience restrictions.Vincent von Büren, a platform engineer at ipt in Switzerland, lived through exactly this incident. In this episode, he breaks down why default Kubernetes ServiceAccount tokens are a quiet security risk hiding in plain sight.You will learn:What's actually inside a Kubernetes ServiceAccount JWT (issuer, subject, audience, and expiry)Why tokens with no audience scoping enable replay attacks across internal and external systemsHow Vault's Kubernetes auth method and JWT auth method compare, and when to choose eachWhat projected tokens are, why they dramatically reduce blast radius, and what's holding teams back from using themPractical steps for auditing which pods actually need API access and disabling auto-mounting everywhere elseSponsorThis episode is sponsored by LearnKube — get started on your Kubernetes journey through comprehensive online, in-person or remote training.More infoFind all the links and info for this episode here: https://ku.bz/LTnB_NtbcInterested in sponsoring an episode? Learn more.
  • GPU Containers as a Service, with Landon Clipp 24.03.2026
    Running GPU workloads on Kubernetes sounds straightforward until you need to isolate multiple tenants on the same server. The moment you virtualize GPUs for security, you lose access to NVIDIA kernel drivers — and almost every tool in the ecosystem assumes those drivers exist.Landon Clipp built a GPU-based Containers as a Service platform from scratch, solving each isolation layer — from kernel separation with Kata Containers + QEMU to NVLink fabric partitioning to network policies with Cilium/eBPF — and shares exactly what broke along the way.In this interview:Why standard NVIDIA tooling (GPU Operator) fails in multi-tenant setups, and how to use CDI with PCI topology scanning to make GPUs visible to Kubernetes without kernel driversHow to partition the NVLink fabric between tenants using a trusted service VM running Fabric Manager, and why the physical PCIe wiring differs between Supermicro HGX and NVIDIA DGX systemsWhy gVisor doesn't work for GPU workloads — NVIDIA's unstable ioctl ABI means Google has to update gVisor for every driver release, and they only support a handful of GPUsWhat caused 8-GPU VMs to take 30+ minutes to boot, and the specific fixes (IOMMUFD, cold plugging, kernel upgrades) that brought it down to minutesHow Cilium network policies enforce tenant isolation at the Kubernetes identity level instead of fragile IP-based rulesWhere Containers as a Service fits best: inference workloads where AI teams want to ship an OCI image without managing infrastructure or signing multi-million dollar cluster contracts.SponsorThis episode is sponsored by LearnKube — get started on your Kubernetes journey through comprehensive online, in-person or remote training.More infoFind all the links and info for this episode here: https://ku.bz/jjK_yJTDzInterested in sponsoring an episode? Learn more.
  • How We Cut Build Debugging Time by 75% with AI, with Ron Matsliah 17.03.2026 20min
    Build failures in Kubernetes CI/CD pipelines are a silent productivity killer. Developers spend 45+ minutes scrolling through cryptic logs, often just hitting rerun and hoping for the best.Ron Matsliah, DevOps engineer at Next Insurance, built an AI-powered assistant that cut build debugging time by 75% — not as a dashboard, but delivered directly in Slack where developers already work.In this episode:Why combining deterministic rules with AI produces better results than letting an LLM guess aloneHow correlating Kubernetes events with build logs catches spot instance terminations that produce misleading errorsWhy integrating into existing workflows and building feedback loops from day one drove adoptionThe prompt engineering lessons learned from testing with real production data instead of synthetic examplesThe takeaway: simple rules plus rich context consistently outperform complex AI queries on their own.SponsorThis episode is sponsored by LearnKube — get started on your Kubernetes journey through comprehensive online, in-person or remote training.More infoFind all the links and info for this episode here: https://ku.bz/PDdYfC00wInterested in sponsoring an episode? Learn more.
  • Migrating Kubernetes Off Big Cloud, with Fernando Duran 10.03.2026 25min
    Managed Kubernetes on a major cloud provider can cost hundreds or even thousands of dollars a month — and much of that spending hides behind defaults, minimum resource ratios, and auxiliary services you didn't ask for.Fernando Duran, founder of SadServers, shares how his GKE Autopilot proof of concept ran close to $1,000/month on a fraction of the CPU of the actual workload and how he cut that to roughly $30/month by moving to Hetzner with Edka as a managed control plane.In this interview:Why Kubernetes hasn't delivered on its original promise of cost savings through bin packing — and what it actually provides insteadA real cost comparison: $1,000/month on GKE vs. $30/month on Hetzner with Edka for the same nominal capacityWhat you need to bring with you (observability, logging, dashboards) when leaving a fully managed cloud providerThe decision comes down to how tightly coupled you are to cloud-specific services and whether your team can spare the cycles to manage the gaps.SponsorThis episode is sponsored by LearnKube — get started on your Kubernetes journey through comprehensive online, in-person or remote training.More infoFind all the links and info for this episode here: https://ku.bz/6nSDbz9m4Interested in sponsoring an episode? Learn more.
  • Migrating to Karpenter: Fun Stories, with Adhi Sutandi 03.03.2026 1t 1min
    Running multiple Kubernetes clusters on AWS with the cluster autoscaler? Every four months, you face the same grind: upgrading Kubernetes versions, recreating auto scaling groups, and hoping instance type changes stick.Adhi Sutandi, DevOps Engineer at Beekeeper by LumApps, shares how his team migrated from the cluster autoscaler to Karpenter across eight EKS clusters — and the hard lessons they learned along the way.In this episode:Why AWS auto scaling groups are immutable and how that creates upgrade bottlenecks at scaleHow the latest AMI tag accidentally turned less critical clusters into chaos engineering environments, dropping SLOs before anyone realized Karpenter was the causeWhy pre-stop sleep hooks solved pod restartability problems that Quarkus's built-in graceful shutdown couldn'tThe case for pod disruption budgets over Karpenter annotations when protecting critical workloads during node rotationsHow Karpenter's implicit 10% disruption budget caught the team off guard — and the explicit configuration that fixed itSponsorThis episode is sponsored by LearnKube — get started on your Kubernetes journey through comprehensive online, in-person or remote training.More infoFind all the links and info for this episode here: https://ku.bz/XyVfsSQPrInterested in sponsoring an episode? Learn more.
  • From ECS to Kubernetes: A Real Migration Story, with Radosław Miernik 24.02.2026 38min
    Migrating from ECS to Kubernetes sounds straightforward — until you hit spot capacity failures, firewall rules silently dropping traffic, and memory metrics that lie to your autoscaler.Radosław Miernik, Head of Engineering at aleno, walks through a real production migration: what broke, what they missed, and the fixes that made it work.In this interview:Running Flux and Argo CD together — Flux for the infra team, Argo CD's UI for developers who don't want to touch YAMLHow the wrong memory metric caused OOM errors, and why switching to jemalloc cut memory usage by 20%Splitting WebSocket and API containers into separate deployments with independent autoscalingFour months of migration, over 100 configuration changes in the first month, and a concrete breakdown of what platform work looks like when you can't afford downtime.SponsorThis episode is sponsored by LearnKube — get started on your Kubernetes journey through comprehensive online, in-person or remote training.More infoFind all the links and info for this episode here: https://ku.bz/x6wFMhVsxInterested in sponsoring an episode? Learn more.

Suosittu maassa

Tämä podcast esiintyy myös näiden maiden podcast-listoilla.