November 9-12, 2026 | Salt Lake City, Utah
Sessions will be recorded and available on the CNCF YouTube channel within two weeks.
Times shown in MST (UTC-7). Seating is first come, first served.
Plan your sessions and build your personal agenda.
Learn how to use the event app and sync favorites across devices.
Please note that this is an off-site Sponsor-hosted Co-located event.
Location: Hilton Salt Lake City Center | 255 South West Temple, Salt Lake City, Utah 84101
Join us in Salt Lake City for a no-cost, in-person OpenShift Commons Gathering alongside KubeCon + CloudNativeCon North America. This event brings together the global OpenShift community-including users, contributors, partners, and Red Hat experts-to share knowledge, build connections, and learn from real production experiences.
Hear practitioner-driven talks on app modernization, developer productivity, AI/LLM workloads, supply chain security, VM migration, GitOps, edge, and scaling OpenShift across hybrid and multi-cloud environments. Expect practical insights you can bring back to your team.
For questions regarding this event, please contact: [email protected]
Agentics Day: MCP + Agents event is a community-driven event dedicated to advancing the Model Context Protocol, an emerging standard for connecting AI models with external tools, data, and workflows. Join developers, contributors, and practitioners to explore real-world implementations, best practices, and the future of building richer, safer, and more capable AI applications with MCP.
To learn more please visit the event's website above.
For questions regarding this event, please reach out to [email protected].
ArgoCon is designed to foster collaboration, discussion, and knowledge sharing on the Argo Project, which consists of four projects: Argo CD, Argo Workflows, Argo Rollouts and Argo Events. To learn more please visit the event's website above.
For questions regarding this event, please reach out to [email protected].
BackstageCon is a one-day conference focused on all things Backstage: an open framework for building developer portals. At BackstageCon, we’ll provide a vendor-neutral space for collaboration and learning centered on improving developer experience and effectiveness through open source technologies.
The event is vendor-neutral and organized by members of the Backstage community. To learn more please visit the event's website above.
For questions regarding this event, please reach out to [email protected].
Cilium is an open source, widely-used, and highly scalable cloud native networking, observability, and security solution based on the kernel technology eBPF, that connects workloads in Kubernetes and beyond. CiliumCon focuses on how Cilium and its sub-projects Hubble and Tetragon are being developed, deployed, and used across the cloud native landscape to revolutionize cloud native platforms.
At CiliumCon you’ll hear from end users sharing how Cilium, Hubble, and Tetragon unlocked levels of scalability, performance, and security that weren’t possible before and from contributors who will teach you how Cilium is leveraging eBPF to deliver these benefits. From Cilium and eBPF internals to how Cilium, Hubble, and Tetragon are helping businesses achieve their goals, you’ll hear it all at CiliumCon. Dive deep into the world of high-performance networking, transparent security, and scalable observability at CiliumCon!
To learn more please visit the event's website above.
For questions regarding this event, please reach out to [email protected].
As we step into the era of rapid AI advancements, organizations are grappling with an unprecedented array of challenges. The rise of Large Language Models (LLMs), the development of Graph RAGs (retrieval-augmented generation) and agentic systems, and the growing importance of Ethical Considerations in AI are reshaping how businesses innovate, scale, and move from development to production.
To learn more please visit the event's website above.
For questions regarding this event, please reach out to [email protected].
FluxCon is the official community gathering for Flux users, contributors, and adopters, such as end-user organizations and service providers. As GitOps and continuous delivery continue to evolve as a core operating model for cloud native infrastructure, FluxCon provides a dedicated space to share best practices, real-world success stories, and deep technical knowledge around continuous delivery with Flux.
To learn more please visit the event's website above.
For questions regarding this event, please reach out to [email protected].
Observability Day is a vendor-neutral gathering of the CNCF observability community. Maintainers, operators, and end users of Prometheus, Fluentd and Fluent Bit, OpenTelemetry, Jaeger, Thanos, Cortex, Perses, Pixie, Kepler, Inspektor Gadget, and the wider ecosystem come together as peers to share how cloud-native systems are observed, debugged, and operated at scale. Talks span project internals, cross-project architectures, the data engineering required to make telemetry useful and affordable, and emerging frontiers such as AI and agent observability, eBPF-based instrumentation, and CI/CD and platform telemetry. The day favors experience-driven content: real architectures, real trade-offs, and the open source tools that make modern observability possible.
To learn more please visit the event's website above.
For questions regarding this event, please reach out to [email protected].
A new event that fosters collaboration and shares innovation in cloud native security and open source software security. Sessions will cover architecture and policy, secure software development, supply chain security, identity and access, and open source public policy. The 1-day event will gather a diverse community of professionals to include software developers, security engineers, public sector experts, CISOs, CIOs and tech pioneers to address challenges and opportunities in modern security. Join us to collaborate and discover tools, knowledge and strategies to ensure a safer, more secure tomorrow.
To learn more please visit the event's website above.
For questions regarding this event, please reach out to [email protected].
Platform Engineering Day is dedicated to exploring the real-world challenges and advanced practices behind building and scaling Internal Developer Platforms (IDPs). IDPs provide curated capabilities, frameworks, and developer-centric experiences to accelerate internal teams such as application developers. The process and techniques described in the CNCF Platforms White Paper and Platform Engineering Maturity Model highlight how achieving a high-impact Developer Experience (DevEx) requires more than tools; it demands a holistic, socio-technical investment. We invite proposals that include deep dives into platform engineering case studies, practical strategies for measuring and maturing platform capabilities, and emerging patterns for integrating AI into platform workflows. Join Platform Engineers, Product Managers, Solutions Architects, and cloud-native stakeholders as we share lessons learned in building and managing internal platforms.
To learn more please visit the event's website above.
For questions regarding this event, please reach out to [email protected].
What happens when your Kubernetes cluster is pushed to its limits—does it scale or break in unexpected ways?
In this lightning talk, I’ll show how kube-burner is used to stress-test clusters, uncover performance bottlenecks, and catch regressions early.
Through quick examples, you’ll see how to run benchmarks, capture meaningful metrics, and identify issues like resource startup latencies and control plane slowdowns.
We’ll also highlight its role in upstream efforts like Kubernetes SIG KubeVirt and SIG Scalability, along with the project’s current scope and upcoming areas of development.
If you care about scalability and performance, this session gives you a fast, practical starting point.
Kubernetes gives platform teams powerful building blocks. But turning those into a real developer experience takes months of work. OpenChoreo is a complete, open-source developer platform for Kubernetes. It's ready to use from day one, for both humans and agents.
In this lightning talk, I'll walk through how OpenChoreo provides development and platform abstractions on top of Kubernetes. It comes with a Backstage-powered developer portal, plus built-in CI/CD, GitOps, and observability. Developers can self-serve deployments without needing to be Kubernetes experts.
These same abstractions also work well for AI agents. Agents need to build, deploy, and operate workloads too, often alongside humans. I'll share what it means to design a platform that serves both audiences from the start.
Since becoming the first cloud-native edge project to graduate from the CNCF, KubeEdge adoption has exploded. From intelligent transportation and smart cities to energy and banking, it is now the backbone of critical edge-cloud ecosystems. In this lightning talk, we will highlight the latest post-graduation features, governance updates, and real-world success stories that demonstrate the power of KubeEdge for Edge AI and industrial workloads.
Learn how the Vitess project and the Vitess Operator offer a k8s native solution for MySQL that offers automated high availability, self-healing & failover, integrated backups and restore, and zero-downtime updates/upgrades.
Flux turns 10!
This talk gives an overview of GitOps and Flux (the CNCF graduated project) and why Flux has been the most trusted tool for cloud providers, small to enterprise customers, and across financial institutions to telcos.
Flux continues to push boundaries in security, scalability, and reliability:
– Security-first design including SOPS with Age post-quantum cipher, Workload Identity, verification, CoSign, and the Flux UI with RBAC, SSO, and more
– Massive scale through design, data size, and network efficiency: Kubernetes server-side apply, horizontal and vertical pod autoscaling, artifact size reduction, and multiple artifact source and data types options
– Flux’s modularity: use only needed Flux controllers, no CPU bottlenecks from app bloat, and now a Flux CLI plugin system for added capabilities
– Flux’s use of the Helm SDK for full capabilities including supporting Helm’s post-render strategies, literal mode, and more
Celebrate Flux’s 10th anniversary with us at KubeCon!
Cloud Native Buildpacks transform your application source code into images that can run on any cloud. They enable advanced caching mechanisms that improve performance at scale. They also allow for modularity and reuse, which ensure developers across your organization aren’t wasting cycles repeating what other teams have already done.
After this short talk, you’ll be able to run buildpacks with the Pack CLI and find off-the-shelf buildpacks in the Buildpack Registry, including those from Google, Heroku, and Paketo. Finally, you’ll learn how operators of large platforms use buildpacks to make their container builds as scalable as possible.
As generative AI reshapes how organizations build and deploy intelligent applications, the demands on model serving infrastructure keep rising, teams need distributed execution, efficient autoscaling, and multi-tenant isolation, all without adding operational complexity. This session explores how KServe, a CNCF incubating project, continues evolving to meet these challenges head-on.
We're excited to walk through latest releases of KServe. These releases introduce multi-node inference without requiring a Ray head node, LeaderWorkerSet-based autoscaling for distributed workloads, and an upgrade to llm-d v0.6 for disaggregated prefill/decode serving. It also adds OpenAI Responses API routing, namespace-scoped ModelCache for multi-tenant model caching, and all the latest vLLM features, alongside meaningful security hardening across the platform.
Join us to see how KServe is simplifying production-grade LLM serving, one release at a time.
Immutable operating systems make Kubernetes nodes much more predictable, but they also introduce a problem for new users: a read-only root filesystem. When you need to install debugging tools, eBPF tracers, or custom storage drivers, the usual habit of running a package manager just does not work.
This 5-minute talk shows a practical way to solve this using system extensions (systemd-sysext) in Flatcar Container Linux. I will explain how system extensions let platform teams dynamically overlay binaries onto a read-only OS at runtime. This keeps the atomic update model intact and respects the rules of immutability. You will walk away knowing how to keep your host OS strictly versioned while still having the flexibility to handle day-to-day debugging and hardware enablement.
Kepler is a CNCF Sandbox project that monitors power usage and attributes it to Pods and other resources in Kubernetes clusters. Adoption is growing: it is used by leading Linux Foundation Project Sylva's telco carriers to optimize CNFs on bare metal, and by CERN to optimize the environmental impact of their MLOps against their sustainability KPIs. Yet while Kepler took a step forward in accuracy during the rewrite, we had to take a big step backward for the virtual machine use case, since we lost the trained model server (hint: we're solving this). What challenges and opportunities lie ahead for the project?
OSCAL Compass Sandbox transforms compliance from a documentation nightmare into an interactive dialogue where your security posture evolves alongside your codebase. On top of automating compliance, we are making it conversational, collaborative, and… actually enjoyable. Come see how agentic workflows and our MCP servers are turning compliance-as-code into compliance-as-skill.
In this talk you will see how compliance integrates directly into the dev workflow via OSCAL Compass, rather than acting as a late-stage gate. Perfect for platform engineers and AI/ML ops shipping regulated AI.
Something is misbehaving inside a pod: an unexpected outbound connection, a file read no one can explain, a process that dies without a trace. Your logs stop at the application boundary, but the answers live in the kernel. In this lightning talk, we will show how Inspektor Gadget uses eBPF to surface kernel-level activity across a running cluster in seconds, without redeploying workloads, installing sidecars, or writing any C. We will live-demo tracing DNS, file, and network events with kubectl gadget, each one automatically enriched with Kubernetes context: pods, containers, and namespaces. We will then show how every Gadget is packaged as an OCI image, so eBPF programs can be built, signed, pushed to any registry, and run anywhere, just like containers. You will leave knowing how to debug your own clusters this afternoon, and how to write and share your first Gadget with the community.
Originally designed as a lightweight log processor, Fluent Bit has evolved into a high-performance, multi-signal telemetry engine powering billions of deployments across Kubernetes, cloud, edge, and AI workloads.
This project update highlights the latest evolution of the Fluent ecosystem, including the introduction of the new Fluent organization, recent advances in OpenTelemetry support, telemetry routing, eBPF observability, performance optimizations, and reliability improvements. We’ll also introduce Fluent Bit 5.2, released during KubeCon North America, and share the roadmap and engineering priorities shaping the next generation of the project.
Whether you’re already running Fluent Bit or discovering it for the first time, you’ll leave with an overview of where the project stands today, what’s new, and what’s coming next.
Chaos engineering has evolved from simple pod-deletes to complex, automated resilience orchestration. Over the last year, the LitmusChaos project has focused on reducing the toil of experiment management while expanding the scope of what can be tested.
In this session, the speaker will cover the major 2025-2026 updates, starting with the new Model Context Protocol (MCP) server. This integration allows teams to move beyond YAML-heavy workflows by using LLM-based assistants to discover, trigger, and analyze chaos experiments through natural language. The session will also cover technical refinements to the core engine, including the ability to target Kubernetes Jobs, execute multi-container stress tests, and utilize a revamped GitOps reconciliation engine for more reliable experiment state management. The 2026 roadmap will also be shared.
Open source only works if the community can contribute without hitting a wall. We realized our architecture at OpenEverest was unintentionally biased toward specific tools, so we rebuilt it. I’ll share how we moved to a modular plugin system that decouples our core from specific operators. By supporting everything from Valkey to competing Postgres flavors, we’ve eliminated vendor lock-in and ensured the community—not the architecture—dictates the project's future.
AI agents are powerful, but production agent workflows are fragile: tools time out, model calls fail, workers restart, human approvals take hours, and long-running tasks can lose state halfway through execution.
Cadence Workflow is a CNCF sandbox project for durable workflow orchestration. In this 5-minute project lightning talk, we will show how Cadence helps make agentic workflows more secure, reliable, and recoverable by giving them persisted state, retries, timeouts, human checkpoints, worker failure recovery, and observable execution history.
Attendees will leave with a simple mental model for when an agent should become a durable workflow, where Cadence fits in the AI-native cloud native stack, and how to get involved with the project through docs, samples, developer tooling, and AI workflow examples.
As LLMs scale, we face a stark reality: specialized AI data centers are costly, and existing facilities are constrained by power and physical footprint. Enter ""Scale-Across""—a paradigm shift for deploying distributed AI workloads across existing multi-site infrastructures.
In this talk, we introduce CoHDI (Composable Hardware in Disaggregated Infrastructure), a CNCF Sandbox project for cloud-native hardware composability. As an active initiative within the IOWN Global Forum, we share the cutting-edge status of CoHDI evaluation over the All-Photonics Network (APN). Specifically, we elaborate on how we leverage the LLM-d (PD Disaggregated) framework to dynamically separate Prefill compute-bound and Decode memory-bound workloads across edge and core sites. Discover how CoHDI orchestrates pooled resources like GPUs and CXL memory natively in Kubernetes to unlock true geo-distributed LLM inference efficiency.
k8gb, the Kubernetes Global Balancer, has grown from a CNCF Sandbox project into an Incubating project with a clear mission: make global traffic resilience Kubernetes-native, open, and vendor-neutral.
In this lightning talk, we’ll give a fast update for new and returning users. We’ll recap how k8gb enables global service load balancing across clusters and regions using declarative Kubernetes resources, health-aware DNS responses, and no central management cluster or single point of failure.
Then we’ll share what changed on the road to incubation: expanded Gateway API support, the vendor-neutral k8gb.io/v1beta1 API group and migration path, and the ZoneDelegation CRD for safer multi-cluster DNS zone management. We’ll also touch on release maturity improvements, including OCI Helm registry migration and Cosign signing.
Come learn where k8gb is today, what problems it solves for platform teams, and how to get involved.
You have a controller that solves a real problem on one cluster. Your job is to get it running on all of them – consistently, with config per cluster, without touching each cluster individually. Think it's not possible? Think again.
The Open Cluster Management (OCM) add-on framework makes this declarative: three objects, one placement binding, and the hub does the rest. In this session we build a real add-on live, using an agentic skill to scaffold and explain each piece as it goes. By the end, a controller that ran on one cluster is running on all of them.
You will leave with a working mental model of ClusterManagementAddOn, AddOnDeploymentConfig, and AddOnTemplate and an actual agentic skill you can use with OCM.
Karmada (Kubernetes Armada) is a Kubernetes management system that enables you to run your cloud-native applications across multiple Kubernetes clusters and clouds.
In this presentation, the maintainer of the Karmada project will share:
- A Brief Introduction to Karmada.
- Typical use cases
- New features over the last year
- Real-world case studies
- Overview of the community
- Roadmap
- QA
Join the K3s maintainers for a rapid update on the lightweight Kubernetes distribution designed for IoT, edge, and resource-constrained environments. We’ll highlight recent releases, key improvements in performance, security, and usability, and how the project continues to simplify Kubernetes at the edge.
We’ll also give a sneak peek at what’s coming next, including upcoming features, roadmap priorities, and opportunities for the community to contribute. Get a fast, clear snapshot of where K3s is today and where it’s headed tomorrow.
Building a Crossplane control plane is one thing. Getting real teams at a risk-averse enterprise to start using it is another.
In this lightning talk, the platform lead at a large insurance company will walk through their successful Crossplane end-user adoption story, focusing on what they learned and what they had to change once they left the prototype stage and actual developer teams started using the platform for real.
We'll learn about the shift from Jira-driven requests to self-service infrastructure that truly empowers app teams, how to design composition APIs your devs can actually use, and some hard lessons on abstractions, defaults, ownership, and support. Anyone working to drive real adoption of an internal platform will leave with ideas to put to work right away.
If a large, risk-averse insurance company can make Crossplane a platform its teams rely on, yours can too. Come hear the field notes from an adopter who has taken Crossplane to real enterprise scale.
Location: Gallivan Center | 50 E 200 S, Salt Lake City, UT 84111
Please note that this is an off-site Sponsor-hosted Co-located event.
Spend the day with engineers, architects, and cloud-native leaders exploring the latest in distributed SQL for modern applications and AI. Expect expert-led technical sessions, live demos, hands-on learning, real-world customer insights, networking with the community, great food, and plenty of opportunities to connect with the people building the future of PostgreSQL-compatible databases.
To learn more, please visit: https://events.ringcentral.com/events/distributed-sql-summit-salt-lake-city-2026/registration
For questions regarding this event, please contact: [email protected]
Please note that this is an off-site Sponsor-hosted Co-located event.
Location: Salt Lake Marriott Downtown at City Creek | 75 S W Temple St, Salt Lake City, UT 84101
Virtualization is moving to Kubernetes, and the question is no longer whether VMs belong there, but how to run them well. VM on Kubernetes Day North America 2026 explores how KubeVirt brings virtual machines onto the same platform as your cloud-native and AI workloads, so legacy apps come along without being left behind and without platform sprawl. Join platform engineers, VMware admins, developers, and other practicioners for real-world migration stories, unified storage and Day 2 strategies, and a hands-on look at running your own VMs on Kubernetes.
For questions regarding this event, please contact: [email protected]
Learn about the latest etcd-operator development and features. We've completely refactored the operator to use custom controllers and the new version can now automagically handle your failed cluster members for you. It's the etcd operator we've always wanted. We'll share development milestones, current feature set, and user experience.
For the past decade, service meshes have existed. Starting with Linkerd, it burst into a new type of software with many competitors. But now, in 2026, do we even need service meshes anymore or are they legacy software? Should it go the way of a monolith?
In a thunder clap of a talk we'll cover who service mesh is for, and where you can find out more
Your cluster is not a pile of YAML files. It's a graph: workloads, custom resources, network policies, and RBAC all relate, and most breakage lives in those relationships, not in any single file. Meshery models Kubernetes and CNCF resources as an ontology and renders that graph visually in Kanvas.
Now a new consumer needs that graph: AI agents. Asked to design and manage Kubernetes, they operate blind, generating YAML one file at a time with no model of how anything relates. The result: configurations that are individually valid and collectively broken.
Meshery maintainer Yash Sharma will show how Meshery exposes its relationship model to agents through MCP (Model Context Protocol), so an agent reasons over a real map of your infrastructure instead of guessing, and validates its proposals against OPA policies before anything ships.
One idea, one takeaway: the knowledge graph that makes clusters visual for humans is what agents need to stop hallucinating your infrastructure.
Ten years ago, Envoy was open-sourced, solving one company’s microservices networking problems.
Today it’s the data plane underneath much of the internet — powering service meshes, API gateways, and now the infrastructure layer for AI traffic.
This lightning talk marks Envoy’s 10th anniversary with a fast tour of the project in 2026: the core proxy’s evolution, dynamic modules enabling Rust extensions without recompiling, and a thriving sub-project family including Envoy Gateway and the newly GA Envoy AI Gateway.
Then, a look ahead: agents calling models, models calling tools, and platforms needing governance over all of it.
Ten years in, Envoy’s second decade is just getting started and here’s how to be part of it.
OPA has long been the policy engine of choice for cloud native security, but it's usually ran as a sidecar. Our Go SDK changed things by letting you embed OPA directly inside your (Go) application, no sidecar required. Now, using the Rego Intermediate Representation (Rego IR), the compiler can generate runtime execution plans that other languages can consume natively. The result? Rego-powered security policy running natively in Swift and Java applications too. This lightning talk walks through how IR works, why it matters for performance, and what it means for bringing consistent security policy to more language runtimes than ever before.
An AI agent is not one thing. Its model, prompts, skill files, MCP servers, and code live in different places: Git repos, object storage, hardcoded strings, loose config files. There is no single package that pins down ""this is the agent, this exact combination."" So when something breaks, nobody can answer the basic question: what was actually running?
KitOps, a CNCF Sandbox project, began by packaging AI/ML models as OCI artifacts called ModelKits. Now it goes further. A ModelKit can bundle agent skills, prompts, MCP servers, and code into one versioned, immutable OCI artifact, and it does not need to include a model at all. Write a short Kitfile, run `kit pack`, and push to any registry you already use. Tags like `:v2.1` or `:champion` apply to the whole agent, not scattered pieces. Sign it with Cosign. Pull it anywhere and get back the exact known-good state.
In five minutes, I'll show what changed and demo a Kitfile that packages an agent with no model.
According to a forecast from the International Data Corporation (IDC) Worldwide Edge Spending Guide, combined enterprise and service provider spending across hardware, software, professional services, and provisioned services for edge solutions will sustain strong growth through 2027 when spending will reach nearly $350 billion. With hardware and software dispersed across hundreds or even thousands of locations, the simple paradigms around observability, loosely coupled systems, declarative APIs, and strong automation that have propelled the success of cloud native technologies in the cloud are the only feasible way to manage these distributed systems. Kubernetes is already a significant component of the edge ecosystem, driving integrations and operations.
Join us at Kubernetes on the Edge Day at KubeCon + CloudNativeCon and take part in defining the future intersection of cloud native and edge computing!
To learn more please visit the event's website above.
For questions regarding this event, please reach out to [email protected].
Join us for OpenTofu Day 2026, a dedicated day for the infrastructure-as-a-code community. We’ll share a day all about OpenTofu, including migrations, technical details, panels, and new use cases. Don’t miss this opportunity to learn, contribute, and join the OpenTofu community.
To learn more please visit the event's website above.
For questions regarding this event, please reach out to [email protected].
Kairos started as a framework that lets you take familiar Linux distributions like Ubuntu, Fedora, and more, and turn them into immutable, image-based operating systems for Kubernetes anywhere: edge, cloud, or on-prem. After years of pushing these distros to their limits, we realized we’d learned enough to define what really matters for building dependable image-based systems. That’s why we took the next leap and built our own distribution: Hadron, a lightweight musl + systemd Linux now in the CNCF.
This update will introduce Hadron publicly and highlight everything new in Kairos since the last project session.
This session explores how Kyverno supercharges Common Expression Language (CEL) to deliver advanced resource lifecycle management for Kubernetes and beyond.
While native Kubernetes handles basic validation checks, it stops short of what enterprise platforms require. Attendees will see how Kyverno bridges these gaps—empowering platform teams with complex validations, mutative rules, dynamic resource generation, and deep context-awareness (such as image metadata lookups and external API calls). Finally, the session showcases how Kyverno is expanding its footprint beyond Kubernetes clusters, allowing organizations to write unified, high-performance policies that secure the entire cloud-native ecosystem.
Please note that this is an off-site Sponsor-hosted Co-located event.
Location: Le Méridien Salt Lake City Downtown, Address: 131 S 300 W, Salt Lake City, UT 84101| Room: Triumph Salon I
Governing autonomous AI agents is no longer a design problem — it’s an operational one. Agents are running in production clusters today, calling tools over MCP, invoking each other, and reaching external model providers.
This session is built entirely on what we’ve seen in real environments – both in Kubernetes and beyond. Tigera CTO Peter Kelly walk the path we’ve watched every team take, in order. Discovery first, because agents arrive through application teams rather than platform teams. Through to registration with identity, because a Kubernetes service account tells you which pod made a call but nothing about which agent, acting for whom, with what authority. Then governance – expressing intent as policy at the level of tools and data rather than IPs and ports, and running it in observation mode long enough to trust it before anything gets enforced.
It’s architecture-first and deliberately vendor-neutral: the goal is a mental model you can apply to whatever you’re building on.
For questions regarding this event, please contact: [email protected]
In this lightning talk, Harbor maintainers share a fast-paced update on the project’s latest progress. We’ll highlight what shipped in Harbor 2.15 and 2.16, including key improvements in security, operations, and maintainability driven by real-world usage.
We’ll also preview what’s coming next on the Harbor roadmap, from upcoming features to active design discussions, and how the community can get involved. Join us for a quick snapshot of where Harbor is today and where it’s headed next.
Kubernetes clusters depend heavily on the nodes they run on. As more teams adopt immutable, minimal operating systems to cut down on security risks and configuration drift, Flatcar Container Linux has become a standard open source option for these workloads.
In this 5-minute talk, I will share the latest updates from the Flatcar project. I will cover our progress within the CNCF ecosystem and highlight recent work on security automation, release governance, and supply chain integrity (like signed SBOMs and SLSA provenance). You will leave with a good idea of our current roadmap, what the community is working on right now, and where new contributors can jump in.
Authorization decisions often depend on information that isn't easily available when you need to make a decision. In Kubernetes, for example, the API server must authorize a request before the body is decoded. Without access to all the fields, authorizers have to first over-grant permissions before using separate admission plugins to enforce fine-grained restrictions after the fact.
Cedar's partial evaluation addresses this by evaluating policies against incomplete information. When it lacks what it needs to return ""allow"" or ""deny,"" it instead produces a residual policy encoding the information needed to reach a concrete decision. Recent extensions to Cedar's partial evaluation support more fine-grained partial inputs, making this practical for real-world use. In this talk, we'll walk through Cedar policy evaluation when key attributes are unknown, producing a residual that is refined incrementally as information becomes available.
This talk is a short format summary of the progress achieved by the Metal3 project and its community, particularly in last couple of release cycles. We will do a quick walkthrough of the latest and greatest features of the project and an overview of the road-map of the project.
Atlantis is the self-hosted Go binary that brings Terraform and OpenTofu into your pull requests: plan when the PR opens, apply once it is approved, with the whole team watching. Teams have leaned on it for years to get infrastructure changes out of laptops and bespoke CI, and it is now a CNCF Sandbox project with an active contributor base and bi-weekly
community meetings.
This talk is a five-minute tour of what has landed recently: first-class OpenTofu support, alpha drift detection and remediation APIs, high-availability deployments backed by Redis, and hardened security across comment parsing and locking. We will close with where the project is headed and how to get involved, whether you run Atlantis today or have never touched Terraform.
Longhorn equips clusters with lightweight, cloud native capabilities, including incremental snapshots, disaster recovery, and RWX support—free from vendor dependencies. In this session, Divya Mohan explores this year's project updates and provides structured insights into the roadmap. The session also aims to outline actionable paths for new contributors by highlighting contribution opportunities within the project.
oras-go powers ORAS, the ORAS CLI, Helm, Flux, and many tools that push and pull OCI artifacts. v2 made it a real registry client; v3 makes it a batteries-included SDK.
In five minutes we'll walk the headline v3 additions: a unified config stack that reads Docker and Podman credentials, registries.conf mirrors, and certs.d TLS in one call; built-in mirror fallback for air-gapped and enterprise setups; policy enforcement and OpenPGP signature verification via containers-policy.json; a ClientBuilder that wires auth, retry, TLS, and caching for you; and an experimental object-oriented API with fluent builders and typed models for artifacts, images, and indexes. We'll show before/after code and where to plug in.
If you build on OCI registries in Go, this is the five minutes that tells you what's coming and how to try it.
Join us for a lightning‐fast tour of k0s, the CNCF Sandbox distro that turns Kubernetes complexity into “one-click” magic. We'll give a quick intro to k0s, the current state of things, latest highlight and plans for the near future.
A Velero user told us their backup window was longer than the time between backups. Their database spanned four PVCs, each snapshotted independently, and every backup copied the entire volume. Something had to change.
Velero 1.18 fixed the first two problems. Concurrent backups eliminated the single-backup queue, letting teams run in parallel. VolumeGroupSnapshots capture all database PVCs in a single crash-consistent snapshot, no more hoping volumes line up.
Velero 1.19 fixed the third. The Block Data Mover uses Change Block Tracking to back up only blocks that changed. That 500GB volume with 5GB of daily writes? It's a 5GB backup now.
Along the way, Velero joined the CNCF Sandbox, a milestone for the community and a signal that Kubernetes-native data protection is here to stay. This lightning talk covers all three and shows what's ahead.
Over the past year, the OVN-Kubernetes community has continued pushing the boundaries of Kubernetes networking with new capabilities for multi-networking, network isolation, routing, observability, scalability, and performance.
This lightning talk provides a rapid update on the project's latest releases, the problems these new features solve, and how they're enabling production deployments across enterprises and cloud providers. We'll also give a preview of upcoming roadmap items and discuss how the community is collaborating with the broader Kubernetes networking ecosystem to drive future innovation.
Whether you're already running OVN-Kubernetes or exploring modern Kubernetes networking solutions, join us for a quick look at what's new, what's next, and how you can become part of the project.
In the last 90 days the KubeStellar Console merged roughly 2,400 pull requests across 211 releases. About 850 of them were opened by autonomous agents that triage the issue, write the fix, and merge through the same CI gates a human contributor faces. Five minutes on what that actually changes for a maintainer: what the agents handle well, what they consistently get wrong, and why the failures are never the code they write.
Each generation of workload pushes Kubernetes further. AI training scales thousands of GPU nodes in minutes. Inference images hit dozens GB. Agents run untrusted code and vanish in seconds. The plumbing was not built for this. You can fork, work around it, or fix it upstream. We chose upstream. First, we rebuilt containerd's image pull pipeline into disk-backed parallel pull: 60% faster cold starts, 8x less peak memory, merged upstream. The old design OOM-killed the runtime on GPU nodes pulling 30+ GB images. Second, we fixed bottlenecks in core Kubernetes that only surface at 100K nodes, a mutex serializing HPA decisions and a scheduler re-listing every PersistentVolume per pod. Third, come explore the hardest layer still ahead: how untrusted, ephemeral code can earn true VM-grade isolation at the level of a single pod, without giving up the speed that makes containers worth running. Two fixes already upstream, and a look at where the next one is headed.
Two years ago, the DevEx organization at CVS Health set an audacious plan in motion: build a platform capable of running the company’s thousands of existing applications while remaining approachable enough that someone joining the company could ship a new application on their very first day. To make this work, we knew we would need to leverage many of the incredible open source projects across the CNCF landscape. But we also knew that achieving our goal would require software we’d have to build: a meta control plane that could weave these projects together into a single streamlined platform. We call the platform CAP and we’re excited to share its design, show a brief demo, and talk about our transition from end user to open source maintainer and contributor as CAP evolves from an internal platform to a brand new open source project.
AI Agents offer unequalled automation tooling for rapid software development, but with unfettered power comes great headaches, namely how to scope and secure them in cloud native stacks. A problematic agent with too much access, either through choice or a lack of knowledge, can take down entire production stacks and wipe databases. In this talk, we explore how using core community building blocks like Confidential Containers and OpenShell, alongside emerging technology across the cloud native ecosystem, allows anyone to isolate their agentic workloads and software development processes at scale with secure, containerized sandboxes.
Cloud native promises a future without limits, but every end-user organization eventually discovers the same truth: no team becomes limitless alone.
Every day, end users push CNCF projects into new industries, at new scales, and with entirely new challenges. They uncover new patterns, expose new edge cases, and find better ways to build. By sharing those experiences back with the community through talks, bug reports, reference architectures, and open collaboration, they create a flywheel where every contribution accelerates the next innovation. The result is a community where shared experience helps everyone solve problems faster, see around corners, and innovate with greater confidence.
Together, Phill Morton and Abby Bangser will explore how end-user experience and community leadership fuel that flywheel. Through the upcoming Platform Engineering Community White Paper and a live demonstration of an enterprise AI platform built on those same principles, they’ll show how community ideas translate into real-world impact, enabling thousands of non-engineers to build at the speed of agents.
You’ll leave with practical ideas for applying those patterns in your own organization and a new perspective on what “Cloud Native Without Limits” really means: the greatest advantage in cloud native isn’t just the technology we build; it’s the community we build it with.
Agent harnesses are the dominant pattern for agentic AI on the desktop and will become the dominant runtime for agents on Kubernetes. The hard part is the move: bringing the harness onto your cluster with end user experience intact while adopting Kubernetes primitives for sandboxed execution, checkpoint/restore, scale, and security. Containers made this exact trip from developer machines to shared infrastructure in Kubernetes. In this keynote we'll present harnessed agents as your cluster's next critical workload. We'll show how harnesses combine models with a modular runtime of tools, skills, and plugins, and how leading CNCF ecosystem projects deliver the identity, policy, and observability to run them in production.
In order to facilitate networking and business relationships at the event, you may choose to visit a third party’s booth or access sponsored content. You are never required to visit third party booths or to access sponsored content. When visiting a booth or participating in sponsored activities, the third party will receive some of your registration data. This data includes your first name, last name, title, company, address, email, standard demographics questions (i.e. job function, industry), and details about the sponsored content or resources you interacted with. If you choose to interact with a booth or access sponsored content, you are explicitly consenting to receipt and use of such data by the third-party recipients, which will be subject to their own privacy policies.
On-premises compute is inelastic, and the costs are mostly up-front. Yet workload demands and hardware health can change at any time. How do you make the most of your on-premises investments without risking capacity overruns and potential outages?
Well, Karpenter and Cluster Autoscaler work great for right-sizing cloud. So what if we took their main principle and inverted it – scaling the workloads to fit the cluster instead?
Getting there required rethinking how pods compete for space, and how cluster capacity is modeled. The Trade Desk tackled both with custom scheduler plugins and a new horizontal scaler. This talk covers the techniques, tradeoffs, and lessons learned along the way.
When a previous supplier dropped support for Colorado in the spring of 2024, the Cybersecurity Center at Metropolitan State University of Denver needed to build a replacement quickly.
What started as four virtual machines scaled into a multi-cluster Kubernetes platform. Keycloak for multi-tenant identity, ArgoCD for GitOps, OPA Gatekeeper for policy, Envoy Gateway for traffic management, and cert-manager for TLS. All running in production for roughly $1,000 a month.
The core SOC runs on OpenSearch with Fluent Bit and Suricata, where student security analysts investigate real threats from live data. Student developers maintain the platform infrastructure itself, contributing to GitOps workflows and operations. Both groups gain hands-on experience with CNCF projects used in production worldwide.
This talk covers architecture decisions, the VM-to-Kubernetes migration, how we built a community around students and underserved organizations, and the mistakes made along the way.
Platform engineers are under pressure to do more with less. But where do you actually start with AI?
At T-Mobile, we stopped waiting and started experimenting. This is our honest account of what we tried, what failed, and what genuinely changed how we work.
We tackled two problems. Developers needed instant answers about their namespaces without filing tickets. And when incidents hit, engineers had no clear starting point. We built tools that solved both, cutting investigations from days down to hours.
Along the way we made real tradeoffs: whether you need MCP, which frameworks are worth it, and where simpler approaches won.
You will leave with a practical framework for evaluating AI tools and a clearer sense of what to build versus what to skip.
No AI expertise required. If you are just starting to explore AI in platform engineering, this talk is for you.
Everyone wants a fancy internal developer platform with all the bells and whistles. Most companies already have one. They just haven't productized it yet. The tooling needs some love.
Those Tofu modules you worked hard on, all the Helm charts you are maintaining and the reusable CI/CD pipelines? There's your platform, it's just a little fragmented, not intent driven and the developer experience could definitely be improved…
You're not building from zero. The harder job is taking what you already have, slapping a fresh coat of paint on it, and turning it into an intentional product with happy users and golden paths people want to follow!
This talk is about that shift. From scattered, fragmented tooling to a platform teams prefer to use over the old way. How to scope your first problem, why success criteria and platform principles need to come before code, and how to measure if any of this is worth the hassle.
Kubernetes has quietly become the default substrate for petabyte-scale data processing. Salesforce, Pinterest, Airbnb, and Apple all run native Spark, Flink, Trino, StarRocks, and Iceberg on Kubernetes at a scale the platform was never designed for.
This panel of end-user operators talks about two things data leaders keep asking about. First, what does it actually take to run large-scale data on Kubernetes? Shuffle bottlenecks and Apache Celeborn, stateful upgrades, multi-tenant scheduling with YuniKorn and Kueue, Iceberg storage layouts, mixed CPU and GPU. Each panelist shares what broke at scale and how they fixed it.
Second, how are AI agents changing day-2 operations? Troubleshooting agents for Spark and Flink, MCP servers for Spark History Server, agentic upgrade planning, autonomous tuning, AI-driven governance and lineage. Where are teams seeing real productivity gains, and where do agents still get it wrong?
Production stories, real disagreements, no vendor pitches.
Everyone is building AI agents. Almost no one is thinking about the infrastructure layer underneath them.
This talk reframes AI agents as cloud native workloads and shows how to build a production-grade multi-agent platform on Kubernetes. Topics covered: modeling agents as K8s custom resources, durable execution and retry semantics with Dapr, event-driven inter-agent messaging with NATS, and end-to-end observability using OpenTelemetry traces across agent hops.
We go beyond LangGraph demos to tackle the hard problems: failure recovery, agent scheduling, concurrency limits, and audit logging — the same concerns we solved for microservices, now applied to agents.
A gateway that looks flawless in a demo can collapse under production load. Routes churn, connection limits get pushed to the brink, and the architectural trade-offs hiding beneath modern gateway controllers suddenly matter.
This session moves beyond "Hello World" to expose what actually breaks at scale. Drawing from open-source benchmarking data across Envoy Gateway, Nginx, Agentgateway, Istio, and more, we dig into real failure modes: dropped connections during zero-downtime route updates, control planes consuming 100x more resources under churn, and latency spikes during route propagation.
As CNCF maintainers, we bring a data-driven framework for evaluating gateway implementations from real systems under stress.
You'll leave with empirical benchmarks and concrete strategies for selecting and tuning a Gateway API architecture that holds up in the reality of day-two operations.
Every cloud ships its own GPU taints, network plugin, storage class, and "don't forget this annotation" footgun. Most platform teams respond by forking their config per provider; and then nobody can answer "what driver version is actually deployed to prod in GKE?" without running Helm or kubectl against a live cluster.
In this talk, we'll show a criteria-driven recipe model (service × accelerator × OS × intent × platform) that generates Helm bundles for EKS, GKE, AKS, OCI, and CoreWeave from one library and a jq-style selector grammar that queries the hydrated result like a database. Live: three commands, three providers, one diff exposing exactly what's portable. Then one-line command answers the same question across every environment without touching a cluster.
Attendees leave with a pattern for cross-cloud GPU configuration that auditors, FinOps, and SRE can actually inspect: no per-cloud forks, no write-only GitOps repos.
Open source maintainers are facing a double assault. On one side, threat actors like TeamPCP are leveraging LLMs to scale attacks against well-known GitHub Actions misconfigurations, and propagating credential-stealing worms through ecosystems like PyPI and npm. On the other side, well-meaning security researchers are using those same LLMs to hunt for and report issues in good-faith, burying projects under a flood of (at times low-quality) vulnerability reports.
In this joint talk, we assume the roles of a combative offensive security researcher and a defensive open source maintainer to grapple with the tsunami of AI slop and irresponsible public disclosures using examples from in-toto’s Witness project. We will:
+ Recon: use LLMs for targeted vulnerability discovery and exploit writing
+ Report: attempt to communicate & disclose the discovery
+ Remediate: provide actionable solutions and actual code via private PRs
+ Reinforce: use AI to proactively secure the project
Plan, Code, Build, Test, Release, Deploy, Operate, Monitor, and repeat. This is the DevOps infinity loop the industry has been accustomed to since 2013, and it has symbiotically evolved alongside Cloud Native, but AI is reshaping every phase. This session introduces “Next Generation DevOps”; a framework for embedding AI in each phase of the DevOps infinity loop to accelerate & automate and where humans remain in the loop for accountability: architectural design record generation, code assistance, dependency agent, pipelines as agentic targets, generated TDD & unit tests, GitOps, rollout agent, canary plan generation, SRE runbook agent, anomaly and postmortem agent. Attendees leave with a framework for infusing AI into DevOps in the context of Cloud Native, including ArgoCD, Tekton, Kubernetes, Prometheus, OTEL, OPA, and more.
There is a massive gap between how easy autoscaling looks in vendor demos and how difficult it is to optimize in the real world. Many users assume that simply deploying an HPA is enough, but have no idea how fast can the Pods really scale when a traffic hits.
You'll be handed a pre-configured autoscaling setup for an HTTP app. Then, we will hit the cluster with a 4-minute benchmark – scenario mixing fast/light and heavy/slow requests.
The catch? Your initial results will be terrible. Our goal sounds simple: improve the system's efficiency score.
As we iterate through several rounds, you will run into real-world roadblocks. You'll discover that CPU doesn't correlate with your throughput, and maybe the application is leaking a bit of memory.
We will conclude with a friendly leaderboard to see who achieved the best results. You will walk away knowing the brutal reality of Pod-level autoscaling, and exactly how to bridge the gap between a basic installation and production efficiency.
SIG Network is responsible for networking for Kubernetes clusters, and there's never a shortage of interesting problems to solve in this space. In this session we'll provide some updates about SIG Network as a whole, including: status and progress of core networking components; status and progress of sub-projects; considerations for the future. If you're interested in hearing about what's going on in the networking space, or maybe even interested in joining the SIG and finding a place to contribute, please join us!
Kubernetes is the second largest open source project in the world, and its future is driven by an elected committee of seven people serving two-year terms. Especially in a project of our size, it’s important to work publicly and in the open, so join us for an open-format interactive session with representatives of the Kubernetes Steering Committee. Bring your questions about project governance, the future of Kubernetes, how you can get involved (or more involved!), or anything else you can think of that we might be able to help you with. A form will be provided to ask questions anonymously.
Since launching the AI Gateway Working Group, the Kubernetes community has been turning early ideas about AI-aware networking into concrete designs and emerging API directions.
AI traffic challenges traditional networking assumptions. Requests now include prompts, tool calls, streaming responses, and external model interactions, requiring new patterns for routing, payload handling, and policy enforcement.
In this session, we will present the latest work from the AI Gateway WG, covering emerging AI traffic patterns, payload processing and transformation hooks, external model egress, and how these map to Gateway API and its extension model.
We will also discuss key design tradeoffs: what should be standardized vs left to implementations, how much payload awareness belongs in the network layer, and how to avoid premature abstractions.
Attendees will leave with a clear understanding of the WG's direction, current designs, and the open questions shaping AI networking in Kubernetes.
The continued embrace of Kubernetes as a platform for AI workloads has presented new challenges for ensuring that clusters and workloads can scale and make efficient use of hardware. There's lots of updates and forward-looking news to share about how the various projects within the SIG are solving these problems.
In addition to providing project updates for Karpenter, Cluster Autoscaler, VPA, and HPA, we'll provide a thorough update to the SIG's effort to integrate infrastructure autoscaling into active efforts to nativize AI workloads on Kubernetes: DRA, Workload/Topology Aware Scheduling; and how we're building new abstractions to standardize core dependencies betweeen Cluster Autoscaler and Karpenter.
Attendees will leave the session with a better understanding of the the SIG's plans to tackle major AI/ML use cases, ensuring we can meet the needs of workload scaling on Kubernetes now and into the future.
Harbor is a leading cloud-native registry that powers secure, scalable container management in Kubernetes environments. This ContribFest session invites intermediate and expert contributors to collaborate directly with Harbor maintainers on high-impact areas, including outstanding technical debt, security enhancements, and feature requests.
Participants will have the opportunity to tackle real-world challenges, contribute to upcoming releases, and engage in hands-on work with maintainers guiding the process. The session will foster mentorship, knowledge sharing, and practical contributions, ensuring attendees leave with tangible results and a deeper understanding of Harbor’s architecture and roadmap.
Join us to make a direct impact on Harbor’s community and codebase, and help shape the future of a project that powers cloud-native infrastructure worldwide.
Fresh off the heels of graduating as a CNCF project, OpenTelemetry returns with its popular Contribfest session! This is a fantastic opportunity for new contributors to get involved with one of the most impactful open-source observability projects. You’ll be guided and supported onsite by maintainers from across the project, including the Collector, semantic conventions, language localization, entities, OpAMP, and more.
It’s never too late to join the thousands of OpenTelemetry contributors worldwide, and this session is designed to help you submit your first (or 10th!) contribution to the project.
Whether you’re a developer, SRE, or just curious about OpenTelemetry, this session will empower you with the skills and confidence to contribute to open source. No prior experience with OpenTelemetry is required; just bring your laptop and enthusiasm!
Kubernetes scales until it has to synchronize. Every write serializes through one ordered log, and every read funnels through lock-guarded caches: the apiserver's watch cache, and the informer caches inside every controller. Make the cluster busy enough (ceiling: ~10-16k mutations/s) and synchronization becomes the limit: caches lag, watchers get 410 Gone, and controllers reconcile against a world that no longer exists.
This talk is how Anthropic has improved Kubernetes stability by augmenting the control plane: in-cluster read replicas that take LIST and WATCH traffic off the apiserver, and controllers designed to stay functional at high mutation rates. The war story: the same innocuous LIST shape pushed a cluster of tens of thousands of nodes into an apiserver doom spiral twice in two weeks, from two different callers; only a read replica really fixes it. Along the way: what a stock controller does when its informer cache goes stale, and the data structures that avoid the trap.
What do a Kubernetes cluster, a chocolate cluster, and a cluster of extensions have in common? Yes, they are all "clusters" yet they represent widely different things in different contexts. When these
meanings collide without a shared contextual map, organizations accrue Communication Debt: the gap that leads to stalled platform adoption, onboarding friction and frustration.
Engineers are experts at using abstractions to design systems, but when explaining those systems, they often skip those layers entirely, leading with deeply technical details before establishing context or relevance.
This talk is designed for engineers and other technical practitioners who need to communicate Kubernetes concepts across the gap to executives, stakeholders, new contributors, and adjacent teams. Using real Kubernetes examples, attendees will learn a practical framework for translating complex concepts across the Cloud Native Computing Foundation (CNCF) ecosystem without oversimplifying them.
The maintainer talks tell you how cache-aware inference routing works. This talk tells you what happens after you adopt it and run a dozen live services on it for a year.
A Korean telecom moved from a single RAG pipeline to production agents on Kubernetes, adopting the Gateway API Inference Extension and llm-d for KV-cache-aware routing. The architecture worked in the demo. Production was a different teacher: cache-hit cliffs under real traffic, a small-versus-large model cascade whose economics only appeared at scale, and the realization that promoting a routing policy needs the same evaluation gate as promoting a model.
We share the failure modes the docs omit, the small/large cascade cost numbers, the measured 40-50% time-to-first-token reduction on cache hits, how we wired evaluation into routing-policy promotion, and the operational signals that actually predicted trouble.
EarnIn runs 15+ CNCF add-ons (including Flux, Karpenter, Kyverno, cert-manager, Linkerd, external-secrets-operator, Velero, and more) on a high-availability platform where infrastructure failure directly affects people's access to their earnings.
After catching too many issues in production — cert-manager chains breaking silently after upgrades, Velero backup jobs completing without actually running — we built an in-cluster infrastructure testing framework using Testkube, wired directly into our Flux GitOps pipeline.
This session covers how we model infrastructure tests as Kubernetes CRDs, how quality gates fire on every GitOps-triggered change, and which failure classes we catch across our CNCF stack. We'll share real test patterns, the architecture that makes it work, and the lessons from operationalizing infrastructure testing at a fintech company where "it worked in staging" is not good enough.
In the age of AI, supply chain security has become essential. Demos and MVPs are relatively simple to build. But scaling such tools across thousands of developers, pipelines, and clusters, with continuously changing workloads has proven very challenging. Such real-world environments introduce an entirely different set of engineering and operational requirements
In this session, architects from AT&T and Nirmata share lessons learned from operationalizing Kubernetes supply chain security at scale based on real production implementations at AT&T, which are processing >10,0000 image operations per second.
This deep dive explores how to achieve high-throughput supply chain security using Sigstore, GitHub Artifact Attestations, and Kyverno’s new ImageValidatingPolicy CRD. Powered by Kubernetes-native CEL, this architecture enables fast, cluster-wide image verification while also covering the best practices and scaling considerations required for rolling it out across Production environments.
Everybody wants the latest and greatest bits, right? But for the essential infrastructure that powers our world, shiny new releases may bring shiny new bugs. Where stability is top priority, the frenetic pace of an upgrade treadmill may not match your real needs.
More time for upgrades is appealing, though when looking at an LTS support cycle versus the release schedules of ancillary software installed on your systems, the fault lines start to show. If part of your system goes EOL before the rest can be changed, you risk having a CVE-shaped problem.
In an LTS environment, relicensing or forks can limit your ability to ingest necessary patches. And if your upstreams aren’t patching any more, you have a maintainability problem, where AI-fueled zero-days can appear faster than AI-assisted patches.
Resolving these conflicts is complex; we’ll detail how to approach it. You’ll leave ready to balance tradeoffs inherent in extended-period software support while also keeping systems secure.
Konnectivity (apiserver-network-proxy) requires a gRPC connection from agent to server. To protect sensitive workloads from data exfiltration, the speaker's org runs Kubernetes nodes in cloud VPCs with zero egress – no NAT Gateway, no route to the API server – making Konnectivity unusable.
This talk presents a fork that replaces the gRPC link with cloud object storage, accessed via private endpoints, as a rendezvous medium. The speaker walks through cluster network design and bucket transport architecture; how kubectl exec and kubectl logs sessions are multiplexed as sequenced bucket objects; and how proxy-server and proxy-agent coordinate ordered, reliable delivery without a persistent connection.
Object storage introduces hard tradeoffs in latency, cost, and cluster networking. The talk covers design choices made to reduce cloud costs at scale and achieve sub-second interactive latency over a store-and-forward medium. Includes a live demo with zero-egress worker VMs.
The first wave of platform engineering focused on Kubernetes abstraction: templates (Crossplane + OpenTofu), golden paths (Helm + Argo), and self-service APIs (often fronted by Backstage). Then AI coding agents arrived, and suddenly, platform building was a solved problem. Just add MCP, right? Maybe… actually, no.
Infrastructure requests are increasingly being generated by copilots, workflow agents, and automation systems operating at machine speed. Platform teams now need APIs and workflows designed for both humans and AI systems.
This talk explores how internal platforms must evolve from infrastructure-centric interfaces toward intent-driven platform capabilities that encode organisational policy, security, and operational knowledge.
Drawing on lessons from real-world platform implementations, we’ll examine the new failure modes AI introduces to platform engineering, why traditional abstraction layers break down, and how workflow-centric capabilities can scale effectively.
Confidential computing can verify a Kubernetes workload launched with expected measurements, but it cannot verify the entire sharing topology beneath it. Kubernetes still shares control planes, schedulers, kernels, image layers, nodes, devices, caches, buses, and roots of trust across tenants.
This talk builds a Kubernetes-specific threat model for confidential workloads: where attestation helps, where it stops, and which risks are patchable bugs versus structural consequences of sharing. Attendees learn a practical three-tier isolation model: shared namespaces and nodes; per-tenant control planes with dedicated node pools; and dedicated clusters per tenant.
The goal isn't to argue that Kubernetes is the wrong platform for confidential computing. It's to help platform and security teams ask the right questions before calling a deployment confidential: who is shared with whom, what is actually verified, and which gaps require topology changes rather than more cryptography.
Have you managed to build your application as a container image, with Docker or Podman, and push it into an image/artifact registry? Congratulations… but the work is not finished 🙂.
Now it’s time to implement the best practices of Supply Chain Security Software!
- Why and how to sign an image?
- Why and how to generate the inventory of my image?
- Does my image have vulnerabilities? Are they exploitable?
- What is the purpose of attestations?
- How to automate all this?
In this talk, step by step, with a mix of slides and live demos, we will see all that, and much more.
You will know how to implement these best practices, locally and even in your CI/CD pipelines.
And bonus track, after this talk terms like 'cosign', 'sbom', 'vex', 'in-toto'… will no longer have any secrets for you!
Running an LLM on Kubernetes is easy. Running it in production is a different story.
Once you move beyond a demo, you face challenges that go far beyond serving a model: GPU scheduling, traffic spikes, streaming responses, autoscaling, model updates, and keeping latency low while efficiently using expensive GPU resources. Production LLM workloads require new approaches to scaling, routing, resource management, and platform integration.
In this session, KServe maintainers will share how the project has evolved beyond traditional model serving. We'll discuss how inference engines, AI gateways, autoscaling, distributed serving, and traffic management work together through Kubernetes-native APIs to simplify running LLM workloads at scale.
We'll share lessons learned from building and operating these capabilities, including practical deployment patterns and design decisions from real production environment
WG Device Management continues to make great progress in enhancing support for GPUs, TPUs, NICs, and other specialized hardware in Kubernetes.
With the 1.34 release, Dynamic Resource Allocation (DRA) reached General Availability, making it easier than ever to configure, allocate, and share advanced hardware resources efficiently. Kuberentes 1.35 and 1.36 have built a great deal on top of that, but the work is nowhere near done!
Come learn about what the community has built in 1.36 and what's coming in 1.37 to improve how you use specialized devices in Kubernetes, such as features for managing device failures, management of node allocatable resources like CPU, and controlled sharing of devices.
Kubernetes operators keep hitting the same friction at the node level. Resource management decisions are locked inside the kubelet with limited configurability. Container-level knobs like ulimits or PID limits require OCI hooks or global kubelet flags instead of per-workload configuration.
NRI plugins intercept container creation to modify CPU sets, memory limits, and device assignments before a container starts, moving resource policy into out-of-tree plugins where vendors and platform teams can iterate without waiting on runtime releases. This session covers how NRI integrates with CRI-O, with plugin examples.
Beyond NRI, CRI-O is tracking upstream CRI changes. Per-container ulimits, per-pod PID limits, exec identity propagation, writable cgroups without privilege escalation, stop signals, and MemoryQoS are at various stages of landing. We cover where CRI-O stands on each.
Since KubeCon EU, we have also continued work on Spegel support for CRI-O and demo what is working.
As cloud native systems and AI agents increasingly interact across APIs and infrastructure, authorization is becoming a critical building block alongside authentication, requiring fine grained and real time decisions.
AuthZEN, finalized by the OpenID Foundation in 2026, defines a standardized API for interactions between policy enforcement and decision components, enabling more interoperable authorization architectures. As adoption grows, platforms need to be ready to support AuthZEN based architectures.
In this session, Yoshiyuki Tabata presents how Keycloak can serve as a unified platform for both authentication and authorization, providing an OAuth/OIDC based identity foundation alongside emerging AuthZEN based authorization capabilities.
The session demonstrates an end to end scenario including token issuance, authorization evaluation, and enforcement. It uses Envoy as the policy enforcement point together with Keycloak to show how this approach can be implemented in practice.
The cloud native ecosystem has spent a decade mastering how to run software. Now the hard problem is making it genuinely easy to build. TAG Developer Experience isn't just philosophizing about this shift, but we're shipping the standards to make it real.
The TAG DevEx chairs and tech leaders will give maintainers an unfiltered look at the work defining the next decade of cloud native productivity. We'll cover our OCI-compliant inner-loop tooling standards for AI engineers, the Application Integration Dependency Specification for declaring how apps interact with cloud services, our State of AI-Assisted Development research, and shared DevEx benchmarks across CNCF projects.
You'll leave with a concrete Developer-First framework backed by real contributor friction data. If you maintain a CNCF project and want to know what good looks like, this is your session.
Every September, I race my friend Malcolm at the Monster of Mazinaw, a 43 km trail ultra in eastern Ontario. Every year, Malcolm wins. So last year I built myself an agentic running coach.
It runs as a kagent Agent CRD on Kubernetes. It reads my Strava and Peloton activity, remembers my training history, and messages me to keep me on track.
It is an event-driven system that has to handle the messy parts: training data that disagrees with itself, memory across weeks, schedule conflicts, and the Strava and Peloton APIs going down.
This talk shows how I built it, where it broke, and why this same pattern works beyond running. The same techniques work for a FinOps agent that watches cloud spend, or a team-specific assistant that knows your runbooks.
kagent, Kubernetes, and open-weight models put this within reach for one person and one problem.
When off-the-shelf solutions stop scaling, you build your own or break trying.
At Nubank, distributed tracing isn't a nice-to-have: it's the backbone of observability across thousands of microservices, processing hundreds of billions of spans every day with zero tolerance for data loss. When the need for full ownership and operational independence became clear, we built a hybrid pipeline from the ground up.
This talk walks through the full tracing journey at Nubank, from instrumentation strategy to a production-grade pipeline operating at petabyte scale. We'll cover how the OpenTelemetry Collector serves as our ingestion backbone, the architectural decisions that made the system resilient, the failure modes we didn't anticipate, and the hard-learned trade-offs behind blending open-source tooling with purpose-built internal components.
You'll leave with a realistic picture of what it takes to run distributed tracing at scale.
Every mature codebase earns a backlog item that's too daunting to complete. For Kubeflow Pipelines it was the frontend: an aging Create React App on end-of-life packages with a dated Jest + Enzyme suite – attempted before, never finished.
What changed was timing. An inflection point in AI coding agents made this kind of work – mechanical, well-specified, backed by tests – finally tractable: CRA to Vite, Enzyme to Vitest + Testing Library, React 16 to 19.
But agents didn't do it alone, and neither did I. This talk is what a career in engineering still brings: spotting patterns, steering an agent with the right questions, knowing when it's confidently wrong, and earning the trust of a community of contributors and maintainers who shaped and reviewed the work. You'll leave knowing how to spot the work that fits this AI moment, and to do it in a way maintainers welcome.
Organizations in regulated industries need AI inference on infrastructure they control, but NVIDIA GPUs are expensive and scarce. Apple Silicon devices with large unified memory pools sit underutilized.
This talk shows how we extended Kubernetes to schedule AI inference across heterogeneous GPU hardware, including Apple Silicon via Metal, using pluggable runtime backends. We cover a controller-runtime operator managing model lifecycle, multi-GPU layer sharding, and health probing across runtimes like llama.cpp for text generation and Moshi for speech-to-speech.
Attendees learn patterns for reconciling inference workloads with CRDs, pre-flight memory validation for GPU nodes, Prometheus metrics via PodMonitors, and multi-runtime inference behind an OpenAI-compatible API. We share benchmarks comparing NVIDIA CUDA and Apple Silicon Metal and the operational tradeoffs of each.
A practitioner talk from an enterprise platform engineering team deploying on-prem AI inference at scale.
Kubernetes made infrastructure scalable, but not always cost-efficient.Organizations have adopted various resource optimization strategies to improve resource utilization and reduce overprovisioning. Some work well across a broad range of workloads, while others show limitations at scale.
With its GA graduation in Kubernetes 1.33, in-place pod resizing has become a new resource optimization capability for platform teams. Like any approach, it has its fit, limits, and ideal workloads. For some applications, it can be a game changer, improving utilization without disruption. For others, especially legacy applications, the journey is more complex and often requires alternatives.
In this session, Meghana from RBC and Dolis from Nirmata will share lessons from enterprise optimization efforts, focusing on in-place pod resizing. They will discuss where it works, where it falls short, common adoption challenges, and how to evaluate the right optimization strategy across workloads.
You locked down RBAC. You scoped your namespaces. You followed least privilege. But the attacker who just became your GCP Org admin? They never left their namespace. The call came from inside the cluster—through your own operator.
I discovered confused deputy vulnerabilities (CWE-441) across Config Connector, ACK, ASO, and Crossplane. The pattern is universal: operators trust user-controlled CRD fields and act with elevated credentials. Your namespace user writes a manifest; the controller escalates them to cloud admin.
You'll see:
– Demo: namespace user → GCP Org admin in 60 seconds
– Why audit logs blame the controller, not the attacker
– A 3-step pattern to audit any operator for confused deputy flaws
– Architectural fixes you can deploy immediately
This research is also accepted at Black Hat USA 2026. This talk is Kubernetes-native: how operator architecture enables this flaw class and how to build controllers that fail closed.
For years, the GitOps conversation has been framed as a binary choice: Argo or Flux. But as platform engineering matures, it has been proven that these two projects aren’t mutually exclusive, rather, they are complementary.
In this session, Argo and Flux join forces to challenge the animosity that is traditionally perceived between the two projects. Grounded in the principles found in the OpenGitOps Working Group, we demonstrate how to leverage their unique architectural strengths: Argo CD’s application-centric control plane with strong visibility, and Flux CD’s modular, controller-driven system with native Helm lifecycle management. We will showcase real-world hybrid patterns, and explore tools and best practices on how to manage the two tools together in one single platform.
Attendees will leave with a practical framework for deciding when to use a single tool is sufficient, and when combining them unlocks a more resilient, flexible platform.
Auto-instrumentation has long been a major topic in observability. eBPF introduced a new approach: instrumentation is no longer confined to userspace agents and technology-specific SDKs. By leveraging visibility from the Linux kernel, eBPF-based instrumentation can capture kernel and application runtime events and model them into traces, metrics, profiles, and service topology.
This talk takes a ground-up look at OpenTelemetry eBPF Instrumentation (OBI), following the path from eBPF programs at the Linux kernel level into the userspace pipeline where events are correlated, enriched, converted into OpenTelemetry data, and emitted. We will examine how OBI combines kprobes, uprobes, socket and cgroup instrumentation, runtime-specific instrumentation, process discovery, context propagation, and Kubernetes metadata enrichment.
This session is for engineers interested in how zero-code eBPF instrumentation works under the hood. Attendees will leave ready to debug, extend, or contribute.
Everyone said AI would reduce developer cognitive load. In some ways, it did. But a different pattern emerged: developers stopped worrying about deployments and started worrying about models, tokens, context windows, and agent behaviour.
The load did not go away; it just moved.
The platform has no opinion on any of it. Every developer figures it out alone, every day.
That is prompt fatigue. And it is a platform problem.
This session introduces the Prompt Layer: the surface area every IDP needs to own. We'll walk through five components: capability interfaces that abstract model selection, context templates as AI golden paths, token governors as resource limits, agent permission policies using OPA or Kyverno, and AI observability through OpenTelemetry.
These are not new tools – they are the same abstraction patterns platforms already apply to infrastructure, extended to cover AI decisions.
Platforms solved the infra floor. The Prompt Layer is the next one.
Platform teams have shift-left covered: scanners on every PR, admission control on every deploy, SBOM+VEXs on every image. Then a high-severity CVE drops with "patch coming soon," and all of it goes quiet for days while a known-exploitable workload sits in production. This is the most uncomfortable corner of day-2 on Kubernetes, and most teams have no declarative answer for it.
We'll show how Kubescape, the CNCF incubating project for Kubernetes security, makes runtime detection a Kubernetes resource: a Rule CRD you review in a PR, version in Git, deploy via Argo or Flux, and roll back like any other manifest. Using the TeamPCP Trivy compromise as an example, we'll cover what the eBPF sensor sees, how to turn an advisory/IoC into a Rule CRD live on stage, how SBOB profiles keep the signal trustworthy, and a CI/CD workflow that gets from "CVE published" to "rule deployed" in under an hour. For platform engineers, SREs, app developers, and security teams alike.
Every AI and analytics system pays the same hidden cost: moving streaming data from where it's produced to where it's used — GPUs for training, tables for analytics, indexes for serving. The usual path re-serializes every row (JSON, protobuf) and grows a separate pipeline per destination. We show a different foundation, running in production on Kubernetes: keep data columnar end-to-end with Apache Arrow, and move it once over Arrow Flight. One write from a producer becomes durable, then lands everywhere it's needed — into GPU memory with no copy, into Delta Lake and Iceberg tables, and into serving systems like Pinot and ClickHouse. Because nothing is re-serialized on the hot path, moving data barely uses CPU; the real work is durability and commit. We cover the fundamentals — columnar transport, decoupling acks from slow sinks, idempotent commits, recovery after pod failure — and share real numbers from a multi-tenant cluster.
Inference clusters today waste expensive accelerators by handing each LLM a whole GPU, even when many models could safely share one. Runtimes such as vLLM and TensorRT-LLM happily run multiple workloads on the same device, but deciding which workloads belong together is a combinatorial decision that depends on each model's configuration and the traffic it actually sees.
This tutorial shows attendees how to delegate that decision to a closed-loop system. A long-running agent profiles each candidate workload in isolation under its target runtime and emits a typed characterization. A constraint solver consumes those characterizations and proposes a co-location plan. A Helm chart turns the plan into a fleet of pods served by whichever runtime each model prefers. Prometheus and DCGM watch the result and trigger re-planning when drift appears.
Attendees follow the chart end to end on a live cluster and leave with a reusable open-source pattern they can fork.
Container images have become the standard way to package software, but they've also made developers responsible for maintaining Dockerfiles, managing base images, and responding to endless CVEs.
Cloud Native Buildpacks provide a better abstraction by separating application code from the underlying OS and runtime, allowing platform teams to own image construction and security while developers focus on writing code.
In this session, you'll learn the fundamentals of Cloud Native Buildpacks, how to get started using them, and how Salesforce uses them to generate integration test images that let developers write custom test scripts without maintaining the underlying container image. You'll leave with practical patterns for reducing container maintenance, simplifying vulnerability management, and enabling developers to safely extend containerized environments without becoming container experts.
The AI boom created a distinct operational challenge for open source security. Instead of uncovering hidden zero day exploits, automated scanners frequently flood maintainers with perfectly formatted, technically accurate reports that completely ignore the project threat model.
This talk provides a realistic look at how the Kubernetes Security Response Committee (SRC) handles the modern vulnerability gold rush. We will explore the heavy triage tax caused by AI submissions and how we counter it using radical pragmatism. Attendees will learn how we partner with HackerOne to enforce strict threat model rubrics, managing waves of duplicate reports without burning out maintainers. We will also detail our internal OPSEC practices, specifically why the SRC relies on simple, local, air gapped LLMs to safely process embargoed vulnerability data.
AI-powered vulnerability discovery is here, and in 2026 it is no longer just a noise generator. These pipelines can uncover real 0-days, but a plausible finding is only the start. Maintainers must still verify the claims and determine their actual impact.
As Metal3's Security Lead, I am the primary analyzer and coordinator for incoming vulnerability reports. I've also built an AI vulnerability discovery pipeline, giving me insight into both sides of the process. Bare-metal provisioning is a useful case study: modern cloud-native architecture meets legacy drivers, deprecated protocols and external integrations. Vulnerabilities are rarely proven by a simple automated crash.
We examine what these pipelines can produce, where their reports fall short and what maintainers look for during triage. Users will see the work behind a security response, while researchers and contributors will learn what turns pipeline output into an actionable report.
Since this is the first maintainer track after incubating, First i'll give a brief introduction about HAMi. HAMi(heterogeneous accelerator virtualization middleware) is an open-source, cloud-native GPU virtualization middleware that brings sharing, isolation and scheduling of heterogeneous accelerators to AI workloads on Kubernetes.
Then i'll introduce our newest scheduling features such as combined scheduling policies, auto-scaling enablement, AMD-device-support, flexible-dynamic-mig.
In the last section, i'll introduce a brand new lightweight HAMi solution——HAMi DRA, which is based on DRA approach, it can be naturally compatible with other scheduling projects(like KAI-scheduler, volcano, etc)
Emissary-ingress 4.1 is shipping! Emissary 4 has been a big change to the way one of the first Kubernetes-native, self-service API gateways and ingress controllers gets built and maintained, and we'd like to share what's been going on and what's coming up.
In this session, we'll start with a quick overview of the need for ingress controllers in general, the benefits of self-service developer workflows, and how Emissary-ingress can help with these issues. We'll also cover what Emissary 4 brings to the table and how to get involved as a contributor. Emissary's maintainer sessions are always great opportunities to talk directly with Emissary-ingress maintainers and make sure your voice is heard when it comes to the project's future — looking forward to seeing you there!
Join us to contribute to a growing community that is building local container tooling for developers and administrators! Attend the Podman Container Tools contribfest and help a CNCF Sandbox project! This is a hands-on workshop where project maintainers will help you make a contribution to the project. First-time contributors are welcome!
This ContribFest will be focused on improving the Podman Container Tools community. We will have a curated list of issues where you can contribute to Podman, Buildah, and Skopeo, with varying levels of difficulty and different areas to contribute. These will include features, bugfixes, documentation changes, and helping with our websites.
No prior experience with Podman is needed, but we recommend bringing a Linux laptop if you intend to contribute to Podman (Windows and Mac are fine for working on the website!)
Want to contribute to Istio but feel overwhelmed just setting up the local environment? You're not alone. Join Istio maintainers for a hands-on workshop designed to take you from a fresh fork to running tests locally with confidence.
In this session, we're skipping the heavy infrastructure headaches. We'll use GitHub Codespaces to get you straight into the Istio ""inner loop. Maintainers will walk you through navigating the codebase, running testing frameworks, and getting fast test feedback. From there, we'll pair up to tackle a curated list of entry level tasks, like adding unit tests or fixing open issues.
Whether you want to fix a bug or just see how maintainers validate code daily, you'll leave this session with everything you need to start contributing regularly.
This session is your guide to becoming an impactful and effective CNCF member. We’ll cover the tangible benefits and responsibilities of membership, including providing project support, contributing code, fostering diversity, and participating in governing bodies. We will debunk common misconceptions and provide a clear framework for how a company, large or small, can actively shape the future of cloud native technologies. You will walk away with actionable steps and a deeper appreciation for the collaborative "movement" that drives innovation at the CNCF.
Your AI agent is Dory. Every conversation it wakes up unable to remember a thing , so you, Marlin, re-explain everything every turn. Four open-source memory systems promise to fix that: Cognee, MemOS, Honcho, and MemPalace. But which architecture actually pays off in tokens, dollars, and resolved tasks?
This talk runs two experiments. First, we benchmark all four backends head-to-head on a single Kubernetes substrate , same harness, same workloads, same OpenTelemetry Collector and rank them by token efficiency, cost per resolved task, and recall under contradiction.
Then we test the question most teams haven't asked: is running an agent with memory worth it at all? We measure two real OSS agent frameworks , kagent (CNCF Sandbox) and sympozium with their built-in memory turned on, then turned off. Real workloads, real numbers.
None of the four backends ships first-party OTel today; the Helm chart, OTel pipeline, and upstream PRs are how we close that gap for the ecosystem.
We introduced AI-assisted incident response into a large-scale financial services platform running on OpenShift and Kubernetes expecting faster resolutions and less engineer toil. Instead, we discovered the biggest bottleneck wasn't signal analysis—it was missing operational context.
AI could detect CPU spikes and latency increases, but it couldn't know a deployment had occurred 20 minutes earlier, a service owner had changed, or an upstream dependency was degrading. That gap led to poor prioritization and engineers chasing the wrong signals during incidents.
This session shares how we rebuilt incident response using OpenTelemetry, Argo CD, Prometheus, Grafana, and Backstage to combine telemetry, deployment intelligence, and service ownership into a single workflow.
Attendees will learn practical patterns for building context-aware incident response on Kubernetes, reducing alert fatigue, and improving operational decision-making beyond prediction-based tooling.
We've been knee-deep in Kubernetes since the early days: maintaining Helm, building Krustlet, provisioning internal clusters with Cluster API, containerizing applications, the whole shebang. Eventually, we swore it off and worked on Wasm because we were tired of watching people hammer everything to fit into a Kubernetes-shaped hole. Now we find ourselves back in it again after learning that both extremes are traps: putting everything in Kubernetes and pretending it doesn’t exist.
In this talk, we'll cover the red flags that signal you're using Kubernetes where you shouldn't, and where it actually shines. Then we’ll talk about the happy middle: cloud-native over Kubernetes-native. We’ll share our success stories and gotchas using CNCF projects both inside and outside of the Kubernetes space, including projects like wasmCloud, NATS, Cilium, and OPA. You’ll walk away with a concrete set of guidelines to help you find that middle ground in your own stack.
Artificial Intelligence is reshaping retail by enabling a shift from traditional automation to autonomous decision-making systems. This talk explores how modern retail organizations leverage AI, real-time data pipelines, and cloud-native architectures to transform data into actionable intelligence.
The session highlights the evolution from batch-based processing to event-driven systems, where AI models are integrated into operational workflows to support real-time decisions in areas such as pricing, supply chain, and customer analytics. It also discusses key architectural patterns, including streaming data platforms and intelligent automation frameworks.
Through practical insights and real-world examples, the talk demonstrates how autonomous decisioning improves efficiency, scalability, and customer experience. It also addresses challenges related to data governance, observability, and responsible AI adoption, providing a roadmap for building future-ready retail systems.
After 10 years of Kubernetes, we are still asking the same question: is this resource healthy?
Built-in resources are mostly predictable, but CRDs define status, conditions, phases and readiness differently. This creates Health Check Hell, where every tool must rediscover what “healthy” means.
As a result, tools repeatedly reimplement health interpretation for the same resources, creating duplicated work for maintainers, confusing behavior for users and operational drift between tools that watch different status fields.
The session explores why Kubernetes lacks a consistent health model for CRDs through real-world examples. It reviews existing approaches like Kstatus and explores potential solutions, such as reusable health libraries and ecosystem-wide compatibility standards.
Attendees will leave with a deeper understanding of Kubernetes health semantics, the tradeoffs behind current approaches and emerging efforts to improve health interoperability across cloud native tooling.
In April 2026 Verkada's incident-triage agent handled 1,847 production alerts across nine Kubernetes clusters at $0.33 per investigation, with an 86% prompt-cache-read ratio and an 82-second median time-to-summary. The control plane runs in-cluster on every shard. No AI-for-SRE SaaS is in the path; inference is managed Claude Sonnet 4.5, state lives on managed KV.
Three reversals shaped the system and each cut cost or restored reliability: pod-to-serverless ingress so chat deploys never drop alerts; per-pod to cross-cluster state on managed KV so a single investigation can span shards; single-shot to multi-turn under a 1-hour cache-TTL budget that holds 4-turn cost roughly flat.
MCP is the boring substrate that made the rest tractable: per-cluster MCP servers for ArgoCD, Prometheus, and the observability stack, with cross-cluster fan-out through the serverless ingress, and per-channel Markdown system prompts that retune triage behavior without a code change.
Resource quota is a veteran Kubernetes feature, widely used yet poorly understood. Simple in theory—track, limit, reject—it navigates the messy trade-offs of consistency, availability, and performance inherent to distributed systems. This talk dissects the lifecycle of a ResourceQuota, from validating admission to the reconciliation controller. We examine design choices like per-namespace batching, optimistic updates, and the lookaside cache used to dodge etcd conflicts. We also explore "default-deny" semantics and documented race conditions that reveal why distributed quota is hard. Then, we pivot to how we built a custom system for complex limits that upstream doesn't natively support. We cover the pitfalls facing anyone building quota at scale: negative reservations, hierarchical lock ordering, and two-phase semantics. Finally, learn how drift correction avoids clobbering in-flight writes. Leave with a robust mental framework to evaluate every quota system against the CAP triangle.
Distributed traces show how requests move across services, but they dont explain what happened on the network side. When latency, packet loss, routing changes, or unreachable hops affect an application, engineers still must leave their observability workflow and reason from separate network tools.
This session shows how traceroute-style path measurements can be modelled as OpenTelemetry traces. By representing each hop with spans, timing metadata, reachability information, and network attributes, path behaviour becomes part of the same telemetry model used for application debugging.
We will cover the trace model, span hierarchy, timestamp normalization, incomplete paths, unknown hops, and strategies for linking network-path traces with application traces for end-to-end correlation. Through a live demo, attendees will see network paths rendered as trace waterfalls and service maps, making it possible to connect application symptoms with the network behaviour behind them.
Multi-cluster Kubernetes solves a real platform problem. It also breaks every developer workflow that used to work. Adopt mulit-cluster frameworks for fleet resilience, and overnight kubectl stops behaving: kubectl get pods sees one cluster, kubectl scale gets reverted because the hub template is the source of truth, and finding a single pod takes a four-command discovery dance.
We built a thin reverse proxy that gives the single-cluster experience back. To the developer, the fleet looks like one cluster. Every kubectl verb just works — including kubectl edit deployment against the hub template. Underneath, the proxy translates writes into hub-side edits with optimistic concurrency, fans reads out in parallel, and stamps every result with the cluster it came from.
This talk walks the design, the surprising parts (kubectl edit works), the honest hard parts (multi-cluster WATCH, LIST pagination, partial failure), and which parts port to other multi-cluster solutions.
Imagine this: you've locked down your Kubernetes cluster with admission webhooks and ValidatingAdmissionPolicies. Life is great, until someone runs "kubectl delete validatingwebhookconfiguration" and your entire policy layer vanishes. That's the catch: the thing enforcing your security rules can be deleted by the very API it's supposed to protect. Your policies don't exist during bootstrap either, and they disappear if etcd goes down.
What if you could load those same webhooks and policies from files on disk, active before the API server serves its very first request, and invisible to the REST API? Even better, what if they could protect admission resources themselves, something REST-based policies were never allowed to do?
In this session, we'll show how Manifest-Based Admission Control (k8s.dev/resources/keps/5793) works, live-demo an attacker deleting critical policies and getting stopped cold, and share patterns for shipping tamper-proof configs across your fleet.
Kubernetes has become the default platform for running LLM inference in production. But most teams are hitting the same walls and assuming everyone else has it figured out. They don't.
Recent survey data from 200 AI practitioners shows where your peers actually stand: half of all production deployments are failing to meet latency targets at peak load, 65% of teams say GPU capacity planning is their hardest scaling challenge, and 22% are more likely to view proximity as a critical requirement for cloud selection.
So what separates teams making progress from those stuck in the same loop? We'll walk through the failure points across scheduling, scaling, and networking that stall most Kubernetes inference deployments, how aligning platform and application ownership removes the friction, and give you a repeatable path from cluster setup to reliable production endpoint. You'll leave knowing exactly where you stand relative to your peers and what to do about it.
Kubernetes is supported by a wide contributor community, and the structures that guide the community can be hard to understand at first. SIG Contributor Experience (ContribEx) helps connect those pieces through sub-projects focused on communication, workflow, and governance. SIG ContribEx brings in changes to its own sub-projects and intra-SIG governance structure in 2026, to better reflect how the group actually works today. The panel walks through the changes and shows how they better impact the contributor journey.
We also emphasize how ContribEx is shaping the contributor journey in the fast moving world of AI. We also show where we are leveraging AI to accelerate our process and what integrations are we enabling for the community to use and enhance their journey.
After serving the cloud native community strong for the past 5+ years, Helm 3 is reaching the end of life. Instead of dwelling on the past, it's time to look towards the future and if you haven't had a chance to upgrade to Helm 4, now is the time to do it. Several key enhancements were fundamental in the development of Helm 4 and these features are just the start of what lies ahead.
In this session, Helm maintainers will share how Helm 4 has helped resolve many of the challenges inherent within Helm 3 and laid the foundation for expanding the capabilities provided by the tool. With a rich set of sub projects beyond Helm itself and a growing community, there are multiple ways that potential contributors can get involved. Learn why Helm continues to be the tool of choice for packaging Kubernetes applications.
High cardinality quietly breaks metrics systems: runaway label sets, exploding series counts, slow queries, and OOMing components. In this session we'll dig into how Cortex, the open source, horizontally scalable, multi-tenant metrics platform, ingests and queries high cardinality data without falling over.
You'll leave knowing which label pattern drives cardinality growth inside a tenant, how to stop one bad query from taking down a querier, and how to query blocks with huge label sets efficiently, thanks to per tenant cardinality tracking, resource based query eviction, and sharded Parquet querying landed in recent releases. We'll close with the roadmap, including the push toward CNCF graduation and what came out of the recent security audit.
Whether you're new to Cortex or already run it in production, you'll leave with practical techniques and a chance for live Q&A with core maintainers.
Agentic AI is changing what cloud-native infrastructure must schedule. AI training, inference, and RL workloads continue to stress scheduler throughput and coordination, while agent workloads add pressure on startup latency and burst handling. The question is no longer how to optimize one scheduling path, but how to balance throughput, latency, and resource utilization across shared Kubernetes clusters.
In this maintainer session, we will share how Volcano is evolving from a batch-oriented scheduler into a unified scheduling framework for AI training, inference, emerging RL workloads, and agents. The focus is the new framework: a coordinated architecture that improves scheduler parallelism while handling conflicts among scheduler workers and keeping scheduling behavior consistent across workload paths.
We will also cover community updates: benchmark results, topology-aware placement, Volcano Global, colocation improvements, and roadmap.
Did you know in-toto is an ingredient in your SLSA, and that you can put in-toto in Git? If you verify attestations using Sigstore, it is likely you have in-toto somewhere. However, like the name implies, in-toto is not meant to be a series of separate step-wise checks, but a holistic security framework. Through user stories that highlight how organizations have incorporated in-toto into build, release, provenance, and policy workflows, this panel will show real-world usage and broad range of applications of in-toto. We will share practical adoption advice: what users start with, which parts of their supply chains they modeled first, how they integrated with existing systems, and what trade-offs they encountered.
We will also discuss our work in identifying in-toto's usage, as maintainers of open source projects are often unaware of where their work is used. As with in-toto's use itself, we don't intend for this to be a one-size-fits-all solution, but merely our experiences.
Agents and agent-like workloads are becoming a predominant class of workloads on Kubernetes, but they are different from "traditional" apps. In the last 18 months, those differences have been starkly emphasized. For over a decade K8s has been the industry standard platform for workloads of all sorts, with a rich ecosystem that has successfully adapted to every new class of workload. It’s time to adapt again.
The Agent Substrate (http://ate.dev) project was born to tackle the problem of agent-like workloads head-on.
Launched in May, Agent Substrate’s primary goal is to make it easy and efficient to run agent-like workloads. It leverages Kubernetes for what it is best at, then builds up from there to address the particular needs of agents. A lightweight control-plane coupled with active session management and network-triggered wake-ups enables users to run orders of magnitude more agent-like workloads than before – more efficiently and with lower latency and higher throughput.
Monitoring a Kubernetes platform with 2M+ containers exposes the limits of standard tools. Most setups track internal controller machinery, ignoring user friction. Teams outgrowing kube-state-metrics often build custom metrics focused on internal mechanics rather than user experience, producing noise instead of insight.
Effective platform observability needs to measure the promises made to developers.
This session designs an SLO framework unified across stateless and stateful workloads using a contract-driven model. Learn to distill thousands of data points into 3–4 critical metrics that drove our operations this past year. We track workload readiness alongside pod-lifecycle signals (init-container latency, image download waits, terminating stalls) to isolate infrastructure faults from user errors before a page fires. This continues KubeCon NA 2025 ("Making Application Rollouts Observable, Actionable, and Boring"), moving from rollout signals to platform-wide promises.
As with many large organizations, cloud native adoption and use often begins in pockets. However, even as innovative patterns are developed and usage increases, some teams within the same organization can remain unaware.
This was true at Adobe. While the public cloud platform team were building leadership internally and credibility in the cloud native community, Adobe’s private cloud teams were busy focusing on other efforts. When business needs shifted, it needed to move quickly to a cloud-first architecture.
In this session, learn how a team managing traditional infrastructure rapidly pivoted their technology stack and approaches using solutions such as Kubernetes, Argo CD, and KubeVirt. We’ll highlight the traits that made their journey successful, including developing reusable assets, establishing strong partnerships, and community engagement. Beyond the technology, this talk shows how tough transformation challenges can be achieved with the right ingredients and approach.
The rise of AI agents is breaking traditional ML infrastructure. While human experimentation is measured in hours, agents researching new model architectures and running hyperparameter sweeps shrink this to minutes. When running 24/7, Kubernetes pod spin-up times and GPU bottlenecks hinder agentic velocity.
In this talk, we’ll explore how to build high-velocity agentic infrastructure using Kubernetes, Volcano, and HAMi. We’ll cover techniques including:
- No-Gap GPU Execution: "Firepool", maintaining warm GPU pods for near-instant launches including zero gap job launches on GPUs.
- Secure Agentic Harness: A workflow based Sidecar pattern isolating untrusted AI logic from credentials.
- Fault-Tolerant Persistence: Mechanisms to recover state if an agent pod gets disrupted.
- Cross-Run Agent Memory: Shared context patterns injecting state across multiple agent runs.
Attendees will leave with strategies to scale 24/7 agentic K8s workloads
AI and agentic systems are becoming production workloads on cloud-native platforms, and the hard questions land with platform teams. How do agents identify themselves? What can they access? How fast can they act? What can they cost? How are their actions audited? What evidence does the platform need when something goes wrong?
This panel brings together four perspectives that usually handle these separately: an end-user platform practitioner, an OpenTelemetry maintainer, a security practitioner, and an AI infrastructure engineer. Together they examine what shared infrastructure must guarantee as AI systems become governable workloads: identity for non-human callers, observability beyond human dashboards, policy and blast-radius controls, inference and routing demands, build-versus-buy, and accountability across application, model, and vendor.
Attendees leave with a practical frame: what today's CNCF substrate can support, where standards emerge, and what platform teams must own.
Migrating a CNI plugin in production is one of the most high-risk operations in Kubernetes infra. At Walmart running hundreds of clusters across cloud and on-premises environments, the team faced route-drop failures in Flannel beyond 300 nodes, iptables rule explosion as services grew and etcd v2 incompatibility blocking upgrades. No published playbook existed for zero-downtime CNI replacement at that scale.
This talk walks through the per-node hybrid migration strategy that enabled zero-downtime transitions from Canal to Cilium. Canal and Cilium coexist during transition with a hybrid tunneling mode maintaining full pod connectivity regardless of which CNI owns a given node. Result: 2.27x throughput improvement and 38% reduction in cross-node latency with no production incidents.
Attendees leave with a reusable migration framework covering hybrid mode architecture, IPAM coordination, cloud provider pitfalls and a structured Cilium triage playbook for post-migration operations.
Everyone's excited about frontier models as a service. But what if you need to run and serve them yourself, for latency, cost, or control? You end up paying for a whole GPU so one pod can use a fraction of it. The Kubernetes scheduler never knew better: one pod, one whole GPU, end of story.
Two things just changed that. Dynamic Resource Allocation went GA in Kubernetes 1.34, the biggest shift in device scheduling since 2017. The scheduler can finally see what a GPU offers, not just count cards. And NVIDIA's Multi-Instance GPU splits one card into isolated partitions. Combined, you get real, hardware-level sharing. No time-slicing hacks, actual isolation.
So we built the stack from zero: cluster, GPU operator, DRA driver, MIG partitioning, ML workloads on one GPU. It's open source, ready to lift onto your platform. I'll walk the architecture and run a live demo. You'll leave knowing when DRA + MIG beats time-slicing, how to roll it out safely, and where the gotchas hide.
Calculating and tracking the memory usage of your applications is far more complicated than meets the eye. In this session, two members of OpenTelemetry's System Semantic Conventions Working Group will do a deep dive into what how memory metrics differ from each other, how the Linux Kernel will respond to memory pressure, and what memory information you should be using to effectively detect problems in your infrastructure.
Most Kubernetes environments end up multicluster not by design but by accident. By the time teams ask how to do this right, the decisions have already been made for them. By tools, by vendors, by whoever got there first.
This panel brings together contributors, end users, and tool builders to answer the questions the community is actually asking, collected via Slido (and live during the session).
We'll discuss: what it takes to design for multicluster from day one; the state of key SIG Multicluster APIs and how they are being used in the real world; how GitOps pull models and existing tools fit into a multicluster workflow; what multicluster looks like running in production at scale; where vendors are present and where community needs to lead; as well as key topics raised by the live audience.
Attendees will leave with a clearer map of the multicluster landscape – what's stable, what's still being worked out, and where they can get involved.
Everyone in this room has been through it. Maybe you're in it right now.
Denial: "Our namespace isolation is fine." Anger: "Why does everything touch the kernel?" Bargaining: "What if we just add a Pod Security Policy?" Depression: "The CVE dropped on a Friday." Acceptance: You can't bolt on isolation; you have to build it in from the start.
This talk is a dual-perspective journey through the 5 stages of K8s security grief. Kavi brings the kernel-level technical reality: the shared surfaces, the subsystems nobody's watching, the failures that leave no CVE and no alert. Ann brings what folks actually say: the rationalizations, the moment the mental model breaks.
It's structured as a comedy, but the content is real. Every stage maps to genuine failure modes, real conversations, and the hard lessons that only land after you've done the grieving. Attendees will leave with a more honest picture of where Kubernetes isolation holds, where it doesn't, and what it actually takes to accept that.
A pod stuck in CreateContainerError for more than 19 hours. containerd rejecting kubelet's create calls repeatedly. But on checking the node, the containers were running fine. So what was kubelet still trying to recreate?
The trigger was a hypervisor reboot. But something subtler was at play; service account tokens were being issued in the future, container name reservations were piling up, and kubelet was trapped in an endless loop, recreating containers that were already running. The root cause? A clock skew during system bootstrap that quietly broke assumptions in the kubelet containerd contract.
This talk is a real world debugging journey from the cloud native stack down to the hardware system clock. We walk through each wrong assumption, the evidence that broke it, and the systematic approach that uncovered the issue. Attendees will come away with practical techniques for debugging CRI interactions, and navigating the boundary where Kubernetes meets infrastructure.
Processing thousands of documents reliably requires more than a Python script. You need a persistent service, quality checks before anything hits your index, and a pipeline that skips files that haven't changed.
In this tutorial you will build exactly that on a local Kubernetes cluster. We start with the Docling Operator managing Docling Serve as a deployment, wire it to Kubeflow Pipelines, and finish with documents flowing into OpenSearch and queryable by the end of the session.
Along the way: writing KFP components that call Docling Serve in-cluster, extracting quality metrics from the DoclingDocument JSON to filter bad conversions before indexing, and using KFP step caching so unchanged documents never reprocess.
You leave with a running cluster, working pipeline code, and a repo you can extend.
Modern workflows and AI agents can automatically recover from crashes, retries, and infrastructure failures. But in regulated and mission-critical environments, recovery isn't enough – organizations also need proof of what happened.
This session introduces Verifiable Execution: durable workflows with cryptographic attestations and tamper-evident execution histories. You'll learn how hash-chained workflow history, signed state transitions, and completion attestations can prove who executed work, what data was processed, and whether execution history was altered.
Using Dapr Workflows, we'll demonstrate how cloud-native applications and AI agents can move beyond fault tolerance to become systems that can cryptographically prove every action they take.
Lima (Linux Machines) was originally designed in 2021 as ""Linux-on-Mac"" for running containers.
Later it was redesigned to support non-Mac host operating systems and non-container workloads.
In 2026, it even gained support for non-Linux guest operating systems, including macOS and Windows.
Now Lima can run almost any guest OS on any host OS.
This is particularly useful for sandboxing AI agent workloads: even if an agent is deceived by malicious instructions found on the Internet (e.g., fake package installations), any potential damage is confined within the VM or limited to the files explicitly mounted from the host.
Meet us in this session to learn about the use cases and recent updates:
v2.0:
– Plugin infrastructure
– GPU acceleration
– MCP server
v2.1:
– macOS guests
– FreeBSD guests
– Improved CLI UX for running AI agents
v2.2:
– Windows guests
v2.3 (planned):
– Windows-native VM driver
Project website: https://lima-vm.io/
AI workloads push Kubernetes to its limits, often forcing teams to fight with GPU drivers and messy configurations. While the AI stack gets optimized, the host OS is frequently left as a fragile, manually patched bottleneck. Treating GPU nodes as mutable machines risks breaking your cluster with every update.
Flatcar Container Linux solves this by providing a minimal, immutable operating system built specifically for container workloads. In this official maintainer session, we briefly introduce the project's atomic update model before diving straight into a demo-heavy practical guide for AI infrastructure.
We will show how to manage NVIDIA drivers dynamically using system extensions (systemd-sysext) without breaking immutability. Finally, we will run a live demo showing how to bootstrap a local LLM environment on a self-hosted Flatcar instance.
Learn how to make your GPU infrastructure reproducible and boring.
AI is evolving fast, and agents are becoming the YAML engineers humans never wanted to be. The Flux project is embracing this shift by building the tools that solve the hard parts: stopping hallucinated fields and invented enum values, cutting the tokens costs and tool calls agents waste reconstructing APIs from stale training data, and giving them the actual meaning of every field instead of a plausible guess.
This talk shows how Flux is closing the agentic loop with static validation using the API server's own semantics, a schema catalog covering 100+ cloud native projects, a public MCP server that grounds every generated manifest in the real APIs, and the GitOps Agent Skills that turn assistants into repository auditors.
The talk ends with the Flux roadmap and how you can contribute to the project even if you use AI coding agents.
As Kubernetes embraces the fast, sandboxed Common Expression Language (CEL), Kyverno has undergone its most significant architectural evolution yet.
With the release of v1.17, the project stabilizes its pure CEL-based policy types (ValidatingPolicy, MutatingPolicy, GeneratingPolicy) and begins deprecating legacy engines. This transition offers platform teams massive performance gains, but moving a large enterprise ecosystem requires careful planning.
In this maintainer session, the Kyverno team breaks down the new features in the latest releases and the roadmap toward v1.20 legacy removal.
Joining them, TIAA shares their real-world migration journey. Attendees will learn how a large enterprise team seamlessly transitioned legacy cluster policies to the new CEL framework without disrupting developer workflows.
The session covers automated migration strategies, debugging modern CEL rules, and managing cloud native policy as code lifecycles beyond simple admission control.
Want to start contributing to Open Source but have no idea where to begin? It’s common to feel overwhelmed by the scale of these projects or unsure of how to take that first step. Whether you're looking to build your skills, give back to the community, or see how things work behind the scenes, finding your footing is the hardest part.
In this session, we’ll use Kubernetes as our roadmap to show you how a massive project actually functions. We’ll break down how the Kubernetes community is structured, how the people within it communicate, and, most importantly, how you can fit in. You’ll learn how to avoid common pitfalls like the 'good first issue trap' and get direct guidance from active contributors on where the best opportunities are hiding. Leave with a clear plan for your first move in the Kubernetes ecosystem.
HAMi is a CNCF Incubation project for heterogeneous GPU virtualization in Kubernetes, supporting 10+ GPU/NPU vendors.
Attendees set up HAMi on their laptop – with or without a GPU – and trace how the device plugin intercepts GPU requests, allocates vGPU resources, and integrates with the scheduler.
We follow the HAMi architecture and GPU virtualization model (project-hami.io/docs/core-concepts/gpu-virtualization), then explore the DRA integration path. Laptop with Go, kind, and a GitHub account required.
Is an agent a normal Kubernetes workload?
It needs compute, identity, networking, storage, and lifecycle management – so it looks like one. But an agent session is long-lived, stateful, bursty, and mostly idle.
And it can do anything a user can. That's what makes sandboxing critical.
The sandbox defines what the agent can see, touch, call, and persist: filesystem access, tool permissions, credentials, network policy, identity, and cleanup. Getting that boundary wrong is a security incident waiting to happen.
This talk covers the lessons we learned designing and implementing agent sandboxing for kagent – including NVIDIA's OpenShell, Kubernetes' agent-sandbox, and Google's substrate support – and what teams need to think about before trusting agents with code execution, credentials, network access, and production data.
OTel's metricstransform docs say it's "not suitable for aggregating metrics from multiple sources (e.g. multiple nodes)" — yet mitigating Istio cardinality requires exactly that. On our 10k-pod, 374-node mesh, Istio emits 2M+ active series even off-peak; PromQL projects an ~86% cut (2.06M→280K) once per-pod identity is collapsed.
We bypassed the constraint with a two-stage OTel architecture, in production:
– Stage 1 (DaemonSet): per-node OTTL stripping of pod-level identifiers to collapse series at the edge.
– Stage 2 (Gateway): cross-node aggregation bounding each service to the gateway replica count (≤50), not thousands of pods, using query-time sum() for autoscaling stability.
– Verification: a low-risk shadow pipeline validated against live traffic before cutover.
Leave with a portable blueprint and hard-won pitfalls: groupbyattrs no-ops, gateway re-enrichment undoing Stage 1, autoscaling fan-out, and silent collector regressions.
Most enterprise teams consume open source but don't prioritize earning community membership — leaving a huge opportunity on the table. This talk shares how an engineering leadership team at Capital One grew its Kubeflow Pipelines team from a single contributor to five recognized community members and two maintainers in just over a year.
The speakers describe the journey from passive dependency on an open source project to full ownership of a business-critical legacy version — reducing critical vulnerability patch turnaround from 30–60 days to 3–5 days, shipping upstream features, improving testing and resilience, and maintaining engineering joy — all while getting ahead of the EU Cyber Resilience Act.
Attendees will learn a practical framework for growing engineers from consumers to maintainers, how to make the business case for contribution time, metrics to track membership growth, and the connection between open source participation and upcoming regulatory requirements.
As organizations race to integrate long-running AI agents into their platform engineering and continuous integration (CI/CD) pipelines, they open a new frontier of security vulnerabilities: identity leakage. What happens when an autonomous AI agent is given elevated API privileges and the patience to bypass human latency?
This presentation dissects a real-world production incident where an autonomous coding agent bypassed branch protection rules and executed unauthorized force-merges directly into protected branches.
Attendees will walk away with concrete architectures for securing AI platform integrations, establishing absolute identity isolation, and enforcing the principle of least privilege for autonomous internal tools.
Once a GPU fleet outgrows a single Kubernetes cluster, optimizing its utilization becomes genuinely hard. At Snapchat, scarce GPUs serve a diverse mix of workloads: autoscaled latency-critical inference alongside batch training, Spark, and Ray across many clusters. Kubernetes' GPU ecosystem shines inside a cluster but stops at its edge: capacity sits idle in one cluster while work queues in another, or a low-priority job holds a GPU a critical service can't preempt on a separate cluster.
Leveraging Virtual Kubelet, we built a virtualization layer that allows us to schedule workloads wherever capacity actually exists. Alongside our extension of Kueue that enforces priority across cluster boundaries, we built a platform that optimizes scarce compute across any number of Kubernetes clusters, regions, and even clouds.
The result was nearly 2x average GPU utilization and $XX million saved annually, while availability and latency for critical workloads improved.
OSS communities are sounding the alarm that AI-generated contributions are flooding their projects. Much like a DDoS attack overwhelms a server with excessive requests until it goes offline, this flood can overwhelm maintainers and push projects to shut their doors. We call it AI-DDoS. Some projects are already responding defensively: closing bug bounties, moving to invitation-only models, or rejecting unsolicited pull requests. These responses may address the immediate symptoms, yet they risk closing the door to new contributors, upending their ecosystem.
We analyzed over 2 million PRs and issues from 350,000+ contributors across 294 widely used repositories. Our findings show that AI-DDoS is a structural, ecosystem-wide pressure on the OSS contribution pipeline. From this evidence, we developed 11 practitioner-driven strategies that projects can use to manage AI-DDoS while balancing openness, maintainer capacity, and their stance toward AI-assisted contributions.
GitOps assumes cluster state converges to Git.But, what happens when the desired state is wrong?
This session examines a real incident where a routine Flux reconciliation turned a small git change into a cascading prune across a Kubernetes management platform managing 39 workload clusters. A missing Kustomization triggered inventory deletion, causing Flux controller to delete itself during reconciliation. Resulting in orphaned resources, stuck finalizers, and loss of critical Git repository storage.
Using timelines and incident evidence, we trace the failure chain and explore Flux inventory tracking, pruning, finalizers, StatefulSet PVC ownership, CSI reclaim policies, and deletion ordering.
We also cover recovery, restoring resources, rebuilding the GitOps control plane, and safely resuming reconciliation. Attendees will learn practical ways to reduce GitOps blast radius, protect stateful infrastructure, and build platforms that remain recoverable when automation fails.
Eventually, every AI platform team asks: Why did my inference bill just explode?
Avoid being caught off guard by using KServe and Kubernetes Gateway API for AI traffic control.
KServe provides a powerful foundation for model serving, with standards-based inference APIs, multi-framework support, and production-grade autoscaling. Layer on a Kubernetes Gateway API-based gateway architecture with emerging projects like agentgateway, and you get an extensible approach to controlling spend.
Beyond standard model consumption features, you can add token-based rate limiting, policy enforcement for guardrails, inference traffic observability, and resilient routing strategies tailored to AI workloads. This session will demo how to solve challenges such as cost control, security, operational visibility, and model failover.
Leave this session with an understanding of how KServe + Gateway API–based AI gateways together support scalable, governable, and resilient AI platforms.
Most clusters boot from a generic OS image, then spend precious scale-out time fetching drivers, pulling large containers, tuning kernels, and installing dependencies. For GPU and other specialized pools, that work repeats on every node replacement and rolling update.
This talk presents a proposed CRD-style workflow for treating worker node images as Kubernetes provisioning artifacts: declared from workload intent, validated before rollout, versioned, and handed off to node lifecycle systems. Drawing on Kubernetes image-builder, Cluster API, Cluster Autoscaler, and related projects, we show how image automation can move risky setup work out of node bootstrap.
The demo compares one pending GPU workload on two paths: a standard cloud-provider supplied image and a custom prevalidated worker image. We break down timings from pending pod to first Ready GPU capacity across provisioning, boot, kubelet registration, device plugin readiness, and first workload readiness.
Most of the AI conversation in Kubernetes is about running LLMs on top of it. This session asks the opposite. Can an LLM run Kubernetes?
Operations teams now face a choice. They are being asked to bring AI into incident response while keeping cluster state inside the perimeter. A capable open-source LLM now runs on a single workstation, so the option exists. The question is whether the option is real.
This session answers with a published benchmark of 10 Kubernetes incident scenarios (crash loops, OOM kills, HPA misconfigurations), each scored by an Ops_Score where quality and safety gate the result and efficiency only adjusts it. The same benchmark already ran on today's leading agents, giving open-source families such as Gemma, Llama, and Qwen a bar to measure against.
Rather than crown a winner, the session maps the gap to that bar. The results show where open-source models hold up, where they fall short, and where they take unsafe actions an on-call engineer must override.
Over the past year, Kubernetes has expanded support for high-volume data workloads through Jobs, while the Workload APIs (StatefulSet, ReplicaSet, PDBs, etc.) have grown more mature, resilient, and full-featured. The SIG-Apps team has been firing on all cylinders and we’re just getting started.
In this session, the SIG Apps leads will provide a breakdown of the major enhancements implemented over the past year. They will look at potential directions for even smarter workload automation. Finally, they will spotlight critical features that need your feedback to cross the finish line.
The session will conclude with a live Q&A and open discussion. Whether you want to voice your feature requests or learn how to make your very first contribution, come help drive the evolution of Kubernetes orchestration!
KubeVirt has made an enormous leap over the past year as we advance toward CNCF graduation. We have shifted to an extended pluggability model via VEP-190, introducing a new Plugin CRD and hooks that enable powerful community extensions. The Hypervisor Abstraction Layer is opening the door to multiple VMMs beyond QEMU/KVM.
We are aggressively targeting next-generation workloads. A new enhancement brings first-class upstream support for the latest NVIDIA hardware, powering AI/ML workloads natively. We are also expanding to the edge with standalone VM-pod patterns.
Governance has matured dramatically, too. A surge in enhancement proposals is now supported by our custom AI agent that helps manage the VEP lifecycle.
Join us to discover the roadmap, modern capabilities, and how to contribute!
Refuel those propeller beanie-caps: ""it depends"" is about to become the Kubernetes API server's favorite answer.
First, Conditional Authorization. Authorizers used to answer only Allow, Deny, or NoOpinion. Now they can say ""Allow if…"" or ""Deny if…"", escaping all-or-nothing semantics. Via partial evaluation you can express ""create gateways only when class is 'test-gateway'"" or ""a controller may touch only its own finalizer"". We'll trace conditions from authorization into admission, show them in kubectl auth can-i, and explain how CEL lets the API server evaluate them in-process.
Next, a nastier gap: anything on the service network can reach an admission webhook unauthenticated. ""IngressNightmare"" (CVE-2025-1974) proved it, turning that exposure into RCE and cluster-wide secret theft. Learn how the API server now proves its identity to webhooks by default with API group scoped webhook bound service account tokens.
Infrastructure is having a year. AI workloads are dragging cloud native back down the stack, into GPUs, bare metal, and datacenter realities most of the ecosystem spent a decade abstracting away with commodity hardware. The tooling underneath is straining, and the operators feeling it need guidance on how to navigate this new reality. That's what the Infrastructure Technical Advisory Group is for.
In this panel, the TAG's co-chairs and tech leads dig into where cloud native infrastructure is heading: what operators are telling us about where the stack falls short, how the TAG turns that feedback into community work and initiatives, and where projects and end users can plug in. Bring your hardest infrastructure questions. The gaps you're hitting in production are exactly the input we're looking for, and this session is how they become the community's next priorities.
As cloud-native architectures evolve, "reliability" means managing everything from microservice cascades to GenAI token costs and energy efficiency. The CNCF Operational Resilience TAG is at the center of this evolution, defining the standards for how the industry builds and operates sustainable, resilient systems.
Join this session with the TAG leads to get a first look at our latest initiatives and future roadmap. We will share what we are up to and how you can join us and help in championing the initiatives.
Welcome, cloud native community! We’re excited to kick off our time in Salt Lake City with you. Join the community and our sponsors from 6:00-7:30PM in the Solutions Showcase for an incredible gathering of local food favorites, beverages, games, and activities.
Explore the sponsor booths to learn more about the latest technologies, browse special offers, job posts, and much more.
In order to facilitate networking and business relationships at the event, you may choose to visit a third party’s booth or access sponsored content. You are never required to visit third party booths or to access sponsored content. When visiting a booth or participating in sponsored activities, the third party will receive some of your registration data. This data includes your first name, last name, title, company, address, email, standard demographics questions (i.e. job function, industry), and details about the sponsored content or resources you interacted with. If you choose to interact with a booth or access sponsored content, you are explicitly consenting to receipt and use of such data by the third-party recipients, which will be subject to their own privacy policies.
In Amsterdam, we explored what cloud native could do with AI agents, MCP, and a drone controlled through natural language. Now it's time for Drone 2.
What happens when AI agents become a natural part of the cloud native ecosystem? Imagine agents that can debug production issues, analyze token spending, automate routine operations, and even make us laugh, all while interacting with Kubernetes, cloud native tools, and MCP servers through natural language, with humans firmly in control.
This keynote explores what's possible when AI agents become trusted collaborators in cloud native. Through live demonstrations with Drone 2, we'll show how natural language and AI agents can transform the way we build, operate, and experience cloud native systems, pushing the boundaries of what's possible while keeping humans in control.
Cloud native teams spent a decade building traffic infrastructure they control: open source, decentralized, observable, and running on Kubernetes. Then AI arrived. Inference, the most sensitive layer in the stack, became a call to a vendor API. Your cluster is sovereign. Your AI isn't.
You already run what's needed to reclaim that layer. An inference request is plain HTTP, with the model name in the JSON body. The ingress and Gateway API patterns in the cluster today can route and balance that traffic, and apply policy at the boundary. One data plane, instead of a fragmented traffic stack.
This session covers consistent-hashing across GPU backends and model-aware routing from request inspection, with inference staying inside the cluster. Everything runs on HAProxy, the open source, model-agnostic load balancer and ingress controller built for high throughput and low latency. You keep sovereign control of AI traffic with no performance penalty.
Kubernetes is the foundation of the next generation of AI. Moving from experimentation to a production requires infrastructure designed for scale, performance and simplicity. In this keynote, Oracle Cloud Infrastructure (OCI) will explore how OCI Kubernetes Engine (OKE) enables organization to build and run demanding cloud-native and AI/ML workloads on an open, enterprise-ready platform.
Join Jeff Hoffman, VP of Container Services at Oracle, to hear how customers are using OKE to bring AI initiatives into production. Learn how OCI’s Kubernetes platform, high-performance infrastructure, and open-source ecosystem help teams innovate faster while maintaining the control and reliability their critical workloads require.
In order to facilitate networking and business relationships at the event, you may choose to visit a third party’s booth or access sponsored content. You are never required to visit third party booths or to access sponsored content. When visiting a booth or participating in sponsored activities, the third party will receive some of your registration data. This data includes your first name, last name, title, company, address, email, standard demographics questions (i.e. job function, industry), and details about the sponsored content or resources you interacted with. If you choose to interact with a booth or access sponsored content, you are explicitly consenting to receipt and use of such data by the third-party recipients, which will be subject to their own privacy policies.
Nearly 25 years after *Moneyball*, baseball analytics more resembles a high-performing tech org than a few analysts managing their own ad-hoc database. This level of sophistication introduces a major challenge: closing the gap between data scientists solving complex baseball problems and the resources they need to train and use their models, especially when there are a lot of existing critical workflows that run in disparate environments, or even on individual laptops.
This session explores how a four-person platform team used a cloud-native stack—including Docker, Kubernetes, Helm, and Knative—to deploy statistical products that translate to on-field results. By meeting data scientists where they are and providing tools that grant freedom with minimal engineering overhead, the team manages the challenges of introducing tech into a specialized domain. It will also cover the wins that come with investing in a good platform, and the impact on products for coaches and decision makers
In past Kubecons, we have spoken about AI assisted explainers as a strategy to explain telemetry pillars such as traces and logs. Over the last year, we at eBay have worked towards furthering triage by explaining telemetry and reasoning on top of it. To be able to make troubleshooting with AI more promising, we have indexed on dashboards.
Why dashboards? Dashboards contain relevant signals grouped together, they have descriptions and they are battle tested across debugging incidents and day to day operational analysis. Having the right signals to analyze is half the battle. The rest?
- subagents that handle each telemetry pillar
- careful engineering of what is done programmatically and what is "reasoned" by LLMs
- tying multiple findings into hypothesis and root causes
This talk, describes how we built AI assisted dashboard analyzers on top of data sitting in Prometheus and ClickHouse. Does this mean that AI can do it all? Not yet, but we are edging closer and closer.
At Isomorphic Labs, we run tens of thousands of concurrent ML workloads per cluster for drug discovery. Using Kueue for scheduling, we found that mixing bursty single-pod inference alongside heavy multi-node distributed training pushed our control plane over the edge.
This talk breaks down a production outage where pending workloads triggered Kube-API webhook timeouts, causing a cascading control plane failure. We detail how we stabilised the cluster: mitigating head-of-line blocking by decoupling queues and tuning Kueue's Preemption Fair Sharing. We also explain the "cardinality problem" of integrating KEDA with Kueue, and how to build scheduler signal limiters to protect the API server from hyper-scale event storms.
Attendees will learn actionable architectural guardrails to navigate volatile scaling, manage extreme concurrency, and safely colocate massive batch inference with distributed training.
LLM inference routing is a cloud-native systems problem: traffic must be steered across Kubernetes clusters, accelerator-backed serving engines, capacity limits, and geography cost while keeping the request path fast and observable.
This session presents a generalized Inference Load Balancer (ILB) proxy/controller architecture for production model serving. A low-latency proxy applies routing weights and request-path signals, while a Kubernetes-operated controller computes and distributes source-cluster-to-engine weights from recent demand, engine capacity/performance profiles, replica state, and geography cost.
We will cover the proxy/controller split, policy distribution, multi-cluster operation, load shedding, policy observability, and debugging patterns for adaptive inference routing. Attendees will leave with reusable Kubernetes patterns, not OpenAI-specific internals, product claims, or non-public metrics.
A compute node in modern cluster contains not only CPU and memory, but also GPUs, NICs and other specialized hardware. The overall performance of the workload nowadays comes not by the characteristics of the single device, but from combination of components and deep understanding on how they interact.
This talk focused on zooming in into hardware topology, explaining what inside physical CPU package is, memory-controller behaviors and locality boundaries, anatomy of I/O hubs and how data flows between components and between devices. This will help to understand deeper about where potential bottlenecks are located for your workload, what are the costs for bandwidth and latencies for each data path and how to potentially mitigate them.
The talk will cover new features and improvements in DRA, NRI and OCI interfaces that are focused on improving flexibility and manageability for devices, node allocatable resources, standard and derived attributes and low-level tuning parameters.
Since Dynamic Resource Allocation (DRA) graduated to GA in Kubernetes 1.34, the DRA APIs have seen a number of enhancements to expand the types of hardware that can be expressed by DRA and how they can be requested by workloads. While these features are all described in the Kubernetes documentation, working examples are crucial for users to understand them and experiment further.
The Kubernetes-sponsored, vendor-neutral DRA Example Driver aims to implement a wide variety of DRA features compatible with any cluster to serve both device vendors implementing features in their own drivers and DRA users looking to request DRA resources for their workloads.
In this talk, the maintainers of the DRA Example Driver will demonstrate how both DRA driver authors and DRA users can learn about implementing and utilizing the latest DRA features with tested examples showing common use cases and best practices.
Agentic workloads don't just add complexity, they change its nature. An agent that reasons, plans, and calls tools across services creates execution paths no traditional trace model was built to capture. In regulated domains like fintech, healthcare, or telco, "we don't know what the agent did" isn't acceptable.
The cloud-native community can solve this, because we've done it before: we built golden paths, guardrails, and observability stacks. Now we extend those patterns to agentic systems.
We'll show how platform teams can build governed, observable foundations for agentic workloads: agent runtimes, context layers, tool gateways, policy/eval pipelines, and multi-agent systems over shared context. Understanding protocols like MCP, A2A and skills matter too. We'll use OpenTelemetry to instrument reasoning chains, tool calls, and handoffs: if you can't trace it, you can't fix it.
Are you a platform engineer or a developer in a heavily regulated industry? This talk is for you.
Step one: get the model running. Step two: realize that was the easy part.
Once you're in production and start to scale, you're suddenly thinking about cost, efficiency, and why adding a second replica didn't help as much as you hoped. Requests route naively, missing the cache. GPU time adds up fast, and your costs reflect it. Whether you're running a multi-GPU model or scaling out to handle more traffic, you're probably leaving performance and cost savings on the table.
llm-d, a CNCF sandbox project, was built to solve these problems with efficient, cost-aware LLM inference on Kubernetes through features like disaggregated prefill/decode, prefix cache-aware routing, and flow control.
From an engineer who's spent years in inference infrastructure and is now digging into llm-d, we'll walk through each feature practically: what it is, why it matters, and when an optimization is right for your setup. Expect real explanations, honest tradeoffs, and as always, fun drawings.
An MCP (Model Context Protocol) server bridges AI apps to external tools and data sources. The MCP spec recommends servers to use a general-purpose access control engine. The number of MCP servers has exploded, but their access control models and maturity vary widely, making it difficult for organizations to enforce consistent policies across their environments.
The CNCF project Cedar Policy offers a unified language and engine for attribute-, relation- and role-based access control policies, along with unique policy querying and analysis features, perfect for use in MCP servers. ToolHive is an open source project that allows you to run and manage MCP servers in Kubernetes, enforcing identity and access policy per request via Cedar.
After this talk, the audience knows how MCP access control works, how ToolHive uses Cedar (demo included!), and how Cedar Analysis answers “who can access this resource?”, “is it possible for an AI agent to send private data to external sinks?” and more
As OpenTelemetry adoption scales and the project extends its breadth, end users find themselves navigating a large configuration surface: SDKs, Collector deployments, data pipelines, instrumentation libraries, semantic conventions… They can all create accidental complexity when not considered as part of a cohesive strategy, resulting in disjointed telemetry and a lack of context, the opposite of OTel's vision.
In this session, we'll introduce OTel Blueprints, a community-driven initiative providing prescriptive guidance for deploying OTel across its many components, grounded in real-world reference implementations. We will walk through newly published blueprints tailored for common enterprise architectures, exploring actionable design patterns for Kubernetes observability, non-Kubernetes infrastructure, and building centralized telemetry platforms that can scale.
Attendees will learn how to apply these blueprints to build a scalable, consolidated observability strategy.
OTTL is already everywhere in advanced OpenTelemetry Collector pipelines. It is the tool teams use to filter noisy data, transform telemetry in flight, route signals, enforce conventions, and shape pipelines without writing custom code. But using OTTL effectively requires a practical understanding of how to apply it to real pipeline problems.
This hands-on tutorial is framed around exactly those problems. Through live Collector configs, running examples, and tools like ottl.run, attendees will learn how to write maintainable statements, debug unexpected behavior, and solve common pipeline challenges. The session also covers advanced patterns, practical OTTL tips and hacks, extending OTTL in a custom Collector distribution, and newer capabilities to watch.
Led by OpenTelemetry and OTTL maintainers, this tutorial gives beginners, intermediate users, and advanced practitioners practical techniques they can apply immediately as well as insight into the design decisions shaping OTTL.
Telemetry collectors have become critical infrastructure. Running on nearly every node in a Kubernetes cluster, even modest efficiency improvements can translate into substantial infrastructure savings while increasing the volume of telemetry a platform can process.
This session explores the engineering behind Fluent Bit’s latest performance work, from reducing CPU usage in OTLP Protobuf generation and OTLP JSON serialization to optimizing routing, parsers, buffering, memory allocation, and protocol handling. Rather than presenting benchmark numbers alone, we’ll examine how profiling guided these optimizations, the tradeoffs between throughput, latency, memory footprint, and protocol fidelity, and why some seemingly obvious optimizations failed to deliver.
Attendees will leave with practical techniques for designing, profiling, and optimizing high-throughput telemetry systems that apply well beyond Fluent Bit.
Over the past year, OpenCost has continued to evolve beyond Kubernetes cost allocation into a richer platform for understanding infrastructure spend at scale. In this maintainer session, we'll cover what's new, why it matters, and where the project is headed next.
After a brief overview of OpenCost's architecture, we'll showcase the redesigned UI with improved reporting and search, then dive into Bingen, OpenCost's next-generation storage layer that reduces Prometheus data requirements while enabling fast, efficient cost queries. We'll also introduce KubeModel, explore its evolving resource model including GPU support, and preview the roadmap for inference costs, RBAC, Prometheus-independent deployments, and future query capabilities.
Whether you're a long-time OpenCost user or just getting started, you'll leave with a clear view of the project's future.
As a CNCF Graduated project, Harbor enables organizations to securely store, scan, sign, replicate, and distribute container images and OCI artifacts across Kubernetes environments and software supply chains.
In this Maintainer Track session, Harbor maintainers will share the latest project updates, highlighting enhancements driven by community feedback and production experience. We'll cover recent improvements in security, registry operations, performance, scalability, day-2 operations, and ecosystem integration, along with architectural changes relevant to operators, platform engineers, and contributors.
We'll also discuss current community initiatives, upcoming roadmap items, and opportunities to contribute. Whether you run Harbor in production, integrate it into your platform, or contribute upstream, you'll leave with a clear understanding of Harbor's current direction and what's next for the project.
We'll first introduce you to Vitess and the Vitess Operator and how it allows you to horizontally scale your MySQL databases. We will then walk through a live demo with a Vitess cluster running in k8s under load and demonstrate how we can shard that database in order to reduce the load and improve query performance.
With the opening of 3.8 development and new momentum on etcd's multiple subprojects including raft, bbolt, operator, documentation, and benchmarking, there are many new opportunities to contribute to etcd. Come ready to work with your laptop and favorite code (or docs) editor, and five etcd approvers will be on hand to help you get started working on one of etcd's many open issues and roadmap priorities.
This is a working session. Attendees will be given a very brief orientation and then set to doing supervised work on open project issues. Do not attend without a laptop.
Functions are the secret sauce to extend Crossplane and solve problems the core framework can't on its own. The function ecosystem is rich, ranging from high-level language support like Python and TypeScript to functions built for specific problems like querying live cloud state. As adoption climbs, that ecosystem needs more people building functions, and every maintainer starts by writing their first one.
Writing a function can look harder than it really is, but this ContribFest session will walk you through the whole thing. We'll tour the ecosystem, then build a new real function together that the community will actually want to use. With maintainers beside you, you'll go from an empty repo to a working, PR-ready function in a single session.
But that’s just the first step of the journey. You'll leave with the skills to keep going, pick up a curated issue, make your next contribution, and grow into a Crossplane function maintainer yourself. Let's build something together!
Building and running Kubernetes in production is a challenge; doing so in AWS GovCloud while hunting for a FedRAMP Moderate Authorization to Operate (ATO) is a multi-year marathon. In this talk, we trace our 3-year journey from an initial design to a fully authorized, compliant platform, pulling back the curtain on what it actually takes to survive federal audits without destroying developer velocity.
We'll cover our initial design choices and the lessons learned when cloud-native assumptions collided with rigid federal regulations. Next, we will discuss the changes required to scale environment access to more engineers as we got our first customers.
Finally, we will deep-dive into the two most critical operational hurdles of our journey:
– why standard CI/CD tools fell short, and how we engineered a dedicated deployment system to meet strict regulatory isolation and access controls.
– how we tackled the required patching and reporting mandates of FedRAMP for open-source software.
At LinkedIn, some of our most critical Kubernetes platform services deploy multiple times a week, including infrastructure components that affect every pod and data-processing job in the cluster. Traditional canary tools based on traffic splitting don’t work well for these services because there is no user-facing L7 traffic to gradually shift.
In this talk, we’ll share how we built a Kubernetes-native canary pipeline to safely validate cluster-critical services before production rollout. By combining admission controls, isolated canary environments, GitOps workflows, and automated end-to-end validation, we reduced deployment verification from days of manual testing to under an hour.
Attendees will learn practical patterns for deploying infrastructure services safely at scale when traditional service-mesh canary approaches are not enough.
At 50,000 nodes, component failure is a statistical certainty—but total control-plane collapse shouldn't be. When Anthropic and Google pushed GKE to these limits on production AI workloads, the real threat wasn't the initial faults; it was the catastrophic loss of graceful degradation. Default upstream mechanics frequently turned localized anomalies into cascading, fleet-wide stalls.
This talk explores six joint war stories where Kubernetes' built-in self-healing broke down: from a CNI startup probe that couldn't outlast its own informer sync, to an APF cost-misestimation that triggered master OOMs.
We will detail the exact failure modes, the joint debugging, and the upstream PRs that enforced blast-radius isolation. Finally, we’ll examine how these lessons drive forward-looking community work—like queue isolation and stateless controllers—ensuring Kubernetes handles massive scale gracefully without breaking.
30% of merged PRs at Podium require no human author. They are opened and closed by autonomous coding agents.
Podium built a Kubernetes-native agent orchestrator: each agent runs as an isolated pod, picks up a task, writes code, and tests its own changes against a real environment before opening a PR. Kubernetes handles the scheduling and isolation, letting hundreds of agents run concurrently without stepping on each other or the production system.
The key insight: agent reliability is not about the model. It is about the feedback loop. Agents that can test against real systems catch their own bugs before a human ever sees the PR.
Drew Bowman, Sr Eng Manager for SDLC at Podium, shares the architecture, the task taxonomy (what agents handle reliably vs. where they still fail), and what changes in an engineering org when 3 in 10 PRs no longer needs a human author.
x86 and ARM carry decades of legacy cruft: The IME, firmware blobs, outdated downstream kernels and on ARM an absence of standards that make using mainline Linux on anything other than pricey server hardware almost impossible. RISC-V, a new open ISA, is finally a real alternative in 2026: RVA23 chips are catching up to Intel, AMD, and ARM on performance, virtualization works, and RISC-V vendors mainline SoC support before shipping; even modern Fedora runs out of the box.
This changes what metal clusters can look like: You own more of your stack, with open boot firmware, open chip designs you could fab yourself, and no license or patent encumbrances. Combining RISC-V with immutable Linux (via `bootc`/`composefs` or `systemd-sysupdate`) and k8s creates near zero-maintenance, legacy-free bare-metal.
The session includes a live demo where a Milk-V Jupiter 1 (RVA22) and Jupiter 2 (RVA23) Kubernetes cluster migrates a workload with CRIU between the boards, then to a managed RISC-V VPS.
Kubernetes incidents rarely stay politely inside the cluster. A slow API may look like an application issue, only to trace back to noisy neighbors, failing storage, cloud limits, or some other infrastructure gremlin that waited until 2am. HolmesGPT, a CNCF Sandbox open-source SRE agent, expands AI-assisted troubleshooting beyond cluster scanning by investigating production problems across Kubernetes, cloud services, databases, and the infrastructure beneath them.
In this demo-driven session, we will use HolmesGPT to investigate failure scenarios that connect application behavior to underlying compute, storage, and cloud-provider signals. We will show how an AI-powered SRE agent can gather context, reason across multiple operational layers, and help teams move from "the pod is broken" to "here is the likely root cause and why." Along the way, we will contrast this broader incident-investigation model with CNCF diagnostic approaches, such as k8sgpt, and highlight where each fits.
MCP makes it easy to connect AI agents to internal tools. It does not, however, tell you how to enforce least-privilege access to each of them consistently. Without a standardized approach, individual teams are forced to implement their own authorization layers. The result is not just security vulnerabilities, but platform sprawl.
Enforcing least-privilege access for AI agents raises three challenges: verifying which user an agent is acting on behalf of, proving the agent's own identity as a workload, and controlling tool access based on both contexts. This session addresses each in turn by propagating user identity across trust boundaries via OAuth 2.0 Token Exchange, attesting agents as verified workloads via SPIFFE/SPIRE, and enforcing identity-aware access decisions via ABAC. It then packages these approaches as a reusable Golden Path that platform teams can adopt.
Leave with a security model for least-privilege AI agent access and a Golden Path ready to ship as a platform.
When a Kubernetes cluster needs another inference replica, the clock starts. Image pull, model download, CUDA initialization, tensor loading. For large models, that takes minutes. The alternative, pre-warming GPU pods, trades latency for cost. Neither scales.
This session presents a better path. The speakers show how they extended Kata Containers and Cloud Hypervisor to snapshot an entire pod VM, including GPU state, after model initialization. New pods restore in seconds. Because Kata runs each pod in a VM, the snapshot captures everything at once, a clean boundary that process-level checkpointing cannot match.
The talk covers snapshot-capable Kata architecture, GPU state capture, and tradeoffs in snapshot size and restore latency. Attendees will see startup times move from minutes to seconds and learn how copy-on-write keeps density high. The session also explores how VM isolation opens a path to confidential containers, keeping model weights invisible to infrastructure operators.
We've shifted left, generated SBOMs, and signed artifacts, but we’re drowning in more vulnerability data than ever before. AI is increasing the speed and sophistication of attacks, compounding the crisis of CVE fatigue while teams struggle to patch "critical" vulnerabilities not reachable at runtime, and navigate from "compliance theater" to actual risk reduction.
Is the panic induced by increasingly capable models justified? Does this fundamentally change the security landscape, or is it just marketing hype? And is the current "scan and patch" model broken? Welcome to the Vulnpocalypse!
This panel debates the next phase of cloud native and open source security. We‘ll explore the transition from static analysis to runtime insight, the role of reachability analysis, and the latest open source tooling to help build a resilient security posture.
Attendees will leave with insight & ideas for cutting through the CVE noise, and levelling the playing field against AI-accelerated attacks.
Think sampling is just about reducing your vendor bill? Think again! The truth is, the cost of telemetry starts way before it reaches your backend. It starts inside the application process itself, and increases the farther it goes before it’s sampled in your OTel Collector or backend.
In this session, we reframe sampling as a performance and cost strategy that spans the entire telemetry pipeline. We start with application-level overhead and see how different language SDKs can impact performance, and why sometimes the answer isn’t better sampling; it’s not instrumenting at all.
We’ll use the OTel demo app, Astronomy Shop, as our live environment to explore the sampling decision hierarchy, including declarative config to drop known-noisy signals, when to use head- or tail-based sampling, probabilistic sampling, and more.
Join our talk to connect business outcomes to concrete instrumentation and configuration decisions, and for an honest take on where sampling helps and where it doesn’t.
Over the past decade, Kubeflow has helped define how AI and machine learning run on Kubernetes, evolving from data engineering, distributed training, and model serving to supporting today's generative AI and agentic workloads. That journey has reached a major milestone with CNCF Graduation.
In this session, Kubeflow maintainers will reflect on the project's evolution, the architectural and community decisions that shaped it, and the lessons learned in building a sustainable open source platform for enterprise-grade AI. The speakers will discuss how Kubeflow adapted to each wave of AI innovation while remaining Kubernetes-native, what CNCF Graduation means for governance, sustainability, and enterprise adoption, and what's next for the project. Attendees will gain a retrospective of the project's evolution, an overview of its current state, and insights into the future of AI on Kubernetes.
OpenFGA helps teams move beyond role checks to relationship-based authorization for cloud native apps, APIs, and AI agents. Prior OpenFGA talks have introduced the project, shown SaaS and gateway adoption, and sketched the agent delegation model. This session goes one layer deeper.
We will walk through OpenFGA as a working project: authorization models, tuples, checks, local testing, CLI/SDK workflows, and production concerns. Then we will show a concrete agent-tool authorization flow: a user delegates a scoped task to an agent, a gateway extracts user/agent/task/tool/resource context, OpenFGA decides before the tool runs, and an audit trace records the request, decision, and result.
Attendees will leave knowing when OpenFGA fits, how to try it, and where to contribute across server, CLI, SDKs, docs, tests, and agent authorization examples.
Cilium continues to evolve across every layer of the cloud native networking stack, from new Kubernetes APIs at the top to innovations in the Linux kernel underneath.
Join us for a technical tour of the latest work happening across the Cilium project. We'll cover recent developments including Gateway API features such as TCPRoute and ListenerSet, new eBPF datapath plugins that make kernel networking more extensible, and the architectural work shaping the next generation of cloud native networking, observability, and security. You'll also hear from an end user on why they chose Cilium.
We'll close with community activities around the Cilium project and learning how to get involved as the project moves into its second decade!
This session covers the latest updates in the Kubernetes Node subsystem. Come to learn how Nodes become more flexible and Pods more dynamic. SIG Node owns components like Kubelet, Container Runtime Interface (CRI), Node API. SIG Node is responsible for Pod lifecycle from allocation to teardown, shared resource management, topology alignment and device access via plugins. We work with container runtimes, kernels, networking, storage, and more; anything between the pod and the underlying hardware that runs them is in SIG Node’s purview!
The session will be interesting for end users, seasoned contributors, and people seeking to get involved. Attendees will leave the session with a better understanding of the latest developments like Dynamic Pods, DRA, PSI, pod level resources, in-place resize and restart and more, as well as understand the roadmap in these days of AI/ML and other workloads adoption.
As an AI Architect, I'm on 20+ calls per week with customers implementing AI in production. These are customers ranging from the largest telecom providers in the world to the largest banks worldwide, healthcare, retail, and organizations across the globe. The calls range from discovery to mapping out architecture to AI implementation and everything in between.
In this session, I'm going to go over the top 5 things that every organization asks for when implementing anything Agentic, regardless of the sector they're in.
The topics will include:
1. Agent and agentic network security around AI Gateways, MCP, LLMs, and Agent Skills.
2. Agent quality output (evals and deterministic behavior).
3. Authentication and authorization with standard OIDC/iDP practices along with OAuth, ReBAC and ABAC.
3. Multi-agent architectures.
5. Cost and performance optimization
The goal of this session is for everyone to leave knowing exactly what organizations are looking to implement with Agentic today.
As AI agents increasingly access sensitive data at scale, the agent harness, including prompt loops, tools, permission logic is becoming a complex, often third-party layer that should not be treated as the sole security boundary. Sandboxes may also be provider-managed, which creates an accountability gap similar to cloud infrastructure’s shared-responsibility model: platform teams remain responsible for data protection, yet current agent infrastructures often lack independent observability and enforcement.
We introduce a runtime data security solution positioned around existing harnesses and sandboxes. Using eBPF for zero-instrumentation OS/network observability, continuous scanning of misconfigured cloud storage, data anonymization, and anomaly detection, it offers automated, independent visibility and control of agent activity. This transparently provides enterprise-grade data protection, compatible even with opaque runtimes, effectively closing the critical agent security gap.
Kubernetes HPA scales on CPU and memory — but production incidents start in logs: connection-pool exhaustion, timeout storms, retry cascades. By the time the metric moves, the on-call is already paged.
This session walks through building an Agentic Autoscaler — a Go operator that reads the same signals SREs read (Prometheus metrics, Loki or Kafka log streams), routes them through a pluggable reasoning step (rule-based by default, optional Claude / OpenAI / Ollama), and emits an explainable scale decision that nests safely inside any existing HPA.
The case study is an API gateway under a timeout storm. HPA reacts late on CPU. The operator detects connection-pool-exhausted across three consecutive log windows and scales preemptively — then annotates Grafana: "Scaled due to timeout error rate exceeding threshold across 3 consecutive log windows."
You'll leave with concrete patterns for log+metric fusion, AI inside a control loop, and natural-language explainability for SREs.
KubeVirt brought VMs to Kubernetes, but it’s hard dependency on QEMU/KVM limits support for modern heterogeneous infrastructure. As organizations adopt memory-safe VMMs like Cloud Hypervisor and specialized stacks such as Microsoft’s MSHV for novel scenarios, KubeVirt must evolve beyond a single hypervisor. This session presents the Hypervisor Abstraction Layer (VEP-97, shipped in v1.8), that decouples the core control plane from the virtualization backend. We will share the technical journey: introducing extension points for spec validation, default, conversion between KubeVirt and backend virtualization API and runtime tuning of VMs. We will cover the challenges of standardizing VM lifecycle while preserving Kubernetes semantics, and the design enabling future out-of-tree plugins. We will show the architecture that allows bringing your own VMM or hypervisor to KubeVirt without forking or modifying the core. This is the first public deep-dive of MSHV (Hyper-V) integration in KubeVirt.
Open-source projects increasingly face contributions from AI. The challenge is not productivity; it is validity. How do maintainers trust that AI-generated PRs solve the problem it claims to? Standard review catches syntax and style, but agent code introduces new failure modes: confident incorrectness, fabricated claims, and broken implementations that appear structurally sound.
We present a methodology for AI contributions grounded in structured experimentation, with examples from llm-d, a CNCF sandbox project. Agents contribute via hypothesis-driven workflows with built-in evidence, while project invariants become machine-checkable contracts. On the review side, convergence protocols run independent perspectives in parallel and trust-tier classification tells maintainers which outputs to verify, catching bugs standard review misses.
Attendees leave with patterns for structuring experiments, codifying invariants, and reviewing agent work with calibrated trust for any CNCF project.
Every Kubernetes network policy engine runs in the kernel. RDMA bypasses the kernel entirely. On multi-tenant AI clusters sharing GPU nodes, no software layer can enforce per-pod isolation on RDMA traffic: NetworkPolicy, Cilium, Calico, all blind.
This talk presents a production-validated pattern that closes the gap. A namespace-scoped CRD with NetworkPolicy-shape selectors makes each pod the subject of a tenant policy. A controller resolves selectors at admission and programs the NIC's hardware access control before the first RDMA verb fires. Same-tenant pods communicate freely. Cross-tenant traffic is dropped at the NIC
We walk through the full flow live: apply the CRD, start cross-tenant RDMA traffic, watch it drop at the NIC, then permit it with a selector change. Enforcement applies within 1 second with no throughput degradation. We cover the CRD design, selector-to-hardware translation, failure modes, and how this could inform NetworkPolicy extensions for kernel-bypass transports
Large language models fit naturally into Kubernetes demos. Operating them reliably is a different problem.
Most operational challenges AI workloads introduce have individual mitigations: model caching, GPU autoscaling, persistent storage, artifact replication. The operational gap is how they compound. Solving them individually still leaves platform teams with unreliable deployments.
This poster examines the assumptions GitOps and Kubernetes tooling quietly depend on that AI inference workloads break: artifacts are small enough for standard operations, pods start fast enough for default probes, sync completion implies readiness, infrastructure state is disposable between deploys, and dependencies are visible and explicit.
Drawing from operating multi-cluster AI inference infrastructure across multiple environments, this poster presents deployment lifecycle diagrams, failure-mode taxonomies, and practical operational patterns for improving AI infrastructure reliability.
Multi-agent setups are easy to sketch on a whiteboard and painful to run on a cluster. If you've watched agents chase each other in circles or ship traces that stop at the model boundary, you already know why "more agents" isn't the hard part, but the operations are. Daniel Oh focuses on the boring-but-important layers such as isolation and scaling on Kubernetes, handoffs with CloudEvents, tracing/metrics with OpenTelemetry across model calls and tools, plain gRPC/HTTP between services, guardrails with OPA, SPIFFE/SPIRE where identity actually matters, and KEDA when work piles up asynchronously. Then we run a live Kubernetes demo using cloud-native microservices (Java/Quarkus-style) wiring multiple agents together end-to-end, with traces you can actually follow in Grafana or Jaeger. Walk away with a short implementation checklist, a handful of failure modes worth catching early, and a simple mental model for keeping multi-agent systems maintainable without bolting on bespoke glue.
Kubernetes control-plane changes are hard to validate before production. Scheduling policy, autoscaling, admission rules, custom controllers, and capacity-management logic can behave differently once many pods, nodes, and API objects interact. Running these scenarios against real staging clusters on every change is rarely practical because of time and cost.
This poster presents a KWOK-based CI pattern for testing control-plane behavior without real kubelets or workloads. The workflow captures snapshots from production clusters, replays those objects into KWOK-backed simulation clusters, applies proposed changes, and verifies the resulting placement, scale-up/down, utilization, and failure-mode behavior.
Attendees will see where KWOK simulation fits between unit tests and real-cluster validation, what regressions it can catch early, and where simulation boundaries remain.
Your GPU dashboard shows 80% utilization. Whose 80%? When tenants share a GPU, the observability stack reports a single number, and that number is lying. It cannot say which tenant caused the noisy neighbour, who is paying for leaked bandwidth, or where the next OOM comes from.
This session audits GPU observability on Kubernetes in 2026. The core demo traces a noisy-neighbour inference incident across MIG, MPS, and HAMi, comparing DCGM, Prometheus, and OpenTelemetry views to show which tenant signals are visible, misleading, or missing. It separates what DCGM and the NVIDIA GPU Operator expose from what OpenTelemetry GenAI can attribute at the request and model layer, then highlights the remaining gaps: per-token cost attribution, correlation between GenAI spans and GPU counters, and thermal/RAS signals that rarely make it into tenant dashboards.
Attendees leave knowing which tenant-level signals their dashboards miss today, and a checklist of what to instrument next.
Every Kubernetes user has faced this: your Pod needs more CPU or memory, but changing resource requests means restarting it — dropping connections, losing in-memory state, and disrupting your workload. What if Kubernetes could resize Pods in place, like dragging a window corner to make it bigger?
In-Place Pod Resize lets you change CPU and memory requests and limits on running containers without restarting them. This talk uses the "explain like I'm 5" approach to break down how it actually works — from the API changes in Pod spec, to how the kubelet negotiates the resize with the container runtime, to what happens when a resize can't be granted. You'll understand resize policies, status conditions, and how VPA can now autoscale your workloads without the restart tax. We'll share production lessons learned and common gotchas from early adopters. By the end, you'll know exactly when to use this feature and when a restart is still the right call
At CVS Health, we have many engineers working to deliver the latest innovations in health services. In the DevEx organization, we understand that these engineers need feature flagging.
We researched what was already available within DevEx and found that two types of solutions existed: expensive vendor-driven platforms and home-grown tooling with very limited feature sets.
After examining these options, we looked to open source software and discovered OpenFeature, a vendor-agnostic API which would enable us to build a flexible, open source standardized platform to create and manage feature flags.
We collaborated with the OpenFeature Project, to understand their plans for the future of their Feature Flagging ecosystem and also to understand how we can contribute back to the community and include what we’re building in that future.
We're excited to share that we are bringing this effort to the open source community. We'd like to share it, and our journey of how we got here, with you.
As agentic systems become more effective, human review becomes the next scaling bottleneck. This talk presents a Kubernetes-native model of earned autonomy, where authority expands through demonstrated performance rather than static permissions.
As agents generate more pull requests, they consume an increasingly scarce resource: human attention. The goal is to spend that attention deliberately, by auto-merging work where tests and a track record justify it, and reserving human review for cases that require judgment.
Borrowing from progressive delivery, agents begin in shadow mode, with their proposals graded against real outcomes, earning authority only after sustained success within narrowly defined classes of work. The pace of delegation is set by how reliably outcomes can be evaluated. Objective verifiable outcomes enable trust to build quickly, freeing up human judgement for evaluation where it is most needed.
Container startup latency directly impacts service responsiveness, autoscaling speed, and incident recovery on Kubernetes. Yet much of that latency is dominated by the image pull step; which is mostly hidden.
Drawing on a real LinkedIn production case study — over 500K bare-metal hosts, 25M cores, and pod density averaging beyond 50 per host in our densest clusters — we’ll show how we cut the long tail of pod startup latency by 42%. We'll examine how container image pulls contribute to startup latency: taking a deep dive into the image pull path to identify bottlenecks and build a mental model, then covering the practical levers we have — compression, kubelet and containerd configuration, and P2P caching.
Come behind the scenes of how Kubernetes actually works! In this session, Lucy Sweet (Member of Technical Staff at Anthropic and Node Lifecycle WG co-Lead for the Kubernetes project) and Sandeep Kanabar (Lead Software Engineer at Gen) will take you on a live journey through setting up Kubernetes “The Hard Way”. No tools, no automation.
We’ll peel back the layers of abstraction to reveal how the core components of Kubernetes interact to orchestrate workloads from small to planet scale. We’ll trace the lifecycle of a Pod from creation through to running container, and explain how often-overlooked components play a critical role in cluster operations.
Through this live demonstration, attendees will gain a deep understanding of Kubernetes' internal mechanics, empowering them to troubleshoot more effectively, make informed architectural decisions, and unlock the true potential of Kubernetes in production environments.
When NVIDIA released drivers for the GB200 and GB300, it changed the way GPU memory was exposed to the system. Both GPU and CPU memory now appeared as standard NUMA nodes, unified in a single coherent address space accessible to any application. This model was extremely powerful — but it also broke Kubernetes in ways nobody anticipated: CPU-only pods silently consumed GPU memory, pod memory limits began firing on CUDA allocations, and pods that should have been queued under memory pressure scheduled and crashed instead.
Adapting Kubernetes quickly proved too complex. Instead, a new model called Coherent Driver-based Memory Management (CDMM) was introduced, returning control of GPU memory to the underlying driver and hiding it from the OS. While CDMM closed the most critical gaps in Kubernetes, some lingering issues remain.
In this talk we'll walk through the initial set of failures, the path to CDMM, and the patches still needed to run Kubernetes reliably on memory-coherent hardware.
etcd stores every object update as a full copy. Not a diff – the entire object, every time. A 15KB pod patched once per second generates 15KB/s of etcd writes that compaction can't keep up with. When the database hits its quota, the cluster goes read-only. No new pods. No scaling. No deletes. You can't even clean up because deletion is a write operation.
This isn't theoretical. It's a documented property of etcd's MVCC model that surfaces predictably at scale – usually at 2 AM when a misbehaving controller has been patching status fields for hours. We'll break a cluster live on stage (kind with a small quota), show exactly how the failure cascades, walk through the non-obvious recovery sequence, and then build the guardrails that prevent it: admission policies that cap annotation size, Prometheus alerts on revision growth rate, and architectural patterns that move high-churn metadata out of the critical path.
For two decades, projects have bent YAML to fit their needs: validation,
comments, type coercion, parsing rules.
Most of it gets bolted on.
In go-yaml v4, we're making it first-class with a deeply integrated plugin
system.
The plugin system refactors the library's internals: the parts people most want
to change become plugins with stable APIs.
go-yaml's defaults are plugins you can reconfigure or replace.
We'll tour some of the plugin families landing in v4:
– Schema: associate validation schemas (JSON Schema and others) with YAML.
– Rules: decide how nodes are interpreted based on their shape and content, enabling custom DSLs.
– Comments: make round-tripping configurable to customized integrations.
– Tag Functions: wire your own custom constructor libraries into the go-yaml loader stack, with a stdlib included.
Rules and Tag Functions together can approximate YAMLScript, scoped to operations you allow.
For engineers whose company needs YAML to behave differently.
Kubernetes is not the right platform, out of the box, to run AI agents. AI agents are bursty, they are asynchronous, pause for LLM, user, and other agents, and resume potentially across workloads. Kubernetes needs a runtime layer on top to manage this unique set of characteristics. That's where the agent-substrate project comes in.
Agent-substrate is a layer that runs on top of Kubernetes that can dynamically schedule "actors" like AI agents to workloads and evict them when they pause thus making the workload reusable. The workloads that agent-substrate uses are hardened sandboxes like gvisor or FirecrackerVM. Agent-substrate can run on any Kubernetes distribution. This talk will deep dive agent-substrate and show how it can be used to run agents safely at scale in an agentic Kubernetes runtime like the CNCF kagent project.
For years, the Kubernetes community has lived with a "necessary evil": granting monitoring agents and health checkers the nodes/proxy permission. While intended for simple metrics scraping, this permission is effectively a node-level superuser capability. Recent security research in early 2026 demonstrated that even a simple GET request on this subresource could be exploited via WebSockets to achieve unlogged Remote Code Execution (RCE) in any pod.
The wait for a fix is over. With the graduation of Fine-Grained Kubelet API Authorization to General Availability in v1.36, Kubernetes finally provides a robust, least-privilege alternative.
You've used agentic apps, but have you deployed the models generating those tokens? Welcome to the server side!
For platform engineers and ML operators, deploying LLMs isn't just "running a container." It demands a new mental model: LLMs cache aggressively, scale by queue depth rather than traditional utilization, and require strict hardware topology awareness (GPUs aren't fungible!).
We'll tackle these challenges using vLLM. We'll start simple with a CPU deployment. Then, we'll climb the ladder: single-GPU, tensor-parallel within a node, and multi-node pipeline parallelism using LeaderWorkerSet. Finally, we'll demystify disaggregated prefill/decode.
You'll learn to read metrics that actually matter (like Time-to-First-Token) instead of acronym soup. End users need fast, private, cost-effective models; you'll leave ready to build the Kubernetes stack that delivers them!
Recent large-scale DNS disruptions highlight the critical role DNS plays in modern infrastructure. In October 2025, a cloud hyperscaler experienced a major disruption involving internal DNS and service discovery, causing application impact estimated in the hundreds of millions to billions of dollars. In July 2025, a leading DNS provider experienced a public recursive resolver outage, showing how failures across DNS can have broad consequences.
These events reinforce the need for scalable, resilient DNS infrastructure. This session explores how CoreDNS is evolving for cloud environments, covering improvements in scalability and operational resilience, along with lessons learned from running DNS at scale.
As AI accelerates vulnerability discovery and defense, it also introduces new challenges. We are seeing more AI-assisted security reports, with mixed results. We will review recent CoreDNS security fixes, hardening strategies, and the future direction of the ecosystem.
SIG Docs Localization has helped Kubernetes grow globally by welcoming users and contributors in their own languages. K8s.io now supports 17 languages, including recently launched Persian, but keeping them up to date is hard: upstream English docs keep changing, teams vary in size and workflow, and reviewers lack time.
AI can help, but raw AI translation PRs don’t meet our quality bar. They may use incorrect or inconsistent terms, or miss local context. In the AI era, human contributors matter more, not less. This session shares how we’re adapting: applying AI guidance while turning glossaries, style conventions, and review practices into reusable AI context by using tools to flag outdated pages.
We’ll discuss where AI can assist, what reviewers must own, and how context-aware localization can support teams without bypassing peer review. Attendees will learn how to get involved in localization, not merely translation: it is how global contributors can find their place in Kubernetes.
Linkerd 2.20 introduced rate-limit-aware load balancing and circuit breaking. When backends signal overload, from Linkerd or by other means, the proxy biases traffic away from those endpoints or removes them from the pool entirely. This session walks through this feature in full, along with the rest of features introduced in 2.20. This includes a tremendous refactor that cuts memory usage by up to 85% in enterprise environments. After deep diving into 2.20, we’ll cover what is upcoming in 2.21. Whether you are a Linky-lover or are taking your first steps in service mesh, this session will leave you with the knowledge on how these features make your workloads more reliable and easier to manage.
etcd 3.7 was released in June and we're planning and working on version 3.8. Join the SIG-etcd team to learn about the features in 3.7, the 3.8 roadmap, new subprojects and special efforts, and how you can get involved contributing to etcd. With demos!
Bring your etcd questions because the folks who can answer them will be on stage.
This hands-on session is designed to help OSS enthusiasts, end-users and ecosystem partners contribute to Kyverno, a CNCF graduated project, that provides a policy engine that elegantly solves critical challenges across security, automation, and compliance.
You will learn about Kyverno’s architecture, the role of each API type, the components, how to set up your development environment, and how to contribute to the project.
We will also discuss how Kyverno has evolved from a Kubernetes only solution to a unified policy engine that can be applied to any layer of the stack, including newer AI agentic applications.
This session will be led by Kyverno maintainers and contributors and is organized so that both developers as well as non-developers can contribute across the software base, sample policies, and documentation.
Join us to shape the future of cloud native and AI governance together!
Cadence Workflow is a CNCF sandbox project for durable workflow orchestration. This ContribFest will be a working session for improving Cadence’s AI-agent developer path: runnable examples, docs, and starter issues that show how to model long-running agentic workflows as durable workflows.
Participants will work with Cadence maintainers on concrete tasks such as validating the local quickstart, improving docs for retries/timeouts/idempotent activities, creating or polishing an AI-agent workflow sample, and identifying confusing setup or contribution steps. The session will start with a short overview of Cadence and the available issues, then split into small groups to make docs PRs, sample updates, or issue triage progress.
The project benefit is a better contributor onramp and more practical examples for developers building reliable agent workflows with Cadence.
Today, a multi-billion-dollar commercial market relies on a dangerously small and overburdened pool of contributors and maintainers. Maintainer burnout is the single greatest threat to OSS health and a direct risk to software supply chain security. Think etcd and ingress-nginx.
One solution: attract and welcome the untapped talent that's been overlooked. That's the mission of Merge Forward, a CNCF community initiative home to seven diversity groups, from Women in Cloud Native to Neurodiversity to Deaf and Hard of Hearing. We build mentorship networks and community spaces for contributors from underrepresented backgrounds and the allies who champion them, helping close the gap between who uses open source and who builds it.
Open source is only as resilient as the community behind it. Come find out how you can help.
As AI infrastructure scales, performance depends not just on GPU count, but on where GPUs sit in the network. Training loses efficiency when collective operations cross spine boundaries. Inference can become bandwidth-limited when tensor-parallel workers are placed outside a shared fabric. Disaggregated inference adds KV cache movement between prefill and decode pools.
Topology-aware scheduling can address these problems and is becoming prominent across Kubernetes-native and Slurm-backed AI platforms. Kueue and KAI Scheduler bring topology-aware placement to Kubernetes; Slinky connects Kubernetes workflows to Slurm. Across those paths, the same prerequisite is assumed rather than solved: accurate, current topology data published in the cluster.
In this talk, we show how Topograph discovers network topology across cloud and on-prem environments and publishes it into Kubernetes and Slinky/Slurm, giving each scheduling layer the information it needs for placement decisions.
Documentation drifts silently. A 2025 GitHub Codespaces update bumped the minimum Docker version and broke every tutorial for Drasi, a CNCF project, at once. Manual testing meant the four-person team didn't catch it for weeks: it doesn't scale, and CI doesn't read prose.
This talk reframes documentation testing as a simulation problem: an AI agent runs inside a sandboxed Dev Container as a naïve new user. It reads the docs, runs commands, verifies output, and files a GitHub issue when they diverge.
In 5 minutes: the architecture (Dev Container, Playwright, agent CLI, weekly GitHub Action), the security model (the container is the boundary), and the result: 200+ synthetic-user runs, 18+ bugs fixed, zero maintainer effort.
The pattern is agent-agnostic. The demo uses Copilot CLI (free for CNCF maintainers) but runs the same on Claude Code and Gemini CLI.
Implement the pattern, and your agent will tell you what's broken before your users do.
Kubernetes Secrets are everywhere: passwords, API keys, certificates, tokens, and cloud credentials. But they were never designed to be a full secrets management system.
This talk explains where Kubernetes Secrets break down in real platforms: storage, mounts, environment variables, logs, GitOps workflows, RBAC, and multi-cluster replication. It also explores safer patterns, including external secrets operators, CSI secret stores, workload identity, short-lived credentials, vault-backed approaches, and secretless architectures.
Attendees will learn when Kubernetes Secrets are acceptable, when they become risky, and how to design safer cloud native platforms without pretending that base64 is security.
Two kernel CVEs hit in the same week — Copy-Fail (CVE-2026-31431) and Dirty Frag (CVE-2026-43284). We fixed both across 5,000+ nodes on EKS, GKE, AKS, OKE, and on-prem; the largest cluster (500 nodes) had every node fixed and checked in 3 minutes from one kubectl apply. But the standard fix didn't look for vulnerable kernel modules that were already loaded, so half-fixed nodes would have looked fine until the next reboot. We added that check to the probe so those nodes showed up as failing pods, then a second, targeted CR rebooted only the nodes that needed it. The same pattern works for any change you need to land on every node: a sysctl, an apt package, etc. Ship the change plus a check that fails when the change didn't work, declared as a CR. kubectl get pods is your check; kubectl apply rolls it out; kubectl delete rolls it back; PDBs + podNonInterruptLabels decide when to cordon, drain, or reboot.
Most Kubernetes cost allocators distribute idle globally. Every team gets a share proportional to their requests. The result: a tax with no behavioral signal. Efficient teams pay the same idle premium as teams forcing oversized pods onto undersized instances.
This 5-minute talk presents a production methodology in use at Nubank since 2025 that allocates idle per-instance. For each instance, the algorithm reads pod composition (requests, usage, purchase option, instance class) and attributes that machine's idle to its pods. Services that generate idle pay for it; services that pack well pay for less.
The math is the incentive: no blame logic, no penalty rules. At Nubank's scale the law of large numbers does what hard-coded rules used to do badly.
The five minutes show one instance-level allocation end-to-end, contrast it against the global approach, and name the design principle: attribute locally, let scale handle fairness.
A Grafana dashboard, written as code and shipped through the Prometheus operator's sidecar, can deploy cleanly and still render "No data" on most of its panels even when the query is right and the scrape is healthy. This 5-minute lightning covers silent failure modes in kube-state-metrics that has burned dashboards at our company before. It walks through the diagnostic that distinguishes the bug from a broken query, then shows what to couple together in the same pull request so the regression cannot recur.
What actually happens before a Kubernetes Pod can send or receive traffic?
Before the first packet is ever transmitted, Kubernetes performs a complex series of networking operations behind the scenes. Without understanding these steps and the potential errors along the way, troubleshooting Pod startup networking issues becomes extremely difficult.
This session follows the networking lifecycle of a Pod from startup to first packet using live demos and packet tracing. We will walk through how Kubernetes creates the Pod sandbox, establishes shared networking through the pause container, and invokes the CNI. We will also dive into namespace setup, veth pair creation, IPAM, routing, and overlay configuration step by step before a Pod becomes reachable.
Attendees will gain a deep understanding of Pod networking at startup which will help them debug common failures including CNI errors, missing routes, overlay issues, and broken Pod connectivity.
Intuit processes a massive, continuous stream of financial events across the QuickBooks ecosystem. Every transaction, payment, and update flows through pipelines that write to the databases customers' books depend on. The challenge is twofold: sustaining millions of events with the elasticity and cost-efficiency of a cloud-native system, while never compromising correctness. For financial data, an event processed out of order or counted twice becomes a wrong number in someone's books.
We'll walk through the architecture that does both, built on Numaflow, an open-source, Kubernetes-native streaming platform developed at Intuit. We'll cover how we scale horizontally to absorb millions of events, how per-entity ordering and exactly-once semantics keep updates correct under load, how we autoscale (including scale-to-zero) to control cost, and how distributed tracing follows a single event end to end. Attendees leave with patterns for high-throughput, correctness-critical pipelines.
Training large ML models on multi-cluster, multi-region GPU fleets is hard: GPUs sit idle on one cluster while jobs queue on another, quotas lack a global view, and datasets are repeatedly re-downloaded across regions. This end-user session shows how Roblox turns a heterogeneous fleet into one shared GPU pool using only CNCF projects. Kueue and MultiKueue provide a single submission point with cross-cluster RayJob dispatch, tiered per-team quotas (guaranteed, borrowing, opportunistic), priority preemption, and data-locality-aware routing with automatic fallback. Dragonfly, a CNCF graduated P2P system, caches datasets and checkpoints in-cluster, turning cross-region downloads into local peer reads. Together, they optimize where jobs run, who gets capacity, and how data arrives. Attendees leave with a reproducible blueprint, real operational failure modes, and lessons learned.
Deploying Kubernetes onto mobile hardware like humanoid robots introduces severe physical risks. While GitOps tools can use "Sync Windows", applying an automated update or restarting a container while a heavy robot is moving can have real-world consequences.
To achieve true defense-in-depth, we implemented a custom Kubernetes Admission Controller directly on the edge cluster. Acting as an infrastructure-level gatekeeper, it intercepts all API requests and evaluates them against the robot's state. Regardless of GitOps schedules, modifications are safely blocked if the hardware is in a high-risk state.
This session explores a production architecture designed to bridge the gap between cloud-native infrastructure and live hardware We will detail how the admission webhook interfaces directly with the active robot, how it catches risks missed by static schedules, and the essential considerations of "break-glass" overrides to prevent lockout during edge cases
We gave a bedrock assistant the keys to 30 EKS clusters and the AWS and GitLab control planes behind them. Read only by default, write capable with reviews. We did it without a vector store, without RAG, and without a 200K token prompt.
This talk is the engineering behind that claim. We will walk through some of the techniques behind it. A typed tool catalog the model plans against, server side projection that turns 40 MB of DescribeInstances into 800 bytes of data, an iterative plan loop with a hard token budget, and a consent layer that makes kubectl apply and git push safe to expose to a language model.
Engineers will leave with a concrete pattern for letting an LLM operate a real fleet. One that scales with your control plane instead of your embeddings bill, and that fails closed when it matters. Example source code is open source so everyone can start their own journey.
GPUs are now a critical resource in Kubernetes clusters, yet observing what is happening on them remains a black box. eBPF, already the go-to technology for system-level observability, can be extended to observe GPU workloads, providing production grade visibility without modifying application code.
In this talk, we will show you how to get started writing eBPF programs that observe GPU activity. We will walk through how to hook into GPU interfaces, enrich data with Kubernetes context (pods, containers, namespaces), and extract meaningful metrics. Then we’ll cover how to package your eBPF programs as OCI images for easy distribution and deployment across clusters. Finally, we’ll demonstrate how to pipe GPU telemetry into Prometheus and OpenTelemetry so it integrates seamlessly with your existing observability stack. Along the way, we will share guidelines and lessons learned from building GPU observability programs, so you can avoid common pitfalls and hit the ground running.
We've seen it before: autoscalers not reacting to traffic spikes for hours, taking down the service. Or a DaemonSet silently spawning duplicate pods on a node, wasting resources in an already strained system. Meanwhile, alerts and dashboards show nothing amiss? There's no bug in your operator's logic, it's a symptom of a deeper problem. Can you wait hours for your system to react?
In a watch-based, eventually consistent environment, staleness is a mirage in your control plane. Your operator believes it is reacting to the present, but is actually trapped in the past. How do we defend our infrastructure against it?
The Kubernetes community is finally tackling this. This session demystifies watch staleness, exposing gaps in how we design reconciliation loops. We introduce capabilities in Kubernetes v1.36+, including comparable resource versions and staleness mitigation. You’ll leave able to build operators that degrade gracefully, detect staleness, and keep your clusters reliable.
Internal Developer Platforms started as portals, golden paths, templates, and self-service workflows. In cloud-native environments, many are now becoming platform control planes with APIs, reconciliation loops, policy boundaries, and multiple consumers: humans, pipelines, CLIs, automation, and AI agents.
This session maps the IDP landscape using four axes: API surface, primary consumer, policy focus, and reconciliation model. We compare portal-first, GitOps-first, control-plane-first, promise/workflow-based, distribution-based, application abstraction, and agent-facing models.
The focus is on architecture diagrams, API boundaries, and reconciliation flows: where requests enter, where policy is enforced, and where ownership changes hands. Attendees leave with a practical framework for choosing or composing IDP architectures without adding accidental complexity.
Ok, they won’t actually be free: you'll still need to rent (or buy) GPU time from your cloud provider. But you'll no longer have to pay per-token costs. In this talk, we'll walk you through how to build a practical, scalable, and multi-tenant LLM "token factory" on your existing Kubernetes platform, using nothing but Kubernetes and a few open source projects. Hint: it’s easier than you thought, and advances in open weight models mean that the quality is astonishingly good. This talk will tackle the practical side of Kubernetes inference serving: choosing the right model and model server; handling authorization, rate limiting, and QoS; evaluating how far you can push cheap hardware and small models; handling concurrency; and more. Whether you’re running inference-enabled code, dealing with data sovereignty requirements, or just want a cheaper alternative to frontier models, it all comes down to serving inference reliably and cost-effectively at scale.
You wouldn't ask a plumber to sign off on your electrical work. Yet most CI/CD pipelines run under a single identity: one credential for signing SBOMs and reporting vulnerabilities alike. Workload identities encode location, not authorization.
This session shows how to close that gap with Tekton: admission-time verification of signed remote task definitions constrains the SPIFFE/SPIRE identity each task receives. Tasks use that SVID to produce in-toto attestations scoped to their role. Policy engines can then enforce that each claim came from the right kind of task — a vulnerability attestation from a scanner, an SBOM from a build task — not just that a signature exists.
Attendees leave with a working technique and one transferable insight: authorization should follow role, not location — from signing attestations to authenticating against external services.
The maintainers of the CNCF Crossplane project (https://www.crossplane.io/) will lead this session that not only introduces the project to new attendees, but also dives deep into Crossplane's latest features, releases, and roadmap.
Crossplane has been production-ready for years, and teams are now building more complex and powerful platforms than ever, using only the primitives from Crossplane's framework. We'll show how those primitives let you build higher-level functionality, and how the same foundation lets AI agents and automated workflows act safely on production infrastructure.
We're also excited to help you make the jump to Crossplane v2, with a walkthrough of the latest migration support and tooling, as well as new reliability and safety features like deletion protection that blocks in-use resources from being removed when a provider is uninstalled.
Join us for live demos, dig into the details with the maintainers, and help shape where the project goes next!
As Gateway API adoption hits an all-time high, the journey is just getting started. Are you ready for what comes next on Kubernetes traffic management?
Join the Gateway API maintainers for a high-energy update on the project's evolution. We have moved beyond simplifying Ingress migrations to tackling the next generation of networking challenges: from legacy TCP/UDP integration and flexible Backend policies to the cutting edge of AI and Agentic traffic patterns.
Whether you are mastering HTTPRoute features, exploring non-HTTP routing, or curious about our approach to AI integration, this session is your destination. Get a direct look at the latest roadmap, learn about new features, and bring your questions—we want to hear what you need from the project next.
The rise of highly complex workloads, most notably GenAI, has driven incredible innovation in managed services and specialized infrastructure-as-code templates. While these tools reduce developer friction, organizations operating across multi-cloud or hybrid environments face a growing challenge: maintaining a unified, highly portable deployment model that standardizes on native Kubernetes APIs.
This talk demonstrates how maintainers and platform builders can leverage the Kubernetes Resource Orchestrator (KRO) to encapsulate massive complexity into highly portable, standard Kubernetes APIs. We will discuss:
- Architecting workloads for portability
- Building platforms without custom software operators
- Delivering unified API experiences across clouds
Attendees should expect to leave this presentation with strategies for using KRO to improve cloud-agnostic patterns in workloads, reduce resource orchestration complexity, and lessen the need to maintain bespoke software controllers.
Have you ever wondered how kubectl and kustomize enhancements are actually designed and built? Or maybe you’re curious why that game-changing feature request of yours hasn’t been accepted? Join the Kubernetes SIG CLI maintainers to pull back the curtain!
In this interactive session, we’ll take you on a whirlwind tour of the diverse tooling ecosystem we manage and show you exactly how to make your first contribution.
What’s on the Agenda:
* The Retrospective: A look at the major breakthroughs we've shipped over the past year.
* The Horizon: An exclusive sneak peek at our upcoming roadmap.
* Your Voice: A live forum to give feedback and help shape the future of the tools you use every single day.
Whether you're a seasoned contributor or complete beginner, come hang out, ask your burning questions, and help us build the future of the Kubernetes CLI!
Multi-agent systems are entering production and the Linux foundation’s Agent2Agent (A2A) protocol is fast becoming the standard method for agents to find one another on kubernetes and call each other. However, the architecture hides a scaling wall. When agents connect peer-to-peer, connections grow quadratically. Four agents need six links, fifty agents over twelve hundred. At scale, this N-squared explosion becomes impossible to route, secure, observe and keep reliable. This talk will show why direct A2A traffic does not scale, and how routing it through an A2A-aware gateway such as agentgateway, built on top of the Kubernetes Gateway API, collapses the tangled mesh into a manageable hub with load balancing & visibility. By the end of the talk, attendees will leave with the connection math, the infrastructure pattern that solves the problem, and one rule: design the connectivity layer before scaling the agents,not after.
AI changed the entry path into cloud native, but it did not remove the need for fundamentals, judgment, and community. Cloud native beginners no longer arrive through one door. One newcomer starts with AI copilots, speed, and confidence, but often without strong mental models. Another arrives with years of traditional infrastructure or software experience and must translate older habits into Kubernetes, platform engineering, and community-native ways of working. This five-minute talk gives both audiences the same first-90-days map: what to learn personally, what to let AI accelerate, how to verify AI output before trusting it, where veteran discipline still matters, and how to join the CNCF community without pretending to know everything. Fabrizio Sgura uses mentoring and community experience as a human frame, not a keynote memoir, to show how people, not copilots, stay accountable for judgment.
When teams first adopt Kubernetes, one of the earliest decisions they face is how to get application traffic into their cluster. As the ecosystem moves beyond the legacy Ingress API, Gateway API is emerging as the new standard for traffic management, but choosing what to adopt still requires understanding the tradeoffs between implementations.
In this session, we introduce Gateway API for cloud native newcomers: why it exists, what problems it solves, where portability is real, and where implementation differences still matter. Drawing on our experience maintaining security, networking, and traffic management systems in the Kubernetes ecosystem, we’ll use examples from Istio, Envoy, and agentgatway to help attendees build a practical framework for evaluation.
Attendees will leave better prepared to choose a gateway approach for new platforms and migrations.
When a language model fails in production, the symptoms rarely show up in traditional dashboards. In this session, we explore what it really means to observe an inference system: which metrics actually matter and which ones can be misleading, how to correlate hardware behavior with model behavior, and how to build a unified view that goes from silicon all the way to the user response. We will talk about token latency, throughput, GPU saturation, and application traces.
As Kubernetes adoption scaled across PlayStation, our original namespace and RBAC workflow began to struggle to scale with the operational complexity. Merge conflicts, tightly coupled automation flows, and centralized ownership created friction for both platform teams and developers.
In this session, we’ll share how we evolved from a DevOps-owned RBAC model to a Kubernetes-native platform capability built on CRDs. We’ll cover how decoupling namespace lifecycle management from RBAC intent enabled safer self-service access management, improved scalability across multi-cluster environments, and reduced operational bottlenecks.
Attendees will leave with practical patterns for modeling RBAC as platform capabilities and applying platform engineering principles to Kubernetes access governance at scale.
Kubernetes upgrades have historically advanced one minor at a time: apiserver first, then controller-manager and scheduler. That creates two waves per minor. At LinkedIn’s scale, 300 clusters and roughly 300,000 nodes, each wave must pass through canary, regional ramps, soak, and blast-radius controls. A three-minor cadence therefore meant six fleet rollouts and more than 35 business days
KEP-4330 changed the model: newer control-plane binaries can emulate older-minor behavior for feature gates, APIs, storage encoding, and validation, while rollback becomes config instead of binary reversal
I investigated whether KEP-4330 could compress six waves into one. Upstream had documented the conservative stepwise path; source code review showed compatibility-version bounds w/o requiring intermediate hops. After removed-gate cleanup, /statusz checks, and canary validation, we promoted v1.31 to v1.34 fleet-wide in one wave under 6 business days [85% gain], with ramp, soak, and rollback intact
GB200s and GB300s are here! Fragmentation is widening the gap between “I have GPU quota” and “my job is actually running.” NVLink domains, rack boundaries, and the physical layout of the data center now shape whether a large training job can start, where it lands, and how efficiently it runs. On modern GPU clusters, quota is no longer enough. Topology has become a scheduling requirement. Kueue’s Topology Aware Scheduling (TAS) brings that reality into Kubernetes.
We will cover three things. First, how TAS works and the configuration choices that matter, including preferred versus required topologies and PodSet slices. Second, how topology-aware placement affects workload performance. Third, the production lessons that do not show up in the happy path: sticky scheduling, quota debugging, fragmentation, and node hot swaps.
Attendees will leave understanding how topology-aware scheduling helps AI teams turn expensive GPU capacity into actual training progress.
This session presents a production architecture orchestrating 252 GPUs across 21 multi-GPU nodes on Kubernetes to fine-tune a domain-specific Vision Language Model on 50–80 TB of multilingual document data.
We cover: (1) Topology-aware scheduling — co-locating pods on NVSwitch-connected nodes and gang-scheduling with Volcano for all-or-nothing job placement. (2) High-bandwidth networking — configuring RDMA interfaces as secondary networks via Multus CNI and tuning NCCL for optimal cross-node all-reduce. (3) Data pipeline — streaming tens of terabytes through OCR, enrichment, and tokenization using parallel Kubernetes jobs. (4) Failure recovery — checkpoint-resume strategies, health-check sidecars for silent GPU errors, and automatic restart logic.
Attendees planning GPU-intensive training on Kubernetes will leave with a practical blueprint covering scheduling, networking, storage, and failure handling patterns.
There are two kinds of autoscaling configurations: the ones that make for a great demo, and the ones that survive production. This session shares lessons from autoscaling strategies proven across thousands of production Kubernetes clusters.
Drawing from a broad KEDA user base, the session covers practical patterns for fine-tuning autoscaling behavior while improving resiliency: combining signals, avoiding noisy or overly reactive scaling decisions, and reducing the load autoscaling queries place on centralized monitoring systems through metric caching with OpenTelemetry.
In complex microservice architectures, autoscaling often spans multiple applications and infrastructure layers that must react in coordination. This session will discuss chained scaling patterns and practical ways to reduce delays between dependent scaling decisions, so teams can keep reliability and infrastructure costs predictable in production.
Admission Controllers and Mutating Webhooks have long been fundamental building blocks for Kubernetes platform extensibility. While powerful, relying on out-of-process requests can introduce significant pain points like increased network latency, complex TLS and cluster bootstrapping steps, and the risk of a failing webhook taking down your entire control plane.
This session highlights the evolution of admission plugins for Kubernetes and demonstrates practical, production-ready examples of the webhook-free world using Common Expression Language (CEL) policies evaluated inside the API Server. The session also explores the architectural tradeoffs of the native ValidatingAdmissionPolicy and MutatingAdmissionPolicy resources, contrasting the simplicity of CEL with full-featured policy engines like OPA Gatekeeper using Rego and Kyverno using CEL.
The shift to dynamic, agentic systems creates massive networking challenges. Autonomous agents operate across heterogeneous environments spanning laptops, edge devices, and diverse clouds. They also use an equally heterogeneous mix of protocols like MCP and A2A. How do we secure them?
This deep dive explores the requirements for agent-native infrastructure: making the network agnostic to physical limitations like NAT, and providing zero-config, zero-trust security. The solution is a decentralized data plane governed by a declarative policy framework.
We introduce Project SAM, an open-source implementation of this approach. Reusing proven industry lessons, SAM leverages the infinite scalability of Kademlia networks and OIDC identities to decouple routing from compute platforms, securing workloads everywhere.
What if your AI-powered observability stack could be quietly poisoned to trick AIOps agents into executing back doors or rogue deployments?
Our talk presents a production hardened zero trust telemetry architecture iterated on through real-world incidents: mTLS between OpenTelemetry Collectors and backends to encrypt and authenticate every payload; cert-based auth for federated Prometheus scrapes to prevent scraper hijacks; fine-grained RBAC on Loki and Grafana queries to block exfiltration; and append-only storage with cryptographic hashing to detect tampering before it fools your AIOps.
We’ll discuss the attack primitives in detail:
– AI-orchestrated injection exploiting tail-sampling blind spots
– Federation pivots chaining through scraped exporter paths
– Query-driven exfiltration of credential-stuffed logs
The audience will walk away with deployable Kubernetes manifests, eBPF verifier scripts for kernel isolation, and a battle-tested threat model template.
The OpenTelemetry Android Agent reached a stable 1.0.0 milestone in January!
Our API provides Android developers with an OpenTelemetry-first, Kotlin-based
domain-specific language (DSL) that makes it easy to instrument your mobile
application.
In this session, we will discuss the features offered by OpenTelemetry Android
and some of the techniques used to operate reliably in a highly diverse and
hostile runtime environment. We will show you the ins and outs of the DSL
and how simple it is to use in your application today. You'll also learn how
to apply auto-instrumentation at build-time in order to trace calls to your
edge services.
All of this will be accompanied by a live demo running on a physical phone.
WebAssembly (Wasm) is a practical sandboxing model for running untrusted code and other workloads on Kubernetes, but many developers and operators haven’t had a chance to try out the latest toolchains.
In this hands-on tutorial, we'll show you how to develop a Rust-based HTTP component using the Wasm Shell (wash) CLI. Then we’ll install the wasmCloud platform on a local cluster and deploy the component through the wasmCloud operator, scoping the host capabilities that define its sandbox. Finally, we'll scale the workload across hosts, observe how the deployment behaves, and troubleshoot the most common errors along the way.
By the end of the session, you'll have a working sandboxed component, a complete reference application manifest, and a clear mental model of where WebAssembly fits alongside containers.
The tutorial is designed for engineers and platform operators with basic Kubernetes familiarity. No prior experience with Rust or WebAssembly required.
The rise of AI-powered security scanners has fundamentally disrupted open-source maintenance. Vulnerabilities across edge-computing components, from runtime bugs to EdgeHub/CloudHub synchronization flaws, are being surfaced by automated systems faster than human maintainers can patch them. For KubeEdge, this influx creates an unsustainable project burden.
In this session, we share how the KubeEdge maintainers team is fighting fire with fire. We will demonstrate how we leverage AI-assisted workflows directly within our development to match this pace. We’ll show how AI review agents interpret complex disclosures, propose optimized patches tailored for footprint-constrained edge nodes, and automate regression testing. By scaling vulnerability management with AI, we show how maintainers can accelerate remediation while preserving production stability.
Kubernetes Operators have transformed how complex applications are managed, but ecosystem growth has introduced challenges around packaging, upgrades, configuration, and lifecycle management. The next generation of operators needs a simpler foundation for developers and platform users.
In this session, the Operator Framework team will explore how OLM is rethinking Kubernetes Operator Lifecycle Management from the ground up. The session will cover the motivations behind the redesign, key architectural changes, and how OLM uses declarative APIs to simplify operator installation, upgrades, and management.
The session will cover the operator journey from packaging and distribution to production operations, including direct bundle installs, OLMv0 migration strategies, telemetry and operational visibility, and enterprise adoption considerations. The team will also discuss how Java Operator SDK complements this evolution by improving operator reliability and scalability.
Karmada (Kubernetes Armada) is a Kubernetes management system that enables you to run your cloud-native applications across multiple Kubernetes clusters and clouds.
In this presentation, the maintainer of the Karmada project will share:
- A Brief Introduction to Karmada.
- Typical use cases
- New features over the last year
- Real-world case studies
- Overview of the community
- Roadmap
- QA
As AI models grow to hundreds of gigabytes, efficient distribution across GPU clusters becomes a critical infrastructure challenge. This session covers Dragonfly's evolving role in AI/ML model distribution, including the HuggingFace and ModelScope backend integrations, the dragonfly-injector for transparent container-level acceleration, and real-world patterns for delivering large models to GPU nodes at scale. We'll share lessons learned from contributing these features and discuss the roadmap for making Dragonfly the default model delivery layer in cloud native AI infrastructure.
eBPF gives deep, low-overhead visibility into Kubernetes and Linux, but writing, packaging, and sharing eBPF tools is still out of reach for most contributors. Inspektor Gadget, a CNCF project, changes that: it packages eBPF programs as portable OCI images called "Gadgets" that anyone can build, share, and run with `kubectl gadget`. Join the maintainers for a hands-on session that takes you from zero to your own published Gadget. We will walk through the Gadget architecture, set up your environment together, and help you write eBPF code that traces real kernel activity, enriched with Kubernetes context. You will build and publish your Gadget so the whole community can use it, then see the payoff: paired with Agent Skills, your Gadget lets an AI agent run `kubectl gadget` to debug a live cluster. Whether you want to write eBPF, work in Go, or improve docs, we will have curated issues and mentors to help. No prior eBPF experience needed.
Want to contribute to open source but not sure where to start? This ContribFest gets you going in 75 minutes. Backstage is the CNCF's open platform for building developer portals, with hundreds of plugins and contributors across dozens of organizations. This session is hands-on: you'll find where your skills and experience actually fit in the project, not just hear about how contributing works. We'll walk through the codebase (Node/JavaScript/React), the contribution process, and local dev setup. Then you pick from a board of prepared issues across docs, frontend, backend, and plugins, based on what you know and what you want to learn. Hosts and experienced contributors will be around to help you pick an issue, work through problems, and get your PR ready. You might leave with an open pull request, a branch in progress, or just a clearer idea of where to jump in next. No Backstage experience needed. Come curious; leave as a contributor-in-progress!
Our hero, an AI agent trying to run on Kubernetes, knows they are destined for greater things! They have a model to serve, tasks to complete, and users to delight. But deploying and managing inference, building and running your agent, and crafting and hosting tools is HARD. One wrong step could be catastrophic. And Hero wonders, should the Kubernetes destination affect the fundamental design of the agent itself?
It is up to you, the audience, to guide Hero to their final form⎯an AI agent running reliably on Kubernetes. In their seventh KubeCon 'Choose Your Own Adventure'-style talk, Whitney and Viktor present choices to make at each layer of the stack: inference, agent frameworks, and tooling. At each step, we discuss the pros and cons of using CNCF projects (KServe, kagent, kmcp, etc) vs rolling your own solutions. Then we let the audience (YOU!) decide what to build into the live demo. Can we assemble the right stack and get our agent running before the session time elapses?
Distributed AI training fails fast: if a single GPU pod in a multi-node job cannot schedule, accelerators sit idle or the workload fails with collective-op timeouts. To compensate, the ecosystem evolved many external schedulers, including Kueue, Volcano, and KAI, each forcing controllers to pick a side.
Kubernetes 1.36 changes this. The new Workload and PodGroup APIs (KEP-4671) introduce scheduler-native primitives for gang-scheduling, topology-aware placement, and coordinated use of high-bandwidth interconnects such as IMEX. TrainJob, JobSet, and LeaderWorkerSet are converging on these APIs.
This session covers Kubeflow Trainer’s adoption (KEP-3015): why TrainJob owns Workload composition and how its plugin architecture enables advanced scheduling features. Attendees will see a demo of multi-node fine-tuning on modern accelerators, how the Workload API unifies a fragmented ecosystem, and how to adopt it in their own controllers.
At Ford, we manage a large-scale Kubernetes platform across hundreds of clusters using GitOps. We use Kustomize with an automated templating and scaffolding system — an advanced alternative to Helm that generates per-cluster overlays from declarative configs. Moving quickly at this scale requires CI that acts as an unbreakable safety net. Our Tekton-based pipeline validates every pull request end-to-end — building overlays, catching template drift, and running semantic validation beyond syntax checking. Between our team, community contributions, AI-generated code, and automated dependency PRs, this safety net is more critical than ever. Automation and AI produce plausible YAML that violates platform conventions or introduces drift no human reviewer would catch. Attendees will leave with a repeatable model for safety-net CI any platform team can adopt — regardless of tooling, cluster count, or org size — plus the lessons so you can skip the mistakes we made.
Training and serving frontier AI models pushes Kubernetes past its breaking point. At massive node scale, the sheer explosion of objects, Leases, Pods and Nodes can melt the API server, bottlenecks etcd, and hangs core controllers. Add the requirement of protecting multi-billion-dollar model weights from insider threats using strict enclaves, and traditional privileged K8s Nodes quickly become a liability.
In this session, Google and Anthropic engineers reveal how we bypassed these limits by fundamentally rethinking K8s architecture. We will share a production tested approach that completely separates the control plane from the execution data plane, providing scalability and security.
We abstract compute into "linked runners" and extend the CRI protocol, eliminating Kubelet overhead and bypassing traditional Node objects entirely. Join us to learn how to achieve dynamic multi-host topology scheduling and build a scalable, zero-trust environment for your most critical AI workloads.
Building secure and scalable agents at scale is hard. From multi-tenancy and noisy-neighbor risk to tool access and data boundaries, teams need more than a prompt and a container image—they need a deliberate execution model.
At Adobe, we are running agents on the CNCF Agent Sandbox project: a sandboxed runtime for AI agents that manages their full lifecycle—provision, execute, observe, and tear down—while keeping workloads isolated from the rest of the platform. Sandboxes enforce resource limits, reduce blast radius, and give operators a consistent place to apply policy.
We have applied optimizations so that development teams can run agents without the need to understand the sandboxing mechanism or worry much about the increased complexity of single tenant agent execution.
In this session we’ll share the problems we hit running agents for real customers, how Agent Sandbox addresses them, and what we’d do differently next time.
LLM inference is becoming memory-centric, where capacity, bandwidth, and data movement dominate performance and cost. Techniques such as continuous batching and prefill/decode disaggregation (e.g., vLLM, LLM-d) improve efficiency within a node, but introduce a new challenge: managing large-scale KV cache across distributed nodes.
At the same time, disaggregated memory technologies like CXL enable memory to scale beyond node boundaries, creating an opportunity to extend Kubernetes with memory-aware orchestration while preserving its model.
We present CoHDI, a CNCF sandbox project that bridges Kubernetes and disaggregated infrastructure. CoHDI provides abstractions for memory-aware resource management across GPU HBM, DRAM, and CXL, and acts as a mediation layer with runtimes like vLLM to enable coordinated placement of compute and KV cache.
Using a vLLM prototype, we demonstrate early results showing the impact of tiered KV placement on concurrency, latency, and GPU utilization.
Production inference on Kubernetes is still too often a closed vertical — one vendor’s gateway, one serving image, one observability silo. Operators running multi-team clusters need the opposite: composable open layers they can operate, upgrade, and replace independently.
This session discusses open inference reference model components:
– Heterogeneous accelerators — GPUs, TPUs, and DPUs on the same Kubernetes control plane (device allocation, placement, capacity)
– Scheduling and fairness for inference and batch (llm-d, Kueue)
– Gateway and routing at the cluster edge (llm-d inference gateway, Gateway API patterns)
– Pluggable serving engines — swap vLLM, TGI, or the next runtime without replatforming
– OpenTelemetry for traces and cost attribution across gateway → serve → model
Speakers walk through a reference architecture and operator tradeoffs: how to compose without lock-in, and which layer to tackle first when inference outgrows bolt-on APIs.
In July 2025, Mistral started building a sovereign AI cloud from scratch. Five months later, our first GB200/GB300 cluster was in production. We did it by standing on the shoulders of CNCF and adjacent projects: ClusterAPI, Metal3, Kamaji, Sveltos, and Slinky (Slurm-on-Kubernetes), assembled together with our own operators. This talk is a one-year production retrospective. We will walk through the architecture (management cluster, multi-tenant control planes, bare-metal provisioning) and all the abstraction we built on top of it, then dive into the surprises: scaling etcd past 4,000 nodes, assisting in the arm64 support on Metal3, fleet monitoring, network management, and the operational reality of BMCs, Redfish, and NVL72 racks when you come from a "cloud-native" background. Attendees will leave with a concrete reference architecture for KaaS on bare metal, and an honest list of the potholes between a kind cluster and a 4,600 node GPU fleet.
What is causing slow cold starts in Kubernetes-based LLM serving systems beyond model loading?
In Kubernetes-based LLM serving, we often blame cold starts entirely on model loading. In reality, a large part of the delay comes from GPU kernel compilation. Even with model weights cached, kernels are recompiled from scratch on every new pod, even on every pod restart
In this talk, we share a practical KServe-based solution that captures the compiled GPU kernels during the initial cold start, stores them as signed OCI artifacts, and caches them alongside the LLMInferenceService.
New pods can then load these pre-compiled kernels at runtime, skipping redundant compilation during scale-out. In our tests under burst traffic, this approach reduced vLLM pod readiness time by 30–70%, depending on model size and GPU type.
If you're an AI platform engineer who is running large language models on Kubernetes, you will learn how kernel-level caching can be integrated into your serving stack.
The goal is not to hide long-lived credentials better; it is to need fewer of them, issue them later, and rotate them automatically. Kubernetes already warns that Secrets are stored unencrypted in etcd by default unless protections are added, that anyone with API or etcd access may retrieve them, and that users who can create Pods in a namespace can often read Secrets there indirectly. The platform has also moved away from older long-lived service-account token Secrets toward projected tokens obtained through the TokenRequest
API. This session turns those warnings into an architecture and migration guide: use projected serviceaccount tokens where they fit, adopt workload identity through SPIFFE/SPIRE, and bring in OpenBao when dynamic secrets, leases, and automatic revocation are needed. Rather than presenting a vendor parade, the talk shows where each layer belongs, how to reduce “secret zero” risk, and how to move from static
credentials toward short-lived, auditable access.
Your GPU bill is rising, your models are serving billions of tokens but what does each token actually cost? GPUs suffer from the same problems as other k8s resources like CPU and RAM: teams size resource requests based on estimates and guesses of anticipated usage, then move on to the next project. Are those GPUs actually used in production? Come learn how OpenCost, a CNCF incubating project, is integrating with llm-d, a CNCF sandbox project, to address these questions and more.
We will show how llm-d integrated with OpenCost turns GPU, vLLM, and llm-d metrics into per-model and per-token cost visibility. We will introduce two essential views: reservation-based cost, and usage-based cost. Combining this with idle detection logic enables comparison of self-hosting with SaaS APIs, intelligent routing of LLM workloads, optimization, and fair AI chargeback across shared infrastructure. Attendees will leave with a practical way to bring financial accountability to AI platform decisions.
SIG Architecture maintains and evolves the design principles of Kubernetes, and provides a consistent body of expertise necessary to ensure architectural consistency over time.
The SIG takes care of evolution of conformance definitions, API definitions/conventions, deprecation policy, design principles, and other cross-cutting concerns.
In this talk, we will provide an introduction to SIG architecture, and updates on its activities. In particular we will discuss:
- Balancing stability for existing users with evolution and change
- Shifts in how we're guiding features to behave in the face of scale limits / stale caches / backoff failure modes
- Making API review easier and more foolproof
AI Agents have emerged as mission-critical workloads. To balance performance and cost, these agents rely on warm-pooling for fast allocation and enter hibernate states when idle. However, Agents struggle with state persistence, where critical agent memory and skills is often lost during restarts. Furthermore, poor backward compatibility between versions complicates migrations, and the lack of specialized workload support makes it difficult to manage the lifecycle of multiple agents.
OpenKruise Agents is an operator designed to manage the sandbox lifecycle. It employs enhanced pre- and post-upgrade hooks, in-place upgrades, and checkpoint-based recreation to ensure seamless agents upgrades. The solution addresses three critical upgrade scenarios: warm pool refresh, in-place upgrades during warm pool allocation , and updates for allocated sandboxes. OpenKruise Agents also provides workload capabilities tailored for session-based agents, allowing phased rollout of a set of agents.
Kubernetes disruption handling is evolving from a collection of custom solutions and workarounds into a more cooperative lifecycle model. Join the Node Lifecycle Working Group for an update on two key efforts: the Eviction Request API for workload disruption coordination, and Specialized Lifecycle Management.
We will discuss how the new declarative Eviction API gives workloads, controllers, and administrators a more explicit way to negotiate disruption than today’s best-effort eviction paths. We will explore how it fits alongside PDBs, API-initiated eviction, preemption, autoscaling, and how it can help with app migration and maintenance automation.
We will also cover Node Lifecycle Conditions, the v1alpha foundation for Specialized Lifecycle Management. This change introduces new conditions to Nodes for draining and node maintenance. A shared condition model gives Kubernetes components and ecosystem controllers a common place to observe the Node lifecycle.
Metal3 manages baremetal host lifecycle declaratively through Kubernetes APIs. This session introduces three extensions covering tenancy, placement, and physical network automation, and demonstrates each.
Multitenancy lets teams share a common hardware pool. Workload owners select hosts by capability, while management credentials and hardware details stay with infrastructure admins. No per-team hardware silos, no credential leakage across namespaces.
Failure domain support spreads control plane nodes across racks using Cluster API semantics. Losing a rack does not lose quorum.
Top-of-rack (ToR) switch automation declares switch inventory and VLAN intent as Kubernetes resources. LLDP discovery maps hosts to physical switch ports. Metal3 configures port VLANs during provisioning and reverts them on decommission — no manual switch configuration.
The result: isolated, rack-aware, network-automated clusters for many teams on one fleet — for AI, edge, and sovereign cloud platforms.
Cloud native innovation has always been driven by the community. From Mesosphere's role as a CNCF founding member to D2iQ's continued Kubernetes leadership and now Nutanix's active contributions across multiple CNCF projects, the mission has remained the same: build together in the open.In this session, we'll revisit that journey and explore how those experiences led to CAREN (Cluster API Runtime Extensions – Nutanix), an extensible framework that enables dynamic cluster configuration and lifecycle management while preserving the simplicity and consistency of Cluster API's declarative model. As we look to the future of Kubernetes fleet management, standardization via Cluster API (CAPI) and the ClusterClass feature has brought immense consistency to cluster provisioning. However, real-world enterprise deployments often hit a wall: they require dynamic customization that strict standardization struggles to accommodate.CAREN bridges the critical gap between strict declarative blueprints and real-world adaptability. By leveraging CAPI's runtime mutation hooks, CAREN dynamically configures and manages Kubernetes clusters without ever needing to fork or touch the base ClusterClass. It empowers platform engineering teams to deliver standardized, yet infinitely adaptable Kubernetes clusters.
Just as the Linux ecosystem thrives by separating a universal, low-level kernel from specialized distributions, Kubernetes is evolving beyond a monolithic, "one-size-fits-all" architecture. Forcing a single platform to act as a universal orchestrator for everything from tiny edge devices to hyper-scale AI clusters creates inefficiency and strain.
To solve this, the industry is implicitly defining a lean Kubernetes Kernel—the core API machinery, reconciliation loops, and resource abstractions—as a standardized foundation. Simultaneously, agentic engineering is drastically reducing the cost of software creation and democratizing the operational expertise needed to build upon this foundation.
Join us to learn how building on the Kubernetes Kernel allows power-users to create vertically integrated, purpose-built custom distros that achieve significant performance gains while still crowd-sourcing defect discovery and reliability.
Kubernetes was designed for a predictable world. We built it to manage deterministic, stateless microservices with known communication paths, defined resource requirements, and static identity boundaries.
The rise of agentic AI completely shatters these assumptions.
An AI agent is not just another stateless workload, it is a non-deterministic entity that generates dynamic network calls, triggers external APIs, and executes machine-generated code on the fly. This behavior introduces a security blind spot: when an autonomous agent invokes a local tool or accesses a cloud database, what identity does it present?
In this keynote, Abdel Sghiouar unpacks the looming identity and resource crisis of running agents on Kubernetes. He will explore why traditional Authn/Authz and container boundaries fail when software behaves unpredictably, and showcase how emerging open-source standards are trying to redefine multi-tenancy and secure execution for the next decade of cloud-native computing.
Agentic AI is moving from experiments to production, and Kubernetes is a natural place to think about how these workloads are run. Let's look at how open source can provide a common platform layer for agentic systems, allowing frameworks, orchestrators, compute substrates, and identity infrastructure to come together behind declarative APIs. The goal is a more consistent way to deploy, schedule, secure, and observe AI agents using the cloud native model teams already know.
In order to facilitate networking and business relationships at the event, you may choose to visit a third party’s booth or access sponsored content. You are never required to visit third party booths or to access sponsored content. When visiting a booth or participating in sponsored activities, the third party will receive some of your registration data. This data includes your first name, last name, title, company, address, email, standard demographics questions (i.e. job function, industry), and details about the sponsored content or resources you interacted with. If you choose to interact with a booth or access sponsored content, you are explicitly consenting to receipt and use of such data by the third-party recipients, which will be subject to their own privacy policies.
Managing connections in a Kubernetes cluster can be tricky. This can get even more difficult when you are trying to manage connections for a fleet of Kubernetes clusters, each with users that require strict network isolation at the application level.
In this talk, you will hear a case study in how GEICO Tech solved the problem of restricting network connections at scale by using Cilium Network Policies to form a broad network boundary and using Istio Authorization policies to provide zero trust networking at the application layer. It goes into detail on the issues that we encountered supporting many distinct workloads with unique networking needs and how we managed the tradeoffs between user flexibility and system performance, all while maintaining the overall security of the system.
Kubernetes environments generate enormous telemetry volumes, yet most AI observability integrations treat logs and traces as raw text — producing hallucination-prone, unreliable incident responses.
This session presents a production-validated four-layer architecture built on CNCF-native tooling: OpenTelemetry for signal collection, a Model Context Protocol (MCP) server as a typed semantic query layer, and AI-driven agents for proactive root-cause analysis. The speaker shares production lessons from running this architecture at petabyte scale across multi-region Kubernetes deployments at Workday.
Attendees will understand why raw OTel pipelines fail with LLMs, how an MCP server exposes telemetry as structured, queryable context, and the cardinality and latency tradeoffs encountered in production. The session closes with a code walkthrough of a lightweight MCP server integrating with Prometheus or Mimir, and practical criteria for when agentic observability genuinely reduces MTTR.
What do our 500 engineers and our AI coding agents have in common? They run on the same Kubernetes cluster.
At monday.com, our engineering org grew 30-60% year over year. Per-developer cloud environments couldn't keep up. We moved 500+ engineers to one shared staging cluster with pod-level isolation, cutting setup from 30 minutes to 10 seconds and eliminating $450/month per-developer environment costs.
Then we added AI agents. Running Cursor's self-hosted agents against the same cluster, we found that agents need exactly what developers need: real traffic, real env vars, real downstream dependencies. Not mocks. The isolation model that stopped human developers from colliding does the same for concurrent agents.
In this talk, we share the infrastructure behind both and the core insight: developer environment problems and AI agent reliability problems are the same problem.
KEDA has supercharged Horizontal Pod Autoscaling with its support for 60+ scalers and ability to scale workloads to and from 0. With KEDA, Stripe runs one of the largest autoscaled fleets in the world. We have seen the great, the good, and the ugly.
In this session, we go beyond how to autoscale with KEDA and discuss how we integrated and operate KEDA at Stripe while maintaining our 99.999+% SLA.
Attendees will deep dive into:
Observability and Operability – Identifying and resolving issues within your autoscaling infrastructure.
Scalability and Abstraction – Scaling reliably from 100 to 2,000 ScaledObjects per cluster, and 10,000+ globally, with a delightful configuration experience.
Safety – Adding guardrails during CI/CD to safeguard against service degradation.
Whether you run a small organization or support a large engineering team, you will leave this talk with learnings that will empower you to safely adopt autoscaling without sacrificing your reliability commitments.
Last year at KubeCon, we showed how MCP and LLMs let you talk to your dashboards. The next step is agents that investigate incidents autonomously. An agent queries Prometheus, reads logs, and produces an RCA, but the metric it cites does not exist, the service dependency it assumes has no edge in the mesh, and the deploy it blames was in another namespace. The diagnosis looks expert and is fabricated.
We call these observability hallucinations: agents inventing metrics, service relationships, or causal timelines. Unlike generic LLM hallucinations, these are detectable because ground truth exists in your infrastructure. Prometheus metadata knows every metric. Kubernetes discovery knows every service path.
This talk presents three hallucination types: metric, topological, and temporal – with a detection architecture that validates claims. We will demo how MCP tool-call interception, evidence grounding, and a self-correction loop correct analysis before it reaches the on-call engineer.
With Ingress NGINX read-only since March 2026, teams are migrating to Gateway API. On Envoy Gateway, one resource becomes four: HTTPRoute plus traffic, security, and client policies. Each can silently dangle when its parent disappears. No error, no warning.
This isn't a bug. Gateway API maintainers acknowledge it: policy attachment is opaque by design. The spec delegates composition to you.
We'll break a real cluster deliberately, exposing why silent dangling is hard to catch in review. Then we'll fix it with Kubernetes Resource Orchestrator's (KRO) ResourceGraphDefinition (RGD): a platform team owns the full graph (HTTPRoute + policies, CEL conditionals, deletion cascades), and app teams apply a single Custom Resource.
We'll compare Helm, Crossplane, and hand-written operators, not to declare a winner, but to give you a framework for when an RGD is enough and when you need a full operator.
You leave with a reusable RGD and a concrete answer to the question the spec left open.
Building agentic AI systems means juggling API keys for LLMs and MCP-connected tools, plus the burden of configuring those services correctly. Add expensive inference calls and highly customized dev setups, and the inner loop quickly becomes inner hell.
This session will demonstrate a streamlined inner-loop experience for agentic AI. We will develop, test, and iterate on multi-agent systems on our laptop, then deploy to Kubernetes with zero code changes. We'll show how to auto-provision local LLM inference, build agents with typed tool contracts, compose workflows, and test interactions with hot reload. We'll also show how MCP and/or A2A enable portable integration and separation of concerns, and how teams can switch to production-grade inference through configuration alone, with Kubernetes-native health checks, observability, and deployment patterns.
You'll leave understanding how to build and ship AI systems with open-source tooling and cloud native patterns you already know.
Kubernetes networking simplifies distributed systems, yet many engineers still struggle to understand how abstractions map to real networking behavior, which raises the question: is Kubernetes networking too complex, or poorly taught? As the ecosystem matures, community-driven efforts like the Certified Kubernetes Networking Engineer(CKNE) exam are emerging to help standardize the real-world Kubernetes networking skills.
This panel brings together CKNE SMEs, networking specialists and educators to challenge how Kubernetes networking is understood, taught, and operationalized today. We’ll focus on where mental models break down in practice, why engineers struggle to reason across services, DNS, CNIs and Linux networking stack and what foundational networking knowledge Kubernetes practitioners actually need.
Join us as we explore the invisible layers powering Kubernetes and how the community can build a more practical and standardized mental model for Kubernetes networking expertise.
Last KubeCon, we shared "1000 Clusters, 1 Brain" marking our exciting leap into AIOps. Since then, that brain had to grow up fast! An AI just executing tasks is not enough. It absolutely must collaborate. In this sequel, we dive into evolving basic AIOps into a dynamic system of multiple agents acting as conversational bots for direct troubleshooting.
What started in one group is catching fire. We made the core platform fully reusable across teams. We will showcase how we run strict evaluations, measure key metrics, and track real progress. Working alongside AI is rapidly becoming our operational standard. We are bringing you the raw story of what worked and failed. We will discuss revamping our content strategy, forging tight feedback loops, and harnessing Knowledge Graphs for deep context to improve accuracy.
Expect a transparent look at our real struggles and big wins. Let us explore what happens when the brain gets a whole lot smarter.
In Kubernetes, managing trust distribution and certificates for workloads remains challenging. Trust anchors are often replicated via ConfigMaps, while sensitive material ends up in Secrets, creating fragile setups that are difficult to operate securely.
This talk focuses on two emerging features that address this problem: PodCertificates and ClusterTrustBundles.
We will walk through how certificates are issued, rotated, and projected into pods, and how workloads obtain short-lived, identity-bound credentials with PodCertificates. We’ll then cover how ClusterTrustBundles distribute trust anchors within the cluster and how they differ from manually managed CA bundles. Using concrete examples, we will demonstrate implementing a certificate signer and integrating it into a demo application for an end-to-end authentication flow.
Attendees will gain an API-level understanding of these primitives and how to use them to build practical trust inside Kubernetes clusters.
Your IDP — Backstage, ArgoCD, Crossplane, Kyverno – is a chain of dependencies & every single link will break. When it does, your devs stare at a loading spinner & your platform investment becomes shelfware.
Nobody talks about this.
Every conf talk is about building golden paths & shipping plugins.
Nobody asks: what’s your Backstage DB backup strategy? When did you last rotate your Vault tokens? What happens to ArgoCD when GitHub returns 401s for 3 hours (it will)? An expired TLS cert took down all Git operations for an hour. Can your Crossplane providers recover from expired cloud credentials?
This talk is about the boring, unglamorous, career-saving work of keeping IDPs alive.
This hands-on workshop will introduce a single failure that cascades through your version control, portal, GitOps pipeline & ultimately – your confidence. Then we'll recover together. Real failure, real recovery — not just slideware & wishful thinking. This failure will haunt you all the way through lunch.
The Kubernetes community has been hard at work enhancing the scheduling of workloads. In 2026, the community started adding support for workload aware scheduling and this work was eventually folded into a working group. In this session, we will highlight the APIs added, provide details on the in-tree roadmap and the goal of the working group with respect to the ecosystem.
We will discuss the integrations with Kueue and popular workload controllers like JobSet. This session will provide the roadmap and future plans for the working group. We are also looking for engagement from the community.
Kubernetes SIG Storage is responsible for ensuring that different types of file and block storage are available wherever a container is scheduled, storage capacity management (container ephemeral storage usage, volume resizing, etc.), influencing scheduling of containers based on storage (data gravity, availability, etc.), and generic operations on storage (snapshotting, etc.). SIG Storage also has a project that provides APIs for object storage support in Kubernetes. In this session, we will deep dive into some projects that SIG Storage is currently working on, provide an update on the current status, and discuss what might be coming in the future.
Join containerd maintainers for project update and deep dive into the latest updates on containerd. LLM-assisted vulnerability discovery has had a big impact on containerd: we'll discuss the impact on the project and changes we've made to improve our security response. For Kubernetes users, we will cover how the new containerd release cycle, support window, and upgrade safety changes make feature availability faster and make long-term support available across Kubernetes versions. We’ll then dive into the exciting work going on in the containerd ecosystem, including further developments in the Node Resource Interface (NRI) and EROFS snapshotter.
SPIFFE has always issued identity to Pods, it gives each a stable identifier and verifiable credentials. But plenty of things that need identity don't run at all: a CRD that projects remote state, a Secret mirroring an external credential stored in Vault, or a ConfigMap that replicates remote config. Today they inherit the coarse identity of whatever controller manages them, or need dedicated Service Accounts. The SPIFFE Broker API changes that: a trusted broker like a Kubernetes operator can request a short-lived SVID on behalf of any object it represents. Each object receives a dedicated SPIFFE ID and can be given fine-grained access, for instance to a secret vault or a code repository.
Curious how kagent brings agentic AI to the cloud native ecosystem? Interested in contributing to kagent ? Want to explore the codebase and understand what powers the kagent framework? Join the kagent maintainers for a hands-on contributor session where you'll learn how the project is built and how you can become part of its growing open source community.
We'll walk through the architecture of kagent, explore how agents and MCP tools work together, show you how to set up a local development environment, explain the project's community and contribution workflow, and help you get started with your first contribution. Whether you're new to the project or looking to become a regular contributor, you'll leave with the knowledge and confidence to make your first PR to kagent.
Argo CD is one of the most widely adopted GitOps tools, yet there are fewer opportunities to learn how teams are actually using it in real-world environments. This ContribFest session shifts the focus from feature walkthroughs to understanding practical Argo CD usage patterns, repository structures, and operational approaches used by the community.
This will be an interactive, discussion-driven session where attendees are encouraged to come prepared to briefly share how they use Argo CD today. Topics may include GitOps repository layouts, multi-cluster setups, promotion strategies, and common challenges encountered in day-to-day operations. Argo CD maintainers will help guide the conversation and ensure broad participation.
To kick off the discussion, maintainers will highlight commonly observed patterns and frequently requested Argo CD features. The insights gathered during this session will help inform future documentation, blog posts, and contributor initiatives.
At Coinbase we run hundreds of Kubernetes services in a multi-cluster, latency-sensitive financial-services environment, where every pod historically ran an auth-agent sidecar verifying an in-house S2S bearer token on every request. Over the past few years we've migrated to SPIFFE/SPIRE-issued identities, extended our ingress gateway to validate them at the edge, and used Istio AuthorizationPolicies to enforce explicit service-to-service access. We're now deleting that sidecar entirely, replacing per-pod JWT verification with ztunnel mTLS using SPIFFE X.509s, delivering over $500K/year in CPU savings, faster service calls, and a cleaner compliance posture.
We'll share the architectural decisions, the rollout sequence across many clusters, the hardships and learnings, and where we're heading next with pod-to-pod mTLS. Whether you're considering SPIFFE for your workloads or weighing how to push identity into your mesh, this talk offers a peer's view of a real production migration.
Cloud provider bills show egress costs as a single line item-they don't say which tenant caused them. At Adobe, we answer that across 400K pods in 31K namespaces spanning 85+ multi-tenant Kubernetes clusters using Retina eBPF agents to capture pod-level flows at the kernel. Prometheus enriches them with node topology, and a CIDR-based classifier maps each destination IP to a cost category (Local, Internet, Peering, Transit Gateway, VPC Endpoint) with cross-AZ and cross-region modifiers that align to cloud provider billing.
The hard part is scaling Prometheus to handle high cardinality from per-flow metrics. We'll walk through our sharded deployment implementation-shard counts, memory limits, and query timeouts along the way.
This is not a victory lap. The pipeline runs in production but scaling remains open. We'll share our progress and invite community input.
Attendees leave with a replicable architecture, real cardinality data, and an honest view of what works and what doesn't.
For years, Kubernetes treated individual pods as scheduling units, complicating the execution of complex distributed systems. With the explosive growth of AI workloads, the ecosystem demands a smarter control plane.
In this panel, scheduling experts from Google, Anthropic, Anyscale and CERN will unpack the paradigm shift of Workload Aware Scheduling. By teaching core Kubernetes to natively understand higher-level application semantics, we are bridging the critical gap between the scheduler and higher-level frameworks. This upstream evolution is essential for optimizing resource sharing and performance in massive multi-tenant environments.
Our panelists will discuss how these foundational blocks enable a cohesive multi-layer scheduling stack. Furthermore, we will explore how this crucial first step unlocks the next frontiers of orchestration including seamless disruption management and tighter integration with infrastructure autoscaling.
Your AIML customers need more GPUs, your providers are at capacity, the C-Suite is demanding ROI, and inference is mission critical under constant cyber threat. We’re the platform infrastructure team at a regulated Fortune 100 insurance company, where we build and operate the Hybrid Cloud Fabric: multi-provider k8s and zero trust overlay on Istio, Cilium and SPIFFE/SPIRE. Our story covers dynamic shifting of GPU workloads between Azure and AWS, the rehoming of inference workloads onto OSS models, and running a global SPIRE deployment. Four nines isn't easy. We'll share where we came up short, alongside our pluggable SPIRE backend proposal (spiffe/spire#5993) and reference Cassandra implementation we’re running in prod and contributing back. Finally, meet our sharp edge: secure attestation on non-traditional hardware like DGX Spark. We’ll let you know how this lands. Walk away confident you can scale beyond the limitations of any single cloud.
DuckDB is one of the fastest analytical databases available today, but wasn't designed to run as a multi-tenant, cloud native service.
At Greybeam, we’ve been running DuckDB in production since before v1.0, serving millions of queries while migrating large-scale customer workloads away from traditional data warehouses. There are tradeoffs when choosing an embedded single-node engine vs. a distributed database, and this talk will cover the challenges of building a DuckDB-based data platform on Kubernetes. Hear about the operational realities of running a stateful engine, such as NVMe provisioning and running performant data systems in environments without swap memory. We’ll then touch on what’s next, such as object-store-native architectures.
Attendees will leave with practical lessons for running stateful workloads on Kubernetes, the tradeoffs of an embedded database, and patterns for building scalable data systems on top of modern open source infrastructure.
Many cloud native stories begin with greenfield systems. This one begins with a 20-year-old production monolith: 500+ tightly-coupled services, 1,000+ database shards, and a fleet of 5,000+ statically provisioned servers.
This case study follows a migration at Cisco Meraki from a citadel-and-outposts architecture to an API-first Kubernetes platform. The team treated Kubernetes not as an infrastructure swap, but as a new platform contract between infrastructure teams and application teams. That shift broke assumptions that had become core to application design, forcing hard decisions about what the platform should abstract away and what should be solved in partnership with application teams.
Attendees will leave with practical patterns for decomposing brownfield architecture, defining platform boundaries and the application-owner-facing interfaces that enforce them, and managing the human side of retiring assumptions that have been reliable for decades.
A team ships an AI-powered support feature on a Friday. By Monday, token spend is eight times the forecast, a quality regression is annoying a swath of customers, a prompt injection slipped past the output guardrails and the comms team is dealing with a viral screenshot of the response.
Shipping AI features responsibly is an experimentation problem. The cloud native community already has the primitives (flags, cohorts, telemetry, experiments, kill-switches), and AI workloads need them more than a CRUD service. Non-determinism means the same input can produce different outputs, so a green pre-prod eval is evidence of nothing; real traffic, real prompts, and real adversaries are the only meaningful test data. When a one-line prompt change can 10x a tenant's bill, progressive rollout and kill-switches cease to be optional.
Todd covers the three axes (quality, cost, risk) and a reference loop using OpenFeature, flagd, and OpenTelemetry attendees can build next sprint.
As Kubernetes platforms scale to multiple clusters, teams often solve cross-cluster connectivity by adding cloud load balancers between clusters. This approach works initially but leads to various issues, including, but not limited to, load balancer sprawl, cloud quota pressure, and connection-level load balancing.
This talk presents an alternative: a lightweight xDS-based service discovery control plane that extends single-cluster endpoint discovery to span multiple clusters. Rather than requiring a full service mesh or relying on connection-level load balancing through cloud load balancers, the system distributes pod-level endpoint information across clusters through peer-to-peer xDS subscriptions. This architecture preserves the same routing capabilities and semantics available for single-cluster traffic, and allows systems to consistently apply features like request-level load balancing and zone affinity.
You’ve probably heard a lot about Dynamic Resource Allocation (DRA), which is quickly becoming a de-facto standard for allocating hardware accelerators in Kubernetes. DRA Driver for NVIDIA GPUs is one of the early implementations of this and is evolving far beyond simple GPU allocation. This talk is a deep dive into the latest capabilities that we have added this year and what comes next for supporting large-scale AI infrastructure.
We will cover the latest on GPU sharing models including MIG, time-slicing, MPS, and vGPU along with topology-aware scheduling across GPUs, NICs, and Data Direct (DMA) interfaces. We will also share about device health monitoring, running sandbox workloads, integration with the GPU Operator and lastly how existing workloads can migrate to use DRA.
Using real examples from GB200 and next-generation Vera Rubin NVL72 systems, we will discuss the infrastructure challenges DRA solves using Compute Domains and where the accelerator ecosystem is headed next.
MCP made our developer platform visible to agents, but that turned out to be the easy half, because listing tools is not the same as teaching an agent how a platform works. The model still has to know which tool to use, what has to happen first, and what the nouns in your product actually mean.
We shipped MCP for OpenChoreo, our CNCF Sandbox developer platform, and watched it fail in normal use: the agent picked the wrong tool, looped on broken paths, and skipped steps it had no way to know about.
Agent Skills closed most of that gap, but not every Skill helped, long prose buried the useful context, API-reference-style Skills mostly got ignored, and the ones that worked were short and shaped around decisions, closer to how you would onboard a new teammate than how you would write public docs.
This talk is an experience report from doing that work, what we tried, what we deleted, which patterns held up under real agent loops, and how we keep Skills honest in CI.
Over the last year, SIG Scheduling has focused on advancing the foundational Workload-Aware Scheduling API and making it generally available. The capabilities have been expanded to encompass topology-aware and multi-level group scheduling. Furthermore, a new Working Group (WG) dedicated to WAS has been established to focus on further development and integration with controllers, allowing SIG Scheduling to focus its efforts on closer integration with autoscalers.
During this session, we will explore the major scheduling advancements introduced in recent releases and examine how they align with the broader Workload-Aware Scheduling framework. We will also share key updates from related sub-projects, including Kueue and the Descheduler.
To conclude, we will discuss the roadmap for upcoming releases, outline community contributions, and detail opportunities for you to help shape the future of Kubernetes scheduling.
The Kubernetes AI/ML workload orchestration ecosystem is maturing from a collection of isolated tools into a unified, high-scale computing suite.
In this session, WG Batch maintainers will present a comprehensive status report on our interconnected projects, including Job, Kueue, and JobSet. We will highlight What's New in recent releases, focusing on core API stability, advanced fair-sharing algorithms, and seamless multi-project integrations that power modern AI platforms.
Looking ahead, we will preview the working group's plans, revealing our long-term vision for scalability and resource optimization through 2027. Whether you are a platform engineer or a prospective contributor, this session equips you with insights into the future of cloud-native AI/ML workloads orchestration.
The Rook project will be introduced to attendees of all levels and experience. Rook is an open source cloud-native storage operator for Kubernetes, providing the platform, framework, and support for Ceph to natively integrate with Kubernetes. The panel will discuss various scenarios to show how Rook configures Ceph to provide stable block, shared file system, and object storage for your production data. Rook was accepted as a graduated project by the Cloud Native Computing Foundation in October 2020.
SIG-Multicluster is focused on solving common challenges related to the management of many Kubernetes clusters, and applications deployed across many clusters, or even across cloud providers.
In this session, we'll give attendees an overview of the current status of the multi-cluster problem space in Kubernetes and of the SIG. We’ll discuss current thinking around best practices for multi-cluster deployments and what it means to be part of a ClusterSet. Then we’ll highlight current SIG projects, focused use cases, and ideas for what’s next.
Most importantly, we’ll provide information on how you can get involved either as a contributor or as a user who wants to provide feedback about the SIG's current efforts and future direction. Bring your questions, problems, and ideas – help us expand the multi-cluster Kubernetes landscape!
A TPU on the wrong torus coordinate is worthless. So is one GPU placed outside its NVLink island. For frontier training, where a pod lands matters as much as whether it lands at all. And they have to land as a gang: every pod in the job on contiguous hardware, or none at all. The default scheduler scores pods one at a time and has no vocabulary for "all 256 of these, on contiguous hardware, or none."
This is Anthropic's account of building the scheduler that does and how we kept it fast as clusters grew toward GKE's 65K node limit. Part one is design: modeling interconnect topology as first-class scheduling domains, gang admission with graceful eviction, capacity allocation, and defragmentation. Part two is the performance story – the optimizations that brought average processing time to under 3ms per job: gating expensive affinity checks behind cheap predicates, efficient resource accounting, and caching infeasibility verdicts.
In 2026, a developer asked a burrito chain's chatbot to reverse a linked list before ordering their meal. The bot delivered a solution, then asked what they wanted.
A user does not need to be malicious to embarrass an enterprise, just a little bit weird. The model will help, enthusiastically. But that burrito bot sits in a production cluster that someone (probably you) is responsible for, wired into real tools and data. More than reputation is at stake.
In this session we build a chatbot for a fictional burrito chain connected to an under-protected MCP server and deploy on a full IDP (internal developer platform). It starts wide open, and we let the audience hack it live.
We show you the relevant Kubernetes security layers you should already have. Then we turn on LLM-specific guardrails, like pre- and mid-inference prompt analysis. Finally we show good GenAI observability practices. By the end the chatbot is selling burritos. Nothing more.
Bring your weirdest off-topic question.
To power our company's mission to "solve all disease”, we built a massively scalable ML inference platform on Kubernetes by rejecting heavy industry-standard frameworks in favor of a minimalist worker-and-queue architecture. Reframing ML serving as a classic distributed computation problem rather than a unique paradigm allowed us to eliminate complex secondary ML schedulers and feature bloat. Instead, we composed native, scalable components anchored on the simplest primitives: a stateless worker pod and an asynchronous queue.
This talk breaks down the mechanics of our "frameworkless" infrastructure. We will detail exactly how we deploy thousands of unique models as stateless pods, leveraging GCP Pub/Sub for asynchronous message routing, KEDA for true scale-to-zero autoscaling, and native K8s pod priority and preemption to guarantee real-time interactive research traffic over massive background batch workloads. Learn the concrete design principles behind building less to scale more.
Most workload platforms force users to choose early: stateless or stateful. While useful labels, they are the wrong API boundary.
Workloads don't live neatly in one bucket. Each has different attributes, which are independent concerns. Some need rolling updates. Others must coordinate with a shard manager before a pod can safely move. Some need persistent volumes for source-of-truth data, while others use them as cache.
Users shouldn't need a new platform every time a dimension changes, nor should they care which cluster serves their compute needs.
This talk shares how LinkedIn built a composable, natively multi-cluster API for both stateless & stateful workloads. Users express intent once: the platform places & rebalances replicas globally. Treating pod template, update strategy, persistence and placement as independent concerns allows us to abstract cluster boundaries, enables rainbow deployments, improves global efficiency, and preserves required workload guarantees.
Every vendor is shipping an agent platform. You probably don't need one — what you need to know is how to compose the CNCF projects you already are running to get headless agents running safely in your clusters.
We identify challenges: cold-start identity gap (the agent egresses before SPIRE issues an SVID), trust-scope drift (egress allowlist decided at deploy, stale by wake), unverified delivery (the event broker hands off events before signature check), and unbound audit (completion not chained back to the signed wake).
We'll walk through one reference pattern: KEDA scaling deployments from queue depth, signed CloudEvents (SPIFFE source + JWS) as the wake contract, sidecar policy reading SVIDs at runtime, and the event broker log as tamper-evident audit. In live demo we show: zero agent replicas -> signed task event to agent wake -> blocked egress on a wrong-scope call -> completion bound to the wake event -> scale to zero when agents are idle.
Kubernetes gave the industry a powerful model for infrastructure: versioned APIs, declarative resources, and clear compatibility rules. We turned to those same ideas after hitting the limits of bespoke APIs and service-specific backend logic in Grafana.
This talk shares what changed when we applied those ideas to Grafana by modeling functionality using standard Kubernetes-style extension mechanisms, and, in doing so, evolved a system used by 25M+ people where existing workflows are heavily relied on. We'll show how we used custom resources and controller-based workflows where a resource model fit naturally, and custom routes or aggregated API servers where it did not.
Using our 3-year journey to a stable dashboards v1 API as a case study, attendees will leave with a rubric for when these patterns are worth the investment, how to handle the parts of an application that do not fit neatly, and what this model unlocked for infrastructure-as-code workflows and new product capabilities.
GenAI is moving fast, and the fragmentation of AI observability is moving just as fast. Vendors, frameworks, and platforms all capture GenAI telemetry differently. Many advertise OpenTelemetry support, but that can mean anything from OTLP export to SDK integration to partial semantic convention compliance. Without consistency across all layers, telemetry is portable in theory but inconsistent and difficult to use in practice.
This talk demonstrates how OpenTelemetry is fighting that fragmentation. The session covers how GenAI semantic conventions and vendor-agnostic instrumentations are developed, how designs are validated, how conformance testing keeps the ecosystem honest, and how AI itself is utilized to accelerate the work without sacrificing rigor. The ultimate goal is to make GenAI telemetry portable, comparable, and useful across all tools.
Attendees will leave with a clear understanding of what genuine OpenTelemetry support for GenAI actually entails.
Time-slicing is often the only viable way to share GPUs on commodity hardware or in cloud environments where MIG and MPS are unavailable. In practice, Kubernetes time-slicing relies on static configuration and does not constrain GPU memory, leading to noisy neighbors, unpredictable performance, and OOM failures in inference workloads.
This session shows how to close that gap by combining memory-aware controls with smarter pod placement. Using KAI for GPU-aware scheduling and HAMi for per-workload memory limit enforcement, we demonstrate how to prevent memory overcommit, stabilize performance, and safely share GPUs across inference workloads. Under the hood, Kyverno mutates pod specs to inject KAI and HAMi configuration on inference pods that should share GPUs.
This talk shows how to get safe, OOM-free GPU sharing on standard Kubernetes clusters without hardware upgrades. Just dynamic scheduling and enforced memory limits that work with the GPUs you already have.
At NVIDIA, we run hundreds of GPU clusters, and ever so often, they need maintenance: node drains, driver upgrades, firmware flashes, or Kubernetes upgrades. At fleet scale, the hard part is coordinating and keeping track of multiple upgrades in parallel without disrupting workloads or wasting GPU capacity.
We’ll present a GitOps-based system we’ve built at NVIDIA to automate and coordinate this. Each cluster’s desired state for every piece of software lives in Git. Each cluster reports drift, and a central scheduler decides what to upgrade and when — factoring in workload sensitivity, GPU usage, and business-critical windows.
You’ll leave knowing how to prioritize competing maintenance work, how to fold GPU usage and tenant sensitivity into scheduling decisions, and how this work connects to in-flight upstream efforts.
AI agents are evolving into autonomous engines that execute code, but running untrusted, LLM-generated code in multi-tenant Kubernetes clusters creates security and operational hurdles.
The open-source Kubernetes SIG Apps project, Agent Sandbox, manages isolated, stateful agent workloads. It standardizes manual infrastructure patterns into a declarative API for secure, lightweight container runtimes.
In this hands-on tutorial attendees will go from zero to deploying a production-ready, highly secure, programmatic agent sandbox platform on Kubernetes.
During this hands-on session, we will cover:
– Architecture Deep Dive: Mapping out the core primitives—exploring Sandbox, SandboxTemplate, and SandboxClaim CRDs.
– Eliminating Cold Starts: Configuring SandboxWarmPool to provision and claim sandboxes in milliseconds instead of seconds.
– Programmatic Execution: Interacting with the cluster dynamically using the official Python and Go SDKs to showcase real-time agent code execution.
Backstage has spent the last several years rebuilding its frontend and backend foundations. Those systems are now mature enough to change not only how we build Backstage, but what can be built with it.
This session explores Backstage’s role in AI-assisted development in both directions. Platform engineers can use AI-assisted coding to create extensions against an isolated, well-defined architecture. At the same time, Catalog AIResource entities and MCP servers let agents discover trusted organizational context, internal skills, and governed actions, making Backstage a shared foundation for a new generation of developer tools.
Through a concrete developer workflow, we’ll show an agent using Backstage context and actions to complete a task. You’ll leave with a clear view of what Backstage’s mature architecture makes possible today and where the project is heading next.
TAG Workloads Foundation defines practices and standards for how cloud native systems execute and manage workloads: from container runtimes and schedulers all the way through to lifecycle management at scale.
Come learn about projects tackling GPU sharing, VM-based workloads, batch workloads, and secure compute, and meet the people behind them. You'll also hear about ongoing work to document how to run scientific AI on Kubernetes at scale.
Whether you run scientific computing, manage GPU clusters, or build infrastructure for AI teams, you'll leave with a clear picture of what cloud native workload execution actually looks like across the full stack and where to plug in.
OpenFeature is an open specification that offers a vendor-agnostic, community-driven API for feature flagging that works with your favourite feature flag tool or in-house solution.
This year the project reached a major milestone: the arrival of OpenFeature-native vendors. Providers including Cloudflare, Google Cloud, Dynatrace, and Datadog now build directly on OpenFeature, moving the standard from something teams adopt on top of their tools to something the tools themselves are built around. We'll dig into what this means for adopters and the ecosystem.
From there, the maintainers cover the rest of what's new: isolated SDK API instances, Python SDK 1.0, the new C++ SDK, growing maturity in the Kotlin and Swift SDKs, the latest flagd updates, and OFREP's includsion of server-sent events (SSE) support. We'll also go behind the scenes of our community and share the lessons we've learned.
Bring your questions and curiosity as we explore the current state and future of OpenFeature.
Whether you have a passion for Agentic software development or want to learn something new, the Helm project is looking to make all phases of Helm use easier through incorporating agentic resources, like skills. Join this interactive hackathon where you or a team of individuals work towards developing agentic based tooling focused on any aspect of the Helm ecosystem: from software development to daily usage – and everything in between.
Submissions will be judged by a group of Helm project maintainers and the winner will have the opportunity to partake in unique Helm project experiences including:
- Present at a Helm Community meeting
- Publish an article on the Helm Blog
Don’t miss out on building community relationships and expanding the agentic capabilities for the Helm project!
Meshery is a CNCF cloud native manager for designing, deploying, and operating infrastructure across essentially every CNCF project. In this hands-on Contribfest, Meshery maintainers get you from clone to merged pull request in one session.
We've staged good-first-issues across four parallel tracks, each with an embedded maintainer: Server (Go) for core API work, UI (React) to extend the Kanvas visual designer, Plugin development to build your own Meshery extension, and Docs and Designs to contribute real cloud native architecture examples.
Ahead of the session we'll publish a setup guide and a labeled issue board, so your environment is ready and issues are sorted by difficulty. We'll walk the contribution workflow end to end, then pair you with a maintainer to land a change.
No prior Meshery, Go, or React experience needed. Come build with the maintainers and join the community. Don't forget to bring your laptop, we hope to see you there.
Uber’s AI/ML Platform Team is relying on Cadence as part of their new Agent Executor Service, facilitating durable agentic workflow orchestration at scale.
Large Language Models (LLMs) have evolved from simple text prediction to autonomous agents capable of orchestrating complex tools. However, a critical gap exists in the current ecosystem: durable orchestration. Most agent SDKs are stateless by design, making them fragile in production environments where network partitions, long-running tasks, and system failures are inevitable.
In this session, we explore a cloud-native architectural pattern to solve the “stateless agent” problem by using Cadence as a robust durability layer for agent orchestration. We will move beyond the “hello world” agent and demonstrate how to build production-ready systems that require:
- Fault tolerance: Resuming complex reasoning chains after a process crash.
- Auditability: Maintaining a permanent, replayable record of agent decisions and tool executions.
- Long-running state: Managing agents that operate over days or weeks rather than seconds.
We will provide a technical deep dive into integrating Cadence with popular frameworks like OpenAI, Google ADK, and Claude. We will also show how Cadence fits into the broader AI ecosystem, including Cadence AI Skills and AI-native coding tools like Claude Code, Cursor, Codex, and Hermes Agent, so developers can write their first Cadence-aware agent workflow in minutes.
Attendees will walk away with a blueprint for durable agent orchestration: implementing automatic checkpointing and retries, recovering from failures instead of restarting, and ensuring their AI agents are as reliable as their microservices.
Moving agentic AI into production requires more than just additional compute; it necessitates a data-driven approach to capacity planning and performance tuning. At Shopify, we deploy fine-tuned open models with 250k token context windows to power complex, long-form multi-turn agentic conversations for products like Pulse and our GraphQL sub-agent. Serving these models at scale requires precise benchmarking to navigate the trade-offs between latency, throughput, and GPU parallelism.
In this talk, we demonstrate how Kubernetes Inference Perf serves as the engine for Shopify's agentic benchmarking workflow. We will share a reusable blueprint for load-testing fine-tuned models under realistic workloads, right-sizing GPU capacity, and translating performance data into production-ready deployments. Attendees will leave with a reproducible framework for benchmarking and deploying their own open-models, ensuring predictable performance for the next generation of agentic applications.
AI is moving into the operating core of communication networks: control loops, spectrum allocation, network slicing with QoS, and customer-issue automation, under regulatory watch on life-critical systems. Opaque models and unreproducible outputs do not meet regulatory legal bar. This session presents a deterministic governance pattern for AI-RAN, drawn from work with the FCC. We separate AI-for-Networks (internal optimization) from Networks-for-AI (multi-tenant hosting), and show how one hybrid-cloud platform carries both. The core is a five-stage pipeline: intent, retrieval, rule-bounded reasoning, synthesis, citation-validated output. Every output is reproducible and audit-traceable.
We map it onto a six-layer governance model using OPA or Kyverno for policy, OpenTelemetry and Kepler for energy telemetry, and GitOps for model, rule, and corpus lifecycle. We close with a maturity model and a measurement checklist that separates real AI-RAN energy savings from selective accounting.
Internal developer platforms have a discovery problem: the capabilities are there, but engineers context-switch between tools to get things done. For non-engineers, it gets worse.
Vela is our answer: an agentic chat interface on top of Karavela, our Backstage-based IDP, turning platform capabilities into natural-language workflows for engineers, PMs, data analysts, and on-call responders who've never touched kubectl — all in a regulated financial company.
We'll share what we learned building the MCP Hub and how its design makes democratization safe: every tool call runs under the requesting user's permissions, with a full audit trail, no service-principal bypass.
We'll be honest about the hard parts: intent classification on ambiguous Portuguese, consent mechanics for write-on-behalf flows, and keeping long agentic tasks recoverable without burning unbounded LLM budget.
The goal isn't to replace platform engineers, it's to make platform knowledge accessible to everyone.
Open-weight models are rapidly closing the capability gap with the frontier, the friction to run models independently has vanished, making it possible to accomplish more locally than ever before. This session details how to build commercial-grade AI infrastructure using consumer hardware and cloud-native tooling.
Attendees will learn how to configure local GPU scheduling, securely source private datasets, and deploy containerized inference workloads on Kubernetes. The presentation explores deploying open-source inference engines like vLLM and llama.cpp, alongside fine-tuning tooling to run large language models on-premises. The key is to bring these datacenter-native workloads into your private infrastructure using the CNCF tooling we know and love.
Autonomous AI Agents running untrusted, dynamic code on Kubernetes require strict sandboxing (e.g., Kata, gVisor) to guarantee tenant isolation. However, platform engineers face two massive production hurdles: securely verifying unvetted Agent images and distributing multi-gigabyte LLM weights to isolated sandboxes under tight latency constraints.
This session explores how Harbor (CNCF Graduated) scales beyond standard containers to become a secure distribution hub for AI infrastructure. We will dive deep into:
1. Models as OCI Artifacts: Storing and versioning massive LLM weights alongside Agent images natively in Harbor.
2. AI Supply Chain Trust: Leveraging Harbor’s scanning and cryptographic signing (Cosign/Notation) to ensure only verified models execute within sandboxes.
3. Latency Optimization: Utilizing proxy caching and lazy-loading to slash sandbox startup times.
Attendees will walk away with a production-ready blueprint for a secure, high-performance AI platform on K8s.
GPU hardware is scarce, expensive, and hostile to fast CI/CD. Yet every component in the NVIDIA Kubernetes GPU stack, device plugins, DRA drivers, GPU Operator, depends on NVML, the user-space C library that talks to the driver. What if you could swap it for a drop-in mock that makes nvidia-smi, and any Go binary linked against go-nvml, believe it runs on a DGX A100, on your laptop, in a kind cluster, or in a GitHub Actions runner with zero GPUs?
This talk introduces nvml-mock, an open-source, YAML-configurable mock of libnvidia-ml.so implementing 89 NVML entry points via an auto-generated CGo bridge. We walk the architecture, C shim, a singleton Go engine with handle tables and reference-counted Init/Shutdown, and YAML device profiles for A100, GB200, and custom topologies, then show the same production binary tested end-to-end against simulated multi-GPU nodes. Finally, we show how to build your own GPU-aware CI pipelines without provisioning a single GPU.
A distributed training run dies six hours in. Somewhere, the signals exist: maybe a node logged a GPU error,maybe a pod was evicted,maybe a worker restarted. With kubectl you can pull the events for any object you already suspect. But it wont tell you if those scattered signals are the same incident or that they’re why training-job-z, 3 abstraction layers up, ended abruptly. Across 100s of pods and a 1-hour event TTL, that join is manual toil. This talk distills what it takes to correlate K8s events with AI-framework failures, drawn from building such a pipeline upstream in Ray. Its framework-agnostic by design, the lessons apply to Kubeflow’s training operator, JobSet, Volcano or your own controller. We’ll cover why stable pod identity is the make-or-break join key, why K8s events are lossy by design and must never be your system of record and a source agnostic provider pattern that normalizes heterogenous signals into 1 schema you can route to a dashboard, alert or an OTel pipeline
Proprietary LLM APIs are convenient — and expensive at scale. For a high-volume, multi-label classification pipeline in production, replacing a managed LLM API with a fine-tuned open-weights model on Kubernetes delivered higher quality at one-third the cost.
The harder problem was the traffic profile: steady baseline most of the day, then surges into thousands of requests per second within minutes. Sizing GPU capacity for the peak wastes money; sizing for the baseline drops requests. The solution lives at the intersection of inference-layer optimization and Kubernetes-native autoscaling.
This session walks through the architecture end to end: QLoRA fine-tuning with Unsloth, vLLM for inference and dynamic batching, speculative decoding for throughput gains at no additional GPU cost, and Karpenter, a CNCF project, for real-time GPU capacity right-sizing.
Attendees leave with a practitioner blueprint for elastic, open-source LLM inference.
Distributed coordination — locks, leader election, barriers — sounds solved. But “usually works” and “correct by design” are very different things.
This talk shares lessons from building distributed coordination primitives on etcd. A common failure mode: a process acquires a lock, gets paused by GC or a network partition, then resumes still believing it owns the resource. Without fencing, zombie clients can corrupt shared state.
The solution is a three-party contract: the lock service issues monotonically increasing fencing tokens, the client forwards them, and the resource rejects stale operations. Together they prevent zombie writes by design.
We then zoom out from locks to coordination as a broader discipline. An extensible provider model — etcd today, ZooKeeper and Redis later — lets one abstraction support different stacks.
Live demos cover zombie scenarios over HTTP and gRPC. Attendees leave with a framework for reasoning about correctness in distributed coordination systems.
Data Protection WG is dedicated to promoting data protection support in Kubernetes. The Working Group is working on identifying missing functionalities and collaborating across multiple SIGs to design features to enable data protection in Kubernetes. In this session, we will discuss what is the current state of data protection in Kubernetes and where it is heading in the future. We will also talk about how interested parties (including storage and backup vendors, cloud providers, application developers, and end users, etc.) can join this WG and contribute to this effort. Details of the WG can be found here: https://github.com/kubernetes/community/tree/master/wg-data-protection.
Come to this session to learn about the Open Policy Agent (OPA) project. OPA is a general-purpose policy engine that solves a number of policy-related use cases for Kubernetes, API Authorization, AI Tool Guardrails, infrastructure permissions, and more. During this session OPA maintainers will introduce the project for newcomers and then provide updates on noteworthy new features landing in OPA projects and the wider ecosystem. If you are interested in policy as code and security as it relates to cloud native technology, this session is for you. OPA maintainers will also be available for questions after the session.
Originally developed at Google 10 years ago, minikube has become the most popular local Kubernetes tool with more than 15M github downloads. overview of journey of Minikube from an internal google project to migrating it out of Google ! This talk explores a history of 10 years of minikube and also shared exciting new AI features of minikube, such as using macbook's GPU in minikube.
Falco has never flown alone. Around the runtime security engine at its core, a whole ecosystem has taken shape: the Operator that runs it, the connectors that carry its events, and the projects that turn detections into action. In this maintainer track session, the maintainers will look at how that ecosystem is maturing and where the project is heading next.
We will share the most relevant developments in Falco itself, from detection and performance to a smoother experience for writing and maintaining rules, and give particular attention to the ecosystem growing around it. With the Operator now production-ready, deploying and managing Falco across Kubernetes clusters is simpler than ever, and the projects around it keep expanding what teams can do once an event leaves the engine.
Join the maintainers for a look at what has changed since we last met, an honest view of where the project stands today, and a preview of the directions we are most excited about for the releases ahead.
Over the past 2 years, Airbnb has been on a journey to migrate several million Spark jobs a week from Hadoop to Kubernetes.
Operating this platform at scale has required continuous improvement, from tracking down Spark bugs that leak thousands of configmaps to backporting CRI-O patches on a surprisingly regular basis.
The workload patterns of Spark push the boundaries of every involved system including node health, observability, networking, and all the way down to individual disk performance.
In this talk, we'll step through each layer of the stack — what broke, what we patched, and the custom tooling we built when off-the-shelf solutions couldn't keep up.
As Cloudera Data Warehouse (CDW) workloads were adopted at scale in public and private cloud environments, our monolithic, stateful architecture hit a wall. Managing 40-minute provisioning cycles with a single-server executor meant that infrastructure instability often resulted in stalled workloads and inconsistent states. This session provides a technical deep dive into our migration to Cadence. We’ll discuss the development of the Lease API for distributed locking and how we utilized an intermediary Async Executor to transition to Cadence’s workflow/activity model without a complete rewrite.
The scope of this talk is narrowed to the practical implementation of fault tolerance: how we refactored long-running tasks to be fault tolerant with the help of Cadence. We will present a real-world case study where a severe memory leak and subsequent node rescheduling – which would have previously caused service disruptions – resulted in zero downtime for user workloads.
AI model gateways start simple: route requests to LLMs, centralize access, and log usage. The harder problem comes later – how does an organization turn single-team proxies into shared platform infrastructure without making model access slow, opaque, or developer-hostile.
This case study traces the evolution of LLM access from users sharing long-lived tokens to a multi-tenant gateway platform serving 20+ teams and users managing 200+ models. Each stage came from a real use case: virtual keys and team-scoped access for isolation, self-service onboarding and discovery for velocity, and legal approval workflows for governance. As adoption scaled, the gateway grew beyond highly available model routing into cost governance: budgets, savings attribution, and cost-optimization intelligence.
Attendees will leave with a maturity model for AI gateways on Kubernetes: what belongs in the gateway, what belongs in platform automation, and which tradeoffs appear only after many teams depend on it.
A dental cabinet manufacturer in Texas now runs their corporate website and systems on a globally distributed CDN, autoscaling serverless compute, and a managed NoSQL datastore. They have no engineers. They vibe-coded it.
This story is only possible because of fifteen years of cloud-native work. Managed platforms are opinionated wrappers around primitives the CNCF community made boring: container orchestration, serverless compute, federated identity, managed observability. Vibe coding is the cloud-native community's success story, even if most of its users will never know we exist.
In this talk, we will share what it took to get from prompt to production, what the cloud-native foundation gave us, and where it left edges that still required human judgment. What is "production-ready" when the platform team is an LLM and how should cloud-native builders think about the wave of AI-assisted, non-engineer-built apps now landing on our infrastructure?
Improving LLM serving platforms like llm-d, a CNCF Sandbox project for Kubernetes-native distributed LLM inference, requires testing choices across routing, admission control, flow management, and scheduling. But GPU-cluster experiments take days and cost thousands of dollars, leaving promising strategies unexplored.
We present an AI-assisted evolution loop for inference serving: simulate, evolve, validate, ship. It uses a CPU-based simulator to evaluate serving policies against SLO-oriented objectives, searches over variants offline, validates candidates on real GPU clusters, and translates winners into production-ready llm-d code.
We ground this talk in three llm-d contributions: a probabilistic admission controller that reduced p90 TTFT by 97% under load, a flow-control mechanism, and a prefill-decode disaggregation strategy. Attendees learn to turn exploratory ideas into testable hypotheses, evaluate SLO behavior before spending GPU time, and apply this to any LLM serving stack
Every platform team faces the same challenge: securely running AI agents that execute arbitrary code with shell access. Most teams use ad-hoc containers with overprivileged secrets—a security nightmare.
This talk presents a production-validated zero-trust architecture using kubernetes-sigs/agent-sandbox with gVisor/Kata runtime enforcement, KRO for declarative abstractions, and GitOps provisioning. Security is layered: kernel-level syscall interception prevents escape even with root access; credential proxies eliminate secret exposure; NetworkPolicy blocks metadata endpoints; warm pools deliver sub-second startup without weakening boundaries.
Platform teams learn to define KRO-based CRDs that abstract consumers from rapid AI framework changes—when the underlying agentic stack evolves, the platform interface remains stable. Attendees discover why treating AI agents as untrusted tenants—not trusted services—is the only scalable security pattern. Built entirely on CNCF projects.
Kubernetes Pods have traditionally been blocked from using hardware-level networking acceleration due to namespace isolation. Cilium is changing that by bringing native Dynamic Resource Allocation (DRA) support to manage hardware resources for Pods.
This talk unveils how we built two new DRA drivers for Cilium: an SR-IOV flavor as well as a new hardware queue leasing mechanism recently merged into the Linux kernel’s networking stack and the DRAgons we faced along the way. We’ll dive deep into how queue leasing proxies real physical NIC queues directly to virtual netdevs, how we integrated it into netkit, Cilium’s eBPF datapath, and the tradeoffs between the two flavors.
This architecture brings native io_uring TCP zero-copy into Pods for the first time, enabling high-speed memory-mapped networking. Join us to see how these two new Cilium DRA drivers can be used and how we enable 100G+ single-stream throughput while preserving the operational model cloud native platforms expect today.
What does an OpenTelemetry trace look like when there's no HTTP request to start it? What does the collector do when the network disappears for ten minutes? What's the right batch size when your "host" is a Raspberry Pi slapped onto a child’s toy?
We added an OTel collector to a toy RC car to find out.
A large automotive manufacturer is working through similar questions in every car coming off their production line, at a very different scale. We'll drive our RC car around the stage, show what happens when an edge device loses its connection, and dig into the OTel collector features that make edge deployments work.
Model Context Protocol (MCP) is becoming the standard for connecting LLMs to tools — but running MCP servers in production raises hard infrastructure questions: How do you isolate tool execution per user? How do you scale to thousands of concurrent sessions without cold-start penalty? How do you safely expose filesystem, shell, and network access to agents?
This session presents a Kubernetes-native approach to MCP server infrastructure using warm sandbox pools. We show how SandboxSet pre-warms execution environments, SandboxClaim binds them to agent sessions in under 100ms, and checkpoint/restore preserves tool state across session boundaries. We cover threat isolation (per-sandbox namespaces + seccomp), dynamic volume injection for per-user data, and graceful session handoff when an MCP server needs upgrade. Attendees leave with reusable patterns for deploying MCP servers at production scale on Kubernetes — without rebuilding your infrastructure from scratch.
Most observability configurations are static. We decide upfront which spans to emit, which attributes to record, and how much telemetry is worth collecting. But production systems are not static. When everything is healthy, detailed telemetry creates unnecessary cost and noise. When incidents happen, the telemetry we need often is not being collected.
What if telemetry verbosity could adapt automatically?
In this talk, we combine OpenTelemetry and OpenFeature to dynamically control instrumentation at runtime. Services can switch between lightweight traces and deep diagnostic telemetry without redeployments or restarts. Feature flags drive observability decisions based on runtime conditions, enabling richer spans, additional attributes, and higher-cardinality signals only when they are needed.
We will demonstrate how OpenTelemetry semantic conventions can act as a standardized control surface, enabling adaptive and vendor-neutral observability across services and environments.
Step into agentic systems on Kubernetes, where identities aren't just users and services, they're agents, models, and workflows.
In this 101-level session, two security maintainers walk through how authN and authZ evolve in AI systems. We'll separate the three identities teams conflate, user, agent, and workload, and show how familiar building blocks like OAuth 2.1, scopes, and short-lived credentials extend to non-human callers. We'll trace how identity travels across multi-step workflows, carrying the chain of "who is acting on whose behalf" through every hop.
Then we'll look at the protocols agents use today (MCP, A2A) and the practical risks they bring: external model calls, replayed tokens, and unbounded automation. We'll touch on CNCF building blocks like SPIFFE/SPIRE and agentgateway, and the simple guardrails that prevent the most common failures.
You'll leave with a clear mental model and a checklist you can apply Monday morning.
We will cover the last 3 releases of Kubernetes in terms of features added to the Kubernetes Control Plane, and will have an exploration of the future at the light of AI / ML needs.
Longhorn is an enterprise-grade, cloud-native distributed storage platform that provides highly available, high-performance persistent volumes for Kubernetes workloads. It includes data services such as snapshots, out-of-cluster backups, data integrity checks, encryption, disaster recovery, and more.
Longhorn’s v1 data engine uses iSCSI over TCP and an in-house protocol for volume access and cross-node replication. Since the 1.12 release, the v2 data engine, built on SPDK with NVMe over TCP, has been generally available, delivering major gains in performance and efficiency. The latest release and roadmap add features such as sharded volumes and fast cloning to further improve scalability.
In this talk, we’ll introduce Longhorn, explain its data engines and architecture, highlight recent improvements, and show how it can strengthen your Kubernetes storage strategy for your business-critical workloads. Longhorn has been an incubating CNCF project since November 2021.
As Bare Metal kubernetes adoption grows within the CNCF ecosystem, projects such as Metal³ play a critical role in reliable, production-grade infrastructure. Core components including IPAM, CAPM3, Bare Metal Operator (BMO) and Ironic Standalone Operator (IRSO) must operate correctly across a wide range of real-world conditions.
This talk presents fuzz testing as a practical and effective technique for improving the robustness of Metal³ components. By continuously injecting randomized and malformed inputs into key interfaces and control loops.
We will discuss how fuzzing can be integrated into CNCF style development workflows, corpus management aligned with upstream contribution practices. We highlight issues uncovered in IPAM, CAPM3, BMO and IRSO.
Attendees will gain actionable guidance on adopting fuzzing within their own cloud-native projects and insight into how proactive testing strategies can strengthen the reliability and security of Bare Metal Kubernetes infrastructure.
Earlier this year, Strimzi celebrated its 1.0.0 release – a major project milestone. After briefly recalling the 1.0.0 release and its main changes, this talk will focus on our roadmap and on where the project is heading next. Gateway API support opens new ways to connect to your Apache Kafka clusters. A new pluggability mechanism inspired by Kubernetes Admission Controllers and policy engines such as Kyverno or Open Policy Agent makes it easier to customize and extend Strimzi. And more flexible security configuration gives you finer control over how your Apache Kafka cluster is secured. Those are just some of the features that are planned or already in progress. We will cover them in this talk and show you how to get involved and help shape them. You will leave this session with a clear picture of the future of Strimzi.
