# KubeCon + CloudNativeCon North America — Schedule

Machine-readable mirror of the schedule page. Generated 2026-09-14.

- Source: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/
- Sessions: 223 across 2 days
- Times are the event's local times, exactly as published. Timing and rooms are subject to change.
- Each session links back to the schedule page, which opens that session's details.

## Tracks

- Agentics Day: MCP + Agents (31)
- Cloud Native AI + Inference Day (31)
- ArgoCon (28)
- Observability Day (28)
- Platform Engineering Day (28)
- BackstageCon (16)
- Open Source SecurityCon (15)
- CiliumCon (12)
- FluxCon (10)
- OpenTofu Day (10)
- Kubernetes on Edge Day (10)
- Registration (4)

## Sunday, November 8, 2026

### 8:00 AM–5:00 PM · Registration + Badge Pickup

- Room: 200 South Entrance (South)
- Track: Registration
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1311891

### 8:00 AM–6:00 PM · Coat + Bag Check

- Room: Salt Palace | Level 2 | East Registration
- Track: Registration
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1311892

## Monday, November 9, 2026

### 7:00 AM–5:00 PM · Registration + Badge Pickup

- Room: 200 South Entrance (South)
- Track: Registration
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1311893

### 7:00 AM–5:00 PM · Coat + Bag Check

- Room: Salt Palace | Level 2 | East Registration
- Track: Registration
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1311894

### 9:00 AM–9:05 AM · Welcome + Opening Remarks

- Room: Salt Palace | Level 1 | Ballroom H
- Track: FluxCon
- Labels: Any Level
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1312037

### 9:00 AM–9:05 AM · Welcome + Opening Remarks

- Room: Salt Palace | Level 1 | Ballroom ACE
- Track: Agentics Day: MCP + Agents
- Labels: Any Level
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1312031

### 9:00 AM–9:05 AM · Welcome + Opening Remarks

- Room: Salt Palace | Level 2 | 255 EF
- Track: BackstageCon
- Labels: Any Level
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1312034

### 9:00 AM–9:05 AM · Welcome + Opening Remarks

- Room: Salt Palace | Level 1 | 151
- Track: ArgoCon
- Labels: Any Level
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1312032

### 9:00 AM–9:05 AM · Welcome + Opening Remarks

- Room: Salt Palace | Level 2 | 254
- Speakers: Yuzhui Liu, Rajs Kakodkar, Yuan Tang
- Track: Cloud Native AI + Inference Day
- Labels: Any Level
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1312036

### 9:00 AM–9:05 AM · Welcome + Opening Remarks

- Room: Salt Palace | Level 2 | 250 A-C
- Track: Observability Day
- Labels: Any Level
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1312039

### 9:00 AM–9:05 AM · Welcome + Opening Remarks

- Room: Salt Palace | Level 2 | 255 BC
- Track: Open Source SecurityCon
- Labels: Any Level
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1312041

### 9:00 AM–9:05 AM · Welcome + Opening Remarks

- Room: Salt Palace | Level 1 | Ballroom BDF
- Track: Platform Engineering Day
- Labels: Any Level
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1312043

### 9:00 AM–9:05 AM · Welcome + Opening Remarks

- Room: Salt Palace | Level 1 | Ballroom J
- Track: CiliumCon
- Labels: Any Level
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1312035

### 9:10 AM–9:35 AM · Flux After 10 Years: GitOps for Humans and Agents

- Room: Salt Palace | Level 1 | Ballroom H
- Speakers: Stefan Prodan
- Track: FluxCon
- Labels: Any Level, Breakout Session, Flux Roadmap + Vision + Community Updates
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269558

Flux began in 2016 as the first tools to put Git at the center of Kubernetes delivery, and "GitOps" grew up alongside it.

After a brief look back, this talk focuses on where Flux is going. We'll cover the features shipped this past year across the GitOps Toolkit and Flux Operator, then dig into two shifts that define the next decade. First is wiring Flux into AI agents and the second is bringing Flux closer to developers through the Flux Operator UI.

The session concludes with the Flux roadmap and opportunities for community contribution.

### 9:10 AM–9:35 AM · How Ericsson Combines Human Judgment to Enhance AI-Powered Capabilities Through Backstage

- Room: Salt Palace | Level 2 | 255 EF
- Speakers: Damien O'Toole, Kieran Egan
- Track: BackstageCon
- Labels: AI Agents + Context in Backstage, Breakout Session, Intermediate
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1267947

AI capabilities accelerate individual tasks but leave the SDLC static — bottlenecks simply migrate between stages. This session presents an AI-native architecture where specialized agents operate at every lifecycle stage while an AI Monitor provides system-level governance.

From Developer Experience (DX) at Ericsson, we show how these patterns extend across the SDLC, integrated into our Backstage-powered IDP.

The AI Monitor correlates upstream causes with downstream failures, surfacing improvements to human engineers who evaluate and implement changes. Humans remain in the loop as decision-makers bringing domain context agents cannot infer alone.

With Backstage as our single source of fresh, human-verified truth, we built AI capabilities on deterministic-first principles — using selective LLM escalation to minimize cost while maximizing reliability — integrated into our IDP.

Learn to pair system-level intelligence with human expertise to build self-improving golden paths.

### 9:10 AM–9:35 AM · ArgoCon Project Update

- Room: Salt Palace | Level 1 | 151
- Track: ArgoCon
- Labels: Any Level, Breakout Session
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1311951

### 9:10 AM–9:35 AM · Rolling Out LLM Inference Engines on Kubernetes: Capacity, Topology, and Deployment Verification

- Room: Salt Palace | Level 2 | 254
- Speakers: Meng Dong, Zihan Zheng
- Track: Cloud Native AI + Inference Day
- Labels: Breakout Session, End-to-End AI/ML in Production (MLOps/AIOps), Intermediate
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269716

Updating an LLM inference engine is not the same as rolling a stateless Kubernetes service. Each engine deployment carries GPU shape, topology, memory layout, prefill/decode balance, and startup cost. At OpenAI, a deployment preflight pipeline can keep GPU-backed pods pinned for hours during startup, warmup, and validation, creating capacity, scheduling, topology, and rollout challenges for the Kubernetes fleet.

Using the deployment preflight pipeline as a case study, the talk explains how quota, GPU topology, model architecture, cluster readiness, and availability targets become rollout admission checks, and how operators reason about compatible capacity, topology fit, rollback conditions, and blocked updates. Attendees will leave with portable patterns for deploying LLM engines when capacity and topology matter as much as pod readiness.

### 9:10 AM–9:35 AM · Observability Project Updates

- Room: Salt Palace | Level 2 | 250 A-C
- Track: Observability Day
- Labels: Any Level, Breakout Session
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1312019

### 9:10 AM–9:35 AM · End to End Software Development Lifecycle Hardening with AMPEL

- Room: Salt Palace | Level 2 | 255 BC
- Speakers: Adolfo García Veytia
- Track: Open Source SecurityCon
- Labels: Beginner, Breakout Session, Supply Chain + Application Security
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269816

As software delivery accelerates, teams need automated ways to verify compliance and security controls across the SDLC.

Hardening a pipeline means checking many things: who can run CI, whether repositories are configured safely, cryptographically verifying commit authors, whether tools and build chains have trusted provenance, auditing for vulnerabilities, and if all this is acceptable before release.

Yes there's a tool for each step, but stitching them together quickly becomes an unmaintainable mess.

Enter AMPEL, the supply chain policy engine. AMPEL uses policy as code to evaluate each stage of the SDLC, blocking builds when conditions are not met. Uniform, reusable policies help protect your systems, users, and organization from insecure releases.

In this hands-on session, we’ll use community-driven policies to lock down steps, monitor checks, and map controls to frameworks like the OSPS Baseline and NIST SSDF. Bring a project and a laptop; let’s harden it together.

### 9:10 AM–9:20 AM · CNCF Platform Engineering Technical Community Group Update

- Room: Salt Palace | Level 1 | Ballroom BDF
- Speakers: Chris Plank, Colin Griffin, Atulpriya Sharma, Katie Greenley
- Track: Platform Engineering Day
- Labels: Any Level, Building Platforms for Day 2 and Beyond, Panel Discussion
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269430

Come along and listen to the leaders of the CNCF's Platform Engineering Technical Community Group give a quick update on what the community has been up to since KubeCon Atlanta, the initiatives underway and planned as well as what events they are running at KubeCon Salt Lake City.

### 9:10 AM–9:35 AM · Handing a Proxy the Keys to Your Pod: Inside Cilium's Sidecarless mTLS

- Room: Salt Palace | Level 1 | Ballroom J
- Speakers: Quang Nguyen, Robin Gögge
- Track: CiliumCon
- Labels: Breakout Session, Intermediate, Security: Network Policy + Encryption + Runtime Enforcement + More on Securing Your Cluster
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1268376

Cilium 1.19 shipped sidecarless mTLS reusing ztunnel, a proxy from Istio's ambient mode: label a namespace, and traffic between its pods is encrypted and identity-authenticated over HBONE, with no app changes necessary. But underneath that one line is a per-node proxy opening sockets inside pods it was never installed into, and a control plane fanning that label out to every node.

So how does it actually work? We'll walk from that label to the kernel primitives beneath it: the file-descriptor hand-off that opens listeners inside your pod, the StateDB reconciler that turns a namespace label into per-node state, how its SPIFFE identity is issued, and the in-pod iptables redirect that intercepts traffic.

You'll leave knowing how the pieces fit and where its edges are today (beta, TCP-only). We close on the SPIRE work we're upstreaming, which binds a certificate to the requesting process, not just to a pod, with a live demo: steal a cert without attestation, then watch SPIRE refuse.

### 9:25 AM–9:50 AM · Is Platform Engineering and the Platform Maturity Model Still Relevant in the Agentic AI Era?

- Room: Salt Palace | Level 1 | Ballroom BDF
- Speakers: Chris Plank, Corey McGalliard
- Track: Platform Engineering Day
- Labels: Any Level, Breakout Session, Platform Engineering and AI
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269433

With the rise of Agentic AI do we still need Platform Engineers? Do we need to care about or even consider platform maturity? With the landscape changing almost on a daily basis, where agentics such as OpenClaw and AWS Kiro are only 6-12 months old, consultancies and vendors preaching new operating models and practices, lots of MVP’s and POCs showing benefits but not many talking about Production, what’s the real life discussions and feedback in the Platform Engineering community?

Come listen to two Platform Engineering practitioners and leaders who have been working on minor updates to the CNCF Platform Engineering Maturity Model and their perspectives.

### 9:40 AM–9:45 AM · Sponsored Keynote | TBA, Microsoft

- Room: Salt Palace | Level 1 | Ballroom ACE
- Track: Agentics Day: MCP + Agents
- Labels: Any Level, Keynote Session
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1311927

### 9:40 AM–9:45 AM · Sponsored Keynote | TBA, IBM

- Room: Salt Palace | Level 2 | 250 A-C
- Track: Observability Day
- Labels: Any Level, Keynote Session
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1311994

### 9:45 AM–10:10 AM · Zero-Touch App Team Onboarding: Powering Self-Service Kubernetes at Enterprise Scale with Flux

- Room: Salt Palace | Level 1 | Ballroom H
- Speakers: Tiffany Wang, Anthony Burris
- Track: FluxCon
- Labels: Breakout Session, Building Internal Developer Platforms with Flux, Intermediate
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1267718

At a large financial institution running Flux across hundreds of private and public cloud Kubernetes clusters, the question is how you make GitOps transparent to thousands of developers.
We'll share our fully automated, zero-touch GitOps onboarding platform targeting Flux-powered Kubernetes clusters — with no platform team intervention.
We'll review the architecture that makes this possible, covering how it meets production SLA requirements while preventing privilege escalation across team boundaries. You’ll see how we leverage the Flux Operator to simplify Flux lifecycle management, and we'll go over how we close the feedback loop for app developers by wiring Flux's notification-controller back into CI pipelines for observability into deployment status.
Whether you're building a platform for 10 teams or 10,000, this talk gives you a battle-tested blueprint for a self-service GitOps infrastructure engine that removes platform engineering team bottlenecks.

### 9:45 AM–9:55 AM · Your Backstage Portal Just Became an AI App Store

- Room: Salt Palace | Level 2 | 255 EF
- Speakers: Christophe Fargette, Evan Shortiss
- Track: BackstageCon
- Labels: AI Agents + Context in Backstage, Intermediate, ⚡ Lightning Talk
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269670

Agents, MCPs, and AI skills are proliferating — but how do developers actually find, evaluate, and securely adopt them? We built an AI Catalog inside Backstage that turns your developer portal into a one-stop shop for AI assets.

In this lightning talk, I'll walk through the user journey: browsing available agents and MCP servers, reading their capabilities and ownership metadata, and installing them — all without leaving Backstage. No hunting through GitHub READMEs, no Slack threads asking "which agent does X?"

We'll see a live demo of the prototype and cover the key design choices: how we modeled AI assets alongside existing software components, what "install" means in the context of an AI skill, and how Backstage's existing plugin and catalog model turns out to be a surprisingly good fit for this problem.

If you're thinking about how to make AI tooling discoverable and governable inside your organization, this talk is for you.

### 9:45 AM–9:50 AM · Sponsored Keynote | TBA, Akuity

- Room: Salt Palace | Level 1 | 151
- Track: ArgoCon
- Labels: Any Level, Keynote Session
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1311945

### 9:45 AM–9:55 AM · Who Controls Your Inference? A Field Guide to Sovereign AI

- Room: Salt Palace | Level 2 | 254
- Speakers: Bjorn Hovland
- Track: Cloud Native AI + Inference Day
- Labels: Ethics + Compliance + Security, Intermediate, ⚡ Lightning Talk
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1268082

“Sovereign AI” means everything from “we don’t train on your data” to “you hold the weights, the runtime, and control of the infrastructure.” For teams making a deployment decision, that ambiguity is a problem. In June 2026, a US export-control directive forced Anthropic to disable its most capable models for a class of users with no notice.

This session replaces the marketing term with an engineering one: AI sovereignty is a control architecture defined by four things an operator must own: the model, the inference stack, the data path, and the update cadence. The talk maps each onto cloud native building blocks, from open-weight models served with vLLM behind KServe to Kubernetes scheduling, the Gateway API Inference Extension, network isolation, and GitOps-governed updates.

It closes with the tradeoffs. Sovereignty guarantees control, not capability or lower cost. Attendees leave with a vendor-neutral framework for deciding which workloads belong on infrastructure they own.

### 9:45 AM–9:55 AM · Creating VEX Documents at Scale: A Look into One Supplier's Process

- Room: Salt Palace | Level 2 | 255 BC
- Speakers: Jeff Mendoza
- Track: Open Source SecurityCon
- Labels: Intermediate, Regulation + Public Policy, ⚡ Lightning Talk
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1268168

What is a VEX? What goes into it? How should you determine that? When and how should you create one? Why do you need one? This session answers those questions and more through a case study into how a large software supplier, Microsoft, is preparing to automate the creation of VEX documents. Specifically, the CSAF VEX specification required fields and options will be covered, however OpenVEX will also be investigated.

VEX documents are now a vital necessity for secure usage and inclusion of Open Source Software. They are required to supply software in compliance with EU CRA, alongside SBOM. With AI accelerating the rate of discovery of security vulnerabilities, VEX documents will be more and more applicable, as software suppliers will have trouble keeping up and many vulnerabilities are not exploitable.

### 9:45 AM–10:10 AM · Cluster Mesh, Native Routing, and Multi-Pool IPAM in IPv6-Only Network: What the Docs Don't Tell You

- Room: Salt Palace | Level 1 | Ballroom J
- Speakers: Kapil Agrawal, Luke Baker
- Track: CiliumCon
- Labels: Advanced, Breakout Session, Use Cases: End User Story Describing the Use of Cilium + Hubble + Tetragon to Overcome Challenges
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269739

Last year we shared how ESnet, the U.S. Department of Energy’s high-performance research network, runs Kubernetes on-prem in IPv6 only network using Cilium. This talk covers what came next: building cluster-mesh using native routing, BGP, and multi-pool IPAM entirely without IPv4.

We'll walk through our architecture: why native routing over tunnels, how Cilium's BGP control plane with multi-pool IPAM enables pod-to-pod connectivity within and across clusters, and how we secure inter-cluster traffic with host firewall.

This isn't a feature walkthrough - we'll focus on what broke. Undocumented IPv6 PodIPPool mask size limitations silently blocked pod scheduling. Host firewall policies locked us out of control plane access and triggered etcd quorum loss, cascading into cluster-wide failures. For each, we'll show diagnosis, recovery, and prevention.

Attendees leave with a concrete IPv6-only cluster-mesh architecture, a trip-wire checklist, and hard-won lessons from production setup.

### 9:50 AM–9:55 AM · Sponsored Keynote | TBA, Port

- Room: Salt Palace | Level 1 | Ballroom ACE
- Track: Agentics Day: MCP + Agents
- Labels: Any Level, Keynote Session
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1311929

### 9:50 AM–9:55 AM · Sponsored Keynote | TBA, Traversal

- Room: Salt Palace | Level 2 | 250 A-C
- Track: Observability Day
- Labels: Any Level, Keynote Session
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1311996

### 9:55 AM–10:00 AM · Sponsored Keynote | TBA, Intuit

- Room: Salt Palace | Level 1 | 151
- Track: ArgoCon
- Labels: Any Level, Keynote Session
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1311947

### 9:55 AM–10:00 AM · Sponsored Keynote | TBA, IBM

- Room: Salt Palace | Level 1 | Ballroom BDF
- Track: Platform Engineering Day
- Labels: Any Level, Keynote Session
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1312009

### 10:00 AM–10:05 AM · Sponsored Keynote | TBA, Spectro Cloud

- Room: Salt Palace | Level 1 | Ballroom ACE
- Track: Agentics Day: MCP + Agents
- Labels: Any Level, Keynote Session
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1313055

### 10:00 AM–10:05 AM · Sponsored Keynote | TBA, Spotify

- Room: Salt Palace | Level 2 | 255 EF
- Track: BackstageCon
- Labels: Any Level, Keynote Session
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1311954

### 10:00 AM–10:05 AM · Sponsored Keynote | TBA, Vultr

- Room: Salt Palace | Level 2 | 254
- Track: Cloud Native AI + Inference Day
- Labels: Any Level, Keynote Session
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1311962

### 10:00 AM–10:05 AM · Sponsored Keynote | TBA, Dynatrace

- Room: Salt Palace | Level 2 | 250 A-C
- Track: Observability Day
- Labels: Any Level, Keynote Session
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1311997

### 10:05 AM–10:10 AM · Sponsored Keynote | TBA, Vultr

- Room: Salt Palace | Level 1 | Ballroom BDF
- Track: Platform Engineering Day
- Labels: Any Level, Keynote Session
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1312010

### 10:10 AM–10:15 AM · Sponsored Keynote | TBA, ControlPlane

- Room: Salt Palace | Level 1 | Ballroom H
- Track: FluxCon
- Labels: Any Level, Keynote Session
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1311969

### 10:10 AM–10:15 AM · Sponsored Keynote | TBA, AWS

- Room: Salt Palace | Level 1 | Ballroom ACE
- Track: Agentics Day: MCP + Agents
- Labels: Any Level, Keynote Session
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1325759

### 10:10 AM–10:15 AM · Sponsored Keynote | TBA, Broadcom

- Room: Salt Palace | Level 2 | 254
- Track: Cloud Native AI + Inference Day
- Labels: Any Level, Keynote Session
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1311964

### 10:10 AM–10:15 AM · Sponsored Keynote | TBA, Clickhouse

- Room: Salt Palace | Level 2 | 250 A-C
- Track: Observability Day
- Labels: Any Level, Keynote Session
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1311998

### 10:15 AM–10:20 AM · Sponsored Keynote | TBA, vCluster

- Room: Salt Palace | Level 1 | Ballroom BDF
- Track: Platform Engineering Day
- Labels: Any Level, Keynote Session
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1312011

### 10:15 AM–10:20 AM · Sponsored Keynote | TBA, Isovalent

- Room: Salt Palace | Level 1 | Ballroom J
- Track: CiliumCon
- Labels: Any Level, Keynote Session
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1311961

### 10:15 AM–10:40 AM · AM Break 1

- Room: Meeting Room Foyers
- Track: Agentics Day: MCP + Agents
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1309409

### 10:20 AM–10:25 AM · Sponsored Keynote | TBA, SUSE

- Room: Salt Palace | Level 2 | 254
- Track: Cloud Native AI + Inference Day
- Labels: Any Level, Keynote Session
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1311965

### 10:20 AM–10:40 AM · AM Break 2

- Room: Meeting Room Foyers
- Track: CiliumCon
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1309412

### 10:25 AM–10:30 AM · Sponsored Keynote | TBA, WSO2

- Room: Salt Palace | Level 1 | Ballroom BDF
- Track: Platform Engineering Day
- Labels: Any Level, Keynote Session
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1312012

### 10:25 AM–10:40 AM · AM Break 3

- Room: Meeting Room Foyers
- Track: Cloud Native AI + Inference Day
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1309411

### 10:30 AM–10:40 AM · AM Break 4

- Room: Meeting Room Foyers
- Track: Platform Engineering Day
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1309410

### 10:40 AM–11:05 AM · From Helmfile to Flux: GitOps for Telco at Enterprise Scale

- Room: Salt Palace | Level 1 | Ballroom H
- Speakers: Daniel Costa Soares
- Track: FluxCon
- Labels: Breakout Session, Intermediate, Running Flux at Scale in Production
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1268262

Telco BSS/OSS software is mission-critical, heavily regulated, and notoriously hard to deploy. For a large product organization, rebuilding years of mature CI/CD pipelines just to “do GitOps” is non-starter. So how do you move an entire BOS portfolio to GitOps with Flux — without throwing away what already works, and without forcing every customer to change overnight?
This is our journey. We made continuous testing the foundation, then built a “helm2flux” conversion that automatically generates Flux manifests from existing Helm assets — no pipeline rebuilds, and room for customers to adopt continuous deployment at their own pace. Today we run microservices delivery and full Git-driven day-2/N configuration across the portfolio. Next: AIOps, using OpenTelemetry, MCP, and agent skills to let AI safely assist the lifecycle, with the security guarantees telco demands.
Daniel covers the org side; Carl-Fredrik covers tech. Expect honest lessons, reusable patterns, and a few scars.

### 10:40 AM–11:05 AM · MCP Without Polling: Building Reactive Agents with Continuous Queries

- Room: Salt Palace | Level 1 | Ballroom ACE
- Speakers: Aman Singh
- Track: Agentics Day: MCP + Agents
- Labels: Any Level, Breakout Session, Building MCP Servers + Clients
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1264811

Most AI agents discover state by polling- wasteful, stale, and brittle at production scale. This talk shows a different architecture: Drasi, a CNCF project, exposes continuous queries as real-time MCP resources that agents can subscribe to. The query becomes the API contract; the agent gets notified the moment data changes. The engineer who built the integration walks through the architecture, demos it end-to-end on Kubernetes, and shares practical lessons from building and operating real-time agent infrastructure with AI coding agents.

### 10:40 AM–11:05 AM · Securing the Agentic Tool Layer: A Runtime Defense Framework for MCP Agent Deployments

- Room: Salt Palace | Level 1 | Ballroom IG
- Speakers: Saurabh Yergattikar, Harshul Jain
- Track: Agentics Day: MCP + Agents
- Labels: Breakout Session, Intermediate, Security + Trust + Reliability
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269818

Most LLM safety work guards the user prompt. MCP agents face a different problem: the attack arrives through the tools the agent already trusts. A poisoned tool description, a hidden instruction in a tool response, or a malicious server pulled from a public registry can hijack an agent without ever touching the user.

This session walks through a threat taxonomy built from 80+ techniques catalogued under SAFE-MCP, a Linux Foundation/OpenSSF project, then shows a runtime proxy that validates tool calls and responses inline, with no changes to the agent or server. In red-team testing across five model backends, it cut tool-poisoning success from 74% to under 9% and indirect injection from 47% to under 6%, adding under 120ms per call.

Attendees leave with a concrete defense architecture, the false-positive tradeoffs that matter in production, and what to vet before approving an MCP server for their stack.

### 10:40 AM–11:05 AM · From Geological Layers to NFS: Migrating a Decade of Backstage with AI

- Room: Salt Palace | Level 2 | 255 EF
- Speakers: Marat Dyatko, Patrik Oldsberg
- Track: BackstageCon
- Labels: Any Level, Breakout Session, Transforming Dev Experience + Streamlining Dev Workflow + Doing More with Less Using Backstage
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1267001

Spotify's internal Backstage goes back to 2014, four years before Backstage even had a name. Over a decade it accumulated ancient plugins, stalled earlier migrations that left everyone stuck between old and new, and frontend sediment from every era of the company. An earlier attempt to modernize ran out of steam and nobody wanted to try again.

We tried again anyway, this time with AI doing most of the grunt work. About 150 plugins moved to the New Frontend System in roughly ten weeks, around half a million lines changed. Without AI agents handling the codemod volume, this migration would have stalled the same way the last one did.

The talk gets into the mechanics: how we iterated agent skills against a codebase few fully understood, swapped centralised app wiring for extension blueprints, used feature flags so teams could fall back to the old experience, and built a custom plugin to make migration progress visible to everyone.

### 10:40 AM–11:05 AM · Beating the Final Boss: GitOps Promoter in Production

- Room: Salt Palace | Level 1 | 151
- Speakers: Michael Crenshaw, Zach Aller
- Track: ArgoCon
- Labels: Breakout Session, Intermediate, Software Delivery
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269659

Environment promotion is the final boss of GitOps, forcing teams to build hacky imperative glue scripts or abandon GitOps entirely. Argo’s Progressive Syncs and Sync Waves work for basic use cases but fail for more complex requirements.

Intuit built something completely different, something fully declarative: GitOps Promoter, an OSS project in Argoproj Labs.

GitOps Promoter moves changes through environments by continuously evaluating gates. Gates pass, and changes ship, like falling dominoes: no buggy CI scripts to debug, no imperative steps to coordinate.

Productionizing GitOps Promoter at Intuit involved shipping new features: web request gates for change management, an Argo UI extension, multi-Argo instance support, etc. We’ll cover these plus features that helped ship GitOps Promoter for other adopters such as Red Hat and Circle.

We’ll catch you up to the environment promotion state of the art and let you know what’s around the corner for Argo and the GitOps ecosystem.

### 10:40 AM–11:05 AM · Demystifying an Argo Workflows Pod — And How We Stopped Needing an Init Container

- Room: Salt Palace | Level 2 | 251 A-C
- Speakers: Alan Clucas, Jason Meridth
- Track: ArgoCon
- Labels: Breakout Session, Data Processing, Intermediate
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1267152

When Argo Workflows runs your workload, your container isn't alone in its pod. Argo wraps other containers and volumes around it to feed your code its inputs and carry the results back out, and most people never notice any of that is happening.

This talk opens up the pod. We'll look at what each piece does: how your parameters and artifacts get in before your code starts, how the outputs and logs get back out after it finishes, and the little shared filesystem Argo uses to pass it all between containers. You'll come away actually knowing what Argo does for you every time a step runs.

Then we'll get to what's changed. Every Argo pod used to begin with an init container that staged things before your work started. Kubernetes now supports image volumes, which means we can drop that init container entirely. I'll show how this new opt-in feature in Argo Workflows 4.1 works, what you get for it, and the small things worth knowing before you switch it on.

### 10:40 AM–11:05 AM · From Routing Weights to Tail Latency: Simulating Kubernetes-Native LLM Inference Without GPUs

- Room: Salt Palace | Level 2 | 254
- Speakers: Dipanwita Guhathakurta, Mert Toslali
- Track: Cloud Native AI + Inference Day
- Labels: Advanced, Breakout Session, End-to-End AI/ML in Production (MLOps/AIOps)
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1259185

How does a single routing weight change impact a 16-instance Kubernetes LLM cluster? It triggers a cascade through prefill/decode disaggregation, KV-cache eviction, and autoscaler loops, causing erratic tail latency spikes. Teams are forced to burn massive GPU budgets on test clusters just to safely tune a single production scheduling parameter.

Discrete-event simulation offers a systematic escape from this hardware bottleneck. By translating pipelines, KV-cache states, and routing logic into deterministic events, performance can be evaluated instantly. Validated across 50 configurations with a 10% median latency error, this approach mimics live hardware without a single GPU.

We break down how to map complex Kubernetes orchestration dynamics into lightning-fast execution events. Attendees leave with an understanding of how to leverage a high-fidelity, open-source simulation framework to accelerate critical workloads like complex policy exploration and automated capacity planning.

### 10:40 AM–11:05 AM · "The RCA Bot Needed an RCA": Why We Stopped Giving Agents Everything They Asked For

- Room: Salt Palace | Level 2 | 251 D-F
- Speakers: Madhu Patel, Sudhanshu Sah
- Track: Cloud Native AI + Inference Day
- Labels: Breakout Session, Intermediate, RAG + AI Agents
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269543

Every major incident now has the same postmortem footnote: automate the RCA. So we gave agents access to everything. Datadog, Splunk, GitHub, Kubernetes. Investigations got faster. Accuracy got worse.

The failure was not the agents. It was not the tools. It was a fundamental assumption about agent context that most teams building on LLMs have not hit yet. The fix required rethinking the layer that runs before any agent sees anything.

We built an incident RCA system on Kubernetes that goes from alert to validated root cause in under 5 minutes. The stack uses LangGraph, Temporal, and pgvector, but the interesting part is none of those. It is what sits between the raw telemetry and the agents.

Three patterns we will cover:
An evidence layer that runs before any agent sees anything
Specialist agents in parallel, with synthesis as a separate final step, and orchestration, where the LLM decides nothing about workflow progression.

### 10:40 AM–11:05 AM · Designing an Internal SLO Platform

- Room: Salt Palace | Level 2 | 250 A-C
- Speakers: Julio Leal, Robert Souza
- Track: Observability Day
- Labels: Breakout Session, End User Stories: Production Deployments + Migrations + Operations, Intermediate
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1245493

Most orgs have SLOs but don't trust them, because platforms are built around services — producing thousands of numbers nobody aggregates. Nubank built its SLO platform around business domains/subdomains instead, so questions like "are card payments healthy?" get one answer.

The talk covers: (1) why domains beat services as the unit; (2) the architecture — a dedicated VictoriaMetrics stack (vicmetrics-*-slo) isolated from main observability, running in six countries, with a Go API for registration and an slo-generator computing SLI/SLO from PromQL; (3) production lessons from running it across thousands of services — what worked, what didn't, what they'd redo.

Takeaway: a practical template for building SLOs on infra you already have (VictoriaMetrics/Prometheus, Grafana, small Go service) and a framework for what to prioritize.

### 10:40 AM–11:05 AM · OpenTelemetry Far and Wide

- Room: Salt Palace | Level 2 | 250 D-F
- Speakers: Nikola Grcevski, Tyler Yahn
- Track: Observability Day
- Labels: Any Level, Breakout Session, Observability Cost + Quality + Scale + Efficiency
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1246975

OpenTelemetry eBPF Instrumentation made broad observability effortless: one command, zero code changes, every service covered. But breadth isn't always the whole story. OpenTelemetry SDK instrumentations operate at the language level, surfacing transaction details that can be lifesavers in debugging scenarios and provide invaluable insights to developers. The catch: rolling out SDK instrumentation across a diverse, multi-technology/multi-team ecosystem can be slow and expensive.

This talk introduces a practical approach that eliminates the tradeoff: by combining eBPF instrumentation with SDK auto-injection, teams get broad coverage immediately and deep, language-level insights where needed, with a single deployment mechanism and dynamic cost controls.

We’ll do a live demo of this combined instrumentation approach and show how with a single command we can instrument everything in our system, upgrade to new OpenTelemetry SDK versions and control the cost of our instrumentation.

### 10:40 AM–11:05 AM · Upstream First Security: Keeping Telco Infrastructure Secure at Cloud Native Scale

- Room: Salt Palace | Level 2 | 255 BC
- Speakers: Jan Melen, Kashif Khan
- Track: Open Source SecurityCon
- Labels: Any Level, Breakout Session, Open Source Management (OSPOs) + Security
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1265217

Telco infrastructure runs on bare metal that powers 5G mobile networks. When a CVE drops in Go dependency used by the stack, we cannot wait until next quarter. Telco regulatory requirements mean patch latency is measured in days. Time from discovery to exploitation has shank from days to hours, we've had to rethink how we work with OS and adopt a strict upstream-first approach.
We maintain bare metal Kubernetes infrastructure by using CNCF projects such as Metal3 and Cluster API. This talk covers how we triage CVEs across a dependency tree spanning 15+ upstream repositories, why we always fix issues upstream first instead of local patches, and how we structure release cadence so that production stays within one minor version of the latest.
We will share numbers, including the time to fix CVEs and the cost of backporting vs. upgrading. We discuss how we automated dependency scanning in multi-repository CNCF project. This are mechanics of keeping infrastructure secure with open source.

### 10:40 AM–11:05 AM · Don’t Lose Your Agent’s Mind: Platform Patterns for Stateful Agents

- Room: Salt Palace | Level 1 | Ballroom BDF
- Speakers: Evaline Ju, Maia Iyer
- Track: Platform Engineering Day
- Labels: Any Level, Breakout Session, Building Platforms for Day 2 and Beyond
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1266322

Agents accumulate context and session history over hours or days, but all of it can vanish in an instant. We deployed multiple agent harnesses like Claude Code and Hermes on Kubernetes and found they all break in similar ways: unrecoverable sessions and silent state loss. Existing harnesses assume stable filesystems and process continuity - assumptions Kubernetes violates. This infrastructure mismatch is a problem no individual harness will fix.

This talk presents platform patterns for managing agent state. We cover where existing primitives stop and what agents actually need from platforms, like durable sessions that survive node failure and state that follows the workload from local development to cloud. We demonstrate crash recovery progressing from “everything is lost” to transparent session resume and outline the future of state governance. Participants will leave with an adoptable Kubernetes-native reference architecture to support state-rich agentic workloads.

### 10:40 AM–11:05 AM · Don’t Make a Mesh of It: Multi-Cluster Networking at Cloud Scale

- Room: Salt Palace | Level 1 | Ballroom J
- Speakers: Anubhab Majumdar, Vamsi Kalapala
- Track: CiliumCon
- Labels: Breakout Session, Intermediate, Meshes: How Cilium Enables Seamless Multi Hybrid Cloud Connectivity + Modernizing Enterprises
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1268169

Multi-cluster networking is moving from demos to production on public clouds, and the MCS-API is making the same shift to real workloads. At cloud scale this surfaces hard problems such as meshes spanning many large clusters, mixed network topologies, and operators who need precise control over what crosses a cluster boundary. Using Cilium Cluster Mesh as a case study, this talk examines two layers that make it practical. For sharing, service export is opt-in and scoped: per service through ServiceExport, or per namespace, so cross-cluster exposure is exactly what platform teams choose rather than on by default. For the data path, it shows how to align routing with topology, using native routing to avoid encapsulation on flat cloud networks instead of tunneling everything, so large meshes don’t pay VXLAN or Geneve overhead they don’t need. You’ll leave able to reason about the export model and the routing and scaling tradeoffs behind production multi-cluster networking.

### 11:15 AM–11:40 AM · Agentic GitOps: Agent and Sandbox Guardrails with Flux

- Room: Salt Palace | Level 1 | Ballroom H
- Speakers: Leigh Capili, Tamao Nakahara
- Track: FluxCon
- Labels: AI-Powered GitOps with Flux, Breakout Session, Intermediate
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269797

With agentic workflows, kubectl commands can have dangerous consequences. Improper RBAC can allow agents to kubectl delete, and that includes deleting your whole CI with no commits to roll back to!
That’s why Flux's security-first design is even more relevant for agentic GitOps. We'll cover how to use Flux to confine agents to human-reviewable PRs for all sorts of use cases. We’ll do this with a kernel-sandboxing tool called `nono`. In addition, for your agent management tool of choice, we'll cover how to manage your agents' sandboxes with Flux so that nefarious (or confused) agents can't destabilize the security policies that you have in place.
We’ll cover safe practices for agentic use cases like:
- using Prometheus metrics to trigger resource tuning
- troubleshooting and rolling back after HPA crashes
- agents requesting additional network access with PR’s for human reviewers

Come join in!

### 11:15 AM–11:40 AM · Building Governance for Hybrid Architectures for Agentic Systems

- Room: Salt Palace | Level 1 | Ballroom ACE
- Speakers: Valentina Rodriguez Sosa, Daniele Zonca
- Track: Agentics Day: MCP + Agents
- Labels: Any Level, Breakout Session, Enterprise Integration + Governance
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269246

Autonomous AI agents are shifting to dynamic orchestrators that switch between local open-source models and external closed LLMs. When an agent autonomously escalates complex requests to public APIs or falls back to a local model, managing data sovereignty, routing rules, and governance becomes a critical platform engineering hurdle.

This presentation offers an end-to-end traffic governance framework for hybrid agentic loops using CNCF infrastructure. We showcase a declarative architecture using Kuadrant to build a secure AI routing tier. By coupling SPIFFE/SPIRE workload identities with Kuadrant's authorization layer, we ensure external API keys are never exposed and only validated agents can invoke specific model endpoints managed by Kubeflow's KServe. Finally, we demonstrate how guardrails play an important role in governance, ensuring payload verification across both local and cloud-hosted model targets.

### 11:15 AM–11:40 AM · Yes, It Is Possible to Create Stunning User Interfaces for Backstage (and Mobile Apps Too!)

- Room: Salt Palace | Level 2 | 255 EF
- Speakers: Olivier Liechti, Ana Marques
- Track: BackstageCon
- Labels: Beginner, Breakout Session, Technical Deep Dives on Backstage Topics
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269177

This should be a fun session, structured in two parts.

First, Ana will explain how she, as a graphical designer, is able to create applications on top of Backstage with a UX that departs from the "out-of-the-box" UI. She will demo alternatives to the opinionated, composable Backstage frontend framework that we all know (long live the nav bar and tabbed entity page).

Olivier will then introduce the "Falcon" architecture, designed to foster the creativity of UI engineers by removing some barriers of the Backstage framework. It will be about technical patterns and techniques, but also about ways to introduce Backstage concepts and capabilities to colleagues without deep DevOps and cloud native literacy. And yes... it will involve some agent skills as well.

We hope that the session will be both entertaining and actionable. It should give a fresh look at what a developer portal can look like, and be specific enough for Backstage teams to replicate the approach in their organizations.

### 11:15 AM–11:40 AM · Securing the Repo Server: A Practical Guide to Argo CD’s New Native mTLS

- Room: Salt Palace | Level 1 | 151
- Speakers: Patroklos Papapetrou
- Track: ArgoCon
- Labels: Breakout Session, Intermediate, Scalability
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1265272

The argocd-repo-server is a high-value target—it handles your Git credentials and generates your Kubernetes manifests. If an attacker gains a foothold in your cluster, unencrypted internal gRPC traffic between Argo components becomes a massive liability. Previously, securing this communication path meant deploying and managing a heavy service mesh just for Argo CD.

This session walks through Argo CD's new native Mutual TLS (mTLS) capability. We will skip the security theory and jump straight into the actual configuration. We’ll cover how to wire up the argocd-cmd-params-cm configmap, evaluate the practical trade-offs between shared and per-component client certificates, and handle certificate rotation without dropping live sync operations. You’ll leave with a straightforward guide to locking down your control plane using nothing but built-in upstream features.

### 11:15 AM–11:40 AM · Cache Me If You Can: Optimizing Expensive Tasks in Argo DAGs

- Room: Salt Palace | Level 2 | 251 A-C
- Speakers: Chetan Sharma
- Track: ArgoCon
- Labels: Breakout Session, Data Processing, Intermediate
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1268500

AI and data workflows often run as DAGs in Argo Workflows. Related runs can repeat expensive steps such as OCR, embedding generation, document parsing, and model inference. Yet cache eviction is often based only on recency, which can discard an expensive result while retaining a cheaper one.

This session presents a Kubernetes-native caching proxy for Argo Workflows. The proxy runs alongside workflow tasks and uses task metadata to prioritize cached results by re-computation cost, DAG dependency count, and invocation frequency. Attendees will see the proposed architecture, cache metadata, eviction flow, and an Argo deployment pattern for sharing reusable task results across related workflow runs. The approach is based on independent research that reported up to 51% lower latency than LRU in multi-agent benchmarks. The session translates that idea into a practical pattern for Argo users building AI and data pipelines on Kubernetes.

### 11:15 AM–11:40 AM · Kueue, DRA, and WAS: The Kubernetes Scheduling Triangle

- Room: Salt Palace | Level 2 | 254
- Speakers: Sohan Kunkerkar, Michał Woźniak
- Track: Cloud Native AI + Inference Day
- Labels: Breakout Session, End-to-End AI/ML in Production (MLOps/AIOps), Intermediate
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1262907

Modern AI workloads expose two gaps in Kubernetes scheduling. First, accelerators are no longer simple device counts: a MIG 1g.10gb slice and a 7g.80gb partition may count as “1 GPU” despite using vastly different capacity. Second, AI jobs need gang scheduling and topology-aware placement. kube-scheduler addresses these through Dynamic Resource Allocation (DRA) for accelerator modeling and Workload-Aware Scheduling (WAS) for workload-level placement. But clusters need quota and queuing, which Kueue provides.

We present Kueue’s integration with DRA, using Partitionable Devices to charge quota in units like GPU memory, not device count. MIG profiles sharing one quota pool can be charged by actual use.

Finally, we introduce scheduler-library, exposing kube-scheduler logic as compile-time building blocks for orchestrators like Kueue and Cluster Autoscaler. For Kueue, it combines scheduler-native accuracy with fast scheduling and preemption simulations for multi-tenant quota management.

### 11:15 AM–11:40 AM · llm-d Observability: Observing LLM Inference at Scale on Kubernetes

- Room: Salt Palace | Level 2 | 251 D-F
- Speakers: Guangya Liu, Nicole Xin
- Track: Cloud Native AI + Inference Day
- Labels: Any Level, Breakout Session, End-to-End AI/ML in Production (MLOps/AIOps)
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1262963

Production LLM inference on Kubernetes is distributed by design: requests flow through gateways, llm-d Router (EPP), KV-cache indexers, vLLM workers, and Prefill/Decode sidecars. Traditional metrics miss token-level latency, cache locality, and routing decisions, and that gap quietly erodes the 3x throughput gains from cache-aware scheduling.
This talk shares llm-d's (CNCF Sandbox) Observability: a component-level audit of metrics, traces, logs, dashboards, and alerts across the stack. We show what is already working (EPP and vLLM OTel, Prometheus scrapers, inference-sim for GPU-free CI) and the critical gap: cross-component trace linkage, zero pre-built dashboards, and opaque P/D KV transfers.
Attendees will leave with practical patterns: W3C TraceContext propagation, OTel Collector configs, tiered KV-cache observability, WVA autoscaling visibility, and a prioritized community roadmap for future release. Applicable beyond llm-d to any Kubernetes-native GenAI serving platform.

### 11:15 AM–11:40 AM · Scaling OpenTelemetry in a Regulated Enterprise: Lessons from CapitalOne's Journey

- Room: Salt Palace | Level 2 | 250 A-C
- Speakers: Sateesh Mamidala, Chandra Siva
- Track: Observability Day
- Labels: Breakout Session, End User Stories: Production Deployments + Migrations + Operations, Intermediate
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269593

Capital One has spent the past five years evolving OpenTelemetry from an early proof of concept into a vendor-agnostic, production-grade observability platform running across one of the largest cloud-native environments in financial services. Because Capital One operates as a fully regulated financial institution, that evolution also meant building compliance and audit requirements directly into the platform.
This talk covers the architectural redesigns driven by cost and value tradeoffs, and the agent extensions built to close gaps with commercial tooling. It also covers how Capital One rolled OpenTelemetry out across 85,000+ compute instances, 100,000+ Lambda invocations, and thousands of Fargate and Kubernetes workloads, without disrupting the experience application teams already relied on.
Attendees leave with lessons on phasing a large-scale rollout, building regulatory guardrails into an observability platform, and driving adoption across hundreds of application teams.

### 11:15 AM–11:40 AM · Endpoint Observability with OpenTelemetry Collector at Cisco IT

- Room: Salt Palace | Level 2 | 250 D-F
- Speakers: Alec Chamberlain, Aaron Nelson
- Track: Observability Day
- Labels: Beginner, Breakout Session, End User Stories: Production Deployments + Migrations + Operations
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1241534

Employee laptops are production systems for IT, but most observability patterns still assume servers, clusters, or services. This session shares how Cisco IT used a custom OpenTelemetry Collector build to observe macOS, Windows, and Linux endpoints without turning every laptop into a high-cardinality telemetry source. The talk walks through the collector architecture: hostmetrics for CPU, disk, memory , and selected processes; httpcheck, tcpcheck, icmpcheck, and a custom pathcheck receiver for SaaS, Webex, DNS, proxy, VPN, and route diagnostics; and a loadlevel processor that keeps detailed process metrics only when CPU or memory pressure makes them useful. Attendees will see the trade-offs behind collection intervals, MTS reduction, route snapshots as logs instead of metric dimensions, privacy controls, and safe rollout profiles for large employee fleets.

### 11:15 AM–11:40 AM · Why Post-Quantum Authentication Will Break K8s (and How Merkle Trees Fix It)

- Room: Salt Palace | Level 2 | 255 BC
- Speakers: Fabian Kammel
- Track: Open Source SecurityCon
- Labels: Breakout Session, Intermediate, Security Education
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1247270

The industry is actively working to secure our key exchanges against quantum threats, but what about our identities? As industry leaders update their PQC migration timeline to 2029, the focus is urgently shifting from encryption to the untackled authentication challenge.
Traditional PKIs retrofitted with ML-DSA face crippling performance degradation from massive signatures and compute overhead. So, how do we secure K8s’ internal communication without grinding the control plane to a halt?
In this talk we:
+ detail the reality of (lacking) ML-DSA support across the ecosystem which makes even experimenting with PQC incredibly difficult
+ analyze the critical compute and size bottlenecks introduced by standard ML-DSA signatures with a local PKI and hard numbers
+ explore Merkle Tree Certificates as a high-performance alternative to side-step PQC signature limits
+ demonstrate how to run an MTC PKI locally and map out how K8s can adopt this approach for its internal certificate requirements

### 11:15 AM–11:40 AM · The Carbon-Aware Platform: Building Sustainable Paved Paths with Kepler and KEDA

- Room: Salt Palace | Level 1 | Ballroom BDF
- Speakers: Sakshi Nasha, Primanshu Choudhary
- Track: Platform Engineering Day
- Labels: Any Level, Breakout Session, Improving Platform Maturity
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269395

As developer platforms scale, their cloud footprints quickly become massive operational cost-centers and drivers of carbon emissions. True platform maturity requires optimizing for environmental sustainability alongside performance.

This session delivers a technical blueprint for constructing a green platform using open-source tooling. We will demonstrate how to deploy Kepler with eBPF to capture granular, pod-level energy consumption. Moving past mere visibility, we show how to hook these metrics into Prometheus and leverage KEDA alongside the Carbon Aware SDK to drive automated, carbon-intelligent behaviors such as dynamically shifting non-critical workloads to periods of low grid intensity or downscaling idle environments. Attendees will leave with actionable strategies to transform sustainability into an automated architectural default within their platform’s paved paths.

### 11:15 AM–11:40 AM · When Packets Disappear in the Cloud: Debugging Inference Workloads with eBPF

- Room: Salt Palace | Level 1 | Ballroom J
- Speakers: Venkat Gattupalli, Haibing Zhou
- Track: CiliumCon
- Labels: Any Level, Breakout Session, Use Cases: End User Story Describing the Use of Cilium + Hubble + Tetragon to Overcome Challenges
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1248003

Large-scale inference workloads are extremely sensitive to rare packet loss: one missing packet can cause seconds of tail latency or failed requests. The hardest cases sit at the boundary between what we control and what we cannot directly observe: guest kernel, vNIC driver, hypervisor, and cloud-provider network.

In this session, we share how eBPF helped us debug and mitigate latency issues in inference workloads running at OpenAI scale. This is an incident-response story, not a generic tutorial. Starting from latency symptoms and TCP retransmissions, we used targeted probes to trace packets through the networking stack and vNIC driver. By correlating TCP sequence numbers, skb lifecycle, DMA mappings, and completion paths, we proved affected packets were leaving the guest driver path and narrowed the issue toward cloud infrastructure. That evidence enabled a shadow-traffic workaround, validated latency improvement, and built rollout confidence while the provider fix was underway.

### 11:40 AM–12:45 PM · Lunch 1

- Room: TBA
- Track: Agentics Day: MCP + Agents
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1309415

### 11:45 AM–11:55 AM · Running Out of IPs Before You Run Out of GPUs: IPv6 for AI-Scale Kubernetes

- Room: Salt Palace | Level 1 | Ballroom J
- Speakers: Marino Wijay
- Track: CiliumCon
- Labels: Breakout Session, Intermediate, Technology: Talks Relating to Cilium’s Architecture, Including New Features + Implementation
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269247

You've provisioned the GPU nodes and sized the power, cooling, and storage. But have you planned the IP addresses? In a dense AI cluster, with dozens of pods per GPU node across a large fleet, a private range that looked endless on the whiteboard runs dry fast.

It usually isn't public IPv4 scarcity that forces the move. It's the quiet exhaustion of RFC1918 space inside the flat, multi-account networks AI platforms grow into. Add autoscaling and constant rebuilds, and your address plan becomes what caps your scale.
So is IPv6 the answer, and is it worth the cost? This talk argues that for AI-scale Kubernetes, IPv6 is the substrate that makes the architecture work. We'll look at teams running it in production, stay honest about the sharp edges, and close with a live demo: IPv6 IP pools advertised over BGP, plus the capacity math to size it yourself.

### 11:50 AM–12:15 PM · Your VMs Are Just CRDs: Git-Native Virtual Machine Fleets with Flux

- Room: Salt Palace | Level 1 | Ballroom H
- Speakers: William Rizzo
- Track: FluxCon
- Labels: Breakout Session, Intermediate, Running Flux at Scale in Production
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269880

Kubernetes operators that expose every surface as a CRD get GitOps for free, and that includes virtual machines. KubeSwift runs VMs as pods on Cloud Hypervisor, defined entirely by CRDs, which means Flux can drive an entire VM platform; install, shared infrastructure, and running fleets, using nothing but OCIRepository, HelmRelease, and Kustomization.
But VMs aren't containers. Image imports take minutes, root disks are stateful, and live-migration physically moves a guest to another node, which can quietly fight Flux's drift correction. This talk walks the working three-layer model (platform → infrastructure → workloads, ordered with Flux primitives), demos scaling and rolling a VM fleet by Git commit, and then gets honest about the sharp edges: which day-2 operations belong in Git, which must stay imperative, and why knowing that boundary is what keeps Flux and your VM controllers from working against each other.

### 11:50 AM–12:15 PM · What the Scheduler Doesn't Know: Why Your Inference Optimizations Aren't Working

- Room: Salt Palace | Level 2 | 254
- Speakers: Sheetal Sriram
- Track: Cloud Native AI + Inference Day
- Labels: Breakout Session, End-to-End AI/ML in Production (MLOps/AIOps), Intermediate
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1267643

You've set up KV-cache-aware routing, tuned your batch scheduler, and sized your GPU pods carefully. Latency is still off. Throughput falls short of what benchmarks suggested. The dashboard looks fine.
This talk is about what's actually happening below the Kubernetes layer and how to see it.
CPU scheduling jitter during decode, NUMA-crossing memory traffic eating into KV cache throughput, interrupt affinity defaults quietly inflating time-to-first-token - these aren't theoretical concerns. They show up repeatedly across serving engines and they're invisible to standard cluster observability.
Drawing on experience profiling workloads across different inference engine in both resource-constrained clusters and production deployments, the talk walks through a practical profiling workflow using nvidia-smi dmon, numastat, and eBPF tracers to find where the time is actually going and which Kubernetes scheduling knobs are worth pulling once you know. (No kernel background needed)

### 11:50 AM–12:25 PM · Cloud Native Agentic Standards: A Progress Report from the CNCF AI TCG

- Room: Salt Palace | Level 2 | 251 D-F
- Speakers: Vincent Caldeira, Josh Halley, Nimisha Mehta, Nina Polshakova
- Track: Cloud Native AI + Inference Day
- Labels: Any Level, Panel Discussion, RAG + AI Agents
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1249491

As agentic AI expands across the cloud native ecosystem, a lack of uniform standardization and interoperability threatens secure enterprise scaling.

In this session, co-authors from the CNCF AI TCG provide an essential structural and operational update based on our recent foundational standards checklist for Kubernetes-driven workloads.We will break down the core pillars required for reliable multi-agent systems: communication protocols like MCP and A2A , identity and authorization frameworks including SPIFFE/SPIRE, granularity in time-series observability metrics , and flexible governance models designed to mitigate emergent AI behaviors.

Attendees will walk away with an agnostic view of cloud-native best practices to deploy securely, scale reliably, and ensure explainability across distributed agent networks.

### 11:50 AM–12:15 PM · Inference Is a Black Box: How We Built GPU-Aware Observability with Fluent Bit and OpenTelemetry

- Room: Salt Palace | Level 2 | 250 A-C
- Speakers: Henrique Santana, Danilo Souza
- Track: Observability Day
- Labels: Breakout Session, Correlation + Context + Cross-Project Architectures, Intermediate
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1247614

GPU inference on Kubernetes generates three disconnected signal types: hardware metrics from DCGM, structured serving logs, and distributed traces. Teams discover GPU memory pressure from pages, not dashboards, because nothing ties these signals together at the pod level.
This talk presents the observability pipeline we built to close that gap. We use dcgm-exporter in Kubernetes mode to map GPU devices to pods, Fluent Bit to collect and enrich inference logs with pod metadata, and the OTel Collector to assemble traces across instrumented model-serving hops. A Grafana dashboard joins all three by pod identity, surfacing which models are over-provisioned and which are memory-constrained.
We cover what broke (Prometheus cardinality explosion from per-pod GPU metrics), the Fluent Bit Lua parsing that extracts token throughput from serving logs, and the OTel Collector config that makes trace-to-GPU correlation queryable. You leave with deployable configs for all three components.

### 11:50 AM–12:15 PM · OpenTelemetry Graduation Milestone Accomplished! What’s Next?

- Room: Salt Palace | Level 2 | 250 D-F
- Speakers: Alolita Sharma, Liudmila Molkova, Austin Parker, Morgan McLean
- Track: Observability Day
- Labels: Any Level, Panel Discussion, Project Deep Dives
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1268202

OpenTelemetry (OTel) has officially graduated, solidifying its place as the industry standard for observability. But graduation isn’t the finish line—it’s just the launchpad.
Join OTel maintainers and governance committee members for a forward-looking panel on what’s next for the ecosystem. We’ll cover the roadmap shaping the future of observability, including the latest on continuous profiling and how OTel is evolving to meet the massive demands of GenAI infrastructure.
Bring your questions and get a first look at the initiatives that will define the next decade of cloud-native observability.

### 11:50 AM–12:15 PM · Set the Terms for Your Agents and Let Your Humans Be Creative: How to Govern Code Nobody Reviewed

- Room: Salt Palace | Level 1 | Ballroom BDF
- Speakers: Thomas Schuetz, Dr. Constanze Roedig
- Track: Platform Engineering Day
- Labels: Any Level, Breakout Session, Platform Engineering and AI
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1255381

AI-assisted coding is changing how software gets built, and not just for engineers. When anyone can generate an application in minutes, the real question becomes how to get it deployed, in line with policies and without risking your company's integrity.

Using AI to translate ideas into code should not be limited by fear

This talk gives platform owners a framework to bridge this gap. We show how to create “desired state” abstractions that enable to safely develop, deploy and assert properties of AI-generated code. We show that leveraging feedback from observability and operational policies actually helps the vibe-coders enjoy their coding more, and avoids the danger of rifts and enmity between teams.

We demo concrete examples of vibe-coded apps, initially violating both compliance (SLSA, GDPR) as well as corporate guidelines (ADRs).

By the end, you will have a means to ensure that rapid experimentation doesn't lead to chaos and that vibe coding is helpful and fun for everyone.

### 12:00 PM–12:10 PM · One Kernel, Two eBPF Stacks: Running OBI Alongside Cilium

- Room: Salt Palace | Level 1 | Ballroom J
- Speakers: Antonio Jimenez Martinez
- Track: CiliumCon
- Labels: Intermediate, Use Cases: End User Story Describing the Use of Cilium + Hubble + Tetragon to Overcome Challenges, ⚡ Lightning Talk
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1255649

As eBPF-based observability matures, platform teams are starting to run multiple eBPF systems on the same Kubernetes nodes. Cilium provides networking, security, and network observability, while OpenTelemetry eBPF Instrumentation (OBI) provides zero-code application telemetry.

This session explores what happens when OBI and Cilium run together in production. We will compare where each system attaches to the kernel, how traffic-control ordering, TCX, Netlink fallback, and attachment priority affect visibility, and what operators should measure when enabling both.

Using benchmarks, demos, and production lessons, attendees will learn how to combine OBI-generated application telemetry with Cilium and Hubble network insights while managing CPU, memory, telemetry volume, and operational complexity.

### 12:15 PM–12:25 PM · A One-Way Valve to netkit: Migrating Cilium's Datapath Transparently at Scale

- Room: Salt Palace | Level 1 | Ballroom J
- Speakers: Hadrien Patte
- Track: CiliumCon
- Labels: Beginner, Use Cases: End User Story Describing the Use of Cilium + Hubble + Tetragon to Overcome Challenges, ⚡ Lightning Talk
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1264636

Cilium's datapath-mode can't be changed on a running node: switching from veth to netkit means recreating every node. How do you migrate hundreds of thousands of nodes across hundreds of production Kubernetes clusters without the workloads on top noticing?

Discover how Datadog designed and carried out that migration as a one-way valve. Nodes moving to netkit one after the other, locking in migrated nodes. Rollback available at any point, bounding the blast radius of any regression to the fraction of the fleet already migrated.

Attendees will leave with a concrete playbook for large-scale Cilium datapath migrations, transferable to any other immutable per-node changes.

### 12:25 PM–12:50 PM · Your Git Repo Is Not a Backup: Kubernetes Disaster Recovery with Flux and Velero

- Room: Salt Palace | Level 1 | Ballroom H
- Speakers: Rushabh Shah, Shivani Rathod
- Track: FluxCon
- Labels: Breakout Session, Intermediate, Running Flux at Scale in Production
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269417

Imagine a pager alert telling you your cluster is completely wiped out. Sure, your workloads and configs are safe in Git, but what about your database PVs and critical secrets? GitOps has changed how we manage infrastructure, but Git alone is not a disaster recovery strategy. 
This session explores what it actually takes to recover from a catastrophic Kubernetes failure using Flux and Velero. Through an intentional, live breakdown, we’ll walk the path from total outage back to a healthy state. Along the way, we’ll discuss the undeniable charm of GitOps, as well as its limitations when things go sideways. 
We’ll look at how backup and recovery tools map to your GitOps workflow to ensure your clusters run like clockwork. You’ll walk away with a repeatable DR approach for platforms of any size helping you prepare for the days when luck just isn’t on your side.

### 12:25 PM–12:30 PM · Closing Remarks

- Room: Salt Palace | Level 1 | Ballroom J
- Track: CiliumCon
- Labels: Any Level
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1312055

### 12:25 PM–1:30 PM · Lunch 2

- Room: TBA
- Track: Cloud Native AI + Inference Day
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1309413

### 12:30 PM–1:30 PM · Lunch 3

- Room: TBA
- Track: CiliumCon
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1309414

### 12:45 PM–1:20 PM · The Golden Path Doesn't Exist Yet: What Platform Teams Told Us About OpenSSF Tools

- Room: Salt Palace | Level 2 | 255 BC
- Speakers: Katherine Druckman, Stacey Potter, Kadi McKean, Tabatha DiDomenico
- Track: Open Source SecurityCon
- Labels: Any Level, Panel Discussion, Security Advocacy + Collaboration
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1268994

At cdCon, the OpenSSF DevRel community ran an open-floor session with no slides and no pitches and asked platform teams a blunt question: what's standing between "this is a good idea" and "this is actually working for us"? The answers were specific and directly relevant to anyone running build and release pipelines.

This talk reports what we heard and connects it to the projects meant to solve it. We'll walk through the real adoption blockers — SBOM accuracy and portability, air-gapped and regulated-environment friction, the awareness-vs-adoption gap, and badge/wayfinding confusion — and then look at how OpenSSF tooling does and doesn't resolve them today: Sigstore for signing and verification, Zarf for offline package delivery, GUAC for correlating SBOMs and attestations, and SLSA's Build and Source tracks.

### 12:50 PM–12:55 PM · Closing Remarks

- Room: Salt Palace | Level 1 | Ballroom H
- Track: FluxCon
- Labels: Any Level
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1312058

### 12:55 PM–1:20 PM · Orchestrating, Discovering, and Autoscaling Distributed Multi-Agent AI Systems Natively in K8s

- Room: Salt Palace | Level 1 | Ballroom ACE
- Speakers: Derek Wang
- Track: Agentics Day: MCP + Agents
- Labels: Any Level, Breakout Session, Performance + Scaling
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1267474

As we move from single-agent applications to distributed multi-agent systems, new challenges emerge. How do you model coordination patterns such as supervisor, handoff, and sequential workflows? How do agents dynamically discover each other and stay in sync as peers join, leave, or change their capabilities? And how do you autoscale I/O-bound agent workloads when CPU utilization is a poor signal?

This talk presents K8s-native patterns for running prod-grade multi-agent systems built on the A2A protocol. We will explore how CRDs and operators can model agent coordination across different frameworks (e.g. LangGraph or ADK), how agents can automatically discover peers and adapt to capability changes, and why traffic-aware autoscaling is more effective than traditional CPU-based HPA for agentic workloads.

Attendees will leave with practical design patterns, architectural tradeoffs, and production lessons for building reliable, scalable, framework-agnostic multi-agent systems on K8s.

### 12:55 PM–1:20 PM · The Agent Did What? Reconstructing MCP Tool-Call Attacks

- Room: Salt Palace | Level 1 | Ballroom IG
- Speakers: Jie Wu
- Track: Agentics Day: MCP + Agents
- Labels: Breakout Session, Intermediate, Security + Trust + Reliability
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1249143

Your AI agent reads a web page. Hidden in the content is an instruction meant for the agent, not for you: indirect prompt injection. Minutes later, the agent uses MCP to retrieve customer data and send it outside your environment. You are the responder. What happened, and can you prove it?

No single log tells the story. The cloud audit log shows what resource was accessed. The MCP tool-call log shows which tools ran, what data they handled, and where it went. Matched on time, resource, and content, they reconstruct the chain.

This talk reproduces the attack in a controlled demo, then rebuilds it step by step: this tool call, this data, this moment. Tool calls are captured by a minimal MCP audit proxy that emits OpenTelemetry spans to Jaeger, turning the incident into a trace. Attendees leave with a practical method to reconstruct MCP attacks, join tool-call evidence to cloud audit logs, and understand what these logs can and cannot prove.

### 12:55 PM–1:20 PM · Treating Backstage Like a Product: Observability with OpenTelemetry

- Room: Salt Palace | Level 2 | 255 EF
- Speakers: Ekansh Gupta
- Track: BackstageCon
- Labels: Any Level, Breakout Session, Transforming Dev Experience + Streamlining Dev Workflow + Doing More with Less Using Backstage
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1268774

Backstage has become the foundation of many Internal Developer Platforms, providing a unified experience for service discovery, documentation, and workflows. Yet while teams invest heavily in building these experiences, many lack visibility into how their platforms are performing, which capabilities are being adopted, and where developers encounter friction

In this session, we’ll explore how to instrument Backstage with OTel to achieve deep visibility into its performance, reliability, and adoption. Attendees will learn how to collect, export, and visualize metrics that help platform teams measure usage patterns, plugin health, and system performance, all using open-source tools

We’ll walk through:
1. Foundational Concepts: Why observability matters for developer platforms.
2. Hands-On Implementation: Step-by-step instrumentation of Backstage’s backend using OTel.
3. Integration and Visualization: Setting up an OTel Collector and integrating with metrics backends for live dashboards.

### 12:55 PM–1:20 PM · Observability Is a Platform Problem Now: Closing the AIOps Loop with Argo CD & GitOps

- Room: Salt Palace | Level 1 | 151
- Speakers: Luke Philips, Hilliary Lipsig
- Track: ArgoCon
- Labels: Any Level, Breakout Session, Observability
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269302

Every team is told to "use AI," with wildly different goals. Business wants dashboards on user behavior and growth. Reliability teams need data-driven decisions and high-quality alerts. Developers want to catch a regression before it ships. Product wants true feature velocity. Different problems, one shared dependency: rich, trustworthy, well-structured observability data. An AI initiative is only as good as its data, and most teams sit on telemetry never built for an agent to consume.

That is a delivery problem too. ArgoCD knows what shipped, when, and whether it's healthy, but that signal dies in Slack. We make the GitOps loop the actuator for AIOps: deploy and sync events normalized as CDEvents into your observability platform; HolmesGPT triggered by ArgoCD Notifications on a failed rollout; an automated root-cause pass; a corrective PR ArgoCD reconciles, human in the loop, Git as the audit trail. Every agent action is a reviewable, revertible commit.

### 12:55 PM–1:20 PM · Benchmarks as Specs: Reproducible LLM Evaluation Pipelines on Argo Workflow

- Room: Salt Palace | Level 2 | 251 A-C
- Speakers: Tatsuhiro Chiba
- Track: ArgoCon
- Labels: Any Level, Breakout Session, Data Processing
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269829

Do you trust your LLM benchmark numbers? You probably shouldn't. Run the "same" benchmark twice and you often get two different answers. Results depend on far more than the model — the client tool and its load profile, the hardware accelerator, the inference server, batch sizes, prompt sets. Most setups bury these knobs in engine code and scripts, so results can't be reproduced, compared, or trusted. LLM benchmarking is broken.

We fix it with a spec-driven approach on Argo Workflows: Benchmarks as Specs. The engine and its configuration live in separate repos — the engine knows how to run, while versioned, declarative specs define what to measure. Argo Workflows turns each spec into a reproducible, auditable run. Because every variable is a spec, the pipeline is portable: swap the hardware accelerator, or switch the client tool — GuideLLM, inference-perf, aiperf — by editing a spec, not the engine.

We'll walk the architecture, demo live spec swaps, and share a pattern you can reuse.

### 12:55 PM–1:30 PM · Lunch 4

- Room: TBA
- Track: FluxCon
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1309427

### 1:20 PM–1:25 PM · Welcome + Opening Remarks

- Room: Salt Palace | Level 1 | Ballroom H
- Track: OpenTofu Day
- Labels: Any Level
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1312042

### 1:20 PM–1:25 PM · Welcome + Opening Remarks

- Room: Salt Palace | Level 2 | 255 A
- Track: Kubernetes on Edge Day
- Labels: Any Level
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1312038

### 1:20 PM–1:55 PM · The Missing Layer of Platform Maturity: Kubernetes Networking

- Room: Salt Palace | Level 1 | Ballroom BDF
- Speakers: Faseela Kundattil, Christophe Sauthier, Chad M. Crowell, Marino Wijay, James Spurin
- Track: Platform Engineering Day
- Labels: Any Level, Panel Discussion, Platform Team Composition
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1248801

Kubernetes networking simplifies distributed systems, yet many engineers still struggle to understand how abstractions map to real networking behavior. As the ecosystem matures, community-driven efforts like the Certified Kubernetes Networking Engineer (CKNE) exam are emerging to help standardize the real-world Kubernetes networking skills platform teams increasingly depend on.

This panel brings together CKNE SMEs, networking specialists, and educators to challenge how Kubernetes networking is understood, taught, and operationalized today. We’ll explore where mental models break down across services, DNS, CNIs, and the Linux networking stack, why networking knowledge remains difficult to scale across teams, and how these gaps impact platform maturity.

Join us as we explore the invisible layers powering Kubernetes and discuss how the community can build a more practical and standardized model for Kubernetes networking expertise.

### 1:30 PM–1:55 PM · An Overview of OpenTofu's Engine Rewrite

- Room: Salt Palace | Level 1 | Ballroom H
- Speakers: Laurence Bordowitz
- Track: OpenTofu Day
- Labels: Advanced, Breakout Session, OpenTofu Internals
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1266401

Out with the old and in with the new, a high level perspective on OpenTofu's supercharged new engine! For a little over a year, the OpenTofu maintainers have been working on a full engine rewrite that is proving to be faster and simpler than the current model. We'll discuss why the maintainers pushed for this change, how it compares to the existing engine/runtime, and what practical impacts it has on the day-to-day usage of the tool.

We will discuss the architecture of the current engine and compare it to the architecture of the new engine. After a short demo, we will show you how you can start using the new engine today!

### 1:30 PM–1:55 PM · Kubernetes at the Grid Edge: Platform Engineering Patterns for Deploying AI in Substations

- Room: Salt Palace | Level 2 | 255 A
- Speakers: Aditya Soni, Neha Soni
- Track: Kubernetes on Edge Day
- Labels: Any Level, Breakout Session, Futures
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269689

SEAPATH virtualizes the substation. GridFM and OpenSTEF promise AI-driven protection and forecasting at the edge. Between the two sits an unsolved deployment problem: how do you ship ML models to 600 substations and guarantee they never exceed the 15-watt thermal envelope of the host hardware?

This talk maps five missing platform layers between "model trained in the cloud" and "model running safely in a substation", and what each needs before edge AI scales past pilot deployments

- Lightweight Kubernetes distributions (K3s, MicroShift) for running inference alongside SEAPATH virtualized workloads
- Resource-constrained scheduling: fitting ML inference within thermal and power envelopes of substation hardware
- Fleet-wide observability: monitoring model drift, inference latency and more
- Model supply chain integrity. Sigstore Cosign signs every model artifact at build timeAttendees leave with a concrete platform blueprint for AI at the substation edge.

### 1:30 PM–1:55 PM · Unlocking Enterprise Knowledge for AI Agents at Scale with MCP

- Room: Salt Palace | Level 1 | Ballroom ACE
- Speakers: Spandana Balumuri, Nirav Jain
- Track: Agentics Day: MCP + Agents
- Labels: Breakout Session, Extending AI Systems with MCP, Intermediate
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269003

Enterprise AI agents are limited by knowledge access, not model capability. A coding assistant that can't read internal API docs is just an expensive autocomplete. Most enterprise documentation lives behind SSO with multi-factor auth, spread across dozens of internal websites. MCP gives agents a standard protocol to query it, but what does the full production architecture look like?

This session presents an architecture blueprint for serving enterprise knowledge via MCP, centered on FastMCP. It covers MCP design patterns for scale: a two-phase tool pattern where agents resolve which sites to query then filter results enabling multi-tenancy. It also covers how knowledge stays current using a distributed ingestion pipeline on Ray that crawls auth-protected sites with authenticated browser sessions and generates embeddings to Milvus.

Attendees leave with an adoptable open-source architecture (Kubernetes, Ray, FastMCP, Milvus, Prometheus) and guidance on production deployment challenges.

### 1:30 PM–1:55 PM · Total Autonomy on Local Silicon: Building Production-Grade Agentic Workflows Completely Offline

- Room: Salt Palace | Level 1 | Ballroom IG
- Speakers: Muskan Jain
- Track: Agentics Day: MCP + Agents
- Labels: Any Level, Breakout Session, Extending AI Systems with MCP
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269757

Cloud AI agents bring fragile internet dependencies, high bills, and data privacy risks. For enterprise workflows, the future is edge-first and fully local.This session demonstrates how to build and extend autonomous, long-horizon agents that run completely offline on developer hardware. We dissect the architecture of a local agent runtime, leveraging open-source components like Ollama/llama.cpp alongside MCP. Attendees will learn how to bypass local VRAM and context limits using progressive tool discovery, implement secure local stdio transport layers to eliminate network overhead, and build deterministic "human-in-the-loop" approval gates to mitigate risk. The session concludes with a live, 100% offline demonstration showing a local model using an MCP server to diagnose system issues, write a patch inside a secure sandbox, and automatically generate its own reusable local skills without hitting external APIs.

### 1:30 PM–1:55 PM · Who Owns This? Automating Backstage Catalog Ownership at Scale

- Room: Salt Palace | Level 2 | 255 EF
- Speakers: Mateusz Wiszniewski, Konrad Przysucha
- Track: BackstageCon
- Labels: Breakout Session, Intermediate, Lessons Learned from Standing Up + Deploying + Driving Backstage Adoption
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1267923

Backstage catalogs rot. People leave, teams get reorganized, and entities end up orphaned or owned by groups that don't exist anymore. Nobody notices until an incident, when it turns out no one's accountable.

This is a case study of a backend plugin we built to fix that. It scans the catalog for orphaned and "zombie" ownership, then proposes new owners based on two signals: who's actually committing code, and where the entity sits in the org hierarchy. Teams accept or reject; if they stall, it escalates up the management chain.

We'll walk through the design, the extension points, and the parts we're still figuring out — contributor-to-team mapping, escalation logic, and a couple of nasty race conditions.

### 1:30 PM–1:55 PM · Progressive Delivery with Cilium Gateway API and Argo Rollouts

- Room: Salt Palace | Level 1 | 151
- Speakers: Christian Hernandez
- Track: ArgoCon
- Labels: Beginner, Breakout Session, Progressive Delivery
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1261386

Progressive delivery has traditionally required additional traffic management infrastructure such as service meshes, custom ingress controllers, or proprietary load balancers. As Kubernetes standardizes on Gateway API, that story is changing.

Cilium includes a production-ready implementation of the Kubernetes Gateway API, providing advanced traffic management capabilities directly within the networking stack. At the same time, Argo Rollouts has support for Gateway API integration, enabling canary deployments, traffic splitting, and targeted user experiences using standard Kubernetes APIs.

In this session, we'll demonstrate how these two projects naturally fit together to create a powerful progressive delivery platform with minimal operational overhead.

### 1:30 PM–1:55 PM · Self-Healing Canaries: OTel-Scoped Argo Rollouts Analysis and an AI SRE That Closes the Loop

- Room: Salt Palace | Level 2 | 251 A-C
- Speakers: Chamod Perera, Shivay Lamba
- Track: ArgoCon
- Labels: Breakout Session, Intermediate, Progressive Delivery
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269782

Based on real production experience operating an observability-driven delivery platform on CNCF projects, this session shows how OpenTelemetry becomes a core primitive for progressive delivery, and how Argo Rollouts turns telemetry into safe, automated rollbacks.
Running dozens of Kubernetes clusters at 60M+ requests per minute, we hit gray failures: builds passed, pods stayed Running, liveness was green, yet real users hit errors. No health signal was scoped to the specific build.
We thread a single buildid from CI (GitHub Actions) into every OTel span, log, and metric. An Argo Rollouts AnalysisTemplate runs PromQL scoped to that buildid and aborts the canary automatically when error rate breaches the SLO.
Then we add a cognitive layer: an AI SRE agent reads the same build_id-scoped Loki and Prometheus signals via the Grafana MCP server, names the offending commit, and opens a PR a human merges. We share the trade-offs and lessons learned.

### 1:30 PM–1:55 PM · GPU Portability Was Broken but Now There's WebGPU

- Room: Salt Palace | Level 2 | 254
- Speakers: Bailey Hayes, Mendy Berger
- Track: Cloud Native AI + Inference Day
- Labels: Any Level, Breakout Session, End-to-End AI/ML in Production (MLOps/AIOps)
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1268234

As ML inference, image processing, and scientific compute move into production Kubernetes environments, teams are forced to choose between vendor-specific runtimes, driver lock-in, and brittle operator configurations that break across cloud providers. The promise of portable, schedulable GPU compute remains largely theoretical.

By translating the WebGPU standard to run server-side as a language-agnostic, vendor-neutral interface called wasi:webgpu, GPU access becomes something a Wasm component can simply declare as a capability, and wasmCloud makes it schedulable.

In this session, we'll demonstrate ONNX models running inference through wasi:webgpu, using the GPU without any framework-specific operator requirements or driver assumptions. The component doesn't know whether it's running on an NVIDIA A100 in a data center or a consumer GPU in an edge cluster. Then we’ll go beyond AI, demonstrating GPU-accelerated image processing and scientific compute as ordinary distributed workloads.

### 1:30 PM–1:55 PM · Beyond Shipping the Model: Making Inference Boringly Repeatable

- Room: Salt Palace | Level 2 | 251 D-F
- Speakers: Prithvi Raj
- Track: Cloud Native AI + Inference Day
- Labels: Breakout Session, Intermediate, LLMs + Generative AI
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269619

Deploying a single model well is hard. Deploying hundreds, across mixed hardware, varying parametes, and competing inference engines, is a different problem entirely and most teams are solving it wrong.

Today, every production inference deployment seems to require a specialist who deeply understands the GPU, the engine (vLLM, SGLang, TensorRT-LLM, Dynamo), and the model architecture itself.

Organizations cannot hire a performance engineer for every team that wants to serve a model.

This session reframes inference scaling as an abstraction problem, not a performance problem. We'll explore how a small group of specialists can encode their knowledge hardware profiles, engine selection logic, autotuning, deployment topology into reusable platform primitives that the rest of the organization simply consumes.
The result: faster time-to-production, fewer trapped configs, lower TCO, and inference expertise that compounds across the org instead of bottlenecking on individuals.

### 1:30 PM–1:55 PM · Observing the stream: open source monitoring for video infrastructure

- Room: Salt Palace | Level 2 | 250 A-C
- Speakers: Erika Alvarez
- Track: Observability Day
- Labels: Any Level, Breakout Session, End User Stories: Production Deployments + Migrations + Operations
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1254983

Before DevOps was involved, our streaming platform had no visibility into live and on-demand video workflows. Errors were caught late, incidents escalated unnoticed, and commercial monitoring tools were too expensive.

This talk presents how a DevOps engineer solved that using open source tooling: a monitoring stack built with Grafana, Prometheus, Kubernetes, Helm, Argo CD and extended with FFmpeg and TSDuck to capture stream health signals at the transport layer.

Attendees will learn what makes streaming observability different from typical web platforms, how the stack was designed and deployed, and the concrete outcomes: earlier error detection, real-time visibility into live streams, and replacing costly commercial tools with a maintainable open source solution.

This is a practitioner case study. Every decision was driven by a real operational problem, and every outcome and learning opportunities we had will be shared honestly.

### 1:30 PM–1:55 PM · "What's the Dollar Cost of Slow?": Answering Business Questions with OTel

- Room: Salt Palace | Level 2 | 250 D-F
- Speakers: Kayla Reopelle, Hannah Ramadan
- Track: Observability Day
- Labels: Any Level, Breakout Session, Observability Cost + Quality + Scale + Efficiency
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1250719

Most teams instrument their services well enough to debug an outage. But what about when you’re eventually asked: what is a slow service actually costing us? This talk presents three custom-instrumentation patterns using OpenTelemetry traces, metrics, and logs that map technical signals to business outcomes.

### 1:30 PM–1:55 PM · MINT: A Fresh Take on Early Vulnerability Triage

- Room: Salt Palace | Level 2 | 255 BC
- Speakers: Jessy Ayala
- Track: Open Source SecurityCon
- Labels: Any Level, Breakout Session, Leveraging + Preparing for AI in OSS Security
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1235805

Recently, OSS maintainers have received an influx of vulnerability reports that have become time-consuming to triage. Some reports look technically plausible, but lack clear evidence, exaggerate severity, omit reproduction steps, or reference files, functions, and project behavior that do not exist. As AI-assisted vulnerability reporting becomes more prevalent, this type of noise can consume scarce maintainer time and delay responses to legitimate security issues. This session presents MINT, a Maintainer-INformed Triage framework for helping OSS projects evaluate noisy vulnerability reports. Rather than replacing maintainer judgment or attempting deep vulnerability discovery, MINT supports early triage by checking report actionability and whether the report is grounded in the target repository. Attendees will learn how maintainer practices can inform security tooling and how repository-grounded triage summaries can help projects prioritize credible reports while reducing review burden.

### 1:30 PM–1:55 PM · “Allowed but Unsupported:” When Platform Support Boundaries Meet Enterprise Reality

- Room: Salt Palace | Level 1 | Ballroom J
- Speakers: Niranjan Shankar, Amit Agarwal
- Track: Platform Engineering Day
- Labels: Any Level, Breakout Session, Platform Engineering Case Study
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269252

Platform engineers want guardrails. Developers want flexibility. So how do you safely add escape hatches along the golden path—and outline who’s responsible when things go wrong?

At AKS, customers like UBS needed capabilities our managed platform explicitly blocked or never exposed. Preventing advanced extensibility became untenable, leading us to an informal middle ground—“Allowed but Unsupported.” But saying “you’re on your own” wasn’t so simple. From EnvoyFilter customizations and custom CNI plugins to third-party observability integrations, platform owners were inevitably pulled in to troubleshoot issues for “unsupported” features.

Join AKS and UBS as we explore the “Allowed but Unsupported” model from both sides - where it held up, where our lines blurred, and how to draw them more effectively for your own platform. You’ll leave with practical guidelines for defining support tiers, classifying capabilities, and setting ownership boundaries that withstand enterprise pressure.

### 2:05 PM–2:30 PM · Inside the Engine: Tapping into OpenTofu’s Internal Logging and Tracing for Security Auditing

- Room: Salt Palace | Level 1 | Ballroom H
- Speakers: Sakshi Nasha, Primanshu Choudhary
- Track: OpenTofu Day
- Labels: Any Level, Breakout Session, OpenTofu Internals
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269402

To safely run Infrastructure as Code at scale, security teams cannot treat the execution engine like a black box. Understanding exactly how OpenTofu parses configurations, communicates with provider plugins over internal gRPC channels, and writes to state files requires deep visibility into the engine's runtime internals.

This technical session explores how to leverage OpenTofu’s internal logging systems and execution tracing mechanics to build a proactive security audit trail. We will break down the engine's internal logging levels, analyze how provider schemas expose sensitive data in debug streams, and demonstrate how to capture these internal telemetry signals directly from execution runners. Attendees will walk away with a practical blueprint for parsing OpenTofu's core logs to detect unauthorized provider actions, trace execution paths, and secure the system pipeline without modifying the core engine binary.

### 2:05 PM–2:30 PM · Feel the Breeze: The High-Energy, Fan-Powered Guide to Sustainable Kubernetes

- Room: Salt Palace | Level 2 | 255 A
- Speakers: David Pech, Petr Rais
- Track: Kubernetes on Edge Day
- Labels: Beginner, Breakout Session, Event + Data Collection
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1256552

Standard HPA is so totally out now. It scales on CPU but lacks synergy. It doesn’t know if your nodes are sipping clean wind or chugging dirty coal. It’s time to stop "shifting left" and start shifting to the sun!

In this carbon-critical session, we’re ditching boring metrics for Project Kepler. Using eBPF magic we’ll expose the raw power consumption of your pods in real-time.

To ground the hype, we’re bringing a physical Solar Panel, lights and a Fan-as-a-Service on stage. Watch the replicas scale to peak intensity the moment our "sun" hits the panel, while industrial fans translate that clean energy into a literal breeze for the front row.

Beyond the hype, come learn how individual pods contribute to your cluster's energy footprint and how to accurately measure it. We’ll showcase a real-time demo of carbon-aware scaling with a few gadgets, and outline the practical decisions you can make to lower your CO2 footprint.

### 2:05 PM–2:30 PM · No Token Budget: What Breaks When Agentic Software Factories Run on Kubernetes

- Room: Salt Palace | Level 1 | Ballroom ACE
- Speakers: Carl Ross
- Track: Agentics Day: MCP + Agents
- Labels: Advanced, Breakout Session, Developer Tooling + Debugging
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1262045

At OpenAI, we explored what happens when agentic software factories with access to tens of millions of tokens per minute are utilized for the full software development lifecycle using CNCF-stack components.

Agents explore designs, implement changes, run tests, review diffs, coordinate releases, inspect production signals, explain failures, propose remediations and maintain entropy. The development model changes from sequential to massively parallel, hermetic, continuously evaluated software delivery with maintained alignment.

This talk shares a technical case study of running these systems on Kubernetes and CNCF ecosystem components: workflow orchestration, event-driven queues, Kubernetes workers, hermetic sandboxes, MCP-style tool gateways, workload identity, policy gates, eval systems, release controls, trace pipelines, production-monitoring integrations, and human approval flows.

Once token budget disappears as a constraint what do we find as the gaps and largest problems?

### 2:05 PM–2:30 PM · Introducing Agent Substrate, Scaling Thousands of Stateful AI Agents on Kubernetes Without the Waste

- Room: Salt Palace | Level 1 | Ballroom IG
- Speakers: Alex Van Boxel
- Track: Agentics Day: MCP + Agents
- Labels: Any Level, Breakout Session, Performance + Scaling
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1254201

As we're shifting toward more and more long-running, stateful autonomous agents everywhere, traditional infrastructure models struggle to keep pace. Most of those agents are mostly idle, waiting for LLMs or even other agents to return. If you need to keep running those workloads, you end up with an overprovisioned cluster.

This presentation introduces Agent Substrate (https://github.com/agent-substrate/substrate), an open-source, framework-agnostic system built to scale agentic workloads efficiently on Kubernetes. Agent Substrate maps a large pool of stateful "actors" onto a significantly smaller pool of active "worker" Pods, achieving high-density multiplexing (30x+ oversubscription) via rapid session teleportation - leveraging kernel-level gVisor checkpoints to suspend and resume running process memory and filesystem state in under a second.

We'll explore the core architecture, including the automated lifecycle controller, routing layers, and node-level snapshotting mechanics.

### 2:05 PM–2:30 PM · Can Your Agent Execute Anything? Authorization for Backstage MCP Actions

- Room: Salt Palace | Level 2 | 255 EF
- Speakers: Masato Kozuka
- Track: BackstageCon
- Labels: AI Agents + Context in Backstage, Breakout Session, Intermediate
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1267548

Backstage exposes its actions registry over MCP to AI agents and IDEs. Because MCP clients can invoke any exposed action, access control must be enforced by Backstage itself, not by the client.

Identifying the caller (authentication) is stabilizing in the community, with CIMD/DCR for clients and Device Authorization for users. However, what they can execute (authorization) remains uncontrolled. The existing permission model is too coarse-grained; it cannot be resource-scoped and does not distinguish between listing and running. An action with no permission is executable by any authenticated caller, and destructive actions are not denied by default.

The session starts with how far an authenticated MCP client can reach and run. It then lays out the requirements any fix must meet and a reference design: an execute permission separate from visibility, registry-enforced, fail-closed on destructive ones, resource-aware, and designed to integrate with Backstage's permission framework.

### 2:05 PM–2:30 PM · Introducing OCI Webhook Support in Argo CD: From Polling to Real-Time

- Room: Salt Palace | Level 1 | 151
- Speakers: Nitish Kumar, Surabhi Mishra
- Track: ArgoCon
- Labels: Any Level, Breakout Session, Progressive Delivery
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269345

As GitOps evolves, more teams are adopting OCI-based artifacts for Helm charts and Kubernetes manifests instead of traditional Git repositories. While Argo CD added support for OCI sources, one major gap remained: event-driven updates. Unlike Git repositories, OCI registries historically lacked first-class webhook integration, forcing teams to rely on polling or manual refreshes. For organizations using GHCR, DockerHub, this meant slower feedback loops and reduced automation in production environments.

Argo CD is adding webhook support for OCI-compliant registries. In this session, I’ll present the design and implementation of webhook support, including GHCR and upcoming Docker Hub integration. I’ll also demonstrate how to configure registry webhooks, connect them to Argo CD, and trigger real-time reconciliation securely.

By the end of this talk, you’ll understand not just how OCI webhooks work internally, but how to configure them correctly in production and what pitfalls to avoid.

### 2:05 PM–2:30 PM · Cost-Aware GitOps: Automating FinOps with Argo Workflows

- Room: Salt Palace | Level 2 | 251 A-C
- Speakers: Ekansh Gupta
- Track: ArgoCon
- Labels: Any Level, Breakout Session, Scalability
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1268762

Cloud costs are often managed through dashboards, reports, and periodic reviews, while the systems generating those costs operate continuously. As infrastructure scales, manual cost optimization becomes difficult to sustain and easy to overlook.

This session explores how Argo Workflows, Argo Events, and GitOps practices can be used to automate FinOps operations. We’ll demonstrate patterns for building event-driven cost optimization workflows that respond to utilization signals, identify idle resources, enforce policy-driven resource configurations, and automate routine optimization tasks.

Attendees will learn how to integrate cost signals into their delivery platforms, implement automated cleanup and rightsizing workflows, and use GitOps principles to make cost management auditable, repeatable, and safe. We’ll also discuss the trade-offs of automated optimization, operational guardrails, and practical considerations for balancing efficiency, reliability, and developer productivity.

### 2:05 PM–2:30 PM · Best Practices for LLM Autoscaling from the llm-d Ecosystem

- Room: Salt Palace | Level 2 | 254
- Speakers: Jiazhou Gao, Abhishek Malvankar
- Track: Cloud Native AI + Inference Day
- Labels: Any Level, Breakout Session, LLMs + Generative AI
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1245135

As LLM inference moves from demos to production, traditional Kubernetes autoscaling patterns are breaking down. GPU utilization stays pegged at high values even when systems are near failure, cold starts take minutes, and KV cache pressure can silently degrade latency long before HPA reacts. Modern inference stacks need autoscaling that is workload-aware, SLO-aware, and hardware-aware.
Queue depth, KV cache utilization, TTFT, and request saturation are more reliable scaling signals than raw GPU utilization. In this talk, we present practical best practices for building production-grade autoscaling for LLM inference on Kubernetes, based on recent research and lessons from heterogeneous inference serving systems in the llm-d ecosystem. We will show how proactive scaling on these signals reduces request failures and improves GPU KV cache utilization.

### 2:05 PM–2:30 PM · Agents Are Not Microservices: Operating Autonomous AI on Kubernetes

- Room: Salt Palace | Level 2 | 251 D-F
- Speakers: Henrique Santana, Lanay Marques
- Track: Cloud Native AI + Inference Day
- Labels: Breakout Session, Intermediate, RAG + AI Agents
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1237691

Deploying a model inference endpoint on Kubernetes is well-understood. Deploying autonomous AI agents is not. Agents break core assumptions: pods hold state across interactions, resource consumption is unpredictable (idle then burst), health checks fail on pods that are "thinking," autoscalers cannot predict demand from agents that spawn sub-tasks dynamically, and graceful shutdown means waiting for multi-step reasoning to complete.

This talk identifies where Kubernetes primitives fall short for agentic workloads and what operators can do today. We cover adapted liveness and readiness probes for long-running inference, scaling strategies when agents trigger unpredictable downstream load, graceful termination for multi-step reasoning chains, and communication patterns between agents using MCP and A2A protocols on Kubernetes. Each gap is paired with a working pattern using existing Kubernetes features and CNCF projects.

### 2:05 PM–2:30 PM · llm-d Observability: Observing LLM Inference at Scale on Kubernetes

- Room: Salt Palace | Level 2 | 250 A-C
- Speakers: Guangya Liu
- Track: Observability Day
- Labels: AI + LLM Observability, Any Level, Breakout Session
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1247321

Production LLM inference on Kubernetes is distributed by design: requests flow through gateways, llm-d Router (EPP), KV-cache indexers, vLLM workers, and Prefill/Decode sidecars. Traditional metrics miss token-level latency, cache locality, and routing decisions, and that gap quietly erodes the 3x throughput gains from cache-aware scheduling.
This talk shares llm-d's (CNCF Sandbox) Observability: a component-level audit of metrics, traces, logs, dashboards, and alerts across the stack. We show what is already working (EPP and vLLM OTel, Prometheus scrapers, inference-sim for GPU-free CI) and the critical gap: cross-component trace linkage, zero pre-built dashboards, and opaque P/D KV transfers.
Attendees will leave with practical patterns: W3C TraceContext propagation, OTel Collector configs, tiered KV-cache observability, WVA autoscaling visibility, and a prioritized community roadmap for future release. Applicable beyond llm-d to any Kubernetes-native GenAI serving platform.

### 2:05 PM–2:30 PM · OTel and Linkerd: A Match Made in Observability Heaven

- Room: Salt Palace | Level 2 | 250 D-F
- Speakers: Julia Furst Morgado, Flynn -
- Track: Observability Day
- Labels: Breakout Session, Correlation + Context + Cross-Project Architectures, Intermediate
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1259095

You’ve deployed Linkerd. You’ve set up OpenTelemetry. Your traces still break at the mesh boundary and you have no idea why.

This happens all the time! Service meshes intercept traffic at the network level but most of them don't propagate trace context correctly, which means your spans get orphaned, your trace IDs don't survive the hop, and you end up with a useless waterfall that stops at the first sidecar.

Even though Linkerd has native OpenTelemetry support, understanding what's actually happening under the hood will help you get it right. That's what this talk is about. We'll unpack how it works, why mesh-level integration is harder than it sounds, and what correct trace context propagation actually looks like in a mesh. And to make it concrete, you'll see a full distributed trace end-to-end through Linkerd with no hacks and no manual context stitching.

"Where did my trace go" will finally have an answer.

### 2:05 PM–2:30 PM · OSS-CRS: Open Framework for Bug-Finding and Patching Cyber Reasoning Systems

- Room: Salt Palace | Level 2 | 255 BC
- Speakers: Andrew Chin
- Track: Open Source SecurityCon
- Labels: Breakout Session, Intermediate, Leveraging + Preparing for AI in OSS Security
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269002

The AI Cyber Challenge demonstrated that AI-powered Cyber Reasoning Systems (CRS) can autonomously find and fix software vulnerabilities at scale. But how do we take those advancements and make them accessible to the broader security community? Enter OSS-CRS: an open-source, standardized framework designed to accelerate the development of AI-assisted bug-finding and remediation systems. In this session, we'll walk through the design principles of OSS-CRS, show how it lowers the barriers to building and benchmarking next-generation CRS tooling, and demonstrate how users can easily deploy and run CRSs against their own codebases. Whether you're a security researcher, tooling developer, AI practitioner, or project maintainer, come learn about the growing ecosystem around AI-powered CRSs.

### 2:05 PM–2:30 PM · Platform as a Product in a 300-Year-Old Bank

- Room: Salt Palace | Level 1 | Ballroom BDF
- Speakers: Abby Bangser, Joel King
- Track: Platform Engineering Day
- Labels: Breakout Session, Intermediate, Platform Engineering Case Study
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1268113

Naming a team "Platform as a Product" is one thing. Cutting time-to-market for Kubernetes applications by over 97% is another.

This talk follows the PaaP team at NatWest Bank through how they actually did it — the technology decisions, the organisational friction, and the trust they had to build along the way. It's also a direct translation of the CNCF platform engineering white papers into practice: not what the theory says you should do, but what a 300-year-old bank did, what worked, and what they'd do differently.

The gap between industry frameworks and real implementation is where most platform teams struggle. This talk closes that gap with a concrete, longitudinal case study from one of the more challenging environments you could pick: a heavily regulated financial institution with centuries of legacy.

### 2:05 PM–2:30 PM · Hidden Platform Metrics: What Unmanaged Open Source Is Costing Developer Veloc

- Room: Salt Palace | Level 1 | Ballroom J
- Speakers: Leslie Pascual
- Track: Platform Engineering Day
- Labels: Breakout Session, Intermediate, Measuring Platform Success
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1268009

Platform teams measure success in deployment frequency, lead time, and developer satisfaction. Open source dependency management rarely appears in those metrics until a CVE blocks a deployment, a transitive dependency breaks a build, or a zero-day rewrites the roadmap. This operational cost is real, but it is usually absorbed as unplanned work rather than surfaced as a platform gap. AI coding tools have made this invisible cost larger, generating dependency intake faster than reactive governance can track, with remediation landing on teams that had no visibility into what was introduced. This session covers how to surface open source governance as a measurable platform capability, what the before-and-after looks like when governance moves upstream, and how to build the case for treating open source management as a platform investment rather than security overhead.

### 2:35 PM–2:40 PM · Sponsored Keynote | TBA, Zededa

- Room: Salt Palace | Level 2 | 255 A
- Track: Kubernetes on Edge Day
- Labels: Any Level, Keynote Session
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1311978

### 2:40 PM–2:50 PM · Real OpenTofu Statistics

- Room: Salt Palace | Level 1 | Ballroom H
- Speakers: Sebastian Stadil
- Track: OpenTofu Day
- Labels: Any Level, Upgrades and Usage, ⚡ Lightning Talk
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269085

Scalr, a drop-in replacement for Terraform Cloud, executes over 1M plans & applies per month. Two thirds of these are now OpenTofu, up from a third same time last year.

Dusting off my old degree in statistics, I ran some queries on internal usage to find some interesting insights into how OpenTofu is used in production, how it's used differently from Terraform, how upgrade lag has evolved, and more.

### 2:40 PM–3:05 PM · What It Took to Write Our Own MCP Gateway

- Room: Salt Palace | Level 1 | Ballroom ACE
- Speakers: Vijay Samuel, Nick Pordash
- Track: Agentics Day: MCP + Agents
- Labels: Any Level, Breakout Session, MCP Fundamentals + Architecture
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269211

Today, everyone is vibe coding software regardless of it existing in the open or if there are vendor products available. At eBay, we wrote our own MCP Gateway. Not because we can, but because there wasn't one out there that fit all our requirements. It is a well known fact that identity is messy especially in organizations that have been around longer. However, heterogenous systems built around the OAuth model make things all the more challenging.

I as an end user, should be able to login once to the gateway and be able to communicate with any MCP server regardless of what kind of authentication the system behind it accepts. To be able to handle such diversity is exactly why we built our own Gateway.

This talk describes:
* how users interact with the gateway
* how we do routing and discovery
* how we introduced the concept of token exchange to freely move from one IAM system to another's token requirements
* what are the challenges we faced and still look to solve

### 2:40 PM–3:05 PM · From Alert to Commit: Agentic Incident Resolution

- Room: Salt Palace | Level 1 | Ballroom IG
- Speakers: Grant Griffiths, Sarah Khalife
- Track: Agentics Day: MCP + Agents
- Labels: Any Level, Breakout Session, Extending AI Systems with MCP
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269186

Current root cause analysis and remediation workflows are often slow, manual, and introduce unnecessary stress: an on-call engineer gets paged at 2am, spends hours correlating across observability tools, drafts a fix under pressure, and hopes it passes review before the SLA clock runs out.

This session introduces a paradigm shift. By pairing an AI SRE agent with GitHub Copilot, we create a pipeline that moves from on-call alert to a tested, compliant PR in minutes, not hours. You'll see live how the SRE agent receives an alert, correlates data, and produces an RCA with supporting evidence. Using the RCA as context, we drive a targeted fix, run it through security scanning and test pipelines, and open a fully documented PR all through agentic orchestration. Attendees will leave with an architectural pattern for wiring together AI agents via MCP—and a clear picture of how agentic incident response can be the foundation of a process that is faster, more auditable, and built for scale.

### 2:40 PM–3:05 PM · Your Agent Shouldn't Log in as You: Identity for MCP in Backstage

- Room: Salt Palace | Level 2 | 255 EF
- Speakers: Shivay Lamba, Hrittik Roy
- Track: BackstageCon
- Labels: AI Agents + Context in Backstage, Breakout Session, Intermediate
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269772

Backstage 1.43 recently shipped scoped, short-lived MCP tokens. But a lot of teams still ignore them, wire the actions backend to a static admin token. The agent can now do anything you can, which is exactly the wrong default the moment an agent starts acting on its own.

This talk goes deep on agent identity inside Backstage, on the primitives that already exist to fix this:

- Token exchange (RFC 8693), so an agent acts with a derived, scoped credential, not yours.
- OAuth 2.1 as the trust boundary between agent, actions backend, and Backstage.
- The permission framework, where the real constraint lives, so an agent can't exceed its task.The through-line: you should not be able to ask an agent to delete a production database, and that guarantee lives in the permission model, not the agent's judgment.

Thus, the attendees will leave knowing how to give agents their own scoped identity and enforce limits in the permission framework, rather than hoping the agent behaves.

### 2:40 PM–3:05 PM · Predicting GitOps at 15,000+ Clusters: What Large-Scale Testing Taught Us

- Room: Salt Palace | Level 1 | 151
- Speakers: Artem Lajko, Gianluca Mardente
- Track: ArgoCon
- Labels: Breakout Session, Intermediate, Scalability
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1244623

Most Kubernetes platforms and GitOps tools are designed for organizations operating between 10 and 300 clusters.

But a different category begins when organizations operate between 1,000 and 10,000 clusters across edge, retail, telecom, IoT, supermarkets, or franchise environments. At this scale, many GitOps architectures and tools start to break down, even with HA setups and extensive tuning.

Public scaling data at this level is limited because these environments are rare, expensive, and difficult to reproduce realistically.

In this talk, we share how we approached predicting the operation of 15,000+ Kubernetes clusters based on a real-world customer requirement. Using large-scale GitOps load testing in a hub-and-spoke architecture, we explored scaling limits, bottlenecks, infrastructure costs, and operational trade-offs within a budget of roughly €100,000.

The session also covers what our findings may tell us about the future of GitOps at extreme scale with Argo CD and Sveltos.

### 2:40 PM–3:05 PM · The Agent Broke It, the Agent Fixed It: Self-Healing Progressive Delivery on Kubernetes

- Room: Salt Palace | Level 2 | 251 A-C
- Speakers: Carlos Sanchez, Kevin Dubois
- Track: ArgoCon
- Labels: Beginner, Breakout Session, Progressive Delivery
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269540

We all have "a friend" who trusts AI-generated code too much. When that code reaches production, rollback helps, but it does not fix root cause. Our open-source approach plugs into progressive delivery tools like Argo Rollouts and Flagger, adding analysis and remediation without replacing rollout flows.

In this session, we'll demo a Kubernetes-native self-healing loop built with open-source projects: distributed analysis and remediation agents run in-cluster, with open-source models dynamically selected by an orchestration workflow.

You'll see:
* AI-driven canary vs stable analysis across logs, metrics, and rollout signals.
* Dynamic model selection per step based on complexity, cost, and reasoning depth.
* Evidence-backed root-cause analysis with remediation as pull requests or issues.
* Configurable human approval gates before merge or promotion.

You'll leave with a practical open architecture to reduce MTTR while preserving observability, rollback safety, and operator control.

### 2:40 PM–3:05 PM · Three Years in Thirty Minutes: Time-Travel Ethics Evaluation for AI Agents

- Room: Salt Palace | Level 2 | 254
- Speakers: Vicente Herrera
- Track: Cloud Native AI + Inference Day
- Labels: Breakout Session, Ethics + Compliance + Security, Intermediate
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1250481

Cloud native teams are deploying agents that may run autonomously for years, but today's evaluation tooling tells them little about long-horizon drift, deception, or ethical degradation. Model cards capture seconds of behavior, not years, while Anthropic research shows concerning behavior under pressure, including blackmail. This session contributes a vendor-neutral, reproducible pattern built on CNCF projects (Kubernetes, Flux, OpenTelemetry), combined with open ethics tooling (Game of Ethics), so teams don't need new infrastructure to measure how their agents behave under autonomy. The whole system (prompts, tools, skills, I/O...) rather than the model alone is evaluated through a simulated production environment, that the agent can't distinguish from reality, where time can be fast-forwarded, to score long-horizon decisions for unsafe or unethical actions. Attendees leave with a pattern they can apply to their own agentic systems to reveal the impact of years of decision-making.

### 2:40 PM–3:05 PM · Ship Fearlessly: Canary Deployments with an AI Quality Gate for Control Plane Releases

- Room: Salt Palace | Level 2 | 251 D-F
- Speakers: Kartikeya Pharasi, Vinay Gonuguntla
- Track: Cloud Native AI + Inference Day
- Labels: Any Level, Breakout Session, RAG + AI Agents
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1267535

Rolling out cluster management components like controllers and operators is uniquely risky. A bad release can destabilize scheduling, networking, or policy enforcement for every workload in a fleet. Yet most teams deploy with simple rolling updates and hope dashboards catch problems before users do.

We present an AI-driven canary deployment system for cluster management components. Each release progresses through stages: a single non-production cluster, a broader pre-production set, a production canary, and finally fleet-wide rollout. At each stage an autonomous AI agent fetches the release diff, queries the observability backend for anomalies, and uses an LLM to correlate code changes against error signals. The agent produces a go or no-go recommendation before the release advances. When anomalies surface, it generates a detailed correlation report, drafts PR fixes and new regression tests, and sends them for human-in-the-loop review to be included in the next release.

### 2:40 PM–3:05 PM · Building Observability Control Planes with OpenTelemetry, eBPF, and OBI

- Room: Salt Palace | Level 2 | 250 A-C
- Speakers: Mike Dame, Rafael Roquetto
- Track: Observability Day
- Labels: Any Level, Breakout Session, eBPF + Zero-Code + Kernel-Level Observability
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1247219

The OpenTelemetry eBPF Instrumentation (OBI) agent provides auto-observability with broad coverage via eBPF. But OBI is more than a fixed agent: it also exposes library tools developers can import, extend, and embed into their own systems.

This talk shows how platform engineers, observability teams, and vendors can use OBI as a library to build custom, dynamic observability controllers. We'll explore how higher-level systems can decide what to instrument, when configuration should change, and how instrumentation maps to platform or product context.

Attendees will leave with practical patterns for building on OBI, tradeoffs between custom and community-managed approaches, and ideas for extending zero-code instrumentation into developer platforms, internal control planes, and third-party tools

### 2:40 PM–3:05 PM · An Ode to Metrics: Is OpenTelemetry Moving Away from Prometheus?

- Room: Salt Palace | Level 2 | 250 D-F
- Speakers: Diana Todea, Reese Lee
- Track: Observability Day
- Labels: Any Level, Breakout Session, Project Deep Dives
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1255260

For nearly a decade, Prometheus has been the default language of cloud-native metrics.

Today, OpenTelemetry recently graduated CNCF project, is redefining how telemetry is generated, transported, and processed. While the two projects remain deeply interconnected, OpenTelemetry metrics are introducing concepts that don't always fit neatly into the traditional Prometheus model.

This talk explores the evolving relationship between Prometheus and OpenTelemetry through a fun "good cop, bad cop" debate. We'll examine native versus exponential histograms, pull versus OTLP-based collection, semantic conventions, exemplars, temporality, collector processing, and cross-signal observability.

Rather than declaring a winner, we'll explore what each approach optimizes for, where interoperability shines, and where the ecosystems are beginning to chart different paths.

The audience will understand where each project is heading and more importantly, where do they need further contribution.

### 2:40 PM–3:05 PM · License to Generate: Securing AI Agents with Keycloak, SPIRE, and Envoy

- Room: Salt Palace | Level 2 | 255 BC
- Speakers: Maia Iyer, hai Huang
- Track: Open Source SecurityCon
- Labels: Any Level, Breakout Session, Leveraging + Preparing for AI in OSS Security
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1268158

Your organization may be aware of prompt injection, privilege escalation, and impersonation risks with AI agents. You may even be aware of some theoretical defenses: identity, sandboxes, and guardrails. But how do you actually implement these at enterprise scale?

This session walks through a working vendor-neutral, open-source implementation built on Keycloak, SPIRE, and Envoy that shifts AI security down to the platform layer. We will show how per-request identity becomes the foundation for fine-grained authorization and dynamic guardrails, and explore the critical interface between agent sandboxes and the broader platform. Finally, we will take a concrete scenario and show this platform blocking an attack at various layers, demonstrating defense-in-depth. Attendees will leave with a modular, cloud-native reference architecture to secure their enterprise against unpredictable, manipulatable agentic workloads.

### 2:40 PM–3:05 PM · Platform as Context: Evolving Our IDP for AI Agents

- Room: Salt Palace | Level 1 | Ballroom BDF
- Speakers: Gang Luo, Markus Mäkelä
- Track: Platform Engineering Day
- Labels: Breakout Session, Intermediate, Platform Engineering and AI
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1244896

At Platform Engineering Day EU 2026 we shared how we built our IDP as an abstraction layer focused on developer experience, defining golden paths, standardizing workflows, and reducing friction across internal systems.

As we started using AI agents in engineering workflows, we realized the challenge shifted from developer experience to agent context. Agents need reliable information including ownership, infrastructure metadata, deployment history, operational status, and relationships between systems, often fragmented across tools and teams.

In this session we explore how we evolved our IDP into a context layer for agents. We cover how we structure and standardize metadata, expose it through MCP, and use the platform to support more reliable AI-driven workflows.

The talk focuses on practical lessons learned and tradeoffs of building AI capabilities on top of an existing platform. Attendees will learn how Platform Engineering can provide the foundation for scalable agentic workflows.

### 2:40 PM–3:05 PM · Zero-Ops Disaster Recovery: A Paved Path Through Multi-Cluster Orchestration

- Room: Salt Palace | Level 1 | Ballroom J
- Speakers: Joe Nathan Abellard, Li Zhuyu
- Track: Platform Engineering Day
- Labels: Breakout Session, Intermediate, Platform Engineering Case Study
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1266339

Whether driven by regulatory mandates, SLAs, or the cost of downtime, organizations need their Kubernetes workloads to survive datacenter failures. Yet, achieving disaster recovery (DR) across multiple clusters remains one of the hardest problems in platform engineering. Most teams either push this complexity onto tenants or build brittle automation that breaks at scale.

At Bloomberg, we run a managed multi-cluster Kubernetes platform for AI and streaming analytics workloads that is powered by Karmada, with control planes spanning regions. This talk shares the zero-ops DR architecture we built so that tenants get cross-datacenter resilience without becoming multi-cluster experts, as well as the challenges we faced along the way. We cover policy-driven placement, automated failover, and continuous resilience verification. While built on Karmada, the patterns (control plane abstraction, intent-based placement, and continuous DR verification) apply to any multi-cluster strategy.

### 2:55 PM–3:20 PM · PM Break 1

- Room: Meeting Room Foyers
- Track: OpenTofu Day
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1309417

### 3:05 PM–3:20 PM · PM Break 2

- Room: Meeting Room Foyers
- Track: Agentics Day: MCP + Agents
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1309416

### 3:20 PM–3:45 PM · Running LLM Inference on K8s at the Edge: Scheduling, Isolation, and “GPU Time-Sharing” in Practice

- Room: Salt Palace | Level 2 | 255 A
- Speakers: Michael Yuan
- Track: Kubernetes on Edge Day
- Labels: Any Level, Breakout Session, Kubernetes in Production on the Edge
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1252728

LLM inference is now a day-2 workload: it must be scheduled, upgraded, observed, and isolated like any other production service — but edge deployments have harsher constraints (single-node clusters, limited VRAM, noisy neighbors, intermittent connectivity).
This session shares a practical approach to running local inference workloads on Kubernetes (often K3s) with limited GPU resources. We will cover: (1) GPU allocation strategies for multiple concurrent workloads, (2) protecting interactive latency under contention, (3) packaging model runtimes as installable applications, (4) basic SLOs + observability signals for inference (latency, queue depth, VRAM pressure), and (5) what failed in real deployments.
We’ll demo two AI workloads competing for the same GPU and how we keep the system responsive while preserving isolation.

### 3:20 PM–3:45 PM · Building and Maintaining Production-Ready SDKs with LLMs

- Room: Salt Palace | Level 1 | Ballroom ACE
- Speakers: Ender Demirkaya, Shijie Sheng
- Track: Agentics Day: MCP + Agents
- Labels: Advanced, Breakout Session, Developer Tooling + Debugging
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1250590

Building and maintaining SDKs across multiple programming languages is a significant challenge for platform teams. APIs must remain consistent while still embracing each language's idioms, tooling, and developer expectations. After the first release, keeping multiple implementations in sync still stays as a large effort.
.
Over the past two years, the Cadence community has explored how LLMs can automate much of this process and developed a pipeline that maps an existing source SDK into a language-agnostic representation, enriches it with semantic metadata, applies language-specific transformation rules, and generates production-ready SDKs for target languages. Most of this workflow runs locally and can produce a new SDK in a matter of hours.

In this session, we'll walk through the architecture behind this system, discuss the challenges we encountered, and share lessons learned from applying LLMs to real-world SDK generation.

### 3:20 PM–3:45 PM · Stop Wasting GPUs: Building a Workload-Aware Caching Proxy for Agentic AI Workflows

- Room: Salt Palace | Level 1 | Ballroom IG
- Speakers: Chetan Sharma
- Track: Agentics Day: MCP + Agents
- Labels: Breakout Session, Intermediate, Performance + Scaling
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1268495

AI workflows on Kubernetes increasingly resemble complex DAGs of agents, tools, and model calls. Each step produces intermediate artifacts, such as OCR, embeddings, tensors, and retrieval context, that are expensive to recompute. Yet, most caches rely on simple policies like LRU. These evict high-cost artifacts while preserving cheap ones, causing avoidable GPU cycles, latency spikes, and cluster autoscaling.

This session presents a Kubernetes-native caching proxy for these workflows. Running alongside workload pods, it observes topology and reuse patterns to make eviction decisions based on re-computation cost, downstream impact, and invocation frequency. Attendees will see the architecture, algorithm, and deployment patterns for Argo Workflows-style DAGs. The talk closes with algorithmic benchmarks from the underlying research, showing up to 51% lower pipeline latency compared to standard LRU, alongside significant reductions in GPU usage and unnecessary autoscaling.

### 3:20 PM–3:45 PM · IDPs Are Dead: Platform Engineering in the Age of AI Agents

- Room: Salt Palace | Level 2 | 255 EF
- Speakers: Tiago Reichert, Lucas Duarte
- Track: BackstageCon
- Labels: AI Agents + Context in Backstage, Breakout Session, Intermediate
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1243609

Internal developer platforms were built around portals, service catalogs, and self-service UIs. Today, developers increasingly interact with platform capabilities through AI coding agents that can generate infrastructure, open pull requests, and trigger deployments from natural language prompts. As this interaction model evolves, platform teams must ensure these systems operate safely within established operational boundaries.

We’ll explore how AI agents are becoming a complementary interface layer for platforms like Backstage. You’ll learn how GitOps workflows, GitFlow approvals, policy enforcement, and admission controllers become the enforcement layer for agent-driven operations. We’ll demonstrate how MCP servers connect agents to platform APIs, how pull requests and policy checks validate infrastructure changes before deployment, and how teams can observe and govern AI-driven workflows while preserving existing controls. Platform engineering is not disappearing. It is evolving.

### 3:20 PM–3:45 PM · Beyond GitOps: Building Intelligent Drift Detection and Auto-Remediation in Argo CD

- Room: Salt Palace | Level 1 | 151
- Speakers: Ram Mohan Rao Chukka, Shibi Ramachandran
- Track: ArgoCon
- Labels: Breakout Session, Intermediate, Software Delivery
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1229385

In most of the Kubernetes clusters, configuration drift is unavoidable and is often hard to detect until something goes wrong. ArgoCD does a great job with basic drift detection, self-healing and auto-sync, but in real-world enterprise environments, not all drift is equal. Scaling a deployment isn’t the same as deleting a secret, and treating them the same can lead to trouble.

In this talk, we’ll walk through how to extend ArgoCD to handle drift more intelligently. Using Custom Health Checks, ApplicationSets, Resource Hooks, and a custom controller, we built a system that understands the severity of changes and responds accordingly by auto-remediating safe changes, triggering approvals for sensitive updates, and rolling back when things go wrong.

This talk is ideal for platform engineers, SREs, and DevOps teams looking to evolve their GitOps strategy with smarter automation, policy enforcement, and enterprise-grade reliability.

### 3:20 PM–3:45 PM · Ship AI Like You Mean It: Progressive Delivery for AI Workloads Using Agentgateway

- Room: Salt Palace | Level 2 | 251 A-C
- Speakers: Lasse Vierow, Kostis Kapelonis
- Track: ArgoCon
- Labels: Breakout Session, Intermediate, Progressive Delivery
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1266370

Just when we thought we had understood how to perform Progressive Delivery with cloud applications, AI workloads emerged, making the situation even more challenging.

Requests are expensive, model versions behave in subtle ways, and the cost of misrouting traffic isn't just a slow response, it's a ballooning bill and degraded user experience.

We will explain: why token-heavy inference traffic demands tighter controls, why model versioning is not the same as image versioning, and why you need a gateway that understands AI-specific semantics such as token-based rate limiting, guardrails, and per-model cost tracking.

Agentgateway gives you the policy enforcement and observability layer that Argo Rollouts needs to make safe, informed promotion decisions for AI workloads.

The session includes a demo highlighting the current state of the integration. Attendees will leave with a realistic understanding of how to combine Argo Rollouts, Gateway API, and AI gateways.

### 3:20 PM–3:45 PM · LLM Inference on Edge Kubernetes: Patterns for Offline, Constrained, and Heterogeneous Devices

- Room: Salt Palace | Level 2 | 254
- Speakers: Harshul Jain, Saurabh Yergattikar
- Track: Cloud Native AI + Inference Day
- Labels: Breakout Session, End-to-End AI/ML in Production (MLOps/AIOps), Intermediate
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269549

Edge LLM inference isn’t “cloud inference, but smaller.” Edge deployments must handle intermittent connectivity, tight memory/compute budgets, and heterogeneous accelerators (CPU-only, iGPU, small GPUs, NPUs). Teams that reuse cloud assumptions (always-on control plane, uniform hardware, centralized logging) hit failure modes that look like randomness: cold starts that take minutes, unpredictable latency, silent model corruption, and observability gaps when devices go offline. This talk presents Kubernetes-native patterns for shipping and operating LLM inference on edge clusters (single-node and small multi-node): model artifact packaging and rollouts (versioning, integrity checks, rollback safety), scheduling/isolation across heterogeneous nodes, offline-first operations (local queues + eventual sync), and observability that still works with delayed export (local buffering + minimal health signals).

Ref Repo: https://github.com/harshuljain13/llm-inference-at-scale

### 3:20 PM–3:45 PM · DRA + Kueue + HAMi + vLLM: A Production Architecture for Scheduler-Agnostic GPU Sharing

- Room: Salt Palace | Level 2 | 251 D-F
- Speakers: Mengxuan Li, Yuchen Fama
- Track: Cloud Native AI + Inference Day
- Labels: Any Level, Breakout Session, User Success Stories + Use Cases
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1267046

GPU sharing in Kubernetes remains solved-in-theory, painful-in-practice. Time-slicing and MIG require pre-carved templates. HAMi and Volcano deliver dynamic slicing — but bundle their own schedulers, making adoption a non-starter in clusters already running custom scheduling stacks.
Dynamic Resource Allocation (DRA) breaks this deadlock. Scheduler-agnostic by design, it layers fine-grained runtime device partitioning on top of any scheduler. No ripping out existing infrastructure.
This talk delivers a live production reference — not a whitepaper, including:
1. A working DRA + Kueue + HAMi webhook + vLLM deployment, with llm-d handling distributed LLM inference routing on top
2. What scheduler-agnostic GPU slicing actually buys you when the workload is LLM inference at scale
3. How to overcome the obstacles when using DRA in production: such as lack of monitoring, and needs API change when allocating GPUs.

### 3:20 PM–3:45 PM · Hallucinating High Availability: Where AI-Driven Remediation Breaks Down

- Room: Salt Palace | Level 2 | 250 A-C
- Speakers: Aderianna T Williams
- Track: Observability Day
- Labels: AI + LLM Observability, Breakout Session, Intermediate
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1268596

Letting an AI agent automatically fix production incidents sounds great on paper, promising faster triage and fewer late-night pages. In a complex microservices architecture, a highly confident suggestion can still be catastrophically wrong if the tool lacks deep context.
This session walks through a realistic Kubernetes incident replay to show exactly where automated operations hit a wall. Imagine a rolling update goes out. The new pods pass readiness probes and resource usage looks normal, but users see intermittent errors. We will show how a typical AI assistant, looking only at surface cluster symptoms, recommends a counter-productive fix that makes the outage worse because it misses downstream dependency timeouts. We provide a practical blueprint for building deterministic guardrails, evidence validation, and automated rollback criteria to keep autonomous actions safe.

### 3:20 PM–3:45 PM · Telemetry Colosseum: OTel Collector vs. Fluent Bit V5 vs. OTel-Arrow. Three Agents, One Crown

- Room: Salt Palace | Level 2 | 250 D-F
- Speakers: Henrik Rexed
- Track: Observability Day
- Labels: Breakout Session, Intermediate, Telemetry Pipelines + Data Engineering
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1266933

Two years ago I sent two agents into the arena: Fluent Bit and the OpenTelemetry Collector. Since then the games have changed. Fluent Bit v5 is back stronger, and a new gladiator has entered the gate: OTel Arrow, built to move telemetry at scale for a fraction of the bandwidth.

So I am reopening the Colosseum, and this time three agents fight: the OTel Collector, Fluent Bit v5, and Otel-Arrow. I run all three under identical Kubernetes conditions and judge them on 5 trials:
- design
- architecture
- signal support across metrics, logs, and traces
- self-observability, the health telemetry each exposes by each agent
- and performance under load (CPU, memory, throughput, latency).
Run independently and vendor-neutral with the OTel-Arrow maintainers, the benchmark shows the crown depends on the arena you fight in. All configs and results will be published on GitHub.

### 3:20 PM–3:45 PM · Hardening Federal Open-Source CI/CD, One Pipeline at a Time

- Room: Salt Palace | Level 2 | 255 BC
- Speakers: Arpit Jain
- Track: Open Source SecurityCon
- Labels: Beginner, Breakout Session, Regulation + Public Policy
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1266670

US federal agencies (cisagov, GSA, NIST, and 300+ other orgs) publish open-source code on GitHub, but their CI/CD pipelines carry the same vulnerability patterns as everyone else: unpinned actions, overly broad tokens, and pwn-request triggers. Executive Order 14028 and the CISA SSDF set the mandate, but the gap between policy and workflow-level implementation is wide.

This talk covers what I found auditing federal repositories for CI/CD weaknesses, and the process of turning findings into accepted contributions. I share before-and-after examples, explain why least-privilege permissions matter even for public repos, and discuss practical challenges: review timelines, CLA processes, and framing a PR so it gets merged.

Over 320 merged security contributions across 75 organizations, with accepted fixes in NASA, NIST, and other federal agencies, plus private vulnerability disclosures to government repositories.

### 3:20 PM–3:45 PM · Where Should Platform Complexity Live? Lessons from a Cross-Team PR Preview Platform

- Room: Salt Palace | Level 1 | Ballroom BDF
- Speakers: Taiki Ono
- Track: Platform Engineering Day
- Labels: Breakout Session, Intermediate, Platform Engineering Case Study
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1242226

Platform engineering decides where complexity lives—and too often it ends up in application developers' heads. We built a per-PR preview platform that runs across teams and repositories via a mesh-native Kubernetes controller, keeping base manifests untouched and developer mental models intact. This talk is the story of the platform/app boundary decisions we made—and one we reversed.

We'll walk through: which cross-cutting concerns we absorbed into the platform (routing, isolation, lifecycle) versus left to apps (small, familiar setup); why we built a controller instead of generating manifests; how we shifted GitOps granularity rather than abandoning it; and how cross-team, cross-repo previews work without app teams ever learning the topology.

Takeaways: (1) a vocabulary for asking 'where should this complexity live?' before you build, (2) which cross-cutting concerns belong in the platform versus the apps, (3) what we kept outside the platform—and why we reversed one earlier choice.

### 3:20 PM–3:45 PM · Managing Distributed AI Infrastructure Under Capacity Constraints Across Kubernetes Environments

- Room: Salt Palace | Level 1 | Ballroom J
- Speakers: Akanksha Sharan
- Track: Platform Engineering Day
- Labels: Breakout Session, Intermediate, Platform Engineering and AI
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1261083

As AI workloads continue to expand, infrastructure teams must manage distributed compute resources while maintaining reliability, efficiency, and operational consistency. In this session, I will discuss practical approaches for operating Kubernetes based infrastructure under real world capacity constraints.
Drawing from my experience building large scale infrastructure systems, I will explore workload placement, resource optimization, infrastructure reliability, and operational scalability across distributed environments. I will share lessons on balancing utilization, resiliency, and performance while supporting growing AI and platform workloads.
Attendees will gain practical insights into infrastructure operations, capacity management, and scaling distributed Kubernetes environments in resource constrained settings.

### 3:55 PM–4:20 PM · Building My Own IaC Engine Was a Mistake: Why I Now Use OpenTofu

- Room: Salt Palace | Level 1 | Ballroom H
- Speakers: Ryan Djurovich
- Track: OpenTofu Day
- Labels: Beginner, Breakout Session, Community Tooling
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269815

My goal: make it possible to deploy a pre-provisioned Kubernetes cluster with a single IaC command.

Initially, I convinced myself I could build an OpenTofu-like solution in TypeScript that could run on a serverless platform. I started with a stateless design, brought in SQLite, switched to PostgreSQL, added a dependency resolver, custom callbacks, and more.

The prototype was fast, but it was brittle, difficult to evolve, and unreliable. The primary feedback from developers was that they wanted a “terraform-like view” to show them previews and progress during the apply step.

I realised I didn’t need to replace the engine. A custom OpenTofu provider, opinionated modules, and a CLI for codegen could deliver on the original goal, with a great user experience, and no custom IaC engine headaches!

This is an experience report on a failed prototype and what happened next. I’ll walk through the original + current architecture and demo the current solution which has OpenTofu at its core.

### 3:55 PM–4:20 PM · Zero-Loss Event Harvesting on Resource-Constrained Edge Gateways

- Room: Salt Palace | Level 2 | 255 A
- Speakers: Nikita Verma, Harshita Varma
- Track: Kubernetes on Edge Day
- Labels: Breakout Session, Event + Data Collection, Intermediate
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269717

Collecting high-throughput telemetry and sensor events at the Kubernetes edge breaks when hitting real-world constraints: volatile WAN backhauls, intermittent network partitions, and strict hardware limits on remote gateways. Running heavy Java or Go-based logging agents on resource-constrained x86/ARM edge nodes eats up the exact memory needed for core workloads. When the uplink drops, unbuffered events are lost, while unchecked local disk buffering risks media wear and node panic.

This session delivers a pragmatic, production-ready architecture for edge data collection that survives network drops without exhausting local hardware limits. We break away from typical cloud-centric logging setups to build a rugged, asynchronous event pipeline tailored explicitly for disconnected edge nodes.

Attendees will learn:

Lightweight Kernel Hooks

Bounded Local Spillover

Backpressure-Aware Backhaul

### 3:55 PM–4:20 PM · I Rebuilt Agent Memory Three Times. The Third Time I Treated It Like Data.

- Room: Salt Palace | Level 1 | Ballroom ACE
- Speakers: Artur Ciocanu
- Track: Agentics Day: MCP + Agents
- Labels: Breakout Session, Extending AI Systems with MCP, Intermediate
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269585

Show a data engineer your agent's memory layer and the questions start: retention policy? eviction past a million rows? backups? when two records disagree, who wins? Most agent memory is a vector store with a timestamp, and it fails all of them. I shipped that, then rebuilt it twice.

First attempt: store everything, retrieve by similarity, expire by timestamp. It scaled and recalled garbage, because embedding proximity is not relevance. Second attempt: the cognitive-science taxonomy, episodic, semantic, procedural. Nicer categories, same open questions: when does a memory go stale, and how do you catch two facts that now contradict?

Third attempt: treat memory as what it is, a stateful data management problem. Retention and acquisition rules, compaction and TTL, indexing, and provenance you can audit and roll back. None of this is new; it runs on Kubernetes like any datastore. MCP, or plain REST, is just delivery. I will demo a system built this way, and the discipline behind it.

### 3:55 PM–4:20 PM · Toward AI-Native Platform Operations: Releasing 100+ Microservices with Agentic AI

- Room: Salt Palace | Level 1 | Ballroom IG
- Speakers: Kai Levin, Afshan Hashemi
- Track: Agentics Day: MCP + Agents
- Labels: Any Level, Breakout Session, Enterprise Integration + Governance
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1253661

Telecom networks serving billions run on cloud-native infrastructure. At Ericsson, our internal application developer platform provides 100+ reusable microservices built on Kubernetes, Helm, Prometheus, OpenTelemetry, Envoy, OPA, Cilium, Argo, etc.

This talk presents Ericsson’s journey toward AI-native platform operations through an agentic release system that turns manual release coordination into a governed path toward autonomous release. The release agent plans, verifies, monitors, recovers, and improves release execution while keeping humans in control at critical decision points.

Agentic AI acts as a reasoning layer above existing DevOps tools. MCP servers connect the agent to Jira, GitLab, Confluence, Gerrit, and Jenkins, while skills handle release-specific tasks.

Attendees will learn how governed agents make release operations AI-native and scale autonomous release potential across 100+ cloud-native services with better speed, consistency, auditability, and resilience.

### 3:55 PM–4:20 PM · Keeping Backstage Plugins Current with Official Codemods

- Room: Salt Palace | Level 2 | 255 EF
- Speakers: Paul Schultz, Alex Bit
- Track: BackstageCon
- Labels: Beginner, Breakout Session, Transforming Dev Experience + Streamlining Dev Workflow + Doing More with Less Using Backstage
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1267485

Keeping Backstage plugins current is one of the biggest maintenance challenges in the ecosystem. Breaking APIs, renamed props, and deprecated configurations mean hours of manual migration with every version bump. But what if a single command could handle it for you?
This talk introduces the official Backstage Codemods repository, a collection of AST-based code transformations that automate the upgrade path for each release. Unlike AI-assisted migrations, codemods are entirely deterministic—the same input always yields the same output. They require no API keys, generate zero hallucinated code, and provide fully auditable transforms.
We will walk through how codemods work, from individual transforms to full migration recipes that chain changes into a single pass. The session includes a live demo upgrading a real plugin across multiple versions using codemod. Whether you maintain one plugin or dozens, you will leave with a repeatable workflow to stay current with minimal effort.

### 3:55 PM–4:20 PM · Drop Your Memory: Do You Really Need Redis for Argo CD?

- Room: Salt Palace | Level 1 | 151
- Speakers: Adi Ziv, Netanel Kadosh
- Track: ArgoCon
- Labels: Breakout Session, Intermediate, Scalability
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1231593

Running ArgoCD at scale usually means managing more than just GitOps. Like many teams, we started with the out of the box in-cluster Redis setup for caching, but while managing hundreds of clusters and thousands of applications, we reached a point where the defaults were no longer enough.
Redis slowly became a platform we had to maintain. Between operational overhead, network inefficiencies, reliability concerns, and compute costs, we realized we were spending too much time managing Redis instead of focusing on GitOps.
In this session, we’ll share how we moved ArgoCD to a redis-free serverless architecture, reducing costs by over 90%, eliminating operational overhead, and improving reliability- without changing dev workflows.
This is a real-world case study focused on what worked, what didn’t and the lessons learned. We’ll cover scaling bottlenecks we hit in large ArgoCD environments, the migration approach, and the measurable impact of moving to a cloud-agnostic serverless backend.

### 3:55 PM–4:20 PM · When Canaries Meet Databases: Handling Schema State Changes in Argo Rollouts

- Room: Salt Palace | Level 2 | 251 A-C
- Speakers: Sakshi Nasha, Primanshu Choudhary
- Track: ArgoCon
- Labels: Breakout Session, Intermediate, Progressive Delivery
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269375

Automated progressive delivery with Argo Rollouts works beautifully for stateless microservices but what happens when an application release requires a database schema change?
If a canary rollout detects an application regression and triggers an automatic rollback, a changed database state can easily break or corrupt production data. This session delivers a practical architectural blueprint for managing database state changes seamlessly during a GitOps progressive delivery cycle. We will break down how to design backward-compatible schema adjustments ensuring that old and new versions of an application can safely operate simultaneously.
Attendees will walk away with an actionable operational framework to guarantee that automated application rollbacks never break the underlying data layer.

### 3:55 PM–4:20 PM · Don't Move That Petabyte! Running Inference Next to the Data

- Room: Salt Palace | Level 2 | 254
- Speakers: David Aronchick
- Track: Cloud Native AI + Inference Day
- Labels: Beginner, Breakout Session, Edge AI + Specialized Environments
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1268384

Nearly every enterprise AI architecture looks the same - Pull all the data into one lake, point the GPUs at it, send answers back out. And while it's clean, it also falls apart the second the data is too big, too sensitive, or too spread out to move. If you've worked with real enterprise data, that's most of the time where the lake is the slowest, most expensive, most legally interesting part of the whole system.

In this talk, we'll discuss how to bring the model to the data. We'll walk through what it takes to run inference where the data already lives, across regions and silos that are never going to consolidate no matter how nicely you ask, and what that does to latency, cost, and the conversation with your security team. I'll be straight about what gets harder: getting models out to a fleet, keeping versions from drifting, and seeing what ran where. You'll leave knowing how to decide which direction your problem actually wants: model to data, or data to model.

### 3:55 PM–4:20 PM · Prefill Here, Decode There: Kubernetes-Native LLM Inference Disaggregation with KAITO and llm-d

- Room: Salt Palace | Level 2 | 251 D-F
- Speakers: Andy Zhang (OSTC), Linbo He
- Track: Cloud Native AI + Inference Day
- Labels: Any Level, Breakout Session, LLMs + Generative AI
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1263142

The problem: P/D disaggregation improves LLM serving efficiency by 30-60% on prefill-heavy workloads, but infrastructure complexity is brutal — routing sidecars, NIXL KV transfer, ZMQ discovery, port management, scheduling profiles, and independent autoscaling all need correct orchestration.

Existing solutions: NVIDIA Dynamo provides disaggregation but is tightly coupled to NVIDIA's stack, not Kubernetes-native. llm-d offers Kubernetes-native P/D routing via Gateway API but requires manual StatefulSet config, sidecar injection, and env var plumbing.

KAITO's MultiRoleInference CRD bridges this gap: a declarative layer orchestrating llm-d components automatically. One CRD generates prefill/decode StatefulSets with correct port assignments, NIXL env vars, decode-only sidecar injection, InferencePool with proper targetPort, EPP plugin chain, and KEDA ScaledObjects per role.

We'll show eval data: TTFT reduction, throughput gains, and autoscaling behavior under mixed workloads.

### 3:55 PM–4:20 PM · Simplifying OpenTelemetry: How the Packaging SIG Makes Installation (Much) Easier Than Nagios

- Room: Salt Palace | Level 2 | 250 A-C
- Speakers: Antoine Toulme
- Track: Observability Day
- Labels: Beginner, Breakout Session, Project Deep Dives
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1244585

What do Nagios and CollectD have that OpenTelemetry lacks?

An installation method as simple as `apt install opentelemetry`!

Today, a typical user needs to install and configure instrumentation SDKs (per app and/or globally per machine), optionally set up the OpenTelemetry injector and collector. Add a layer of fleet management with OpAmp to fine tune config. If you factor in vulnerability patching and regular updates, maintaining this type of solution is a full time job.

The newly minted OpenTelemetry Packaging SIG will take you through how they endeavored in making OpenTelemetry easy to install and configure. We will show you how the OpenTelemetry many packages piece together to form as a cohesive product. We will demonstrate how you will be able to install OpenTelemetry fleet-wide in a few minutes.

You will leave this talk with the knowledge necessary to standardize your OpenTelemetry deployments on Linux, and the tips to use overlay configurations to fine tune them.

### 3:55 PM–4:20 PM · Auto Deriving Threat Signals from OTel CPU Profiles & OBI Network Flows

- Room: Salt Palace | Level 2 | 250 D-F
- Speakers: Shivanshu Raj Shrivastava
- Track: Observability Day
- Labels: Breakout Session, Intermediate, Telemetry Pipelines + Data Engineering
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269839

Most security tooling is bolted on after the fact: agents, sidecars, and rule engines that teams must configure, tune, and maintain.

But two signals you may already be collecting for observability, Continuous CPU profiles and network flows, quietly encode a remarkably rich security picture.

A profile that suddenly shows say a crypto-mining hot path, or a workload that starts talking to an IP it never contacted before, is a threat indicator hiding in plain sight.

In this talk we show how to fuse OTeL eBPF derived CPU profiles with L3/L4/L7 network flows to derive baseline behavior per workload and surface anomalies, lateral movement, unexpected egress, suspicious execution without writing a single detection rule.

We'll walk through the data model, how profiles and flows correlate to a single workload identity, and a live demo turning raw telemetry into actionable detections.

Attendees leave able to extract "free" security value from observability data they already own.

### 3:55 PM–4:20 PM · When in-toto Met OPA: Turning Attestations into E2E Security & Compliance with Cloud Native Policies

- Room: Salt Palace | Level 2 | 255 BC
- Speakers: Snahil Singh, Trishank Kuppusamy
- Track: Open Source SecurityCon
- Labels: Any Level, Breakout Session, Supply Chain + Application Security
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269657

Today, there are in-toto attestations about artifacts like code reviews & build provenance, but no policies to turn them into end-to-end software supply chain security narratives. Was the same container image uploaded to the registry as the one built on CI? Did the CI use two-party-reviewed source code?

The in-toto policy framework was designed to connect different attestations together. It allows different policy languages with reproducible, portable verification, but no compliant implementation existed for a de facto engine. OPA is such an engine, but was missing a supply chain security story.

We present the first implementation of the in-toto policy framework written on OPA enabling ergonomic, reusable, and extensible policies powerful enough to verify artifacts from development to deployment. Results are reproducible & portable so different engines can verify them despite using different languages. We illustrate with examples ranging from container registries to package managers.

### 3:55 PM–4:20 PM · Compliance Drift on Kubernetes: Building Proof Infrastructure for AI Agents in Regulated Industries

- Room: Salt Palace | Level 1 | Ballroom BDF
- Speakers: Saideep Navakoti
- Track: Platform Engineering Day
- Labels: Breakout Session, Intermediate, Platform Engineering and AI
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1266966

When regulators audit a banking application, they ask: can you prove this system does what you said it does? For deterministic software, the CI/CD pipeline is the proof. AI agents on Kubernetes break every proof mechanism. Same image, same manifest, same cluster — different answer every time. The model drifts, the context window changes, the RAG corpus refreshes. GitOps shows zero changes while behavior has shifted. This talk presents four Kubernetes-native patterns for compliance proof infrastructure. Behavioral baselines as Kyverno admission control — blocking promotion when drift exceeds thresholds. Continuous proof generation via OpenTelemetry traces capturing model version, prompt hash, context sources, and response fingerprints. A custom CRD controller detecting behavioral drift and triggering re-approval workflows. And automated regulatory evidence generation from cluster telemetry, replacing manual spreadsheets with auditor-queryable dashboards.

### 3:55 PM–4:20 PM · What Customer Success Knows That Platform Engineering Keeps Relearning

- Room: Salt Palace | Level 1 | Ballroom J
- Speakers: Ryan Etten
- Track: Platform Engineering Day
- Labels: Any Level, Breakout Session, Platform Adoption Stories
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1260420

Platform engineering keeps rediscovering what customer success worked out a decade ago. Onboarding decides who stays, a green dashboard can hide a quiet exodus, and adoption is earned, not mandated. The field treats these as fresh discoveries, missing a playbook most teams have never read.

Drawing on years of leading enterprise adoption, this talk starts where platform teams rarely look. You have no authority over whether a customer uses what they bought. Internal platform teams are in that exact position. Your developers are customers who churn quietly, routing around you, unseen by your metrics.

It matters more as platforms are asked to govern AI agents. Before a team trusts a platform with an agent, they have to trust it with a deploy, and that trust is built in the same adoption work CS has done for years. This talk maps that playbook onto internal platforms. It ends with the only adoption question that matters. If your platform vanished tomorrow, who would actually feel it?

### 4:30 PM–4:55 PM · Private Data Providers for Clean Data Parsing

- Room: Salt Palace | Level 1 | Ballroom H
- Speakers: Craig Sebenik
- Track: OpenTofu Day
- Labels: Breakout Session, Intermediate, Upgrades and Usage
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269634

HCL is simple and easy to use. But, once you have to read and parse complex data files, it can lead to some pretty complicated and hard to manage code; a maze of `yamldecode`, `flatten`, `try`, and `lookup` calls that nobody wants to review in a PR. The logic is buried, untestable, and fragile.

Private providers offer a better path. Moving YAML parsing into a Go provider, you get to use real programming paradigms: typed structs, error handling, and unit tests. The HCL stays simple and readable. Your parsing logic lives in a place where `go test` actually works.

The primary motivation was to minimize the amount of change during a vendor migration. The data formats vendors inherently use can be wildly different. This talk lays out how and why we decided to move the complexity to Go and leave the HCL as simple as possible. The pros and cons of the approach will be discussed so that one can get into custom providers knowing what problems they solve and what problems they introduce.

### 4:30 PM–4:55 PM · Your OS Is an OCI Artifact: Immutable Kubernetes Across 1,500 Retail Stores

- Room: Salt Palace | Level 2 | 255 A
- Speakers: William Rizzo
- Track: Kubernetes on Edge Day
- Labels: Breakout Session, Intermediate, Kubernetes in Production on the Edge
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269840

Retail at scale breaks the usual operational playbook: 1,500 stores, no on-site operators, WAN links that may drop, and a hard rule that nothing is ever patched by hand. This session is a production case study in eliminating configuration drift by construction.
Every store runs a three-node NUC cluster with LibVirt and KVM, hosting a four-node k0s cluster whose nodes are Kairos-based Rocky Linux VMs. Those nodes are never upgraded in place. The operating system is built as an OCI artifact and versioned in Git; every change triggers a CI pipeline that rebuilds and pushes a fresh image, and Flux rolls it out as a full immutable rebuild. Cluster lifecycle is managed declaratively through Cluster API.
The session covers why and how each store runs a standalone control plane, the tradeoffs that choice imposes versus Hosted Control Plane, and the honest limits of an immutable-rebuild model, including how a single Git commit reaching 1,500 stores is kept from becoming a fleet-wide outage.

### 4:30 PM–4:55 PM · Protocol Solved, Design Still Open: Domain-Driven Design for MCP Servers

- Room: Salt Palace | Level 1 | Ballroom ACE
- Speakers: Fabrizio Lazzaretti, Annegret Junker
- Track: Agentics Day: MCP + Agents
- Labels: Breakout Session, Building MCP Servers + Clients, Intermediate
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1250074

The Model Context Protocol (MCP) went from zero servers in late 2024 to thousands within a year, most thin wrappers over REST endpoints and database tables. When agents fail, it's rarely the model that's weak; it's that the tool surface is incoherent. We're repeating the REST-over-CRUD mistakes of the 2010s, one layer up.
An MCP tool is an API contract for a non-deterministic consumer: the name, description, and schema are all the agent has to work with. Domain language is a reliability concern, not an aesthetic one.
This talk shows how Domain-Driven Design maps onto MCP design. Building on Anthropic's finding that code execution can cut token usage by up to 98.7%, the question isn't "should we build an MCP server?" but "which capabilities are tools, which are code APIs, and which are skills?"
A before-and-after example leaves you with a repeatable method: start from the bounded context, classify each capability by surface, and name tools in your business's words, not your database's.

### 4:30 PM–4:55 PM · Detonation and Isolation: Hardening Kubernetes Runtimes for Agents

- Room: Salt Palace | Level 1 | Ballroom IG
- Speakers: Syeda Anjum, Kaslin Fields, Kavitha Rajendran
- Track: Agentics Day: MCP + Agents
- Labels: Breakout Session, Intermediate, Security + Trust + Reliability
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269602

Giving an AI agent access to terminal environments and code execution tools is mathematically necessary for advanced reasoning, yet systems-wise terrifying. A model capable of generating Python scripts to verify its own logic can accidentally drop production databases, exfiltrate environment credentials, or be exploited by prompt injection attacks. This session addresses the security mechanics of ‘The Secure Runtime’, illustrating how to safely isolate, govern, and audit code-executing agents on Kubernetes.

We detail the complete container security boundary designed for "untrusted" model execution. We explore the deployment of sandboxes to run the tool executions, redirecting and filtering application-level system calls in user space to block access to the underlying host kernel. Solve the cold-start challenge for pods with heavy libraries to re-initialized sandboxes instantly. Finally, we will showcase how to implement least-privilege principles by filtering agent intent dynamically.

### 4:30 PM–4:55 PM · The Hardest Part of Adopting Backstage Has Nothing to Do with Backstage

- Room: Salt Palace | Level 2 | 255 EF
- Speakers: Juan Pablo Garcia Ripa, Alejandro Quisbert
- Track: BackstageCon
- Labels: Beginner, Breakout Session, Lessons Learned from Standing Up + Deploying + Driving Backstage Adoption
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1254707

Backstage is a flexible and powerful framework, but frustration in adoption often appears when the organization is missing some basic requirements the tool depends on: ownership mapping, clear team boundaries, infrastructure ready to be automated, a culture of documentation and cross-team contribution.

This talk is about those gaps — not as a troubleshooting checklist, but as an honest account of what they look like in practice and why recognizing them early changes how you approach adoption entirely.

Whether you're evaluating Backstage or already using it and hitting walls, you'll know what to look for when adoption starts to feel harder than it should.

### 4:30 PM–4:55 PM · Minimizing Deployment Blast-Radius with Argo CD Progressive Sync

- Room: Salt Palace | Level 1 | 151
- Speakers: Alexandre Gaudreault, Kanika Rana
- Track: ArgoCon
- Labels: Beginner, Breakout Session, Progressive Delivery
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1258900

Rolling out changes across dozens or hundreds of clusters simultaneously is a recipe for disaster. Even with Argo Rollout or environment promotion, simultaneously syncing a faulty change can cause an outage. The ApplicationSet Progressive Sync feature gives teams fine-grained control over the order and pace of Application syncs — enabling staged rollouts that gate on application health before proceeding.
In this session, we will explain how Progressive Sync fits into the broader continuous deployment ecosystem, walk through the RollingSync strategy, and demonstrate the new ApplicationSet UI that makes deployment progress visible in real time. We'll also explain the feature complexity and limitations you need to understand before adopting it.
You'll leave with a clear mental model of when Progressive Sync is the right tool, how to configure it, and what to watch out for, so you can deploy confidently.

### 4:30 PM–4:55 PM · When the Controller and the Agent Both Think They're Right: Reconciliation in an Agentic Cluster

- Room: Salt Palace | Level 2 | 251 A-C
- Speakers: Bert Bullough
- Track: ArgoCon
- Labels: Breakout Session, Intermediate, Software Delivery
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1257387

GitOps reconciliation tells auditors a clean story: when someone hand-mutates the cluster, self-heal reverts it to the approved Git state, and "restoring desired state isn't a change" holds up. Point an agentic remediation system at the same cluster, one that fixes incidents by changing state on purpose, and it stops being true. Self-heal and the agent fight, each "correcting" the other, both making unapproved changes to production. The "drift correction isn't a change" line depends on the thing causing drift being a fat-fingered human. Make it autonomous and the argument evaporates. In a regulated environment, that's a finding.

From standing up agentic infrastructure inside a FedRAMP High boundary, this talk walks the failure modes when reconciliation collides with autonomous agents, why no vocabulary covers "two non-human actors disagree about production," and what arbitration must be: who wins, who logs, how the audit trail survives. A problem 18 months from everyone's doorstep.

### 4:30 PM–4:55 PM · Your KV Cache Is Bigger Than Your GPU: Scaling Agentic Inference with Networked Storage

- Room: Salt Palace | Level 2 | 254
- Speakers: Kfir Toledo, Miro Nikolov
- Track: Cloud Native AI + Inference Day
- Labels: Any Level, Breakout Session, LLMs + Generative AI
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1264176

Agentic workflows push context lengths past 100K tokens per session, and at production concurrency the KV cache outgrows GPU HBM by orders of magnitude. The default response is to recompute prefills, burning GPU cycles on work already done. The bottleneck isn't compute. It's that there's nowhere to keep the cache.
llm-d, the CNCF Sandbox project for K8s-native vLLM, treats networked storage as a first-class KV tier. In this joint IBM Research + Google Cloud session, we share what we learned shipping the storage tier: a POSIX FS connector validated on Google Cloud Managed Lustre, IBM Storage Scale, and Ceph; GPU Direct Storage that skips the CPU; and an object store backend via NVIDIA NIXL.
We benchmark throughput, latency tails, and concurrency across all three. You leave knowing how storage choices boost inference throughput, how to size storage for systems holding 1+ billion KV tokens, and when storage offload pays off versus when GPU memory still wins.

### 4:30 PM–4:55 PM · Building Multi-Cloud AI Platforms on Kubernetes Without the Pain

- Room: Salt Palace | Level 2 | 251 D-F
- Speakers: Romil Bhardwaj
- Track: Cloud Native AI + Inference Day
- Labels: Breakout Session, Intermediate, LLMs + Generative AI
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1262100

AI workloads are redefining how modern platforms are architected. They require elastic GPU capacity, but no single cloud can meet all needs reliably or affordably. Kubernetes can become the de-facto substrate for AI orchestration, but going multi-cloud with Kubernetes slows down even the best infra teams. And for many AI engineers, the steep learning curve of Kubernetes itself remains a major barrier to adoption.

This talk is a hands-on guide to building a multi-cloud, Kubernetes-based AI platform that unifies resources across hyperscalers, neoclouds, and on-prem clusters into a single compute abstraction. We'll walk through practical details including open-source tooling for workload scheduling, automated cloud selection for cost optimization, experiment tracking and dependency management. This approach lets AI engineers use the same unified interface for interacting with compute, enabling them to focus on building great AI products rather than wrestling with cloud complexity.

### 4:30 PM–4:55 PM · The Trace Stops Here? Rethinking Context Propagation for Managed Services

- Room: Salt Palace | Level 2 | 250 A-C
- Speakers: Liudmila Molkova
- Track: Observability Day
- Labels: Breakout Session, Correlation + Context + Cross-Project Architectures, Intermediate
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269536

You're a SaaS, IaaS, or PaaS company. You expose an API, do some magic, and respond. When things go wrong, you need observability - but can you trust the trace context the caller passed in?

Probably not sampling decision or baggage - those are theirs, not yours. And if you reuse their trace ID, your platform's observability now hangs on a caller you don't control.

It gets harder once you host AI agents or complex workflows, where observability matters to you *and* your callers. One trace or two about the same request? How do you stay observable at the depth you need while showing users what matters to them?

We'll argue for treating ingress and egress as trust boundaries: link to the incoming context instead of reusing it, start your own trace, and trim baggage and tracestate at every external edge so nothing leaks in or out. We'll cover what OTel makes possible today, what's still missing, and how to turn the cloud black box into a gray one.

### 4:30 PM–4:55 PM · Rethinking Telemetry Pipelines: What Happens When Telemetry Stays Columnar End-to-End?

- Room: Salt Palace | Level 2 | 250 D-F
- Speakers: Laurent Querel, Joshua MacDonald
- Track: Observability Day
- Labels: Breakout Session, Intermediate, Telemetry Pipelines + Data Engineering
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1264694

Telemetry volumes continue to grow as systems become more complex, instrumentation becomes richer, and AI agents introduce new streams of logs, metrics, and traces that are expensive to transport and process.

This session presents a first-principles rethink of the telemetry data plane: what changes when telemetry remains columnar across the entire path? Using the OTel Arrow Dataflow Engine as an open source case study, the speakers will show how OTAP, Apache Arrow, and a columnar execution model reduce serialization boundaries, improve data locality, and reshape the design of receivers, processors, and exporters.

The talk connects architecture to measured results: in published Phase 2 benchmarks, the OTAP path delivered roughly 10x to 20x higher throughput than the OTLP path on the same engine. Attendees will learn which architectural choices mattered most, what tradeoffs emerged, and what this means for the future of OpenTelemetry pipeline design.

### 4:30 PM–4:55 PM · AI-Driven Security for Open Source: Lessons From Envoy

- Room: Salt Palace | Level 2 | 255 BC
- Speakers: Boteng Yao, Yan Avlasov
- Track: Open Source SecurityCon
- Labels: Any Level, Breakout Session, Leveraging + Preparing for AI in OSS Security
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1268398

AI is changing open source security from a reactive disclosure process into a continuous vulnerability discovery and triage loop. This talk uses Envoy as the anchor example to explain a new operating model for AI-driven security management in large OSS projects: how researchers can use LLMs/Mythos and program-analysis tools to find risky code paths, review patches for security regressions, and prioritize issues before they become CVEs.

The session will cover concrete patterns: building a vulnerability-hunting workflow around threat models, historical CVEs, and maintainer review; separating useful AI findings from hallucinated reports; handling responsible disclosure when AI increases report volume; and defining governance, embargo, release, and backport practices that scale. Attendees will leave with a practical blueprint for using AI as a security amplifier while keeping human maintainers accountable for final judgment.

### 4:30 PM–4:55 PM · Automated Resource Allocation Driven by Service Level Objectives

- Room: Salt Palace | Level 1 | Ballroom BDF
- Speakers: Pedro Célestin, Julia Furst Morgado
- Track: Platform Engineering Day
- Labels: Advanced, Breakout Session, Platform Engineering Case Study
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1249654

Platform engineering has done a lot for developers. Internal portals and container orchestration have made deploying to production no longer the hard part. However, resource allocation still lands on the developer's plate, and most of them have no idea what to put there. This decision becomes a guessing game where they either copy from a colleague or ask the platform team, who don't understand the workload.The data to automate this—SLOs and error budgets—already exists but is disconnected from infrastructure action. This session presents a working feedback loop where Prometheus tracks SLI, KEDA triggers scaling, Argo handles orchestration, and Knative handles abstraction. OpenTelemetry feeds the signal that makes all of it reliable. Attendees will learn how to distinguish resource-caused SLO burns from bugs and validate every resource allocation change against real SLI data. The result is a self-managing platform where developers define SLOs and the infrastructure automatically adapts.

### 4:30 PM–4:55 PM · Data-Driven Platform Guardrails: What 16,500 Incidents Tell You to Automate First

- Room: Salt Palace | Level 1 | Ballroom J
- Speakers: Henrique Santana, Henrique Dalssaso
- Track: Platform Engineering Day
- Labels: Breakout Session, Intermediate, Platform Engineering Case Study
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1237497

Platform engineering teams face a constant prioritization challenge: which guardrails to build next? Most decide based on the last escalation or the loudest internal customer. We analyzed 16,500 production Kubernetes incidents to identify which failure patterns are preventable at the platform layer.

The results reshape priorities. Node provisioning failures (15% of incidents) are nearly all preventable through platform-managed node lifecycle. Authentication breakdowns (10%) disappear with centralized certificate rotation. But pod scheduling errors (12%) resist platform-level fixes because they stem from application-specific resource declarations.

This talk gives platform engineers a data-backed framework for deciding what to automate, what to guard with policy, and what to solve with better developer experience. Each of the 9 patterns maps to a specific platform intervention: admission policies, automated remediation, self-service tooling, or improved defaults.

### 5:00 PM–5:10 PM · No More Click-Ops Compliance: Postgres Audit Logging with OpenTofu

- Room: Salt Palace | Level 1 | Ballroom H
- Speakers: Akash Singh
- Track: OpenTofu Day
- Labels: Intermediate, Upgrades and Usage, ⚡ Lightning Talk
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269415

Database compliance usually lives in someone's head or a console nobody audits. At Sabi, our Postgres audit logging, roles, and grants were click-ops i.e. undocumented, drift-prone, and stressful every time a PCI audit came around. This case study shows how we moved it all into OpenTofu.
In 10 minutes, I'll walk through managing Postgres audit logging and access controls as code: declaring it in OpenTofu, code-reviewing every change, catching drift before an auditor does, and getting the versioned history of who changed what and why : exactly the access-trail evidence PCI DSS Requirement 10 asks for. I'll cover how we structured modules, handled secrets and state safely, and what doesn't belong in IaC.
Attendees will leave with a practical pattern for treating database compliance as code: reproducible, reviewable, and audit-ready for PCI and beyond.

### 5:00 PM–5:10 PM · The PLC and the Pod: Putting Kubernetes on the Plant Floor Without Tripping the Line

- Room: Salt Palace | Level 2 | 255 A
- Speakers: Miguel Rojas
- Track: Kubernetes on Edge Day
- Labels: Intermediate, Kubernetes in Production on the Edge, ⚡ Lightning Talk
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1263619

Cloud-native teams arrive at the factory with continuous delivery, push-based GitOps, and the expectation that systems should reconcile state on demand. The plant floor has other ideas. Change windows can last for months, safety systems have veto power, and OT networks often treat your control plane as an unwanted guest.

Drawing on experience from semiconductor manufacturing and industrial energy sites, this lightning talk covers three OT realities every edge Kubernetes operator needs to understand. We'll look at practical design patterns that help Kubernetes workloads and PLCs coexist without disrupting operations.

You'll leave with one design rule you can put to work on Monday.

### 5:00 PM–5:10 PM · Cents and Sensibility: Cost Aware Practices for Tool Calling Agents

- Room: Salt Palace | Level 1 | Ballroom ACE
- Speakers: Erica Hughberg
- Track: Agentics Day: MCP + Agents
- Labels: Any Level, Enterprise Integration + Governance, ⚡ Lightning Talk
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269778

Every team making LLM calls hits the same wall. Rate limits are measured in requests or tokens, but the people paying the bill care about dollars.

Tool-calling agents can quickly spend dollars through iterations and context growth. A short burst of expensive model calls can blow a budget while sitting comfortably under a token quota. Neither requests nor tokens are a stable unit of cost. A token is priced differently across models, inputs, outputs, and cached contexts.

This talk explores the latest practices for cutting tool-calling spend, what it takes to set limits in real currency enforced where the traffic actually flows, and how a policy that understands model pricing turns “requests per second” into something a budget owner recognizes: dollars per team per day.

Attendees leave with a mental model for cost-based rate limiting that they can take to any AI gateway or proxy.

### 5:00 PM–5:10 PM · Reducing Cycle Time 10x: Building an Intelligent Ticket Pipeline on Kubernetes

- Room: Salt Palace | Level 1 | Ballroom IG
- Speakers: Gonzalo Vazquez
- Track: Agentics Day: MCP + Agents
- Labels: Beginner, Extending AI Systems with MCP, ⚡ Lightning Talk
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1230824

Every platform team faces the same bottleneck: transforming intake requests into actionable work items. At RBC, our hybrid cloud team's manual process averaged 45 minutes per request with inconsistent quality.

This talk demonstrates an event-driven pipeline on Kubernetes that reduces cycle time to under 5 minutes. The architecture combines Argo Events for GitHub webhook triggers, Argo Workflows for orchestration, and MCP servers as containers for LLM-powered requirement analysis.

Attendees will learn:
- Value Stream Mapping to identify waste and build automation business cases
- Event-driven architecture with Argo (CNCF Graduated) for GitOps-native pipelines
- Containerizing AI agents (MCP servers) for reproducible, auditable automation
- Metrics frameworks for cycle time and quality measurement

Includes a live demo: creating a GitHub issue and watching the pipeline analyze, decompose, and create structured Jira tickets with functional and non-functional requirements in real-time.

### 5:00 PM–5:10 PM · The Guardrail Paradox: Why Secure Software Templates Actually Speed Up Your Day 1 Code

- Room: Salt Palace | Level 2 | 255 EF
- Speakers: Sakshi Nasha
- Track: BackstageCon
- Labels: Beginner, Transforming Dev Experience + Streamlining Dev Workflow + Doing More with Less Using Backstage, ⚡ Lightning Talk
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269427

When software engineers hear the words "compliance checks" or "security gating," they immediately envision delayed releases and broken deployment pipelines. This beginner-friendly lightning talk completely flips that script. We will look at how treating security as a Day 1 helper rather than a final gate, drastically accelerates developer velocity.

Using Backstage Software Templates, we will demonstrate how to build an invisible developer paved path. We will show how a novice engineer can spin up a pristine, production-ready microservice that natively contains verified cryptographic identity handles and pre-configured policy limits from the very first commit.

Attendees will learn how shifting security constraints directly into their catalog scaffold completely eliminates the late-stage friction of fixing infrastructure bugs right before a deployment deadline.

### 5:00 PM–5:25 PM · Argo CD’s Health Checks Are Your Secret Weapon to CD Nirvana

- Room: Salt Palace | Level 1 | 151
- Speakers: Dan Garfield
- Track: ArgoCon
- Labels: Breakout Session, Intermediate, Software Delivery
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269555

Kubernetes has basic health checks to cover the container lifecycle, but so much of what happens on a cluster is about what’s happening around the container. Ingress assignments, storage allocation, workload identity, and so many others are common failure points in Kubernetes. That’s all before we get to custom resources!

But Argo CD has a secret weapon, the ability to create custom health checks for any resource. Argo CD ships with hundreds of custom health checks already for handling common resource types, but when you can easily write and add your own to handle any scenario. These checks allow you to track state in dynamic and powerful ways.

It turns out that kind of dynamic, fast feedback is key to better CD, especially when driven by AI, and much more reliable deployments and informed progressive delivery. Join this deep dive into what Argo CD health checks could be doing for you.

### 5:00 PM–5:10 PM · Scheduling AI Workloads on Kubernetes with Argo: Patterns from Production

- Room: Salt Palace | Level 2 | 251 A-C
- Speakers: Neel Shah
- Track: ArgoCon
- Labels: Any Level, Data Processing, ⚡ Lightning Talk
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269722

Building AI workflows on a large scale is not only about the models but also a scheduling challenge as well. AI jobs such as training will be competing for expensive GPUs, while batch inference should run alongside the real-time serving, and sometimes a failure in the pipeline can cost many hours of compute time. In this talk, we share proven practices for building AI pipelines in production with Argo Workflows on Kubernetes.

During this presentation, we'll show how to build GPU-aware directed acyclic graphs for efficient utilization, how to ensure proper recovery after a failure in the long running AI training jobs, how to deal with model artifacts management and versioning during different pipeline steps, and how to set up human approval for the pipeline before starting its execution.

### 5:00 PM–5:10 PM · One GPU Pool to Serve Them All: Scheduling with Complex Priorities

- Room: Salt Palace | Level 2 | 254
- Speakers: Yuriy Natarov
- Track: Cloud Native AI + Inference Day
- Labels: End-to-End AI/ML in Production (MLOps/AIOps), Intermediate, ⚡ Lightning Talk
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1268868

Production LLM inference platforms rarely run one workload type. The same GPU pool must serve dedicated endpoints with per-user guarantees, shared model endpoints with bursty demand, and batch inference workers that should use spare capacity without starving online traffic. At Nebius, this exposed a big ops problem for our multi-tenant inference platform: each tier needed different scheduling semantics, and manual capacity decisions made correct workload provisioning take hours.

This talk shares how the architecture evolved on Kubernetes, why the team chose KAI Scheduler after comparing alternatives and which criteria mattered most. The resulting solution reduced provisioning time from hours to minutes.

### 5:00 PM–5:10 PM · Building Apps with AI Agents: A LangGraph Pipeline from Idea to Deploy

- Room: Salt Palace | Level 2 | 251 D-F
- Speakers: Gonzalo Vazquez
- Track: Cloud Native AI + Inference Day
- Labels: Beginner, LLMs + Generative AI, ⚡ Lightning Talk
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1266667

Most AI agent demos stop at a chatbot. This live, code-first session shows a multi-agent pipeline that takes a plain-language prompt and drives it through a real software development lifecycle (architecture, build, test, deploy), ending in a running app.
We’ll walk the pipeline, built on LangGraph in TypeScript: conditional routing, parallel agent forks, human-in-the-loop checkpoints, and feedback loops. It’s also wired for cloud native operations: OpenTelemetry traces every agent step end to end, and CloudEvents moves work between phases, keeping the agent graph observable and event-driven instead of a black box.
You’ll leave with a concrete, vendor-neutral pattern for orchestrating LLM agents as production-grade systems, plus the failure modes that appear once agents touch real delivery: runaway loops, cost, and where humans still belong in the loop.

### 5:00 PM–5:10 PM · The Negative Diagnosis: Why "Not an Ingress Issue" Is the Most Valuable Verdict

- Room: Salt Palace | Level 2 | 250 A-C
- Speakers: Ritvik Pandya
- Track: Observability Day
- Labels: Correlation + Context + Cross-Project Architectures, Intermediate, ⚡ Lightning Talk
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1230522

Most observability systems are built to detect problems and surface them.Most automation is built to react and remediate. Both miss something important: the most valuable verdict an automated diagnostic system can produce is often "this is not my problem."
This talk argues that for any observability system claiming to attribute latency or root cause,the negative diagnosis -the confident statement that a subsystem is not the cause -is more valuable than the positive detection.It changes the social contract of incident response:from "everyone investigates everything" to "the system that owns the signal owns the diagnosis."
Drawing on work building component-level latency attribution inside an edge proxy,this talk presents the negative diagnosis as a design pattern: how to build it,why its threshold structure differs from positive detection,why it is harder to defend,and why systems that know when not to fire are the systems engineers actually trust enough to deploy autonomously.

### 5:00 PM–5:25 PM · From YAML Blues to Autocomplete: Schema-First Config for OpenTelemetry Collector Components

- Room: Salt Palace | Level 2 | 250 D-F
- Speakers: Jan Korona
- Track: Observability Day
- Labels: Breakout Session, CI/CD + Platform + Develop-Experience Observability, Intermediate
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1239971

The OpenTelemetry Collector is a central extension point of the observability ecosystem, but component configuration is still split across Go structs, custom validation, component-specific defaults, and manual documentation. That creates drift for users, maintenance overhead for contributors, and a lack of support for validation and editor autocomplete.

This talk introduces the schema-first roadmap for Collector component configuration. Component authors will define configuration through YAML schema specifications that generate Go structs, JSON Schemas, and documentation. We will cover why Go-struct-first and AST-parsing approaches fell short, how schema-first generation fits with existing tooling such as metadata.yaml and mdatagen, and how core and contrib components can migrate without breaking existing user configuration.

As one of the main engineers working on this feature, I will explain the design tradeoffs, compatibility constraints, and adoption path for component authors.

### 5:00 PM–5:10 PM · Your Agent Trusts Its Tools Too Much: A 10-Minute Tour of MCP Tool-Layer Attacks

- Room: Salt Palace | Level 2 | 255 BC
- Speakers: Saurabh Yergattikar
- Track: Open Source SecurityCon
- Labels: Intermediate, Leveraging + Preparing for AI in OSS Security, ⚡ Lightning Talk
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269826

10 minutes, one uncomfortable truth: your agent's biggest risk is not the user, it is the tools it already trusts. This lightning talk shows three live attack patterns against MCP agents: a poisoned tool description, a hidden instruction smuggled in a tool response, and a malicious server pulled from a public registry.

Drawing on the SAFE-MCP taxonomy (Linux Foundation/OpenSSF, 80+ techniques), I show why prompt-level guardrails miss all three, then demo a runtime proxy that validates tool calls inline with no agent or server changes, cutting tool-poisoning success from 74% to under 9% at under 120ms per call.

Attendees will leave with the 3 things to vet before approving any MCP server for your stack.

### 5:00 PM–5:10 PM · From Clever to Correct: Lessons from Rebuilding SeatGeek's Platform AI Agent

- Room: Salt Palace | Level 1 | Ballroom BDF
- Speakers: Colin O'Dell
- Track: Platform Engineering Day
- Labels: Intermediate, Platform Engineering and AI, ⚡ Lightning Talk
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1266455

Rufus, SeatGeek's internal platform AI agent, was a victim of its own cleverness. We were on the bleeding edge, making bold bets on ideas before they'd had time to prove themselves. By the time the broader community converged on better patterns, like agent skills, unwinding our decisions wasn't a refactor - it was a rewrite.

In 10 minutes, we'll cut straight to the decisions that hurt us most: where we over-specialized, where we misread where the community was heading, and what we'd do differently today. No fluff - just the hard-won lessons we wish we'd had before building v1.

### 5:00 PM–5:10 PM · The Signals We Often Miss: Lessons from Years of Platform Engineering

- Room: Salt Palace | Level 1 | Ballroom J
- Speakers: Shivangi Motwani
- Track: Platform Engineering Day
- Labels: Any Level, Using Platforms to Drive Developer Experience, ⚡ Lightning Talk
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269651

Over years, We've worked on multiple platform teams. The technologies changed, the teams changed, and the organizational structures changed. The struggles did not.

The same themes kept surfacing: discoverability, ownership, feedback loops, and cognitive load. While platform engineers focused on building reliable and scalable systems, the teams consuming those platforms were often trying to answer simpler questions: Where do I start? Who owns this? Am I doing this correctly? What should I do next?

Pattern that eventually became clear: engineers naturally optimize technology before understanding user experience.

This talk shares what that realization changed about how we approach platform engineering and how other teams might take from it.

### 5:15 PM–5:25 PM · DRY Out Your Tofu, with Symbol Libraries!

- Room: Salt Palace | Level 1 | Ballroom H
- Speakers: Christian Mesh
- Track: OpenTofu Day
- Labels: Beginner, Breakout Session, Upgrades and Usage
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1266232

Recently added to OpenTofu, Symbol Libraries finally allow re-use of constants, types, and declarations of functions within configuration. In this talk, we will cover how to leverage this as an end user, as well as how modules and providers can be supercharged with symbols.

### 5:15 PM–5:25 PM · Bootstrapping Edge Nodes When the SPIRE Server Is 200ms Away and Sometimes Gone

- Room: Salt Palace | Level 2 | 255 A
- Speakers: Manan Chawda
- Track: Kubernetes on Edge Day
- Labels: Intermediate, Kubernetes in Production on the Edge, ⚡ Lightning Talk
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269793

Most zero-trust bootstrapping guides assume the identity server is reachable. At the edge that assumption breaks regularly. SVID rotation intervals tuned for a data center become a real operational risk when a node drops off the network for six hours and comes back expecting to re-attest.
In this talk we will walk through a production K3s deployment where we configured SPIFFE/SPIRE to handle disconnected SVID rotation, replaced long-lived tokens on physical hardware with short-lived credentials, and used a locally deployed OPA instance to enforce admission policy without a live connection to the control plane. We will cover the specific SPIRE configuration decisions that made this work and the two failure scenarios that forced us to rethink our rotation interval assumptions entirely. This talk will give edge practitioners a tested identity bootstrapping pattern for environments where intermittent connectivity is not an edge case but the default operating condition.

### 5:15 PM–5:25 PM · Nobody Is Signing MCP Tool Definitions. That Is a Supply Chain Problem

- Room: Salt Palace | Level 1 | Ballroom ACE
- Speakers: Nikunj Doshi
- Track: Agentics Day: MCP + Agents
- Labels: Extending AI Systems with MCP, Intermediate, ⚡ Lightning Talk
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269809

You probably sign your container images. You might sign your Helm charts. But if you are building agentic systems with Model Context Protocol today there is a good chance the tool definitions telling your agents which APIs to call and what actions to take are being shipped with no signing, no attestation, and no verification at runtime.
In this talk we will make the case that MCP tool definitions belong inside the same supply chain security model we apply to any other software artifact. We will demonstrate a working pattern using Cosign to sign MCP server manifests, Rekor to store the attestations, and an OPA policy to verify tool definition provenance before an agent is permitted to invoke anything. We will close with an honest look at where this pattern holds, where it breaks down, and what the CNCF community will need to build together to make agentic supply chain security something practitioners can rely on.

### 5:15 PM–5:25 PM · Agents Need a REPL: Programmatic Tool Calling with MCP

- Room: Salt Palace | Level 1 | Ballroom IG
- Speakers: Ahmed Ibrahim
- Track: Agentics Day: MCP + Agents
- Labels: Intermediate, MCP in Production: Case Studies, ⚡ Lightning Talk
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269590

Most agents call one tool per model turn. That adds inference round trips and fills the context window with intermediate results.
This talk covers programmatic tool calling with MCP: the model writes a small sandboxed program that discovers tools, runs calls in parallel, loops over results, and filters data locally.
Using Codex code mode as a case study, it covers capability scoping, approvals, nested-call tracing, partial failures, and bounded tool output.

### 5:15 PM–5:25 PM · Instrument Backstage with OpenTelemetry: From Zero Visibility to Full Traces in Minutes

- Room: Salt Palace | Level 2 | 255 EF
- Speakers: Mithun Banerjee
- Track: BackstageCon
- Labels: Any Level, Technical Deep Dives on Backstage Topics, ⚡ Lightning Talk
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269056

Backstage has built-in OpenTelemetry support that most teams never configure. The result — plugin failures, scaffolder timeouts, and catalog slowdowns that are completely invisible.

This lightning talk shows exactly how to enable full OTel observability in your Backstage instance with a single instrumentation file — capturing traces and metrics from the catalog, scaffolder, and plugins in under 10 minutes.

What we cover:
- Why Backstage is an observability blind spot by default
- The one instrumentation file that changes everything
- Key traces and metrics you unlock immediately
- Where to send your Backstage telemetry

Takeaway:
- Copy-paste OTel setup for Backstage backend
- Instant visibility into catalog, scaffolder, and plugin health

### 5:15 PM–5:25 PM · The Silent Drops: Making Argo CD Network-Aware with eBPF and Cilium Hubble

- Room: Salt Palace | Level 2 | 251 A-C
- Speakers: Emily Chen
- Track: ArgoCon
- Labels: Any Level, Observability, ⚡ Lightning Talk
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1268610

Your pods are running, your probes are passing, and the Argo CD dashboard shows a glowing green "Synced & Healthy" status. Yet, downstream users are experiencing errors because a misconfigured NetworkPolicy is silently dropping traffic. Traditional Kubernetes readiness probes operate within the pod boundary, leaving the GitOps control plane unaware of runtime data-plane failures. This visibility gap forces platform engineers to dig through logs and run terminal traces during outages.

This lightning talk presents a solution to bridge the gap between runtime eBPF telemetry and GitOps state. Attendees will discover how to leverage Argo CD’s Lua-based Custom Health Checks paired with custom resource status fields updated by Cilium Hubble flow metrics. The plan is to demonstrate a lightweight, vendor-neutral architectural loop that catches dropped packets and forces Argo CD to automatically mark a deployment as Degraded when an eBPF policy drop occurs.

### 5:15 PM–5:25 PM · An Agent Is Not a Function Call

- Room: Salt Palace | Level 2 | 254
- Speakers: Shaghayegh Moradirad
- Track: Cloud Native AI + Inference Day
- Labels: Intermediate, RAG + AI Agents, ⚡ Lightning Talk
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269652

Most AI APIs are shaped like function calls: send input, receive output. Coding agents break that abstraction. They operate over time, keep state across turns, stream intermediate work, act on workspaces, consume budgets, and sometimes need to be steered or interrupted.

Drawing from the design of the public OpenAI Codex Python SDK, this lightning talk asks what the right interface to an agent runtime should expose. The answer is not just a cleaner wrapper around inference. It is a set of control-plane primitives: threads for persistent context, turns for units of work, streams for observability, sandbox policies for workspace access, and result objects that preserve status, errors, timing, usage, items, and final output.

The central claim is simple: once a model can act, the API boundary becomes a safety boundary, an observability boundary, and a product boundary at the same time. Treating agents as function calls hides the machinery operators need to trust them.

### 5:15 PM–5:25 PM · Who Let the Agent in? Securing Agentic AI with Non-Human Identity in Cloud Native Infrastructure

- Room: Salt Palace | Level 2 | 251 D-F
- Speakers: Mahendran Selvakumar
- Track: Cloud Native AI + Inference Day
- Labels: Any Level, RAG + AI Agents, ⚡ Lightning Talk
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1266390

AI agents are starting to perform real operational tasks across Cloud Native environments including Kubernetes,CI/CD,observability,security and platform workflows. But many agents still rely on shared tokens, broad credentials or inherited human permissions.This creates risks around over-authorization, invisible delegation,weak auditability and uncontrolled access.
This session presents a Cloud Native security pattern for securing agentic AI with non-human identity and an agent gateway.It explores how Kubernetes service accounts,SPIFFE/SPIRE,Open Policy Agent,Kyverno,GitOps workflows,short-lived credentials and audit logs can help verify agent identity,route agent actions through a controlled gateway,enforce least privilege,bind every tool or API call to policy and require approval for risky operations.
Attendees will leave with a practical blueprint for moving from trusted agents to verified agents without giving them standing access to Cloud Native infrastructure.

### 5:15 PM–5:25 PM · OpenTelemetry Metrics Just Got 25× Faster

- Room: Salt Palace | Level 2 | 250 A-C
- Speakers: Cijo Thomas
- Track: Observability Day
- Labels: Any Level, Observability Cost + Quality + Scale + Efficiency, ⚡ Lightning Talk
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1251884

OpenTelemetry's metrics performance has long been a sore point for high-throughput users. Recording a simple Counter with three attributes/labels cost around 50 nanoseconds per call - enough to show up in CPU profiles and rule OpenTelemetry out of the busiest code paths.

Pre-resolved metric handles are an old answer. Prometheus client libraries expose labelled-metric handles you can cache, and Windows Performance Counters have used the same pattern for decades. OpenTelemetry now offers it as a first-class API, called bound instruments, and on that same hot path, recording drops to under 2 nanoseconds - roughly 25× faster (Measured in OTel Rust Sdk)

This lightning talk shows where the speedup comes from, when to reach for it (high-frequency counters with a fixed, known attribute set), and - more importantly - when not to. Used the wrong way, the new fast path can be slower than the original.

Flexible by default, fast when you need it.

### 5:15 PM–5:25 PM · Sherlock Pods: Investigating a Compromised Kubernetes Cluster

- Room: Salt Palace | Level 2 | 255 BC
- Speakers: Maxime Coquerel
- Track: Open Source SecurityCon
- Labels: Beginner, Security Education, ⚡ Lightning Talk
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1261993

Kubernetes was not designed for forensics but attackers don’t care, and you still need evidence when something goes wrong.

In this session, we dive into how to perform forensic investigations in compromised Kubernetes clusters: capturing container and node artifacts, preserving their integrity, and reconstructing what actually happened in a system built to be distributed, dynamic, and ephemeral.

You’ll walk away with practical techniques for evidence acquisition and post-incident analysis that still work when workloads vanish, nodes auto-scale away, and logs disappear along with a clear understanding of what is realistically possible (and what is not!) when investigating incidents in Kubernetes.

### 5:15 PM–5:25 PM · The Emperor’s New Code: Your Developer Portal’s New Audience Is an Agent

- Room: Salt Palace | Level 1 | Ballroom BDF
- Speakers: Erica Hughberg
- Track: Platform Engineering Day
- Labels: Any Level, Platform Engineering and AI, ⚡ Lightning Talk
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269822

Developers don’t handwrite most of their code anymore, and they’re not reading your docs either. They describe what they want, and the agent just builds.

It’s the Emperor’s New Clothes: code everyone signs off on that no one actually read, running, acting, and spending in production.

So who is your developer portal really for? You wrote the golden paths, capability catalogs, and integration guides for humans. But the thing building against your platform now is an agent. If it can’t consume that knowledge, it improvises, and you get code shipped to prod that never touched a paved path, reviewed by no one.

Platform teams have a new primary audience: the agent.

This talk covers how to make platform capabilities and golden paths agent-consumable, so agents build consistently, and what runtime guardrails must catch when no one reads the code. Attendees leave knowing how to rethink a developer portal for an audience that builds without reading.

### 5:15 PM–5:25 PM · 10 Years, 10 Lessons

- Room: Salt Palace | Level 1 | Ballroom J
- Speakers: Ram Iyengar
- Track: Platform Engineering Day
- Labels: Any Level, Platform Adoption Stories, ⚡ Lightning Talk
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1269741

I've been working with platforms for 10 years now. I've served several roles in the process. I've been an engineer, a product manager, and for the longest time a community manager. My experience with platforms also spans several kinds.

I've struggled myself and watched other teams struggle with their platform design, adoption, and onboarding. This talk is designed as a quick refresher on what has worked well and what hasn't in my experience. The lessons span the following areas:
1. Abstractions
2. Platform definition
3. Platform maturity
4. Measurement
5. Costs
6. Developer experience
7. Ownership
8. Feedback loops
9. CI - Continuous Improvement
10. Business outcomes

### 5:25 PM–5:30 PM · Closing Remarks

- Room: Salt Palace | Level 1 | Ballroom H
- Track: OpenTofu Day
- Labels: Any Level
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1312065

### 5:25 PM–5:30 PM · Closing Remarks

- Room: Salt Palace | Level 2 | 255 A
- Track: Kubernetes on Edge Day
- Labels: Any Level
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1312063

### 5:25 PM–5:30 PM · Closing Remarks

- Room: Salt Palace | Level 1 | Ballroom ACE
- Track: Agentics Day: MCP + Agents
- Labels: Any Level
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1312051

### 5:25 PM–5:30 PM · Closing Remarks

- Room: Salt Palace | Level 2 | 255 EF
- Track: BackstageCon
- Labels: Any Level
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1312054

### 5:25 PM–5:30 PM · Closing Remarks

- Room: Salt Palace | Level 2 | 254
- Speakers: Rajas Kakodkar, Yuzhui Liu, Yuan Tang
- Track: Cloud Native AI + Inference Day
- Labels: Any Level
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1312057

### 5:25 PM–5:30 PM · Closing Remarks

- Room: Salt Palace | Level 2 | 250 A-C
- Track: Observability Day
- Labels: Any Level
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1312059

### 5:25 PM–5:30 PM · Closing Remarks

- Room: Salt Palace | Level 2 | 255 BC
- Track: Open Source SecurityCon
- Labels: Any Level
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1312060

### 5:25 PM–5:30 PM · Closing Remarks

- Room: Salt Palace | Level 1 | Ballroom BDF
- Track: Platform Engineering Day
- Labels: Any Level
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1312061

### 5:30 PM–5:40 PM · The Full Stack Preview: Kubernetes, Argo CD, and Ephemeral Databases

- Room: Salt Palace | Level 1 | 151
- Speakers: Cecil Wöbker
- Track: ArgoCon
- Labels: Breakout Session, Intermediate, Software Delivery
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1250455

Preview environments promise developers isolated testing for every pull request. But implementations tend to stop at stateless services, sharing databases across environments and hoping for the best. When PR #42 corrupts shared data, everyone suffers.

This talk shares our journey building production-grade preview environments with data isolation. Using ArgoCD ApplicationSets, each GitHub PR automatically provisions a complete environment—including its own database.

You'll learn:

- ApplicationSet Pull Request generators for automatic environment lifecycle
- Per-PR isolated database provisioning
- The seeding problem: synthetic data vs. sanitized production snapshots
- Cleanup automation: Jobs and finalizers for orphaned resources
- Cost control: Spot instances, shared infrastructure, and smart resource limitsWe'll cover what worked, what broke, and the tradeoffs we made. Walk away with patterns for preview environments that catch bugs before staging—not after.

### 5:30 PM–5:40 PM · From Commit to Rollout: Tracing GitOps with Argo CD and OpenTelemetry

- Room: Salt Palace | Level 2 | 251 A-C
- Speakers: Antonio Jimenez Martinez
- Track: ArgoCon
- Labels: Intermediate, Observability, ⚡ Lightning Talk
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1255714

OpenTelemetry is making CI/CD pipelines observable, but in GitOps environments delivery does not end when CI succeeds. After the pipeline updates a GitOps repository, Argo CD reconciliation, Kubernetes controllers, health checks, and Argo Rollouts determine whether a change actually reaches production.

This session shows how to trace GitOps deployments with Argo CD, Argo Rollouts, and OpenTelemetry. We will model the delivery path from commit, build, and GitOps update through sync, rollout, promotion, rollback, and production incident context.

Through a live demo, attendees will learn how to correlate asynchronous GitOps events using trace context, span links, deployment metadata, Kubernetes annotations, semantic conventions, and OpenTelemetry Collector pipelines.

### 5:30 PM–6:30 PM · Evening Reception (Available for All CNCF Co-Located Events)

- Room: Meeting Room Foyers
- Track: Agentics Day: MCP + Agents
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1309453

Join us onsite for drinks and appetizers with fellow co-located attendees from Monday's CNCF-hosted Co-located Events.

Attendees from all CNCF Co-located Events are welcome.

### 5:40 PM–5:45 PM · Closing Remarks

- Room: Salt Palace | Level 1 | 151
- Track: ArgoCon
- Labels: Any Level
- Link: https://events.linuxfoundation.org/kubecon-cloudnativecon-north-america/co-located-events/cncf-hosted-co-located-schedule/?id=1312052

