Mark Collier
Executive Director
PyTorch Foundation
Mark has a proven track record of fostering organizational growth in neighboring open source foundations. Mark co-founded the Open Stack Foundation (now OpenInfra Foundation), and has played a key role in the development and global success of OIF. Over the past 13 years, he established it as a leading cloud platform and one of the most dynamic projects in the history of open source computing. Since its modest beginnings under Mark’s leadership in 2012 OSF/OIF has expanded into a community of 8 platinum, 10 gold and 66 silver corporate members. Mark’s robust knowledge of AI, including a central role in shaping the definition of open source AI through the Open Source Initiative, makes him an ideal leader for LF AI & Data and the PyTorch Foundation.
Conference Sessions
- Keynote: Welcome to PyTorch Conference North America Oct 20, 2026 • 9:00 AM-9:10 AM Grand Ballroom (Concourse Level)
- Keynote: Welcome Back Oct 21, 2026 • 9:00 AM-9:10 AM Grand Ballroom (Concourse Level)
Brian Stevens
Senior Vice President and AI CTO
Red Hat
Brian is SVP and AI CTO for Red Hat, joining following the acquisition of Neural Magic where he was CEO. At Red Hat he focuses on the strategic direction of the AI portfolio, including open-source efforts and ecosystem development. A tech veteran, Brian has a rich history of building high-impact companies and driving disruptions that transform the industry. In his role at Neural Magic, Brian led the vision of democratizing Generative AI for enterprises, through open source LLMs together with an open source inference stack. Neural Magic was the top developer of the industry’s de facto open source inference server, vLLM, as well as the leader in LLM compression techniques. In his prior role as VP of Product and CTO of Google Cloud, Brian shifted and led the product and business transformation to focus on the needs of enterprise organizations. He also defined and brought to market the industry’s first cloud AI services leveraging the pioneering work of the Google Brain team. Prior to Google, Brian served as Red Hat’s CTO and EVP of Worldwide Engineering. Brian led the shift toward enterprise, including the creation of the enterprise subscription model which together transformed Red Hat’s business. Brian currently serves on the board of directors of Nutanix, Genpact and Zus Health. Brian holds a master’s degree in computer systems from Rensselaer Polytechnic Institute and a bachelor’s degree in computer science from the University of New Hampshire. In his personal life, Brian is an avid runner and cycler, and together with his wife Beth has a passion for restoring historic properties.
Conference Sessions
- Keynote: Open Source Inference at Production Scale Oct 20, 2026 • 10:20 AM-10:27 AM Grand Ballroom (Concourse Level)
Bill Jia
VP of Engineering, Core ML/AI
Google Cloud
Bill Jia is a VP of Engineering at Google’s Core ML/AI org, responsible for making AI models (e.g. Gemini, Veo, Nano Bananas) and AI Infrastructure (e.g. TPU), from research to production, great at Google and for the world.
Conference Sessions
-
Keynote: Workload Fungibility in the Age of Agents
Oct 21, 2026 • 9:20 AM-9:30 AM
Grand Ballroom (Concourse Level)
PyTorch is fundamentally designed to embrace hardware heterogeneity. From the flexibility of eager mode to the extreme performance of torch.compile and TorchDynamo, the PyTorch stack empowers developers to build truly hardware-agnostic models. Coupled with advanced distributed APIs like FSDP2, extensive parallelism strategies, TorchTitan, inference engines such as vLLM and SGLang, and Helion, PyTorch provides the definitive ecosystem for scaling AI workloads across any silicon. Today, we are thrilled to showcase this architectural vision in action with a deep dive and some real-world use cases of TorchTPU, now publicly available in OSS. Co-developed in close partnership with PyTorch core maintainers, TorchTPU showcases PyTorch’s inherent workload fungibility and a production-ready lowering path. By seamlessly integrating with the core PyTorch ecosystem – from modeling in HuggingFace Transformers, training in TorchTitan, and serving in both vLLM and SGLang – TorchTPU allows teams to achieve out-of-the-box pretraining, fine-tuning, and serving performance using the PyTorch code they already know. It is already powering production workloads at top-tier AI labs globally. To further accelerate PyTorch’s multi-hardware reality and leading position among AI practitioners, we are also introducing a new paradigm of agentic developer experiences. We will demonstrate how long-horizon agentic workflows can seamlessly migrate complex model workloads from GPUs to TPUs. Beyond migration, we’ll explore how you can leverage these agents for autonomous performance hill-climbing—tackling complex, low-level optimizations like quantization, custom kernel generation, and advanced sharding strategies to profile, debug, and extract maximum hardware utilization with minimal manual intervention. Join us as we explore the frontier of PyTorch hardware heterogeneity, and discover how TorchTPU and agent-driven development can radically simplify your path to high-performance AI.
Sara Hooker
CEO & Co-Founder
Adaption
Sara Hooker is a co-founder of Adaption, which builds intelligence that continuously evolves. Sara leads a large team of AI researchers and engineers that build extremely efficient, adaptable systems. Sara Hooker was previously VP of Research at Cohere, a $6.8 billion frontier AI company focused on generative AI for enterprise. Prior to Cohere, she built large systems in computer vision and NLP at Google Deepmind. Her work has been featured in mainstream news outlets including Techcrunch, New York Times, Washington Post, Axios, MIT Technology, The Atlantic. Sara is a frequent expert advisor to AI research and policy initiatives around the world: she is currently on Kaggle's ML Advisory Research Board and serves on the World Economic Forum council on the Future of Artificial Intelligence and the Future of Data Frontiers. She has been listed as one of AI's top 13 innovators by Fortune and one of Time100 Most Influential People in AI.
Conference Sessions
-
Keynote: Beyond Brute Force: The Era of Adaptive Intelligence.
Oct 20, 2026 • 9:55 AM-10:03 AM
Grand Ballroom (Concourse Level)
The next generation of truly intelligent systems won't be defined by scale. The next unlock is architectural: systems that keep learning after deployment, closing the gap between a model's frozen training distribution and the evolving world it operates in. In this talk, Sara Hooker, Co-founder of Adaption, discusses moving past static AI toward systems that are malleable by design. She'll dig into the technical foundations of continual, gradient-free learning: how models can update behavior without full retraining, without catastrophic forgetting, and without the compute cost of repeated fine-tuning cycles. Built so that AI adapts to people rather than the reverse.
Rachad Alao
Vice President of Engineering
Cohere
As a seasoned technology leader, Rachad brings over two decades of experience in the high-tech industry, with a focus on responsible artificial intelligence, operating systems and embedded systems. Currently, he serves as Vice President of Engineering at Cohere, leading Cohere's North organization, which focuses on building Cohere's Agentic AI platform. Rachad's professional journey spans multiple roles and companies. He was previously Senior Director of Engineering at Meta Superintelligence Labs, leading Meta's AI Trust and Safety organization focused on building safe and responsible AI foundation models, including Llama open-source and Muse Spark models. Before rejoining Meta, he was Senior Director of Engineering at Google, where he led the Applied Responsible AI effort as well as Google Maps User-Generated Content, Trust and Safety AI engineering teams. Previously, he was Director of Engineering at Facebook AI, leading the Facebook Responsible AI effort centered around AI fairness, privacy, security, robustness, and safety, as well as transparency and control. Rachad's experience also includes six years at Google as an Engineering Director, helping Android scale to 3 billion devices. He led Android's multimedia efforts, including Android OS Audio, Video, and Camera teams, contributing to the development of Google's Pixel Camera and a more secure Android multimedia stack. Earlier in his career, Rachad spent nine years at Bluestreak Technology in Montreal, serving as Chief Technology Officer and other engineering leadership roles. As CTO, he helped build the company's embedded Flash and HTML technology, deploying it to millions of devices for Orange France and Time Warner Cable US. Rachad holds a Master's degree in Telecommunication Engineering and Computer Architecture from Telecom Paris in Paris, France. He was born in Benin, West Africa and grew up in Côte d'Ivoire before moving to France at the age of 16 to pursue his studies.
Conference Sessions
-
Keynote: Open Science and AI Foundation Models: Allies in the Pursuit of Knowledge
Oct 21, 2026 • 10:05 AM-10:13 AM
Grand Ballroom (Concourse Level)
This talk will discuss how open science provides a powerful paradigm to ensure new AI foundation models are developed and deployed in a way that is transparent and secure, while enabling speed and scale of innovation across the ecosystem.
Alban Desmaison
Meta
Alban has been working on PyTorch since nearly its inception, first during his PhD at the University of Oxford and now at Meta. He is focused on maintaining core components, designing a wide breadth of features and fostering the PyTorch Community.
Conference Sessions
- Keynote: PyTorch Updates Oct 20, 2026 • 9:20 AM-9:35 AM Grand Ballroom (Concourse Level)
-
Meet the Developers of PyTorch and Safetensor
Oct 20, 2026 • 10:35 AM-11:10 AM
Community Expo (Concourse Level)
Join PyTorch experts for an interactive “Meet the Developers” session – a chance to learn in a small group setting and ask your most pressing questions. No sign-up required – just drop in and get real answers from the people building the tools.
Mazin Gilbert
Executive Director
Agentic AI Foundation
Mazin Gilbert is the Executive Director of the Agentic AI Foundation, part of the Linux Foundation. An IEEE Fellow with a Ph.D. in Artificial Intelligence and a Wharton MBA, Mazin brings 25+ years of experience driving AI innovation from research to global-scale deployment. His career spans senior leadership at Google as Director of Engineering for Google Distributed Cloud AI/ML, and at AT&T as VP of Network Analytics and Automation. Mazin has built production-grade GenAI, machine learning, and agentic AI platforms at scale. He co-founded open-source landmarks including ONAP, Akraino, and Acumos, and helped found the LF Networking, LF Edge, and LF AI. He has served on the boards of the Open Network Foundation, NSF MLWINS, and the International Computer Science Institute at Berkeley. Mazin holds 260+ U.S. patents, has authored 100+ research papers, and became an IEEE Fellow in 2012 for pioneering contributions to speech processing.
Conference Sessions
- Keynote: Mazin Gilbert, Executive Director, Agentic AI Foundation Oct 21, 2026 • 9:10 AM-9:15 AM Grand Ballroom (Concourse Level)
Jana van Greunen
Director of PyTorch Engineering
Meta
Jana van Greunen is the Director of PyTorch Engineering at Meta, where she leads efforts to ensure PyTorch remains the leading AI/ML framework for researchers and developers worldwide. With deep expertise in distributed systems, large-scale infrastructure, and over 15 years of experience building high-performance systems, Jana drives PyTorch's evolution to meet the demands of cutting-edge AI workloads—including partnerships like TorchTPU.
Conference Sessions
-
Keynote: PyTorch at Meta – From Internal Workloads to Multi-Silicon Ecosystem
Oct 20, 2026 • 10:05 AM-10:10 AM
Grand Ballroom (Concourse Level)
Jana van Greunen shares how Meta uses PyTorch to power its frontier AI research and production workloads — and how that work feeds back into the open-source ecosystem. She'll cover Meta's latest PyTorch initiatives, including Helion (the kernel layer) and TorchTitan's expansion into reinforcement learning, as well as Meta's deepening collaboration with hardware partners to bring PyTorch-native support to the next generation of AI chips.
Robert Hundt
VP Distinguished Engineer
Amazon
Robert Hundt is the Senior Tech Lead for Amazon's Neuron SW stack. Prior to this new role, Robert spent the last 20 years at Google, working on many compiler and performance project. Specifically, Robert was the senior Tech Lead for the TPU SW stack, including AI/ML Compilers (XLA), Performance (xprof), and TorchTPU. He also wrote the book "Quantum Computing for Programmers", published by Cambridge University Press.
Conference Sessions
-
Keynote: Trainium’s Journey to Native PyTorch
Oct 20, 2026 • 9:50 AM-9:55 AM
Grand Ballroom (Concourse Level)
PyTorch now runs natively on Trainium with no code changes required. AWS customers are already training frontier models using standard PyTorch workflows on Trainium. In this keynote, we share how we got here: enabling eager mode and torch.compile, integrating with TorchTitan, TorchAO, and HuggingFace Transformers v5, and making our tools like Neuron Explorer work natively with PyTorch. We'll cover our upstream contributions to PyTorch and TorchTitan, and our ongoing collaboration with the PyTorch team on what comes next.
Carlos Costa
Distinguished Engineer
IBM
Conference Sessions
- Keynote: Open Source Inference at Production Scale Oct 20, 2026 • 10:20 AM-10:27 AM Grand Ballroom (Concourse Level)
Abdullah Gharaibeh
Senior Staff Software Engineer
Google Cloud
Conference Sessions
- Keynote: Open Source Inference at Production Scale Oct 20, 2026 • 10:20 AM-10:27 AM Grand Ballroom (Concourse Level)
Ujval Kapasi
VP, AI & HPC Frameworks and Libraries
NVIDIA
Ujval Kapasi is a senior leader at NVIDIA, helping to build software and systems that make large-scale AI and computing possible and practical. Throughout his career, Ujval has led teams working on AI infrastructure, math libraries, developer tools, and performance optimization. A particular focus for his team is to help accelerate major open source AI frameworks, such as PyTorch, JAX, vllm, and sglang.
Conference Sessions
-
Keynote: PyTorch Performance for Agents on Vera Rubin
Oct 20, 2026 • 10:30 AM-10:40 AM
Grand Ballroom (Concourse Level)
In this keynote, Ujval will explore what agentic workloads need from AI systems and how the PyTorch stack can address these needs. Agentic systems introduce new bottlenecks and optimization opportunities. A single request can trigger long chains of reasoning, tool calls, code execution, data movement, and distributed inference. Training and inference serving software needs to be optimized across CPUs, GPUs, memory, and networking. Using Vera Rubin as a case study, he will show how co-design across hardware and software can improve end-to-end performance, efficiency, and developer experience for agents at scale.
Chris Lattner
Executive Vice President of Advanced AI Software and Platforms
Qualcomm
Conference Sessions
-
Keynote: MAX and Mojo: The Best Way to Extend PyTorch
Oct 20, 2026 • 9:35 AM-9:40 AM
Grand Ballroom (Concourse Level)
PyTorch is where the industry builds models. Chris Lattner shows what happens when you extend that foundation with MAX and Mojo. Mojo is a systems programming language in the Python family. MAX adds model authoring plus a graph compiler that targets hardware from multiple vendors, so a kernel you write once runs across them instead of being rewritten per backend. Together, they give PyTorch developers an entrypoint into systems programming and performance optimization, so they can get more out of the hardware they already have.
Cody Mazza-Anthony
Senior Staff ML Engineer
Shopify
Conference Sessions
-
Keynote: Shopify's Continual Learning Loop with PyTorch and vLLM
Oct 20, 2026 • 9:10 AM-9:15 AM
Grand Ballroom (Concourse Level)
The flywheel turns Shopify's product expertise into robust evaluations, then mines, critiques, and repairs low-scoring conversations to train better models. Using Buyer Profile as a case study, we show how small models can outperform frontier models on well-scoped tasks at a fraction of the cost, with lower latency and higher throughput.
Simon Mo
Co-Founder and CEO
Inferact
Simon is the co-founder of Inferact, a startup founded by creators and core maintainers of vLLM, the most popular open-source LLM inference engine. Inferact's mission is to grow vLLM as the world's AI inference engine and accelerate AI progress by making inference cheaper and faster. Simon has co-led the vLLM community effort from UC Berkeley since 2023. Simon previously worked at Anyscale and Character AI on serving systems.
Conference Sessions
-
Keynote: vLLM Update: Scaling Open Frontier Inference Infrastructure
Oct 20, 2026 • 10:15 AM-10:20 AM
Grand Ballroom (Concourse Level)
As LLMs grow in size, context length, and architectural complexity, vLLM must evolve to meet new performance and scalability challenges. This talk presents key improvements in vLLM's core architecture and highlights major optimizations in KV cache management and GPU kernels.
Priya Nagpurkar
Vice President, AI-native Systems Research
IBM
Priya Nagpurkar is the Vice President of AI-native Systems research at I.B.M.'s T.J. Watson Research Center in New York. She leads a global team of systems and AI experts working on all interesting aspects of the AI platform — from scalable, performant, and efficient deployment of AI to the use of AI to transform how systems are built and managed in the future. Her team has worked on multiple influential open-source projects such as Istio, Knative, PyTorch, vLLM, llm-d and contributed to several impactful IBM offerings. Priya’s early work at IBM Research focused on characterizing and optimizing emerging workloads across different layers of the application and system stack, with an emphasis on language runtimes and processor architecture and design. Priya holds a Ph.D. in Computer Science from the University of California, Santa Barbara. She is passionate about making computing fast and consumable.
Conference Sessions
-
Keynote: From Silicon to Models at Scale
Oct 21, 2026 • 9:30 AM-9:35 AM
Grand Ballroom (Concourse Level)
"Enabling new hardware, supporting thousands of models, and scaling inference takes progress across the open-source AI stack. This keynote shares how IBM and Red Hat are contributing across the PyTorch ecosystem to make that happen. We’ll highlight work bringing IBM Spyre to the PyTorch compiler and vLLM, how AI tools are helping scale model enablement to thousands of models, and a look ahead at next-generation Spyre hardware. We’ll also cover Red Hat’s contributions to PyTorch and CI/CD infrastructure for out-of-tree accelerators, helping hardware backends keep pace with the community. Finally, we will touch on our journey with the llm-d community to enable efficient distributed inference at scale."
Mark Saroufim
Co-Founder
Core Automation
Mark Saroufim is a former PyTorch maintainer, cofounder of GPU MODE, and cofounder at Core Automation. His work focuses on GPU programming, ML systems, and making high-performance computing more accessible through open-source tools, competitions, and technical communities.
Conference Sessions
-
Keynote: Linear Algebra for the Age of Research
Oct 21, 2026 • 10:25 AM-10:33 AM
Grand Ballroom (Concourse Level)
This talk introduces the linear algebra kernels we are working on and why these long-standing problems still deserve attention. I will discuss what the kernels do, where the performance bottlenecks are, and why accelerating these routines matters for research. The talk also highlights how AI tools and open technical communities are helping people make faster progress on problems studied for decades.
Yiwei Song
Senior Director of AI Model Cycle
Crusoe
Yiwei is the Sr. Director of AI Model Cycle at Crusoe. He previously served as a Director at Coupang and an Applied Science Manager at Amazon, where he specialized in training large-scale models to drive significant business impact. He holds a Ph.D. in Information Theory.
Conference Sessions
-
Keynote: AI Model Customization Platform
Oct 21, 2026 • 10:15 AM-10:20 AM
Grand Ballroom (Concourse Level)
We are sharing our vision and progress of building an LLM post-training platform which allows the customers to customize the open weights models for their business usecase, to improve accuracy and reduce inference cost.
Ion Stoica
Professor / Executive Chairman
UC Berkeley / Anyscale
Ion Stoica is a Professor in the EECS Department and Xu Bao Chancellor Chair at the UC Berkeley. He is the Director of the Sky Computing Lab and Executive Chairman of Databricks and Anyscale. His current research focuses on AI systems and cloud computing, and his work includes several open-source projects such as vLLM, SGLang, Chatbot Arena, SkyPilot, Ray, and Apache Spark. He has also co-founded several companies, including LMArena (2025), Anyscale (2019), Databricks (2013), and Conviva (2006).
Conference Sessions
- Keynote: Evolving Ray and Kubernetes Together for the AI Era Oct 20, 2026 • 9:40 AM-9:45 AM Grand Ballroom (Concourse Level)
-
Crossing the Divide: Co-Evolving Kubernetes and Ray for the AI Era
Oct 20, 2026 • 2:50 PM-3:15 PM
210B (Concourse Level)
The AI workload orchestration landscape is currently split across two massive, parallel universes: the CNCF (the bedrock of cloud native and modern infrastructure) and the PyTorch Foundation (the epicenter of AI/ML innovation). While these foundations operate independently, the end user does not have the luxury of choosing just one. PyTorch Foundation contains PyTorch, vLLM, Ray, and more, while CNCF owns Kubernetes, Envoy, OpenTelemetry, llm-d, containerd and more. To build, deploy, and scale modern AI applications, users require a seamless integration of projects from both ecosystems. In this session, Jago Macleod (Google) and Ion Stoica (Anyscale / Databricks / UC Berkeley) will explore the critical bridge connecting these communities: the co-evolution of Kubernetes and Ray into a unified OSS AI Stack. We will discuss how Google and Anyscale, together with the broader community, are actively collaborating to build an open, vertically integrated stack that prevents ecosystem fragmentation.
Edward Yang
Research Engineer
Meta
Edward Yang is a research engineer at Meta who has worked on PyTorch since the very beginning of the project, with involvement in nearly all aspects of the project. Lately, he has been working on framework features targeted at large scale LLM training, such as distributed communication, activation checkpointing and profiling. One of this big interests in the age of AI coding is how deep learning frameworks and compilers should evolve in an era of abundant tokens.
Conference Sessions
- Keynote: PyTorch Updates Oct 20, 2026 • 9:20 AM-9:35 AM Grand Ballroom (Concourse Level)
-
Contributing to PyTorch with AI agents
Oct 20, 2026 • 10:35 AM-11:05 AM
Community Expo (Concourse Level)
How should AI agents be used productively to contribute to PyTorch? What are good things to know? How can you get someone to review your PR? Come talk about these things. A good pre-read is this devlog: https://docs.pytorch.org/devlogs/ai-agents/2026-05-30-ai-coding-playbook/
-
Meet the Developers of PyTorch and Safetensor
Oct 20, 2026 • 10:35 AM-11:10 AM
Community Expo (Concourse Level)
Join PyTorch experts for an interactive “Meet the Developers” session – a chance to learn in a small group setting and ask your most pressing questions. No sign-up required – just drop in and get real answers from the people building the tools.
Zhipeng Wang
Senior Staff Software Engineer
Zhipeng is the Senior Staff Software Engineer and Manager at Google, where he leads ML Systems and AI Infra work for supporting Large Foundation Models training and inference. He also serves as the code maintainer and Technical Steering Committee of DeepSpeed, which is among the most popular open-source LLM training optimization libraries. DeepSpeed’s mission is to democratize Large model training and make it more efficient, effective and easy-to-use. Zhipeng has co-led DeepSpeed community efforts since 2025. He previously worked at LinkedIn, AWS, Google[X] and Apple on leading AI research efforts in the areas of training and serving systems, multimodal models and LLM post-training.
Conference Sessions
-
Keynote: Scaling Large Frontier Models Training with DeepSpeed
Oct 21, 2026 • 9:55 AM-9:58 AM
Grand Ballroom (Concourse Level)
DeepSpeed is among the most popular open-source training optimization libraries that enables memory efficient, scalable and most competitive distributed training performance for large models up to trillion parameters. DeepSpeed’s pioneering work on ZeRO (Zero Redundancy Optimizer) introduced a new paradigm for eliminating memory redundancy in distributed training. The techniques have since been broadly adopted across all the training ecosystem and have become foundational to the training of today’s frontier models. In this talk, I will share some of the latest technical advances from the DeepSpeed team that continue to push the frontier of large-scale model training. In particular, I will highlight our work on model-systems co-design and optimizations, which enable DeepSpeed to deliver high performance across diverse model architectures, workloads, and accelerator platforms. I will also discuss the continued development of the DeepSpeed open-source community and how we collaborate with model developers, hardware vendors, and AI infrastructure partners to expand DeepSpeed’s adoption and enable emerging large-scale AI workloads.
DE LA SOUL
Conference Sessions
-
AI Community Bash featuring DE LA SOUL
Oct 21, 2026 • 6:00 PM-8:30 PM
Concert Venue (Concourse Level)
Close out PyTorch Conference with a night worth sticking around for. We’re bringing the AGNTCon + MCPCon and PyTorch communities together for an evening of food, drinks, great company—and a live concert from legendary hip-hop group DE LA SOUL. Celebrate with fellow developers, maintainers, founders and contributors from across the AI ecosystem, then stick around as DE LA SOUL takes the stage. One last night. Two AI communities. One unforgettable show. Let us know you’re joining when you register for the PyTorch Conference so we hold you a spot! Space is limited!