PyTorch Conference North America
A crowd of attendees smiling and laughing.

Poster Schedule

Poster Presentations

Posters will be displayed in the Community Expo during the entire event. Presentations will take place during the Flare Party, Tuesday, October 21 6:00-7:30 PM PDT.

Filter by
Tuesday October 20
6:00 PM
"Are We There Yet?" (AWTY): A Framework for Assessing Small Model Viability on Local Hardware
Community Expo Ian Ballantyne
Applications
2 in 1: Deep Dive into PyTorch Certified Associate (PTCA) and PyTorch Ambassador Program
Community Expo Sahdev Zala
Introduction
Accelerating Helion Autotuning with Warm-Start and Shared Caches
Community Expo Alessandro Sangiorgi
Kernel Engineering
Accelerating Whisper Inference on Arm CPUs through PyTorch Ecosystem Optimizations
Community Expo Elham Harirpoush, Fadi Arafeh
Inference
Agentic Kernel Optimization at the Scale of the Web
Community Expo Joshua Lochner
Kernel Engineering
Agents as Distributed Processes: A Native PyTorch Design Pattern for Multi-Agent Orchestration
Community Expo Amit R
Applications
AI Edge Quantizer: Streamlining PyTorch Model Optimization for LiteRT On-Device Inference
Community Expo Maria Lyubimtseva, Weiyi Wang
Inference
AI-Driven Supply Chain Intelligence with PyTorch-Based Predictive Models
Community Expo Nishant Verma
Responsible AI
AOTI with CUDA Graph
Community Expo Kevin Fu
Inference
AugVein: Vein Detection through Augmented Reality with PyTorch
Community Expo Vishal Jaimin Vakil
Applications
Automatic Graph-Based Activation Checkpointing and Offloading for TorchTitan Frontier Model Training
Community Expo Nurlan Nazaraliyev
Training
Batch Cost Lookahead: Coordination-Free Workload Regularization for Packed Variable-Length Training
Community Expo Tao Huang
Training
Before the Loss Spike: Training Alignment for Precise PyTorch Debugging at Scale
Community Expo Ziming Zhou
Training
Better Together: Get More Concurrent Users with Heterogeneous Multi-Agent Workload optimization
Community Expo Weifeng Yao, Louie Tsai
Inference
Beyond SPMD: Compiling Triton Kernels for Dataflow AI Architectures
Community Expo Takuya Nakaike, Mori Ohara
Kernel Engineering
Beyond the Execution Layer: Integrating Custom Accelerators with vLLM
Community Expo Burkhard Ringlein
Inference
Beyond the Learner: The CPU’s Role in Scaling PyTorch Multi-Agent Reinforcement Learning
Community Expo Sagar Surendran, Na Li
Training
Breakable CUDA Graphs: Eager Breaks in CUDA Graph Capture for PyTorch LLM Inference
Community Expo Frederik Gossen
Inference
Breaking torch.distributed on Purpose: Chaos Engineering for Large-Scale PyTorch Training
Community Expo Paulo Aragao
Training
Bridging PyTorch Export to Edge GPU Backends: Enabling High-Performance grid_sample with ExecuTorch,
Community Expo Baris Demir
Inference
Bringing High-Performance Inference Serving to Trainium with an Out-of-Tree vLLM Plugin
Community Expo Raj Thakur, Finn Thompson
Inference
Bringing SmolLM2-135M to Resource-Constrained Arm Edge Devices with ExecuTorch
Community Expo Xingguo Li
Inference
Bringing vime RL Post-Training to ROCm: End-to-End Training and Rollout on AMD Instinct
Community Expo Joy Song
Training
Building On-Device AI Agents on Exynos with ExecuTorch
Community Expo Jiseong Oh, Hoon Choi
Applications
Composable Kernel HSTU Attention Kernel for Generative Recommenders
Community Expo Chunxing Yin, Hao Wu
Kernel Engineering
Connectionless Scale-Out EP: GPU-Initiated MoE Dispatch/Combine for vLLM
Community Expo Naveen Ravi, Nathan Wichmann
Inference
CUDA Graph on Large Scale Recommender Systems
Community Expo Shuai Yang
Training
cuxray: Optimize Kernels Without a GPU
Community Expo Kareem Khaled Fareed
Kernel Engineering
Day-0 RL at Scale for Emerging LLM and VLM Architectures
Community Expo Shuang Yu, Huiying Li
Training
Day-1 Cooperative Matrix Support in ExecuTorch: Accelerating On-Device Inference on Mobile GPUs
Community Expo Yanwen Xu
Kernel Engineering
Debuggable and Profileable Intel GPU Software Stack for AI Computing with PyTorch
Community Expo Jianyi Zhang, Eikan Wang
Introduction
Demystifying Nondeterministic PyTorch Model Edge Inference with the Google AI Edge Debugger
Community Expo Qidong Zhao, Xiaoming Hu
Inference
Don't Just Trust the Model, Test the Physics: Evaluating PyTorch Models with PhysicsNeMo-CFD
Community Expo Kaustubh Tangsali
Applications
Dynamo Snapshot: Autoscaling and Failure Recovery in Seconds Using GMS
Community Expo Schwinn Saereesitthipitak, Vikram Sharma Mailthody
Inference
Efficient Inference of Masked Diffusion Language Models on AMD XDNA2 NPU
Community Expo Ravi Gupta, Rishi Teja Madduri
Inference
Efficient, Large-Scale LoRA and Fine-Tuning for Diffusion Models in PyTorch
Community Expo Pranav Prashant Thombre US
Training
Elastic Endpoint P2P Transfers for Inference with NIXL
Community Expo Samuel Nordmann
Inference
Empirically-guided Tensor Placement in AutoParallel
Community Expo Frost Mitchell, Syed Shahbaaz Ahmed
Training
Enabling APC for hybrid models in vLLM
Community Expo Thomas Parnell
Inference
Encryption Hides the Data, Not the Agent: Trace-Privacy for PyTorch LLM Agents
Community Expo Rahul Vishwakarma, Shrey Modi
Responsible AI
ERRC: Entropy-Reinvested Residual Correction for Tensor-Parallel LLM Communication
Community Expo Abdulsalam Bande
Inference
Escape Hatches for torch.compile
Community Expo Yidi Wu
Core PyTorch
Evaluating TorchComms Adoption in vLLM on XPU
Community Expo Jerome Mitchell
Inference
EvoTuner: Evolutionary Search with a Learned Cost Model for Triton Kernel Autotuning
Community Expo Chun-Lin Huang, Wei-Shen Huang
Kernel Engineering
ExecuTorch WebGPU: PyTorch's First GPU-Accelerated Inference in the Browser
Community Expo Julian Ng-Thow-Hing
Inference
Exploiting CUDA Locality Domains in PyTorch on Blackwell and Beyond
Community Expo Matthias Jouanneaux
Core PyTorch
Extending OpenSHMEM for AI and MoE Communication Patterns
Community Expo Nathan Wichmann, Naveen Ravi
Core PyTorch
Fail Fast or Restart? Automatically Triaging Distributed PyTorch Training Failures
Community Expo Sreyashi Chatterjee
Training
Fenceline: a self-hosted prompt firewall for locally served open-weights models
Community Expo Krishna Patel, Rahul Vishwakarma
Responsible AI
Firetap: Model Internal Telemetry for Large-Scale ML Training
Community Expo Tanvi Gupta, Chris McGillicuddy
Training
Five Bits on Metal: Adding Q5_K Quantization to ExecuTorch's MLX Backend
Community Expo Suvrakamal Das, Shivay Lamba
Core PyTorch
FlagQuantum: A PyTorch-Native Runtime for Differentiable and Distributed Quantum Computing
Community Expo Wei Liu
Core PyTorch
FlagRelease: A Self-Verifying Agent Platform for Model Migration Across Heterogeneous Accelerators
Community Expo Venly wu, Kevin Zhao
Applications
FlexBudget: Interchangeable KV Importance Signals under One Compiled FlexAttention Budget
Community Expo Huanran Wang
Inference
Flint: Compiler Enabled Cluster-Free Design Space Exploration for Distributed ML
Community Expo Jinsun Yoo
Core PyTorch
FP8-RL: A Practical and Stable Low-Precision Stack for LLM Reinforcement Learning
Community Expo Zhaopeng Qiu, Jingqi Zhang
Training
From Bottlenecks to Autonomy: Profiling RLHF Pipelines with PyTorch TorchTItan
Community Expo Manpreet Sokhi
Core PyTorch
From Client Communications to Competing Risks: A PyTorch Transformer for Client Attrition Model
Community Expo Witold Czubala
Applications
From Model Integrity to Runtime Integrity: Trusted MCP Agents on ARM Edge Devices
Community Expo Parichay Das
Responsible AI
From MP4 to Tensor: Optimizing Multi-Camera Video Data Loading for PyTorch Training
Community Expo Kyle Huang, Mihai Aldén
Training
From PT2E to Arm/TOSA: Dynamic W8A8 Linear Quantization in ExecuTorch
Community Expo Elena Zhelezina
Inference
From Scheduled to Real Overlap: All-Gather/Reduce-Scatter in PyTorch Distributed
Community Expo Jeffrey Mahou
Training
Graphing the Ungraphable: Making MoE Training CUDA-Graphable in PyTorch
Community Expo Syed Ahmed, Vasudevan Rengasamy
Training
Grouped GEMM on AMD ROCm in PyTorch
Community Expo Jagadish Krishnamoorthy
Kernel Engineering
Hardware-Accelerated Strided KV-Cache Transfers for Asymmetric LLM Inference
Community Expo Michal Shalev
Inference
High-Performance CPU Inference in PyTorch: TorchInductor, oneDNN, and vLLM Ecosystem Integration
Community Expo Aditya Tewari
Core PyTorch
Importance Aware Attention for Faster PyTorch Inference on Arm CPUs
Community Expo Radu Salavat, Nikhil Gupta
Inference
Improving PyTorch Efficiency on ARM with SVE-Accelerated Memory Primitives
Community Expo N Maajid Khan, Ashish Chopra
Inference
Improving the torch.compile user experience
Community Expo Rob Timpe
Introduction
Inductor-TV: Formal Methods for the Pytorch Compiler
Community Expo Abhilash Majumder
Responsible AI
Inference Performance Analysis using TraceLens
Community Expo Deval Shah
Inference
Integrating Out-of-Tree Accelerators with PyTorch Profiler
Community Expo Shreya Srivastava, Lucas Hendren
Applications
Integrating RCCL GPU-Initiated Networking into TorchTitan’s MoE Communication
Community Expo Rishi Sinha
Training
KernelFoundry: Hardware-Aware Evolutionary GPU Kernel Optimization
Community Expo Benjamin Ummenhofer, Sameer Sheorey
Kernel Engineering
Layout-Aware Compilation Cache Validation for Custom AI Accelerators in PyTorch
Community Expo KC Hema Prasanna, Pradipta Ghosh
Core PyTorch
Maestro: Benchmarking Concurrent GPU Operations as They Run in Real Workloads
Community Expo Eyal Chocron
Applications
MAKCI: A Multi-Abstraction AI Kernel Cost Instrumentor for Performance Modeling in AI Compiler
Community Expo Shubham Jain, Ritik Raj
Kernel Engineering
Memory Management in CUDA Graphs
Community Expo Frank Lin
Core PyTorch
Moral Tensors and DecisionProofs: Compiling Language into an Auditable, Grounded Safety Layer
Community Expo Andrew Bond
Responsible AI
More Than Meets the Eyes: Mirroring CPython's Object Protocol in Dynamo
Community Expo Guilherme Leobas
Core PyTorch
MoRI + vLLM: Wide Expert Parallelism and RDMA KV-Cache Transfer for Disaggregated MoE Serving on AMD
Community Expo Shiksha Patel, Rishi Teja Madduri
Inference
Multi-Chip KV-Cache Transfer for Disaggregated vLLM Serving with FlagCX
Community Expo Tao Chang
Inference
Multi-Teacher On-Policy Distillation for SOTA Agentic AI models
Community Expo Wenwen Gao
Training
Network fault tolerance for AI workloads
Community Expo Evgeny Leksikov
Inference
New Activation Memory APIs in PyTorch
Community Expo Jeffrey Wan
Core PyTorch
Offloading KV Cache to Secondary Memory: A Three-Tier Hierarchy Emulator using vLLM
Community Expo Mengmei Ye, Chen Wang
Inference
One Policy, Two Engines: Closing Training–Rollout Drift in DeepSeek-V4 Flash RL
Community Expo Xinyu Kang, Yuankai Chen
Training
Optimizing Streaming ASR on Arm Edge Devices with ExecuTorch
Community Expo Josiah Davis, Kshitij Sisodia
Inference
OSU HPC-AI: Distributed Pre-training, Reinforcement Learning, and Inference with MPI Communication
Community Expo Quentin Anthony, Dhabaleswar K (DK) Panda
Core PyTorch
Perch 2.0 and Perch-Hoplite in Pure PyTorch: An End-to-End TensorFlow-Free Bioacoustics Pipeline
Community Expo Duane Edgington
Applications
Private Pool Memory Visualization for CUDA Graphs
Community Expo Shangdi Yu, Boyuan Feng
Core PyTorch
Production FP8 Training for Multimodal Transformers in PyTorch
Community Expo Wei Chen US, Iain Weissburg
Training
Quantize-Then-Refine: Two-Stage Scoring for Memory-Efficient GPU Retrieval
Community Expo Ronak Kaoshik, Pratik Dixit
Applications
Race-Free TorchInductor Artifact Reuse for Elastic Training
Community Expo Christine Cheng, Kwanghoon An
Training
Real Shapes, Real Bugs: Model-Derived Operator Tests for Quick Out-of-Tree Accelerator Bring-Up
Community Expo Kazuaki Ishizaki
Inference
Rebuilding the Build: PyTorch's Move to a Standards-Based Packaging Backend
Community Expo Klaus Zimmermann
Core PyTorch
Red-Teaming Self-Hosted Agents: What Each Injection Defence Stops, Breaks, and Costs
Community Expo Manad Desai
Responsible AI
Regional Inductor: Selective Inductor Compilation within torch.compile
Community Expo Shangdi Yu
Core PyTorch
Regional-AOTI: A PyTorch-Native Partial Compilation Framework For High-Performance Inference
Community Expo Tzu-Hsin Yang, Yidi Wu
Inference
Reinforcement Learning for LLMs with Intel XPU and PyTorch on ALCF Aurora
Community Expo Filippo Simini, Guoqiong Song
Training
Resilient PyTorch Distributed with NCCL Shrink and Revoke
Community Expo Bruce Chang
Core PyTorch
Resolving Numerical Differences in Attention for RL
Community Expo Angel Li
Core PyTorch
Resource-Efficient Secure and Privacy-Preserving LLM Inference via GPU Virtualization
Community Expo Gilliean Lee, Chun Tao
Responsible AI
Retrieval as Inference: Building a Production-Scale Torch-Native Retrieval Engine
Community Expo Dhritiman Das, Vishal Shah
Inference
Retrieval-Augmented Autotuning: Navigating the Quality–Readiness Frontier with Historical Evidence
Community Expo Ishan Aryendu, Angela Yi
Kernel Engineering
Safe Checkpoint Provenance using Sigstore + Safetensors
Community Expo Harshita Varma, Nikita Verma
Responsible AI
Scaling Large-scale Distributed Training with DeepSpeed using Muon Optimizer
Community Expo Zhipeng Wang
Training
Scaling LLM Inference with Megakernels
Community Expo Krutarth Bhatt, Bhavya Shah
Inference
Scaling Multi-Model PyTorch Workloads for Closed-Loop Robotics Evaluation
Community Expo Nan Zhu
Applications
Scaling PyTorch Knowledge Through Community: Lessons from PyTorch Korea
Community Expo Jiho Kim, Junghwan Park
Introduction
Scaling PyTorch Validation for a Multi-Accelerator Future
Community Expo Riya Punia, Mansi Agarwal
Core PyTorch
Serving Interpretability Probes on Open Models
Community Expo Aishwarya Ramasethu
Responsible AI
Serving Video Diffusion on Trainium by fusing Communication, Attention, and Quantization
Community Expo Aneesh Shetty, Yide Zou
Inference
SGLang-Plugin-FL: A Unified Multi-Chip Backend for SGLang Inference
Community Expo John Zhou
Inference
Shape-Stable Dynamic Control Flow in PyTorch CUDA Graphs
Community Expo Thomas Ortner, Daniel Galvez
Core PyTorch
StageFrontier: Always-On Stage Accounting for Distributed PyTorch Training
Community Expo Boram Yoon, Aaron Meng
Training
Stage-Specialized E-PD Disaggregation for SLO-Aware Multimodal Serving on Intel Arc Pro B50/B70
Community Expo Rahul Unnikrishnan Nair, MinSung Kim
Inference
Support Helion Flow to MLIR Linalg for RISC-V Vector Extension
Community Expo Tai-Hsiang Peng, Chao-Lin Lee
Kernel Engineering
TabTune: One Interface for the Tabular Foundation Model Lifecycle
Community Expo Aditya Tanna, Vinay Kumar Sankarapu
Applications
TCCL: Low-latency Communication Backend for Apple Silicon
Community Expo Radim Urban
Core PyTorch
Tensor KV Cache: A PyTorch-Native Approach to KV Cache Offload
Community Expo Qi Bao, Ugur Kaynar
Inference
Testing the Edge: Scalable Validation for ExecuTorch
Community Expo Riya Punia, Arkadip Maitra
Inference
The Graph Break You Didn't Expect: Making torch.compile and FSDP2 Work Together
Community Expo Mansi Agarwal
Core PyTorch
The Power of Graph-level Optimizations for LLM Inference on Edge CPUs: Efficient SDPA with YNNPACK
Community Expo Volodymyr Kysenko
Inference
TorchInductor for SGLang Inference on RISC-V CPUs
Community Expo Jui-Ting Chen, Jenq-Kuen Lee
Kernel Engineering
TorchTitan GraphTrainer: LLM Distributed Training with PyTorch’s Compiler Toolkit
Community Expo Sherlock Huang
Training
TorchTitan on Intel XPU: From Desktop to Supercomputer – A Guide to Large-Scale LLM Training
Community Expo Subrata Goswami
Training
TorchTitan Performance Optimizations
Community Expo Elfie Guo Guo, Syed Ahmed
Training
TorchTPU x Hugging Face: Native TPU Acceleration Across the HF Ecosystem
Community Expo Alvaro Moran, Jingya Huang
Applications
Train Across Your Entire GPU Fleet: Heterogeneous Mixed-Generation Training in Pure PyTorch
Community Expo Roy Allela, Aravind Neelakantan
Training
Train More with Less: Fewer GPUs, Less Data Movement, Better Cluster Utilization
Community Expo Hyungyo Kim, Apoorve Mohan
Training
Trident: Compiling Triton Kernel Launch Paths from Python to LLVM via MLIR-Based JIT Compilation
Community Expo Jinjie Liu, Chunlei Men
Inference
Triton-Based Operator and Compiler Optimizations for LLM Inference Across AI Accelerators
Community Expo Hang Xiao
Inference
Upgrading CPU Preprocessing Modules from TorchScript to PT2: IR Design, Developer Experience, and Pr
Community Expo Georgia Phillips, Quanqi Hu
Inference
Using HugePages with PyTorch for Faster LLM Inference
Community Expo Moein Khazraee
Inference
vLLM IR: Graph Compiler Intermediate Representation for Fast LLM Inference
Community Expo Luka Govedič
Inference
What Privacy Actually Costs: An Accuracy-vs-Epsilon Curve for LoRA Fine-Tuning
Community Expo Shrinath Thube, Vaibhav Tupe
Responsible AI
What Should SRE Watch? Calibration-Gated Decisions and the Missing Golden Signals for GPU Clusters
Community Expo Andrew Espira
Responsible AI
When Does Offloading FSDP's Reduce-Scatter to a DPU Actually Pay Off?
Community Expo Venkatesh Bellale
Training
When Should I Disaggregate? Portable Diagnostics for LLM Inference on vLLM and llm-d
Community Expo Sean McGovern, Arkadip Maitra
Inference
Where the Blocks Live: Tier-Aware KV-Cache Routing for vLLM at Scale
Community Expo Mengmei Ye, Bongwoo Bak
Inference
Your Agent Passed the Benchmark. Then Production Broke It.
Community Expo Shikhar Mathur
Applications
Your GPU Is Starving: The Hidden Cost of Loading LLMs at the Edge
Community Expo Rini Susan V S
Inference
Zero-Code Model Portability: Running HuggingFace Transformers & Diffusers on Any Chip via Torch-FL
Community Expo Yufeng Lyu
Applications
6:10 PM
TPU Model Performance Auto-optimization
Community Expo Aleksey Vlasenko
Training
6:30 PM
Training Embedding Models Resiliently for Multimodal Model Inference Routing
Community Expo Haichen Zhang, Huamin Chen
Training

Sponsors

DIAMOND

PLATINUM

GOLD

SILVER

BRONZE

Startup + Non-Profit + VC