Running large Mixture-of-Experts models across pods creates two major challenges: expert traffic and KV-cache transfer.
At #PyTorchCon, Ravi Gupta, Rishi Madduri, and Shiksha Patel (AMD) will present how MoRI + vLLM tackles both.

Running large Mixture-of-Experts models across pods creates two major challenges: expert traffic and KV-cache transfer.
At #PyTorchCon, Ravi Gupta, Rishi Madduri, and Shiksha Patel (AMD) will present how MoRI + vLLM tackles both.
High-Velocity GPU Kernel Authoring with CUTLASS Python
At PyTorch Conference North America 2026, Guray Ozen from NVIDIA will present recent improvements to CUTLASS for Python.
Join us on October 20-21 in San Jose: https://hubs.la/Q04v4SL60
Optimizing kernels can be time-consuming and require deep expert knowledge. KernelAgent is a multi-agent harness that can write and optimize kernels for you.
At #PyTorchCon, Kaiming Cheng (Meta) will present hardware-guided GPU kernel optimization.
Bridging the ExecuTorch Adoption Gap Through Developer Learning
Join Matt Cossins from Arm at PyTorch Conference North America 2026 in San Jose, CA for a Birds of a Feather discussion session on ExecuTorch.
Join us on October 20-21: https://hubs.la/Q04v4SL60
#PyTorchCon Day 0 event: PyTorch, Powering the Enterprise with Red Hat, NVIDIA, and IBM on October 19th will explore how PyTorch has begun delivering essential enterprise-level features and enhancements, specifically tailored to enterprise use cases: events.linuxfoundation.org/pytorch-conf...
Involved in the AI ecosystem? There's something for everyone at #PyTorchCon North America. @thesteve0.bsky.social of Red Hat will be giving a talk on the use of PyTorch and PyTorch-based libraries from the ecosystem to do remote sensing for satellite imagery. Join us: hubs.la/Q04v4SL60
Think DeepSpeed is only good for ZeRO and its usability, not training throughput?
At #PyTorchCon North America 2026, Masahiro Tanaka will cover tensor, sequence, and expert parallelism in DeepSpeed and walk through benchmarks.
Red Hat, NVIDIA AI, and IBM are hosting “PyTorch, Powering the Enterprise,” a PyTorch Conference Day 0 event on Oct. 19 covering PyTorch, vLLM, Ray, Helion, and llm-d: https://bit.ly/4znEDgl
Register for #PyTorchCon NA: https://bit.ly/3U6WZ69
We’re heading to #PyTorchCon NA 2026! 🚀
RX-M’s @ronaldpetty.bsky.social will be in San Jose, October 20–21, joining the open source AI community!
If you’re attending it, keep an eye out for Ron and come say hello.
pytorch.org/event/pytorc...
PyTorch Foundation Ambassador Abdulsalam Bande will present a poster on ERRC (Entropy-Reinvested Residual Correction for Tensor-Parallel LLM Communication) at PyTorch Conference North America 2026.
Join us in San Jose, Oct 20-21: https://hubs.la/Q04v88dJ0
Modern AI workloads rely on specialized GPU kernels, and no single backend is optimal across every operator.
At #PyTorchCon, Liz Li and Jiahui Cao (AMD) will present how FlyDSL integrates with TorchInductor while preserving the torch.compile experience. https://hubs.la/Q04v4SL60
In his talk at PyTorch Conference North America 2026, Suvaditya Mukherjee from Hugging Face will share guidance on how to use the PyTorch Profiler on your models to extract maximum performance efficiency.
Join us in San Jose, Oct 20-21: https://hubs.la/Q04v4SL60
How does the PyTorch Ambassador Program work, and how can you position yourself to become one?
At #PyTorchCon North America 2026, Sahdev Zala (IBM) will cover what ambassadors have built, challenges they’ve faced, and ways to engage more deeply.
vLLM is a high-throughput, memory-efficient inference and serving engine for LLMs. At #PyTorchCon North America, Simon Mo’s keynote and vLLM technical sessions cover KV cache management, disaggregated serving, attention, and more.
Communication overhead remains one of the largest bottlenecks in scaling distributed training.
Join us at PyTorch Conference North America in San Jose on October 20-21 to explore the future of distributed training infrastructure: https://hubs.la/Q04v4SL60
In her upcoming session on Hardware-Aware AI at #PyTorchCon NA 2026, Kavya Sri Chennoju from Arm will demonstrate how PyTorch, ExecuTorch, and Arm Device Connect work together to deploy AI agents safely onto edge hardware.
Join us in San Jose, October 20-21: https://hubs.la/Q04v4SL60
Debugging LLM training in production is notoriously challenging.
Join us at PyTorch Conference North America in San Jose, October 20 to 21, to learn how bitwise alignment delivers faster, more precise LLM training debugging: https://hubs.la/Q04v4SL60
At PyTorch Conference North America, Maajid Khan from Fujitsu Research India will explore efficient MoE LLM inference using vLLM and OpenVINO, sharing strategies for running these models effectively on ARM architectures.
Join us in San Jose, October 20-21: hubs.la/Q04v4SL60
At #PyTorchCon North America 2026, Liz Li (AMD) and Shekhar Pandey (AMD) will discuss scalable MXFP8 pretraining on MI355X GPUs, including TorchAO kernel optimization and TorchTitan integration for an end-to-end upstream PyTorch training stack.
At #PyTorchCon North America 2026, Alessandro Sangiorgi, Senior Software Engineer (Red Hat), will show how warm starts let Helion autotuning reuse cached configurations for faster iteration instead of repeated cold searches.