PyTorch Tensors and neural networks in Python with strong hardware acceleration.

Hardware heterogeneity is no longer an edge case - it is the default.Join us at PyTorch Conference North America 2026 (O...
09/05/2026

Hardware heterogeneity is no longer an edge case - it is the default.

Join us at PyTorch Conference North America 2026 (Oct 20–21, San Jose) for deep dives into kernel engineering, custom compiler backends, and performance optimization across GPUs, TPUs, NPUs, and custom ASICs.

Featured speakers include:

AWS / AWS Annapurna Labs / Annapurna Labs: Maen Suleiman, Yahav Biran, Esha Lakhotia
Annapurna ML: Pinak Panigrahi
Arm: Kavya Sri Chennoju, Thomas Cottenier
Google: Sandeep Pokkunuri, Chris Jones, Rob Mulla, Claudio Basile
Huawei: Jiahao Chen, Jiahao Tan
Hugging Face: Michael Benayoun
IBM/ IBM Research: Takeshi Yoshimura, Olivier Tardieu, Matthew Arnold, Thomas Parnell, Prasanth Chatarasi, Bardia Mahjour, Viji Srinivasan, David Grove, Antoni Viros Martin, Avery Blanchard
Intel: Yu Guangye, Eikan Wang, Panagiotis Kourdis, Tanima Dey, Huma Abidi, Ashok Emani, Qiacheng Li
Meta: Richard Zou, Mergen Nachin, Digant Desai
Red Hat: Markell Rawls
Beijing Academy of Artificial Intelligence: Yonghua Lin
Fujitsu Research of India: Abhishek Jain, N Maajid Khan

Register here and stay for AGNTCon + MCPCon that same week: https://bit.ly/464FTbp

Read the Hardware Acceleration & Compute Infrastructure Sessions guide: https://bit.ly/4yoK5z2

09/05/2026

Ever wondered why type systems don't automatically track tensor shapes across your neural network modules?

In his upcoming talk at PyTorch Conference North America 2026 (October 20–21 in San Jose, CA), Avik Chaudhuri from Meta breaks down how PyreFly - a static type checker for Python - enables end-to-end static shape coverage for PyTorch models.

Register for PyTorchCon North America today: https://bit.ly/4qStJvL

From research prototypes to production serving across heterogeneous hardware, PyTorch Conference North America 2026 brin...
09/04/2026

From research prototypes to production serving across heterogeneous hardware, PyTorch Conference North America 2026 brings the core developer community together this October 20-21 in San Jose.

Sessions will cover native PyTorch hardware integrations, agentic workflows, dynamic shape dynamic compiler optimization, high-performance kernel design and more.

Featured speakers include:

AWS : Maen Suleiman
AMD : Liz Li, Xiaohu Guo, Shekhar Pandey
Google Cloud /Google: Bill Jia, Qi Zhou
Huawei: Jing Li, Qi Guo
Hugging Face: Sayak Paul
Intel: Whitney Tsang, Artur Fierka
Meta: Oguz Ulgen, Dunfan Lu, Jason Ansel, Elias Ellison, Driss Guessous, Colin Taylor, Angela Yi, Jane Xu, Simon Layton, Angel Li, Richard Zou, Yidi Wu, Laith Sakka
NVIDIA: Michael Goldfarb, Guray Ozen, Daniel Galvez
Rebellions: Minwook Ahn
Red Hat: Sean McGovern, Chris Leonard
Core Automation: Mark Saroufim
Crusoe AI: Suman Debnath, JanakiRam Goteti
d-Matrix: Tristan Webb
Lemurian Labs: Jay Dawani
Modular: Stefan Lindall

Today is the last chance to register for a standard ticket - join us for PyTorch Conference North America 2026 and stay for AGNTCon + MCPCon that same week: https://bit.ly/464FTbp

Read the Hardware Acceleration & Compute Infrastructure Sessions guide: https://bit.ly/4yoK5z2

09/04/2026

🚨 Final call, AI community!

Ticket prices for PyTorch Conference North America go up TONIGHT at 11:59 PM PDT, making this your final opportunity to save $200.

Join AI pioneers, researchers, developers, and open source innovators October 20-21 in San Jose, CA for two days of technical learning, big ideas, and community-powered progress.

Register before tonight’s deadline. Future you will call it an intelligent decision:
https://bit.ly/4sh3DSw

The PyTorch AOTI backend delivered a 1.14x–1.28x speedup over the Python backend in NVIDIA’s HSTU inference tests when d...
09/04/2026

The PyTorch AOTI backend delivered a 1.14x–1.28x speedup over the Python backend in NVIDIA’s HSTU inference tests when deployed with Triton Inference Server.

With the PyTorch AOTI backend and KV cache, NVIDIA reports a 2.20x–2.38x speedup in an ideal all-GPU cache-hit scenario.

The results come from NVIDIA’s recsys-examples repository, a collection of examples demonstrating best practices for training and deploying generative recommenders on NVIDIA GPUs using PyTorch. It includes optimized HSTU and Semantic ID implementations covering training and inference workflows.

For HSTU inference, recsys-examples supports PyTorch AOTInductor to execute the model in the Torch C++ runtime. The NVIDIA developer blog also covers nv-embedding-cache, which provides PyTorch-compatible modules for accelerating large-scale embedding lookups.

🔗 Read the full post:

Recommender systems (RecSys) are one of the most ubiquitous machine learning problems in the consumer internet industry yet notoriously difficult to train and serve at scale. The advent of LLMs has…

In PyTorch 2.14, fault tolerance becomes a first-class c10d concept, with in-place process-group reconfiguration.When a ...
09/03/2026

In PyTorch 2.14, fault tolerance becomes a first-class c10d concept, with in-place process-group reconfiguration.

When a rank fails in a large job, the usual recovery is to tear down the process group and restart, which discards warm state across the whole cluster. Backend and ProcessGroup now expose reconfiguration interfaces so a group can be rebuilt in place, with abort hooks and pre and post collective hooks wired through the same path.

Gloo gains fault-tolerance support alongside nccl2, and the reconfigure APIs are documented. The release blog marks the API as unstable.

🔗 Read the PyTorch 2.14 release blog: https://pytorch.org/blog/pytorch-2-14-release-blog/

09/03/2026

At PyTorch Conference North America 2026, hear directly from and connect with the people working on PyTorch, vLLM, DeepSpeed, Ray, Helion, and Safetensors, alongside experts working across the AI stack.

At NA 2025, Richard Liaw (Ray, Anyscale) spoke about the Ray team’s focus on building something “useful and critical to the AI infrastructure space” and about “how important Ray can be for the stack.”

PyTorch conferences are the open source AI community’s town square, where what’s next gets decided.

Register by September 4 (tomorrow) to save on your conference pass: https://hubs.la/Q04tDcHQ0

Discover how open source agentic search and hardware-guided workflows are unlocking massive speedups across GPUs and TPU...
09/03/2026

Discover how open source agentic search and hardware-guided workflows are unlocking massive speedups across GPUs and TPUs at PyTorch Conference North America (Oct 20–21 in San Jose).

Speakers will cover hardware-guided multi-agent kernel optimization and LLM autotuning to accelerate performance, multimodal world action models and ExecuTorch runtimes for physical AI, stateful guardrails and governance frameworks to ensure production accountability.

Featured speakers include:

AMD: Anshu Raina, Peyman Razaghi
Arm: Thomas Cottenier, Kavya Sri Chennoju
Amazon Web Services: Sai Charan Teja Gopaluni, Aaresh Sharma
Google: Sandeep Pokkunuri, Chris Jones
Intel: Xiaogang Gu, Qun Yang
Meta: Kaiming Cheng, Laura Wang, Jongsok Choi, Ethan Che, Oguz Ulgen, Dunfan Lu, Jason Ansel, Sanket Jayant Purandare, Aditya Venkataraman, Jacob Szwejbka, Mergen Nachin, Digant Desai
NVIDIA: Susie Xia, Ruijie Zheng, George Kurian
: Purva Chiniya
Apple: Prakshal Doshi, Aditi Mewada
Boson AI: Mu Li, Huapeng Zhou, Lindsey Allen
Capital One: Ravi Teja Prabhala Venkata
Rerun: Gijs de Jong, Peter Prettenhofer

Register for your standard ticket by September 4: https://bit.ly/464FTbp
Read the Agentic AI and Next-Gen Sessions guide: https://bit.ly/4xL2vK4

09/03/2026

Tomorrow is the final day to save $200 on registration for PyTorch Conference North America.

Join the researchers turning bold hypotheses into breakthroughs, the developers translating ideas into code, and the AI pioneers helping accelerate open intelligence.

It all happens October 20-21 in San Jose, CA.

Register by September 4 at 11:59 PM PDT before the price goes up:
https://bit.ly/4sh3DSw

Lablup Inc. has joined PyTorch Foundation as a Silver Member!The PyTorch Foundation is the vendor-neutral home for the o...
09/03/2026

Lablup Inc. has joined PyTorch Foundation as a Silver Member!

The PyTorch Foundation is the vendor-neutral home for the open source intelligence layer developers use for training, optimizing, serving, orchestrating, and running models on any chip in any cloud for any agent. Our membership reflects an alignment with these principles and engagement with the open source AI ecosystem.

Lablup builds the software that powers AI—from personal desktops to data centers. With Backend.AI and its ecosystem, we enable anyone to choose, connect, and operate AI models at the right cost and scale. We're working toward a world where intelligence flows like electricity, reaching wherever it's needed. Lablup is here to supply intelligence for humanity.

"Lablup has developed Backend.AI as an open source project since our founding. Today Backend.AI runs AI workloads every day across the globe, and PyTorch sits at the center of that work. We have long been beneficiaries of the PyTorch ecosystem; joining the PyTorch Foundation is our way of standing on the contributing side of that foundation. We have always believed that keeping the core software of AI open is what keeps AI accessible to everyone."

— Jeongkyu Shin, founder and CEO of Lablup

Address

San Francisco, CA

Alerts

Be the first to know and let us send you an email when PyTorch posts news and promotions. Your email address will not be used for any other purpose, and you can unsubscribe at any time.

Shortcuts

Share