top of page

W3

Hardware-software co-design for accelerators for foundation models
and event-based neural processing

ROOM Granados II (9:00 - 11:00) and ROOM Victoria (14:00 - 17:45)

CHAIRS

Kavya Sreedhar (Google, US)

Thierry Tambe (Stanford University, US)

Melika Payvand (University of Zurich and ETH Zurich, CH)

Jae-sun Seo (Cornell Tech, US)

Matheus Moreira (Meta, US)

Francesco Conti (University of Bologna, IT)

ABSTRACT

This workshop brings together the communities of accelerators for foundation models and event-based neural processing under a unifying theme: hardware–software co-design for efficient intelligence at scale. As foundation models continue to grow in size and capability, their deployment increasingly depends on architectural specialization, memory hierarchy redesign, and co-optimized software stacks. In parallel, neuromorphic and event-based systems offer radically different computational primitives which are sparse, data-driven, and temporally structured, challenging the conventional boundaries of AI hardware. By bridging these two worlds, the workshop aims to identify shared bottlenecks, complementary design philosophies, and emerging opportunities in memory technologies, model representation, learning modes, and cross-layer optimization. We invite experts from both domains to cross-polinate ideas across the two sister communities.

PROGRAM

9:00 - 9:10

Introduction

Neuromorphic Cross-stack Co-design for Efficient Intelligent Computing 

ROOM Granados II

session chairs

Melika Payvand (University of Zurich and ETH Zurich, CH)

Jae-sun Seo (Cornell Tech, US)

 

9:10-09:37

From Attention to KANs: Energy-Efficient In-Memory Computing for Next-Gen AI

John Paul Strachan (RWTH Aachen University, DE)

​This talk explores how in-memory computing (IMC) architectures can remove key performance and energy bottlenecks in modern deep learning—especially in transformer-based large language models (LLMs) and emerging architectures for scientific computing. First, we introduce an analog IMC self-attention mechanism built on charge-based “gain-cell” memories that store token projections in-place and execute dot-products in parallel within memory, substantially reducing the data-movement overhead inherent to attention with KV caching. However, analog gain-cell circuits impose non-idealities preventing direct use of pre-trained models, and thus we developed a light-weight initialization strategy that recovers GPT-2-level performance without training from scratch.

We then broaden to work addressing other major bottlenecks—including softmax operations, where precision demands and computational structure become increasingly problematic as sequence lengths scale. We present an IMC-optimized reformulation that shifts complexity away from sequence length, fuses softmax sub-operations with adjacent VMMs, and supports both efficient digital and spike-encoded analog implementations with negligible accuracy loss. Finally, we extend IMC beyond transformers to Kolmogorov–Arnold Networks (KANs), which trade linear layers for trainable nonlinear functions and can substantially reduce model size for scientific workloads. We present a flexible, cross-layer-optimized IMC accelerator that computes arbitrary nonlinear functions efficiently using read-optimized memristive arrays and a single-read evaluation scheme, delivering large improvements in energy and energy–delay compared to CPUs and prior accelerators.

 

​09:38 - 10:05

In-Memory Computing Architectures for AI at the Edge

Irem Boybat (IBM Zurich, CH)

The growing computational demands of AI inference are pushing edge platforms beyond the limits of conventional architectures, where data movement between memory and processing units dominates energy consumption. In-memory computing (IMC) offers a compelling path forward by performing computation directly within memory arrays, enabling significant gains in energy efficiency.

This talk presents a cross-stack perspective on heterogeneous IMC-based architectures for edge AI acceleration. At the system level, we present an architecture integrating IMC arrays with programmable multi-core accelerators, including ISA extensions for efficient attention execution. At the tile level, we discuss area-efficient periphery designs for distributed IMC architectures, addressing quantization-aware execution, partial-sum accumulation, and lightweight activation support. Spanning both levels, we present mapping and scheduling strategies that co-optimize workload placement across heterogeneous compute resources. Together, these contributions outline a path toward efficient and flexible deployment of AI inference at the edge.

10:06 - 10:33

Cross-stack optimization through a brain-inspiration lens

Melika Payvand (University of Zurich and ETH Zurich, CH )

Embedded and edge systems must interpret continuous sensory streams under strict power and memory limits. I will present how brain-inspired computational principles, such as event-driven communication, localized connectivity, and multi-timescale dynamics, enable efficient temporal processing on constrained hardware. Through examples from neuromorphic architectures and hardware–algorithm co-design, I will show how NeuroAI can guide the development of scalable, low-power computing substrates for real-time intelligence.

10:33-11:00

Nonlinear In-memory Computing for SNNs at the Edge

Arindam Basu (City University of Hong Kong, HK)

With the rapid increase in parameters and concomitant energy cost of modern Deep Neural Networks (DNN), it has become imperative to explore more energy-efficient architectures, especially for deployment on energy and area constrained edge devices such as smart glasses, hearables or medical implantable devices for brain interfaces. In this context, brain-inspired spiking neural networks (SNN) provide a fascinating new option with low-energy operation due to sparse binary activations, that reduce communication overhead and replace synaptic multiplications with additions.

Similar to ANN counterparts, In-memory computing (IMC) can be used in SNNs as well to remove the memory access bottlenecks in terms of throughput and energy. However, with synaptic operations becoming more efficient, the nonlinear operations such as dendritic computing and winner take all constitute a new “nonlinearity wall”. In this talk, I will show our recent work in embedding such scalar and vector nonlinear functions efficiently within memory leveraging nonlinear in-memory analog-digital converters. The architectures are general and have been demonstrated in both SRAM and RRAM substrates. Further, I will demonstrate how co-designing the algorithm and hardware lead to immense gains in energy and throughput with little sacrifice in task accuracy. I will also show the possible applications of these methods in improving the implementation of conventional ANN architectures like Recurrent neural networks and Transformers.

We envision the embedding of nonlinear operations within memory as an exciting new frontier for IMC research with potential to impact both SNN and ANN architectures.

11:00 - 11:30

Coffee break

11:30 - 13:00

(no session) 

13:00 - 14:00

Lunch

Hardware/software co-design for accelerators for foundation models

ROOM Victoria 

session chairs

Kavya Sreedhar (Google, US)

Thierry Tambe (Stanford University, US)

 

 

14:00 - 14:35

TPU Hardware/Software Co-Design at Google

Kavya Sreedhar (Google, US)

AI is growing at an unprecedented scale, with the compute required to train AI models doubling more often than every five months. Given the scale of this problem, and that Moore’s Law is slowing down, there is a need to explore cross-stack optimization opportunities to design efficient AI chips. This talk will overview hardware/software co-design considerations at Google when designing Tensor Processing Units (TPUs) for AI applications.

 

14:35 - 15:10

A Structured Approach to the Hardware/Software Co-Design of an AI-Accelerator

Joel Emer (MIT / NVIDIA, US)

Over the past few years, efforts to address the challenges of the end of Moore's Law has led to significant rise in AI-focused accelerators. However, the well-understood pipeline stages used in the past for designing and optimizing general-purpose processors does not apply to these accelerators. As a result,there has not been a systematic way to express the range of design features used by these accelerators. This makes it difficult to understand the impact of each design choice and compare or extend the state-of-the-art.

In this talk, I will present a novel separation of concerns taxonomy to better understand and design AI-accelerators. One aspect of this approach is the use tensor algebraic notation to describe and optimize the computation. Second, in an analogous fashion to our prior work that categorized DNN dataflows into patterns like weight stationary and output stationary, this talk will try to provide a set of options that characterize the other attributes of tensor accelerators. Thus, rather than just presenting a single specific design, I will present a generalized framework for describing computations, dataflows and the placement of activity. In this framework, this separation of concerns is intended to better understand design choices and facilitate the exploration of the wide design space of tensor accelerators. Thus, using this common language, I will show how the choices invoked were used to generate our high-performance transformer accelerator.

15:30 – 16:00

Coffee break

16:00 - 16:35

Leveraging Natural Human Behavior for Efficient and Intelligent AR and VR Systems

Sai Zhang (NYU, US)

​Augmented and virtual reality (AR/VR) systems are emerging as a critical computing platform in modern life, with increasing impact across fields like education, healthcare and industrial applications. Despite their growing importance, AR and VR devices operate under strict constraints on latency, energy consumption, and computational resources, making efficient system design a fundamental challenge. A defining characteristic that distinguishes AR and VR from conventional edge devices is their direct and continuous interface with the human user, where perception, attention, and intention fundamentally shape system behavior. By leveraging natural human behavior such as gaze, head motion, and hand interaction as first class signals, AR and VR systems can adapt their computation to what truly matters to the user, enabling selective processing and more efficient use of limited resources.

 

16:35 - 17:10

Algorithm-Hardware Co-Design of Retention-Aware Differentiated Memory Architectures

Thierry Tambe (Stanford University, US)

Modern AI systems are increasingly limited not by arithmetic, but by memory. As frontier AI models become more capable, they require far more data to be moved, stored, and accessed efficiently. These workloads systematically generate large volumes of short-lived data that are written in memory, consumed, and quickly discarded, as well as long-lived data that must be retained reliably across much longer time scales. Conventional memory systems are poorly optimized to this behavior, as they are typically designed as one-size-fits-all storage, resulting in excessive energy consumption and increasingly limited density scaling. We propose to address this mismatch by developing a retention-aware computing stack that treats data persistence as a central design consideration by statically and dynamically matching short-lived and long-lived data to differentiated memory architectures and technologies, each optimized for the appropriate retention window. The result is a more efficient and sustainable foundation for accelerating large-scale AI systems in both datacenter and edge settings.

17:10 - 17:45

Disruptive Architectures via Diverse, Dense, and Deeply 3D Integrated Memory + Logic

Robert Radway (UPenn, US)

​Pervasive AI/ML demands ever-higher on-chip memory + logic capacity to minimize costly on- and off-chip data movement. Three techniques pave the way for the next generations of future system scaling: 1) Ultra-dense 3D integration to combine heterogenous memory + logic technologies (e.g., SRAM, ReRAM, Gain Cells, and beyond), matching distinct device characteristics to application needs for reliable operation, 2) New architectural design points leveraging dataflows unlocked by ultra-dense 3D memory-logic connectivity, 3) Illusion for energy-efficient multi-chiplet systems via new fine-grained mapping, scheduling, and power management techniques.

BIOSKETCHES

John Paul Strachan currently directs the Peter Grünberg Institute on Neuromorphic Compute Nodes at Forschungszentrum Jülich and is W3 Professor at RWTH Aachen.  John Paul has degrees in electrical engineering and physics from MIT, and received the PhD from Stanford University. Previously he worked in industry for 13 years, leading the Emerging Accelerators team as a Distinguished Technologist at Hewlett Packard Enterprise (previously HP). His teams develop circuit IP, prototype ICs, and explore the co-design of special purpose architectures with algorithms. Interests span applications in AI, optimization, genomics, and security. He has over 60 patents, has authored or co-authored over 110 peer-reviewed papers. He developed sensing systems for precision agriculture in a company he co-founded, and was awarded the Falicov Award for studies of magnetic memory technologies.

 

Irem Boybat is a Staff Research Scientist at IBM Research Europe, Zurich, Switzerland. She received her Ph.D. degree in Electrical Engineering from Ecole Polytechnique Federale de Lausanne (EPFL), Switzerland, in 2020, her M.Sc. degree in Electrical Engineering from EPFL, Switzerland, in 2015, and her B.Sc. degree in Electronics Engineering from Sabanci University, Turkey, in 2013. Her research focuses on in-memory computing-based architectures for low-power edge and data center environments for AI acceleration, application-hardware co-design, model optimization, and deployment strategies. She has co-authored over 65 scientific papers in journals and conferences, received five best conference presentation/paper/poster awards and holds 8 granted patents. She was a co-recipient of the 2018 IBM Pat Goldberg Memorial Best Paper Award and 2020 EPFL PhD Thesis Distinction in Electrical Engineering.

 

Melika Payvand is an Assistant Professor at the Institute of Neuroinformatics, University of Zurich and ETH Zurich and leads the Emerging Intelligent Substrates lab. She received her PhD in Electrical and Computer Engineering at the University of California Santa Barbara. Her research interest is in developing intelligent learning systems on physical substrates, inspired by the hierarchical structure-function correlate in the biological brain. In 2023, she received the prestigious Swiss National Science Foundation Starting Grant. 

She is an active member of the Neuromorphic Engineering community who has co-coordinated the European project NEUROTECH (neurotechai.eu), served as the co-chair of the International Conference on Neuromorphic Systems (ICONS) (https://iconsneuromorphic.cc) for several years, and has co-organized the scientific program of the Capocaccia Neuromorphic Intelligence workshop (https://capocaccia.cc) from 2019-2023. She also serves in the Technical Committee of the IEEE Circuits and Systems and European Solid State Circuits societies.

Arindam Basu received the B.Tech and M.Tech degrees in Electronics and Electrical Communication Engineering from the Indian Institute of Technology, Kharagpur in 2005, the M.S. degree in Mathematics and PhD. degree in Electrical Engineering from the Georgia Institute of Technology, Atlanta in 2009 and 2010 respectively. Dr. Basu received the Prime Minister of India Gold Medal in 2005 from I.I.T Kharagpur. He is a Professor in City University of Hong Kong in the Department of Electrical Engineering and is a Fellow of IEEE, AAIS and HKIE. He has served as IEEE CAS Distinguished Lecturer in the past and holds several Editorial roles in leading journals.  Dr. Basu received the best student paper award at Ultrasonics symposium, 2006, best live demonstration at ISCAS 2010, a finalist position in the best student paper contest at ISCAS 2008 and the Industry Choice Award in BioCAS 2015. He was awarded MIT Technology Review's TR35 Asia Pacific award in 2012 and inducted into Georgia Tech Alumni Association's 40 under 40 class of 2022.

​​​

Kavya Sreedhar is a Senior Research Scientist in the TPU HW/SW Co-design team at Google. At Stanford, she received her PhD in electrical engineering, advised by Mark Horowitz, in 2025, and her MS in electrical engineering in 2020. She received a BS in electrical engineering and a BS in business, economics, and management from Caltech in 2019. Her graduate school research was supported by Stanford's Knight-Hennessy Graduate Fellowship and the Quad Fellowship. She previously interned with Meta Reality Labs, NVIDIA's Architecture Research Group, Apple, Microsoft, and Intel. Her research interests are broadly in hardware design and performance modeling for machine learning and cryptography applications.

​​

Joel S. Emer is a Professor of Electrical Engineering and Computer Science at MIT. He also is a Senior Distinguished Research Scientist at Nvidia. Prior to joining Nvidia, he worked at Intel where he was an Intel Fellow and Director of Microarchitecture Research. Previously he worked at Compaq and Digital Equipment Corporation (DEC). His current research focus is on the systematic design and modeling of accelerators. In the past, he has made architectural contributions to a number of VAX, Alpha and X86 processors and is recognized as one of the developers of the widely employed quantitative approach to processor performance evaluation. He has also been recognized for his contributions in the advancement of deep learning accelerator design, spatial and parallel architectures, processor reliability analysis, memory dependence prediction, pipeline and cache organization, performance modeling methodologies and simultaneous multithreading. He earned a doctorate in electrical engineering from the University of Illinois in 1979. He received a bachelor's degree with highest honors in electrical engineering in 1974, and his master's degree in 1975 -- both from Purdue University. Among his honors, he is a Fellow of both the ACM and IEEE, and a member of the NAE. He also received both the Eckert-Mauchly award and the B. Ramakrishan Rau award for lifetime Contributions in computer architecture.

 

Sai Qian Zhang is an Assistant Professor of Electrical Engineering and Computer Science at New York University (NYU). Prior to joining NYU, he spent two years at Reality Labs at Meta, working on smart sensor design for next-generation AR/VR device. Sai earned his Ph.D. in Computer Science from Harvard University in 2021 and holds both M.A.Sc and B.A.Sc degrees in Electrical Engineering from the University of Toronto. His research interests include AR/VR computing, machine learning algorithms and system codesign. His work has been published in top tier conferences including ASPLOS, ISCA, MICRO, HPCA.

Thierry Tambe is an Assistant Professor of Electrical Engineering and, by courtesy, of Computer Science at Stanford University. His research centers on co-designing algorithms and hardware—from high-level models down to custom silicon—to enable efficient execution of AI and data-intensive workloads, with memory efficiency as a central theme. In simple terms, his lab studies how to represent data more compactly, move it less, and place it intelligently across differentiated memory systems. Previously, Thierry was a visiting research scientist at NVIDIA and a senior engineer at Intel. He received his PhD in Electrical Engineering from Harvard University.

 

Robert M. Radway is an Assistant Professor of Electrical and Systems Engineering at the University of Pennsylvania.  He received his Ph.D. in Electrical Engineering from Stanford University and M.Eng. and B.S. in Electrical Engineering and Computer Science from MIT. His research focuses on building hardware systems leveraging heterogeneous memory, logic, and 3D integration for large benefits. His research contributions include the first published non-volatile Resistive RAM (RRAM) systems, the first edge AI/ML chips using foundry RRAM for full on-chip inference and training, multi-chiplet systems to efficiently scaleup to larger models, and the first foundry monolithic 3D hardware delivering large benefits. His honors include the Stanford Graduate Fellowship, and his papers have received the Symposium on VLSI Circuits, the Symposium on VLSI Technology, and the IEDM Roger A. Haken Best Student Paper Awards.

bottom of page