top of page

PARALLEL SESSION

C1L-E

LLMs, Transformers, and Generative AI Accelerators Session

SALA TURINA

THURSDAY- 10 September | 10:40 - 12:20

CHAIRS

Jae-Sun Seo, Guilherme Paim

SESSION PROGRAM

10:40 - 11:00

DiFlash: A 13.70TOPS/W 156.05Pixels/mJ Video Generation Accelerator with Bubble-Free Online-Quantized FlashAttention and Query-Reuse Sliding-Tile Attention

Jun Liu, Shiwei Liu, Li Ding, Shuaiheng Li, Tianlang Zhao, Yuxuan Zhang, Xinhao Li, Shan Huang, Jinhao Li, Qi Luo, Ya...

11:00 - 11:20

A 28nm 97.23 TFLOPS/W Pipelined Two-Stage-Aggregation Compute-in-Memory Macro for Edge-LLM Inference and Fine-Tuning

Yi Yang, Jinwu Chen, Yutong Zhang, Yucheng Du, Defa Wu, Yingxuan Zhou, Yanqi Zhang, Tianhui Jiao, Xiaoxue Zhong, Xing...

11:20 - 11:40

DRAGON: A 28nm 5.12Ms-Response and 42.83Token/S Digital CiM-Based Multi-Phase Accelerator for Hybrid-Intensive Retrieval-Augmented Generation

Zihan Wu, Yiqi Wang, Zhen He, Huiming Han, Yang Hu, Fengbin Tu, Shouyi Yin

11:40 - 12:00

BumbleBee: A 3 mW Fused-Kernel Flash Attention Edge AI Co-Processor in 18 nm FD-SOI CMOS

Joaquin A. Cornejo, Filipe Pouget, Amelie Poullot, Raphael Gras, Dominique Bousquet, Nicolas Lebouleux, Tarun Chawla,...

12:00 - 12:20

SSTP: A 13.32 TFLOPS/W End-to-End Simultaneous Speech Translation Processor on Edge Devices

Bokyoung Seo, Ghangmin Yun, Chaeyoon Kim, Jueun Jung, Kyuho Lee

bottom of page