PARALLEL SESSION
C1L-E
LLMs, Transformers, and Generative AI Accelerators Session
SALA TURINA
THURSDAY- 10 September | 10:40 - 12:20
CHAIRS
Jae-Sun Seo, Guilherme Paim
SESSION PROGRAM
10:40 - 11:00
DiFlash: A 13.70TOPS/W 156.05Pixels/mJ Video Generation Accelerator with Bubble-Free Online-Quantized FlashAttention and Query-Reuse Sliding-Tile Attention
Jun Liu, Shiwei Liu, Li Ding, Shuaiheng Li, Tianlang Zhao, Yuxuan Zhang, Xinhao Li, Shan Huang, Jinhao Li, Qi Luo, Ya...
11:00 - 11:20
A 28nm 97.23 TFLOPS/W Pipelined Two-Stage-Aggregation Compute-in-Memory Macro for Edge-LLM Inference and Fine-Tuning
Yi Yang, Jinwu Chen, Yutong Zhang, Yucheng Du, Defa Wu, Yingxuan Zhou, Yanqi Zhang, Tianhui Jiao, Xiaoxue Zhong, Xing...
11:20 - 11:40
DRAGON: A 28nm 5.12Ms-Response and 42.83Token/S Digital CiM-Based Multi-Phase Accelerator for Hybrid-Intensive Retrieval-Augmented Generation
Zihan Wu, Yiqi Wang, Zhen He, Huiming Han, Yang Hu, Fengbin Tu, Shouyi Yin
11:40 - 12:00
BumbleBee: A 3 mW Fused-Kernel Flash Attention Edge AI Co-Processor in 18 nm FD-SOI CMOS
Joaquin A. Cornejo, Filipe Pouget, Amelie Poullot, Raphael Gras, Dominique Bousquet, Nicolas Lebouleux, Tarun Chawla,...
12:00 - 12:20
SSTP: A 13.32 TFLOPS/W End-to-End Simultaneous Speech Translation Processor on Edge Devices
Bokyoung Seo, Ghangmin Yun, Chaeyoon Kim, Jueun Jung, Kyuho Lee
