Information for Paper ID 8206
Paper Information:
Paper Title: BumbleBee: A 3 mW Fused-Kernel Flash Attention Edge AI Co-Processor in 18 nm FD-SOI CMOS 
Student Contest: Yes 
Affiliation Type: Industry 
Keywords: Co-processor, Transformers, RISC-V, Systolic Array, Addressable MAC Array, Flash Attention, Mixed-Precision 
Abstract: The attention mechanism is the core operation in Transformer models but remains difficult to deploy at the edge due to high memory access and data movement requirements. This work presents BumbleBee, an 18nm FD-SOI CMOS co-processor implementing the Flash Attention mechanism for the first time. It achieves 3mW power consumption and reduces bus transfers by 5 through a fully dataflow-driven architecture. Beyond energy efficiency, this dataflow also reduces latency by minimizing memory stalls, enabling fast token-level processing. The architecture integrates a dedicated attention head combining an addressable MAC array together with systolic arrays, and leverages binarized representations to further reduce computational cost. It achieves 20 and 248mJ/inference for BERT-Mini and DETR workloads, respectively. 
Track ID: 13 
Track Name: Architectures and Circuits for AI and ML 
Final Decision: Accept as Lecture 
Session Name: LLMs, Transformers, and Generative AI Accelerators (Lecture)