Fast convolutional neural networks on FPGAs with hls4ml
Name
Aarrestad_2021_Mach._Learn.__Sci._Technol._2_045015.pdf
Description
Published version
Size
2.18 MB
Format
Adobe PDF
Checksum (MD5)
e592d10c0a10c59d597bd2bf054c7975
Author(s) • • • • • • • • •
Aarrestad, Thea
Loncar, Vladimir
Ghielmetti, Nicolò
Pierini, Maurizio
Summers, Sioni
Ngadiuba, Jennifer
Petersson, Christoffer
Linander, Hampus
Iiyama, Yutaro
Di Guglielmo, Giuseppe
Date Issued
2021
Journal
Machine Learning: Science and Technology
Publisher
IOP Publishing
Citation
Aarrestad, Thea, Loncar, Vladimir, Ghielmetti, Nicolò, Pierini, Maurizio, Summers, Sioni et al. 2021. "Fast convolutional neural networks on FPGAs with hls4ml." Machine Learning: Science and Technology, 2 (4).
Version
Final published version
Abstract
Abstract
We introduce an automated tool for deploying ultra low-latency, low-power deep neural networks with convolutional layers on field-programmable gate arrays (FPGAs). By extending the hls4ml library, we demonstrate an inference latency of 5 µs using convolutional architectures, targeting microsecond latency applications like those at the CERN Large Hadron Collider. Considering benchmark models trained on the Street View House Numbers Dataset, we demonstrate various methods for model compression in order to fit the computational constraints of a typical FPGA device used in trigger and data acquisition systems of particle detectors. In particular, we discuss pruning and quantization-aware training, and demonstrate how resource utilization can be significantly reduced with little to no loss in model accuracy. We show that the FPGA critical resource consumption can be reduced by 97% with zero loss in model accuracy, and by 99% when tolerating a 6% accuracy degradation.
We introduce an automated tool for deploying ultra low-latency, low-power deep neural networks with convolutional layers on field-programmable gate arrays (FPGAs). By extending the hls4ml library, we demonstrate an inference latency of 5 µs using convolutional architectures, targeting microsecond latency applications like those at the CERN Large Hadron Collider. Considering benchmark models trained on the Street View House Numbers Dataset, we demonstrate various methods for model compression in order to fit the computational constraints of a typical FPGA device used in trigger and data acquisition systems of particle detectors. In particular, we discuss pruning and quantization-aware training, and demonstrate how resource utilization can be significantly reduced with little to no loss in model accuracy. We show that the FPGA critical resource consumption can be reduced by 97% with zero loss in model accuracy, and by 99% when tolerating a 6% accuracy degradation.
MIT Department
Massachusetts Institute of Technology. Department of Physics
Terms of Use
Creative Commons Attribution 4.0 International license
Persistent DSpace Link
DOI of Published Version
https://doi.org/10.1088/2632-2153/AC0EA1