On-Chip Memory Access Reduction for Energy-Efficient Dilated Convolution Processing
Date
Advisor
Volume
Issue
Journal
Series Titel
Book Title
Publisher
Supplementary Material
Other Versions
Link to publishers' Version
Abstract
Dilated convolutions have recently become increasingly popular in deep neural networks. However, the inference of these operations on hardware accelerators is not mature enough to reach the efficiency of standard convolutions. Therefore, we extended a dedicated accelerator for dilated convolutions to reduce the number of energy-intensive accesses to the on-chip memory. We achieve this by applying the principle of feature map decomposition to an output-stationary compute array with a strided feature loading. Our solution shows a 50% reduction in memory accesses for an unpadded 3×3 kernel and a dilation rate of 9 compared to a recently proposed dilated convolution accelerator. We also support flexible parameter selection for kernel sizes and dilation rates to meet the requirements of modern neural networks. The energy consumption of the additional hardware modules is less than the savings achieved by the reduced memory accesses. This results in a relative energy saving by a factor of 4.77 for dilated convolutions with unpadded 3×3 kernels.
Description
Keywords
Keywords GND
Conference
Publication Type
Version
Collections
License
Dieses Dokument darf im Rahmen von § 53 UrhG zum eigenen Gebrauch kostenfrei heruntergeladen, gelesen, gespeichert und ausgedruckt, aber nicht auf anderen Webseiten im Internet bereitgestellt oder an Außenstehende weitergegeben werden.
