Design and Implementation of Parallel MAC Unit for FPGA-Based Deep Learning Algorithms Using VHDL

Authors

  • A.Naganjanamma1 , G.Lakshmi Bharath2 Author

DOI:

https://doi.org/10.62643/

Keywords:

MAC, GOPS, CNN, FPGA

Abstract

Deep neural network algorithms have proven their enormous capabilities in wide range of artificial intelligence applications, especially in Printed/Handwritten text recognition, Multimedia processing, Robotics and many other high-end technological trends. The most challenging aspect nowadays is to overcome the extremely computational processing demands in applying such algorithms, especially in real-time systems. Recently, the Field Programmable Gate Array (FPGA) has been considered as one of the optimum hardware accelerator plat form for accelerating the deep neural network architectures due to its large adaptability and the high degree of parallelism it offers. In this paper, the proposed 8-bits fixed-point parallel multiply accumulate (MAC) unit architecture aimed to create a fully-customize MAC unit for the Convolutional Neural Networks (CNN) instead of depending on the conventional DSP blocks and embedded memories units on the FPGAs architecture silicon fabrics. The proposed 8-bits fixed-point parallel multiply-accumulate (MAC) unit architecture is designed using VHDL language and can performs a computational speed up to 4.17 Giga Operation per Second (GOPS) using high-density FPGAs

Downloads

Published

20-05-2026

How to Cite

Design and Implementation of Parallel MAC Unit for FPGA-Based Deep Learning Algorithms Using VHDL. (2026). International Journal of Engineering Research and Science & Technology, 22(2(1), 2412-2416. https://doi.org/10.62643/