FPGA Based High Performance and Scalable Block LU Decomposition Architecture

Manish Kumar Jaiswal, Nitin Chandrachoodan
Indian Institute of Technology, Madras
IEEE Transactions on Computers, 2011


   title={FPGA Based High Performance and Scalable Block LU Decomposition Architecture},

   author={Jaiswal, M.K. and Chandrachoodan, N.},

   journal={IEEE Transactions on Computers},


   publisher={Published by the IEEE Computer Society}


Source Source   



Decomposition of a matrix into lower and upper triangular matrices (LU decomposition) is a vital part of many scientific and engineering applications, and the block LU decomposition algorithm is an approach well suited to parallel hardware implementation. This paper presents an approach to speed up implementation of the block LU decomposition algorithm using FPGA hardware. Unlike most previous approaches reported in the literature, the approach does not assume the matrix can be stored entirely on-chip. The memory accesses are studied for various FPGA configurations, and a schedule of operations for scaling well is shown. The design has been synthesized for FPGA targets and can be easily re-targeted. The design outperforms previous hardware implementations, as well as tuned software implementations including the ATLAS and MKL libraries on workstations.
No votes yet.
Please wait...

* * *

* * *

HGPU group © 2010-2024 hgpu.org

All rights belong to the respective authors

Contact us: