AUT Journal of Electrical Engineering

AUT Journal of Electrical Engineering

Efficient Dense Matrix Multiplication on UPMEM Processing-in-Memory Architecture: Optimization and Performance Analysis

Document Type : Research Article

Author
Department of Computer Science and Engineering, Shahid Beheshti University, 1983969411 Tehran, Iran
10.22060/eej.2026.26237.6085
Abstract
Matrix multiplication is a fundamental operation in image processing and many data-intensive applications, and its growing computational and memory demands increasingly challenge conventional CPU- and GPU-based systems. Processing-near/in-Memory (PnM/PiM) architectures address this challenge by reducing data movement between memory and computation units. UPMEM is a commercially available DRAM-based PIM platform that integrates lightweight processors within memory chips; however, efficiently mapping compute-intensive kernels such as dense matrix multiplication onto this architecture remains non-trivial due to limited arithmetic support, restricted on-chip memory, and low per-core frequency. In this work, we investigate the feasibility and performance of dense matrix multiplication on UPMEM. Starting from naive implementation, we apply loop reordering and tiling optimizations tailored to the UPMEM execution model. We analyze the impact of key architectural parameters, including the number of DPUs and tasklets, and compare UPMEM-based implementations against optimized CPU baselines under identical input sizes and precision constraints. Experimental results show that exploiting tasklet-level parallelism and DPU scalability yields significant performance improvements, with the tiling-based UPMEM implementation achieving up to 2.76x speedup over CPU tiling and outperforming loop-reordered CPU implementations for large matrices using 32-bit arithmetic.
Keywords
Subjects


Articles in Press, Accepted Manuscript
Available Online from 16 August 2026