Department of Computer Science and Engineering, Shahid Beheshti University, 1983969411 Tehran, Iran
10.22060/eej.2026.26237.6085
Abstract
Matrix multiplication is a fundamental operation in image processing and many data-intensive applications, and its growing computational and memory demands increasingly challenge conventional CPU- and GPU-based systems. Processing-near/in-Memory (PnM/PiM) architectures address this challenge by reducing data movement between memory and computation units. UPMEM is a commercially available DRAM-based PIM platform that integrates lightweight processors within memory chips; however, efficiently mapping compute-intensive kernels such as dense matrix multiplication onto this architecture remains non-trivial due to limited arithmetic support, restricted on-chip memory, and low per-core frequency. In this work, we investigate the feasibility and performance of dense matrix multiplication on UPMEM. Starting from naive implementation, we apply loop reordering and tiling optimizations tailored to the UPMEM execution model. We analyze the impact of key architectural parameters, including the number of DPUs and tasklets, and compare UPMEM-based implementations against optimized CPU baselines under identical input sizes and precision constraints. Experimental results show that exploiting tasklet-level parallelism and DPU scalability yields significant performance improvements, with the tiling-based UPMEM implementation achieving up to 2.76x speedup over CPU tiling and outperforming loop-reordered CPU implementations for large matrices using 32-bit arithmetic.
Cheshmikhani,E . (2026). Efficient Dense Matrix Multiplication on UPMEM Processing-in-Memory Architecture: Optimization and Performance Analysis. (e6155). AUT Journal of Electrical Engineering, (), e6155 doi: 10.22060/eej.2026.26237.6085
MLA
Cheshmikhani,E . "Efficient Dense Matrix Multiplication on UPMEM Processing-in-Memory Architecture: Optimization and Performance Analysis" .e6155 , AUT Journal of Electrical Engineering, , , 2026, e6155. doi: 10.22060/eej.2026.26237.6085
HARVARD
Cheshmikhani E. (2026). 'Efficient Dense Matrix Multiplication on UPMEM Processing-in-Memory Architecture: Optimization and Performance Analysis', AUT Journal of Electrical Engineering, (), e6155. doi: 10.22060/eej.2026.26237.6085
CHICAGO
E Cheshmikhani, "Efficient Dense Matrix Multiplication on UPMEM Processing-in-Memory Architecture: Optimization and Performance Analysis," AUT Journal of Electrical Engineering, (2026): e6155, doi: 10.22060/eej.2026.26237.6085
VANCOUVER
Cheshmikhani E. Efficient Dense Matrix Multiplication on UPMEM Processing-in-Memory Architecture: Optimization and Performance Analysis. AUT J Electr Eng. 2026;():e6155. doi: 10.22060/eej.2026.26237.6085