Hoai Luan Pham

dblp:236/3499 · also Hoai-Luan Pham, Pham Hoai Luan · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
0000-0002-4272-0132ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MergeSlide: Continual Model Merging and Task-to-Class Prompt-Aligned Inference for Lifelong Learning on Whole Slide Images
abstract
Lifelong learning on Whole Slide Images (WSIs) aims to train or fine-tune a unified model sequentially on cancer-related tasks, reducing the resources and effort required for data transfer and processing, especially given the gigabyte-scale size of WSIs. In this paper, we introduce MergeSlide, a simple yet effective framework that treats lifelong learning as a model-merging problem by leveraging a vision–language pathology foundation model. When a new task arrives, it is ❶ defined with class-aware prompts, ❷ fine-tuned for a few epochs in a classifier-free manner, and ❸ merged into a unified model using an orthogonal continual-merging strategy that preserves performance and mitigates catastrophic forgetting. For inference under the class-incremental learning (CLASS-IL) setting, where task identity is unknown, we introduce Task-to-Class Prompt-aligned (TCP) inference. Specifically, TCP first identifies the most relevant task using task-level prompts and then applies the corresponding class-aware prompts to generate predictions. To evaluate MergeSlide, we conduct experiments on a stream of six TCGA datasets. The results show that MergeSlide outperforms both rehearsal-based continual learning and vision-language zero-shot baselines. Code and data are available at https://github.com/caodoanh2001/MergeSlide.
Doanh C. Bui, Ba Hung Ngo, Hoai Luan Pham, Khang Nguyen 0001, Maï K. Nguyen, Yasuhiko Nakashima
WACV3
2026 An innovative HLS framework for all network architectures: From Python to SoC
Minh Tan Ha, Xuan Thao Tran, Ngoc Quoc Tran, Vu Trung Duong Le, Hoai Luan Pham
Integr.6
2026 UPVM-ASR: An Ultralight Parallel Dual-Stream Visual Mamba U-Net With Sub-Million Parameters for Audio Super-Resolution
abstract
Audio super-resolution (ASR) restores missing high-frequency content from low-bandwidth speech, which frequently arises in IoT acoustic sensing and edge voice interfaces due to power, bandwidth, and latency constraints. Recent Mamba-based models improve long-range modeling efficiency but are often adapted from vision templates and are not optimized for ASR deployment. We propose UPVM-ASR, an ultralight dual-stream U-Net built on VMamba visual state-space (VSS) blocks to reconstruct magnitude and phase spectrograms in parallel. A channel-parallel, weight-shared selective-scan design (PVSS) reduces parameter redundancy while retaining global-local spectro-temporal modeling with linear complexity. UPVM-ASR is highly compact, achieving 0.41M parameters and 1.00 GFLOP, and remains competitive across 16 kHz and 48 kHz targets on the VCTK dataset. UPVM-ASR attains the lowest log-spectral distance (LSD) in the most challenging settings (from 2 kHz to 16 kHz and from 8 kHz to 48 kHz) while maintaining strong ViSQOL and STOI scores. These results indicate that optimized lightweight state-space U-Nets can deliver high-fidelity ASR under tight compute and communication budgets. Furthermore, to assess deployment and robustness, we additionally provide a deployment-oriented pure C implementation and benchmark runtime on embedded platforms, as well as report results under noisy/out-of-domain conditions on VCTK-DEMAND and NISQA datasets.
Phuong Quang Nguyen, Hoai Luan Pham, Shanq-Jang Ruan, Lieber Po-Hung Li
IEEE Internet Things J.2
2025 CTFE: A High-Efficient Heterogeneous Cryptographic CGRA for Diverse Security Applications
abstract
Nowadays, cryptographic computation across various security applications necessitates the development of hardware that is not only fast and power-efficient but also flexible enough to support a range of cryptographic algorithms. Unfortunately, existing computing platforms for cryptography struggle to balance high flexibility, high performance, and low power consumption. To address these issues, this article introduces the crypto-tailored flexible engine (CTFE), a next-generation coarse-grained reconfigurable array (CGRA) for cryptography. Concretely, the CTFE incorporates four innovative ideas to achieve high flexibility and performance with high hardware efficiency: 1) processing element array (PEA) with dual-buffer lanes and multiplexer optimization; 2) high flexibility Hyper-ALU; 3) heterogeneous PEA; and 4) bi-tiered pipeline coordination. Real-time evaluation results on Xilinx ZCU102 FPGA at the System-on-Chip (SoC) level demonstrate that the CTFE is 1.13–14.3 times better in throughput and 55.3–14 232 times better in energy efficiency than state-of-the-art CPUs. Experiments on an ASIC 45 nm CMOS technology show that the CTFE consumes the power of 1.06 W, occupies an area of$2.77 \; \text {mm}^{{2}}$, and operates at the frequency of 510 MHz. In comparison to existing CGRA solutions, CTFE outperforms 1.63–20.23 times in throughput and 1.61–73.2 times in area efficiency.
Vu Trung Duong Le, Hoai Luan Pham, Thi Hong Tran, Van Duy Tran, Tuan Hai Vu 0001, Yasuhiko Nakashima
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.2
2025 MINA: A Hardware-Efficient and Flexible Mini-InceptionNet Accelerator for ECG Classification in Wearable Devices
abstract
Classification is a crucial aspect of cardiovascular-related challenges, requiring thorough research and optimization to develop effective solutions for both patients and doctors. Recently, rapid advancements in artificial intelligence, particularly Convolutional Neural Networks (CNNs), have introduced numerous effective methods, significantly improving disease classification in Electrocardiogram (ECG) analysis. However, existing CNN-based accelerators often encounter challenges such as high parameter counts, limited flexibility in handling diverse CNN configurations, and inefficient hardware utilization. To address these issues, this paper proposes the Mini InceptionNet Accelerator (MINA), a hardware-efficient and flexible accelerator designed specifically for one-dimensional (1-D) CNN-based ECG classification. First, a novel 1-D CNN model, Mini InceptionNet, reduces the parameter count by 41.6% compared to the smallest existing 1-D CNN, minimizing memory requirements while maintaining high classification accuracy. Second, a flexible Processing Element Array (PEA) is designed with a Sharing Buffer Allocator (SBA) to support dynamic data coordination across various network topology parameters. Third, each Processing Element (PE) is equipped with four Local Data Memories (LDMs) and an ALU, enabling efficient intermediate data storage and versatile operations for modern CNN models. To demonstrate its effectiveness, MINA has been successfully implemented and verified on the ZCU102 FPGA at the system-on-chip level. FPGA evaluations show that MINA achieves 1.3×-2.9× higher energy efficiency (GOP/s/MeLUT) than state-of-the-art 2-D CNN accelerators. Compared to existing 1-D CNN accelerators, MINA achieves at least 1.53× improvement in the area-delay product (ADP). Additionally, weight pruning is discussed as a supporting strategy, achieving up to 3× faster inference time and a 2.13× improvement in ADP at 70% sparsity.
Hoai Luan Pham, Vu Trung Duong Le, Yasuhiko Nakashima
IEEE Trans. Circuits Syst. I Regul. Pap.1
2024 LiCryptor: High-Speed and Compact Multi-Grained Reconfigurable Accelerator for Lightweight Cryptography
abstract
Emerging modern internet-of-things (IoT) systems require hardware development to support multiple 8/32/64-bit lightweight cryptographic (LWC) algorithms with high speed and energy efficiency to ensure diverse security requirements. Accordingly, a coarse-grained reconfigurable array (CGRA) is considered the most effective architecture for achieving high speed, low power, and high flexibility for implementing LWC algorithms. However, existing CGRA designs for cryptography focus only on improvements to outdated 8/32-bit algorithms, suffer from large area requirements, and have long compilation times. To address these issues, this paper proposes a new CGRA-based accelerator named LiCryptor to support various 8/32/64-bit LWC algorithms with high speed and small area. Three innovative ideas are proposed to enable LiCryptor to achieve these goals: a compact multi-grained processing element array (M-PEA), a shared 8/32/64-bit arithmetic logic unit (ALU), and an assembly-like inline directive (AID) mapping method. The LiCryptor has been successfully implemented and verified on the Xilinx ZCU102 FPGA. Real-time performance evaluation across various LWC algorithms on FPGA shows that LiCryptor is 1.33 to 4 times better in execution time and 3.4 to 153 times better in power-delay products (PDP) compared to today’s most powerful CPUs. Notably, evaluation of AID mapping on the ARM Cortex-A53 CPU of the ZCU102 FPGA shows that its compilation time is less than 1.5 ms for most LWC algorithms, at least 2,333 times faster than CFG mapping in current CGRAs. Moreover, experimental results on 45nm ASIC technology show that the LiCryptor significantly outperforms existing CGRAs and other reconfigurable designs in terms of throughput and area efficiency.
Hoai Luan Pham, Vu Trung Duong Le, Van Duy Tran, Tuan Hai Vu 0001, Yasuhiko Nakashima
IEEE Trans. Circuits Syst. I Regul. Pap.1
2021 BCA: A 530-mW Multicore Blockchain Accelerator for Power-Constrained Devices in Securing Decentralized Networks
abstract
Blockchain distributed ledger technology (DLT) has widespread applications in society 5.0 because it improves service efficiency and significantly reduces labor costs. However, employing blockchain DLT entails considerable energy consumption in the mining process. This paper proposes a blockchain accelerator (BCA) with ultralow power consumption and a high processing rate to address the problem. The BCA focuses on accelerating the double secure hash algorithm (SHA) 256 function required in the mining process at a system-on-chip (SoC) level. We propose three ideas, namely, multiple local memories (multimem), double-cell processing element (D-PE), and nonce autoupdate (NAU), to reduce the external data transfer time and improve the BCA hardware efficiency. We propose a cascaded multiple BCA chip model to enhance the system throughput by several-fold. Our experiments on an ASIC and FPGA prove that the proposed BCA successfully performs the mining process for multiple blockchain networks with much lower power consumption than that of the state-of-the-art CPUs and GPUs. The BCA is laid out with Renesas 65 nm technology with a chip area of$25~mm^{2}$and consumes$530~mW$at 100MHz. The power efficiency of the layout chip is improved by 2428 and 143 times compared with that of the fastest CPU Intel i9-10940X and GPU RTX 3090, respectively.
Thi Hong Tran, Hoai Luan Pham, Tri Dung Phan, Yasuhiko Nakashima
IEEE Trans. Circuits Syst. I Regul. Pap.2