EDBT 2026 Demo / reviewers in the wild / expert
Xuan Truong Nguyen
dblp:217/0717
· DBLP profile ↗
26ranked-venue papers
1as first author
23since 2021 · last 2026
0000-0002-7527-6971ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 21 · 19 since 2021Software engineering, systems software and programming languages · 5 · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Communication-Aware Hybrid Parallelism Mapping for Low-Cost MCM-based DNN AcceleratorsabstractThe growing scale of deep neural networks has surpassed the capacity of single-chip accelerators, particularly pin cost-sensitive edge devices. Multi-Chip-Module (MCM) architectures enable scalability but rely on bandwidth-limited chip-to-chip (C2C) interfaces, causing substantial inter-chip communication overhead. Among model-parallel strategies, tensor parallelism (TP) offers high concurrency at the cost of communication overhead, while pipeline parallelism (PP) reduces it at the cost of lower compute utilization inherent to pipeline execution. This work presents Stitch, a two-phase rebalancing framework for hybrid model-parallel mapping in low-cost MCM-based CNN accelerators. Phase I mitigates TP’s C2C-induced communication overhead by jointly optimizing partitioning and datapath through a layer-wise C2C–DRAM selection solved via dynamic programming. Since TP alone cannot fully minimize communication, Phase II extends the design space by combining TP and PP at the package level. Guided by simulated annealing, Stitch selects layer groups, tunes pipeline stages, and balances communication–utilization trade-offs. Evaluation on a cycle-accurate simulator shows that Stitch reduces the energy–delay product by up to 42.8% compared to prior TP-based methods, demonstrating its effectiveness under practical C2C bandwidth constraints. Jicheon Kim, Chunmyung Park, Xuan Truong Nguyen |
DATE | 3 |
| 2026 | STEAM-SSM: A Streaming-Tiled Efficient Accelerator for Mamba-2 SSM on FPGA
Sehoon Chon, Xuan Truong Nguyen |
ISCAS | 2 |
| 2026 | PISA: Parameter-Efficient Image Super-Resolution Acceleration for Multiplierless Architecture
Xuan Truong Nguyen |
ISLPED | 2 |
| 2026 | TTX: Towards Autotuning Triton Kernels via Latency Prediction with XGBoostabstractGPU kernel optimization is essential for highperformance machine learning systems, but it remains laborintensive. Triton simplifies GPU kernel development with a Python-based interface, yet autotuning still requires searching a large configuration space, which is costly. We present TTX, a Triton tuning framework with an XGBoost-based performance predictor. By using input shapes, tuning parameters, and IR-level features, TTX accurately estimates kernel latency and quickly selects promising candidates for further search. Experiments show that TTX achieves around $10 \%$ MAPE on operators such as matrix multiplication, batched matrix multiplication, and convolution across multiple GPUs. It reaches about $80 \%$ of the best performance with the top-1 candidate and over $95 \%$ with the top- 50 candidates. Van-Sang Pham, Tien Son Pham, Tuan-Duc Chu, Xuan Truong Nguyen, Thanh Tuan Dao |
ISPASS | 4 |
| 2025 | DEAR-PIM: Processing-in-Memory Architecture with Disaggregated Execution of All-bank RequestsabstractEmerging transformer-based large language models (LLMs) involve many low-arithmetic intensity operations, which result in sub-optimal performance on general-purpose CPUs and GPUs. Processing-in-Memory (PIM) has shown promise in enhancing performance by reducing data movement bottlenecks. Commodity near-bank PIMs enable in-memory computation through bank-level compute units and typically rely on all-bank commands, which simultaneously operate the compute units of all banks to maximize internal bandwidth and parallelism. However, activating all banks simultaneously before issuing all-bank commands generally requires high peak power, which may exceed system power limit, when stacking multiple PIM devices for LLM inference. Additionally, under a DRAM power constraint, all-bank commands are only issued after all banks are fully activated through a sequence of single-bank activations, incurring bubble cycles and degrading overall performance. To address these shortcomings, this study proposes DEAR-PIM, a novel PIM architecture with Disaggregated Execution of All-bank Requests. DEAR-PIM incorporates disaggregated command queue, allowing it to buffer all-bank commands and provide them to each bank sequentially without waiting to complete all-bank activations. However, since all banks must finish their disaggregated execution before simultaneous post-processing, synchronization between early-activated and last-activated banks is necessary. To tackle the issue, DEAR-PIM introduces a column-aware synchronization command scheme that inserts no-op-like commands into unused columns without modifying the memory controller. Experiments demonstrate that DEAR-PIM achieves a speedup of 2.03-3.33′ over an A100 GPU and improves performance by 1.11-1.52′ compared to the sequential activation scheme. DEAR-PIM also reduces the peak power consumption by 21.3-41.7% compared to the simultaneous activation scheme. Jungi Hyun, Seongho Jeong, Xuan Truong Nguyen |
DATE | 5 |
| 2025 | Leveraging Hot Data in a Multi-Tenant Accelerator for Effective Shared Memory ManagementabstractMulti-tenant neural networks (MTNN) have been emerging in various domains. To effectively handle multi-tenant workloads, modern hardware systems typically incorporate multiple compute cores with shared memory systems. While prior works have intensively studied compute- and bandwidth-aware allocation, on-chip memory allocation for MTNN accelerators has not been well studied. This work identifies two key challenges of on-chip memory allocation in MTNN accelerators: on-chip memory shortages, which force data eviction to off-chip memory, and on-chip memory underutilization, where memory remains idle due to coarse-grained allocation. Both issues lead to increased external memory accesses (EMAs), significantly degrading system performance. To address these challenges, we propose HotPot, a novel multi-tenant accelerator with a runtime temperature-aware memory allocator. HotPot prioritizes hot data for global on-chip memory allocation, reducing unnecessary EMAs and optimizing memory utilization. Specifically, HotPot introduces a temperature score that quantifies reuse potential and guides runtime memory allocation decisions. Experimental results demonstrate that HotPot improves system throughput (STP) by up to 1.88 × and average normalized turnaround time (ANTT) by 1.52 × compared to baseline methods. Chunmyung Park, Jicheon Kim, Eunjae Hyun, Xuan Truong Nguyen |
DATE | 4 |
| 2025 | Live Demonstration: DVS-CIS Sensor Fusion System for Real-Time DNN-Based Object DetectionabstractThis demonstration presents a high-speed, energy-efficient sensor fusion system integrating CMOS Image Sensors (CIS) and Dynamic Vision Sensors (DVS) for advanced image recognition. Using CIS for high-res imaging and DVS for rapid event-driven capture, the FPGA-implemented architecture with an NPU running YOLOv3-Tiny achieves 18 ms inference latency with a minimal 2.78% mAP loss. Selective NPU activation based on DVS-detected regions yielded 31.5% power savings, while a custom receiver module efficiently fused DVS (13,900 fps) and CIS (60 fps) data. The system uses a power of 6.977 W on a Xilinx Zynq+ ZCU106 board. Mincheol Cha, Keehyuk Lee, Bobaro Chang, Soosung Kim 0003, Xuan Truong Nguyen, Tae Sung Kim, Hyunsurk Ryu |
ISCAS | 6 |
| 2025 | A DVS-CIS Sensor Data Receiver on FPGA with a 10 Gbps MIPI ControllerabstractFusing a dynamic vision sensor (DVS) and a CMOS image sensor (CIS) is promising in real-time vision applications. However, unlike common CIS, DVS typically come with a custom data format due to their naturally sparse data, which becomes a challenge to fuse DVS and CIS data streams on a general-purpose CPU. To address this problem, this work proposes a DVS-CIS sensor stream receiver on FPGA. The proposed receiver incorporates a cost-effective address decoder and an inline transpose to receive and store a DVS stream on DRAM effectively. At a system level, a host PC can stream the DVS-CIS stream from FPGA via PCIe and display streams on a monitor. Experimental results demonstrate that our architecture can decode up to 13,900fps of DVS frames without incurring any frame drops while concurrently streaming frames at 60fps from a CIS. The design only uses 135 BRAM, 38 DSPs, 69489 LUTs, and 86626 FFs on a Xilinx Zynq+ ZCU106 FPGA board and consumes a power of 6.977 W. Mincheol Cha, Keehyuk Lee, Bobaro Chang, Soosung Kim 0003, Xuan Truong Nguyen, Tae Sung Kim, Hyunsurk Ryu |
ISCAS | 6 |
| 2025 | An Energy-Efficient Daily Surveillance System with DVS-CIS Sensor Fusion and Event-based NPU TriggeringabstractThis study presents a daily surveillance system based on dynamic vision sensors (DVS) and CMOS image sensors (CIS) to enable real-time image recognition with low energy consumption. In such a system, a neural processing unit (NPU) - which executes a DNN model to detect objects on a given CIS image - may consume a lot of energy when always-on. To address the problem, this work introduces a system with a DVS-based region of interest (ROI) detector and an event-based NPU trigger for energy savings. Based on DVS, the ROI detector effectively recognizes scene changes in dynamic environments, e.g., low-light scenes at midnight, which serves as a trigger to invoke the NPU for object detection. Our system prototype was built on a host PC and two Xilinx Zynq+ ZCU106 FPGA boards, one for the DVS-CIS receiver and the other for our NPU. The experimental results demonstrated that Over a 24-hour testing period, our system achieved a 31.5% reduction in energy usage. Operating a YOLOv3-Tiny object detector at 200 MHz, our NPU achieves a latency of just 18 ms, enabling seamless real-time monitoring capabilities. Mincheol Cha, Keehyuk Lee, Bobaro Chang, Soosung Kim 0003, Daniel Moon, Tae Sung Kim, Hyunsurk Ryu, Xuan Truong Nguyen |
ISCAS | 10 |
| 2025 | Cost-Effective Reconfigurable MCM with Common-Value Elimination and AlignmentabstractMultiple constant multiplication (MCM) generates area/energy-efficient circuit with additions and shifts for FIR filters and DNN inference acceleration. Reconfigurable MCM (ReMCM) enables time-multiplexing execution of MCM circuits for multiple constant sets by reusing adders across different sets. Finding a shift-add-mux circuit in ReMCM, however, generally relies on a rigid order of constants, which may cause unnecessary topology conflicts and thereby incur a power/area overhead. To address the problem, this work proposes an Adaptive Grouping Compatible Graph Synthesis (AG-CGS) framework for ReMCM, which reorders constants to enhance power/area efficiency. Employing an (index, value) data structure for a constant, AG-CGS introduces two novel techniques in ReMCM: intra-set common value elimination (Intra-SCE) and inter-set common value alignment (Inter-SCA). Consequently, AG-CGS enables searching for a cost-effective circuit associated with reordered coefficients to reduce redundant computations. The experimental results demonstrate that AG-CGS reduces area by up to 58.1%, 16.3%, and 13.8%, and energy by up to 63.1%, 15.2%, and 12.1%, compared to [1], [2], and [3], respectively. Xuan Truong Nguyen, Dongsuk Jeon |
ISCAS | 2 |
| 2025 | Live Demonstration: A Scalable CNN Accelerator SoC With a Cost-Effective Chip-to-Chip AdapterabstractIn this demonstration, we present a system-on-chip (SoC) designed to support scalable CNN acceleration at a low cost. The SoC features a cost-effective chip-to-chip adapter that enables scalable performance improvements with minimal design costs. This adapter manages chip-to-chip synchronization and data scheduling, effectively reducing chip-to-chip latency overhead. The SoC is implemented using the Silterra 130-nm CMOS technology with a chip size of 6.3×6.3 mm2. Jicheon Kim, Chunmyung Park, Eunjae Hyun, Xuan Truong Nguyen |
ISCAS | 4 |
| 2025 | Mixed INT4-INT8 LLM Quantization via Progressive Layerwise Assignment with Dynamic Sensitivity EstimationabstractQuantization is essential for optimizing large language models (LLMs) by reducing memory usage and computational load. However, traditional low-bit quantization methods applied uniformly across all layers can significantly degrade accuracy, especially on modern hardware architectures. We introduce a novel, adaptive quantization strategy leveraging Layer Sensitivity, which assigns bit-width to each layer based on its sensitivity to quantization. This method includes both Static Sensitivity Estimation—a single-pass sensitivity measurement to rank layers by quantization tolerance—and Dynamic Sensitivity Estimation, which iteratively re-evaluates layer sensitivity after each quantization step, ensuring optimal bit-width allocation. Tested on models such as GPT2, OPT, and LLaMA-3.2, our approach minimizes accuracy loss and outperforms uniform quantization methods. This scalable solution effectively balances computational efficiency and accuracy, offering a robust path for deploying low-precision LLMs. Xuan Truong Nguyen |
ISCAS | 3 |
| 2025 | ENAF: A Multi-Exit Network with an Adaptive Patch Fusion for Large Image Super ResolutionabstractTo accelerate single image super-resolution (SISR) networks on large images (2K-8K), many recent approaches decompose an image into small patches and dynamically determine an execution path according to its difficulty (referred to as a dynamic network). To quantify the hardness of a patch, they mainly rely on a handcrafted assessment score, e.g., edge, which weakly associates a patch's texture with the computational complexity of a SISR model. To address the problem, we introduce ENAF - a dynamic network for SISR with an adaptive patch fusion. Built on top of a backbone, ENAF incorporates multiple early exits (EEs) to tackle the over-parameterized SISR model. More importantly, ENAF plugs a tiny network that estimates PSNR to associate data texture with a computation cost at an EE. Based on the scores, ENAF effectively assigns image patches to an exit, enhancing the quality-complexity trade off. Extensive experiments on common datasets with popular SISR backbones demonstrate the effectiveness of ENAF in various settings. The source code will be available. Manh Duong Nguyen, Tuan Nghia Nguyen 0002, Xuan Truong Nguyen |
WACV | 3 |
| 2025 | FDAM: Filter-Dedicated Approximate Multiplier Design for Real-Time CNN AccelerationabstractVideo super-resolution (VSR) is widely used in various high-definition applications, such as HDTVs and smartphones, requiring a dedicated upscaling technique for real-time full-HD generation. To reduce on-chip buffers for large-size output feature maps, a streaming VSR accelerator may employ an output stationary dataflow, leading to a large energy consumption caused by frequent filter switching. To mitigate this, we introduce a new filter-dedicated multiplier design for real-time VSR acceleration. We replace costly multipliers with adders, shifters, and multiplexers (MUXes), referred to as unified multiple constant multiplications (UMCM). The conventional UMCM, however, may incur considerable area/power overhead due to the unified topology constraint among different filter sets. To address this problem, we propose a new approximated MCM (AMCM) problem to relax the constraint and an approximate compatible graph synthesis (A-CGS) framework to efficiently solve AMCM by jointly searching for approximated filters and constructing a unified graph. Additionally, we suggest a lightweight fine-tuning method by freezing approximated filters and only fine-tuning biases, which can recover the original model’s accuracy within a few epochs. Experimental results with synthetic data demonstrate that AMCM reduces the area by up to 49.8%, 44.8%, and 40.3% when considering constant sets of 2, 4, and 8, respectively. Our designs with SR application achieve up to a 73.3% reduction in energy consumption. Experiments with the Set5 and Set14 datasets show that our model with bias correction achieves similar restoration performance compared to the eight-bit models. Xuan Truong Nguyen, Dongsuk Jeon |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 2025 | FACET: On-the-Fly Activation Compression for Efficient Transformer TrainingabstractTraining Transformer models, known for their outstanding performance in various tasks, can be challenging due to extensive training times and substantial memory requirements. One promising approach to minimize the memory footprint and accelerate training is compressing activations, which account for more than half of the total memory usage for large batch sizes. However, conventional compression schemes, such as FP8 quantization, may not adequately represent the dynamic range of Transformer activations, potentially leading to unsatisfactory accuracy. To address this issue, we propose FACET, a lightweight yet effective Transformer activation compressor and its corresponding hardware design. This compressor comprises a base-delta compression (BDC) and a bit-plane compression (BPC), targeting the exponent and sign/mantissa of activation data, respectively. The bitstreams generated by BDC and BPC are concatenated and then truncated to a target size, e.g., 8 bits per data. Experimental results with popular Transformer models (BERT, GPT-2, and T5) indicate that FACET reduces activation memory by 2-$4\times $, with negligible accuracy degradation. We implemented our compressor in hardware and synthesized it using the 45nm TSMC process library. The encoder and decoder require 16K and 12K gate counts, respectively, exhibiting a 2.2-$3.8\times $smaller overhead compared to other compressors. We also propose a system that integrates our compressor within memory with minimal system modifications, leveraging the small overhead of the compressor. Seungyong Lee 0003, Geonu Yun, Xuan Truong Nguyen |
IEEE Trans. Circuits Syst. I Regul. Pap. | 3 |
| 2024 | IANUS: Integrated Accelerator based on NPU-PIM Unified Memory SystemabstractAccelerating end-to-end inference of transformer-based large language models (LLMs) is a critical component of AI services in datacenters. However, the diverse compute characteristics of LLMs' end-to-end inference present challenges as previously proposed accelerators only address certain operations or stages (e.g., self-attention, generation stage, etc.). To address the unique challenges of accelerating end-to-end inference, we propose IANUS - Integrated Accelerator based on NPU-PIM Unified Memory System. IANUS is a domain-specific system architecture that combines a Neural Processing Unit (NPU) with a Processing-in-Memory (PIM) to leverage both the NPU's high computation throughput and the PIM's high effective memory bandwidth. In particular, IANUS employs a unified main memory system where the PIM memory is used both for PIM operations and for NPU's main memory. The unified main memory system ensures that memory capacity is efficiently utilized and the movement of shared data between NPU and PIM is minimized. However, it introduces new challenges since normal memory accesses and PIM computations cannot be performed simultaneously. Thus, we propose novel PIM Access Scheduling that manages not only the scheduling of normal memory accesses and PIM computations but also workload mapping across the PIM and the NPU. Our detailed simulation evaluations show that IANUS improves the performance of GPT-2 by 6.2× and 3.2×, on average, compared to the NVIDIA A100 GPU and the state-of-the-art accelerator. As a proof-of-concept, we develop a prototype of IANUS with a commercial PIM, NPU, and an FPGA-based PIM controller to demonstrate the feasibility of IANUS. Xuan Truong Nguyen, Seok Joong Hwang, Yongkee Kwon, Guhyun Kim, Chanwook Park, Ilkon Kim, Jaehan Park, Jeongbin Kim 0001, Woojae Shin, Jongsoon Won, Haerang Choi, Kyuyoung Kim, Daehan Kwon, Chunseok Jeong, Yongseok Choi, Wooseok Byun, Seungcheol Baek, John Kim 0001 |
ASPLOS (3) | 2 |
| 2024 | A Resource-Constrained Spatio-Temporal Super Resolution ModelabstractThis paper proposes a resource-constrained spatio-temporal super resolution (SR) model, which effectively enhances both the spatial resolution and frame rate of input videos. Replacing the entire deep learning model for spatio-temporal SR on devices that already have spatial SR capability is a challenging task. This is especially true for compact devices like image sensors that are composed of hardware modules. There is a need to enable spatio-temporal SR with minimal hardware overhead on devices that already have the SR module. The proposed model demonstrates an example of hardware implementation by combining independent hardware-friendly spatial SR and frame interpolation (FI). This configuration allows for seamless support of spatial SR, temporal SR, and spatio-temporal SR functionalities through data flow reconfiguration. Moreover, we propose schemes that leverage the flow estimation module to further reduce the computational burden of spatial SR. The experimental results show that the proposed model achieves competitive quality with state-of-the-art (SOTA) methods, while utilizing very limited computational resources. Da Hyeon Jung, Min-Wu Jeong, Xuan Truong Nguyen, Chae-Eun Rhee |
ISCAS | 3 |
| 2024 | A Scalable Multi-Chip YOLO Accelerator With a Lightweight Inter-Chip AdapterabstractMulti-chip-module (MCM) technology offers a promising solution for designing large-scale deep-learning inference systems while concurrently minimizing fabrication and design costs. Nevertheless, compared to monolithic dies, MCMs often incur resource and performance overhead due to interchip communication. Addressing these challenges, this paper introduces a scalable MCM-based DNN accelerator that incorporates a lightweight chip-to-chip adapter (C2CA) and an effective multi-chip dataflow. Inspired by the on-chip bus architecture, the C2CA efficiently shares data and address channels, achieving nearly optimal throughput with significantly reduced pin count and hardware costs. Additionally, the proposed design adopts a layer-wise dataflow within a ring-based C2C topology to fully utilize the constrained C2C bandwidth, mitigating both performance and communication overhead. When implemented on the Xilinx ZCU104 FPGA board, the system demonstrates significant throughput improvements compared to a single-chip configuration, yielding 1.92x and 3.57x enhancements for 2-chip and 4-chip configurations, respectively, on the YOLOv3-Tiny. Jicheon Kim, Chunmyung Park, Eunjae Hyun, Xuan Truong Nguyen |
ISCAS | 4 |
| 2024 | USDN: A Unified Sample-wise Dynamic Network with Mixed-Precision and Early-ExitabstractTo reduce computation in deep neural network inference, a promising approach is to design a network with multiple internal classifiers (ICs) and adaptively select an execution path based on the complexity of a given input. However, quantizing an input-adaptive network, a must-do task for network deployment on edge devices, is a non-trivial task due to jointly allocating its computation budget along with network layers and IC locations. In this paper, we propose Unified Sample-wise Dynamic Network (USDN) with a mixed-precision and early-exit framework that obtains both the optimal location of ICs and layer-wise bit configurations under a given computation budget. The proposed USDN comprises multiple groups of layers, with each group representing a varying degree of complexity for input samples. Experimental results demonstrate that our approach reduces computational cost of the previous work by 12.78% while achieving higher accuracy on ImageNet dataset. Ji-Ye Jeon, Xuan Truong Nguyen, Soojung Ryu |
WACV | 2 |
| 2024 | A Low-Latency FPGA Accelerator for YOLOv3-Tiny With Flexible Layerwise Mapping and DataflowabstractObject detection models have demonstrated outstanding performance in terms of accuracy. However, mapping convolutional neural network-based object-detection models to memory and computing-constrained devices is still challenging, which commonly leads to accuracy degradation and long latency. To address the problem, this work presents a design methodology to map the YOLOv3-tiny model onto a small FPGA board, in this case the Nexys A7-100T, which only has 0.5 MB on-chip SRAM and 240 DSPs. First, we design four identical MAC arrays to maximize the throughput by utilizing both DSPs and LUTs. Second, to exploit the MACs fully, we propose a dynamic data reuse scheme that handles inter-layer and intra-layer executions effectively under a small on-chip SRAM footprint. To this end, the proposed accelerator achieves an inference speed of 76.75 frames per second and throughput of 95.08 GOPs at 100MHz and consumes power of 2.203W. Specifically, it achieves a hardware utilization rate of 82.53%, thus significantly outperforming current YOLOv3-tiny accelerators. Minsik Kim 0006, Kyoungseok Oh, Youngmock Cho, Hojin Seo, Xuan Truong Nguyen |
IEEE Trans. Circuits Syst. I Regul. Pap. | 5 |
| 2023 | ViT-P3DE∗: Vision Transformer Based Multi-Camera Instance Association with Pseudo 3D Position EmbeddingsabstractMulti-camera instance association, which identifies identical objects among multiple objects in multi-view images, is challenging due to several harsh constraints. To tackle this problem, most studies have employed CNNs as feature extractors but often fail under such harsh constraints. Inspired by Vision Transformer (ViT), we first develop a pure ViT-based framework for robust feature extraction through self-attention and residual connection. We then propose two novel methods to achieve robust feature learning. First, we introduce learnable pseudo 3D position embeddings (P3DEs) that represent the 3D location of an object in the world coordinate system, which is independent of the harsh constraints. To generate P3DEs, we encode the camera ID and the object's 2D position in the image using embedding tables. We then build a framework that trains P3DEs to represent an object's 3D position in a weakly supervised manner. Second, we also utilize joint patch generation (JPG). During patch generation, JPG considers an object and its surroundings as a single input patch to reinforce the relationship information between two features. Ultimately, experimental results demonstrate that both ViT-P3DE and ViT-P3DE with JPG achieve state-of-the-art performance and significantly outperform existing works, especially when dealing with extremely harsh constraints. Xuan Truong Nguyen |
IJCAI | 3 |
| 2023 | Live Demonstration: Layer-wise Configurable CNN Accelerator with High PE UtilizationabstractWe demonstrate two end-to-end frameworks, ShortcutFusion [1] and ShortcutFusion++ [2], that effectively map many well-known deep neural networks, such as YOLO-v3, MobileNet-v2, EfficientNet-B0, and Resnet-50, to a generic CNN accelerator on FPGA. The experimental results show that ShortcutFusion++ achieves a processing element utilization of 80.95% for the well-known object detector YOLO-v3. Chunmyung Park, Eunjae Hyun, Jicheon Kim, Xuan Truong Nguyen |
ISCAS | 4 |
| 2022 | An Energy-Efficient YOLO Accelerator Optimizing Filter Switching ActivityabstractConvolutional neural network (CNN) based object detectors such as the you-only-look-once (YOLO) achieve remarkable performance but come with high computing complexity and a large memory bandwidth. Therefore, it is challenging to design an accelerator for such object detectors on edge devices, which have a limited power budget and relatively small on-chip memory footprint. In this paper, we propose an energy-and memory-efficient CNN accelerator for YOLO. First, we propose a novel dataflow which reduces the amount of filter switching by 99.56% on average. Second, we propose a layer-wise on-chip memory reuse scheme in which multi-bank on-chip buffers are efficiently utilized for both feature maps (FMs) and filters without the need to access external memory for FMs. The proposed design is implemented on a Xilinx ZC706 FPGA and consumes power of 5. 05W but achieves a throughput rate of 370.5 GOPS for Tiny-YOLOv2 while using only 640 DSPs and 322.5 BRAMs. Our design achieves energy efficiency of 73.39 GOPS/W, thus outperforming those in previous works. Kyeongjong Lim, Gyuri Kim, Taehyung Park, Xuan Truong Nguyen |
ISCAS | 4 |
| 2020 | An Efficient Sampling Algorithm With a K-NN Expanding Operator for Depth Data Acquisition in a LiDAR SystemabstractThe spatial resolution of a depth-acquisition device, such as a Light Detection and Ranging (LiDAR) sensor, is limited because of the slow acquisition. To accurately reconstruct a depth image from limited spatial resolution, a two-stage sampling process has been widely used. However, two-stage sampling uses an irregular sampling pattern for the sampling operation, which requires complex computation for reconstruction and additional memory space for storage. A mathematical formulation of a LiDAR system demonstrates that two-stage sampling does not satisfy its timing constraint for practical use. To overcome the drawbacks of two-stage sampling, this paper proposes a new sampling method that reduces the computational complexity and memory requirements by generating the optimal representatives of a sampling pattern in down-sample data. A sampling pattern can be derived from a k-NN expanding operation from the downsampled representatives. The proposed algorithm is designed to preserve the object boundary by restricting the expansionoperation only to the object boundary or complex texture. In addition, the proposed algorithm runs in linear-time complexity and reduces the memory requirements using a down-sampling ratio. The experimental results demonstrate that the proposed sampling outperforms grid sampling by at most 7.92 dB. Consequently, the proposed sampling achieves reconstructed quality similar to that of optimal sampling, while substantially reducing the computation time and memory requirements. Xuan Truong Nguyen, Hyun Kim 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2019 | A Novel FPGA Implementation of a Time-to-Digital Converter Supporting Run-Time Estimation and CompensationabstractTime-to-digital converters (TDCs) are widely used in applications that require the measurement of the time interval between events. In previous designs using a feedback loop and an extended delay line, process-voltage-temperature (PVT) variation often decreases the accuracy of measurements. To overcome the loss of accuracy caused by PVT variation, this study proposes a novel design of a synthesizable TDC that employs run-time estimation and compensation of PVT variation. A delay line consisting of a series of buffers is used to detect the period of a ring oscillator designed to measure the time interval between two events. By comparing the detected period and the system clock, the variation of the oscillation period is compensated at run-time. The proposed TDC is successfully implemented by using a low-cost Xilinx Spartan-6 LX9 FPGA with a 50-MHz oscillator. Experimental results show that the proposed TDC is robust to PVT variation with a resolution of 19.1 ps. In comparison with previous design, the proposed TDC achieves about five times better tradeoff in the area, resolution, and frequency of the reference clock. Dinh Van Luan, Xuan Truong Nguyen |
ACM Trans. Reconfigurable Technol. Syst. | 2 |
| 2018 | A New FPGA Implementation of a Time-to-Digital Converter Supporting Run-Time Estimation of Operating Condition VariationabstractA Time-to-Digital Converter (TDC) is widely used in applications that need to measure the time interval between events. Previous designs based on a feedback loop and an extended delay line suffers from poor accuracy caused by Process, Voltage, and Temperature (PVT) variations of the feedback path. This paper proposes a novel design of a synthesizable TDC that can estimate the operating event at run-time. The proposed TDC includes a ring oscillator of which oscillation period is measured at run-time to detect any change of operating event. The proposed TDC is implemented by using Xilinx Spartan-6 LX9 FPGA with 50MHz oscillator and it achieves about 19ps resolution. For 3ns time interval, the TDC detects it as 2.989ns on average with the standard deviation of about 148ps at 70°C. Dinh Van Luan, Xuan Truong Nguyen |
ISCAS | 2 |