Dongjun Xu

dblp:149/3936 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
3since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 6 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 A High-Accuracy and Low-Resource Design for Nonlinear Functions in AI Accelerators
abstract
In deep learning models, the increasing complexity of activation functions makes hardware optimization crucial. This paper proposes an efficient computation framework for 4 elementary functions (exponential, logarithmic, reciprocal and square root). By leveraging the floating-point representation and employing piecewise linear (PWL) approximation for the mantissa, our method substantially reduces multiplication operations and storage overhead. Our framework supports the execution of prevalent activation functions through combinations of these functions, further enhancing its versatility. The proposed design is validated on Xilinx Zynq X7CZ020 FPGA. Specifically, the FP16 4-segment pipelined module achieves a mean relative error of 2.26e-3 across all valid input ranges, but only utilizes 725 LUTs and 129 FFs without using DSP. It reaches a maximum working frequency of 163.64 MHz, and requires a 4 clock cycles latency, demonstrating significant advantages in resource utilization and computation efficiency.
Dongjun Xu
ISCAS4
2024 Study on Heat Transfer and Flow Characteristics of Supercritical CO2 of SLM Additive Manufacturing Heat Exchangers with Mini-Channels
abstract
In recent years, additive manufacturing has gained significant attention in the production of high-temperature and high-pressure mini-channel heat exchangers. Among the various techniques, selective laser melting (SLM) is widely utilized in metal additive manufacturing. This study experimentally investigates the flow and heat transfer performance of rectangular and circular mini-channel heat exchangers produced by SLM. The experiments were conducted under pressure conditions ranging from 8 to 10 MPa and temperatures between 200 and 310°C. Heat transfer correlations for supercritical carbon dioxide (SCO2) in both rectangular and circular mini-channels were refined. Additionally, flow characteristics without heat transfer were analyzed, revealing that the Darcy friction factor in both types of mini-channels initially increases with the Reynolds number (Re), followed by a decrease, and stabilizes at higher Re values.
Dechao Liu, Dongjun Xu, Qiyuan Ma
INDIN3
2024 Spiking-Hybrid-YOLO for Low-Latency Object Detection
abstract
Compared to conventional deep neural networks (DNNs), spiking neural networks (SNNs) have gained significant attention in recent years owing to their low power consumption, superior inference efficiency, and high biological plausibility. Although SNNs have demonstrated success in simple classification tasks, they are not frequently employed in complex regression tasks such as object detection. During the inference process, the multi-time-step nature of SNNs can lead to multiple storage accesses and computational operations, thereby increasing the power consumption and inference latency. In this study, we introduce Spiking-HybridYOLO, a hybrid architecture designed for low-time-step object detection. Distinct from the conventional DNN-to-SNN conversion approaches, Spiking-Hybrid-YOLO can be trained directly using surrogate gradient methods. After examining the computational complexity of each layer in YOLO, we selectively replace certain ANN layers with spiking layers to strike a balance between execution time, energy consumption, and accuracy. Experimental results show that Spiking-HybridYOLO achieves low-power object detection in a single timestep, consuming only 2.77% of the power used by ANNs.
Mingxin Guo, Dongjun Xu, Liang Chen 0003
ISCAS2
2016 A Q-Learning Based Self-Adaptive I/O Communication for 2.5D Integrated Many-Core Microprocessor and Memory
abstract
A self-adaptive output-voltage swing adjustment is introduced in the design of energy-efficient I/O communication for 2.5D integrated many-core microprocessor and memory. Instead of transmitting signal with large voltage swing, a Q-learning based I/O management is deployed to adaptively adjust the I/O output-voltage swing under constraints of both communication power and bit error rate (BER). Simulation results show that the proposed adaptive 2.5D I/Os (in 65 nm CMOS) can achieve an average of 12.5 mW I/O power, 4 GHz bandwidth and 3.125 pJ/bit energy efficiency for one channel under 10-6BER. With the use of conventional Q-learning and further accelerated Q-learning, we can achieve 12.95 and 18.89 percent power reduction and 14 and 15.11 percent energy efficiency improvement when compared to the use of uniform output-voltage swing based I/O communication.
Sai Manoj Pudukotai Dinakarrao, Hao Yu 0001, Hantao Huang, Dongjun Xu
IEEE Trans. Computers4
2014 Reinforcement learning based self-adaptive voltage-swing adjustment of 2.5D I/Os for many-core microprocessor and memory communication
abstract
A reinforcement learning based I/O management is developed for energy-efficient communication between many-core microprocessor and memory. Instead of transmitting data under a fixed large voltage-swing, an online reinforcement Q-learning algorithm is developed to perform a self-adaptive voltage-swing control of 2.5D through-silicon interposer (TSI) I/O circuits. Such a voltage-swing adjustment is formulated as a Markov decision process (MDP) problem solved by model-free reinforcement learning under constraints of both power budget and bit-error-rate (BER). Experimental results show that the adaptive 2.5D TSI I/Os designed in 65nm CMOS can achieve an average of 12.5mw I/O power, 4GHz bandwidth and 3.125pJ/bit energy efficiency for one channel under 10-6BER, which has 18.89% power saving and 15.11% improvement of energy efficiency on average.
Hantao Huang, Sai Manoj Pudukotai Dinakarrao, Dongjun Xu, Hao Yu 0001, Zhigang Hao
ICCAD3
2014 An energy-efficient 2.5D through-silicon interposer I/O with self-adaptive adjustment of output-voltage swing
abstract
A self-adaptive output swing adjustment is introduced for the design of energy-efficient 2.5D through-silicon interposer (TSI) I/Os. Instead of transmitting signal with large voltage swing, Q-learning based self-adaptive adjustment is deployed to adjust I/O output-voltage swing under constraints of both power budget and bit error rate (BER). Experimental results show that the adaptive 2.5D TSI I/Os designed in 65nm CMOS can achieve an average of 13mW I/O power, 4GHz bandwidth and 3.25pJ/bit energy efficiency for one channel under 10-6 BER, which has ~21.42% reduction of power and ~14.47% energy efficiency improvement.
Dongjun Xu, Sai Manoj Pudukotai Dinakarrao, Hantao Huang, Ningmei Yu, Hao Yu 0001
ISLPED1