Dong Xu 0015

dblp:09/3493-15 · DBLP profile ↗
← Back
5ranked-venue papers
0as first author
5since 2021 · last 2023
0009-0008-8307-3972ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 4 · 4 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
YearPublicationVenuePosition
2023 Graph Representation Learning for Microarchitecture Design Space Exploration
abstract
Design optimization of modern microprocessors is a complex task due to the exponential growth of the design space. This work presents GRL-DSE, an automatic microarchitecture search framework based on graph embeddings. GRL-DSE uses graph representation learning to build a compact and continuous embedding space. Multi-objective Bayesian optimization using an ensemble surrogate model conducts microarchitecture design space exploration in the graph embedding space to efficiently and holistically optimize performance-power-area (PPA) objectives. Experimental studies on RISC-V BOOM show that GRLDSE outperforms previous techniques by 74.59% on Pareto front quality and outperforms manual designs in terms of PPA.
Xiaoling Yi, Jialin Lu, Xiankui Xiong, Dong Xu 0015, Fan Yang 0001
DAC4
2023 TPNoC: An Efficient Topology Reconfigurable NoC Generator
abstract
With the core count increasing in Chip to support various data-intensive workloads, Network-on-chip (NoC) has become the better solution for addressing on-chip interconnection. Various data-intensive workloads have different traffic patterns that require NoC with different topologies and microarchitectures. On the one hand, topology type selection has a great influence on the final performance, area, and energy. However, it is difficult to change the topology type in the traditional NoC RTL design process once it is determined. On the other hand, NoC platforms have many tunable micro-architecture design parameters, which require careful design space exploration to trade off performance advantages and overhead. Designing and validating each microarchitecture of NoCs to account for various trade-offs will greatly exacerbate the design cost issue.
Jiangnan Yu, Fan Yang 0001, Xiaoling Yi, Chixiao Chen, Jun Tao 0001, Dong Xu 0015, Xiankui Xiong
ACM Great Lakes Symposium on VLSI6
2023 Luminance-Preserving Visible and Near-Infrared Image Fusion Network with Edge Guidance
abstract
Near-infrared (NIR) images and visible (VIS) images can provide mutually complementary information for each other, thus the fusion of the two modalities can create images of high quality even in adverse conditions. However, the luminance of NIR and VIS images may be inconsistent in some regions, resulting in color distortion and unrealistic appearance in the fused images. The existing methods perform poorly at luminance retention. Aiming at the problem and based on deep learning framework, we propose an edge-guided method which can be applied to the image fusion network. Edge maps are utilized as prior knowledge of images to boost the performance of the neural network. Additionally, we propose a luminance-preserving loss function combined with max-edge loss to further improve the image quality. Experimental results show the superiority of our method.
Ruoxi Zhu, Yi Ling, Xiankui Xiong, Dong Xu 0015, Xuanpeng Zhu, Yibo Fan
ICIP4
2022 A 11.6μ W Computing-on-Memory-Boundary Keyword Spotting Processor with Joint MFCC-CNN Ternary Quantization
abstract
This paper presents an ultra-low-power keyword spotting processor using an algorithm-architecture co-design approach. Joint MFCC-CNN ternary weight quantization is proposed to reduce power consumption. The Mel filter and the DCT module are merged into one matrix multiplication. The merged coefficients and the weights of the rest NN classifier are ternary-quantized, causing less than 3% accuracy loss but 39 × energy efficiency improvement. Moreover, a Computing-on-Memory-Boundary macro is adopted to store the quantized coefficients and weights, and perform matrix multiplications. Compared to the existing computing-in-memory technology, the proposed technique can reduce power consumption due to the higher utilization ratio. To verify the proposed techniques, a keyword spotting processor prototype is designed with 28nm CMOS technology. Simulation results show that the prototype achieves power consumption of 11.6μ W under a power supply of 0.72V and a clock frequency of 250KHz.
Xinru Jia, Haozhe Zhu, Yunzheng Wang, Jinshan Zhang 0006, Xiankui Xiong, Dong Xu 0015, Chixiao Chen, Qi Liu 0010
ISCAS7
2022 NNASIM: An Efficient Event-Driven Simulator for DNN Accelerators with Accurate Timing and Area Models
abstract
In this paper, we propose NNASIM, an efficient timing and area accurate event-driven simulator for custom DNN accelerators. NNASIM is a highly-modular and highly parameterized modeling framework. We build accurate timing and area models for common accelerator modules like GEMM, ALU array, and crossbar using ASIC synthesis flows. These models are fed into the event-driven simulator for fast simulation. NNASIM is integrated with a RISC-V simulator. This approach guarantees the functional correctness of the accelerator simulation at the instruction level. The experimental results show that our model evaluates the performance and area of DNN accelerators with less than 0.76% and 2.83% error, respectively, compared to RTL implementations. NNASIM allows designers to model the performance and area of the accelerator at a high level, and thus enables the systematic microarchitecture design space exploration of the custom accelerators. Index Terms accelerators.
Xiaoling Yi, Jiangnan Yu, Xiankui Xiong, Dong Xu 0015, Chixiao Chen, Jun Tao 0001, Fan Yang 0001
ISCAS5