Zexu Zhang

dblp:91/7576 · DBLP profile ↗
← Back
25ranked-venue papers
2as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 10 · 1 first-author · 7 since 2021Artificial intelligence and machine learning · 8 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 RouterAcc: FPGA Acceleration for VLSI Detailed Router via Hierarchical Storage Mapping
abstract
Detailed routing constitutes a critical phase in the very large-scale integration (VLSI) physical design, widely regarded as the most time-consuming and computationally intensive step in the back-end design process. Due to its iterative nature and strong data dependencies, conventional parallel acceleration techniques often suffer from limited scalability and effectiveness. To address these challenges, we propose RouterAcc, an FPGA-based software–hardware co-design acceleration framework tailored for VLSI detailed routing. RouterAcc incorporates an access analysis mechanism and a termination condition strategy to accelerate convergence. Furthermore, we employ a hierarchical storage mapping scheme and a flexible dimension-partitioning architecture to alleviate memory bottlenecks and enhance data locality. Additionally, RouterAcc leverages a hierarchical comparison pipeline with fully parallelized computing units and a data preprocessing strategy to maximize computational efficiency. Experimental results on the ISPD’18 benchmarks demonstrate that RouterAcc achieves consistent speedups of 2.1×–2.3× over TritonRoute with less than 1% quality degradation. With further co-optimization, RouterAcc attains speedups of 2.7×–11.8× while maintaining routing quality comparable to TritonRoute and surpassing Dr.CU 2.0 as well as the state-of-the-art (SOTA) FPGA-based approaches.
Ruiyuan Guo, Zexu Zhang, Da Tang, Weiqi Shen, Haodong Lu 0001, Xiqiong Bai, Kun Wang 0005, Jianli Chen, Jun Yu 0010
DATE2
2026 HCSL: Rumor Detection by Integrating Intra-Sample Curriculum Learning and Hierarchical Semantic Learning
Shufeng Hao, Xiaoning Hao, Zexu Zhang, Usman Naseem
WWW5
2026 Self-Chained Dynamic Context Perception to Tracking by Natural Language Specification
abstract
Vision-language cross-modal learning has significantly improved Tracking by Natural Language specification (TNL). Most existing TNL methods follow a Siamese-like matching paradigm, where visual search-region features and language-query features are aligned with the aid of pre-trained image-text representations. Although such representations provide strong static semantic cues, they are often less effective in explicitly modeling target-state changes described by action-related phrases in natural language queries. As a result, dynamic linguistic cues, such as verbs and motion-related descriptions, may be insufficiently emphasized during cross-modal matching. To address this issue, we propose Self-Chained Dynamic Context Perception (SeDCP), a self-chained framework for explicit dynamic query modulation and language-guided visual refinement in TNL. Specifically, SeDCP consists of two coupled chains. First, the Forward Chain performs visual-evidence-guided dynamic query modulation by injecting trajectory-aware spatiotemporal cues into the language representation, thereby enhancing phrases that describe target-state changes. Second, the Backward Chain uses the dynamically enhanced query representation to refine visual spatiotemporal features, strengthening the alignment between language cues and target-state evolution. In addition, we introduce sequence-level matching rather than isolated pairwise matching to better exploit temporal dynamics, and design a Global-Local enhanced video Transformer to capture both long-range contextual dependencies and fine-grained target details. Extensive experiments on seven standard TNL benchmarks and an additional unseen $\mathrm {LaSOT}_{\mathrm {ext}}$ benchmark demonstrate that SeDCP consistently outperforms state-of-the-art methods and generalizes well to unseen categories and video characteristics.
Ding Ma 0001, Zexu Zhang, Xiangqian Wu 0002
IEEE Trans. Image Process.2
2025 End-to-end Compilation is All FPGAs Need: A Unified Overlay-based FPGA Compiler for Deep Learning
abstract
Field-Programmable Gate Array (FPGA) has shown great application potential in deploying Neural Networks (NNs) due to the characteristics of programmability, low power consumption, etc. However, deploying NNs on FPGA is non-trivial because (1) Mainstream NNs pose significant FPGA architecture design challenges due to their large number of parameters, complex operations, and the need for data optimization, and (2) Supporting the deployment of different machine learning frameworks to FPGA requires significant manual effort, consuming a large amount of time. In this paper, we propose AutoCompiler, a unified compiler for mapping NNs to different FPGAs, along with overlay techniques to enable fast and efficient implementation. To the best of our knowledge, we are the first work to support both Deep Neural Networks (DNNs) and Transformer-based networks for overlay-based FPGA deployment. AutoCompiler comprises three integrated enablers: (1) Model Translator, built on top of a topology-based NNs representation, which can optimize the topology and data representation of the models from an algorithmic level based on different hardware configurations, e.g., DSP utilization, (2) Instruction Generator, which generates pipeline data streams according to various FPGA resource configurations by manipulating the instruction set at the upper level rapidly, and (3) End-to-end optimization, which moves as much of the computational processes as possible onto the FPGA chip and minimizes the interaction between CPU and FPGA. Extensive experiments on various Xilinx FPGAs show that AutoCompiler outperforms state-of-the-art overlay-based compiler by 1.2× - 1.35× and same-level GPUs by 1.15× - 1.59× for classic DNN models, and ViT inference, respectively.
Haodong Lu 0001, Yinqiu Liu, Zexu Zhang, Kun Wang 0005
ASP-DAC4
2025 A Generative Adversarial Network Method with Multi-Scale Hybrid Attention for Remote Sensing Image Super-Resolution
abstract
With the development of urbanization, remote sensing images often face the problem of insufficient resolution due to hardware and environmental limitations, making it difficult to meet the needs of urban safety governance. Existing superresolution methods hardly balance detail reconstruction and spectral fidelity in complex scenes. This paper proposes a Generative Adversarial Network (HADP-GAN) that fuses multiscale hybrid attention and dynamic divide-and-conquer strategy. First, a multi-scale pyramidal attention mechanism captures hierarchical features, while integrated channel-spatial attention amplifies critical information representation. Second, areas are partitioned according to gradient complexity, with lightweight residual, dense connection, and multi-level attention modules selectively deployed to optimize computational efficiency. Finally, gradient alignment constraints are introduced to suppress spectral distortion. Experiments show that HADP-GAN achieves an average PSNR of 36.66 dB and SSIM of 0.8992 on the GF1 satellite dataset, which are 2.92 dB to 7.78 dB and 0.064 to 0.1771 higher than Bicubic, Real-ESRGAN, and SwinIR. Moreover, it outperforms comparative methods in texture details and spectral fidelity, effectively enhancing the super-resolution reconstruction of remote sensing images in complex ground object scenes.
Zexu Zhang, Qiyong Tang, Yangzhou Long, Cuifeng Zhang, Jiawang Ge, Jie She
HPCC1
2025 LUT-HD: Accelerating Hyperdimensional Computing Inference via Efficient Table Lookup
abstract
Hyperdimensional computing (HDC) has emerged as a promising cognitive computing paradigm, offering exceptional robustness and energy efficiency for intelligent applications. However, the computational demands of HDC, particularly during the encoding and associative search phases, pose significant challenges due to their time and resource intensity. In this paper, we propose LUT-HD, a software-hardware co-design framework that accelerates HDC inference by leveraging efficient table lookup techniques. First, we introduce a binary code quantization (BCQ) algorithm based on a lookup table (LUT) that transforms costly matrix-vector multiplications in HDC into simple table lookups using precomputed results. Next, we propose a custom FPGA-based accelerator tailored for LUT-based HDC to strike a balance between accuracy and efficiency. This accelerator incorporates a performance-optimized pipeline for encoding and associative search, enhancing computational speed and resource utilization. Experimental results demonstrate that LUT-HD achieves up to 14.6 × inference speedup and reduces 97.3% energy consumption compared to the GPU platform. In addition, compared to state-of-the-art (SOTA) HDC solutions, LUT-HD offers a 5.5× speedup with negligible accuracy loss and reduces 44.8% energy consumption.
Haodong Lu 0001, Da Tang, Xiqiong Bai, Zexu Zhang, Kun Wang 0005
ICCAD4
2025 GS2Poly: Textured Polygonal Building Reconstruction Guided by Gaussian Opacity Fields
abstract
Compact low-poly building models with concise structures and texture fidelity are essential infrastructure for digital twin cities. Traditional point cloud-based reconstruction methods often rely on surface normals, and the presence of missing data and noise poses significant challenges for accurate reconstruction. In this paper, we propose GS2Poly, a textured polygonal mesh reconstruction method for buildings based on the 3D Gaussian Splatting (3DGS) framework. Firstly, 3DGS of the building scene is reconstructed under planar structure constraints. A density-weighted Gaussian sampling method is utilized to sample high-quality surface point clouds and extract planar primitives from 3DGS reconstruction results. Next, GS2Poly applies an adaptive spatial partitioning strategy to generate a set of candidate convex polyhedra. Finally, guided by the Gaussian opacity field, a Markov random field is constructed to extract the polygonal mesh surface, followed by high-fidelity texture mapping using an optimal rendering strategy. Experimental results across diverse building scenarios demonstrate that GS2Poly exhibits higher geometric fidelity than spatial partitioning-based or 3DGS-based mesh simplification methods. Additionally, the proposed texture mapping strategy effectively avoids typical texture artifacts such as occlusion, seams and distortions.
Xinyi Liu 0002, Weiwei Fan, Yongjun Zhang 0002, Zexu Zhang, Yi Wan 0001, Dongdong Yue, Jiachen Zhong
IEEE Trans. Geosci. Remote. Sens.4
2024 AutoHammer: Breaking the Compilation Wall Between Deep Neural Network and Overlay-based FPGA Accelerator
abstract
Field-Programmable Gate Array (FPGA) has shown great potential in accelerating Deep Neural Networks (DNNs) due to its characteristics of programmability and high power efficiency. In address the compilation challenges between DNNs and FPGA, we propose AutoHammer, an automated compiler for mapping DNNs to different FPGAs. Specifically, AutoHammer leverages overlay techniques to enable fast and effective implementation. Moreover, three enablers are integrated into AutoHammer. First, the Model Translator optimizes the topology and predicts a DNN's results based on different hardware configurations, built on top of a topology-based representation of DNNs. Second, the Instruction Generator generates pipeline data streams in various FPGA resource configurations by manipulating the instruction set at the upper level rapidly. Last, we realize the End-to-end Optimization, moving the whole computational processes onto the FPGA. Extensive experimental results show that AutoHammer improves great deployment efficiency when validated by 14 types of DNN models on 3 companies' (Xilinx, Fudan Micro, and Pango Micro) mainstream FPGA chips.
Yinqiu Liu, Haodong Lu 0001, Zexu Zhang, Ruiqiu Chen, Kun Wang 0005
FPGA5
2024 TransLib: An Extensible Graph-Aware Library Framework for Automated Generation of Transformer Operators on FPGA
Yang Liu 0376, Zexu Zhang, Jun Yu 0010, Kun Wang 0005
ICCAD4
2023 Efficient Implementation of Activation Function on FPGA for Accelerating Neural Networks
abstract
In this paper, we present the Integer Lightweight Softmax (ILS) algorithm for approximating the Softmax activation function. The accurate implementation of Softmax on FPGA can be huge resource-intensive and memory-hungry. Then, we present the implementation of ILS on a Xilinx XCKU040 FPGA to evaluate the effectiveness of ILS. Evaluations on CIFAR 10, CIFAR 100 and ImageNet show that ILS achieves up to$2.47\times, 40\times$and$323\times$speedup over CPU implementation, and$4\times, 63\times$and$51\times$speedup over GPU implementation, respectively. In comparison to previous FPGA-based Softmax implementations, ILS strikes a better balance between resource consumption and precision accuracy.
Yinqiu Liu, Zexu Zhang, Kun Wang 0005
ISCAS3
2023 Joint Dual-UAV Trajectory and RIS Design for ARIS-Assisted Aerial Computing in IoT
abstract
Reconfigurable intelligent surface (RIS), as an emerging technology, has recently been applied to expand the range of mobile-edge computing (MEC) networks and improve wireless environments. However, current terrestrial RIS-assisted MEC networks have some limitations, such as severe signal attenuation and inflexible equipment deployment. To take full advantage of the superiority of the RIS, this article considers an aerial RIS (ARIS)-assisted aerial computing scheme, where the ARIS and the other unmanned aerial vehicle (UAV) equipped with a MEC server are employed to facilitate offloading computing tasks from Internet of Things (IoT) user equipments (UEs) to the access point (AP). With the flexibility of the dual-UAV, we can mitigate the Non-Line-of-Sight (NLoS) air–ground paths caused by obstacles. In the proposed scenario, to improve the system energy efficiency while ensuring the UEs receive high-quality wireless services, we intend to jointly optimize the trajectories of the two UAVs, the phase shift of the ARIS, the computation offloading strategy, and computation resource allocation. The issue is formulated as a mixed nonconvex optimization problem, so it is difficult to solve it in time for adapting different environments by using conventional convex optimization methods. However, we develop a double deep$Q$-network (DDQN)-based algorithm to obtain the near-optimal online decision-making solution. Simulation findings indicate that the proposed DDQN-based algorithm can effectively increase the energy efficiency of the proposed dual-UAV cooperative MEC system in comparison to the benchmark schemes.
Bin Duo, Maolin He, Qingqing Wu 0001, Zexu Zhang
IEEE Internet Things J.4
2023 Probabilistic Network Topology Prediction for Active Planning: An Adaptive Algorithm and Application
abstract
This article tackles the problem of active planning to achieve cooperative localization for multirobot systems under measurement uncertainty in GNSS-limited scenarios. Specifically, we address the issue of accurately predicting the probability of a future connection between two robots equipped with range-based measurement devices. Due to the limited range of the equipped sensors, edges in the network connection topology will be created or destroyed as the robots move with respect to one another. Accurately predicting the future existence of an edge, given imperfect state estimation and noisy actuation, is therefore a challenging task. An adaptive power series expansion (or APSE) algorithm is developed based on current estimates and control candidates. Such an algorithm applies the power series expansion formula of the quadratic positive form in a normal distribution. Finite-term approximation is made to realize the computational tractability. Further analyses are presented to show that the truncation error in the finite-term approximation can be theoretically reduced to a desired threshold by adaptively choosing the summation degree of the power series. Several sufficient conditions are rigorously derived as the selection principles. Finally, extensive simulation results and comparisons, with respect to both single and multirobot cases, validate that a formally computed and therefore more accurate probability of future topology can help improve the performance of active planning under uncertainty.
Zexu Zhang, Roland Siegwart, Jen Jen Chung
IEEE Trans. Robotics2
2023 Robust Optimal Control of Uncertain Discrete-Time Multiagent Systems With Digraphs
abstract
This article studies the distributed robust optimal control for discrete-time linear multiagent systems (MASs) with parametric uncertainties, where digraphs that only contain a directed spanning tree are allowed. Using the linear quadratic regulator approach, an optimal control protocol is presented. The presented controller is fully distributed, since the global information of graphs is unneeded for the design and implementation of the presented controller. The global performance index of MASs can be minimized by using the presented control protocol, and the optimal solution is independent with the information of parametric uncertainties. Finally, some simulated examples are provided to show the effectiveness of the proposed approaches.
Zhuo Zhang 0006, Yang Shi 0001, Zexu Zhang, Shouxu Zhang, Huiping Li 0003, Bing Xiao 0001, Weisheng Yan
IEEE Trans. Syst. Man Cybern. Syst.3
2022 Robust Cooperative Optimal Sliding-Mode Control for High-Order Nonlinear Systems: Directed Topologies
abstract
This article is concerned with the robust cooperative optimal control of nonlinear multiagent systems (MASs) with external disturbances and modeling uncertainties. Using the super-twisting algorithm, a continuous sliding-mode control protocol is presented for high-order nonlinear MASs with multiple inputs. The sliding-mode dynamics is modeled by the Takagi-Sugeno fuzzy approach, and the nominal control protocol that guarantees the robust optimization of the cost function is designed. Directed topologies are allowed using the presented protocol, and many assumptions about topologies are removed. Finally, three numerical examples are reported to demonstrate the effectiveness and improved performance of the presented protocol.
Zhuo Zhang 0006, Yang Shi 0001, Shouxu Zhang, Zexu Zhang, Weisheng Yan
IEEE Trans. Cybern.4
2020 A Connectivity-Prediction Algorithm and its Application in Active Cooperative Localization for Multi-Robot Systems
abstract
This paper presents a method for predicting the probability of future connectivity between mobile robots with range-limited communication. In particular, we focus on its application to active motion planning for cooperative localization (CL). The probability of connection is modeled by the distribution of quadratic forms in random normal variables and is computed by the infinite power series expansion theorem. A finite-term approximation is made to realize the computational feasibility and three more modifications are designed to handle the adverse impacts introduced by the omission of the higher order series terms. On the basis of this algorithm, an active and CL problem with leader-follower architecture is then reformulated into a Markov Decision Process (MDP) with a one-step planning horizon, and the optimal motion strategy is generated by minimizing the expected cost of the MDP. Extensive simulations and comparisons are presented to show the effectiveness and efficiency of both the proposed prediction algorithm and the MDP model.
Zexu Zhang, Roland Siegwart, Jen Jen Chung
ICRA2
2019 New Results on Sliding-Mode Control for Takagi-Sugeno Fuzzy Multiagent Systems
abstract
This paper investigates the sliding-mode control (SMC) problem of Takagi-Sugeno (T-S) fuzzy multiagent systems (MASs). A cooperative fuzzy-based dynamical sliding-mode (SM) controller is designed and the overall closed-loop T-S fuzzy MAS is constructed. A new model transformation method for T-S fuzzy MASs is presented to transform the fuzzy weighting matrix into a set of fuzzy weighting scalars. By applying the method of linear matrix inequality, a general stability analysis approach for T-S fuzzy MASs is proposed. Moreover, the energy-cost constraint problem is studied by using the linear quadratic regulator method. Finally, numerical examples are provided to illustrate the effectiveness of the proposed theoretical approaches and the improved performance compared to existing results.
Zhuo Zhang 0006, Yang Shi 0001, Zexu Zhang, Weisheng Yan
IEEE Trans. Cybern.3
2017 Dynamical sliding-mode control of leader-following multi-agent systems
abstract
This paper studies the leader-following problem of general linear multi-agent systems (MASs). A leader-following model is constructed, and then the sliding-mode control (SMC) technique is utilized to guarantee that followers can track the leader. In order to eliminate the chattering, the sliding-mode variable is designed as a dynamical one, that is, the sign function exist in the derivative of the variable. Moreover, the energy-cost constraint problem is considered as well. Finally, the multi-spacecraft tracking problem is taken as an example to demonstrate the effectiveness of the theoretical approach.
Zhuo Zhang 0006, Zexu Zhang
IECON2
2017 Surrounding control in cooperative second-order agent networks
abstract
In this paper, a surrounding control problem for second-order agent system is investigated. A set of followers composed by the second-order agent networks is applied to surround a team of static leaders into a convex hull spanned by these followers. The problem is considered under a decentralized estimation-and-control framework by using tools from algebraic graph theory and Lyapunov stable theory. Under some basic assumptions, the surrounding problem is solved with the followers converge to their desired position with zero velocities and control inputs, even if the geometric center of the leaders can only be obtained from estimators. Simulation results are provided to illustrate the effectiveness of proposed methods.
Zexu Zhang, Yandi Qiao, Xiangquan Wei
IECON2
2015 Decentralized robust attitude tracking control for spacecraft networks under unknown inertia matrices
Zhuo Zhang 0006, Zexu Zhang, Hui Zhang 0019
Neurocomputing2
2015 Robust reduced-order l2-l∞ filtering for network-based discrete-time linear systems
Zhuo Zhang 0006, Zexu Zhang
Signal Process.2
2015 Finite-Time H∞ Filtering for T-S Fuzzy Discrete-Time Systems With Time-Varying Delay and Norm-Bounded Uncertainties
abstract
In this paper, we investigate the filtering problem of discrete-time Takagi-Sugeno (T-S) fuzzy uncertain systems subject to time-varying delays. A reduced-order filter is designed. With the augmentation technique, a filtering error system with delayed states is obtained. In order to deal with time delays in system states, the filtering error system is first transformed into two interconnected subsystems. By using a two-term approximation for the time-varying delay, sufficient delay-dependent conditions of finite-time boundedness and H∞performance of the filtering error system are derived with the Lyapunov function. Based on these conditions, the filter design methods are proposed and the filter gain matrices can be obtained by calculating a set of linear matrix inequalities. A numerical example is used to illustrate the effectiveness of the proposed approaches.
Zhuo Zhang 0006, Zexu Zhang, Hui Zhang 0019, Peng Shi 0001, Hamid Reza Karimi
IEEE Trans. Fuzzy Syst.2
2013 On mode-dependent H∞ filtering for network-based discrete-time systems
Lin Li 0037, Zexu Zhang, Jingcheng Xu
Signal Process.3
2010 Exponential stability on stochastic neural networks with discrete interval and distributed delays
abstract
This brief addresses the stability analysis problem for stochastic neural networks (SNNs) with discrete interval and distributed time-varying delays. The interval time-varying delay is assumed to satisfy 01¿ d(t) ¿ d2and is described asd(t) =d1+h(t) with 0 ¿h(t) ¿d2-d1. Based on the idea of partitioning the lower boundd1, new delay-dependent stability criteria are presented by constructing a novel Lyapunov-Krasovskii functional, which can guarantee the new stability conditions to be less conservative than those in the literature. The obtained results are formulated in the form of linear matrix inequalities (LMIs). Numerical examples are provided to illustrate the effectiveness and less conservatism of the developed results.
Rongni Yang, Zexu Zhang, Peng Shi 0001
IEEE Trans. Neural Networks2
2009 Stochastic stability of Markovian jumping Hopfield neural networks with constant and distributed delays
Lin Zhao 0009, Zexu Zhang, Yan Ou
Neurocomputing3
2009 New passivity criteria for neural networks with time-varying delay
Zexu Zhang, Shaoshuai Mou, James Lam, Huijun Gao
Neural Networks1