VLDB 2026 Research / reviewers in the wild / expert
Wei Zuo
dblp:13/6778
· DBLP profile ↗
22ranked-venue papers
10as first author
5since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 10 · 3 first-author · 1 since 2021Artificial intelligence and machine learning · 6 · 2 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-authorTheory of computation · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Hamiltonian cycle and path of folded hypercube with faulty matchings
Wei Zuo, Huazhong Lü |
Discret. Appl. Math. | 1 |
| 2025 | MikuDance: Animating Character Art With Mixed Motion DynamicsabstractWe propose MikuDance, a diffusion-based pipeline incorporating mixed motion dynamics to animate stylized character art. MikuDance consists of two key techniques: Mixed Motion Modeling and Mixed-Control Diffusion, to address the challenges of high-dynamic motion and reference-guidance misalignment in character art animation. Specifically, a Scene Motion Tracking strategy is presented to explicitly model the dynamic camera in pixel-wise space, enabling unified character-scene motion modeling. Building on this, the Mixed-Control Diffusion implicitly aligns the scale and body shape of diverse characters with motion guidance, allowing flexible control of local character motion. Subsequently, a Motion-Adaptive Normalization module is incorporated to effectively inject global scene motion, paving the way for comprehensive character art animation. Through extensive experiments, we demonstrate the effectiveness and generalizability of MikuDance across various character art and motion guidance, consistently producing high-quality animations with remarkable motion dynamics. Xianfang Zeng, Xin Chen 0040, Wei Zuo, Gang Yu 0002, Zhigang Tu 0001 |
ICCV | 4 |
| 2022 | Pan-location mapping and localization for the in-situ science exploration of Zhurong Mars rover
Xingguo Zeng, Xingye Gao, Wangli Chen, Wei Zuo |
Sci. China Inf. Sci. | 8 |
| 2022 | DML: Dynamic Partial Reconfiguration With Scalable Task Scheduling for Multi-Applications on FPGAsabstractFor several new applications, FPGA-based computation has shown better latency and energy efficiency compared to CPU or GPU-based solutions. We note two clear trends in FPGA-based computing. On the edge, the complexity of applications is increasing, requiring more resources than possible on today's edge FPGAs. In contrast, in the data center, FPGA sizes have increased to the point where multiple applications must be mapped to fully utilize the programmable fabric. While these limitations affect two separate domains, they both can be dealt with by using dynamic partial reconfiguration (DPR). Thus, there is a renewed interest to deploy DPR for FPGA-based hardware. In this work, we presentDoing More with Less (DML)– a methodology for scheduling heterogeneous tasks across an FPGA's resources in a resource efficient manner while effectively hiding the latency of DPR. With the help of an integer linear programming (ILP) based scheduler, we demonstrate the mapping of diverse computational workloads in both cloud and edge-like scenarios. Our novel contributions include: enabling IP-level pipelining and parallelization to exploit the parallelism available within batches of work in our scheduler, and strategies to map and run multiple applications simultaneously. We consider the application of our methodology on real world benchmarks on both small (a Zedboard) and large (a ZCU106) FPGAs, across different workload batching and multiple-application scenarios. Our evaluation proves the real world efficacy of our solution, and we demonstrate an average speedup of 5X and up to 7.65X on a ZCU106 over a bulk-batching baseline via our scheduling strategies. We also demonstrate the scalablity of our scheduler by simultaneously mapping multiple applications to a single FPGA, and explore different approaches to sharing FPGA resources between applications. Ashutosh Dhar, Edward Richter, Mang Yu, Wei Zuo, Xiaohao Wang, Nam Sung Kim, Deming Chen |
IEEE Trans. Computers | 4 |
| 2021 | Reconstruct Anomaly to Normal: Adversarially Learned and Latent Vector-Constrained Autoencoder for Time-Series Anomaly Detection
Chunkai Zhang, Wei Zuo, Shaocong Li, Xuan Wang 0002, Peiyi Han, Chuanyi Liu |
PRICAI (2) | 2 |
| 2017 | Accurate High-level Modeling and Automated Hardware/Software Co-design for Effective SoC Design Space ExplorationabstractA desirable feature of a development tool for SoC design is that, given the important applications in the domain to be targeted by the SoC, a powerful hardware-software partitioning engine is available to determine which function(s) shall be mapped to hardware. However, to provide high-quality partitioning, this engine must be able to consider a rich design space of possible alternate hardware and software implementations for each program region candidate for hardware acceleration, in turn making the task of finding the optimal mapping very difficult given the number of design points to consider and the need for accurate modeling of latency, power and area. Wei Zuo, Louis-Noël Pouchet, Andrey Ayupov, Chung-Wei Lin, Shinichi Shiraishi, Deming Chen |
DAC | 1 |
| 2017 | Machine learning on FPGAs to face the IoT revolutionabstractFPGAs have been rapidly adopted for acceleration of Deep Neural Networks (DNNs) with improved latency and energy efficiency compared to CPU and GPU-based implementations. High-level synthesis (HLS) is an effective design flow for DNNs due to improved productivity, debugging, and design space exploration ability. However, optimizing large neural networks under resource constraints for FPGAs is still a key challenge. In this paper, we present a series of effective design techniques for implementing DNNs on FPGAs with high performance and energy efficiency. These include the use of configurable DNN IPs, performance and resource modeling, resource allocation across DNN layers, and DNN reduction and re-training. We showcase several design solutions including Long-term Recurrent Convolution Network (LRCN) for video captioning, Inception module for FaceNet face recognition, as well as Long Short-Term Memory (LSTM) for sound recognition. These and other similar DNN solutions are ideal implementations to be deployed in vision or sound based IoT applications. Xiaofan Zhang 0001, Anand Ramachandran 0001, Chuanhao Zhuge, Di He 0004, Wei Zuo, Zuofu Cheng, Kyle Rupnow, Deming Chen |
ICCAD | 5 |
| 2017 | Machine learning on FPGAs to face the IoT revolutionabstractFPGAs have been rapidly adopted for acceleration of Deep Neural Networks (DNNs) with improved latency and energy efficiency compared to CPU and GPU-based implementations. High-level synthesis (HLS) is an effective design flow for DNNs due to improved productivity, debugging, and design space exploration ability. However, optimizing large neural networks under resource constraints for FPGAs is still a key challenge. In this paper, we present a series of effective design techniques for implementing DNNs on FPGAs with high performance and energy efficiency. These include the use of configurable DNN IPs, performance and resource modeling, resource allocation across DNN layers, and DNN reduction and re-training. We showcase several design solutions including Long-term Recurrent Convolution Network (LRCN) for video captioning, Inception module for FaceNet face recognition, as well as Long Short-Term Memory (LSTM) for sound recognition. These and other similar DNN solutions are ideal implementations to be deployed in vision or sound based IoT applications. Xiaofan Zhang 0001, Anand Ramachandran 0001, Chuanhao Zhuge, Di He 0004, Wei Zuo, Zuofu Cheng, Kyle Rupnow, Deming Chen |
ICCAD | 5 |
| 2017 | New advances of high-level synthesis for efficient and reliable hardware design
Keith A. Campbell, Wei Zuo, Deming Chen |
Integr. | 2 |
| 2016 | Designing high-quality hardware on a development effort budget: A study of the current state of high-level synthesisabstractHigh-level synthesis (HLS) promises high-quality hardware with minimal development effort. In this paper, we evaluate the current state-of-the-art in HLS and design techniques based on software references and architecture references. We present a software reference study developing a JPEG encoder from pre-existing software, and an architecture reference study developing an AES block encryption module from scratch in SystemC and SystemVerilog based on a desired architecture. Additionally, we develop micro-benchmarks to demonstrate best-practices in C coding styles that produce high-quality hardware with minimal development effort. Finally, we suggest language, tool, and methodology improvements to improve upon the current state-of-the-art in HLS. Zelei Sun, Keith A. Campbell, Wei Zuo, Kyle Rupnow, Swathi T. Gurumani, Frederic Doucet, Deming Chen |
ASP-DAC | 3 |
| 2016 | Parallel code-specific CPU simulation with dynamic phase convergence modeling for HW/SW co-designabstractWhile SystemC models provide a promising solution to the complex problem of HW/SW co-design within the system-on-chip paradigm, such requires a detailed annotation of transaction level energy and performance data within the model. While this data can be obtained through source code profiling of an application running on the target processor, accomplishing such when the target CPU hardware is not actively available typically requires time-consuming CPU simulation, which is often too slow to practically consider for large programs. Additionally, while the use of SystemC modeling with TLM 2.0 standard is widely adopted for the SoC modeling, the process of transforming C/C++ code to SystemC code with TLM 2.0 functionality remains non-trivial. Herein we propose an automated framework that: 1. Enables high speed code-specific CPU profiling support for both Sniper and gem5 using parallelized dynamic steady state phase convergence modeling, providing automatic annotation of energy and latency within source code. 2. Provides an automated C to SystemC TLM 2.0 code generation flow that utilizes the back-annotated source code to produce a SystemC module for seamless incorporation into the virtual prototype. Maximum speedups obtained using Sniper and gem5 are 105.78× and 562× respectively, while average results obtained speedups of 42.7× and 323.1×. Sniper results maintain an average accuracy of 0.64% for latency and 0.10% for energy, while gem5 achieves average accuracies of 4.16% and 2.87% for latency and energy respectively. Warren Kemmerer, Wei Zuo, Deming Chen |
ICCAD | 2 |
| 2015 | A Polyhedral-based SystemC Modeling and Generation Framework for Effective Low-power Design Space ExplorationabstractWith the prevalence of System-on-Chips there is a growing need for automation and acceleration of the design process. A classical approach is to take a C/C++ specification of the application, convert it to a SystemC (or equivalent) description of hardware implementing this application, and perform successive refinement of the description to improve various design metrics. In this work, we present an automated SystemC generation and design space exploration flow alleviating several productivity and design time issues encountered in the current design process. We first automatically convert a subset of C/C++, namely affine program regions, into a full SystemC description through polyhedral model-based techniques while performing powerful data locality and parallelism transformations. We then leverage key properties of affine computations to design a fast and accurate latency and power characterization flow. Using this flow, we build analytical models of power and performance that can effectively prune away a large amount of inferior design points very fast and generate Pareto-optimal solution points. Experimental results show that (1) our SystemC models can evaluate system performance and power that is only 0.57% and 5.04% away from gate-level evaluation results, respectively; (2) our latency and power analytical models are 3.24% and 5.31% away from the actual Pareto points generated by SystemC simulation, with 2091x faster design-space exploration time on average. The generated Pareto-optimal points provide effective low-power design solutions given different latency constraints. Wei Zuo, Warren Kemmerer, Jong Bin Lim, Louis-Noël Pouchet, Andrey Ayupov, Kyungtae Han, Deming Chen |
ICCAD | 1 |
| 2014 | Change-centric Model for Web Service EvolutionabstractWeb service is subject to frequent changes during its lifecycle. Web service evolution is a widely discussed topic. Many related problems have also been generated from Web service evolution such as Web service adaptation, Web service versioning and Web service change management. To treat with these issues efficiently, a complete evolution model for Web service should be built. In this paper, we introduce our change-centric model for Web service evolution and how we use it to design, execute, and adapt to the changes during Web service evolution. Wei Zuo, Aïcha-Nabila Benharkat, Youssef Amghar |
ICWS | 1 |
| 2014 | Holistic and Change-centric Model for Web Service EvolutionabstractUnder the constantly evolving requirements from the consumers and competition pressure from the peers, Web Service providers are always striving to improve their services through publishing new versions. As more enterprises choose to embrace SOA, the frequent updates of Web services and increasing distributed environments have resulted in major challenges for all stakeholders to address the evolution of the Web service As a result, lots of solutions have been proposed to deal with the issues caused by Web Service evolution such as models, monitor, analysis, versioning, adaptation, and execution. However, few of them concentrate on the solution that covers all the evolution-related issue under one holistic model which explains 1) what has been changed, 2) when the changes occur, 3) how to apply changes, and 4) how to react to the changes. In this article, we present a change-centric model for Web Service evolution and explain how it deals with the evolution-related issues. Wei Zuo, Aïcha-Nabila Benharkat, Youssef Amghar |
SERVICES | 1 |
| 2013 | Improving high level synthesis optimization opportunity through polyhedral transformationsabstractHigh level synthesis (HLS) is an important enabling technology for the adoption of hardware accelerator technologies. It promises the performance and energy efficiency of hardware designs with a lower barrier to entry in design expertise, and shorter design time. State-of-the-art high level synthesis now includes a wide variety of powerful optimizations that implement efficient hardware. These optimizations can implement some of the most important features generally performed in manual designs including parallel hardware units, pipelining of execution both within a hardware unit and between units, and fine-grained data communication. We may generally classify the optimizations as those that optimize hardware implementation within a code block (intra-block) and those that optimize communication and pipelining between code blocks (inter-block). However, both optimizations are in practice difficult to apply. Real-world applications contain data-dependent blocks of code and communicate through complex data access patterns. Existing high level synthesis tools cannot apply these powerful optimizations unless the code is inherently compatible, severely limiting the optimization opportunity. In this paper we present an integrated framework to model and enable both intra- and inter-block optimizations. This integrated technique substantially improves the opportunity to use the powerful HLS optimizations that implement parallelism, pipelining, and fine-grained communication. Our polyhedral model-based technique systematically defines a set of data access patterns, identifies effective data access patterns, and performs the loop transformations to enable the intra- and inter-block optimizations. Our framework automatically explores transformation options, performs code transformations, and inserts the appropriate HLS directives to implement the HLS optimizations. Furthermore, our framework can automatically generate the optimized communication blocks for fine-grained communication between hardware blocks. Experimental evaluation demonstrates that we can achieve an average of 6.04X speedup over the high level synthesis solution without our transformations to enable intra- and inter-block optimizations. Wei Zuo, Yun Liang 0001, Peng Li 0031, Kyle Rupnow, Deming Chen, Jason Cong |
FPGA | 1 |
| 2012 | A New Multimodal Particle Swarm Optimization Algorithm Based on Greedy AlgorithmabstractA new multimodal particle swarm optimization algorithm based on greedy algorithm (GAPSO) is proposed in this paper. A ring topology is employed in each niche instead of a fully connected topology in GAPSO. More importantly, GAPSO does not depend on any niching parameter, which improves its effectiveness and efficiency greatly. A range of standard test functions are employed to test this proposed algorithm's performance. Experimental results demonstrate that GAPSO performs as well as existed optimization algorithms, and needs smaller number of evaluations to locate all the global optima for simple problems. For complex problems, GAPSO can achieve better and more consistent performance, and can locate more global optima than SPSO, ANPSO and rpso, which makes it possible to be widely used in the real world. Yu Liu 0035, Mingwei Lv, Wei Zuo |
Int. J. Comput. Intell. Appl. | 3 |
| 2010 | A New Iterative Learning Controller Using Variable Structure Fourier Neural NetworkabstractA new iterative learning control approach based on Fourier neural network (FNN) is presented for the tracking control of a class of nonlinear systems with deterministic uncertainties. The proposed controller consists of two loops. The inner loop is a feedback control action that decreases system variability and reduces the influence of random disturbances. The outer loop is an FNN-based learning controller that generates the system input to suppress the error caused by system nonlinearities and deterministic uncertainties. The FNN employs orthogonal complex Fourier exponentials as its activation functions. Therefore, it is essentially a frequency-domain method that converts the tracking problem in the time domain into a number of regulation problems in the frequency domain. Through a novel phase compensation technique, this model-free method makes it possible to use higher-frequency components in the FNN to improve the tracking performance. In addition, the structure of the FNN can be reconfigured according to the system output information to make the learning more efficient and increase the convergent speed of the tracking error. Experiments on both a commercial gear box and a belt-driven positioning table are conducted to show the effectiveness of the proposed controller. Wei Zuo, Lilong Cai |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2009 | Fourier-Neural-Network-Based Learning Control for a Class of Nonlinear Systems With Flexible ComponentsabstractThis paper considers an output feedback learning control for a class of uncertain nonlinear systems with flexible components. The distinct time delay caused by system flexibility leads to the phase lag phenomenon and low system bandwidth. Therefore, the tracking problem of such systems is very difficult and challenging. To improve the tracking performance of such systems, an iterative learning control scheme using the Fourier neural network (FNN) is presented in this paper. This scheme uses only local output information for feedback. FNN employs orthogonal complex Fourier exponentials as its activation functions and the physical meaning of its hidden-layer neurons is clear. The FNN-based learning controller introduced here relies on the frequency-domain method, which converts the tracking problem in the time domain into a number of regulation problems in the frequency domain. A novel phase compensation method is introduced to deal with the phase lag phenomenon, so that the bandwidth of the closed-loop system is increased. Experiments on a belt-driven positioning table are conducted to show the effectiveness of the proposed controller. Wei Zuo, Lilong Cai |
IEEE Trans. Neural Networks | 1 |
| 2008 | Adaptive-Fourier-Neural-Network-Based Control for a Class of Uncertain Nonlinear SystemsabstractAn adaptive Fourier neural network (AFNN) control scheme is presented in this paper for the control of a class of uncertain nonlinear systems. Based on Fourier analysis and neural network (NN) theory, AFNN employs orthogonal complex Fourier exponentials as the activation functions. Due to the clear physical meaning of the neurons, the determination of the AFNN structure as well as the parameters of the activation functions becomes convenient. One salient feature of the proposed AFNN approach is that all the nonlinearities and uncertainties of the dynamical system are lumped together and compensated online by AFNN. It can, therefore, be applied to uncertain nonlinear systems without any a priori knowledge about the system dynamics. Derived from Lyapunov theory, a novel learning algorithm is proposed, which is essentially a frequency domain method and can guarantee asymptotic stability of the closed-loop system. The simulation results of a multiple-input-multiple-output (MIMO) nonlinear system and the experimental results of an X - Y positioning table are presented to show the effectiveness of the proposed AFNN controller. Wei Zuo, Lilong Cai |
IEEE Trans. Neural Networks | 1 |
| 2007 | Automatic Video Object Segmentation using Graph CutabstractThis paper presents an algorithm for automatic video object planes extraction from coarse to fine. For the case of single moving object in a scene, block-based segmentation defines regions of foreground, background and boundary blocks. Then, the segmentation problem is formulated as an energy minimization problem which is settled by using graph cut algorithm. Automatic segmentation can be realized by obtaining prior knowledge from foreground and background blocks and computation complex is reduced by restricting the refined segmentation region to boundary blocks. Experimental results show the effectiveness of proposed algorithm. It is can be implemented in the head-and-shoulder video sequence segmentation applications. Hong Zhang 0018, Helong Wang, Wei Zuo |
ICIP (3) | 4 |
| 2005 | Tracking control of a belt-driving system using improved Fourier series based learning controllerabstractThe flexible joints in robotic manipulators may lower the bandwidth of the robotic system. Therefore, it is difficult to achieve good control performance on robots with flexible joints by the conventional control schemes. In this paper, we presented the implementation and improvement of the Fourier series based learning controller for tracking control of a belt-driving system which is one type of flexible joints. Experimental results demonstrate the effectiveness of the applied methodology. Wei Zuo, Lilong Cai |
IROS | 2 |
| 2004 | Knowledge-driven segmentation of the central sulcus from human brain MR imagesabstractThis paper presents a knowledge-driven algorithm to identify and segment the central sulcus (CS) from human brain MR images. The dataset is reformatted along the anterior and posterior commissures (AC-PC) plane first. Then, the 3D region within the two coronal planes passing through the AC and PC is defined as the region of interest (ROl) to search for all the sulci within it. The CS is the sulcus with the largest volume within the ROI. Together with the sulci, grey matter (GM) is included for the region growing in order to deal with the partial volume effect. The GM is removed through skeletonization. Experimental results are given. Wei Zuo, Qingmao Hu, Aamer Aziz, Kia-Fock Loe, Wieslaw Lucjan Nowinski |
ICIP | 1 |