Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Zhonglei Wang

dblp:74/1643 · DBLP profile ↗
← Back
21ranked-venue papers
9as first author
6since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 9 · 6 first-authorSoftware engineering, systems software and programming languages · 6 · 5 first-authorComputer networks · 3Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Theory of computation · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Learning theory · 70% Efficient and distributed learning · 30%
Theoretical computer science
1 paper
Mathematical optimization · 100%
Computer architecture, parallel and distributed computing, and storage systems
2 papers
Embedded and real-time systems · 60% Electronic design automation · 40%

Topics — the 11 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
active learning
0.912025
Balanced Active Inference · NeurIPS 2025
Machine learning › Learning theory › matrix completion
inductive matrix completion
0.912025
Group-Sparse Inductive Matrix Completion Through Transfer Learning · IEEE Trans. Inf. Theory 2025
Machine learning › Learning theory
matrix completion
0.912025
Group-Sparse Inductive Matrix Completion Through Transfer Learning · IEEE Trans. Inf. Theory 2025
Mathematical optimization
experimental design
0.912025
Balanced Active Inference · NeurIPS 2025
Machine learning › Learning theory › statistical learning theory
asymptotic analysis
0.312025
Balanced Active Inference · NeurIPS 2025
Embedded and real-time systems
automotive systems
0.112009
SysCOLA: a framework for co-development of automotive software and system platform · DAC 2009
Embedded and real-time systems
cyber-physical system platforms
0.112009
SysCOLA: a framework for co-development of automotive software and system platform · DAC 2009
Electronic design automation
hardware/software co-design
0.112009
SysCOLA: a framework for co-development of automotive software and system platform · DAC 2009
Embedded and real-time systems › worst-case execution time analysis
software performance estimation
0.112009
An efficient approach for system-level timing simulation of compiler-optimized embedded software · DAC 2009
Electronic design automation
system-level design
0.112009
SysCOLA: a framework for co-development of automotive software and system platform · DAC 2009
Requirements engineering and software design
software architecture
0.012009
SysCOLA: a framework for co-development of automotive software and system platform · DAC 2009

Methods — techniques the papers use, named apart from their topics

cube method · 1.7balanced constraints · 1.7inductive matrix completion · 0.9group sparsity · 0.9virtual prototyping · 0.2systemc · 0.2intermediate source code · 0.2back-annotation · 0.2COLA · 0.2
YearPublicationVenuePosition
2026 Bridging land and sea: A latent diffusion framework for high-resolution ocean floor mapping
Duo Shuai, Qiang Deng, Yang Lu 0009, Qixian Zhong, Zhonglei Wang
Expert Syst. Appl.6
2025 Transfer Learning via Functional Balancing in Reproducing Kernel Hilbert Spaces
abstract
As the availability of different data sources increases, transfer learning has become popular for improving estimation efficiency by incorporating those sources. In this paper, we propose a functional balancing transfer learning algorithm for observational studies integrating external summary information. We achieve efficiency improvement even when the model for the external summary information is misspecified, making the proposed algorithm more robust practically. The asymptotic properties of the proposed algorithm are investigated, and numerical studies demonstrate its advantages compared to alternative methods.
Boyan Gu, Xiaojun Mao, Zhonglei Wang
ICASSP4
2025 Balanced Active Inference
abstract
Limited labeling budget severely impedes data-driven research, such as medical analysis, remote sensing and population census, and active inference is a solution to this problem. Prior works utilizing independent sampling have achieved improvements over uniform sampling, but its insufficient usage of available information undermines its statistical efficiency. In this paper, we propose balanced active inference, a novel algorithm that incorporates balanced constraints based on model uncertainty utilizing the cube method for label selection. Under regularity conditions, we establish its asymptotic properties and also prove that the statistical efficiency of the proposed algorithm is higher than its alternatives. Various numerical experiments, including regression and classification in both synthetic setups and real data analysis, demonstrate that the proposed algorithm outperforms its alternatives while guaranteeing nominal coverage.
Zhixiang Zhou, Liuhua Peng, Zhonglei Wang
NeurIPS4
2025 Group-Sparse Inductive Matrix Completion Through Transfer Learning
abstract
The emergence of big data has enabled the creation of significant models by allowing the storage of large data volumes. Transfer learning is a machine learning technique that transfers knowledge between different domains by utilizing pretrained models from the source domain to optimize the target domain. In contrast, inductive matrix completion is a method that leverages side information from multiple sources to improve task performance. This paper explores inductive matrix completion within the transfer learning framework, with our proposed approach assuming group sparsity for the difference between the core matrices of the target and source domains. Theoretical guarantees of our method are investigated to demonstrate the gains achieved through transfer learning compared with standard inductive matrix completion. Several synthetic experiments are conducted to evaluate the performance of the proposed approach and existing methods, demonstrating that our method outperforms others.
Xiaojun Mao, Hengfang Wang, Zhonglei Wang
IEEE Trans. Inf. Theory3
2024 Collective Matrix Completion via Graph Extraction
abstract
Collective matrix completion (CMC) offers a straightforward approach to dealing with data with entries from various sources. Benefiting from the joint structure in the collective matrix, CMC often achieves fast convergence. However, since CMC conducts matrix-level operations, it neglects the entry-wise information that can potentially be very useful for matrix completion. In this paper, to capture the entry-wise information, we propose a method called graph collective matrix completion (GCoMC). Specifically, our method integrates a graph pattern extraction module into CMC via a relational graph convolutional network. Experiments on simulated and real-world datasets show that our method significantly outperforms some existing counterparts.
Tong Zhan, Xiaojun Mao, Jian Wang 0016, Zhonglei Wang
IEEE Signal Process. Lett.4
2023 Transductive Matrix Completion with Calibration for Multi-Task Learning
abstract
Multi-task learning has attracted much attention due to growing multi-purpose research with multiple related data sources. More- over, transduction with matrix completion is a useful method in multi-label learning. In this paper, we propose a transductive matrix completion algorithm that incorporates a calibration constraint for the features under the multi-task learning framework. The proposed algorithm recovers the incomplete feature matrix and target matrix simultaneously. Fortunately, the calibration information improves the completion results. In particular, we provide a statistical guarantee for the proposed algorithm, and the theoretical improvement induced by calibration information is also studied. Moreover, the proposed algorithm enjoys a sub-linear convergence rate. Several synthetic data experiments are conducted, which show the proposed algorithm out-performs other methods, especially when the target matrix is associated with the feature matrix in a nonlinear way.
Hengfang Wang, Yasi Zhang, Xiaojun Mao, Zhonglei Wang
ICASSP4
2013 DANCE: distributed application-aware node configuration engine in shared reconfigurable sensor networks
abstract
Wireless sensor networks (WSNs) are often primarily tailored to single applications to achieve one specific mission. Considering that the same physical phenomenon can be used by multiple applications, the benefit of sharing the WSN infrastructure is obvious in terms of development and deployment cost. However, allocating the tasks to the WSNs to meet the requirements of all applications while keeping the energy efficiency is very challenging. Introducing reconfigurable nodes in the shared sensor networks can improve the performance, the energy efficiency and the flexibility but it increases the system complexity. In this paper, we propose a biologically inspired node configuration scheme in shared reconfigurable sensor network named DANCE, which can adapt to the changing environment and efficiently utilize WSN resources. Our experiments show that our scheme reduces the energy consumption by up to 76%.
Chih-Ming Hsieh, Zhonglei Wang, Jörg Henkel
DATE2
2013 Fast and accurate cache modeling in source-level simulation of embedded software
abstract
Recently, source-level software models are increasingly used for software simulation in TLM (Transaction Level Modeling)-based virtual prototypes of multicore systems. A source-level model is generated by annotating timing information into application source code and allows for very fast software simulation. Accurate cache simulation is a key issue in multicore systems design because the memory subsystem accounts for a large portion of system performance. However, cache simulation at source level faces two major problems: (1) as target data addresses cannot be statically resolved during source code instrumentation, accurate data cache simulation is very difficult at source level, and (2) cache simulation brings large overhead in simulation performance and therefore cancels the gain of source level simulation. In this paper, we present a novel approach for accurate data cache simulation at source level. In addition, we also propose a cache modeling method to accelerate both instruction and data cache simulation. Our experiments show that simulation with the fast cache model achieves 450.7 MIPS (million simulated instructions per second) on a standard x86 laptop, 2.3x speedup compared with a standard cache model. The source-level models with cache simulation achieve accuracy comparable to an Instruction Set Simulator (ISS). We also use a complex multimedia application to demonstrate the efficiency of the proposed approach for multicore systems design.
Zhonglei Wang, Jörg Henkel
DATE1
2012 Accurate source-level simulation of embedded software with respect to compiler optimizations
abstract
Source code instrumentation is a widely used method to generate fast software simulation models by annotating timing information into application source code. Source-level simulation models can be easily integrated into SystemC based simulation environment for fast simulation of complex multiprocessor systems. The accurate back-annotation of the timing information relies on the mapping between source code and binary code. The compiler optimizations might make it hard to get accurate mapping information. This paper addresses the mapping problems caused by complex compiler optimizations, which are the main source of simulation errors. To obtain accurate mapping information, we propose a method called fine-grained flow mapping that establishes a mapping between sequences of control flow of source code and binary code. In case that the code structure of a program is heavily altered by compiler optimizations, we propose to replace the altered part of the source code with functionally-equivalent IR-level code which has an optimized structure, leading to Partly Optimized Source Code (POSC). Then the flow mapping can be established between the POSC and the binary code and the timing information is back-annotated to the POSC. Our experiments demonstrate the accuracy and speed of simulation models generated by our approach.
Zhonglei Wang, Jörg Henkel
DATE1
2012 A Reconfigurable Hardware Accelerated Platform for Clustered Wireless Sensor Networks
abstract
With the advent of the FPGA technology, a reconfigurable platform becomes possible to enhance the capability and adaptability of wireless sensor nodes while reducing the energy consumption. In this paper, we investigate the application of reconfigurable platform in wireless sensor networks. We focus on using hardware accelerated lossless compression to reduce the energy consumption of the data aggregation in clustered networks. Our study is based on real-world measurement on our integrated reconfigurable sensor node and the measured data is then used in the network simulation with a proper channel model. Our experiments show that the use of hardware acceleration on our platform can reduce the energy consumption by up to 35%. In addition, with proper parameters of the network, the run-time reconfiguration becomes beneficial and this opens a lot of opportunities for the optimizations according to the field situations.
Chih-Ming Hsieh, Zhonglei Wang, Jörg Henkel
ICPADS2
2012 An adaptive data gathering strategy for target tracking in cluster-based wireless sensor networks
Zhonglei Wang, Jörg Henkel
ISCC2
2012 An adaptive data gathering strategy for target tracking in cluster-based wireless sensor networks
abstract
In a typical cluster-based sensor network, the data is usually gathered and fused on cluster heads. In target tracking applications, a target is often detected by sensor nodes in multiple clusters, leading to redundant data transmissions through multiple paths from the cluster heads to the data sink. To reduce such redundant data transmissions and thus to save energy, this paper proposes an adaptive data gathering strategy, called ADGS. Our novel idea is to adaptively select one node with the most residual energy and the least communication cost from the active nodes around the target. This node is responsible for gathering and aggregating the data from the other active nodes and is therefore called Aggregation Node (AN). The aggregated data is then transmitted only from the AN to the sink. Our experiments demonstrate that the proposed approach achieves a significant reduction in power consumption for data transmission and prolongs the network lifetime by 857.6% and 85.8% compared to two state-of-the-art data gathering approaches.
Zhonglei Wang, Jörg Henkel
ISCC2
2012 ECO/ee: Energy-aware Collaborative Organic execution environment for wireless sensor networks
abstract
This paper presents an energy-aware execution environment, called Energy-aware Collaborative Organic Execution Environment (ECO/ee), for Wireless Sensor Networks (WSN). ECO/ee provides an energy-aware and autonomous data routing scheme. It can dynamically adapt to the user's queries and efficiently determine the energy-abundant delivery paths. The fundamental concepts are inspired by the forage and task allocation mechanisms of honey bees as well as the interaction between the honey bee colony and the bee keeper. The experiments show that our approach outperforms the state-of-the-art in terms of energy efficiency and adaptability to the user requirements.
Chih-Ming Hsieh, Zhonglei Wang, Jörg Henkel
WCNC2
2011 An approach to improve accuracy of source-level TLMs of embedded software
abstract
Virtual Prototypes (VPs) based on Transaction Level Models (TLMs) have become a de-facto standard for design space exploration and validation of complex software-centric multicore or multiprocessor systems. The most popular method to get timed software TLMs is to annotate timing information at the basic-block level granularity back into application source code, called source code instrumentation (SCI). The existing SCI approaches realize the back-annotation of timing information based on mapping between source code and binary code. However, optimizing compilation has a large impact on the code mapping and will lower the accuracy of the generated source-level TLMs. In this paper, we present an efficient approach to tackle this problem. We propose to use mapping between source-level and binary-level control flows as the basis for timing annotation instead of code mapping. Software TLMs generated by our approach allow for accurate evaluation of multiprocessor systems at a very high speed. This has been proven by our experiments with a set of benchmark programs and a case study.
Zhonglei Wang, Kun Lu 0005, Andreas Herkersdorf
DATE1
2011 RDTS: A Reliable Erasure-Coding Based Data Transfer Scheme for Wireless Sensor Networks
abstract
Information redundancy using erasure coding is an efficient way to increase the reliability of data transmission in communication systems. In Wireless Sensor Networks (WSNs), erasure encoding and decoding are performed on the source node and sink node, respectively, and a large amount of redundant data is generated according to the quality of the whole path and transmitted through multiple hops. In this paper, we propose a reliable data transfer scheme, RDTS, where erasure coding is performed in a hop-by-hop manner, which means that each intermediate node is able to perform erasure coding and adaptively calculates the number of redundant packets for the next hop. Usually, only a small amount of redundant data is needed for reliable transmission over a single hop. Therefore, using RDTS, the network load caused by redundant data is significantly reduced and also well balanced, leading to a longer network lifetime. In addition, hop-by-hop coding has also the advantage of low coding overhead. We further reduce the coding time by proposing a partial coding scheme. Our experimental results show that RDTS achieves up to 69.7% less network load and 153.8% longer lifetime, and meanwhile, the coding overhead is reduced by up to 78.1%, compared with a state-of-the-art erasure-coding based approach.
M. Sammer Srouji, Zhonglei Wang, Jörg Henkel
ICPADS2
2010 Software performance simulation strategies for high-level embedded system design
Zhonglei Wang, Andreas Herkersdorf
Perform. Evaluation1
2009 An efficient approach for system-level timing simulation of compiler-optimized embedded software
abstract
Software accounts for more than 80% of embedded system development efforts, so software performance estimation is a very important issue in system design. Recently, source level simulation (SLS) has become a state-of-the-art approach for software simulation in system level design. However, the simulation accuracy relies on the mapping between source code and binary code, which can be destroyed by compiler optimizations. This drawback strongly limits the usability of this technique in practical system design. We introduce an approach to overcome this limitation by converting source code to a low level representation, called intermediate source code (ISC). ISC has accounted for most compiler optimizations and has a structure close to binary code, so it allows for accurate back-annotation of timing information from the binary level. To show the benefits of our approach, we present a quantitative comparison of the related techniques with the proposed one, using a set of benchmarks.
Zhonglei Wang, Andreas Herkersdorf
DAC1
2009 SysCOLA: a framework for co-development of automotive software and system platform
abstract
A modeling language with formal semantics is able to capture a system's functionality unambiguously, without concerning implementation details. Such a formal language is well-suited for a design process that employs formal techniques and supports hardware/software synthesis. On the other hand, SystemC is a widely used system level design language with hardware-oriented modeling features. It provides a desirable simulation framework for system architecture design and exploration. This paper presents a design framework, called SysCOLA, that makes use of the unique advantages of both a new formal modeling language, COLA, and SystemC, and allows for parallel development of application software and system platform. In SysCOLA, function design and architecture exploration are done in the COLA based modeling environment and the SystemC based virtual prototyping environment, respectively. Our concepts of abstract platform and virtual platform abstraction layer facilitate the orthogonalization of functionality and architecture by means of mapping and integration in the respective environments. As SysCOLA is targeted at the automotive domain, the whole design approach is showcased using a case study of designing an automotive system.
Zhonglei Wang, Andreas Herkersdorf, Wolfgang Haberl, Martin Wechs
DAC1
2009 Flow Analysis on Intermediate Source Code for WCET Estimation of Compiler-Optimized Programs
abstract
Many WCET analysis tools developed in academia integrate WCET analysis into program compilation, either to transform flow information extracted from the source code level to the object code level, or to perform flow analysis on a special intermediate representation. This integration increases analysis complexity, forces software developers to use a special compiler, and thus, strongly limits the usability of the tools in practice. Motivated by this limitation in the existing flow analysis approaches, this paper presents a more efficient approach, that performs flow analysis on the intermediate source code (ISC), transformed from the original source code. ISC retains the functional behavior and executability of the original source code but has a structure close to the object code. This low level structure facilitates the transformation of the flow facts, extracted from the ISC, down to the object code level. In the whole approach, no modification of standard tools is needed.
Zhonglei Wang, Andreas Herkersdorf
RTCSA1
2008 A Model Driven Development Approach for Implementing Reactive Systems in Hardware
abstract
To deal with the increasing complexity of digital systems, the model driven development approach has proven to be beneficial. This paper presents a model driven hardware design process that is dedicated to reactive embedded systems. The approach is based on the component language (COLA), a synchronous data flow language with formal semantics. COLA follows the hypothesis of perfect synchrony. Models thus do not assume specific timing properties and remain deterministic as long as data flow requirements are retained. This is an essential feature for modeling safety-critical systems. Further, the well-defined semantics not only allows that the resulting models can be formally reasoned about, but is also the key to translation to domain-specific languages. This paper describes the approach of translating the models to VHDL descriptions from their graphical representations. As COLA is well-adapted to both data flow description and control automata, the generated VHDL code can be synthesized to very efficient FPGA circuits, comparable to that synthesized from hand-written VHDL code according to our case study.
Zhonglei Wang, Andreas Herkersdorf, Stefano Merenda, Michael Tautschnig
FDL1
2008 A Simulation Approach for Performance Validation during Embedded Systems Design
Zhonglei Wang, Wolfgang Haberl, Andreas Herkersdorf, Martin Wechs
ISoLA1