Zheyuan Wang

dblp:147/8581 · DBLP profile ↗
← Back
17ranked-venue papers
7as first author
14since 2021 · last 2027
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 7 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2027 Auxiliary particle filtering for nonlinear inequality constrained systems via iterative linearization and truncation
Zheyuan Wang, Yunqi Chen, Zhibin Yan
Signal Process.1
2026 Contrastive Decoupling: Dynamic Regularization for Enhanced Fine-Grained Image Classification
Zheyuan Wang, Tingyao Li, Bin Sheng 0001
IEEE Trans. Circuits Syst. Video Technol.1
2026 Attention-driven feedback for adaptive semantic augmentation in fine-grained visual classification
Zheyuan Wang
Vis. Comput.1
2026 Mask-guided anatomy-aware region mixing for fine-grained medical image classification: VUR grading on VCUG
Shengwei Tian, Yuyin Ma, Zheyuan Wang
Vis. Comput.5
2025 MetaPrior: Meta-Learning Guided Prior Injection for Few-Shot Antibody Affinity Prediction
abstract
Accurate antibody affinity prediction is crucial for drug discovery. However, it remains challenging due to data scarcity, which limits the performance of deep learning models in few-shot, antigen-specific scenarios. To address this challenge, we propose MetaPrior, a novel meta-learning framework designed to learn robust prior knowledge before fine-tuning. The core of MetaPrior is a teacher-student mechanism where a Meta Network (teacher) is trained via bilevel optimization to generate high-quality labels for a diverse set of pseudo-samples and a binding affinity predictor (student) learns from these labels. The teacher is optimized by the student's performance on a small and experimentally measured guide set, enabling the pseudo-samples to approximate the true data distribution under conditions of data scarcity. Extensive experiments demonstrate that MetaPrior is compatible with popular network architectures (e.g., MLPs, CNN s and Transformers) and significantly surpasses standard fine-tuning strategies, particularly in data-scarce scenarios. Across several antigen-specific datasets, such as VEGF, MetaPrior achieves a notable performance gain even with only 40% of the available training data, increasing the Pearson correlation coefficient from 0.38 to 0.58 (a 53% improvement) and enhancing training stability by reducing the standard deviation from 0.42 to 0.19 (a 54.8% reduction in variability). These results confirm the robustness and generalizability of our proposed method. By jointly optimizing pseudo-labels and affinity learning, MetaPrior aligns synthetic supervision with the true data distribution, offering a principled solution for few-shot antibody affinity prediction.
JiaShu, Tingyao Li, Zheyuan Wang, Dezhi Wu, Yihao Song, Yikai Wu 0006, Tobias Plötz, Karin Hrovatin, Stephanie M. Linker, Alexander V. Hopp, Mathias Winkel, Philipp H. P. Harbach
BIBM3
2025 A Progressive Local Variance-guided Strategy for Improving Data Augmentation Reliability
abstract
Recently, CutMix-based augmentation has emerged as a promising strategy for providing regularization to deep neural networks. However, the randomness in cropping may result in uninformative or non-representative regions being selected, resulting in a synthesized image without the desired features. To address these issues, we propose a simple, flexible, and effective augmentation strategy called Progressive LOcal Variance-guided Mix (PlovMix). PlovMix indicates the effective regions based on the information density distribution of the image, which maintains the consistency between synthetic images and the corresponding labels, and further improves the reliability of the augmented data. Additionally, our method generates idiosyncratic shape-free mask for image, which helps the network learn more appropriate feature distributions from the diverse synthetic data. Finally, Experimental results demonstrate PlovMix significantly improves the generalization performance of popular deep networks on various datasets, such as CIFAR-10, CIFAR-100, Tiny ImageNet, and FGVC-Aircraft.
Zheyuan Wang, Dezhi Wu, Haoran Liao, Jiajia Li 0004
ICASSP1
2025 Generative artificial intelligence for ophthalmic images: developments, applications and challenges
Tingyao Li, Zheyuan Wang, Zehua Jiang, Huaiqin Zhong
Vis. Comput.2
2025 HRDC challenge: a public benchmark for hypertension and hypertensive retinopathy classification from fundus images
Xiangning Wang, Zhouyu Guan, An-ran Ran, Tingyao Li, Zheyuan Wang, Xinming Shu, Jinyang Xie, Shichang Liu, Guanyu Xing, Julio Silva-Rodríguez, Riadh Kobbi, Ping Li 0016, Tingli Chen, Lei Bi 0001, Jinman Kim, Weiping Jia, Huating Li, Harry Qin, Ping Zhang 0016, Ching Yu Cheng, Pheng-Ann Heng, Tien Yin Wong, Carol Y. Cheung, Nadia Magnenat-Thalmann, Bin Sheng 0001
Vis. Comput.7
2025 MSPAN: lightweight image super-resolution with multi-semantic guidance
Zheyuan Wang
Vis. Comput.1
2024 Lightweight Remote Sensing Image Super-Resolution via Background-Based Multiscale Feature Enhancement Network
abstract
In the field of remote sensing image super-resolution (RSISR), most methods based on convolutional neural networks (CNNs) tend to focus on high-weight features in the convolutional kernels, thus overlooking low-weight background features. This bias may result in the neglect of some important information in the background. To address this challenge, we propose a background-based multiscale feature enhancement network (BMFENet), which can extract and supplement missing features from different scale backgrounds to improve the reconstruction of remote sensing images (RSIs). Specifically, we constructed a large kernel feature supplement block (LFSB). The LFSB uses large kernel attention mechanism and multiscale mechanism to expand the receptive field, aggregating global information. Meanwhile, it generates background feature weights to increase the attention to neglected information, thereby reducing the distortion of detailed features. Furthermore, to enhance the nonlinear expression capability of the model, we designed a lattice gated unit (LGU). The LGU removes redundant information through a gating mechanism, efficiently aggregates useful channel information through interchannel interactions and attention mechanisms, and introduces directional convolution to make the model more adaptable to super-resolution (SR) tasks in complex scenes. We validated our method on two remote sensing and four SR benchmark datasets, and the results show that our approach achieves a good balance between performance and complexity.
Tianren Wu, Rundong Zhao, Ming Lv, Zhenhong Jia, Liangliang Li 0001, Zheyuan Wang, Hongbing Ma
IEEE Geosci. Remote. Sens. Lett.6
2024 Heterogeneous Policy Networks for Composite Robot Team Communication and Coordination
abstract
High-performing human–human teams learn intelligent and efficient communication and coordination strategies to maximize their joint utility. These teams implicitly understand the different roles of heterogeneous team members and adapt their communication protocols accordingly. Multiagent reinforcement learning (MARL) has attempted to develop computational methods for synthesizing such joint coordination–communication strategies, but emulating heterogeneous communication patterns across agents with different state, action, and observation spaces has remained a challenge. Without properly modeling agent heterogeneity, as in prior MARL work that leverages homogeneous graph networks, communication becomes less helpful and can even deteriorate the team's performance. In the past, we proposed heterogeneous policy networks (HetNet) to learn efficient and diverse communication models for coordinating cooperative heterogeneous teams. In this extended work, we extend HetNet to support scaling heterogeneous robot teams. Building on heterogeneous graph-attention networks, we show that HetNet not only facilitates learning heterogeneous collaborative policies, but also enables end-to-end training for learning highly efficient binarized messaging. Our empirical evaluation shows that HetNet sets a new state-of-the-art in learning coordination and communication strategies for heterogeneous multiagent teams by achieving an 5.84% to 707.65% performance improvement over the next-best baseline across multiple domains while simultaneously achieving a 200× reduction in the required communication bandwidth.
Esmaeil Seraj, Rohan R. Paleja, Luis Pimentel, Kin Man Lee, Zheyuan Wang, Matthew Sklar, John Z. Zhang, Zahi M. Kakish, Matthew C. Gombolay
IEEE Trans. Robotics5
2023 A2M: An Amplification-Arbitrary Module for Remote Sensing Image Super-Resolution
abstract
Remote-sensing (RS) image super-resolution (SR) aims to recover high-resolution (HR) images from the corresponding low-resolution (LR) images. In recent years, the SR methods based on convolutional neural networks (CNNs) have achieved incredible performance in case of fixed scale factors (e.g., ×2, ×3, and ×4). However, these methods need to train a single model for each scale factor, and fail to directly reconstruct the HR image of decimal factors. To solve the lack of research on arbitrary scale of RS image SR, we propose a novel amplification module called amplification-arbitrary module (A2M). A2M can be easily embedded in the tail of the previous SR networks, so that the previous networks can also achieve end-to-end arbitrary scale SR. Specifically, we first utilize the combination of convolutional and pixelshuffle layers to zoom in the deep feature matrix 2×, 3×, and 4× along spatial dimension. Information cross transmission (ICT) is then utilized to gather information of multiple spatial sizes. ICT is not only beneficial to enrich the diversity of information, but also can avoid training only a single branch in the training stage. To make better use of multi-scale features, we designed an efficient signal weighting unit (SWU) to generate a correlation matrix at a small cost, and then the signals of multi-scale features at the same position are fused according to the correlation matrix. Experimental results on RS and generic datasets demonstrate that our method with single pre-training model can perform well at any scale factors.
Yuan Xue 0007, Zheyuan Wang, Liangliang Li 0001, Hongbing Ma
IEEE Geosci. Remote. Sens. Lett.2
2022 Learning Coordination Policies over Heterogeneous Graphs for Human-Robot Teams via Recurrent Neural Schedule Propagation
abstract
As human-robot collaboration increases in the workforce, it becomes essential for human-robot teams to coordinate efficiently and intuitively. Traditional approaches for human-robot scheduling either utilize exact methods that are intractable for large-scale problems and struggle to account for stochastic, time varying human task performance, or application-specific heuristics that require expert domain knowledge to develop. We propose a deep learning-based framework, called HybridNet, combining a heterogeneous graph-based encoder with a recurrent schedule propagator for scheduling stochastic human-robot teams under upper- and lower-bound temporal constraints. The HybridNet's encoder leverages Heterogeneous Graph Attention Networks to model the initial environment and team dynamics while accounting for the constraints. By formulating task scheduling as a sequential decision-making process, the HybridNet's recurrent neural schedule propagator leverages Long Short-Term Memory (LSTM) models to propagate forward consequences of actions to carry out fast schedule generation, removing the need to interact with the environment between every taskagent pair selection. The resulting scheduling policy network provides a computationally lightweight yet highly expressive model that is end-to-end trainable via Reinforcement Learning algorithms. We develop a virtual task scheduling environment for mixed human-robot teams in a multi-round setting, capable of modeling the stochastic learning behaviors of human workers. Experimental results showed that HybridNet outperformed other human-robot scheduling solutions across problem sizes for both deterministic and stochastic human performance, with faster runtime compared to pure-GNN-based schedulers.
Batuhan Altundas, Zheyuan Wang, Joshua Bishop, Matthew C. Gombolay
IROS2
2022 FeNet: Feature Enhancement Network for Lightweight Remote-Sensing Image Super-Resolution
abstract
In the field of remote sensing, due to memory consumption and computational burden, the single-image super-resolution (SISR) methods based on deep convolution neural networks (CNNs) are limited in practical application. To address this problem, we propose a lightweight feature enhancement network (FeNet) for accurate remote-sensing image super-resolution (SR). Considering the existence of equipment with extremely poor hardware facilities, we further design a lighter FeNet-baseline with about 158K parameters. Specifically, inspired by lattice structure, we construct a lightweight lattice block (LLB) as a nonlinear feature extraction function to improve the expression ability. Here, channel separation operation makes the upper and lower branches of the LLB only responsible for half of the features, and the weight coefficients calculated through the attention mechanism enable the upper and lower branches to communicate efficiently. Based on LLB, the feature enhancement block (FEB) is designed in a nested manner to obtain expressive features, where different layers are responsible for the features with different texture richness, and then features from different layers are sequentially fused from deep to shallow. Model parameters and multi-adds operations are used to evaluate network complexity, and extensive experiments on two remote-sensing and four SR benchmark test datasets show that our methods can achieve a good tradeoff between complexity and performance. Our code will be available athttps://github.com/wangzheyuan-666/FeNet.
Zheyuan Wang, Liangliang Li 0001, Yuan Xue 0007, Chenchen Jiang, Kaipeng Sun, Hongbing Ma
IEEE Trans. Geosci. Remote. Sens.1
2020 Exploiting Multi-Layer Grid Maps for Surround-View Semantic Segmentation of Sparse LiDAR Data
abstract
In this paper, we consider the transformation of laser range measurements into a top-view grid map representation to approach the task of LiDAR-only semantic segmentation. Since the recent publication of the SemanticKITTI data set, researchers are now able to study semantic segmentation of urban LiDAR sequences based on a reasonable amount of data. While other approaches propose to directly learn on the 3D point clouds, we are exploiting a grid map framework to extract relevant information and represent them by using multi-layer grid maps. This representation allows us to use well-studied deep learning architectures from the image domain to predict a dense semantic grid map using only the sparse input data of a single LiDAR scan. We compare single-layer and multi-layer approaches and demonstrate the benefit of a multi-layer grid map input. Since the grid map representation allows us to predict a dense, 360° semantic environment representation, we further develop a method to combine the semantic information from multiple scans and create dense ground truth grids. This method allows us to evaluate and compare the performance of our models not only based on grid cells with a detection, but on the full visible measurement range.
Frank Bieder, Sascha Wirges, Johannes Janosovits, Sven Richter, Zheyuan Wang, Christoph Stiller
IV5
2016 Object detection capability evaluation for SAR image
abstract
The existing SAR image quality assessment method could not be effectively used for assessing the performance of object detection. Thus, it is difficult to select SAR images and corresponding detection algorithms for SAR object detection. By analyzing the relationship between object detection results and basic image quality indicators, this paper studies the image quality assessment for image object detection. Based on the concept of “application suitability”, basic quality indicators including radiometric resolution, spatial resolution, PSLR and ISLR are integrated into a single indicator called Detection Index, which is able to comprehensively evaluate the degree to which SAR image is suitable for object detection tasks. Experimental results on aircraft detection with single scene show the effectiveness of the proposed model for SAR image capability evaluation in object detection applications.
Zheyuan Wang, Fangjie Yu, Wenxian Yu, Zhuhui Jiang, Yongke Ding
IGARSS1
2014 Cancer Classification Using Ensemble of Error Correcting Output Codes
Kunhong Liu 0001, Zheyuan Wang
ICIC (3)3