Wenjia Meng

dblp:194/2961 · DBLP profile ↗
← Back
14ranked-venue papers
7as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 5 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 MTRL-CG: Multi-Task Reinforcement Learning Method with Spectral Clustering-Based Task Grouping
abstract
Multi-task reinforcement learning (RL) aims to enhance agent performance across multiple tasks by enabling effective knowledge transfer. However, these methods adopt a fully shared policy across all tasks without explicitly distinguishing between related and conflicting ones, making them suffer from negative interference issue, where updates beneficial to one task adversely affect others and lead to degraded overall performance. In this paper, we propose a multi-task reinforcement learning method with spectral clustering-based task grouping (MTRL-CG), which leverages spectral clustering to group related tasks and separate conflicting ones, enabling group-wise policy learning to mitigate negative interference. We first quantify inter-task affinity by measuring the influence of task-specific updates on others within a shared model, and construct an affinity matrix to capture these relationships. Spectral clustering is then applied to partition tasks via spectral embedding and k-means clustering. Each task group is trained with a dedicated policy network to promote focused learning. Built upon the Soft Actor-Critic (SAC) algorithm, MTRL-CG can be readily integrated into existing SAC-based multi-task RL methods. Extensive experiments on the Meta-World benchmark demonstrate the effectiveness of the proposed MTRL-CG method.
Wenjia Meng, Haoliang Sun, Yilong Yin
AAAI1
2026 CausalCOMRL: Context-based offline meta-reinforcement learning with causal representation
abstract
Context-based offline meta-reinforcement learning (OMRL) methods have achieved appealing success by leveragingpre-collected offline datasets to develop task representations that guide policy learning. However, current context-based OMRL methods often introduce spurious correlations, where task components are incorrectly correlated due to confounders. These correlations can degrade policy performance when the confounders in the test taskdiffer from those in the training task. To address this problem, we propose CausalCOMRL, a context-based OMRL method that integrates causal representation learning. This approach uncovers causal relationships among the task components and incorporates the causal relationships into task representations, enhancing the generalizability of RL agents. We further improve the distinction of task representations from different tasks by using mutual information optimization and contrastive learning. Utilizing these causal task representations, we employSAC to optimize policies on meta-RL benchmarks. Experimental results show that CausalCOMRL achieves better performance than other methods on most benchmarks.
Zhengzhe Zhang, Wenjia Meng, Haoliang Sun, Gang Pan 0001
Neural Networks2
2025 A Gaussian Filter-Based 3D Registration Method for Series Section Electron Microscopy
abstract
Series Section Electron Microscopy (ssEM) is a crucial technique for visualizing three-dimensional (3D) biological structures, which involves collecting electron microscopy images from a series of biological sections along the z-axis and reconstructing the 3D structure. 3D registration is an essential step in ssEM, designed to eliminate axial misalignment and nonlinear distortions introduced during sample sectioning. A significant challenge in 3D registration is eliminating nonlinear distortions while preserving natural deformations. In this paper, we present a new formulation of the 3D registration problem from a frequency domain perspective and propose a Gaussian filtering-based 3D registration method, which defines 3D registration as a superposition problem of high-frequency and low-frequency components. We extend the concept of a one-dimensional Gaussian filter to three-dimensional image stacks and integrate it with optical flow networks to consolidate the deformation field within the receptive field. Extensive experiments demonstrate that our method can successfully decouple nonlinear distortions and natural deformations in the frequency domain, proving superior to existing methods in rapidly and accurately eliminating nonlinear distortions and restoring biological structures, and has the potential to be extended to large datasets.
Zhenbang Zhang, Hongjia Li 0001, Wenjia Meng, Renmin Han
AAAI4
2025 SeqMvRL: A Sequential Fusion Framework for Multi-view Representation Learning
abstract
Multi-view representation learning integrates multiple observable views of an entity into a unified representation to facilitate downstream tasks. Current methods predominantly focus on distinguishing compatible components across views, followed by a single-step parallel fusion process. However, this parallel fusion is static in essence, overlooking potential conflicts among views and compromising representation ability. To address this issue, this paper proposes a novel Sequential fusion framework for Multi-view Representation Learning, termed SeqMvRL. Specifically, we model multi-view fusion as a sequential decision-making problem and construct a pairwise integrator (PI) and a next-view selector (NVS), which represent the environment and agent in reinforcement learning, respectively. PI merges the current fused feature with the selected view, while NVS is introduced to determine which view to fuse subsequently. By adaptively selecting the next optimal view for fusion based on the current fusion state, SeqMvRL thereby effectively reduces conflicts and enhances unified representation quality. Additionally, an elaborate novel reward function encourages the model to prioritize views that enhance the discriminability of the fused features. Experimental results demonstrate that SeqMvRL outperforms parallel fusion schemes in classification and clustering tasks.
Ren Wang 0011, Haoliang Sun, Yuxiu Lin, Chuanhui Zuo, Yongshun Gong, Yilong Yin, Wenjia Meng
CVPR7
2025 Pixel-wise Single Image Reflection Removal Method Based on Reinforcement Learning
abstract
Single image reflection removal is particularly important in improving image quality. However, existing single image reflection removal methods cannot remove reflection in a pixel-by-pixel manner, significantly reducing their effectiveness. To address this issue, we propose a pixel-wise single image reflection removal method based on reinforcement learning. Specifically, we formalize the image reflection removal process as a sequential decision-making process and remove the single image reflection pixel by pixel. We first introduce the single image reflection removal method at each time step. We then describe the reinforcement learning setup for image reflection, including state, action, reward, and agent network design. Finally, we conducted experiments on benchmark datasets, and the results show that our method outperforms other single image reflection removal methods.
Xueshi Yu, Zhengzhe Zhang, Xiankai Lu, Yilong Yin, Wenjia Meng
ICME7
2025 Few-shot classification of Cryo-ET subvolumes with deep Brownian distance covariance
abstract
Few-shot learning is a crucial approach for macromolecule classification of the cryo-electron tomography (Cryo-ET) subvolumes, enabling rapid adaptation to novel tasks with a small support set of labeled data. However, existing few-shot classification methods for macromolecules in Cryo-ET consider only marginal distributions and overlook joint distributions, failing to capture feature dependencies fully. To address this issue, we propose a method for macromolecular few-shot classification using deep Brownian Distance Covariance (BDC). Our method models the joint distribution within a transfer learning framework, enhancing the modeling capabilities. We insert the BDC module after the feature extractor and only train the feature extractor during the training phase. Then, we enhance the model's generalization capability with self-distillation techniques. In the adaptation phase, we fine-tune the classifier with minimal labeled data. We conduct experiments on publicly available SHREC datasets and a small-scale synthetic dataset to evaluate our method. Results show that our method improves the classification capabilities by introducing the joint distribution.
Xueshi Yu, Renmin Han, Haitao Jiao, Wenjia Meng
Briefings Bioinform.4
2025 A noise-robust classification method for cryo-ET subtomograms with out-of-distribution detection
abstract
MOTIVATION: Cryogenic electron tomography (cryo-ET) enables high-resolution 3D reconstruction of biological samples, with accurate subtomogram classification critical for structural analysis. However, current subtomogram classification methods often struggle with out-of-distribution (OOD) data issue, causing misclassification and mismatched structures. RESULTS: To solve this problem, we propose a unified subtomogram classification framework that incorporates OOD detection to distinguish unknown (OOD) from known (in-distribution, ID) classes and predict labels for ID data, thereby enhancing existing subtomogram classification methods. Within this framework, we develop a noise-robust classification method that integrates a 3D discrete wavelet transform-based encoder to reduce high-frequency noise and extract robust features. Additionally, we incorporate a Mahalanobis distance-based OOD detector with a reliable metric for 3D subtomograms and introduce an adaptive classifier that adjusts to accommodate datasets of varying scales. The experimental and visualization results demonstrate that our noise-robust method improves subtomogram classification accuracy and effectively models features while enhancing OOD detection. AVAILABILITY AND IMPLEMENTATION: Our code is available at https://github.com/yxs1137/Subtomo-Classification-with-OOD.git. The real data used in this study can be accessed through CryoET Data Portal.
Wenjia Meng, Xueshi Yu, Renmin Han
Bioinform.1
2025 Cross-graph meta matching correction for noisy graph matching
Fangkai Li, Feiyu Pan, Wenjia Meng, Haoliang Sun, Xiushan Nie, Yilong Yin, Xiankai Lu
Comput. Vis. Image Underst.3
2025 LAC-PS: A Light Direction Selection Policy Under the Accuracy Constraint for Photometric Stereo
abstract
Photometric stereo (PS) methods recover surface normals from appearance changes under varying light directions, excelling in tasks like 3D surface reconstruction and defect inspection. However, collecting the illumination images is expensive, and current PS methods cannot obtain the light direction set that satisfies the pre-defined accuracy constraint, limiting their adaptability to various applications with varying accuracy requirements. To address this issue, we propose the LAC-PS, a light direction selection policy under the accuracy constraint for photometric stereo, which optimizes the light direction set to meet target reconstruction accuracy. In our method, we develop an accuracy assessment network that estimates reconstruction accuracy without ground truth. With this estimated accuracy, we put forward a reinforcement learning-based method that can utilize policy to sequentially select light directions and obtain the light directions satisfying the desired PS recovery accuracy constraint. Experimental results on real and synthetic datasets demonstrate that our method effectively selects light directions that satisfy accuracy constraints.
Wenjia Meng, Huimin Han, Xiankai Lu, Yilong Yin, Gang Pan 0001
IEEE Trans. Circuits Syst. Video Technol.1
2025 Off-OAB: Off-Policy Policy Gradient Method With Optimal Action-Dependent Baseline
abstract
The policy-based methods have achieved remarkable success in solving challenging reinforcement learning (RL) problems. Among these methods, the off-policy policy gradient (OPPG) methods are particularly important because they can benefit from off-policy data. However, these methods suffer from the high variance of the OPPG estimator, which results in poor sample efficiency during training. In this article, we propose an off-policy policy gradient method with the optimal action-dependent baseline (Off-OAB) to mitigate this variance issue. Specifically, this baseline maintains the OPPG estimator's unbiasedness while theoretically minimizing its variance. To enhance practical computational efficiency, we design an approximated version of this optimal baseline. Utilizing this approximation, our method (Off-OAB) aims to decrease the OPPG estimator's variance during policy optimization. We evaluate the proposed Off-OAB method on six representative tasks from OpenAI Gym and MuJoCo, where it demonstrably surpasses the state-of-the-art methods on the majority of these tasks.
Wenjia Meng, Long Yang 0004, Yilong Yin, Gang Pan 0001
IEEE Trans. Neural Networks Learn. Syst.1
2023 Off-Policy Proximal Policy Optimization
abstract
Proximal Policy Optimization (PPO) is an important reinforcement learning method, which has achieved great success in sequential decision-making problems. However, PPO faces the issue of sample inefficiency, which is due to the PPO cannot make use of off-policy data. In this paper, we propose an Off-Policy Proximal Policy Optimization method (Off-Policy PPO) that improves the sample efficiency of PPO by utilizing off-policy data. Specifically, we first propose a clipped surrogate objective function that can utilize off-policy data and avoid excessively large policy updates. Next, we theoretically clarify the stability of the optimization process of the proposed surrogate objective by demonstrating the degree of policy update distance is consistent with that in the PPO. We then describe the implementation details of the proposed Off-Policy PPO which iteratively updates policies by optimizing the proposed clipped surrogate objective. Finally, the experimental results on representative continuous control tasks validate that our method outperforms the state-of-the-art methods on most tasks.
Wenjia Meng, Gang Pan 0001, Yilong Yin
AAAI1
2022 An Off-Policy Trust Region Policy Optimization Method With Monotonic Improvement Guarantee for Deep Reinforcement Learning
abstract
In deep reinforcement learning, off-policy data help reduce on-policy interaction with the environment, and the trust region policy optimization (TRPO) method is efficient to stabilize the policy optimization procedure. In this article, we propose an off-policy TRPO method, off-policy TRPO, which exploits both on- and off-policy data and guarantees the monotonic improvement of policies. A surrogate objective function is developed to use both on- and off-policy data and keep the monotonic improvement of policies. We then optimize this surrogate objective function by approximately solving a constrained optimization problem under arbitrary parameterization and finite samples. We conduct experiments on representative continuous control tasks from OpenAI Gym and MuJoCo. The results show that the proposed off-policy TRPO achieves better performance in the majority of continuous control tasks compared with other trust region policy-based methods using off-policy data.
Wenjia Meng, Gang Pan 0001
IEEE Trans. Neural Networks Learn. Syst.1
2020 Qualitative Measurements of Policy Discrepancy for Return-Based Deep Q-Network
abstract
The deep Q-network (DQN) and return-based reinforcement learning are two promising algorithms proposed in recent years. The DQN brings advances to complex sequential decision problems, while return-based algorithms have advantages in making use of sample trajectories. In this brief, we propose a general framework to combine the DQN and most of the return-based reinforcement learning algorithms, named R-DQN. We show that the performance of the traditional DQN can be significantly improved by introducing return-based algorithms. In order to further improve the R-DQN, we design a strategy with two measurements to qualitatively measure the policy discrepancy. We conduct experiments on several representative tasks from the OpenAI Gym and Atari games. The state-of-the-art performance achieved by our method with this proposed strategy validates its effectiveness.
Wenjia Meng, Long Yang 0004, Pengfei Li 0005, Gang Pan 0001
IEEE Trans. Neural Networks Learn. Syst.1
2018 A Unified Approach for Multi-step Temporal-Difference Learning with Eligibility Traces in Reinforcement Learning
abstract
Recently, a new multi-step temporal learning algorithm Q(σ) unifies n-step Tree-Backup (when σ = 0) and n-step Sarsa (when σ = 1) by introducing a sampling parameter σ. However, similar to other multi-step temporal-difference learning algorithms, Q(σ) needs much memory consumption and computation time. Eligibility trace is an important mechanism to transform the off-line updates into efficient on-line ones which consume less memory and computation time. In this paper, we combine the original Q(σ) with eligibility traces and propose a new algorithm, called Qπ(σ,λ), where λ is trace-decay parameter. This new algorithm unifies Sarsa(λ) (when σ = 1) and Qπ (λ) (when σ = 0). Furthermore, we give an upper error bound of Qπ(σ,λ) policy evaluation algorithm. We prove that Qπ (σ, λ) control algorithm converges to the optimal value function exponentially. We also empirically compare it with conventional temporal-difference learning methods. Results show that, with an intermediate value of σ, Qπ(σ,λ) creates a mixture of the existing algorithms which learn the optimal value significantly faster than the extreme end (σ = 0, or 1).
Long Yang 0004, Minhao Shi, Wenjia Meng, Gang Pan 0001
IJCAI4