Zhijie Wen

dblp:47/9638 · DBLP profile ↗
← Back
28ranked-venue papers
7as first author
14since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 5 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 since 2021Systems, architecture and hardware · 1 · 1 first-author
YearPublicationVenuePosition
2026 High-Precision Tracking Control of Multi-Axis Electro-Optical Systems Based on Accurate Identification of Dynamic Parameters
abstract
High-precision tracking control is essential for multi-axis electro-optical systems (EOSs) to defense low-slow-small (LSS) moving targets. The inherent rotational inertia and nonlinear joint friction degrade the accuracy in both dynamic and low-speed tracking scenarios, necessitating model-based feedforward controllers that require precise parameter identification. To solve this issue, this study proposes a motor-coupled dynamic parameters set (MCDPS) identification framework that incorporates physical feasibility constraints, specifically designed for EOSs that lack joint torque sensors. An asymmetric nonlinear friction model is proposed to characterize the nonlinearity and direction-dependence of joint friction. Then, an Iterative Composite Least Squares (ICLS) algorithm is employed to identify the inertial and nonlinear friction parameters through iterative linear regression. A backpropagation neural network (BPNN) using joint position and load data as inputs is also proposed to further reduce the modeling error of joint nonlinearities. Experiments are performed using a multi-axis EOS to track typical joint trajectory and LSS simulated target. Results demonstrate that the proposed ICLS algorithm combined with the position- and load-related BPNN outperforms traditional identification methods. Additionally, the model-based feedforward controller is compared with proportional integral derivative (PID), extended kalman filter (EKF) and extended state observer (ESO) with sliding-mode control (SMC) approaches, showing a 21.29% improvement in tracking accuracy.
Kehui Xu, Mubang Xiao, Zhijie Wen, Shixun Fan, Dapeng Fan
IEEE Trans Autom. Sci. Eng.4
2026 Re-Visible Dual-Domain Self-Supervised Deep Unfolding Network for MRI Reconstruction
abstract
Magnetic Resonance Imaging (MRI) is widely used in clinical practice, but suffers from prolonged acquisition time. Although deep learning methods have been proposed to accelerate acquisition and demonstrate promising performance, they rely on high-quality fully-sampled datasets for training in a supervised manner. However, such datasets are time-consuming and expensive-to-collect, which constrains their broader applications. On the other hand, self-supervised methods offer an alternative by enabling learning from under-sampled data alone, but most existing methods rely on further partitioned under-sampled k-space data as model's input for training, which causes an input distribution shift between the the training stage and the inference stage. Additionally, their models have not effectively incorporated comprehensive image priors, leading to degraded reconstruction performance. In this paper, we propose a novel re-visible dual-domain self-supervised deep unfolding network to address these issues when only under-sampled datasets are available. Specifically, by incorporating re-visible dual-domain loss, all under-sampled k-space data are utilized during training to mitigate the input distribution shift caused by further partitioning. This design enables the model to implicitly adapt to all under-sampled k-space data as input. Additionally, we design a Deep Unfolding Network based on Chambolle and Pock Proximal Point Algorithm (DUN-CP-PPA) to achieve end-to-end reconstruction. By employing a Spatial-Frequency Feature Extraction (SFFE) block to capture both global and local representations, the model effectively integrates imaging physics with comprehensive image priors to enhance reconstruction performance. Experiments on both single-coil and multi-coil datasets demonstrate that our method outperforms state-of-the-art approaches in terms of reconstruction performance and generalization capability.
Hao Zhang 0026, Qi Wang 0128, Jian Sun 0009, Zhijie Wen, Jun Shi 0004, Shihui Ying
IEEE J. Biomed. Health Informatics4
2025 Mitigating noisy labels in long-tailed image classification via multi-level collaborative learning
Xinyang Zhou, Zhijie Wen, Yuandi Zhao, Jun Shi 0004, Shihui Ying
Appl. Intell.2
2025 Open set label noise learning with robust sample selection and margin-guided module
Yuandi Zhao, Qianxi Xia, Zhijie Wen, Liyan Ma, Shihui Ying
Knowl. Based Syst.4
2025 Deep unfolding network with spatial alignment for multi-modal MRI reconstruction
Hao Zhang 0026, Qi Wang 0128, Jun Shi 0004, Shihui Ying, Zhijie Wen
Medical Image Anal.5
2024 Spatial and Modal Optimal Transport for Fast Cross-Modal MRI Reconstruction
abstract
Multi-modal magnetic resonance imaging (MRI) plays a crucial role in comprehensive disease diagnosis in clinical medicine. However, acquiring certain modalities, such as T2-weighted images (T2WIs), is time-consuming and prone to be with motion artifacts. It negatively impacts subsequent multi-modal image analysis. To address this issue, we propose an end-to-end deep learning framework that utilizes T1-weighted images (T1WIs) as auxiliary modalities to expedite T2WIs' acquisitions. While image pre-processing is capable of mitigating misalignment, improper parameter selection leads to adverse pre-processing effects, requiring iterative experimentation and adjustment. To overcome this shortage, we employ Optimal Transport (OT) to synthesize T2WIs by aligning T1WIs and performing cross-modal synthesis, effectively mitigating spatial misalignment effects. Furthermore, we adopt an alternating iteration framework between the reconstruction task and the cross-modal synthesis task to optimize the final results. Then, we prove that the reconstructed T2WIs and the synthetic T2WIs become closer on the T2 image manifold with iterations increasing, and further illustrate that the improved reconstruction result enhances the synthesis process, whereas the enhanced synthesis result improves the reconstruction process. Finally, experimental results from FastMRI and internal datasets confirm the effectiveness of our method, demonstrating significant improvements in image reconstruction quality even at low sampling rates.
Qi Wang 0128, Zhijie Wen, Jun Shi 0004, Qian Wang 0001, Dinggang Shen, Shihui Ying
IEEE Trans. Medical Imaging2
2024 Histopathology Image Classification With Noisy Labels via The Ranking Margins
abstract
Clinically, histopathology images always offer a golden standard for disease diagnosis. With the development of artificial intelligence, digital histopathology significantly improves the efficiency of diagnosis. Nevertheless, noisy labels are inevitable in histopathology images, which lead to poor algorithm efficiency. Curriculum learning is one of the typical methods to solve such problems. However, existing curriculum learning methods either fail to measure the training priority between difficult samples and noisy ones or need an extra clean dataset to establish a valid curriculum scheme. Therefore, a new curriculum learning paradigm is designed based on a proposed ranking function, which is named The Ranking Margins (TRM). The ranking function measures the 'distances' between samples and decision boundaries, which helps distinguish difficult samples and noisy ones. The proposed method includes three stages: the warm-up stage, the main training stage and the fine-tuning stage. In the warm-up stage, the margin of each sample is obtained through the ranking function. In the main training stage, samples are progressively fed into the networks for training, starting from those with larger margins to those with smaller ones. Label correction is also performed in this stage. In the fine-tuning stage, the networks are retrained on the samples with corrected labels. In addition, we provide theoretical analysis to guarantee the feasibility of TRM. The experiments on two representative histopathologies image datasets show that the proposed method achieves substantial improvements over the latest Label Noise Learning (LNL) methods.
Zhijie Wen, Haixia Wu, Shihui Ying
IEEE Trans. Medical Imaging1
2023 JSMix: a holistic algorithm for learning with label noise
Zhijie Wen, Shihui Ying
Neural Comput. Appl.1
2022 Multitask transfer learning with kernel representation
Shihui Ying, Zhijie Wen
Neural Comput. Appl.3
2022 Long Time Series Deep Forecasting with Multiscale Feature Extraction and Seq2seq Attention Mechanism
Xin Wang 0084, Yixian Luo, Zhijie Wen, Shihui Ying
Neural Process. Lett.4
2021 Few-Shot Classification With Intra-Class Unrelated Multi-Prototype Representation and Episode Adaptation Strategy
abstract
The episode training strategy, which trains models by many episodes to recognize unseen object categories using one or a few samples, is used by many existing approaches to solve the few-shot classification. An episode can be regarded as a classification task for the few-shot classification. However, this training strategy will be tricky when the task feature differences between the training episodes are big. To improve on this shortcomimg, we propose an Episode Adaptation Loss (EAL) to reduce these gaps for optimizing the models efficiently. Furthermore, to enhance the ability of each class feature representation and preserve diversities of the intra-class features, we propose Multi-Prototypical Representations (MPR) and a Multi-Prototype Unrelated Loss (MPUL). In addition, we introduce the multi-prototypes inductive inference as a classification strategy for our method. We evaluate our method on miniImageNet and Fewshot-CIFAR100 benchmarks. Experimental results demonstrate that our method outperforms the baseline approaches. The ablation study validates that the components of the proposed method all provide positive effects on few-shot learning.
Zhijie Wen, Liyan Ma
ICTAI1
2021 Learning Discriminative Representations for Fine-Grained Diabetic Retinopathy Grading
abstract
Diabetic retinopathy is one of the leading causes of blindness. However, no specific symptoms of early DR lead to a delayed diagnosis, which results in disease progression in patients. To determine the disease severity levels, ophthalmologists need to focus on the discriminative parts of the retinal images. In recent years, deep learning has achieved great success in medical image analysis. However, most works directly employ algorithms based on convolutional neural networks (CNNs), which ignore the fact that the difference among classes is subtle and gradual. Hence, we consider automatic image grading of DR as a fine-grained classification task, and construct a bilinear model to identify the pathologically discriminative areas. In order to leverage the ordinal information among classes, we put the soft labels with ordinal information among classes into the loss function rather than the most commonly used one-hot labels for the diabetic retinopathy classification. In addition, other than only using a categorical loss to train our network, we also introduce the metric loss to learn a more discriminative feature space which is beneficial to locate the finer discriminative lesion parts. Experimental results demonstrate the superior performance of the proposed method on publicly available IDRiD, DeepDRiD and FGADR datasets.
Liyan Ma, Zhijie Wen, Shaorong Xie, Yupeng Xu
IJCNN3
2021 Topology-preserving nonlinear shape registration on the shape manifold
Zhijie Wen, Zhongyi Hu 0001
Multim. Tools Appl.2
2021 GCSBA-Net: Gabor-Based and Cascade Squeeze Bi-Attention Network for Gland Segmentation
abstract
Colorectal cancer is the second and the third most common cancer in women and men, respectively. Pathological diagnosis is the "gold standard" for tumor diagnosis. Accurate segmentation of glands from tissue images is a crucial step in assisting pathologists in their diagnosis. The typical methods for gland segmentation form a dense image representation, ignoring its texture and multi-scale attention information. Therefore, we utilize a Gabor-based module to extract texture information at different scales and directions in histopathology images. This paper also designs a Cascade Squeeze Bi-Attention (CSBA) module. Specifically, we add Atrous Cascade Spatial Pyramid (ACSP), Squeeze Position Attention (SPA) module and Squeeze Channel Attention module (SCA) to model semantic correlation and maintain the multi-level aggregation on the spatial pyramid with different dilations. Besides, to solve the imbalance of data distribution and boundary blur, we propose a hybrid loss function to response the object boudary better. The experimental results show that the proposed method achieves state-of-the-art performance on the GlaS challenge dataset and CRAG colorectal adenocarcinoma dataset, respectively.
Zhijie Wen, Ru Feng, Jingxin Liu 0005, Ying Li 0028, Shihui Ying
IEEE J. Biomed. Health Informatics1
2020 Wavelet U-Net for Medical Image Segmentation
Ying Li 0028, Tuo Leng, Zhijie Wen
ICANN (1)4
2020 ACE-Net: Adaptive Context Extraction Network for Medical Image Segmentation
Tuo Leng, Ying Li 0028, Zhijie Wen
ICANN (1)4
2020 Multi-metric Joint Discrimination Network for Few-Shot Classification
Zhijie Wen, Liyan Ma, Shihui Ying
PRCV (3)2
2020 Residual network with detail perception loss for single image super-resolution
Zhijie Wen, Jiawei Guan, Tieyong Zeng, Ying Li 0028
Comput. Vis. Image Underst.1
2020 A 91-Channel Hyperspectral LiDAR for Coal/Rock Classification
abstract
During the mining operation, it is a critical task in coal mines to significantly improve the safety by precision coal mining sorting and rock classification from different layers. It implies that a technique for rapidly and accurately classifying coal/rock in-site needs to be investigated and established, which is of significance for improving the coal mining efficiency and safety. In this letter, a 91-channel hyperspectral LiDAR (HSL) using an acousto-optic tunable filter (AOTF) as the spectroscopic device is designed, which operates based on the wide-spectrum emission laser source with a 5-nm spectral resolution to tackle this issue. The spectra of four-type coal/rock specimens collected by HSL are used to classify with three multi-label classifiers: naive Bayes (NB), logistic regression (LR), and support vector machine (SVM). Furthermore, we discuss and explore whether Gaussian fitting (GF) method and calibration with the reference whiteboard (RB) can enhance the classification accuracy. The experimental results show that the GF technique not only improves the accuracy of range measurement but also optimizes the classification performance using the spectra collected by the HSL. In addition, calibration with RB can improve classification accuracy as well. In addition, we also discuss methods to improve the calibration-free classification accuracy preliminarily.
Yuwei Chen 0005, Zhirong Yang, Changhui Jiang, Wei Li 0095, Haohao Wu, Zhijie Wen, Eetu Puttonen, Juha Hyyppä
IEEE Geosci. Remote. Sens. Lett.7
2019 Manifold Alignment and Distribution Adaptation for Unsupervised Domain Adaptation
abstract
Unsupervised domain adaptation is a problem which exploits the knowledge learned from the resource-rich domain to obtain an accurate classifier for the resource-poor domain. Most of the existing methods lift performance by reducing the differences between distributions, such as the difference between marginal probability distributions, the difference between conditional probability distributions, or both. However, all these methods consider the two distributions to be equally important, which could lead to poor classification performance in practical applications. Therefore, a balanced factor is required to weigh the two distributions to compensate for the degraded performance. In this paper, we first introduce this balance factor to weigh the distribution importance. On this base, we utilize the marginal distribution, introduce the ideas of manifold regularization, and then preserve the neighboring structures of the data sets, with the dimension reduction as much as possible. By this way, we propose the manifold alignment and balanced distribution adaptation algorithm. A large number of experiments have also been conducted, showing that our algorithm behaves much better than the previous ones.
Ying Li 0028, Yaxin Peng, Zhijie Wen, Shihui Ying
ICME4
2019 Asymmetric Local Metric Learning with PSD Constraint for Person Re-identification
abstract
Person re-identification is one of the key issues in both machine learning and video monitor application. In particular, defining an appropriate distance metric between the person images is very important. Existing metric learning approaches used in person re-identification either learn a single measure, or ignore the positive semi-definite (PSD) of measurement matrix, at the same time, since the number of negative sample pairs largely exceeds the number of positive sample pairs, some metric learning methods are largely influenced by the sample imbalance. Considering the above issues, we propose a new adaptive local metric learning method with positive semi-definite (PSD) constraint. Unlike existing metric learning methods which learn a single distance metric, we use an approximation error bound of a smooth metric matrix function over the data manifold to learn local metrics as linear combinations of basis metrics defined on anchor points over different regions of the instance space. Besides, we develop an efficient two stage algorithm that first learns the anchor points and the linear combinations of each instance, then learns the metric matrices of the anchor points. We employ the fast iterative shrinkage-thresholding algorithm which is a fast first-order optimization algorithm in the learning process of the linear combinations as well as the basis metrics of the anchor points. Our metric learning method has excellent performance. We firstly apply the proposed method on 5 UCI databases, which are widely used in machine learning, to test and evaluate the effectiveness of the proposed method. Then the proposed approach is applied for person re-identification, achieving better performance on three challenging databases (GRID, VIPeR, CUHK01) than the existing methods. The experimental results show that the proposed method can prvide the theoretical and practical support for the person re-identification problem.
Zhijie Wen, Ying Li 0028, Shihui Ying, Yaxin Peng
ICRA1
2018 Feasibility Study of Ore Classification Using Active Hyperspectral LiDAR
abstract
Recently, a major effort has been made to develop methods or tools for rock characterization and mineral content mapping. Light detection and ranging (LiDAR) is an efficient active remote sensing technique for collecting geometry information about rock surfaces. However, traditional LiDAR sensors work with a single-wavelength laser source, and it is unfeasible to obtain spectral information using one LiDAR sensor. The combination of hyperspectral imaging and LiDAR techniques is an emerging method for acquiring spatial and spectral information simultaneously that allows remote mapping of high-resolution mineral content and distributions and identifies subtle chemical variations. Unfortunately, spatial and spectral data registration, which introduces additional complicated data processing, is an inevitable and essential issue for this method. In this letter, first, we investigate the feasibility of ore classification applications with hyperspectral LiDAR (HSL). HSL consists of 17 spectral channels covering the visible–shortwave infrared (SWIR) spectral range. Spatial and spectral information about seven different ore samples is obtained under a controlled laboratory environment using HSL. The standard deviation of the distance measurements is less than 1.1 cm for different spectral channels, and the classification accuracy can reach 100% if all 17 spectral measurements are used. To optimize the system design with lower cost and system complexity, a spectral band selection criterion is built based on the feature contribution degree (FCD), which is calculated using the normalized variance of the reflectance values for different ore samples at each wavelength. Two different strategies of FCD selection are tested to generate vectors: ascending sequences and descending sequences. Feature vectors with descending sequences have better classification accuracy. In addition, the results show that the classification accuracy can reach 100% with the feature vector of the seven largest FCD values compared to 59.57% for the feature vector with the seven smallest FCD values. Moreover, we find that the channels with high FCD values are primarily centered in SWIR bands. This result could be a reference for optimizing the hardware design of HSL for ore classification or mineral identification.
Yuwei Chen 0005, Changhui Jiang, Juha Hyyppä, Shi Qiu 0002, Zheng Wang 0054, Mi Tian 0005, Wei Li 0095, Eetu Puttonen, Hui Zhou 0013, Yuming Bo, Zhijie Wen
IEEE Geosci. Remote. Sens. Lett.12
2018 Manifold Preserving: An Intrinsic Approach for Semisupervised Distance Metric Learning
abstract
In this paper, we address the semisupervised distance metric learning problem and its applications in classification and image retrieval. First, we formulate a semisupervised distance metric learning model by considering the metric information of inner classes and interclasses. In this model, an adaptive parameter is designed to balance the inner metrics and intermetrics by using data structure. Second, we convert the model to a minimization problem whose variable is symmetric positive-definite matrix. Third, in implementation, we deduce an intrinsic steepest descent method, which assures that the metric matrix is strictly symmetric positive-definite at each iteration, with the manifold structure of the symmetric positive-definite matrix manifold. Finally, we test the proposed algorithm on conventional data sets, and compare it with other four representative methods. The numerical results validate that the proposed method significantly improves the classification with the same computational efficiency.
Shihui Ying, Zhijie Wen, Jun Shi 0004, Yaxin Peng, Hong Qiao
IEEE Trans. Neural Networks Learn. Syst.2
2017 Fabric defect inspection using prior knowledge guided least squares regression
Junjie Cao 0001, Jie Zhang 0056, Zhijie Wen, Xiuping Liu
Multim. Tools Appl.3
2016 Compute Karcher means on SO(n) by the geometric conjugate gradient method
Shihui Ying, Han Qin, Yaxin Peng, Zhijie Wen
Neurocomputing4
2016 Nonlinear 2D shape registration via thin-plate spline and Lie group representation
Shihui Ying, Yuanwei Wang, Zhijie Wen, Yuping Lin
Neurocomputing3
2016 Virus image classification using multi-scale completed local binary pattern features extracted from filtered images by multi-scale principal component analysis
Zhijie Wen, Zhuojun Li, Yaxin Peng, Shihui Ying
Pattern Recognit. Lett.1
2011 Iwasawa decomposition: a new approach to 2D affine registration problem
Shihui Ying, Yaxin Peng, Zhijie Wen
Pattern Anal. Appl.3