Shijie Yu

dblp:51/8789 · DBLP profile ↗
← Back
12ranked-venue papers
2as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Evidential Uncertainty Modulated Adaptive Predictive Contrastive Learning for Multimodal Fusion
abstract
Multimodal learning has achieved remarkable success by integrating heterogeneous information from multiple sources. However, most existing methods either treat all samples uniformly or rely on deterministic prediction correctness to guide cross-modal alignment, overlooking the varying degrees of epistemic uncertainty inherent in each modality. Such assumptions often lead to noise propagation, especially when models exhibit over-confident yet unreliable predictions. To address this limitation, we propose Adaptive Predictive Contrastive Learning (AdaPCL), an evidential uncertainty modulated framework designed to regulate cross-modal interactions based on modality-specific reliability. Specifically, we first employ evidential deep learning to explicitly quantify modality-specific epistemic uncertainty, establishing a dual-view assessment that calibrates discriminative confidence with evidential reliability. Building upon this, we introduce a continuous reliability-modulated mechanism that synthesizes these dual perspectives to assign soft, instance-specific weights. This design enables AdaPCL to formulate a unified and adaptive contrastive objective encompassing three complementary alignment strategies: (i) Symmetric Alignment for mutually reliable modality pairs, (ii) Directional Distillation where a reliable modality provides pseudo-supervision for its uncertain counterpart, and (iii) Reliability-Aware Slack Regularization that adaptively attenuates the influence of mutually unreliable samples without enforcing rigid geometric constraints. Extensive experiments on multiple benchmark datasets demonstrate that AdaPCL consistently outperforms baseline multimodal classification methods. Code is available at https://github.com/yuhongcqupt/AdaPCL.
Qiuyu Mei, Hong Yu 0007, Shijie Yu
ICMR3
2026 Self-supervised multi-view multi-label learning with attention mechanisms
Qiuyu Mei, Hong Yu 0007, Shijie Yu, Guoyin Wang 0001
Multim. Syst.3
2025 Orthogonal Constrained Minimization with Tensor \(\ell_{2,{p}}\) Regularization for HSI Denoising and Destriping
abstract
Abstract. Hyperspectral images (HSIs) are often contaminated by a mixture of noise such as Gaussian noise, dead lines, stripes, and so on. In this paper, we propose a multiscale low-rank tensor regularized [Formula: see text] (MLTL2p) approach for HSI denoising and destriping, which consists of an orthogonal constrained minimization model and an iterative algorithm with convergence guarantees. The model of the proposed MLTL2p approach is built based on a new sparsity-enhanced Multiscale Low-rank Tensor regularization and a tensor [Formula: see text] norm with [Formula: see text]. The multiscale low-rank regularization for HSI denoising utilizes the global and local spectral correlation as well as the spatial nonlocal self-similarity priors of HSIs. The corresponding low-rank constraints are formulated based on independent higher-order singular value decomposition with sparsity enhancement on its core tensor to prompt more low-rankness. The tensor [Formula: see text] norm for HSI destriping is extended from the matrix [Formula: see text] norm. A proximal block coordinate descent algorithm is proposed in the MLTL2p approach to solve the resulting nonconvex nonsmooth minimization with orthogonal constraints. We show any accumulation point of the sequence generated by the proposed algorithm converges to a first-order stationary point, which is defined using three equalities of substationarity, symmetry, and feasibility for orthogonal constraints. In the numerical experiments, we compare the proposed method with state-of-the-art methods, including a deep learning based method, and test the methods on both simulated and real HSI datasets. Our proposed MLTL2p method demonstrates outperformance in terms of metrics such as mean peak signal-to-noise ratio as well as visual quality.
Shijie Yu, Jian Lu 0002, Xiaojun Chen 0001
SIAM J. Imaging Sci.2
2024 EHIR: Energy-based Hierarchical Iterative Image Registration for Accurate PCB Defect Detection
Shuixin Deng, Xiangze Meng, Ting Sun 0005, Baohua Chen, Zhixiang Chen 0003, Yusen Xie, Hanxi Yin, Shijie Yu
Pattern Recognit. Lett.10
2024 AirGeoNet: A Map-Guided Visual Geo-Localization Approach for Aerial Vehicles
abstract
Aerial vehicles (AVs) commonly operate in vast environments, presenting a persistent challenge in achieving high-precision localization. The contemporary popular global positioning methods have their inherent limitations. For instance, the precision of GPS is susceptible to decline or even complete failure when the signal is disrupted or absent. Furthermore, the precision of image retrieval techniques is inadequate. The construction of 3-D models is a time-consuming and storage-intensive endeavor. In addition, scene coordinate regression necessitates retraining to adapt to varying scenarios, which presents challenges when attempting to generalize across expansive environments. Addressing these challenges, we propose a network named AirGeoNet, which integrates satellite images and semantic maps to achieve high-precision efficient localization. In the first phase, we introduce the foundation model DINOV2 to extract features from satellite and aerial images, employ a vector of locally aggregated descriptor (VLAD) for image retrieval to get coarse position, and, finally, significantly enhance retrieval accuracy by combining sequential images with particle filters. Subsequently, AirGeoNet matches aerial images with semantic maps to determine the three degrees of freedom in pose, including position and orientation. The semantic maps utilized by AirGeoNet are sourced from OpenStreetMap and our self-produced QMap, and training is conducted in a supervised manner using real camera poses. Our AirGeoNet method is highly efficient, requiring only a 1546-D feature vector per image for image retrieval and 240k storage for a 0.9-$\text {km}^{2}$semantic map while achieving state-of-the-art accuracy with single-frame localization errors of 2.854 m on semantically rich datasets and 11 m in complex scenarios. Our code is publicly available athttps://github.com/mxz520mxz/AirGeoNet.git
Xiangze Meng, Wulong Guo, Ting Sun 0003, Shijie Yu
IEEE Trans. Geosci. Remote. Sens.6
2023 COCAS+: Large-Scale Clothes-Changing Person Re-Identification With Clothes Templates
abstract
Recent years person re-identification (ReID) has been developed rapidly due to its broad practical applications. Most existing benchmarks assume that the same person wears the same clothes across captured images, while, in real-world scenarios, person may change his/her clothes frequently. Thus the Clothes-Changing person ReID (CC-ReID) problem is introduced and several related benchmarks are established. CC-ReID is a very difficult task as the main visual characteristics of a human body, clothes, are different between query and gallery, and clothes-irrelevant features are relatively weak. To promote the research and applications of person ReID in clothes-changing scenarios, in this paper, we introduce a new task called Clothes Template based Clothes-Changing person ReID (CTCC-ReID), where the query image is enhanced by a clothes template which shares similar visual patterns with the clothes of the target person image in the gallery. So, ReID methods are encouraged to jointly consider the original query image and the given clothes template for retrieval in the proposed CTCC-ReID setting. To facilitate research works on CTCC-ReID, we construct a novel large-scale ReID dataset named ClOthes ChAnging person Set Plus (COCAS+), which contains both realistic and synthetic clothes-changing person images with manually collected clothes templates. Furthermore, we propose a novel Dual-Attention Biometric-Clothes Transfusion Network (DualBCT-Net) for CTCC-ReID, which can effectively learn to extract biometric features from the original query person image and clothes features from the given clothes template and then fuse them through a Dual-Attention Fusion Module. Extensive experimental results show that the proposed CTCC-ReID setting and COCAS+ dataset can help greatly push the performance of clothes-changing ReID toward practical applications, and synthetic data is impressively effective for CTCC-ReID. What’s more, the proposed DualBCT-Net shows significant improvements over state-of-the-art methods on the CTCC-ReID task. COCAS+ and code of DualBCT-Net will be released inhttps://github.com/Chenhaobin/COCAS-plus.
Shihua Li 0006, Shijie Yu, Zhiqun He, Feng Zhu 0006, Rui Zhao 0001, Jie Chen 0012, Yu Qiao 0001
IEEE Trans. Circuits Syst. Video Technol.3
2021 Layerwise Optimization by Gradient Decomposition for Continual Learning
abstract
Deep neural networks achieve state-of-the-art and sometimes super-human performance across various domains. However, when learning tasks sequentially, the networks easily forget the knowledge of previous tasks, known as "catastrophic forgetting". To achieve the consistencies between the old tasks and the new task, one effective solution is to modify the gradient for update. Previous methods enforce independent gradient constraints for different tasks, while we consider these gradients contain complex information, and propose to leverage inter-task information by gradient decomposition. In particular, the gradient of an old task is decomposed into a part shared by all old tasks and a part specific to that task. The gradient for update should be close to the gradient of the new task, consistent with the gradients shared by all old tasks, and orthogonal to the space spanned by the gradients specific to the old tasks. In this way, our approach encourages common knowledge consolidation without impairing the task-specific knowledge. Furthermore, the optimization is performed for the gradients of each layer separately rather than the concatenation of all gradients as in previous works. This effectively avoids the influence of the magnitude variation of the gradients in different layers. Extensive experiments validate the effectiveness of both gradient-decomposed optimization and layer-wise updates. Our proposed method achieves state-of-the-art results on various benchmarks of continual learning.
Shixiang Tang, Dapeng Chen, Jinguo Zhu, Shijie Yu, Wanli Ouyang
CVPR4
2021 Complementary Relation Contrastive Distillation
abstract
Knowledge distillation aims to transfer representation ability from a teacher model to a student model. Previous approaches focus on either individual representation distillation or inter-sample similarity preservation. While we argue that the inter-sample relation conveys abundant information and needs to be distilled in a more effective way. In this paper, we propose a novel knowledge distillation method, namely Complementary Relation Contrastive Distillation (CRCD), to transfer the structural knowledge from the teacher to the student. Specifically, we estimate the mutual relation in an anchor-based way and distill the anchor-student relation under the supervision of its corresponding anchor-teacher relation. To make it more robust, mutual relations are modeled by two complementary elements: the feature and its gradient. Furthermore, the low bound of mutual information between the anchor-teacher relation distribution and the anchor-student relation distribution is maximized via relation contrastive loss, which can distill both the sample representation and the inter-sample relations. Experiments on different benchmarks demonstrate the effectiveness of our proposed CRCD.
Jinguo Zhu, Shixiang Tang, Dapeng Chen, Shijie Yu, Yakun Liu, Mingzhe Rong, Aijun Yang, Xiaohua Wang 0001
CVPR4
2020 COCAS: A Large-Scale Clothes Changing Person Dataset for Re-Identification
abstract
Recent years have witnessed great progress in person re-identification (re-id). Several academic benchmarks such as Market1501, CUHK03 and DukeMTMC play important roles to promote the re-id research. To our best knowledge, all the existing benchmarks assume the same person will have the same clothes. While in real-world scenarios, it is very often for a person to change clothes. To address the clothes changing person re-id problem, we construct a novel large-scale re-id benchmark named Clothes Changing Person Set (COCAS), which provides multiple images of the same identity with different clothes. COCAS totally contains 62,382 body images from 5,266 persons. Based on COCAS, we introduce a new person re-id setting for clothes changing problem, where the query includes both a clothes template and a person image taking another clothes. Moreover, we propose a two-branch network named Biometric-Clothes Network (BC-Net) which can effectively integrate biometric and clothes feature for re-id under our setting. Experiments show that it is feasible for clothes changing re-id with clothes templates.
Shijie Yu, Shihua Li 0006, Dapeng Chen, Rui Zhao 0001, Yu Qiao 0001
CVPR1
2016 The monitoring of land use and land cover change of Sichuan province and Chengdu district, China
abstract
Land use and land cover change (LUCC) is necessary to explore the factors leading to heavy drought and rainy-flood disaster in some districts of Sichuan province. A method based RS, GIS, GPS and Google earth (GE) is presented to establish LUCC database in Sichuan province and Chengdu district. At first, LUCC is interpreted based on the new temporal images and the land use and land cover database from TM in 2000.Secondly, some ground objects, which could not be identified in the new temporal images, were interpreted utilizing GE with some higher spatial resolution images. Thirdly, the new interpreted LUCC was validated in the field with GPS handheld receiver. Then, LUCC of Sichuan province was updated. A comparative analysis of LUCC between in Sichuan province and in Chengdu district was conducted and the result showed: (1) a large amount of farmland in Sichuan Province was occupied from 2000 to 2005 and the area is 84 573 ha. While construction land gained obviously and the area was 35 828 ha. The dynamic degree of construction land was 111.100/00from 2000 to 2005. The LUCC demonstrated that the economy of Sichuan province continued to develop, the cities were overspreading and the urban heat island effect was deteriorated from 2000 to 2005. (2) A large amount of farmland was also occupied in Chengdu district from 2000 to 2005, the area amounted to 12 989 ha. The farmland lost was mainly changed to construction land, amounting to 93%. And the dynamic degree was 117.410/00from 2000 to 2005, which was bigger than that in Sichuan province.
Shijie Yu, Zezhong Zheng, Wunian Yang, Mingcang Zhu, Yong He 0007, Zhenlu Yu, Shengli Wang, Jiang Li 0001
IGARSS1
2016 The manifold learning for dimensionality reduction with hyperspectral image
abstract
Hyperspectral remote sensing image (HSI) consists of hundreds of bands that contain rich space, radiation and spectral information. The high-dimensional data can also lead to the curse of dimensionality problem making it difficult to be used effectively. In this paper, we proposed a manifold learning algorithm to reduce the dimensionality for HSI data. For high dimensional datasets with continuous variables, it is often the case that the data points are arranged along with low dimensional structures, named manifolds, in the high dimensional space. Manifold learning aims to identifying those special low dimensional structures for subsequent usage such as classification or regression. However, many manifold learning algorithms perform an eigenvector analysis on a data similarity matrix whose size is N×N, where N is the number of data points. The memory complexity of the analysis is at least O(N2) that is not feasible for a regular computer to compute or storage for very large datasets. To solve this problem, we used statistical sampling methods to sample a subset of data points as landmarks. A skeleton of the manifold was then identified based on the landmarks. The remaining data points were then inserted into the skeleton by Locally Linear Embedding (LLE). We tested our algorithm on AVIRIS Salinas-A data set. The experimental results showed that the HSI dataset could be reduced to a lower-dimensional space for land use classification with good performance, and the main structure was preserved well.
Zezhong Zheng, Pengxu Chen, Mingcang Zhu, Zhiqin Huang, Yong He 0007, Yicong Feng, Yufeng Lu, Zhenlu Yu, Shijie Yu, Shengli Wang, Jiang Li 0001
IGARSS9
2016 The tradeoff of accuracy with different landmarks with manifold learning
abstract
High-dimensional data such as hyperspectral images contain abundant information of surface radiation. But the massive redundant information makes it complex to be utilized conveniently. To solve this problem, a manifold learning dimensionality reduction framework for hyperspectral image is proposed. Firstly, statistical sampling methods were used to sample a subset of data points as landmarks. A skeleton of the manifold was then identified basing on the landmarks. The remaining data points were then inserted into the skeleton by Locally Linear Embedding algorithm. At last, original data sets and data sets reduced with different manifold learning approaches were classified by KNN classifier to evaluate the performance of the proposed framework. The framework was tested on AVIRIS Salinas-A dataset. The experimental results showed that the tradeoff of accuracy with different landmarks is of great significant. Insufficient landmarks lead to low accuracy and excess landmarks may spend a considerable amount of time.
Zezhong Zheng, Chengjun Pu, Mingcang Zhu, Zhiqin Huang, Yong He 0007, Yicong Feng, Yufeng Lu, Zhenlu Yu, Shengli Wang, Shijie Yu, Jiang Li 0001
IGARSS10