Chen Wang 0026

dblp:82/4206-26 · DBLP profile ↗
← Back
21ranked-venue papers
6as first author
15since 2021 · last 2026
0000-0002-1728-4076ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 3 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Dual-label noise filtering for weakly supervised person search
Huadong Lin, Pengcheng Zhang 0003, Xiaohan Yu 0001, Xiao Bai 0001, Chen Wang 0026
Neurocomputing5
2025 Guest Editorial: Multi-view representation learning for computer vision
abstract
Object recognition and scene analysis in single-view images may face difficulties such as occlusion and incomplete information, while multi-view learning can address this limitation. When an object or scene is observed from multiple views, information on target objects can be significantly enriched to improve the performance of computer vision tasks. For this reason, multi-view has become one of the important forms of data representation, which leads to the emerging of new research topics on complete or in-complete multi-view learning. Multi-view learning enables the use of multi-source information, nevertheless, the heterogeneous characteristics of data make it difficult to reliably associate information from different views, especially in a complex environment. It remains challenging for tasks to make effective use of the consistent and complementary information between different complete views and to enhance the completeness of potential representation. A wide variety of research is being conducted to explore and discover possible challenges and opportunities to exploit multi-view representation learning for computer vision. The purpose of this Special Issue is to collect high-quality articles on the recent development and trend of multi-view representation learning in computer vision, publish new ideas, theories, solutions and insights on this topic, and showcase their applications. In this Special Issue, we have received 36 papers, all of which underwent peer review. Of the 36 originally submitted papers, 10 have been accepted, which cover a variety of fields, such as person re-identification, gait recognition, 3D object recognition, and behaviour recognition. These accepted papers are mainly divided into three categories. The first category covers the incomplete multi-view data learning theoretics and methods. The papers in this category are of He et al., Kun et al., Fan et al. and Wang et al. The last two categories are both multi-view applications. One of which is 3D-related applications. The papers in this category are of Qi et al. and Sun et al. The other category is about 2D recognition. The papers in this category are of Zhang et al., Huang et al., Zheng et al. and Zhang et al. A brief presentation of each of the paper follows. He et al. present an innovative multi-view subspace clustering method with incomplete graph information. Specifically, they separate one shared and multiple specific graphs from multiple raw graph data, and exploit the mask fusion strategy and block diagonal regulariser to obtain the inherent category information. The clustering results on six real-world datasets show that the method outperforms a series of classic incomplete multi-view clustering methods. Kun et al. propose a new method for low-rank-based multi-view subspace clustering based on low-rank correlation analysis. To overcome the limitations of unreliable low-rank structure and imprecise graphs caused by multi-view noise and outliers, they introduce the canonical correlation analysis strategy and a dual regularisation term to characterise the connections between different views adaptively. Experimental results reveal the method's superiority over compared state-of-the-art (SOTA) methods in accuracy, normalised mutual information, and F-score evaluation metrics. Fan et al. address the challenge of partial mapping between the views in multi-view clustering, and propose a self-inferring incomplete multi-view clustering algorithm to explore the information hidden in the local geometric structure and recover missing instances through mining the information hidden in existing instances. Experimental results show that the method can improve the clustering performance compared with the SOTA methods. Wang et al. propose a semi-paired semi-supervised deep hashing to solve the large-scale multimedia retrieval task. The method is an end-to-end deep neural network model with high-order affinity. To maintain the consistency within the modalities, they introduce a common representation that combines with the labelled information to associate different modalities. Experimental results demonstrate the superior performance of proposed method. Qi et al. propose a double-weighting convolution neural network based on the L2-S grouping mechanism for multi-view 3D object recognition. The goal of the proposed L2-S grouping mechanism is to calculate the discrimination score of views and group views more reasonably. Results of the experiments show that the method can achieve SOTA performance. Sun et al. present a dual-matching method with cross-attention mechanism to address the limitations of matching-based methods caused by a preset fixed disparity range on depth estimation task. To tackle the mismatches on edges and details, they introduce an exquisite module based on left-right consistency. The method is proved to be competitive and effective by experiments conducted under popular benchmarks. Zhang et al. want to answer the following two questions: (1) does a query image with higher resolution than that of the gallery image also affect the pedestrian re-identification performance? If so, and (2) how does it affect performance? So, they propose an end-to-end trainable resolution independent person re-identification network that is composed of a cross-resolution Generative Adversarial Networks and embedding batch normalisation layers. The results demonstrate that the proposed method outperforms the SOTA methods in the pedestrian re-identification task on their expanded benchmark dataset. Huang et al. address the limitation of current gait-based age and gender recognition methods under multi-view scene, and propose an attention-aware spatio–temporal learning framework that employs silhouette sequence as an input to learn essential spatial–temporal gait representation. The proposed method has produced results that outperformed the benchmarks with an Mean Absolute Error of 6.68 years for age estimation and a Correct Classification Rate of 97% for gender classification. Zheng et al. apply deep learning to multi-view classroom behaviour detection. First, they propose an improved detection model based on YOLOv5 to improve the convergence speed of the prediction box. Second, they establish a quantitative evaluation standard for students' classroom attention, and then conduct training and verification by collecting multi-view classroom datasets. Finally, they increase the environment variation in the training model phase to make the model have better generalisation ability. Experiments demonstrate that the method can effectively identify and detect students' behaviours in the classroom from different views. Zhang et al. propose a method for multi-dimensional video anomaly detection, which uses the Object-meta instead of video frames as the input, and the Memory Search Guided Autoencoder with Memory Pools (MSGAE-MP) to reconstruct. The multi-dimensional information carried by the input can be strengthened via Object-meta. The MSGAE-MP construct multi-level memory pools, so as to reconstruct Object-meta in different dimensions. Experiments show that the method is feasible and has achieved excellent results. All of the papers published in this Special Issue show that multi-view representation learning theoretics have developed very fast in recent years. In addition, it is very promising to solve traditional computer vision tasks under multi-view setting, including but not limited to 3D object recognition, person re-identification, gait-based age and gender estimation, and depth estimation. Xin Ning and Chen Wang are responsible for the writing of Proposal and Editorial materials; Jun Zhou is responsible for the processing of articles; and Jing Wu, Lin Gu and Jian Cheng are responsible for the solicitation and publicity of the special issue. Firstly, we would like to thank all the authors for their innovative contributions and all the reviewers for their professional and crucial, yet constructive comments. Also, we wish to express our thanks to Mr Hang Ran, PhD students at Institute of Semiconductors, Chinese Academy of Sciences, for his assistance in this process. Last, we wish to express our gratitude to the editorial team of IET Computer Vision for their support throughout this venture. We hope you enjoy this collection of papers and that the Special Issue can stimulate further research and development in this area. This work is supported by the National Natural Science Foundation of China (Grant no. 61901436). National Natural Science Foundation of China, Grant/Award Number: 61901436. Data sharing is not applicable to this article as no new data were created or analyzed in this study. Xin Ning (SMIEEE) received a B.S. degree in software engineering in 2012, and a Ph.D. degree in electronic circuit and system from the university of Chinese Academy of Sciences, in 2017. He is currently an associate professor with the Laboratory of Artificial Neural Networks and High Speed Circuits, Institute of Semiconductors, Chinese Academy of Sciences. His current research interests include neural networks, intelligent systems and computer vision. He has published as the first or corresponding author in more than 45 papers in journals and refereed conferences. Now he serves as the young associated editor of CAAI Transactions on Intelligent Systems, the guest editor of Elsevier Journal on DISPLAYS. He is also the guest editor of CONNECTION SCIENCE and CONCURR COMP-PRACT E. He was the Website Chair of the IEEE HPBD&IS 2020 and the Publication Chair of the IEEE HPBD&IS 2021. Jun Zhou received a B.S. degree in computer science and a B.E. degree in international business from the Nanjing University of Science and Technology, Nanjing, China, in 1996 and 1998, respectively, an M.S. degree in computer science from Concordia University, Montreal, QC, Canada, in 2002, and a Ph.D. degree in computing science from the University of Alberta, Edmonton, AB, Canada, in 2006. He was a research fellow with the Research School of Computer Science, The Australian National University, Canberra, ACT, Australia, and a researcher with the Canberra Research Laboratory, National Information and Communications Technology Australia, Canberra. In 2012, he joined the School of Information and Communication Technology, Griffith University, Nathan, QLD, Australia, where he is currently a reader. His research interests include pattern recognition, computer vision, and spectral imaging and their applications in remote sensing and environmental informatics. He is the associate editor for the journal of Pattern Recognition and IEEE Trans. on Remote Sensing. Jian Cheng is a professor of Institute of Automation, Chinese Academy of Sciences. He received the B.S. and M.S. degrees in Mathematics from Wuhan University in 1998 and 2001, respectively. After that, he received a Ph.D degree in pattern recognition and intelligent systems from Institute of Automation, Chinese Academy of Sciences in 2004. His current major research interests include deep learning, computer vision, chip design, etc. Jing Wu is now a postdoc at the school of computer science, Beihang University. He received his B.E. degree from the school of computer science, Northwestern Polytechnical University in 2013 and received his PhD. degree from the school of computer science, Beihang University in 2021. His research interests include computer vision, stereo matching, 3D reconstruction and camera localization. Chen Wang is now a postdoc at the school of computer science, Beihang University. He received his B.E. degree from the school of computer science, Northwestern Polytechnical University in 2013 and received his PhD. degree from the school of computer science, Beihang University in 2021. His research interests include computer vision, stereo matching, 3D reconstruction and camera localization. Lin Gu received a B.Eng. degree from Shanghai University, Shanghai, China, in 2009, and a Ph.D. degree in computer vision from Australian National University in 2014. After Ph.D. graduation from the Australian National University, he worked as a post-doctoral researcher at A*STAR, Singapore. Then, he was a project researcher with the National Institute of Informatics, Japan, and also a visiting scholar with Kyoto University, Japan. He is currently a research scientist at RIKEN AIP, Japan, and a special researcher with the University of Tokyo, Japan. He is also an in-charge of a Moonshot and an ACT-X Project to improve artificial intelligence by simulating the human brain. His primary research interests lie in machine learning, medical imaging, and computational photography.
Xin Ning 0001, Jun Zhou 0001, Jian Cheng 0001, Jing Wu 0004, Chen Wang 0026, Lin Gu 0003
IET Comput. Vis.5
2024 Joint discriminative representation learning for end-to-end person search
Pengcheng Zhang 0003, Xiaohan Yu 0001, Xiao Bai 0001, Chen Wang 0026, Xin Ning 0001
Pattern Recognit.4
2024 Towards effective person search with deep learning: A survey from systematic perspective
Pengcheng Zhang 0003, Xiaohan Yu 0001, Chen Wang 0026, Xin Ning 0001, Xiao Bai 0001
Pattern Recognit.3
2023 Adaptive Cost Aggregation in Iterative Depth Estimation for Efficient Multi-view Stereo
Xiang Wang 0014, Xiao Bai 0001, Chen Wang 0026
ICIG (2)3
2023 MFE-ResNet: A new extraction framework for land cover characterization in mining areas
Chen Wang 0026, Tao Chen 0004, Antonio Plaza
Future Gener. Comput. Syst.1
2023 Makeup transfer: A review
abstract
Abstract Makeup transfer (MT) aims to transfer the makeup style from a given reference makeup face image to a source image while preserving face identity and background information. In recent years, MT has attracted the attention of many scholars, and it has a wide range of application prospects and research value. Since then, many methods have been proposed to accomplish MT, most of which are based on Generative Adversarial Network methods. A taxonomy of existing algorithms in the field of MT is first proposed. Then, evaluation methods are proposed, existing methods are analysed, and existing datasets are introduced. This paper finally discusses the current problems in the field of MT and the trend of future research.
Feng He 0008, Kai Bai, Yixin Zong, Yimai Jing, Guoqiang Wu, Chen Wang 0026
IET Comput. Vis.7
2023 ACGAN: Age-compensated makeup transfer based on homologous continuity generative adversarial network model
abstract
Abstract The authors focus on the makeup transformation problem, which refers to the transfer of makeup from a reference face to a source face image while maintaining the source makeup‐free face image. In recent years, makeup transformation has become a hot issue and a lot of research has been conducted on this basis, but there are some limitations in the existing methods, mainly due to the lack of consideration of age factor, which makes the final generated face makeup images appear not natural and lack appearance attractiveness. In order to further solve this problem, an age‐compensated makeup transformation framework based on homology continuity is proposed. In order to achieve a stable and controllable age‐compensation effect, the authors design a new coding module that can map the face makeup semantic vector into the higher feature space and achieve age compensation by adjusting the direction of the semantic vector. Finally, in order to comprehensively evaluate the effectiveness of the authors’ proposed method, a large number of qualitative and quantitative experiments have been conducted, and the experimental results show that the authors’ proposed framework outperforms existing methods.
Guoqiang Wu, Feng He 0008, Yimai Jing, Xin Ning 0001, Chen Wang 0026, Bo Jin 0018
IET Comput. Vis.6
2023 A smoothing Group Lasso based interval type-2 fuzzy neural network for simultaneous feature selection and system identification
Tao Gao 0003, Chen Wang 0026, Guoqiang Wu, Xin Ning 0001, Xiao Bai 0001, Jian Wang 0010
Knowl. Based Syst.2
2023 Learning consistent region features for lifelong person re-identification
Jinze Huang, Xiaohan Yu 0001, Dong An 0001, Yaoguang Wei, Xiao Bai 0001, Chen Wang 0026, Jun Zhou 0001
Pattern Recognit.7
2022 Revisiting Domain Generalized Stereo Matching Networks from a Feature Consistency Perspective
abstract
Despite recent stereo matching networks achieving impressive performance given sufficient training data, they suffer from domain shifts and generalize poorly to unseen domains. We argue that maintaining feature consistency between matching pixels is a vital factor for promoting the generalization capability of stereo matching networks, which has not been adequately considered. Here we address this issue by proposing a simple pixel-wise contrastive learning across the viewpoints. The stereo contrastive feature loss function explicitly constrains the consistency between learned features of matching pixel pairs which are observations of the same 3D points. A stereo selective whitening loss is further introduced to better preserve the stereo feature consistency across domains, which decorrelates stereo features from stereo viewpoint-specific style information. Counter-intuitively, the generalization of feature consistency between two viewpoints in the same scene translates to the generalization of stereo matching performance to unseen domains. Our method is generic in nature as it can be easily embedded into existing stereo networks and does not require access to the samples in the target domain. When trained on synthetic data and generalized to four real-world testing sets, our method achieves superior performance over several state-of-the-art networks. The code is available online11https://github.com/jiaw-z/FCStereo.
Xiang Wang 0014, Xiao Bai 0001, Chen Wang 0026, Lei Huang 0015, Lin Gu 0003, Jun Zhou 0001, Tatsuya Harada, Edwin R. Hancock
CVPR4
2022 A modified interval type-2 Takagi-Sugeno fuzzy neural network and its convergence analysis
Tao Gao 0003, Xiao Bai 0001, Chen Wang 0026, Liang Zhang 0044, Jian Wang 0010
Pattern Recognit.3
2022 Uncertainty estimation for stereo matching based on evidential deep learning
Chen Wang 0026, Xiang Wang 0014, Liang Zhang 0044, Xiao Bai 0001, Xin Ning 0001, Jun Zhou 0001, Edwin R. Hancock
Pattern Recognit.1
2021 Relation-Aware Reasoning with Graph Convolutional Network
Lei Zhou 0008, Yang Liu 0357, Xiao Bai 0001, Xiang Wang 0014, Chen Wang 0026, Liang Zhang 0044, Lin Gu 0003
ICIG (1)5
2021 Self-Supervised Multiscale Adversarial Regression Network for Stereo Disparity Estimation
abstract
Deep learning approaches have significantly contributed to recent progress in stereo matching. These deep stereo matching methods are usually based on supervised training, which requires a large amount of high-quality ground-truth depth map annotations that are expensive to collect. Furthermore, only a limited quantity of stereo vision training data are currently available, obtained either by active sensors (Lidar and ToF cameras) or through computer graphics simulations and not meeting requirements for deep supervised training. Here, we propose a novel deep stereo approach called the "self-supervised multiscale adversarial regression network (SMAR-Net)," which relaxes the need for ground-truth depth maps for training. Specifically, we design a two-stage network. The first stage is a disparity regressor, in which a regression network estimates disparity values from stacked stereo image pairs. Stereo image stacking method is a novel contribution as it not only contains the spatial appearances of stereo images but also implies matching correspondences with different disparity values. In the second stage, a synthetic left image is generated based on the left-right consistency assumption. Our network is trained by minimizing a hybrid loss function composed of a content loss and an adversarial loss. The content loss minimizes the average warping error between the synthetic images and the real ones. In contrast to the generative adversarial loss, our proposed adversarial loss penalizes mismatches using multiscale features. This constrains the synthetic image and real image as being pixelwise identical instead of just belonging to the same distribution. Furthermore, the combined utilization of multiscale feature extraction in both the content loss and adversarial loss further improves the adaptability of SMAR-Net in ill-posed regions. Experiments on multiple benchmark datasets show that SMAR-Net outperforms the current state-of-the-art self-supervised methods and achieves comparable outcomes to supervised methods. The source code can be accessed at: https://github.com/Dawnstar8411/SMAR-Net.
Chen Wang 0026, Xiao Bai 0001, Xiang Wang 0014, Xianglong Liu 0001, Jun Zhou 0001, Xinyu Wu 0001, Hongdong Li, Dacheng Tao
IEEE Trans. Cybern.1
2020 HMFlow: Hybrid Matching Optical Flow Network for Small and Fast-Moving Objects
abstract
In optical flow estimation task, coarse-to-fine warping strategy is widely used to deal with the large displacement problem and provides efficiency and speed. However, limited by the small search range between the first images and warped second images, current coarse-to-fine optical flow networks fail to capture small and fast-moving objects which has disappeared at coarse resolution levels. To address this problem, we introduce a lightweight but effective Global Matching Component (GMC) to grab global matching features. We propose a new Hybrid Matching Optical Flow Network (HMFlow) by integrating GMC into existing coarse-to-fine networks seamlessly. Besides keeping in high accuracy and small model size, our proposed HMFlow can apply global matching features to guide the network to discover the small and fast-moving objects mismatched by local matching features. We also build a new dataset, named SFChairs, for evaluation. The experimental results show that our proposed network achieves considerable performance, especially at regions with small and fast-moving objects.
Suihanjin Yu, Youmin Zhang 0005, Chen Wang 0026, Xiao Bai 0001, Liang Zhang 0044, Edwin R. Hancock
ICPR3
2019 Discriminative Features Matter: Multi-layer Bilinear Pooling for Camera Localization
Xiang Wang 0014, Chen Wang 0026, Xiao Bai 0001, Jing Wu 0004, Edwin R. Hancock
BMVC3
2019 Multiscale Visual Attention Networks for Object Detection in VHR Remote Sensing Images
abstract
Object detection plays an active role in remote sensing applications. Recently, deep convolutional neural network models have been applied to automatically extract features, generate region proposals, and predict corresponding object class. However, these models face new challenges in VHR remote sensing images due to the orientation and scale variations and the cluttered background. In this letter, we propose an end-to-end multiscale visual attention networks (MS-VANs) method. We use skip-connected encoder-decoder model to extract multiscale features from a full-size image. For feature maps in each scale, we learn a visual attention network, which is followed by a classification branch and a regression branch, so as to highlight the features from object region and suppress the cluttered background. We train the MS-VANs model by a hybrid loss function which is a weighted sum of attention loss, classification loss, and regression loss. Experiments on a combined data set consisting of Dataset for Object Detection in Aerial Images and NWPU VHR-10 show that the proposed method outperforms several state-of-the-art approaches.
Chen Wang 0026, Xiao Bai 0001, Shuai Wang 0049, Jun Zhou 0001, Peng Ren 0001
IEEE Geosci. Remote. Sens. Lett.1
2019 Self-Supervised deep homography estimation with invertibility constraints
Chen Wang 0026, Xiang Wang 0014, Xiao Bai 0001, Yun Liu 0014, Jun Zhou 0001
Pattern Recognit. Lett.1
2018 Predicting the popularity of topics based on user sentiment in microblogging websites
Xiang Wang 0015, Chen Wang 0026, Zhaoyun Ding, Jiumin Huang
J. Intell. Inf. Syst.2
2017 Deep Residual Convolutional Neural Network for Hyperspectral Image Super-Resolution
Chen Wang 0026, Yun Liu 0014, Xiao Bai 0001, Wenzhong Tang, Jun Zhou 0001
ICIG (3)1