VLDB 2026 Research / reviewers in the wild / expert
Xin Ning 0001
dblp:28/8118-1
· DBLP profile ↗
100ranked-venue papers
18as first author
92since 2021 · last 2027
0000-0001-7897-1673ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 56 · 10 first-author · 52 since 2021Graphics, computer vision, multimedia, augmented reality and games · 20 · 5 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 14 · 14 since 2021Systems, architecture and hardware · 7 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Computer networks · 3 · 3 since 2021Security and privacy · 1 · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | From sparse cues to rich semantics: Occluded person re-identification via non-local interaction and structural consistency learning
Enhao Ning, Wenfa Li, Sheng Xie, Yibo Lv, Liangtao Shi, Deepak Kumar Jain 0003, Libin Wu, Xin Ning 0001 |
Inf. Process. Manag. | 9 |
| 2026 | Deep reinforcement learning for energy-efficient workflow scheduling in edge computing
Mengyao Wen, Xiufeng Liu 0001, Xin Ning 0001, Cong Liu 0012, Jiawei Nian, Long Cheng 0003 |
Comput. Networks | 3 |
| 2026 | Occluded person re-identification in multi-scenarios: A Synergistic Interaction Framework with Perception-Aware Optimization
Enhao Ning, Junfeng Miao, Sheng Xie, Haifei Ma, Xin Ning 0001 |
Eng. Appl. Artif. Intell. | 5 |
| 2026 | Efficient large-scale road surface reconstruction via curvature-based frame selection
Hongjia Xing, Changshuo Wang 0001, Zaiyang Yu, Tingran Wang, Yuanjin Fang, Liping Zhang 0014, Xin Ning 0001 |
Eng. Appl. Artif. Intell. | 8 |
| 2026 | Counterfactual distribution intervention for few-shot class-incremental learning
Jicheng Yuan, Wenfa Li, Lusi Li, Liping Zhang 0014, Enhao Ning, Xingyu Gao 0001, Xin Ning 0001 |
Knowl. Based Syst. | 7 |
| 2026 | Advancing oral leukoplakia progression recognition: A benchmark with dataset, method, and application
Linfei Feng, Qiankun Li 0004, Hao Wang 0260, Xuanyu Li, Feng He 0008, Xin Ning 0001, Prayag Tiwari |
Neural Networks | 7 |
| 2026 | Flexible-Weighted Chamfer Distance: Enhanced Objective Function for Point Cloud CompletionabstractThe Chamfer Distance (CD) is a cornerstone objective function for point cloud completion, yet its inherent symmetric weighting mechanism limits the quality of the generated results. By penalizing local detail deviations and global coverage deficiencies equally, standard CD often causes structural defects such as point aggregation and incomplete spatial structures. We introduce the Flexible-weighted Chamfer Distance (FCD), which decouples CD into local precision and global completeness sub-objectives. FCD employs an asymmetric weighting strategy that prioritizes global structural integrity, steering the optimization away from sub-optimal solutions. As a plug-and-play module with negligible overhead, extensive experiments on state-of-the-art networks demonstrate that FCD significantly enhances global distribution metrics while preserving local precision. Specifically, on the ShapeNet55 benchmark using AdaPoinTr, FCD reduces the Density-aware Chamfer Distance (DCD) by approximately 12.4% (from 0.613 to 0.537), effectively mitigating point clustering. Similarly, on the PCN dataset, the proposed method reduces the Earth Mover's Distance (EMD) from 23.79 to 21.40, demonstrating superior global uniformity compared to the standard CD baseline. Furthermore, FCD demonstrates excellent generalization. When applied to diverse tasks and datasets, including real-world scans (KITTI), industrial components (ABC), and point cloud upsampling (PU-GAN), it yields significant quantitative gains and produces visually more uniform and structurally complete point clouds. These results underscore FCD's potential as a versatile objective function for the broader point cloud generation domain. Shengwei Tian, Long Yu 0001, Xin Ning 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2026 | DNGaussian++: Improving Sparse-View Gaussian Radiance Fields With Depth NormalizationabstractSynthesizing novel views from sparse views has achieved impressive advances with radiance fields, yet prevailing methods suffer from high consumption or insufficient refinement capability. This paper introduces DNGaussian, a depth-regularized framework based on 3D Gaussian Splatting, offering real-time and high-quality few-shot novel view synthesis at low costs. Our motivation stems from the remarkable advancement of recent 3D Gaussian Splatting, despite it will encounter a geometry degradation when input views decrease. In the Gaussian radiance fields, we find this degradation in scene geometry primarily lined to the positioning of Gaussian primitives and can be mitigated by depth constraint. Consequently, we propose a Hard and Soft Depth Regularization to restore accurate scene geometry under coarse monocular depth supervision while maintaining a fine-grained color appearance. To further refine detailed geometry, we introduce Global-Local Depth Normalization, enhancing the focus on small local depth changes. Although DNGaussian shows impressive performance, its patch-wise regularization obscures the inconsistency in cross-patch errors. Additionally, primitives can still be irreversibly trapped in local minima under sparse views, even if depth regularization is applied. In this paper, we propose an extended version, DNGaussian++. First, a Geometry Instance Regularizer is developed to enable depth regularization for continuous consistency by exploiting reliable instance-level depth cues. Leveraging the depth gradient guidance, we then propose a Depth-Guided Geometry Reorganization to address the aforementioned local minima problem with high representation efficiency. Extensive experiments show that DNGaussian++ exhibits state-of-the-art performance in multiple datasets and scenarios with high efficiency, and the broad applicability and effectiveness are verified on various backbones and tasks. Jiahe Li 0007, Xiaohan Yu 0001, Xiao Bai 0001, Xin Ning 0001, Lin Gu 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2026 | Neural Network Optimization Reimagined: Decoupled Techniques for Scratch and Fine-TuningabstractWith the accumulation of resources in the era of Big Data and the rise of pre-trained models in deep learning, optimizing neural networks for various tasks often involves different strategies for fine-tuning pre-trained models versus training from scratch. However, existing optimizers primarily focus on reducing the loss function by updating model parameters, without fully addressing the unique demands of these two major paradigms. In this paper, we propose DualOpt, a novel approach that decouples optimization techniques specifically tailored for these distinct training scenarios. For training from scratch, we introduce real-time layer-wise weight decay, designed to enhance both convergence and generalization by aligning with the characteristics of weight updates and network architecture. For more importantly fine-tuning, we integrate weight rollback with the optimizer, incorporating a rollback term into each weight update step. This ensures consistency in the weight distribution between upstream and downstream models, effectively mitigating knowledge forgetting and improving fine-tuning performance. Additionally, we extend the layer-wise weight decay to dynamically adjust the rollback levels across layers, adapting to the varying demands of different downstream tasks. Extensive experiments across diverse tasks, including image classification, object detection, semantic segmentation, and instance segmentation, demonstrate the broad applicability and state-of-the-art performance of DualOpt. Xin Ning 0001, Qiankun Li 0004, Xiaolong Huang 0001, Qiupu Chen, Feng He 0008, Weijun Li 0002, Prayag Tiwari, Xinwang Liu 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2026 | HybridEditDif: Text and exemplar guided image editing with diffusion models
Xuemei Fu, Long Cheng 0003, Jungong Han, Catarina Moreira, Xin Ning 0001, Xiao Bai 0001 |
Pattern Recognit. | 7 |
| 2026 | Causally Invariant Video Anomaly Detection via Counterfactual Reasoning and prototype intervention
Enhao Ning, Weixuan Gao, Jicheng Yuan, Shuyao He, Xin Ning 0001 |
Pattern Recognit. | 7 |
| 2026 | ABM: An Automatic Body Measurement framework via body deformation and topology-aware B-spline approximation
Xin Ning 0001, Limin Jiang, Liping Zhang 0014, Tingran Wang, Weijun Li 0002, Pengjiang Qian |
Pattern Recognit. | 1 |
| 2026 | FIRE: Fourier-series Implicit Neural Representations for high-fidelity continuous signal modeling
Jufeng Han, Shu Wei, Xin Ning 0001, Lusi Li, Hong Qin 0007, Weijun Li 0002 |
Pattern Recognit. | 4 |
| 2026 | Distribution entropy regularized multimodal subspace support vector data description for anomaly detection
Chuang Wang 0011, Xin Ning 0001, Pengjiang Qian, Jian Yao 0005, E. Y. K. Ng, Khin Wee Lai, Shitong Wang 0001 |
Pattern Recognit. | 2 |
| 2026 | Beyond discriminative features: Invariant Representation Learning for Few-Shot Class-Incremental Learning
Jicheng Yuan, Wenfa Li, Lusi Li, Liping Zhang 0014, Jijie Wu, Enhao Ning, Xin Ning 0001 |
Pattern Recognit. | 8 |
| 2026 | Energy-Efficient Online Continual Learning for Time Series Classification in Nanorobot-Based Smart HealthabstractNanorobots have been used in smart health to collect time series data such as electrocardiograms and electroencephalograms. Real-time classification of dynamic time series signals in nanorobots is a challenging task. Nanorobots in the nanoscale range require a classification algorithm with low computational complexity. First, the classification algorithm should be able to dynamically analyze time series signals and update itself to process the concept drift (CD). Second, the classification algorithm should have the ability to handle catastrophic forgetting (CF) and classify historical data. Most importantly, the classification algorithm should be energy-efficient to use less computing power and memory to classify signals in real-time on a smart nanorobot. To solve these challenges, we design an algorithm that can Prevent Concept Drift in Online continual Learning for time series classification (PCDOL). The prototype suppression item in PCDOL can reduce the impact caused by CD. It also solves the CF problem through the replay feature. The computation per second and the memory consumed by PCDOL are only 3.572 M and 1 KB, respectively. The experimental results show that PCDOL is better than several state-of-the-art methods for dealing with CD and CF in energy-efficient nanorobots. Le Sun 0003, Qingyuan Chen, Xin Ning 0001, Deepak Gupta 0002, Prayag Tiwari |
IEEE J. Biomed. Health Informatics | 4 |
| 2026 | PDR-FSCIL: prompt decoupling regularization with an optimized classifier for few-shot class-incremental learning
Meilan Hao, Yiren Cai, Yizhan Gu, Xin Ning 0001, Deepak Kumar Jain 0003 |
J. Supercomput. | 4 |
| 2025 | TFMSCNet: Dual Time-Frequency Modeling with Multi-Scale Feature Enhancement for EIT-Based Pulmonary Disease DiagnosisabstractEffective pulmonary disease monitoring and diagnosis are crucial for improving patient outcomes and guiding clinical interventions. Pulmonary Electrical Impedance Tomography (EIT) faces challenges due to its nonlinearity and complex data, which has led researchers to directly use raw EIT boundary voltage data for disease classification. However, the data's high non-stationarity, multi-scale time-frequency features, and long-term temporal dependencies limit traditional classification methods. To address these challenges, we propose the Time-Frequency Multi-Scale Convolutional Network (TFMSCNet), a novel architecture that simultaneously captures temporal dynamics and frequency patterns through parallel processing branches. The integrated Multi-Scale Feature Enhancement (MSFE) module adaptively fuses features across different temporal resolutions, enabling robust representation learning from complex EIT signals. Experiments show that TFMSCNet achieves 99.50% accuracy on a private EIT dataset, outperforming eleven baseline methods. Validation on ten public datasets confirms its strong generalization, highlighting its potential for clinical pulmonary disease diagnosis. Dashuang Zhu, Pengjiang Qian, Chuang Wang 0011, Xin Ning 0001, Jiafeng Yao |
BIBM | 5 |
| 2025 | Point Clouds Meets Physics: Dynamic Acoustic Field Fitting Network for Point Cloud UnderstandingabstractWhile existing pre-training-based methods have enhanced point cloud model performance, they have not fundamentally resolved the challenge of local structure representation in point clouds. The limited representational capacity of pure point cloud models continues to constrain the potential of cross-modal fusion methods and performance across various tasks. To address this challenge, we propose a Dynamic Acoustic Field Fitting Network (DAF-Net), inspired by physical acoustic principles. Specifically, we represent local point clouds as acoustic fields and introduce a novel Acoustic Field Convolution (AF-Conv), which treats local aggregation as an acoustic energy field modeling problem and captures fine-grained local shape awareness by dividing the local area into near field and far field. Furthermore, drawing inspiration from multi-frequency wave phenomena and dynamic convolution, we develop the Dynamic Acoustic Field Convolution (DAF-Conv) based on AF-Conv. DAF-Conv dynamically generates multiple weights based on local geometric priors, effectively enhancing adaptability to diverse geometric features. Additionally, we design a Global Shape-Aware (GSA) layer incorporating EdgeConv and multi-head attention mechanisms, which combines with DAF-Conv to form the DAF Block. These blocks are then stacked to create a hierarchical DAFNet architecture. Extensive experiments demonstrate that DAFNet significantly outperforms existing methods across multiple tasks. Changshuo Wang 0001, Shuting He, Jiawei Han 0008, Zhonghang Liu, Xin Ning 0001, Weijun Li 0002, Prayag Tiwari |
CVPR | 6 |
| 2025 | Unleashing Foundation Vision Models: Adaptive Transfer for Diverse Data-Limited Scientific DomainsabstractIn the big data era, the computer vision field benefits from large-scale datasets such as LAION-2B, LAION-400M, and ImageNet-21K, Kinetics, on which popular models like the ViT and ConvNeXt series have been pre-trained, acquiring substantial knowledge.
However, numerous downstream tasks in specialized and data-limited scientific domains continue to pose significant challenges.
In this paper, we propose a novel Cluster Attention Adapter (CLAdapter), which refines and adapts the rich representations learned from large-scale data to various data-limited downstream tasks. Specifically, CLAdapter introduces attention mechanisms and cluster centers to personalize the enhancement of transformed features through distribution correlation and transformation matrices. This enables models fine-tuned with CLAdapter to learn distinct representations tailored to different feature sets, facilitating the models' adaptation from rich pre-trained features to various downstream scenarios effectively. In addition, CLAdapter's unified interface design allows for seamless integration with multiple model architectures, including CNNs and Transformers, in both 2D and 3D contexts.
Through extensive experiments on 10 datasets spanning domains such as generic, multimedia, biological, medical, industrial, agricultural, environmental, geographical, materials science, out-of-distribution (OOD), and 3D analysis, CLAdapter achieves state-of-the-art performance across diverse data-limited scientific domains, demonstrating its effectiveness in unleashing the potential of foundation vision models via adaptive transfer.
Code is available at https://github.com/qklee-lz/CLAdapter. Qiankun Li 0004, Feng He 0008, Huabao Chen, Xin Ning 0001, Kun Wang 0056, Zengfu Wang |
NeurIPS | 4 |
| 2025 | Real-time workflow scheduling in hybrid clouds with privacy and security constraints: A deep reinforcement learning approach
Haoyang He, Yang Hu 0009, Fang Fang 0007, Xin Ning 0001, Long Cheng 0003 |
Expert Syst. Appl. | 5 |
| 2025 | Uniformity and deformation: A benchmark for multi-fish real-time tracking in the farming
Jinze Huang, Xiaohan Yu 0001, Dong An 0001, Xin Ning 0001, Jincun Liu, Prayag Tiwari |
Expert Syst. Appl. | 4 |
| 2025 | FefDM-Transformer: Dual-channel multi-stage Transformer-based encoding and fusion mode for infrared-visible images
Junwu Li, Yaomin Wang, Xin Ning 0001, Wenguang He, Weiwei Cai 0001 |
Expert Syst. Appl. | 3 |
| 2025 | Guest Editorial: Multi-view representation learning for computer visionabstractObject recognition and scene analysis in single-view images may face difficulties such as occlusion and incomplete information, while multi-view learning can address this limitation. When an object or scene is observed from multiple views, information on target objects can be significantly enriched to improve the performance of computer vision tasks. For this reason, multi-view has become one of the important forms of data representation, which leads to the emerging of new research topics on complete or in-complete multi-view learning. Multi-view learning enables the use of multi-source information, nevertheless, the heterogeneous characteristics of data make it difficult to reliably associate information from different views, especially in a complex environment. It remains challenging for tasks to make effective use of the consistent and complementary information between different complete views and to enhance the completeness of potential representation. A wide variety of research is being conducted to explore and discover possible challenges and opportunities to exploit multi-view representation learning for computer vision. The purpose of this Special Issue is to collect high-quality articles on the recent development and trend of multi-view representation learning in computer vision, publish new ideas, theories, solutions and insights on this topic, and showcase their applications. In this Special Issue, we have received 36 papers, all of which underwent peer review. Of the 36 originally submitted papers, 10 have been accepted, which cover a variety of fields, such as person re-identification, gait recognition, 3D object recognition, and behaviour recognition. These accepted papers are mainly divided into three categories. The first category covers the incomplete multi-view data learning theoretics and methods. The papers in this category are of He et al., Kun et al., Fan et al. and Wang et al. The last two categories are both multi-view applications. One of which is 3D-related applications. The papers in this category are of Qi et al. and Sun et al. The other category is about 2D recognition. The papers in this category are of Zhang et al., Huang et al., Zheng et al. and Zhang et al. A brief presentation of each of the paper follows. He et al. present an innovative multi-view subspace clustering method with incomplete graph information. Specifically, they separate one shared and multiple specific graphs from multiple raw graph data, and exploit the mask fusion strategy and block diagonal regulariser to obtain the inherent category information. The clustering results on six real-world datasets show that the method outperforms a series of classic incomplete multi-view clustering methods. Kun et al. propose a new method for low-rank-based multi-view subspace clustering based on low-rank correlation analysis. To overcome the limitations of unreliable low-rank structure and imprecise graphs caused by multi-view noise and outliers, they introduce the canonical correlation analysis strategy and a dual regularisation term to characterise the connections between different views adaptively. Experimental results reveal the method's superiority over compared state-of-the-art (SOTA) methods in accuracy, normalised mutual information, and F-score evaluation metrics. Fan et al. address the challenge of partial mapping between the views in multi-view clustering, and propose a self-inferring incomplete multi-view clustering algorithm to explore the information hidden in the local geometric structure and recover missing instances through mining the information hidden in existing instances. Experimental results show that the method can improve the clustering performance compared with the SOTA methods. Wang et al. propose a semi-paired semi-supervised deep hashing to solve the large-scale multimedia retrieval task. The method is an end-to-end deep neural network model with high-order affinity. To maintain the consistency within the modalities, they introduce a common representation that combines with the labelled information to associate different modalities. Experimental results demonstrate the superior performance of proposed method. Qi et al. propose a double-weighting convolution neural network based on the L2-S grouping mechanism for multi-view 3D object recognition. The goal of the proposed L2-S grouping mechanism is to calculate the discrimination score of views and group views more reasonably. Results of the experiments show that the method can achieve SOTA performance. Sun et al. present a dual-matching method with cross-attention mechanism to address the limitations of matching-based methods caused by a preset fixed disparity range on depth estimation task. To tackle the mismatches on edges and details, they introduce an exquisite module based on left-right consistency. The method is proved to be competitive and effective by experiments conducted under popular benchmarks. Zhang et al. want to answer the following two questions: (1) does a query image with higher resolution than that of the gallery image also affect the pedestrian re-identification performance? If so, and (2) how does it affect performance? So, they propose an end-to-end trainable resolution independent person re-identification network that is composed of a cross-resolution Generative Adversarial Networks and embedding batch normalisation layers. The results demonstrate that the proposed method outperforms the SOTA methods in the pedestrian re-identification task on their expanded benchmark dataset. Huang et al. address the limitation of current gait-based age and gender recognition methods under multi-view scene, and propose an attention-aware spatio–temporal learning framework that employs silhouette sequence as an input to learn essential spatial–temporal gait representation. The proposed method has produced results that outperformed the benchmarks with an Mean Absolute Error of 6.68 years for age estimation and a Correct Classification Rate of 97% for gender classification. Zheng et al. apply deep learning to multi-view classroom behaviour detection. First, they propose an improved detection model based on YOLOv5 to improve the convergence speed of the prediction box. Second, they establish a quantitative evaluation standard for students' classroom attention, and then conduct training and verification by collecting multi-view classroom datasets. Finally, they increase the environment variation in the training model phase to make the model have better generalisation ability. Experiments demonstrate that the method can effectively identify and detect students' behaviours in the classroom from different views. Zhang et al. propose a method for multi-dimensional video anomaly detection, which uses the Object-meta instead of video frames as the input, and the Memory Search Guided Autoencoder with Memory Pools (MSGAE-MP) to reconstruct. The multi-dimensional information carried by the input can be strengthened via Object-meta. The MSGAE-MP construct multi-level memory pools, so as to reconstruct Object-meta in different dimensions. Experiments show that the method is feasible and has achieved excellent results. All of the papers published in this Special Issue show that multi-view representation learning theoretics have developed very fast in recent years. In addition, it is very promising to solve traditional computer vision tasks under multi-view setting, including but not limited to 3D object recognition, person re-identification, gait-based age and gender estimation, and depth estimation. Xin Ning and Chen Wang are responsible for the writing of Proposal and Editorial materials; Jun Zhou is responsible for the processing of articles; and Jing Wu, Lin Gu and Jian Cheng are responsible for the solicitation and publicity of the special issue. Firstly, we would like to thank all the authors for their innovative contributions and all the reviewers for their professional and crucial, yet constructive comments. Also, we wish to express our thanks to Mr Hang Ran, PhD students at Institute of Semiconductors, Chinese Academy of Sciences, for his assistance in this process. Last, we wish to express our gratitude to the editorial team of IET Computer Vision for their support throughout this venture. We hope you enjoy this collection of papers and that the Special Issue can stimulate further research and development in this area. This work is supported by the National Natural Science Foundation of China (Grant no. 61901436). National Natural Science Foundation of China, Grant/Award Number: 61901436. Data sharing is not applicable to this article as no new data were created or analyzed in this study. Xin Ning (SMIEEE) received a B.S. degree in software engineering in 2012, and a Ph.D. degree in electronic circuit and system from the university of Chinese Academy of Sciences, in 2017. He is currently an associate professor with the Laboratory of Artificial Neural Networks and High Speed Circuits, Institute of Semiconductors, Chinese Academy of Sciences. His current research interests include neural networks, intelligent systems and computer vision. He has published as the first or corresponding author in more than 45 papers in journals and refereed conferences. Now he serves as the young associated editor of CAAI Transactions on Intelligent Systems, the guest editor of Elsevier Journal on DISPLAYS. He is also the guest editor of CONNECTION SCIENCE and CONCURR COMP-PRACT E. He was the Website Chair of the IEEE HPBD&IS 2020 and the Publication Chair of the IEEE HPBD&IS 2021. Jun Zhou received a B.S. degree in computer science and a B.E. degree in international business from the Nanjing University of Science and Technology, Nanjing, China, in 1996 and 1998, respectively, an M.S. degree in computer science from Concordia University, Montreal, QC, Canada, in 2002, and a Ph.D. degree in computing science from the University of Alberta, Edmonton, AB, Canada, in 2006. He was a research fellow with the Research School of Computer Science, The Australian National University, Canberra, ACT, Australia, and a researcher with the Canberra Research Laboratory, National Information and Communications Technology Australia, Canberra. In 2012, he joined the School of Information and Communication Technology, Griffith University, Nathan, QLD, Australia, where he is currently a reader. His research interests include pattern recognition, computer vision, and spectral imaging and their applications in remote sensing and environmental informatics. He is the associate editor for the journal of Pattern Recognition and IEEE Trans. on Remote Sensing. Jian Cheng is a professor of Institute of Automation, Chinese Academy of Sciences. He received the B.S. and M.S. degrees in Mathematics from Wuhan University in 1998 and 2001, respectively. After that, he received a Ph.D degree in pattern recognition and intelligent systems from Institute of Automation, Chinese Academy of Sciences in 2004. His current major research interests include deep learning, computer vision, chip design, etc. Jing Wu is now a postdoc at the school of computer science, Beihang University. He received his B.E. degree from the school of computer science, Northwestern Polytechnical University in 2013 and received his PhD. degree from the school of computer science, Beihang University in 2021. His research interests include computer vision, stereo matching, 3D reconstruction and camera localization. Chen Wang is now a postdoc at the school of computer science, Beihang University. He received his B.E. degree from the school of computer science, Northwestern Polytechnical University in 2013 and received his PhD. degree from the school of computer science, Beihang University in 2021. His research interests include computer vision, stereo matching, 3D reconstruction and camera localization. Lin Gu received a B.Eng. degree from Shanghai University, Shanghai, China, in 2009, and a Ph.D. degree in computer vision from Australian National University in 2014. After Ph.D. graduation from the Australian National University, he worked as a post-doctoral researcher at A*STAR, Singapore. Then, he was a project researcher with the National Institute of Informatics, Japan, and also a visiting scholar with Kyoto University, Japan. He is currently a research scientist at RIKEN AIP, Japan, and a special researcher with the University of Tokyo, Japan. He is also an in-charge of a Moonshot and an ACT-X Project to improve artificial intelligence by simulating the human brain. His primary research interests lie in machine learning, medical imaging, and computational photography. Xin Ning 0001, Jun Zhou 0001, Jian Cheng 0001, Jing Wu 0004, Chen Wang 0026, Lin Gu 0003 |
IET Comput. Vis. | 1 |
| 2025 | Fully Decoupled End-to-End Person Search: An Approach without Conflicting Objectives
Pengcheng Zhang 0003, Xiaohan Yu 0001, Xiao Bai 0001, Xin Ning 0001, Edwin R. Hancock |
Int. J. Comput. Vis. | 5 |
| 2025 | FIPNet: Self-supervised low-light image enhancement combining feature and illumination priors
Qijie Zou, Xin Ning 0001, Gang Wang 0023 |
Neurocomputing | 3 |
| 2025 | Secure and Efficient Authentication Protocol for Supply Chain Systems in Artificial-Intelligence-Based Internet of ThingsabstractWith the rapid development of globalization and technology, the industrial structure has undergone profound changes and market competition is becoming increasingly fierce. In order to enhance their competitive advantage, enterprises are integrating resources through supply chain systems (SCS) and establishing strategic partnerships to achieve rapid market response and optimized resource allocation. Artificial Intelligence (AI) -based Internet of Things (IoT) has brought unprecedented changes to supply chain systems. AIoT not only improves the transparency and efficiency of the supply chain, but also significantly reduces carbon emissions, making a positive contribution to green and sustainable development. However, with the widespread application of AIoT in SCS, its security issues have become increasingly prominent, posing a serious threat to the stable operation of the supply chain. Therefore, this article proposes a secure and efficient authentication protocol for supply chain systems in AIoT. This protocol implements user authentication, ensures forward security of session key and guarantees secure data transmission. Through security analysis, this article proves that the protocol can resist various known attacks. In addition, through performance analysis, the protocol has low overhead and can meet the efficiency requirements of SCS. Junfeng Miao, Xin Ning 0001, Shuangxi Hong, Lanlan Wang |
IEEE Internet Things J. | 2 |
| 2025 | Simultaneous outlier detection and elimination in hyperspectral unmixing via weighted non-negative matrix tri-factorizationabstractAbstract Hyperspectral unmixing (HU) involves separating mixed pixel spectra into pure endmember spectra and their corresponding abundance fractions. However, it faces significant challenges due to outliers in the hyperspectral data, which often appear as pixel and band anomalies. Outliers in pixels could result in incorrect classification and inaccurate quantification of materials, while outliers in bands could alter spectral characteristics, leading to misidentifying endmembers and incorrect estimates of abundance. To tackle these issues, this paper introduces a new approach, named simultaneous outlier detection and elimination via weighted non-negative matrix tri-factorization (SODE-WNMTF), which offers an efficient means of addressing the impact of outliers in the unmixing process. Leveraging the co-clustering property of NMTF, SODE-WNMTF introduces a novel weighting matrix, which involves simultaneous clustering of both pixels and spectral bands to effectively detect and mitigate the negative impact of both pixel and band outliers during the unmixing process. At the same time, the inherent structure of the hyperspectral image (HSI) is utilized through the examination of local and global connections among pixels and spectral bands, consequently improving the co-clustering procedure. In addition, SODE-WNMTF proposes a spatial weighting factor, which utilizes the similarity of adjacent pixels, to promote piecewise smoothness in abundance maps while mitigating the impact of outliers. Moreover, since pixels in regions dominated by a single endmember exhibit spectra closely resembling that endmember, SODE-WNMTF incorporates a sparse estimation technique for endmember signatures. Finally, to verify the performance of SODE-WNMTF, a series of experiments is conducted on both synthetic and real HSIs, with outcomes proving its superiority against other cutting-edge approaches. The source code is also available at https://github.com/yasinhashemi/SODE-WNMTF . Yasin Hashemi-Nazari, Farid Saberi Movahed, Azita Tajaddini, Catarina Moreira, Xin Ning 0001, Prayag Tiwari |
Mach. Learn. | 5 |
| 2025 | A prompt regularization approach to enhance few-shot class-incremental learning with Two-Stage Classifier
Meilan Hao, Yizhan Gu, Kejian Dong, Prayag Tiwari, Xiaoqing Lv, Xin Ning 0001 |
Neural Networks | 6 |
| 2025 | Cross-modal knowledge transfer for 3D point clouds via graph offset prediction
Long Yu 0001, Guoqi Wang, Shengwei Tian, Zaiyang Yu, Weijun Li 0002, Xin Ning 0001 |
Pattern Recognit. | 7 |
| 2025 | Dual Guidance Enabled Fuzzy Inference for Enhanced Fine-Grained RecognitionabstractIn the field of fine-grained visual recognition (FGVR), the ability to resolve minute and often subtle differences between highly similar object categories is paramount. The advent of vision transformers (ViTs) has marked a significant advancement in this domain, primarily due to their capacity to model the intricate interdependencies among object parts represented as image patches. However, their inherent single-scale processing limitation hampers their effectiveness in FGVR tasks. Furthermore, the challenge of uncertainty inherent in FGVR tasks remains unresolved, necessitating the development of methods that bolster the robustness of these models, particularly across varying scales of visual features. We introduce a new plug-in module that can be seamlessly integrated into ViT, called dual guidance enabled fuzzy inference (DGEFI), which combines fuzzy inference with dual guidance mechanisms. Dual guidance includes scale-aware guidance and probability guidance. The former strengthens the model's focus on salient scales, and the latter refines the distinction between similar categories by optimizing intraclass compactness and interclass separability. Fuzzy inference enables the model to adaptively tweak the influence of distinct scales in the final decision-making phase, thereby enhancing the overall accuracy of recognition tasks. We demonstrate the versatility and efficacy of our DGEFI module by integrating it into several leading ViT backbones, including ViT, Swin, Mvitv2, and EVA-02. Empirical results exhibit exceptional performance gains, with the integration of DGEFI into EVA-02 remarkable accuracy improvements, reaching 93.6% on the CUB-200-2011 dataset and 94.5% on the NA-Birds dataset, respectively, improving over the state-of-the-art method 0.5% and 1.5%. Qiupu Chen, Feng He 0008, Gang Wang 0023, Xiao Bai 0001, Long Cheng 0003, Xin Ning 0001 |
IEEE Trans. Fuzzy Syst. | 6 |
| 2025 | Dynamic Personalized Federated Learning for Cross-Spectral Palmprint RecognitionabstractPalmprint recognition has recently garnered attention due to its high accuracy, strong robustness, and high security. Existing deep learning-based palmprint recognition methods usually require large amounts of data for centralized training, facing the challenge of privacy disclosure. In addition, the non-independent and identically distributed (non-IID) issue in the multi-spectral palmprint images generally leads to the degradation of recognition performance. To tackle these problems, this paper proposes a dynamic personalized federated learning model for cross-spectral palmprint recognition, called DPFed-Palm. Specifically, for each client's local training, we present a new combination of loss functions to enforce the constraints of local models and effectively enhance the feature representation capability of models. Subsequently, DPFed-Palm aggregates the above-trained local models by using the combined aggregation strategies of the Federated Averaging (FedAvg) and Personalized Federated Learning (PFL) to obtain the best personalized global model of each client. For the selection of the best personalized global model, we develop a dynamic weight selection strategy to obtain the optimal weights of the local and global models by cross-spectral (cross-client) testing. Extensive experimental results on three public PolyU multispectral, IITD, and CASIA datasets show that the proposed method outperforms the existing techniques in privacy-preserving and recognition performance. Shuyi Li 0003, Jianian Hu, Bob Zhang 0001, Xin Ning 0001, Lifang Wu |
IEEE Trans. Image Process. | 4 |
| 2025 | Practical and Secure Authentication Protocol for Vehicle to Grid in Intelligent Transportation SystemsabstractWith the gradual increase in market share of electric vehicle (EV), Vehicle to Grid(V2G) has become a new research hotspot in the field of intelligent transportation systems. Its goal is to avoid overloading the power grid due to the simultaneous charging of a large number of electric vehicles. However, when EV is connected to the power grid, V2G will involve a large amount of privacy data exchange. Once these data are leaked, the privacy and security of users will be threatened. Ensuring the secure transmission of user privacy information in V2G is crucial. Therefore, this paper proposes a practical and secure authentication protocol for V2G in intelligent transportation systems. This protocol ensures user login security through three-factor authentication mechanism and then implements authentication based on Chebyshev chaotic maps. Finally, secure communication is carried out through the established key. Security analysis shows that this protocol is secure and can ensure the privacy and security of V2G. Informal security analysis shows that this protocol can meet various security attributes. Functional comparison and performance analysis indicate that the protocol not only has high security but also has low computation and communication overhead. Junfeng Miao, Zhaoshun Wang, Xin Ning 0001, Achyut Shankar, Carsten Maple, Joel J. P. C. Rodrigues |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2025 | Unsupervised Recognition of Unknown Objects for Open-World Object DetectionabstractOpen-world object detection (OWOD) extends object detection problem to a realistic and dynamic scenario, where a detection model is required to be capable of detecting both known and unknown objects and incrementally learning newly introduced knowledge. Current OWOD models detect the unknowns that exhibit similar features to the known objects, but they suffer from a severe label bias problem, i.e., they tend to detect all regions (including unknown object regions) that are dissimilar to the known objects as part of the background. To eliminate the label bias, this article proposes a novel module, namely reconstruction error-based Weibull (REW) model, that learns an unsupervised discriminative model for recognizing true unknown objects based on prior knowledge of object occurrence frequency via Weibull modeling. The resulting model can be further refined by another module of our method, called REW-enhanced object localization network (ROLNet), which iteratively extends pseudo-unknown objects to the unlabeled regions. Experimental results show that our method 1) significantly outperforms the prior SOTA in detecting unknown objects while maintaining competitive performance of detecting known object classes on the MS COCO dataset and 2) achieves better generalization ability on the LVIS and Objects365 datasets. Code is available at https://github.com/frh23333/mepu-owod. Ruohuan Fang, Guansong Pang, Wenjun Miao, Xiao Bai 0001, Xin Ning 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2025 | Brain-Inspired Fast- and Slow-Update Prompt Tuning for Few-Shot Class-Incremental LearningabstractFew-shot class-incremental learning (FSCIL) aims to learn new classes incrementally with a limited number of samples per class. Foundation models combined with prompt tuning showcase robust generalization and zero-shot learning (ZSL) capabilities, endowing them with potential advantages in transfer capabilities for FSCIL. However, existing prompt tuning methods excel in optimizing for stationary datasets, diverging from the inherent sequential nature in the FSCIL paradigm. To address this issue, taking inspiration from the "fast and slow mechanism" of the complementary learning systems (CLSs) in the brain, we present fast- and slow-update prompt tuning FSCIL (FSPT-FSCIL), a brain-inspired prompt tuning method for transferring foundation models to the FSCIL task. We categorize the prompts into two groups: fast-update prompts and slow-update prompts, which are interactively trained through meta-learning. Fast-update prompts aim to learn new knowledge within a limited number of iterations, while slow-update prompts serve as meta-knowledge and aim to strike a balance between rapid learning and avoiding catastrophic forgetting. Through experiments on multiple benchmark tests, we demonstrate the effectiveness and superiority of FSPT-FSCIL. The code is available at https://github.com/qihangran/FSPT-FSCIL. Hang Ran, Xingyu Gao 0001, Lusi Li, Weijun Li 0002, Songsong Tian, Gang Wang 0023, Hailong Shi, Xin Ning 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2024 | DNGaussian: Optimizing Sparse-View 3D Gaussian Radiance Fields with Global-Local Depth NormalizationabstractRadiance fields have demonstrated impressive performance in synthesizing novel views from sparse input views, yet prevailing methods suffer from high training costs and slow inference speed. This paper introduces DNGaussian, a depth-regularized framework based on 3D Gaussian radiance fields, offering real-time and high-quality few-shot novel view synthesis at low costs. Our motivation stems from the highly efficient representation and surprising quality of the recent 3D Gaussian Splatting, despite it will encounter a geometry degradation when input views decrease. In the Gaussian radiance fields, we find this degradation in scene geometry primarily lined to the positioning of Gaussian primitives and can be mitigated by depth constraint. Consequently, we propose a Hard and Soft Depth Regularization to restore accurate scene geometry under coarse monocular depth supervision while maintaining a fine-grained color appearance. To further refine detailed geometry reshaping, we introduce Global-Local Depth Normalization, enhancing the focus on small local depth changes. Extensive experiments on LLFF, DTU, and Blender datasets demonstrate that DNGaussian outperforms state-of-the-art methods, achieving comparable or better results with significantly reduced memory cost, a 25 × reduction in training time, and over 3000 × faster rendering speed. Code is available at: https://github.com/Fictionarry/DNGaussian. Jiahe Li 0007, Xiao Bai 0001, Xin Ning 0001, Jun Zhou 0001, Lin Gu 0003 |
CVPR | 5 |
| 2024 | TalkingGaussian: Structure-Persistent 3D Talking Head Synthesis via Gaussian Splatting
Jiahe Li 0007, Xiao Bai 0001, Xin Ning 0001, Jun Zhou 0001, Lin Gu 0003 |
ECCV (10) | 5 |
| 2024 | GPSFormer: A Global Perception and Local Structure Fitting-Based Transformer for Point Cloud Understanding
Changshuo Wang 0001, Meiqing Wu, Siew-Kei Lam, Xin Ning 0001, Shangshu Yu, Ruiping Wang 0005, Weijun Li 0002, Thambipillai Srikanthan |
ECCV (8) | 4 |
| 2024 | Prompting Continual Person SearchabstractThe development of person search techniques has been greatly promoted in recent years for its superior practicality and challenging goals. Despite their significant progress, existing person search models still lack the ability to continually learn from increasing real-world data and adaptively process input from different domains. To this end, this work introduces the continual person search task that sequentially learns on multiple domains and then performs person search on all seen domains. This requires balancing the stability and plasticity of the model to continually learn new knowledge without catastrophic forgetting. For this, we propose a Prompt-based Continual Person Search (PoPS) model in this paper. First, we design a compositional person search transformer to construct an effective pre-trained transformer without exhaustive pre-training from scratch on large-scale person search data. This serves as the fundamental for prompt-based continual learning. On top of that, we design a domain incremental prompt pool with a diverse attribute matching module. For each domain, we independently learn a set of prompts to encode the domain-oriented knowledge. Meanwhile, we jointly learn a group of diverse attribute projections and prototype embeddings to capture discriminative domain attributes. By matching an input image with the learned attributes across domains, the learned prompts can be properly selected for model inference. Extensive experiments are conducted to validate the proposed method for continual person search. The source code is available at https://github.com/PatrickZad/PoPS. Pengcheng Zhang 0003, Xiaohan Yu 0001, Xiao Bai 0001, Xin Ning 0001 |
ACM Multimedia | 5 |
| 2024 | A blockchain-enabled privacy-preserving authentication management protocol for Internet of Medical Things
Junfeng Miao, Zhaoshun Wang, Zeqing Wu, Xin Ning 0001, Prayag Tiwari |
Expert Syst. Appl. | 4 |
| 2024 | Occluded person re-identification with deep learning: A survey and perspectives
Enhao Ning, Changshuo Wang 0001, Xin Ning 0001, Prayag Tiwari |
Expert Syst. Appl. | 4 |
| 2024 | MARP: A Cooperative Multiagent DRL System for Connected Autonomous Vehicle PlatooningabstractIn modern urban areas, inefficiency traffic management is one of the main causes of road congestion, leading to reduced fuel efficiency and increased traffic safety hazards. Traditional researches typically focus only on enhancing the throughput of intersections by optimizing traffic signals or individual vehicle trajectories. However, these methods often overlook the dynamic nature of the traffic system and the potential benefits of vehicle platooning, limiting their effectiveness in complex traffic environments. Addressing this challenge, this article presents MARP, a Cooperative Multiagent deep reinforcement learning (DRL) System for connected autonomous vehicle (CAV) Platooning. Utilizing vehicle to vehicle (V2I) and vehicles to infrastructure (V2V) technologies, MARP integrates sensing, computing, and communication to collect and process real-time data on traffic conditions, thereby achieving dynamic synchronization between traffic signal controllers and CAV platoons. By constructing platoons that collaborates with the infrastructure through a multiagent DRL collaboration model, MARP adapts to real-time traffic flow changes, significantly optimizing the fluidity and efficiency of the entire traffic network. Detailed experiments show that MARP effectively reduces traffic congestion, shortens intersection travel times, and cuts fuel consumption and emissions, surpassing the state-of-the-art approach. Shuhong Dai, Shike Li, Haichuan Tang, Xin Ning 0001, Fang Fang 0007, Yunxiao Fu, Qingle Wang, Long Cheng 0003 |
IEEE Internet Things J. | 4 |
| 2024 | Learning optimal inter-class margin adaptively for few-shot class-incremental learning via neural collapse-based meta-learningabstractFew-Shot Class-Incremental Learning (FSCIL) aims to learn new classes incrementally with a limited number of samples per class. It faces issues of forgetting previously learned classes and overfitting on few-shot classes. An efficient strategy is to learn features that are discriminative in both base and incremental sessions. Current methods improve discriminability by manually designing inter-class margins based on empirical observations, which can be suboptimal. The emerging Neural Collapse (NC) theory provides a theoretically optimal inter-class margin for classification, serving as a basis for adaptively computing the margin. Yet, it is designed for closed, balanced data, not for sequential or few-shot imbalanced data. To address this gap, we propose a Meta-learning- and NC-based FSCIL method, MetaNC-FSCIL, to compute the optimal margin adaptively and maintain it at each incremental session. Specifically, we first compute the theoretically optimal margin based on the NC theory. Then we introduce a novel loss function to ensure that the loss value is minimized precisely when the inter-class margin reaches its theoretically best. Motivated by the intuition that “learn how to preserve the margin” matches the meta-learning’s goal of “learn how to learn”, we embed the loss function in base-session meta-training to preserve the margin for future meta-testing sessions. Experimental results demonstrate the effectiveness of MetaNC-FSCIL, achieving superior performance on multiple datasets. The code is available at https://github.com/qihangran/metaNC-FSCIL. Hang Ran, Weijun Li 0002, Lusi Li, Songsong Tian, Xin Ning 0001, Prayag Tiwari |
Inf. Process. Manag. | 5 |
| 2024 | ICGNet: An intensity-controllable generation network based on covering learning for face attribute synthesis
Xin Ning 0001, Feng He 0008, Xiaoli Dong, Weijun Li 0002, Fayadh Alenezi, Prayag Tiwari |
Inf. Sci. | 1 |
| 2024 | MV-ReID: 3D Multi-view Transformation Network for Occluded Person Re-Identification
Zaiyang Yu, Prayag Tiwari, Luyang Hou, Lusi Li, Weijun Li 0002, Limin Jiang, Xin Ning 0001 |
Knowl. Based Syst. | 7 |
| 2024 | Adaptively identify and refine ill-posed regions for accurate stereo matching
Changlin Liu, Linjun Sun, Xin Ning 0001, Weijun Li 0002 |
Neural Networks | 3 |
| 2024 | Enhancement, integration, expansion: Activating representation of detailed features for occluded person re-identification
Enhao Ning, Changshuo Wang 0001, Xin Ning 0001 |
Neural Networks | 5 |
| 2024 | A survey on few-shot class-incremental learningabstractLarge deep learning models are impressive, but they struggle when real-time data is not available. Few-shot class-incremental learning (FSCIL) poses a significant challenge for deep neural networks to learn new tasks from just a few labeled samples without forgetting the previously learned ones. This setup can easily leads to catastrophic forgetting and overfitting problems, severely affecting model performance. Studying FSCIL helps overcome deep learning model limitations on data volume and acquisition time, while improving practicality and adaptability of machine learning models. This paper provides a comprehensive survey on FSCIL. Unlike previous surveys, we aim to synthesize few-shot learning and incremental learning, focusing on introducing FSCIL from two perspectives, while reviewing over 30 theoretical research studies and more than 20 applied research studies. From the theoretical perspective, we provide a novel categorization approach that divides the field into five subcategories, including traditional machine learning methods, meta learning-based methods, feature and feature space-based methods, replay-based methods, and dynamic network structure-based methods. We also evaluate the performance of recent theoretical research on benchmark datasets of FSCIL. From the application perspective, FSCIL has achieved impressive achievements in various fields of computer vision such as image classification, object detection, and image segmentation, as well as in natural language processing and graph. We summarize the important applications. Finally, we point out potential future research directions, including applications, problem setups, and theory development. Overall, this paper offers a comprehensive analysis of the latest advances in FSCIL from a methodological, performance, and application perspective. Songsong Tian, Lusi Li, Weijun Li 0002, Hang Ran, Xin Ning 0001, Prayag Tiwari |
Neural Networks | 5 |
| 2024 | Deformation depth decoupling network for point cloud domain adaptation
Xin Ning 0001, Changshuo Wang 0001, Enhao Ning, Lusi Li |
Neural Networks | 2 |
| 2024 | Learning adversarial semantic embeddings for zero-shot recognition in open worlds
Guansong Pang, Xiao Bai 0001, Lei Zhou 0008, Xin Ning 0001 |
Pattern Recognit. | 6 |
| 2024 | Joint discriminative representation learning for end-to-end person search
Pengcheng Zhang 0003, Xiaohan Yu 0001, Xiao Bai 0001, Chen Wang 0026, Xin Ning 0001 |
Pattern Recognit. | 6 |
| 2024 | Towards effective person search with deep learning: A survey from systematic perspective
Pengcheng Zhang 0003, Xiaohan Yu 0001, Chen Wang 0026, Xin Ning 0001, Xiao Bai 0001 |
Pattern Recognit. | 5 |
| 2024 | Practical and secure multifactor authentication protocol for autonomous vehicles in 5GabstractAbstract Autonomous vehicles (AV) can not only improve traffic safety and congestion, but also have strategic significance for the development of the transportation industry. With the continuous updating of core technologies such as artificial intelligence, sensor detection, synchronous positioning, and high‐precision mapping, the development of AV has been promoted. When 5G network is combined with Internet of Vehicles, the problems of AV can be solved by taking advantage of 5G ultra‐large bandwidth, low latency and high reliability. However, when the user controls the vehicle remotely, a real‐time and reliable authentication process is needed, while minimizing the overhead of security protocols. Therefore, this article proposes a practical and secure multifactor user authentication protocol for AV in 5G network. By introducing non‐interactive zero‐knowledge proof technology and physical uncloning function, the protocol completes mutual authentication and key agreement without revealing any sensitive information. The article proves the security of the protocol through BAN logic and the simulation of Scyther. And it can resist malicious attacks and provide more security features. The informal security analysis shows that the protocol can meet the proposed security requirements. Finally, we evaluate the efficiency of the protocol, and the results show that the protocol can provide better performance. Junfeng Miao, Zhaoshun Wang, Xin Ning 0001, Weiwei Cai 0001, Ruimin Liu |
Softw. Pract. Exp. | 3 |
| 2024 | A Recognizable Expression Line Portrait Synthesis Method in Portrait Rendering RobotabstractAn artistic line portrait robot can generate, process, and draw line portraits. Compared to real face images, line portraits lose some recognizable information. Maintaining recognizability during the process of expression edition of line portraits is an important challenge for artistic portrait robots. A recognizable expression line portrait synthesis method based on a triangle coordinate system (TCS) is proposed. First, based on public facial expression databases [JAFFE, Oulu CASIA, RaFD, and Cohn-Kanade (CK)], by studying the feature deviations between different expressions of the same person, an expression deformation constraint criterion (EDCC) that is conducive to maintaining recognizable features is proposed. Then, by comparing features between the source line portrait and reference expression portrait, the expression features are calculated. Finally, under the EDCC, based on expression features, a recognizable expression line portrait is generated through image topological deformation based on TCS. In addition, we can synthesize different degrees of expression line portraits. On the public face datasets (FHHQ, CelebA-HQ, and CK), we implemented qualitative and quantitative contrast experiments. Experimental results demonstrate that this method can automatically synthesize an expression line portrait with reference expression, where the expression degree of the reference expression is controllable, and the generated expression portrait still has high recognizability. The expression samples generated by the proposed method are used for face authentication on the CK dataset, and only 0.22% of the samples fail to pass the authentication. Xiaoli Dong, Xin Ning 0001, Weijun Li 0002, Liping Zhang 0014 |
IEEE Trans. Comput. Soc. Syst. | 2 |
| 2024 | 3D Person Re-Identification Based on Global Semantic Guidance and Local Feature AggregationabstractPerson re-identification (Re-ID) has played an extremely crucial role in ensuring social safety and has attracted considerable research attention. 3D shape information is an important clue to understand the posture and shape of pedestrians. However, most existing person Re-ID methods learn pedestrian feature representations from images, ignoring the real 3D human body structure and the spatial relationship between the pedestrians and interferents. To address this problem, our devise a new point cloud Re-ID network (PointReIDNet), designed to obtain 3D shape representations of pedestrians from point clouds of 3D scenes. The model consists of modules, namely global semantic guidance module and local feature extraction module. The global semantic guidance module is designed by enhancing the point cloud feature representation in similar feature neighborhoods and to reduce the interference caused by 3D shape reconstruction or noise. Further, to provide an efficient representation of point clouds, we propose space cover convolution (SC-Conv), which efficiently encodes information on human shapes in local point clouds by constructing anisotropic geometries in the coordinate neighborhoods. Extensive experiments are conducted on four holistic person Re-ID datasets, one occlusion person Re-ID dataset and one point cloud classification dataset. The results exhibit significant improvements over point-cloud-based person Re-ID methods. In particular, the proposed efficient PointReIDNet decreases the number of parameters from 2.30M to 0.35M with an insignificant drop in performance. The source code is available at: https://github.com/changshuowang/PointReIDNet. Changshuo Wang 0001, Xin Ning 0001, Weijun Li 0002, Xiao Bai 0001, Xingyu Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Pedestrian 3D Shape Understanding for Person Re-Identification via Multi-View LearningabstractRecent development in computing power has resulted in performance improvements on holistic(none-occluded) person Re-Identification (ReID) tasks. Nevertheless, the precision of the recent research will diminish when a pedestrian is obstructed by obstacles. Within the realm of 2D space, the loss of information from obstructed objects continues to pose significant challenges in the context of person ReID. Person is a 3D non-grid object, and thus semantic representation learning in only 2D space limits the understanding of occluded person. In the present work, we propose a network based on 3D multi-view learning, allowing it to acquire geometric and shape details of an occluded pedestrian from 3D space. Simultaneously, it capitalizes on advancements in 2D-based networks to extract semantic representations from 3D multi-views. Specifically, the surface random selection strategy is proposed to convert images of 2D RGB into 3D multi-views. Using this strategy, we build four extensive 3D multi-view data collections for person ReID. After that, Pedestrian 3D Shape Understanding for Person Re-Identification via Multi-View Learning(MV-3DSReID), is proposed for identifying the person by learning person geometry and structure representation from the groups of multi-view images. In comparison to alternative data formats (e.g., 2D RGB, 3D point cloud), multi-view images complement each other’s detailed features of the 3D object by adjusting rendering viewpoints, thus facilitating a more comprehensive understanding of the person for both holistic and occluded ReID situations. Experiments on occluded and holistic ReID tasks demonstrate performance levels comparable to state-of-the-art methods, validating the effectiveness of our proposed approach in tackling challenges related to occlusion. The code is available at https://github.com/hangjiaqi1/MV-TransReID. Zaiyang Yu, Lusi Li, Jinlong Xie, Changshuo Wang 0001, Weijun Li 0002, Xin Ning 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2024 | Robust and Sparse Least Square Regression for Finger Vein and Finger Knuckle Print RecognitionabstractDue to their high reliability, security, and anti-counterfeiting, finger-based biometrics (such as finger vein and finger knuckle print) have recently received considerable attention. Despite recent advances in finger-based biometrics, most of these approaches leverage much prior information and are non-robust for different modalities or different scenarios. To address this problem, we propose a structured Robust and Sparse Least Square Regression (RSLSR) framework to adaptively learn discriminative features for personal identification. To achieve the powerful representation capacity of the input data, RSLSR synchronously integrates robust projection learning, noise decomposition, and discriminant sparse representation into a unified learning framework. Specifically, RSLSR jointly learns the most discriminative information from the original pixels of the finger images by introducing the$l_{2,1}$norm. A sparse transformation matrix and reconstruction error are simultaneously enforced to enhance its robustness to noise, thus making RSLSR adaptable to multi-scenarios. Extensive experiments on five contact-based and contactless-based finger databases demonstrate the clear superiority of the proposed RSLSR in terms of recognition accuracy and computational efficiency. Shuyi Li 0003, Bob Zhang 0001, Lifang Wu, Ruijun Ma 0001, Xin Ning 0001 |
IEEE Trans. Inf. Forensics Secur. | 5 |
| 2024 | DCNet: A Self-Supervised EEG Classification Framework for Improving Cognitive Computing-Enabled Smart HealthcareabstractCognitive computing endeavors to construct models that emulate brain functions, which can be explored through electroencephalography (EEG). Developing precise and robust EEG classification models is crucial for advancing cognitive computing. Despite the high accuracy of supervised EEG classification models, they are constrained by labor-intensive annotations and poor generalization. Self-supervised models address these issues but encounter difficulties in matching the accuracy of supervised learning. Three challenges persist: 1) capturing temporal dependencies in EEG; 2) adapting loss functions to describe feature similarities in self-supervised models; and 3) addressing the prevalent issue of data imbalance in EEG. This study introduces the DreamCatcher Network (DCNet), a self-supervised EEG classification framework with a two-stage training strategy. The first stage extracts robust representations through contrastive learning, and the second stage transfers the representation encoder to a supervised EEG classification task. DCNet utilizes time-series contrastive learning to autonomously construct representations that comprehensively capture temporal correlations. A novel loss function, SelfDreamCatcherLoss, is proposed to evaluate the similarities between these representations and enhance the performance of DCNet. Additionally, two data augmentation methods are integrated to alleviate class imbalances. Extensive experiments show the superiority of DCNet over the current state-of-the-art models, achieving high accuracy on both the Sleep-EDF and HAR datasets. It holds substantial promise for revolutionizing sleep disorder detection and expediting the development of advanced healthcare systems driven by cognitive computing. Yiyang Zhang 0008, Le Sun 0003, Deepak Gupta 0002, Xin Ning 0001, Prayag Tiwari |
IEEE J. Biomed. Health Informatics | 4 |
| 2024 | A UAV-Assisted Authentication Protocol for Internet of VehiclesabstractAs a component of the Intelligent Transportation System (ITS), Internet of Vehicles (IoV) is becoming increasingly important in the management and construction of urban transportation as it can provide users with a range of applications related to traffic accident warnings, entertainment information, collaborative driving and real-time road information through communication devices on vehicles. However, with the increasing variety of services in the IoV, the growing demand for user traffic and the advances in Unmanned Aerial Vehicle (UAV) technology, UAV is introduced into the IoV as a solution, which can relieve the pressure on the communication infrastructure in the network, provide emergency communication services and improve the performance of network services. Due to the openness of IoV and the high-speed movement of vehicles, authentication and privacy issues are among the most pressing issues in IoV. Therefore, the paper proposes a secure and effective authentication protocol for UAV-assisted IoV. The protocol utilises elliptic curve cryptography to assure the security of the authentication. The protocol undergoes proof of security, Burrows-Abadi-Needham (BAN) logic analysis and informal security analysis to ensure secure and mutual authentication, and have a good resistance to known attacks. Furthermore, performance analysis and comparison are conducted to evaluate the efficiency of our protocol. The results indicate that our protocol has superior advantages in overhead. Junfeng Miao, Zhaoshun Wang, Xin Ning 0001, Achyut Shankar, Carsten Maple, Joel J. P. C. Rodrigues |
IEEE Trans. Intell. Transp. Syst. | 3 |
| 2024 | PointGT: A Method for Point-Cloud Classification and Segmentation Based on Local Geometric TransformationabstractRecently, three-dimensional (3D) point-cloud analysis has been extensively utilized in the domain of machine vision, encompassing tasks include shape classification and segmentation. However the inherent disorder in point clouds poses a challenge in capturing relationships among points, particularly when dealing with mutilated and occluded data. To this end, We propose the Point Geometry Transformation (PointGT) method for 3D point-cloud classification and part segmentation, by exploring the underlying geometric structure in the local and global of points. Specifically, the efficacy of PointGT arises from the integration of a local abstraction (LA) module and an optimization strategy. The LA module is tailored to address the localized features inherent to point clouds. This module encapsulates the multidimensional attributes of local edge and inside points. The bi-directional cross-attention mechanism amalgamates these two constituents into the native channel with the primary objective of optimizing the exploitation of edge and inside delineations, thereby judiciously mitigating noise artifacts. Ultimately, the channel residual connections disseminate the postdownsampling point attributes, thereby inheriting the edge and inside delineations gleaned via post bi-directional attention. The effectiveness of the proposed method was verified through the validation of point-cloud classification and segmentation datasets. The empirical findings confirmed the efficacy of PointGT; accuracies of 93.2% and 87.8% were achieved for the ModelNet40 and ScanObjectNN datasets, respectively. Changshuo Wang 0001, Long Yu 0001, Shengwei Tian, Xin Ning 0001, Joel J. P. C. Rodrigues |
IEEE Trans. Multim. | 5 |
| 2023 | Blind image quality assessment based on the multiscale and dual-domains features fusionabstractAbstract Image quality assessment is to simulate subjective human visual perception and realize image quality inference automatically. Although deep neural networks have achieved great success, the majority of them do not fully consider perception characteristics. Therefore, according to the human visual scale characteristics, we proposed an image quality assessment algorithm based on multiscale and dual domains fusion. Firstly, the original image and its phase congruency respectively input into two branches, feature pyramid and channel attention mechanism are adopted to extract multiscale features. After that, bilinear pool is used to aggregate the spatial and frequency domain characteristics of the corresponding scales, and allows arbitrary scale input to ensure that the features are extracted from the inherent quality images. Finally, the single quality score is obtained through learned weights of each scale. Comparative experiments between our approach and state‐of‐the‐art are conducted on five public databases, the results demonstrate that the proposed algorithm is not only robust to different types and across database, but also sensitive to scale. Yaxuan Lu, Weijun Li 0002, Xin Ning 0001, Xiaoli Dong, Liping Zhang 0014, Linjun Sun, Chuantong Cheng |
Concurr. Comput. Pract. Exp. | 3 |
| 2023 | Multi-view frontal face image generation: A surveyabstractAbstract Face images from different perspectives reduce the accuracy of face recognition, and the generation of frontal face images is an important research topic in the field of face recognition. To understand the development of frontal face generation models and grasp the current research hotspots and trends, existing methods based on 3D models, deep learning, and hybrid models are summarized, and the current commonly used face generation methods are introduced. Dataset, and compare the performance of existing models through experiments. The purpose of this paper is to fundamentally understand the advantages of existing frontal face generation, sort out the key issues of such generation, and look toward future development trends. Xin Ning 0001, Fangzhe Nan, Shaohui Xu, Liping Zhang 0014 |
Concurr. Comput. Pract. Exp. | 1 |
| 2023 | A review of research on co-trainingabstractSummary Co‐training algorithm is one of the main methods of semi‐supervised learning in machine learning, which explores the effective information in unlabeled data by multi‐learner collaboration. Based on the development of co‐training algorithm, the research work in recent years was further summarized in this article. In particular, three main steps of relevant co‐training algorithms are introduced: view acquisition, learners' differentiation, and label confidence estimation. Finally, we summarized the problems existing in the current co‐training methods, gave some suggestions for improvement, and looked forward to the future development direction of the co‐training algorithm. Xin Ning 0001, Shaohui Xu, Weiwei Cai 0001, Liping Zhang 0014, Wenfa Li |
Concurr. Comput. Pract. Exp. | 1 |
| 2023 | ACGAN: Age-compensated makeup transfer based on homologous continuity generative adversarial network modelabstractAbstract The authors focus on the makeup transformation problem, which refers to the transfer of makeup from a reference face to a source face image while maintaining the source makeup‐free face image. In recent years, makeup transformation has become a hot issue and a lot of research has been conducted on this basis, but there are some limitations in the existing methods, mainly due to the lack of consideration of age factor, which makes the final generated face makeup images appear not natural and lack appearance attractiveness. In order to further solve this problem, an age‐compensated makeup transformation framework based on homology continuity is proposed. In order to achieve a stable and controllable age‐compensation effect, the authors design a new coding module that can map the face makeup semantic vector into the higher feature space and achieve age compensation by adjusting the direction of the semantic vector. Finally, in order to comprehensively evaluate the effectiveness of the authors’ proposed method, a large number of qualitative and quantitative experiments have been conducted, and the experimental results show that the authors’ proposed framework outperforms existing methods. Guoqiang Wu, Feng He 0008, Yimai Jing, Xin Ning 0001, Chen Wang 0026, Bo Jin 0018 |
IET Comput. Vis. | 5 |
| 2023 | 3D human pose and shape estimation via de-occlusion multi-task learning
Hang Ran, Xin Ning 0001, Weijun Li 0002, Meilan Hao, Prayag Tiwari |
Neurocomputing | 2 |
| 2023 | Continuous transfer of neural network representational similarity for incremental learningabstractThe incremental learning paradigm in machine learning has consistently been a focus of academic research. It is similar to the way in which biological systems learn, and reduces energy consumption by avoiding excessive retraining. Existing studies utilize the powerful feature extraction capabilities of pre-trained models to address incremental learning, but there remains a problem of insufficient utilization of neural network feature knowledge. To address this issue, this paper proposes a novel method called Pre-trained Model Knowledge Distillation (PMKD) which combines knowledge distillation of neural network representations and replay. This paper designs a loss function based on centered kernel alignment to transfer neural network representations knowledge from the pre-trained model to the incremental model layer-by-layer. Additionally, the use of memory buffer for Dark Experience Replay helps the model retain past knowledge better. Experiments show that PMKD achieved superior performance on various datasets and different buffer sizes. Compared to other methods, our class incremental learning accuracy reached the best performance. The open-source code is published at https://github.com/TianSongS/PMKD-IL . Songsong Tian, Weijun Li 0002, Xin Ning 0001, Hang Ran, Hong Qin 0007, Prayag Tiwari |
Neurocomputing | 3 |
| 2023 | Quantum detectable Byzantine agreement for distributed data trust management in blockchainabstractNo system entity within a contemporary distributed cyber system can be entirely trusted. Hence, the classic centralized trust management method cannot be directly applied to it. Blockchain technology is essential to achieving decentralized trust management, its consensus mechanism is useful in addressing large-scale data sharing and data consensus challenges. Herein, an n-party quantum detectable Byzantine agreement (DBA) based on the GHZ state to realize the data consensus in a quantum blockchain is proposed, considering the threat posed by the growth of quantum information technology on the traditional blockchain. Relying on the nonlocality of the GHZ state, the proposed protocol detects the honesty of nodes by allocating the entanglement resources between different nodes. The GHZ state is notably simpler to prepare than other multi-particle entangled states, thus reducing preparation consumption and increasing practicality. When the number of network nodes increases, the proposed protocol provides better scalability and stronger practicability than the current quantum DBA. In addition, the proposed protocol has the optimal fault-tolerant found and does not rely on any other presumptions. A consensus can be reached even when there are n−2 traitors. The performance analysis confirms viability and effectiveness through exemplification. The security analysis also demonstrates that the quantum DBA protocol is unconditionally secure, effectively ensuring the security of data and realizing data consistency in the quantum blockchain. Zhiguo Qu, Zhexi Zhang, Prayag Tiwari, Xin Ning 0001, Khan Muhammad 0001 |
Inf. Sci. | 5 |
| 2023 | A smoothing Group Lasso based interval type-2 fuzzy neural network for simultaneous feature selection and system identification
Tao Gao 0003, Chen Wang 0026, Guoqiang Wu, Xin Ning 0001, Xiao Bai 0001, Jian Wang 0010 |
Knowl. Based Syst. | 5 |
| 2023 | Correction to: Underwater target detection with an attention mechanism and improved scale
Long Yu 0001, Shengwei Tian, Pengcheng Feng, Xin Ning 0001 |
Multim. Tools Appl. | 5 |
| 2023 | Hyper-sausage coverage function neuron model and learning algorithm for image classificationabstractRecently, deep neural networks (DNNs) promote mainly by network architectures and loss functions; however, the development of neuron models has been quite limited. In this study, inspired by the mechanism of human cognition, a hyper-sausage coverage function (HSCF) neuron model possessing a high flexible plasticity. Then, a novel cross-entropy and volume-coverage (CE_VC) loss is defined, which compresses the volume of the hyper-sausage to the hilt, and helps alleviate confusion among different classes, thus ensuring the intra-class compactness of the samples. Finally, a divisive iteration method is introduced, which considers each neuron model as a weak classifier, and iteratively increases the number of weak classifiers. Thus, the optimal number of the HSCF neuron is adaptively determined and an end-to-end learning framework is constructed. In particular, to improve the classification performance, the HSCF neuron can be applied to classical DNNs. Comprehensive experiments on eight datasets in several domains demonstrate the effectiveness of the proposed method. The proposed method exhibits the feasibility of boosting DNNs with neuron plasticity and provides a novel perspective for further developments in DNNs. The source code is available at https://github.com/Tough2011/HSCFNet.git . Xin Ning 0001, Weijuan Tian, Feng He 0008, Xiao Bai 0001, Le Sun 0003, Weijun Li 0002 |
Pattern Recognit. | 1 |
| 2023 | Corrigendum to' HCFNN: High-order coverage function neural network for image classification' Pattern Recognition. Volume 131(2022) 108873
Xin Ning 0001, Weijuan Tian, Zaiyang Yu, Weijun Li 0002, Xiao Bai 0001, Yuebao Wang |
Pattern Recognit. | 1 |
| 2023 | Three-dimensional Softmax Mechanism Guided Bidirectional GRU Networks for Hyperspectral Remote Sensing Image Classification
Guoqiang Wu, Xin Ning 0001, Luyang Hou, Feng He 0008, Hengmin Zhang, Achyut Shankar |
Signal Process. | 2 |
| 2023 | Hierarchical Domain Adaptation Projective Dictionary Pair Learning Model for EEG Classification in IoMT SystemsabstractEpilepsy recognition based on electroencephalogram (EEG) and artificial intelligence technology is the main tool of health analysis and diagnosis in Internet of medical things (IoMT). As a distributed learning framework, federated learning can train a shared model from multiple independent edge nodes using local data, which has greatly promoted the development of IoMT. One of the main challenges of EEG-based epilepsy recognition in IoMT is that EEG records show varying distributions in different devices, different times, and different people. This nonstationary characteristic of EEG reduces the accuracy of the recognition model. To improve the classification performance in IoMT, a hierarchical domain adaptation projective dictionary pair learning (HDA-PDPL) model is developed in the study. HDA-PDPL integrates EEG signals from different domains (person, edge nodes, devices, etc.) into a set of hierarchical subspace and simultaneously learns synthesis and analysis dictionary pairs in each layer. Specifically, a nonlinear transform function is introduced to seek hierarchical feature projection. The domain adaptation term on sparse coding builds a connection between different domains. Thus, the shared synthesis and analysis dictionaries can encode domain-invariant representation and discrimination knowledge from different domains. Besides, the local preserved term of projective codes is introduced to capture the potential discriminative local structures of samples. The experimental results on two EEG epilepsy classifications verified that the HDA-PDPL model can outperform other comparisons by utilizing more shared knowledge of different domains. Weiwei Cai 0001, Ming Gao 0026, Yizhang Jiang, Xiaoqing Gu, Xin Ning 0001, Pengjiang Qian, Tongguang Ni |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2023 | Stereo Attention Cross-Decoupling Fusion-Guided Federated Neural Learning for Hyperspectral Image ClassificationabstractFederated learning is a promising solution in several industries for co-training models among distributed clients via centralized servers without leaving private user data on the devices. Thus, federated learning can be seen as a stimulus for the edge computing paradigm as it supports collaborative learning and model optimization. In view of the strict requirements for data security and system reliability of hyperspectral classification techniques for surveillance, aerospace, and military missions, this paper proposes a novel stereo attention cross-decoupling fusion-guided federated neural learning algorithm for hyperspectral image classification, which first trains client devices using a scalable federated learning approach consisting of master server, secure aggregator and edge client devices of a certain size.The distributed devices train local models of the neural network for classifying hyperspectral images and send them to the secure aggregator, which aggregates the local models using a weighted averaging strategy and sends them to the master server for iteration. In addition, the stereo attention cross-decoupling fusion module is used to mine the multidimensional spatial details of the hyperspectral images, specifically by first extracting the most discriminative features from different directions (horizontal, vertical, and spatial) using the attention mechanism, and then using the decoupling fusion strategy to classify the original feature map into three levels: significant, minor, and redundant, and use them to model the multidimensional spatial relationships, thus strengthening the capability to represent features. Extensive experiments on several public datasets have shown that the proposed method provides competitive performance and, more importantly, is effective in enhancing privacy and reliability for hyperspectral image classification. Weiwei Cai 0001, Ming Gao 0026, Yao Ding 0010, Xin Ning 0001, Xiao Bai 0001, Pengjiang Qian |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | A Novel Hyperspectral Image Classification Model Using Bole Convolution With Three-Direction Attention Mechanism: Small Sample and Unbalanced LearningabstractCurrently, the use of rich spectral and spatial information of hyperspectral images (HSIs) to classify ground objects is a research hotspot. However, the classification ability of existing models is significantly affected by its high data dimensionality and massive information redundancy. Therefore, we focus on the elimination of redundant information and the mining of promising features and propose a novel Bole convolution (BC) neural network with a tandem three-direction attention (TDA) mechanism (BTA-Net) for the classification of HSI. A new BC is proposed for the first time in this algorithm, whose core idea is to enhance effective features and eliminate redundant features through feature punishment and reward strategies. Considering that traditional attention mechanisms often assign weights in a one-direction manner, leading to a loss of the relationship between the spectra, a novel three-direction (horizontal, vertical, and spatial directions) attention mechanism is proposed, and an addition strategy and a maximization strategy are used to jointly assign weights to improve the context sensitivity of spatial–spectral features. In addition, we also designed a tandem TDA mechanism module and combined it with a multiscale BC output to improve classification accuracy and stability even when training samples are small and unbalanced. We conducted scene classification experiments on four commonly used hyperspectral datasets to demonstrate the superiority of the proposed model. The proposed algorithm achieves competitive performance on small samples and unbalanced data, according to the results of comparison and ablation experiments. The source code for BTA-Net can be found athttps://github.com/vivitsai/BTA-Net. Weiwei Cai 0001, Xin Ning 0001, Guoxiong Zhou, Xiao Bai 0001, Yizhang Jiang, Wei Li 0032, Pengjiang Qian |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2023 | Graph-Structured Convolution-Guided Continuous Context Threshold-Aware Networks for Hyperspectral Image ClassificationabstractAlthough convolutional neural networks (CNNs) have shown superior performance to traditional machine learning algorithms for hyperspectral image classification tasks, the ability of traditional CNNs to model remote dependencies in the spatial orientation of HSIs is still limited, and they always extract similar low-level features, leading to feature redundancy. To cope with this limitation, this paper proposes a novel multi-order statistical representation-guided graph convolution and continuous context threshold-aware network for the classification of hyperspectral images with limited training samples. Initially, the spectral spatial information is separately modeled using first-order features and second-order pooling operators. Secondly, we propose graph-structuring the patch’s features. By employing a random walk transition probability matrix, graph-structured convolution can mine more discriminative direction features. In addition, we design a continuous context threshold-aware network to model multidimensional spatial relationships, thereby enhancing the representation of graph features. Specifically, the cross-attention mechanism is used to calculate the attention weights in the vertical and horizontal directions, and the features are divided into two levels—important and secondary—by solving the cosine distance between feature vectors, and the former is retained and the latter is punished. Extensive experiments on multiple HSIs datasets demonstrated that the proposed method delivers competitive performance. The code will be available at: https://github.com/vivitsai/GSC-CCTA. Weiwei Cai 0001, Pengjiang Qian, Yao Ding 0010, Meiqiao Bi, Xin Ning 0001, Danfeng Hong, Xiao Bai 0001 |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2022 | Harnessing semantic segmentation masks for accurate facial attribute editingabstractSummary In recent years, with the rapid development of adversarial learning technology, facial attribute editing has made great success in a number of areas. Realistic visual effect, invariant identity information, and accurate editing area are the three key issues of facial attribute editing. Unfortunately, most researches focus on the former two problems. However, lack of awareness of the accurate editing area in the task is the main reason for damaging attribute‐irrelevant details. To address this issue, this article proposes a novel facial attribute editing algorithm—a generative adversarial network (GAN) with semantic masks—from the perspective of editing location accuracy. By generating the mask with respect to attribute‐related areas, the semantic segmentation network can only constrain the manipulation in the target region while not harming any attribute‐irrelevant details. The GAN is then combined with the semantic segmentation network to formulate the entire framework, which is referred to as SM‐GAN. Extensive experiments on the public datasets CelebA and LFWA prove that the presented method can not only ensure that the attribute manipulation is realistic, but also allow attribute‐irrelevant regions to remain unchanged. Moreover, it can also simultaneously edit multiple facial attributes. Xiaoli Dong, Linjun Sun, Weijun Li 0002, Xin Ning 0001, Guojun Wang 0005, Ziheng Chen 0002 |
Concurr. Comput. Pract. Exp. | 7 |
| 2022 | Conditional generative adversarial networks based on the principle of homologycontinuity for face agingabstractAbstract Age is one of the most important biological characteristics of the human face. The increase of age coincides with the increase of the aging degree of the face. Face aging synthesis is attracting increasingly more attention from domestic and overseas scholars in the computer vision and computer graphics fields, and it can be integrated into the basic research of face correlation, such as cross‐age face analysis and age estimation. At present, some achievements have been made in face aging synthesis research; however, it is still an urgent problem to reduce the number of parameters and computational complexity of the network while ensuring the aging effect. Therefore, a new face aging algorithm is proposed in this article. Unlike the previous methods of aging process simulation, we introduce an assisted age classification network based on the principle of homology continuity, which is more in line with the human cognition process. After pretraining, the result of age classification is improved, and the pretraining model is then added to the framework of aging face generation for fine‐tuning to constrain the generated aging face, which can improve the aging accuracy of the generated image. Furthermore, we reconstruct the input face by using the age tag of the input face and the synthesized aging face and maintain the identity invariance in the face aging process by minimizing the reconstruction loss. The experimental results show that the method proposed in this article produces a considerable effect of face aging and significantly reduces the number of parameters and the complexity of computational. Xin Ning 0001, Duoduo Gou, Xiaoli Dong, Weijuan Tian, Chuansheng Wang |
Concurr. Comput. Pract. Exp. | 1 |
| 2022 | AGCNN: Adaptive Gabor Convolutional Neural Networks with Receptive Fields for Vein Biometric RecognitionabstractSummary In recent years, finger vein recognition has attracted more attention and research as a secure method of identification. Convolutional neural networks have achieved great success in the field of finger vein recognition, yet they suffer from high computational complexity, large parameters, and other challenges. To solve these problems, we propose a Gabor convolutional neural network with receptive fields. We use Gabor filters with receptive field properties to design Gabor convolutional layers. Then we replace the conventional convolutional layer with the Gabor convolutional layer; analyze the influence of different loss functions, convolution kernel size, and feature size on the network model; and choose the most suitable model parameters and loss function. Finally, we systematically investigate comparative performance using AGCNN and CNNs in different finger vein databases. Experimental results show that the parameter complexity of AGCNN is significantly less than that of CNNs with a slight performance decrease. Yakun Zhang 0002, Weijun Li 0002, Liping Zhang 0014, Xin Ning 0001, Linjun Sun, Yaxuan Lu |
Concurr. Comput. Pract. Exp. | 4 |
| 2022 | Hybrid Dilated Convolution Guided Feature Filtering and Enhancement Strategy for Hyperspectral Image ClassificationabstractWith the increasing maturity of optics and photonics, hyperspectral technology has also greatly advanced. Hyperspectral images composed of hundreds of adjacent bands and containing useful information can be easily obtained. However, unlike ordinary remote sensing images, each sample in hyperspectral remote sensing images has high-dimensional features and contains rich spatial and spectral information, which greatly increases the difficulty of feature selection and mining, increases the computational complexity, and limits the recognition accuracy of the model. Therefore, in this letter, a novel hybrid dilated-convolution-guided feature filtering and enhancement strategy (HDCFE-Net) model is proposed to classify hyperspectral images. Dilated convolution can reduce the spatial feature loss without reducing the receptive field and can obtain distant features. It can also be combined with the traditional convolution without losing its original information. We propose a feature filtering and enhancement strategy that eliminates redundant features and reduces computational complexity. The core concept is to set a threshold feature value, like the rounding method, to filter and enhance features. Experiments on three well-known hyperspectral datasets—Indian Pines (IPs), Pavia University (PU), and Salinas—show that in less than 1% (IPs: 5%) of the training samples, the overall accuracy (OA) of our method is 77%, 89%, and 91%, respectively, which is superior to several well-known methods. The experiments demonstrated the effectiveness and superiority of HDCFE-Net. Runmin Liu, Weiwei Cai 0001, Guangjun Li, Xin Ning 0001, Yizhang Jiang |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2022 | HCFNN: High-order coverage function neural network for image classification
Xin Ning 0001, Weijuan Tian, Zaiyang Yu, Weijun Li 0002, Xiao Bai 0001, Yuebao Wang |
Pattern Recognit. | 1 |
| 2022 | Uncertainty estimation for stereo matching based on evidential deep learning
Chen Wang 0026, Xiang Wang 0014, Liang Zhang 0044, Xiao Bai 0001, Xin Ning 0001, Jun Zhou 0001, Edwin R. Hancock |
Pattern Recognit. | 6 |
| 2022 | Multi-Source Domain Transfer Discriminative Dictionary Learning Modeling for Electroencephalogram-Based Emotion RecognitionabstractCognitive computing is dedicated to researching a computing principle and method that can simulate the intelligence ability of human brain. Human emotion is the basic component of human cognitive activities. Electroencephalogram (EEG) computer signals obtained from a brain computer interface are difficult to conceal, and using machine learning methods to analyze EEG emotion is a hot topic in artificial intelligence. However, the EEG signal is non-stationary, making it difficult to select sufficient data from the same person to train a classifier for a subject. To promote the performance of emotion recognition methods, a multi-source domain transfer discriminative dictionary learning modeling (MDTDDL) is proposed in this study. The method integrates transfer learning and dictionary learning in a learning model, including the concepts of subspace learning, manifold smoothness, margin-based discriminant embedding, and large margin. The domain-specific transformation matrix projects EEG signals from various domains into the transfer subspace. The domain-invariant dictionary can find potential connections between multiple source domains and target domain. The manifold smoothness and margin-based discriminant embedding term further improve the model’s learning ability. The alternating optimization technique is used in model solving to efficiently compute model parameters. Experiments on the SEED and DEAP datasets demonstrate the effectiveness of MDTDDL. Xiaoqing Gu, Weiwei Cai 0001, Ming Gao 0026, Yizhang Jiang, Xin Ning 0001, Pengjiang Qian |
IEEE Trans. Comput. Soc. Syst. | 5 |
| 2022 | Learning Discriminative Features by Covering Local Geometric Space for Point Cloud AnalysisabstractAt present, effectively aggregating and transferring the local features of point cloud is still an unresolved technological conundrum. In this study, we propose a new space-cover convolutional neural network (SC-CNN) for tasks such as point cloud classification and segmentation. The core of this network is space-cover convolution (SC-Conv), which implements depthwise separable convolution on the point cloud. In addition, a newly designed space-cover operator (SCOP) replaces depthwise convolution. The key to SC-Conv is constructing anisotropic spatial geometry in the local point cloud. The SCOP achieves this by utilizing the positional and feature relationships to learn the high-order relationship expression between points. First, data-driven adaptive learning from the 3-D coordinate relationship between the local points is used to determine the weight of the SCOP. Then, the edge feature of the neighboring point relative to the sampling point is used as the input of the SCOP. Finally, a deformable spatial geometry is constructed in the feature space between local points to aggregate the local high-order features. By stacking SC-Conv to construct SC-CNN with a hierarchical network structure for point cloud analysis, we can better perceive the shape information of point cloud and improve network robustness. Finally, we provide numerous experiments to verify that SC-CNN parallels or even outperforms advanced methods in shape classification, part segmentation, and large-scale indoor scene segmentation tasks. The open-source code was published athttps://github.com/changshuowang/SC-CNN. Changshuo Wang 0001, Xin Ning 0001, Linjun Sun, Liping Zhang 0014, Weijun Li 0002, Xiao Bai 0001 |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2022 | An Adaptive Clustering-Based Algorithm for Automatic Path Planning of Heterogeneous UAVsabstractDue to the high maneuverability and strong adaptability, autonomous unmanned aerial vehicles (UAVs) are of high interest to many civilian and military organizations around the world. Automatic path planning which autonomously finds a good enough path that covers the whole area of interest, is an essential aspect of UAV autonomy. In this study, we focus on the automatic path planning of heterogeneous UAVs with different flight and scan capabilities, and try to present an efficient algorithm to produce appropriate paths for UAVs. First, models of heterogeneous UAVs are built, and the automatic path planning is abstracted as a multi-constraint optimization problem and solved by a linear programming formulation. Then, inspired by the density-based clustering analysis and symbiotic interaction behaviours of organisms, an adaptive clustering-based algorithm with a symbiotic organisms search-based optimization strategy is proposed to efficiently settle the path planning problem and generate feasible paths for heterogeneous UAVs with a view to minimizing the time consumption of the search tasks. Experiments on randomly generated regions are conducted to evaluate the performance of the proposed approach in terms of task completion time, execution time and deviation ratio. Jinchao Chen, Ying Zhang 0060, Lianwei Wu, Tao You, Xin Ning 0001 |
IEEE Trans. Intell. Transp. Syst. | 5 |
| 2022 | Beyond Triplet Loss: Person Re-Identification With Fine-Grained Difference-Aware Pairwise LossabstractPerson Re-IDentification (ReID) aims at re-identifying persons from different viewpoints across multiple cameras. Capturing the fine-grained appearance differences is often the key to accurate person ReID, because many identities can be differentiated only when looking into these fine-grained differences. However, most state-of-the-art person ReID approaches, typically driven by a triplet loss, fail to effectively learn the fine-grained features as they are focused more on differentiating large appearance differences. To address this issue, we introduce a novel pairwise loss function that enables ReID models to learn the fine-grained features by adaptively enforcing an exponential penalization on the images of small differences and a bounded penalization on the images of large differences. The proposed loss is generic and can be used as a plugin to replace the triplet loss to significantly enhance different types of state-of-the-art approaches. Experimental results on four benchmark datasets show that the proposed loss substantially outperforms a number of popular loss functions by large margins; and it also enables significantly improved data efficiency. Guansong Pang, Xiao Bai 0001, Changhong Liu, Xin Ning 0001, Lin Gu 0003, Jun Zhou 0001 |
IEEE Trans. Multim. | 5 |
| 2022 | Encoder-X: Solving Unknown Coefficients Automatically in Polynomial Fitting by Using an AutoencoderabstractModeling, prediction, and recognition tasks depend on the proper representation of the objective curves and surfaces. Polynomial functions have been proved to be a powerful tool for representing curves and surfaces. Until now, various methods have been used for polynomial fitting. With a recent boom in neural networks, researchers have attempted to solve polynomial fitting by using this end-to-end model, which has a powerful fitting ability. However, the current neural network-based methods are poor in stability and slow in convergence speed. In this article, we develop a novel neural network-based method, called Encoder-X, for polynomial fitting, which can solve not only the explicit polynomial fitting but also the implicit polynomial fitting. The method regards polynomial coefficients as the feature value of raw data in a polynomial space expression and therefore polynomial fitting can be achieved by a special autoencoder. The entire model consists of an encoder defined by a neural network and a decoder defined by a polynomial mathematical expression. We input sampling points into an encoder to obtain polynomial coefficients and then input them into a decoder to output the predicted function value. The error between the predicted function value and the true function value can update parameters in the encoder. The results prove that this method is better than the compared methods in terms of stability, convergence, and accuracy. In addition, Encoder-X can be used for solving other mathematical modeling tasks. Guojun Wang 0005, Weijun Li 0002, Liping Zhang 0014, Linjun Sun, Xin Ning 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2021 | Costly Features Classification using Monte Carlo Tree SearchabstractIn many real-world tasks, acquiring features requires a certain cost, which gives rise to the costly features classification problem. In this study, We formulate the problem in the reinforcement learning framework and sequentially select the subset of features to make a balance between the classification error and the feature cost. Specifically, advantage actor critic algorithm is firstly used to solve it. Furthermore, to improve the learned policy and make it explainable, we employ the Monte Carlo Tree Search to update the policy iteratively. During the procedure, we also consider its performance on imbalanced datasets. Our empirical evaluation shows that our method performs well in comparison with other traditional methods. Ziheng Chen 0002, Jin Huang 0010, Hongshik Ahn, Xin Ning 0001 |
IJCNN | 4 |
| 2021 | JWSAA: Joint weak saliency and attention aware for person re-identification
Xin Ning 0001, Weijun Li 0002, Liping Zhang 0014 |
Neurocomputing | 1 |
| 2021 | MC-Net: Multiple max-pooling integration module and cross multi-scale deconvolution network
Hongfeng You, Long Yu 0001, Shengwei Tian, Yan Xing 0004, Xin Ning 0001, Weiwei Cai 0001 |
Knowl. Based Syst. | 6 |
| 2021 | Underwater target detection with an attention mechanism and improved scale
Long Yu 0001, Shengwei Tian, Pengcheng Feng, Xin Ning 0001 |
Multim. Tools Appl. | 5 |
| 2021 | Feature Refinement and Filter Network for Person Re-IdentificationabstractIn the task of person re-identification, the attention mechanism and fine-grained information have been proved to be effective. However, it has been observed that models often focus on the extraction of features with strong discrimination, and neglect other valuable features. The extracted fine-grained information may include redundancies. In addition, current methods lack an effective scheme to remove background interference. Therefore, this paper proposes the feature refinement and filter network to solve the above problems from three aspects: first, by weakening the high response features, we aim to identify highly valuable features and extract the complete features of persons, thereby enhancing the robustness of the model; second, by positioning and intercepting the high response areas of persons, we eliminate the interference arising from background information and strengthen the response of the model to the complete features of persons; finally, valuable fine-grained features are selected using a multi-branch attention network for person re-identification to enhance the performance of the model. Our extensive experiments on the benchmark Market-1501, DukeMTMC-reID, CUHK03 and MSMT17 person re-identification datasets demonstrate that the performance of our method is comparable to that of state-of-the-art approaches. Xin Ning 0001, Weijun Li 0002, Liping Zhang 0014, Xiao Bai 0001, Shengwei Tian |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2020 | Continuous Learning of Face Attribute SynthesisabstractThe generative adversarial network (GAN) exhibits great superiority in the face attribute synthesis task. However, existing methods have very limited effects on the expansion of new attributes. To overcome the limitations of a single network in new attribute synthesis, a continuous learning method for face attribute synthesis is proposed in this work. First, the feature vector of the input image is extracted and attribute direction regression is performed in the feature space to obtain the axes of different attributes. The feature vector is then linearly guided along the axis so that images with target attributes can be synthesized by the decoder. Finally, to make the network capable of continuous learning, the orthogonal direction modification module is used to extend the newly-added attributes. Experimental results show that the proposed method can endow a single network with the ability to learn attributes continuously, and, as compared to those produced by the current state-of-the-art methods, the synthetic attributes have higher accuracy. Xin Ning 0001, Weijun Li 0002, Xiaoli Dong, Shaohui Xu, Fangzhe Nan, Yuanzhou Yao |
ICPR | 1 |
| 2020 | A Local Descriptor with Physiological Characteristic for Finger Vein RecognitionabstractLocal feature descriptors exhibit great superiority in finger vein recognition due to their stability and robustness against local changes in images. However, most of these are methods use general-purpose descriptors that do not consider finger vein-specific features. In this work, we propose a finger vein-specific local feature descriptors based physiological characteristic of finger vein patterns, i.e., histogram of oriented physiological Gabor responses (HOPGR), for finger vein recognition. First, a prior of directional characteristic of finger vein patterns is obtained in an unsupervised manner. Then the physiological Gabor filter banks are set up based on the prior information to extract the physiological responses and orientation. Finally, to make the feature robust against local changes in images, a histogram is generated as output by dividing the image into non-overlapping cells and overlapping blocks. Extensive experimental results on several databases clearly demonstrate that the proposed method outperforms most current state-of-the-art finger vein recognition methods. Liping Zhang 0014, Weijun Li 0002, Xin Ning 0001, Linjun Sun, Xiaoli Dong |
ICPR | 3 |
| 2020 | Real-Time 3D Face Alignment Using an Encoder-Decoder Network With an Efficient Deconvolution LayerabstractIn the field of 3D face alignment, most researchers have focused on improving the prediction accuracy of algorithms and ignored the portability for practical applications. To this end, this study presents a real-time 3D face-alignment method that uses an encoder-decoder network with an efficient deconvolution layer. The fusion of the encoding and decoding feature adds more abundant features to this network. An efficient deconvolution layer at the decoding stage applies the L1 norm to select useful features and generate abundant ones through linear operations. Experimental results using the standard AFLW2000-3D and AFLW-LFPA datasets show that our algorithm has low prediction errors with real-time applicability. Xin Ning 0001, Pengfei Duan 0003, Weijun Li 0002 |
IEEE Signal Process. Lett. | 1 |
| 2019 | Faster Real-Time Face Alignment Method on CPU
Pengfei Duan 0003, Xin Ning 0001, Weijun Li 0002 |
PRCV (1) | 2 |
| 2018 | Deep Adaptive Update of Discriminant KCF for Visual Tracking
Xin Ning 0001, Weijun Li 0002, Weijuan Tian, Xuchi, Dongxiaoli, Zhangliping |
ICONIP (6) | 1 |
| 2018 | Face Anti-spoofing based on Deep Stack Generalization Networks
Xin Ning 0001, Weijun Li 0002, Meili Wei, Linjun Sun, Xiaoli Dong |
ICPRAM | 1 |
| 2018 | The Principle of Homology Continuity and Geometrical Covering Learning for Pattern RecognitionabstractHomology Continuity is a fundamental property of the nature, but few of the traditional pattern recognition algorithms were aware of it. Firstly, this paper gives a brief description to the Principle of Homology Continuity (PHC), and tries to mathematically redefine it. Then, we introduce a PHC-based pattern learning method — Geometrical Covering Learning (GCL), following the Hyper sausage neural network as an instance of GCL. Lastly, we propose a GCL solution to the “two-spirals” pattern recognition problem. The final experimental results show that the new method is feasible and efficient. Xin Ning 0001, Weijun Li 0002 |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2018 | BULDP: Biomimetic Uncorrelated Locality Discriminant Projection for Feature Extraction in Face RecognitionabstractThis paper develops a new dimensionality reduction method, named Biomimetic Uncorrelated Locality Discriminant Projection (BULDP), for face recognition. It is based on unsupervised discriminant projection and two human bionic characteristics: principle of homology continuity and principle of heterogeneous similarity. With these two human bionic characteristics, we propose a novel adjacency coefficient representation, which does not only capture the category information between different samples, but also reflects the continuity between similar samples and the similarity between different samples. By applying this new adjacency coefficient into the unsupervised discriminant projection, it can be shown that we can transform the original data space into an uncorrelated discriminant subspace. A detailed solution of the proposed BULDP is given based on singular value decomposition. Moreover, we also develop a nonlinear version of our BULDP using kernel functions for nonlinear dimensionality reduction. The performance of the proposed algorithms is evaluated and compared with the state-of-the-art methods on four public benchmarks for face recognition. Experimental results show that the proposed BULDP method and its nonlinear version achieve much competitive recognition performance. Xin Ning 0001, Weijun Li 0002, Bo Tang 0011, Haibo He |
IEEE Trans. Image Process. | 1 |