Lin Cao 0003

dblp:00/1183-3 · DBLP profile ↗
← Back
17ranked-venue papers
3as first author
17since 2021 · last 2027
0000-0003-0875-1549ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 2 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Security and privacy · 1 · 1 since 2021
YearPublicationVenuePosition
2027 PP-LLMs: A progressive pruning approach with medium-granularity for large language models
Kangning Du, Yinkai Wang, Benkui Zhang, Jinxiao Wang, Lin Cao 0003
Expert Syst. Appl.6
2026 Dual Student Discrepancy Correction for Semi-Supervised Medical Image Segmentation
abstract
ABSTRACT Semi‐supervised medical image segmentation (SSMIS) has proven to be an effective solution that leverages limited labelled data and abundant unlabeled data, thereby significantly reducing the labour and cost associated with manual annotation. However, most of the existing teacher‐student frameworks are prone to suffer from confirmation bias during training, adversely affecting the performance of SSMIS. To address this challenge, we propose the Dual Student Discrepancy Correction framework (DSDC), which extends the Mean Teacher (MT) framework by incorporating an additional student model with identical architecture but independently updated parameters. This design mitigates the parameter coupling issue that may arise when updating the teacher model via Exponential Moving Average (EMA) in conventional single‐student paradigms. Moreover, the prediction discrepancy between the two student models is leveraged for error detection and correction, enabling the network to identify and rectify its own cognitive biases, ultimately enhancing segmentation accuracy. Comprehensive experiments on two public benchmarks, an MRI dataset (LA) and a CT dataset (Pancreas‐NIH), reveal that our DSDC framework surpasses current State‐of‐the‐Art (SOTA) approaches across all evaluation metrics. These findings substantiate the framework's effectiveness in SSMIS tasks. Code is accessible at https://github.com/Sangfugui/DSDC .
Zhenfu Sang, Chong Fu 0001, Lin Cao 0003, Chiu-Wing Sham
Expert Syst. J. Knowl. Eng.4
2026 DEG-NER: dynamic graph-enhanced diffusion model for named entity recognition
Lin Cao 0003, Zhongxue Jia, Benkui Zhang, Huanyu Bian, Fan Zhang 0037, Kangning Du, Yanan Guo 0003
Expert Syst. Appl.1
2026 Deep supervised anomaly detection for generalized face forgery detection
abstract
Nowadays, face forgery poses a significant threat to societal security, making the development of effective countermeasures imperative. Though most existing methods adopt neural networks to automatically extract discriminative features for forgery detection and have achieved promising results, significant challenges remain. Namely, when detecting forgery faces generated by unseen forgery methods, the detection performance degrades significantly, indicating poor generalization capability. To address such limitation, a novel deep supervised anomaly detection for generalized face forgery detection (DAGFD) is proposed in this paper. Specifically, the artifact map detector optimized by triplet focal loss and metric-softmax loss is first used to locate the forgery regions and obtain artifact maps. Next, forgery detection is reformulated from the supervised anomaly detection perspective, and the artifact map score is calculated to detect forgery videos. Furthermore, mean square error (MSE) loss is used to minimize the artifact map score of real samples while increase the one of forgery samples to generalize well to unseen forgery methods. Also, circle loss is used for auxiliary classifier to learn more discriminative artifact features. Finally, the experimental results demonstrate that the proposed method’s detection accuracy is better than other state-of-the-art methods.
Fan Zhang 0037, Lin Cao 0003, Kangning Du, Yanan Guo 0003, Peiran Song, Chen Shao
Pattern Recognit.3
2026 Dual-Granularity Contrastive Learning for DeepFake Detection
abstract
In recent years, contrastive learning has made significant progress in DeepFake detection. However, existing methods emphasize class granularity, and it is difficult to distinguish between the real instance and its forgery counterparts effectively. Furthermore, the diversity of forgery cues produced by different manipulation methods cannot be effectively clustered by class granularity alone. Thus, the model’s generalization capability is limited. To tackle the above problems, a Dual-Granularity Contrastive Learning (DGCL) for DeepFake detection is proposed in this paper. Specifically, Class Granularity Contrastive Learning (CGCL) and Instance Granularity Contrastive Learning (IGCL) are designed. Firstly, for semantic aggregation at the class level, CGCL incorporates the class prototype, which encourages anchor approaches to the prototype of the positive class, thereby pulling the intra-class features closer. Secondly, for distinguishing between real and fake instances, Real Instance Granularity Contrastive Learning (RIGCL) and Fake Instance Granularity Contrastive Learning (FIGCL) are proposed based on the instance characteristics. RIGCL endeavors to distinguish fake instances from original real instances by expanding the differentiation in the feature space. Meanwhile, FIGCL extracts consistent forgery features from various manipulation methods using cosine similarity constraints. Finally, the superiority and generalizability of DGCL are validated by the experimental results on CELEBDF, DFD, and DFDC datasets.
Fan Zhang 0037, Chen Shao, Kangning Du, Yanan Guo 0003, Peiran Song, Lin Cao 0003, Xin Yuan 0004
IEEE Trans. Inf. Forensics Secur.6
2025 Multi-View Normal and Distance Guidance Gaussian Splatting for Surface Reconstruction
abstract
3D Gaussian Splatting (3DGS) achieves remarkable results in the field of surface reconstruction. However, when Gaussian normal vectors are aligned within the single-view projection plane, while the geometry appears reasonable in the current view, biases may emerge upon switching to nearby views. To address the distance and global matching challenges in multi-view scenes, we design multi-view normal and distance-guided Gaussian splatting. This method achieves geometric depth unification and high-accuracy reconstruction by constraining nearby depth maps and aligning 3D normals. Specifically, for the reconstruction of small indoor and outdoor scenes, we propose a multi-view distance reprojection regularization module that achieves multi-view Gaussian alignment by computing the distance loss between two nearby views and the same Gaussian surface. Additionally, we develop a multiview normal enhancement module, which ensures consistency across views by matching the normals of pixel points in nearby views and calculating the loss. Extensive experimental results demonstrate that our method outperforms the baseline in both quantitative and qualitative evaluations, significantly enhancing the surface reconstruction capability of 3DGS.
Bo Jia, Yanan Guo 0003, Ying Chang, Benkui Zhang, Kangning Du, Lin Cao 0003
IROS7
2025 IRAGKR:Iterative retrieval augmented generation with fine-grained knowledge refinement
Kangning Du, Benkui Zhang, Fan Zhang 0037, Lin Cao 0003, Yanan Guo 0003
Neurocomputing6
2025 CTIDRNet: Cross-Temporal Interaction With Difference Refinement Network for Remote Sensing Image Change Detection
abstract
Remote sensing change detection (RSCD) has achieved creditable success in recent years. However, the challenge of identifying changed objects with shape details persists in RSCD. In this letter, we proposed a cross-temporal interaction with difference refinement network (CTIDRNet) to solve interference-caused fake change and incomplete irregular change shape in RSCD tasks. Specifically, by combining cross-attention and self-attention to steer the temporal feature interaction of each input, we design a temporal feature attention (TFA) module to excavate the potential relation of change areas and suppress the unchanged object interference. Afterward, a deformable convolution is used to design a difference feature refinement (DFR) architecture to capture temporal difference information at diverse feature levels. At last, we proposed a multiscale-guided fusion (MGF) module to fuse pyramid features, thereby dealing with scaling changes. Experimental results on three datasets show that CTIDRNet can extract irregularly changed areas effectively, and the evaluation result outperforms other SOTA methods, with an improvement of 1.79%–19.82%, 2.9%–11.07%, and 0.97%–8.91% in terms of F1 for CDD, SYSU, and LEVIR datasets, respectively. The demo code of this work is publicly available athttps://github.com/lucyjiong/CTIDR.
Kangning Du, Xian Sun 0001, Lin Cao 0003, Shu Tian
IEEE Geosci. Remote. Sens. Lett.4
2025 MSSI-Net: Multiscale Semantic-Guided Synergistic Interaction Network for Remote Sensing Image Change Detection
abstract
Remote sensing change detection (RSCD) has become an essential tool in observing and analyzing geographical information. However, existing deep learning approaches dependent solely on visual modalities may encounter challenges in discerning subtle variations amidst noise interference. To overcome these issues, we propose a multiscale semantic-guided synergistic interaction network (MSSI-Net), which utilizes the advanced multimodal semantic representations for enhancing the capacity to perceive hierarchical changes. Specifically, we first devise a multiscale interaction module (MIM) which leverages multiscale attention mechanism to guide the interaction between the coarse and fine stages of different visual features. The fine-grained visual features subsequently complement the semantic features through scale weight reassignment to enhance the discriminative capability of vision-language features. Furthermore, driven by the semantic-guided synergistic interaction mechanism, our developed cross-modal feature fusion module (CFFM) exploits both homogeneous and heterogeneous features among modalities. This ensures that the generated vision-language features are semantically representative. Finally, we formulate a manifold differential perception head (MDPH) to optimize the detection of changes by efficiently fusing diverse differential feature representations, achieving comprehensive performance enhancement. Extensive experiments conducted on four benchmark datasets (LEVIR-CD, CDD, SYSU-CD and WHU-CD) indicate that the designed MSSI-Net achieves state-of-the-art performance compared to existing methods.
Shu Tian, Jiyuan Shen, Lin Cao 0003, Lihong Kang, Xian Sun 0001, Xiangwei Xing, Chunzhuo Fan, Kangning Du, Chong Fu 0001, Ye Zhang 0008
IEEE Trans. Geosci. Remote. Sens.3
2024 Cross-Modal Dual Matching and Comparison for Text-to-Image Person Re-identification
Lin Cao 0003, Yanan Guo 0003, Shoujing Wang, Boqian Lv
PRCV (5)1
2024 Joint object contour points and semantics for instance segmentation
abstract
Abstract The edges of objects are of great significance to the task of instance segmentation. However, most of the current popular deep neural networks do not pay much attention to the object edge information. More importantly, using the down‐sampling pooling layer in the deep learning network, the edge detail information of the object will be lost. To address this issue, inspired by the manual annotation process, we propose Mask Point R‐CNN aiming at promoting the neural network's attention to the object boundary. Specifically, we introduce the auxiliary task of object contour point detection on the Mask R‐CNN framework, which can effectively improve the gradient flow between different tasks by multi‐task learning and repairing objects' boundary information via feature fusion. Consequently, the model can be more sensitive to the edges of the object and capture more geometric features. Quantitatively, the experimental results show that our Mask Point R‐CNN outperforms vanilla Mask R‐CNN by 3.8% on the Cityscapes dataset and 0.8% on the COCO dataset.
Wenchao Zhang 0001, Chong Fu 0001, Mai Zhu, Lin Cao 0003, Ming Tie, Chiu-Wing Sham
Expert Syst. J. Knowl. Eng.4
2024 A modality separation approach for facial sketch synthesis
abstract
Abstract The technology for face‐to‐sketch synthesis transforms optical face images into a sketch‐style format. However, traditional style losses are insufficient to discern the modal differences between optical and sketch domain images, leading to unclear images. At the same time, generated images lack clarity due to traditional approaches' disregard for high‐frequency texture. To address these issues, a modality separation approach for facial sketch synthesis is proposed. First, a modality separation structure is proposed, using a quicksort algorithm to merge features of optical and sketch images as target modality (positive samples), ensuring the generated images' feature distribution matches real sketches. By controlling the Euclidean distance between generated images (anchors) and both target and filtered modality (positive and negative samples), irrelevant information is effectively filtered out. Next, an edge‐promoting module feeds processed blurry sketch images into the discriminator to enhance robustness. Lastly, a detail optimization module uses Laplacian filtering to extract high‐frequency texture from optical face images for local enhancement. Experimental validation on CUHK, AR, and XM2VTS datasets shows that this method outperforms mainstream sketch face synthesis methods in terms of Fréchet inception distance and learned perceptual image patch similarity, producing more realistic and natural images with richer texture details.
Kangning Du, Lin Cao 0003, Yanan Guo 0003
IET Image Process.3
2024 STOD: toward semi-supervised tiny object detection
Yanan Guo 0003, Kangning Du, Lin Cao 0003
Neural Comput. Appl.4
2024 An efficient chaotic image encryption scheme using simultaneous permutation-diffusion operation
Qingxin Sheng, Chong Fu 0001, Zhaonan Lin, Junxin Chen 0001, Lin Cao 0003, Chiu-Wing Sham
Vis. Comput.5
2023 Sketch face recognition based on light semantic Transformer network
abstract
Abstract Sketch face recognition has a wide range of applications in criminal investigation, but it remains a challenging task due to the small‐scale sample and the semantic deficiencies caused by cross‐modality differences. The authors propose a light semantic Transformer network to extract and model the semantic information of cross‐modality images. First, the authors employ a meta‐learning training strategy to obtain task‐related training samples to solve the small sample problem. Then to solve the contradiction between the high complexity of the Transformer and the small sample problem of sketch face recognition, the authors build the light semantic transformer network by proposing a hierarchical group linear transformation and introducing parameter sharing, which can extract highly discriminative semantic features on small–scale datasets. Finally, the authors propose a domain‐adaptive focal loss to reduce the cross‐modality differences between sketches and photos and improve the training effect of the light semantic Transformer network. Extensive experiments have shown that the features extracted by the proposed method have significant discriminative effects. The authors’ method improves the recognition rate by 7.6% on the UoM‐SGFSv2 dataset, and the recognition rate reaches 92.59% on the CUFSF dataset.
Lin Cao 0003, Jianqiang Yin, Yanan Guo 0003, Kangning Du, Fan Zhang 0037
IET Comput. Vis.1
2022 CODH++: Macro-semantic differences oriented instance segmentation network
Wenchao Zhang 0001, Chong Fu 0001, Lin Cao 0003, Chiu-Wing Sham
Expert Syst. Appl.3
2022 Protection of image ROI using chaos-based encryption and DCNN-based object detection
Chong Fu 0001, Yu Zheng 0021, Lin Cao 0003, Ming Tie, Chiu-Wing Sham
Neural Comput. Appl.4