Canghong Shi

dblp:188/7866 · DBLP profile ↗
← Back
23ranked-venue papers
3as first author
16since 2021 · last 2026
0000-0002-5188-6230ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 8 since 2021Security and privacy · 2Software engineering, systems software and programming languages · 2Databases, data management, data science and information retrieval · 2 · 2 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 AMOS: Absent minority oversampling neural network for imbalanced data classification
Zhan ao Huang, Canghong Shi, Jia He 0003, Xiaojie Li 0001, Xi Wu 0004
Inf. Sci.2
2026 KFMF: A Keyframe-Matching Framework for long duration audio copy-move forgery detection
Canghong Shi, Xiaojie Li 0001, Minfeng Shao, Xianhua Niu
Speech Commun.2
2025 CAPAST: Content Affinity Preserved Arbitrary Style Transfer
abstract
Balancing the consistency of style and the integrity of content is the main challenge in arbitrary style transfer domain. Currently, local style details can be effectively captured by attention mechanism but easily produce distorted style patterns and inconsistent content structure. In this paper, we propose a Content Affinity Preserving Arbitrary Style Transfer (CAPAST) framework to ensure style features can be stably integrated into the content structure. Considering the local feature learning ability of CNN and the global feature representation advantage of transformer, a dual encoder is proposed to capture local and global features of images with the combination between transformer and CNN. In addition, a channel and spatially aligned attention (CSAA) is introduced to generate high-quality results by stably fusing style features and content features. In experiments, we demonstrated the superior performance of our method in preventing content structure distortion and maintaining consistency between style and content. Codes are available at https://github.com/miaopashi-zxy/CAPAST.
Xinyuan Zheng, Xiaojie Li 0001, Canghong Shi, Jia He 0003, Zhan ao Huang, Xian Zhang 0008, Imran Mumtaz
ICASSP3
2025 Wave Height Prediction: 3D Spatiotemporal FourCastNet Method with Multi-Factor
abstract
Accurate wave height prediction is essential for various marine operations. However, the complexity of the marine environment, influenced by numerous factors, underscores the importance of effectively leveraging available data. This paper introduces a novel spatiotemporal model, FourCastNet, which serves as a baseline for capturing wave height trends through spectral analysis. To refine spatiotemporal representations, we incorporate 3D convolution to extract local features from the data via nonlinear transformations. Employing Fourier transform, we convert the features into the frequency domain, filtering out high-frequency noise while enhancing the visibility of spatiotemporal patterns. Finally, we fuse various features and capture complex patterns via channel mixing, thereby enhancing the model’s predictive accuracy. Iterative forecasting techniques are applied to minimize prediction errors. Experiments using French wave reanalysis data, projected up to 24 hours at three-hour intervals, demonstrate that our method achieves approximately a 10% improvement in accuracy over existing state-of-the-art techniques, particularly for short-term forecasts within the first 6 hours.
Chenchen He, Zhanao Huang, Canghong Shi, Xiaojie Li 0001, Xi Wu 0004
IJCNN4
2025 Memo-UNet: Leveraging historical information for enhanced wave height prediction
Teng Fang, Xiaojie Li 0001, Canghong Shi, Xian Zhang 0008, Yi Kou, Imran Mumtaz, Zhan ao Huang
Neurocomputing3
2025 S-Faster R-CNN: Intraspectral Similarity Learning for Audio Copy-Move Forgery Localization in IoT Security
abstract
In the Internet of Audio Things, communication security of the audio control terminal is vulnerable to copy-move threats, and detecting and locating audio copy-move forgery remains challenging nowadays. The forgery detection method based on deep learning achieves higher detection accuracy but fails to localize forged regions. To address this issue, this article proposes an S-Faster R-CNN model for audio copy-move forgery detection and localization. We integrate a novel Similarity Computation Module (SCM) into the Faster R-CNN framework, forming the S-Faster R-CNN model. Obtaining the integration of the SCM, which allows the S-Faster R-CNN to precisely localize forgery regions within the spectrogram. Finally, the image coordinate transformation algorithm is used to map these forged regions to the corresponding locations of the original audio waveform, thus completing the audio copy-move forgery detection and localization. Evaluated on three datasets, our method achieves an average recall of 90%, an average precision of 84%, and an average F1-score of 87%, respectively. Experimental results indicate that the S-Faster R-CNN outperforms state-of-the-art methods in both forgery detection accuracy and especially in localization. Moreover, the proposed method shows good robustness under multiple post-processing.
Canghong Shi, Xiaojie Li 0001, Sani M. Abdullahi
IEEE Internet Things J.2
2025 Robust copy-move detection and localization of digital audio based CFCC feature
Xiaojie Li 0001, Canghong Shi, Xianhua Niu, Ling Xiong, Hanzhou Wu, Qing Qian 0001
Multim. Tools Appl.3
2025 An Explanation Method Based on Interpretable Linear Model With Four Key Characteristics
abstract
For the interpretability of deep neural networks (DNNs) in visual-related tasks, existing explanation methods commonly generate a saliency map based on the linear relation between output results and input features. However, when the explanation conflicts with a human visual examination, these methods do not provide further evidence to analyze the saliency explanation. Most may fail to provide feature attribution with identifiable semantics or produce misleading explanations due to their insufficient robustness. In this paper, we first propose four key characteristics (richness, adaptivity, exclusiveness, and fairness) to evaluate the existing linear relation-based explanation method, and then construct an interpretable linear model to satisfy them. We formalize the characteristics and develop a novel explanation method based on this. We extract and reconstruct key exclusive semantic features from the feature map using the Nonnegative Matrix Factorization (NMF) algorithm, utilize the information entropy model to determine the number of features adaptively and their richness, and then linearly combine each feature with fairly assigned weights using an approximate Shapley algorithm to generate the saliency map. Compared with the state-of-the-art methods, our explanations of different datasets and DNNs are more convincing and robust in terms of Average drop (AD), Average increase (AI), Deletions (Del), and Insertions (Ins). Our supplementary experiments provide sufficient evidence that the four characteristics guarantee the feasibility of feature attribution analysis and enhance the quality of the resulting explanations.
Yuecan Yuan, Zhan ao Huang, Ying Fu 0003, Xuemin Zhao, Canghong Shi, Xiaojie Li 0001, Xi Wu 0004
IEEE Trans. Image Process.6
2024 CAGAN: Classifier-augmented generative adversarial networks for weakly-supervised COVID-19 lung lesion localisation
abstract
Abstract The Coronavirus Disease 2019 (COVID‐19) epidemic has constituted a Public Health Emergency of International Concern. Chest computed tomography (CT) can help early reveal abnormalities indicative of lung disease. Thus, accurate and automatic localisation of lung lesions is particularly important to assist physicians in rapid diagnosis of COVID‐19 patients. The authors propose a classifier‐augmented generative adversarial network framework for weakly supervised COVID‐19 lung lesion localisation. It consists of an abnormality map generator, discriminator and classifier. The generator aims to produce the abnormality feature map M to locate lesion regions and then constructs images of the pseudo‐healthy subjects by adding M to the input patient images. Besides constraining the generated images of healthy subjects with real distribution by the discriminator, a pre‐trained classifier is introduced to enhance the generated images of healthy subjects to possess similar feature representations with real healthy people in terms of high‐level semantic features. Moreover, an attention gate is employed in the generator to reduce the noise effect in the irrelevant regions of M . Experimental results on the COVID‐19 CT dataset show that the method is effective in capturing more lesion areas and generating less noise in unrelated areas, and it has significant advantages in terms of quantitative and qualitative results over existing methods.
Xiaojie Li 0001, Xin Fei, Hongping Ren, Canghong Shi, Xian Zhang 0008, Imran Mumtaz, Xi Wu 0004
IET Comput. Vis.5
2024 Robust audio watermarking algorithm resisting cropping based on SIFT transform
Xiangyi Liu, Xiaojie Li 0001, Xianhua Niu, Canghong Shi, Ling Xiong, Qian Qing
Multim. Tools Appl.4
2024 A novel SVD-based adaptive robust audio watermarking algorithm
Xiangyi Liu, Xiaojie Li 0001, Canghong Shi, Xianhua Niu, Ling Xiong
Multim. Tools Appl.3
2023 Pluralistic Face Inpainting With Transformation of Attribute Information
abstract
Most face-inpainting methods perform well in face repair. However, these methods can only complete a single face image per input. Although existing various image-inpainting methods can achieve pluralistic image inpainting, they typically produce faces with distorted structures or the same texture. To resolve these shortcomings and achieve high-quality diverse face inpainting, we propose PFTANet, a two-stage pluralistic face-inpainting network that transforms attribute information. In the first stage, the face-parsing network is fine-tuned to obtain semantic facial region information. In the second stage, a generator consisting of SNBlock, CF_ShiftBlocks, and CF_MergeBlock, which ensures that high-quality pluralistic face results are generated, is used. Specifically, CF_ShiftBlocks completes pluralistic face generation by transforming the attribute information from the conditional face extracted by the attribute extractor and ensuring the consistency of the attribute information between the conditional and generated faces. CF_MergeBlock ensures structural consistency between the masked and background regions of the generated face using facial region semantic information. A multi-patch discriminator is used to enhance facial detail generation. Experimental results for the CelebA and CelebA-HQ datasets indicated that PFTANet achieved pluralistic and visually realistic face inpainting.
Yang Zhang 0155, Xian Zhang 0008, Canghong Shi, Xi Wu 0004, Xiaojie Li 0001, Jing Peng 0003, Kunlin Cao, Jiancheng Lv 0001, Jiliu Zhou
IEEE Trans. Multim.3
2022 Multistage semantic-aware image inpainting with stacked generator networks
abstract
Deep learning has been widely applied into image inpainting. However, traditional image processing methods (i.e., patch-based and diffusion-based methods) generally fail to produce visually natural contents and semantically reasonable structures due to ineffectively processing the high-level semantic information of images. To solve the problem, we propose a stacked generator networks assisted by patch discriminator for image inpainting by multistage. In the proposed method, our generator network mainly consists of three-layer stacked encoder-decoder architecture, which could fuse different level feature information and achieve image inpainting via a coarse-to-fine hierarchical representation. Meanwhile, we split the masked image into different patches in each layer, which could effectively enlarge the receptive field and extract more useful features of images. Moreover, the patch discriminator is introduced to judge the patches of inpainting image are real or fake. In this way, our network can effectively utilize the semantic information to complete a fine result. Furthermore, both perceptual loss and style loss are used to improve the inpainting results in verse. Experimental results on Places2 and Paris StreetView illustrate that our approach could generate high-quality inpainting results, and our method is more effective than the existing image inpainting methods.
Yongpeng Ren, Hongping Ren, Canghong Shi, Xian Zhang 0008, Xi Wu 0004, Xiaojie Li 0001, Jiancheng Lv 0001, Jiliu Zhou, Imran Mumtaz
Int. J. Intell. Syst.3
2022 DE-GAN: Domain Embedded GAN for High Quality Face Image Inpainting
Xian Zhang 0008, Xin Wang 0045, Canghong Shi, Xiaojie Li 0001, Bin Kong 0001, Siwei Lyu, Bin B. Zhu, Jiancheng Lv 0001, Youbing Yin, Qi Song 0001, Xi Wu 0004, Imran Mumtaz
Pattern Recognit.3
2022 Image outpainting guided by prior structure information
Canghong Shi, Yongpeng Ren, Xiaojie Li 0001, Imran Mumtaz, Zhiheng Jin, Hongping Ren
Pattern Recognit. Lett.1
2021 A novel NMF-based authentication scheme for encrypted speech in cloud computing
Canghong Shi, Hongxia Wang 0001, Xiaojie Li 0001
Multim. Tools Appl.1
2020 Robust geodesic based outlier detection for class imbalance problem
Canghong Shi, Xiaojie Li 0001, Jiancheng Lv 0001, Jing Yin, Imran Mumtaz
Pattern Recognit. Lett.1
2020 Image segmentation of nasopharyngeal carcinoma using 3D CNN with long-range skip connection and multi-scale feature pyramid
Canghong Shi, Xiaojie Li 0001, Xi Wu 0004, Jiliu Zhou, Jiancheng Lv 0001
Soft Comput.2
2019 AMCNet: Attention-Based Multiscale Convolutional Network for DCM MRI Segmentation
abstract
For patients with dilated cardiomyopathy (DCM), fast and accurate diagnosis is important to save lives. MRI is a non-invasive, effective medical imaging method that allows doctors to diagnose DCM. However, manual and semi-automatic segmentation is subjective, non-reproducible and time-consuming task. In this paper, a new attention-based convolutional encoder-decoder network is proposed to automatically segment my-ocardium in DCM, which assisting the doctor to quickly diagnose. In the proposed method, the attention mechanism module is used, which is able to fully highlight useful features that facilitate segmentation while suppress useless features that are not conducive to segmentation. Combining with the multi-scale convolution, our encoder-decoder network can accurately segment the my-ocardium in DCM. We verified our approach on 1155 myocardial MRI. Our network achieves the most advanced segmentation performance on the cardiac DCM dataset. Experiment results demonstrate the effectiveness of the proposed method.
Canghong Shi, Xian Zhang 0008, Jing Peng 0003, Xiaojie Li 0001, Yucheng Chen 0003
COMPSAC (2)2
2019 TL-GAN: Generative Adversarial Networks with Transfer Learning for Mode Collapse (S)
abstract
Image generation based on the generative adversarial network (GAN) has been widely used in the field of computer vision.It helps generate images similar to the given data by learning their distribution.However, in many tasks, training on small datasets of scenes may lead to mode collapse, such that the generated images are often blurred and almost the same.To solve this problem, we propose a generative adversarial network with transfer learning for mode collapse called TL-GAN.Owing to the size of the training dataset, we introduce transfer learning (VGG pre-training network) to extract more useful features from the underlying pixels and add them to the discriminator, which can be used to calculate the distance between samples, and to provide the discriminator with a new training target.The discriminator thus learns the best features that can distinguish between real data and generated data using the proposed model.This also enhances the learning capability of the generator, which learn further about the distribution of real data.Meanwhile, generator can produce new images more realistic.The results of experiments show that the TL-GAN can guarantee the diversity of samples.A qualitative comparison with several prevalent methods confirmed its effectiveness.
Xianyu Wu, Shihao Feng, Xiaojie Li 0001, Jing Yin, Jiancheng Lv 0001, Canghong Shi
SEKE6
2018 Outlier Detection Based on the Data Structure
abstract
Outlier detection is one of the most frequently demanded task for optimizing results. Distance-based methods are a popular approach. They require no prior assumptions about the data generating distribution and are uncomplicated to implement. However, related methods have different parameters that are difficult to determine such that the identification results are generally unstable. Presenting related techniques without sacrificing stability is a challenging task. In this paper, we propose a new distance-based method that depends on the data structure to detect such points. In the proposed method, a global binary tree is constructed and the local distance score of a point is calculated to evaluate to what degree the observation is an outlier. The greater the value of the distance score, the more likely the point is an outlier point. Unlike typical distance-based methods, our algorithm has good scalability. Even when the dimension of the data points increases, the performance of our algorithm does not diminish. To reduce extra parameters, the top-p ranked points can be identified as outliers. Experimental results on synthetic and real-world datasets demonstrate the effectiveness and stability of our method.
Canghong Shi, Xiaojie Li 0001, Jia He 0003, Xi Wu 0004
IJCNN2
2018 Reference Sharing Mechanism-Based Self-Embedding Watermarking Scheme with Deterministic Content Reconstruction
abstract
This paper presents a reference sharing mechanism-based self-embedding watermarking scheme. The host image is embedded with watermark bits including the reference data for content recovery and the authentication data for tampering location. The special encoding matrix derived from the generator matrix of selected systematic Maximum Distance Separable (MDS) code is adopted. The reference data is generated by encoding all the representative data of the original image blocks. On the receiver side, the tampered image blocks can be located by the authentication data. The reference data embedded in one image block can be shared by all the image blocks to restore the tampered content. The tampering coincidence problem can be avoided at the extreme. The maximal tampering rate is deduced theoretically. Experimental results show that, as long as the tampering rate is less than the maximal tampering rate, the content recovery is deterministic. The quality of recovered content does not decrease with the maximal tampering rate.
Dongmei Niu, Hongxia Wang 0001, Minquan Cheng, Canghong Shi
Secur. Commun. Networks4
2016 Speech Authentication and Recovery Scheme in Encrypted Domain
Qing Qian 0001, Hongxia Wang 0001, Sani M. Abdullahi, Huan Wang 0010, Canghong Shi
IWDW5