Lihong Cao

dblp:119/9232 · DBLP profile ↗
← Back
19ranked-venue papers
3as first author
15since 2021 · last 2026
0000-0002-3866-2734ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 1 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 6 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Computer networks · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 A method for extracting emotion-cause pairs based on bidirectional machine reading comprehension
Guorui Li, Yaxin Wen, Cong Wang 0009, Lihong Cao, Sancheng Peng
Eng. Appl. Artif. Intell.4
2026 Generalized Transferable Attack Across Datasets
abstract
Existing transferable attack methods commonly assume that the attacker knows the training set (e.g., the label set, the input size) of the black-box victim models, which is usually unrealistic because in some cases the attacker cannot know this information. In this paper, we define a Generalized Transferable Attack (GTA) problem where the attacker operates without prior knowledge of these specifics and must attack randomly encountered images, potentially from unknown datasets. To solve the challenging GTA problem, we propose a novel Image Classification Disruptor (ICD), designed to train a particular attack to disrupt classification information of any images from arbitrary datasets. Experiments across several datasets demonstrate that ICD clearly outperforms existing transferable attacks on GTA, and show that ICD uses similar texture-like noises to perturb different images from different datasets. Moreover, we observed that ICD noise across images mainly consists of three specific-frequency sine waves for the R, G, and B channels. Inspired by this interesting finding, we also design another novel Sine Attack (SA) method directly optimizes the three sine waves. Experiments show that SA performs comparably to ICD, revealing a notable vulnerability in CNNs under the GTA setting.
Yunxiao Qin, Yuanhao Xiong, Jinfeng Yi, Lihong Cao, Cho-Jui Hsieh
IEEE Trans. Circuits Syst. Video Technol.4
2025 Enabling scale and rotation invariance in convolutional neural networks with retina like transformation
Jiahong Zhang, Guoqi Li 0002, Qiaoyi Su, Lihong Cao, Yonghong Tian 0001, Bo Xu 0002
Neural Networks4
2024 A Brain-inspired Method for Occluded 3D Object Recognition
abstract
Recently, 3D object recognition has been widely applied in various practical scenes and several methods were proposed for this task. Though these methods have made great progresses in 3D object recognition, they perform much poorly for occluded 3D objects. As a comparison, humans recognize 3D objects well, no matter the object is occluded or non-occluded. Some brain science researches demonstrated that humans recognize 3D objects by the process from sensory input to proto-objects representation, object-file representation, and finally, recognition. We therefore propose a novel brain-inspired 3D multi-view object recognition network (BiMVNet) to realize this process by mimicing the human visual system with visual working memory and knowledge. Specifically, BiMVNet imitates the human visual cortex to extract the proto-objects representation, the human visual working memory and knowledge to extract the object-file representation, and finally recognizes 3D objects based on the object-file representation. Furthermore, we propose an occluded 3D object dataset Occ-ModelNet40 based on the popular ModelNet40 dataset to thoroughly evaluate BiMVNet. Experiments on Occ-ModelNet40 and ModelNet40 show that our BiMVNet outperforms existing methods on both occluded and non-occluded object recognition.
Zining Wan, Jiahong Zhang, Lihong Cao
IJCNN3
2024 Prompt-based Continual Learning for Extending Pretrained CLIP Models' Knowledge
abstract
CLIP model has demonstrated remarkable performance and strong zero-shot capabilities through its training on text-image datasets using contrastive learning. This has sparked interest in developing continuous learning methods based on CLIP to extend its knowledge to new datasets. However, traditional continuous learning approaches often involve modification to the original parameters of the pretrained CLIP model and consequently compromise its zero-shot capabilities. To tackle these challenges, we propose Image Text (IT-)Prompt, which leverages the inherent correlation between visual and textual information to train discrete prompts dedicated to individual tasks, serving as repositories for task-specific knowledge. By employing discrete textual prompts as guidance, we ensure the uniqueness of each task's prompt and prevent interference among tasks, thus alleviating catastrophic forgetting during continuous learning. While retaining the pretrained parameters of CLIP, our approach introduces only a small number of additional trainable parameters. This allows us to enhance training efficiency and preserving the original zero-shot capabilities of CLIP. Comparative experiments show that IT-Prompt achieves a performance improvement of at least 10% compared to state-of-the-art methods. The implementation code can be available at https://github.com/jiaolifengmi/IT-Prompt. © 2024 Copyright held by the owner/author(s). Publication rights licensed to ACM.
Lihong Cao
MMAsia2
2024 Textual emotion classification using MPNet and cascading broad learning
Lihong Cao, Sancheng Peng, Aimin Yang 0002, Jianwei Niu 0002, Shui Yu 0001
Neural Networks1
2024 Self-Supervised Video Representation Learning via Capturing Semantic Changes Indicated by Saccades
abstract
In this paper, we propose a self-supervised video representation learning (video SSL) method by taking inspiration from cognitive science and neuroscience on human visual perception. Different from previous methods that focus on the inherent properties of videos, we argue that humans learn to perceive the world through the self-awareness of the semantic changes or consistency in the input stimuli in the absence of labels, accompanied by representation reorganization during the post-learning rest periods. To this end, we first exploit the presence of saccades as an indicator of semantic changes in a contrastive learning framework, mimicking self-awareness in human representation learning. The saccades are generated by alternating the fixations following the predicted scanpath. Second, we model the semantic consistency in eye fixation by minimizing the prediction error between the predicted and the true state of another time point. Finally, we incorporate prototypical contrastive learning to reorganize the learned representations to enhance the associations among perceptually similar ones. Compared to previous video SSL solutions, our method can capture finer-grained semantics from video instances and further associate similar ones together. Experiments show that the proposed bio-inspired video SSL method significantly improves the Top-1 video retrieval accuracy on UCF101 and achieves superior performance on downstream tasks such as action recognition under comparable settings.
Qiuxia Lai, Ailing Zeng, Ye Wang 0011, Lihong Cao, Yu Li 0007, Qiang Xu 0001
IEEE Trans. Circuits Syst. Video Technol.4
2023 BIFRNet: A Brain-Inspired Feature Restoration DNN for Partially Occluded Image Recognition
abstract
The partially occluded image recognition (POIR) problem has been a challenge for artificial intelligence for a long time. A common strategy to handle the POIR problem is using the non-occluded features for classification. Unfortunately, this strategy will lose effectiveness when the image is severely occluded, since the visible parts can only provide limited information. Several studies in neuroscience reveal that feature restoration which fills in the occluded information and is called amodal completion is essential for human brains to recognize partially occluded images. However, feature restoration is commonly ignored by CNNs, which may be the reason why CNNs are ineffective for the POIR problem. Inspired by this, we propose a novel brain-inspired feature restoration network (BIFRNet) to solve the POIR problem. It mimics a ventral visual pathway to extract image features and a dorsal visual pathway to distinguish occluded and visible image regions. In addition, it also uses a knowledge module to store classification prior knowledge and uses a completion module to restore occluded features based on visible features and prior knowledge. Thorough experiments on synthetic and real-world occluded image datasets show that BIFRNet outperforms the existing methods in solving the POIR problem. Especially for severely occluded images, BIRFRNet surpasses other methods by a large margin and is close to the human brain performance. Furthermore, the brain-inspired design makes BIFRNet more interpretable.
Jiahong Zhang, Lihong Cao, Qiuxia Lai, Yunxiao Qin
AAAI2
2023 Extracting the Brain-Like Representation by an Improved Self-Organizing Map for Image Classification
abstract
Backpropagation-based supervised learning has achieved great success in computer vision tasks. However, its biological plausibility is always controversial. Recently, the bioinspired Hebbian learning rule (HLR) has received extensive attention. Self-Organizing Map (SOM) uses the competitive HLR to establish connections between neurons, obtaining visual features in an unsupervised way. Although the representation of SOM neurons shows some brain-like characteristics, it is still quite different from the neuron representation in the human visual cortex. This paper proposes an improved SOM with multi-winner, multi-code, and local receptive field, named mlSOM. We observe that the neuron representation of mlSOM is similar to the human visual cortex. Furthermore, mlSOM shows a sparse distributed representation of objects, which has also been found in the human inferior temporal area. In addition, experiments show that mlSOM achieves better classification accuracy than the original SOM and other state-of-the-art HLR-based methods. The code is accessible at https://github.com/JiaHongZ/mlSOM.
Jiahong Zhang, Lihong Cao, Moning Zhang
ICASSP2
2023 Multi-source domain adaptation method for textual emotion classification using deep and broad learning
Sancheng Peng, Lihong Cao, Jianwei Niu 0002, Chengqing Zong, Guodong Zhou 0001
Knowl. Based Syst.3
2022 A Multi-Head Convolutional Neural Network with Multi-Path Attention Improves Image Denoising
Jiahong Zhang, Meijun Qu, Ye Wang 0011, Lihong Cao
PRICAI (3)4
2022 NHNet: A non-local hierarchical network for image denoising
abstract
Abstract With the fast development of deep learning models, hierarchical convolutional neural networks have achieved great success in image denoising tasks. To further boost the performance of image denoising, a novel non‐local hierarchical network (NHNet) is proposed. Unlike existing U‐Net‐based hierarchical methods, which mainly focus on downsampling operations, NHNet adopts an initial resolution path and a high resolution path. Specifically, the high‐resolution features are obtained through upsampling, where the non‐local mechanism is adopted to capture the self‐similarity properties, which contribute to a better denoising performance. Cross connections and channel attention layers are added between the two paths to integrate features in different resolutions. Compared with other U‐Net‐based hierarchical networks, NHNet requires fewer parameters. Experiments show that NHNet achieves state‐of‐the‐art performance in Gaussian denoising tasks and gets competitive results when dealing with real image denoising.
Jiahong Zhang, Lihong Cao, Weiheng Shen
IET Image Process.2
2021 Learning How to Zoom In: Weakly Supervised ROI-Based-DAM for Fine-Grained Visual Classification
Shuang Ran, Lihong Cao
ICANN (2)4
2021 A Biologically Plausible Audio-Visual Integration Model for Continual Learning
abstract
The problem of catastrophic forgetting has a history of more than 30 years and has not been completely solved yet. Since the human brain has natural ability to perform continual lifelong learning, learning from the brain may provide solutions to this problem. In this paper, we propose a novel biologically plausible audio-visual integration model (AVIM) based on the assumption that the integration of audio and visual perceptual information in the medial temporal lobe during learning is crucial to form concepts and make continual learning possible. Specifically, we use multi-compartment Hodgkin-Huxley neurons to build the model and adopt the calcium-based synaptic tagging and capture as the model's learning rule. Furthermore, we define a new continual learning paradigm to simulate the possible continual learning process in the human brain. We then test our model under this new paradigm. Our experimental results show that the proposed AVIM can achieve state-of-the-art continual learning performance compared with other advanced methods such as OWM, iCaRL and GEM. Moreover, it can generate stable representations of objects during learning. These results support our assumption that concept formation is essential for continuous lifelong learning and suggest the proposed AVIM is a possible concept formation mechanism.
Fengtong Du, Ye Wang 0011, Lihong Cao
IJCNN4
2021 Aspect-level sentiment analysis using context and aspect memory network
Yanxia Lv, Fangna Wei, Lihong Cao, Sancheng Peng, Jianwei Niu 0002, Shui Yu 0001, Cuirong Wang
Neurocomputing3
2018 Predicting spikes with artificial neural network
Lihong Cao, Jiamin Shen, Ye Wang 0011
Sci. China Inf. Sci.1
2018 Influence analysis in social networks: A survey
Sancheng Peng, Yongmei Zhou, Lihong Cao, Shui Yu 0001, Jianwei Niu 0002, Weijia Jia 0001
J. Netw. Comput. Appl.3
2017 Social influence modeling using information theory in mobile social networks
Sancheng Peng, Aimin Yang 0002, Lihong Cao, Shui Yu 0001, Dongqing Xie
Inf. Sci.3
2012 An Auction Approach to Resource Allocation in OFDM-Based Cognitive Radio Networks
abstract
We study a repeated auction for the resource allocation problem in OFDM-based cognitive radio networks (CRNs), in which secondary users (SUs) share the primary spectrum under the interference constraints of primary users (PUs). With the inter-cell interference and mutual interference between PUs and SUs, the resource allocation problem is formulated as a non-convex optimization problem. Auction performs well in solving non-convex problems, therefore the interference auction with cooperative bidding is proposed. Moreover, with the theoretical analysis of equilibrium, an implementation algorithm for the auction is developed and the convergence is proved. Simulation results show that the interference auction obtains a good spectrum efficiency improvement and a rapid convergence rate.
Lihong Cao, Wenjun Xu 0001, Jiaru Lin, Kai Niu 0001, Zhiqiang He 0001
VTC Spring1