Ge Cao

dblp:116/7297 · DBLP profile ↗
← Back
13ranked-venue papers
5as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 3 first-author · 3 since 2021Human-computer interaction and ubiquitous computing · 3 · 1 first-author · 3 since 2021Artificial intelligence and machine learning · 1Systems, architecture and hardware · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 Online detection of hardware Trojan enabled packet tampering attack on network-on-chip: A Bayesian approach
Xiaohang Wang 0001, Ge Cao, Yingtao Jiang, Amit Kumar Singh 0002, Mei Yang 0001, Liang Wang 0020, Fen Guo
Integr.2
2026 Optimal Proxy Mining Contrastive Network for Unsupervised Person Re-Identification
abstract
Unsupervised person re-identification (Re-ID) performance enhancement hinges on extracting the most informative features from unlabeled person datasets. In recent approaches, proxy-based contrastive learning with awareness of camera labels has been adopted for model training, thereby achieving highly promising results. However, inappropriate selections of contrastive pairs can significantly degrade the performance of these models. To address this issue, we propose the Optimal Proxy Mining Contrastive Network (OPMCN), a novel framework designed to strategically optimize the selection of proxies for positive and negative pair formation, thus enhancing the efficacy of contrastive training. The OPMCN framework proposes two specific contrastive losses: Hardest Camera Proxy Mining (HCPM) and False Negative Proxies Mining (FNPM), each essential for enhancing model performance in unsupervised settings. The HCPM loss targets proxies from the most challenging cameras to maximize semantic differences between pairs while ensuring minimal background shifts. In contrast, the FNPM loss counters noise in pseudo labels by prioritizing similarity rankings over clustering results to effectively identify and correct false negatives among proxies. Moreover, we have developed the Pyramid Kernel Global Context (PKGC) block, which employs an attention mechanism that focuses on identity-invariant semantic cues in instances. This module utilizes optimally sized convolutional kernels to enhance identity recognition consistency across camera-based variations, thereby improving the precision of feature extraction. Experimental results on several popular datasets prove that our work surpasses existing unsupervised person Re-ID approaches to a remarkable extent.
Ge Cao, Qing Tang 0004, Xuan-Thuy Vo, Adri Priadana, Kang-Hyun Jo
IEEE Trans. Circuits Syst. Video Technol.1
2025 Large Vision-Language Models with PEFT for Generating Descriptive Annotations in Person Re-Identification
abstract
Generating descriptive annotations for person re-identification (Re-ID) images is essential for bridging vision and language domain, improving both interpretability and cross-modal retrieval performance. However, large vision-language models (LVLMs), which trained on broad web-scale corpora, often struggle to generate accurate, context-relevant descriptions for ReID samples due to inherent challenges such as occlusions, low resolution, varying illumination, and diverse viewpoints. In this paper, we propose to apply Parameter-Efficient Fine-Tuning (PEFT) via Low-Rank Adaptation (LoRA) to tune Qwen2-VL for ReID-specific captioning tasks. Leveraging existing Re-ID datasets with paired image-text annotations, our fine-tuned model generates domain-aligned and discriminative captions. Experiments show significant improvements in caption relevance and identity descriptiveness, highlighting the potential of PEFT-tuned LVLMs for real-world ReID applications.
Ge Cao, Qing Tang 0004, Adri Priadana, Tran Tien Dat, Ashraf Uddin Russo, Kang-Hyun Jo
HSI1
2025 Efficiency-Accuracy Trade-Off of Facial Attribute Classifier Supporting Human-Robot Interaction
abstract
The advancement of robotics has been driven by the integration of artificial intelligence, machine learning, and sophisticated sensing technologies, enabling more seamless Human-Robot Interaction (HRI). Facial Attribute Classifier (FAC) plays a crucial role in HRI by helping robots understand human emotions, intentions, and social cues, fostering personalized and intuitive interactions. However, while existing methods achieve high accuracy, their computational complexity limits real-time applications on low-cost or CPU-based devices, highlighting the need for lightweight models that balance accuracy and efficiency. This work proposes an Efficient Network (ENet) designed to achieve an optimal trade-off between efficiency and accuracy of FAC. ENet introduces an Enhanced Sequential Efficient Attention Module (ESEAM) to improve the quality of feature maps while maintaining high efficiency. Accordingly, ENet demonstrates a compromise between efficiency and accuracy on the CelebA and LFWA datasets. The proposed ENet is computationally efficient, generating a few parameters, making it well-suited for CPU-based applications. When combined with a face detector, the optimized FAC achieves a processing speed of 25.88 frames per second (FPS) on an Intel Core i7-9750H CPU, demonstrating its suitability for real-time use.
Adri Priadana, Duy-Linh Nguyen, Xuan-Thuy Vo, Ge Cao, Jehwan Choi, Kang-Hyun Jo
HSI4
2025 A High-Accuracy and Faster Face Recognizer Supporting Biometric Continuous Authentication for Smart Factory Workers
abstract
Smart factories require secure and sustainable worker authentication for safe operations. Biometric continuous authentication based on facial recognition is one of the most convenient mechanisms. This method applies a face recognition task to verify the captured face as an authorized user. However, existing methods that employ large networks for high-accuracy face recognition incur high computational costs and slow down the process, rendering them unsuitable for continuous operation. This work proposes an efficient and rapid face recognizer with high accuracy. It offers a faster face residual network, containing efficient FasterFace blocks and efficient channel spatial attention for improved feature extraction. As a result, the proposed network achieves 97.08% based on average accuracy, outperforming the other networks on five benchmark datasets. It performs faster at 19.91 frames per second in real time on CPU-based hardware when integrated with a face detector, showcasing its capacity to support real-time biometric continuous authentication for smart factory workers.
Adri Priadana, Duy-Linh Nguyen, Xuan-Thuy Vo, Muhamad Dwisnanto Putro, Ge Cao, Kang-Hyun Jo
IEEE Trans. Ind. Informatics5
2025 Local Self-Attention With Mixing Abstract Tokens for Urban Autonomous Driving
abstract
Although local self-attentions exhibit translation equivariance and locality similar to convolution, the model has limited receptive fields and weak modeling ability. The main reason is that self-attention is computed within nonoverlapped windows. To overcome this issue, common methods need further operations to communicate the information across windows, such as window shifting, and sliding. These operations are memory unfriendly, not well supported, and optimized by modern deep-learning frameworks. Alternatively, this article exchanges information across nonoverlapped windows via efficiently mixing abstract tokens (MAT). The MAT block includes the following steps. First, the image tokens are partitioned into windows and each window is merged with an abstract token. Second, in each window, interactions of image tokens and the abstract token to image tokens are performed. Third, because the abstract token learns abstract information from each corresponding window, mixing all abstract tokens via transformer encoder helps to exchange information between local windows and result in global context modeling. Fourth, the global information of the mixed tokens is propagated back to the image tokens through transformer decoder. The MAT block is efficient and easy to implement, only containing matrix multiplications. In addition, this article also proposes a bilinear patch embedding that samples relevant regions of the input tokens based on learned offsets. Extensive experiments are conducted and evaluated with various tasks such as image classification, object detection, and segmentation. As a result, our method achieves promising performances across tasks. For example, MAT-2 accomplishes79.0%top-1 accuracy on ImageNet-1 K with0.7GFLOPs and outperforms the baseline Swin-0.7 G by4.6%while reducing15.2 mson CPU and0.53 mson GPU devices. The MAT-4 surpasses Swin-T by1.8%mIoU with only70%GFLOPs.
Xuan-Thuy Vo, Duy-Linh Nguyen, Adri Priadana, Ge Cao, Jehwan Choi, Kang-Hyun Jo
IEEE Trans. Ind. Informatics4
2024 Inverted Residual Bottlenecks with Large Kernel Attention for Remote Scene Classification
abstract
Remote sensing image classification plays a piv-otal role in environmental monitoring and urban planning, yet it faces the challenge of accurately interpreting complex and high-resolution images with fast inference speed for real time applications. To address this, we introduce the Mobile Large Kernel Attention Network (MLKANet), which integrates MobileNetV2's inverted residual structures with the large kernel attention mechanism from the Visual Attention Network (VAN). Our proposed MLKANet achieves a compelling balance of computational efficiency and sophisticated feature extraction, while maintaining the speed from the MobileNetV2 baseline. This study evaluates MLKANet's performance against state-of-the-art models using the Aerial Image Dataset (AID), demonstrating superior accuracy and efficiency. The architecture's effectiveness is further evidenced through an ablation study highlighting the scalability of our approach and class-wise performance analysis that showcases MLKANet's proficiency across various scene types.
RussoMohammadAshraf Uddin, Adri Priadana, Ge Cao, Kang-Hyun Jo
HSI3
2024 Enhancing Unsupervised Domain Adaptive Person Re-identification Clustering Through Parsing-Based Attribute Labeling
Ge Cao, Kang-Hyun Jo
ICIC (4)1
2020 Aggregated Deep Saliency Prediction by Self-attention Network
Ge Cao, Qing Tang 0004, Kang-Hyun Jo
ICIC (3)1
2020 Accurate and Efficient Traffic Sign Detection with a Guided Region Enlarging Algorithm
Qing Tang 0004, Ge Cao, Kang-Hyun Jo
ICIC (1)2
2020 Operational Risk Evaluation of Active Distribution Networks Considering Cyber Contingencies
abstract
A deeper coupling of information and power would make active distribution networks (ADNs) evolve into typical cyber-physical systems (CPSs). The unreliability of cyber systems can affect the security of ADNs. In this article, a methodology based on cyber-power joint analysis is presented to quantitatively evaluate the operational risk of ADNs caused by cyber contingencies. First, a typical CPS structure of ADNs is demonstrated and the influence ways of cyber contingencies are discussed. Then, a CPS modeling framework is proposed to analyze the information flow and cyber contingencies of cyber systems. To determine the best data transmission paths in cyber networks, a routing algorithm based on the link-state protocol is presented. Furthermore, in view of the response of ADNs to physical failures, an optimization model is developed to minimize the load shedding that results from the over-limited voltage. Finally, the expected load curtailment and the expected fault coverage used to quantify potential system losses are calculated using Monte Carlo simulation. The simulation results explain the effectiveness and practicability of the proposed models and methods.
Ge Cao, Wei Gu 0004, Peixin Li, Wanxing Sheng, Keyan Liu, Lijing Sun, Zhihuang Cao
IEEE Trans. Ind. Informatics1
2014 Multi-source transfer ELM-based Q learning
Xuesong Wang 0001, Yuhu Cheng 0001, Ge Cao
Neurocomputing4
2012 Reinforcement Learning Based on Extreme Learning Machine
Xuesong Wang 0001, Yuhu Cheng 0001, Ge Cao
ICIC (3)4