Guiping Cao

dblp:250/6126 · DBLP profile ↗
← Back
7ranked-venue papers
5as first author
6since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Image recognition and object detection · 44% Deep learning architectures and training · 23% Efficient and distributed learning · 12%

Topics — the 16 heaviest of 17, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection › object detection
detection transformer
1.622025
DS-Det: Single-Query Paradigm and Attention Disentangled Learning for Flexible Object Detection · ACM Multimedia 2025
MLP-DINO: Category Modeling and Query Graphing with Deep MLP for Object Detection · IJCAI 2024
Computer vision › Image recognition and object detection
object detection
1.622025
DS-Det: Single-Query Paradigm and Attention Disentangled Learning for Flexible Object Detection · ACM Multimedia 2025
MLP-DINO: Category Modeling and Query Graphing with Deep MLP for Object Detection · IJCAI 2024
Machine learning › Deep learning architectures and training
attention mechanism
0.912025
DS-Det: Single-Query Paradigm and Attention Disentangled Learning for Flexible Object Detection · ACM Multimedia 2025
Machine learning › Trustworthy machine learning
calibration
0.912025
H-Calibration: Rethinking Classifier Recalibration With Probabilistic Error-Bounded Objective · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Computer vision › Vision and language
cross-modal alignment
0.912025
Open-Det: An Efficient Learning Framework for Open-Ended Detection · ICML 2025
Machine learning › Deep learning architectures and training › attention mechanism › transformer attention
disentangled attention
0.912025
DS-Det: Single-Query Paradigm and Attention Disentangled Learning for Flexible Object Detection · ACM Multimedia 2025
Machine learning › Efficient and distributed learning
efficient training
0.912025
Open-Det: An Efficient Learning Framework for Open-Ended Detection · ICML 2025
Computer vision › Image recognition and object detection › object detection › open-world object detection
open-set object detection
0.912025
Open-Det: An Efficient Learning Framework for Open-Ended Detection · ICML 2025
Machine learning › Trustworthy machine learning › calibration
post-hoc calibration
0.912025
H-Calibration: Rethinking Classifier Recalibration With Probabilistic Error-Bounded Objective · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Computer vision › Image recognition and object detection › object detection
small object detection
0.912025
Cross-DINO: Cross the Deep MLP and Transformer for Small Object Detection · IEEE Trans. Multim. 2025
Machine learning › Efficient and distributed learning › efficient training
training acceleration
0.912025
Open-Det: An Efficient Learning Framework for Open-Ended Detection · ICML 2025
Machine learning › Deep learning architectures and training › feedforward neural network
multilayer perceptron
0.812024
MLP-DINO: Category Modeling and Query Graphing with Deep MLP for Object Detection · IJCAI 2024
Computer vision › Image recognition and object detection
image classification
0.712023
Strip-MLP: Efficient Token Interaction for Vision MLP · ICCV 2023
Natural language and speech › Language models and text generation
large language model
0.312025
Open-Det: An Efficient Learning Framework for Open-Ended Detection · ICML 2025
Computer vision › Image recognition and object detection › object detection
query-based detection
0.312025
DS-Det: Single-Query Paradigm and Attention Disentangled Learning for Flexible Object Detection · ACM Multimedia 2025
Computer vision › Image recognition and object detection › object recognition › category recognition
object category modeling
0.212024
MLP-DINO: Category Modeling and Query Graphing with Deep MLP for Object Detection · IJCAI 2024

Methods — techniques the papers use, named apart from their topics

vision-language model · 0.9proper scoring rules · 0.9prompt distillation · 0.9probabilistic learning framework · 0.9pocoo loss · 0.9masked alignment loss · 0.9contrastive learning · 0.9boost loss · 0.9binning-based calibration · 0.9attention mechanism · 0.9
YearPublicationVenuePosition
2025 Open-Det: An Efficient Learning Framework for Open-Ended Detection
abstract
Open-Ended object Detection (OED) is a novel and challenging task that detects objects and generates their category names in a free-form manner, without requiring additional vocabularies during inference. However, the existing OED models, such as GenerateU, require large-scale datasets for training, suffer from slow convergence, and exhibit limited performance. To address these issues, we present a novel and efficient Open-Det framework, consisting of four collaborative parts. Specifically, Open-Det accelerates model training in both the bounding box and object name generation process by reconstructing the Object Detector and the Object Name Generator. To bridge the semantic gap between Vision and Language modalities, we propose a Vision-Language Aligner with V-to-L and L-to-V alignment mechanisms, incorporating with the Prompts Distiller to transfer knowledge from the VLM into VL-prompts, enabling accurate object name generation for the LLM. In addition, we design a Masked Alignment Loss to eliminate contradictory supervision and introduce a Joint Loss to enhance classification, resulting in more efficient training. Compared to GenerateU, Open-Det, using only 1.5% of the training data (0.077M vs. 5.077M), 20.8% of the training epochs (31 vs. 149), and fewer GPU resources (4 V100 vs. 16 A100), achieves even higher performance (+1.0% in APr). The source codes are available at: https://github.com/Med-Process/Open-Det.
Guiping Cao, Wenjian Huang 0001, Xiangyuan Lan, Jianguo Zhang 0001, Dongmei Jiang
ICML1
2025 DS-Det: Single-Query Paradigm and Attention Disentangled Learning for Flexible Object Detection
abstract
Popular transformer detectors have achieved promising performance through query-based learning using attention mechanisms. However, the roles of existing decoder query types (e.g., content query and positional query) are still underexplored. These queries are generally predefined with a fixed number (fixed-query), which limits their flexibility. We find that the learning of these fixed-query is impaired by Recurrent Opposing in Teractions (ROT) between two attention operations: Self-Attention (query-to-query) and Cross-Attention (query-to-encoder), thereby degrading decoder efficiency. Furthermore, "query ambiguity" arises when shared-weight decoder layers are processed with both one-to-one and one-to-many label assignments during training, violating DETR's one-to-one matching principle. To address these challenges, we propose DS-Det, a more efficient detector capable of detecting a flexible number of objects in images. Specifically, we reformulate and introduce a new unified Single-Query paradigm for decoder modeling, transforming the fixed-query into flexible. Furthermore, we propose a simplified decoder framework through attention disentangled learning: locating boxes with Cross-Attention (one-to-many process), deduplicating predictions with Self-Attention (one-to-one process), addressing ''query ambiguity'' and ''ROT'' issues directly, and enhancing decoder efficiency. We further introduce a unified PoCoo loss that leverages box size priors to prioritize query learning on hard samples such as small objects. Extensive experiments across five different backbone models on COCO2017 and WiderPerson datasets demonstrate the general effectiveness and superiority of DS-Det. The source codes are available at https://github.com/Med-Process/DS-Det/.
Guiping Cao, Xiangyuan Lan, Wenjian Huang 0001, Jianguo Zhang 0001, Dongmei Jiang, Yaowei Wang 0001
ACM Multimedia1
2025 H-Calibration: Rethinking Classifier Recalibration With Probabilistic Error-Bounded Objective
abstract
Deep neural networks have demonstrated remarkable performance across numerous learning tasks but often suffer from miscalibration, resulting in unreliable probability outputs. This has inspired many recent works on mitigating miscalibration, particularly through post-hoc recalibration methods that aim to obtain calibrated probabilities without sacrificing the classification performance of pre-trained models. In this study, we summarize and categorize previous works into three general strategies: intuitively designed methods, binning-based methods, and methods based on formulations of ideal calibration. Through theoretical and practical analysis, we highlight ten common limitations in previous approaches. To address these limitations, we propose a probabilistic learning framework for calibration called $h$h-calibration, which theoretically constructs an equivalent learning formulation for canonical calibration with boundedness. On this basis, we design a simple yet effective post-hoc calibration algorithm. Our method not only overcomes the ten identified limitations but also achieves markedly better performance than traditional methods, as validated by extensive experiments. We further analyze, both theoretically and experimentally, the relationship and advantages of our learning objective compared to traditional proper scoring rule. In summary, our probabilistic framework derives an approximately equivalent differentiable objective for learning error-bounded calibrated probabilities, elucidating the correspondence and convergence properties of computational statistics with respect to theoretical bounds in canonical calibration. The theoretical effectiveness is verified on standard post-hoc calibration benchmarks by achieving state-of-the-art performance. This research offers valuable reference for learning reliable likelihood in related fields.
Wenjian Huang 0001, Guiping Cao, Jiahao Xia 0001, Jingkun Chen, Hao Wang 0230, Jianguo Zhang 0001
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Cross-DINO: Cross the Deep MLP and Transformer for Small Object Detection
abstract
Small Object Detection (SOD) poses significant challenges due to limited information and the model's low class prediction score. While Transformer-based detectors have shown promising performance, their potential for SOD remains largely unexplored. In typical DETR-like frameworks, the CNN backbone network, specialized in aggregating local information, struggles to capture the necessary contextual information for SOD. The multiple attention layers in the Transformer Encoder face difficulties in effectively attending to small objects and can also lead to blurring of features. Furthermore, the model's lower class prediction score of small objects compared to large objects further increases the difficulty of SOD. To address these challenges, we introduce a novel approach calledCross-DINO. This approach incorporates the deep MLP network to aggregate initial feature representations with both short and long range information for SOD. Then, a new Cross Coding Twice Module (CCTM) is applied to integrate these initial representations to the Transformer Encoder feature, enhancing the details of small objects. Additionally, we introduce a new kind of soft label named Category-Size (CS), integrating the Category and Size of objects. By treating CS as new ground truth, we propose a new loss function called Boost Loss to improve the class prediction score of the model. Extensive experimental results on COCO, WiderPerson, VisDrone, AI-TOD, and SODA-D datasets demonstrate that Cross-DINO efficiently improves the performance of DETR-like models on SOD. Specifically, our model achieves36.4%AP$_{S}$on COCO for SOD with only 45M parameters, outperforming the DINO by+4.4%AP$_{S}$(36.4% vs. 32.0%) with fewer parameters and FLOPs, under 12 epochs training setting.
Guiping Cao, Wenjian Huang 0001, Xiangyuan Lan, Jianguo Zhang 0001, Dongmei Jiang, Yaowei Wang 0001
IEEE Trans. Multim.1
2024 MLP-DINO: Category Modeling and Query Graphing with Deep MLP for Object Detection
Guiping Cao, Wenjian Huang 0001, Xiangyuan Lan, Jianguo Zhang 0001, Dongmei Jiang, Yaowei Wang 0001
IJCAI1
2023 Strip-MLP: Efficient Token Interaction for Vision MLP
abstract
Token interaction operation is one of the core modules in MLP-based models to exchange and aggregate information between different spatial locations. However, the power of token interaction on the spatial dimension is highly dependent on the spatial resolution of the feature maps, which limits the model’s expressive ability, especially in deep layers where the feature are down-sampled to a small spatial size. To address this issue, we present a novel method called Strip-MLP to enrich the token interaction power in three ways. Firstly, we introduce a new MLP paradigm called Strip MLP layer that allows the token to interact with other tokens in a cross-strip manner, enabling the tokens in a row (or column) to contribute to the information aggregations in adjacent but different strips of rows (or columns). Secondly, a Cascade Group Strip Mixing Module (CGSMM) is proposed to overcome the performance degradation caused by small spatial feature size. The module allows tokens to interact more effectively in the manners of within-patch and cross-patch, which is independent to the feature spatial size. Finally, based on the Strip MLP layer, we propose a novel Local Strip Mixing Module (LSMM) to boost the token interaction power in the local region. Extensive experiments demonstrate that Strip-MLP significantly improves the performance of MLP-based models on small datasets and obtains comparable or even better results on ImageNet. In particular, Strip-MLP models achieve higher average Top-1 accuracy than existing MLP-based models by +2.44% on Caltech-101 and +2.16% on CIFAR-100. The source codes will be available at https://github.com/Med-Process/Strip_MLP.
Guiping Cao, Shengda Luo, Wenjian Huang 0001, Xiangyuan Lan, Dongmei Jiang, Yaowei Wang 0001, Jianguo Zhang 0001
ICCV1
2020 Deep Reinforcement Active Learning for Medical Image Classification
Yuguang Yan, Yubing Zhang, Guiping Cao, Ming Yang 0039, Michael Kwok-Po Ng
MICCAI (1)4