EDBT 2026 Demo / reviewers in the wild / expert
Guiping Cao
dblp:250/6126
· DBLP profile ↗
7ranked-venue papers
5as first author
6since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author · 4 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
6 papers |
Image recognition and object detection · 44% Deep learning architectures and training · 23% Efficient and distributed learning · 12% |
Topics — the 16 heaviest of 17, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Image recognition and object detection › object detection
detection transformer |
1.6 | 2 | 2025 | DS-Det: Single-Query Paradigm and Attention Disentangled Learning for Flexible Object Detection · ACM Multimedia 2025 MLP-DINO: Category Modeling and Query Graphing with Deep MLP for Object Detection · IJCAI 2024 |
Computer vision › Image recognition and object detection
object detection |
1.6 | 2 | 2025 | DS-Det: Single-Query Paradigm and Attention Disentangled Learning for Flexible Object Detection · ACM Multimedia 2025 MLP-DINO: Category Modeling and Query Graphing with Deep MLP for Object Detection · IJCAI 2024 |
Machine learning › Deep learning architectures and training
attention mechanism |
0.9 | 1 | 2025 | DS-Det: Single-Query Paradigm and Attention Disentangled Learning for Flexible Object Detection · ACM Multimedia 2025 |
Machine learning › Trustworthy machine learning
calibration |
0.9 | 1 | 2025 | H-Calibration: Rethinking Classifier Recalibration With Probabilistic Error-Bounded Objective · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Computer vision › Vision and language
cross-modal alignment |
0.9 | 1 | 2025 | Open-Det: An Efficient Learning Framework for Open-Ended Detection · ICML 2025 |
Machine learning › Deep learning architectures and training › attention mechanism › transformer attention
disentangled attention |
0.9 | 1 | 2025 | DS-Det: Single-Query Paradigm and Attention Disentangled Learning for Flexible Object Detection · ACM Multimedia 2025 |
Machine learning › Efficient and distributed learning
efficient training |
0.9 | 1 | 2025 | Open-Det: An Efficient Learning Framework for Open-Ended Detection · ICML 2025 |
Computer vision › Image recognition and object detection › object detection › open-world object detection
open-set object detection |
0.9 | 1 | 2025 | Open-Det: An Efficient Learning Framework for Open-Ended Detection · ICML 2025 |
Machine learning › Trustworthy machine learning › calibration
post-hoc calibration |
0.9 | 1 | 2025 | H-Calibration: Rethinking Classifier Recalibration With Probabilistic Error-Bounded Objective · IEEE Trans. Pattern Anal. Mach. Intell. 2025 |
Computer vision › Image recognition and object detection › object detection
small object detection |
0.9 | 1 | 2025 | Cross-DINO: Cross the Deep MLP and Transformer for Small Object Detection · IEEE Trans. Multim. 2025 |
Machine learning › Efficient and distributed learning › efficient training
training acceleration |
0.9 | 1 | 2025 | Open-Det: An Efficient Learning Framework for Open-Ended Detection · ICML 2025 |
Machine learning › Deep learning architectures and training › feedforward neural network
multilayer perceptron |
0.8 | 1 | 2024 | MLP-DINO: Category Modeling and Query Graphing with Deep MLP for Object Detection · IJCAI 2024 |
Computer vision › Image recognition and object detection
image classification |
0.7 | 1 | 2023 | Strip-MLP: Efficient Token Interaction for Vision MLP · ICCV 2023 |
Natural language and speech › Language models and text generation
large language model |
0.3 | 1 | 2025 | Open-Det: An Efficient Learning Framework for Open-Ended Detection · ICML 2025 |
Computer vision › Image recognition and object detection › object detection
query-based detection |
0.3 | 1 | 2025 | DS-Det: Single-Query Paradigm and Attention Disentangled Learning for Flexible Object Detection · ACM Multimedia 2025 |
Computer vision › Image recognition and object detection › object recognition › category recognition
object category modeling |
0.2 | 1 | 2024 | MLP-DINO: Category Modeling and Query Graphing with Deep MLP for Object Detection · IJCAI 2024 |
Methods — techniques the papers use, named apart from their topics
vision-language model · 0.9proper scoring rules · 0.9prompt distillation · 0.9probabilistic learning framework · 0.9pocoo loss · 0.9masked alignment loss · 0.9contrastive learning · 0.9boost loss · 0.9binning-based calibration · 0.9attention mechanism · 0.9
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Open-Det: An Efficient Learning Framework for Open-Ended DetectionabstractOpen-Ended object Detection (OED) is a novel and challenging task that detects objects and generates their category names in a free-form manner, without requiring additional vocabularies during inference. However, the existing OED models, such as GenerateU, require large-scale datasets for training, suffer from slow convergence, and exhibit limited performance. To address these issues, we present a novel and efficient Open-Det framework, consisting of four collaborative parts. Specifically, Open-Det accelerates model training in both the bounding box and object name generation process by reconstructing the Object Detector and the Object Name Generator. To bridge the semantic gap between Vision and Language modalities, we propose a Vision-Language Aligner with V-to-L and L-to-V alignment mechanisms, incorporating with the Prompts Distiller to transfer knowledge from the VLM into VL-prompts, enabling accurate object name generation for the LLM. In addition, we design a Masked Alignment Loss to eliminate contradictory supervision and introduce a Joint Loss to enhance classification, resulting in more efficient training. Compared to GenerateU, Open-Det, using only 1.5% of the training data (0.077M vs. 5.077M), 20.8% of the training epochs (31 vs. 149), and fewer GPU resources (4 V100 vs. 16 A100), achieves even higher performance (+1.0% in APr). The source codes are available at: https://github.com/Med-Process/Open-Det. Guiping Cao, Wenjian Huang 0001, Xiangyuan Lan, Jianguo Zhang 0001, Dongmei Jiang |
ICML | 1 |
| 2025 | DS-Det: Single-Query Paradigm and Attention Disentangled Learning for Flexible Object DetectionabstractPopular transformer detectors have achieved promising performance through query-based learning using attention mechanisms. However, the roles of existing decoder query types (e.g., content query and positional query) are still underexplored. These queries are generally predefined with a fixed number (fixed-query), which limits their flexibility. We find that the learning of these fixed-query is impaired by Recurrent Opposing in Teractions (ROT) between two attention operations: Self-Attention (query-to-query) and Cross-Attention (query-to-encoder), thereby degrading decoder efficiency. Furthermore, "query ambiguity" arises when shared-weight decoder layers are processed with both one-to-one and one-to-many label assignments during training, violating DETR's one-to-one matching principle. To address these challenges, we propose DS-Det, a more efficient detector capable of detecting a flexible number of objects in images. Specifically, we reformulate and introduce a new unified Single-Query paradigm for decoder modeling, transforming the fixed-query into flexible. Furthermore, we propose a simplified decoder framework through attention disentangled learning: locating boxes with Cross-Attention (one-to-many process), deduplicating predictions with Self-Attention (one-to-one process), addressing ''query ambiguity'' and ''ROT'' issues directly, and enhancing decoder efficiency. We further introduce a unified PoCoo loss that leverages box size priors to prioritize query learning on hard samples such as small objects. Extensive experiments across five different backbone models on COCO2017 and WiderPerson datasets demonstrate the general effectiveness and superiority of DS-Det. The source codes are available at https://github.com/Med-Process/DS-Det/. Guiping Cao, Xiangyuan Lan, Wenjian Huang 0001, Jianguo Zhang 0001, Dongmei Jiang, Yaowei Wang 0001 |
ACM Multimedia | 1 |
| 2025 | H-Calibration: Rethinking Classifier Recalibration With Probabilistic Error-Bounded ObjectiveabstractDeep neural networks have demonstrated remarkable performance across numerous learning tasks but often suffer from miscalibration, resulting in unreliable probability outputs. This has inspired many recent works on mitigating miscalibration, particularly through post-hoc recalibration methods that aim to obtain calibrated probabilities without sacrificing the classification performance of pre-trained models. In this study, we summarize and categorize previous works into three general strategies: intuitively designed methods, binning-based methods, and methods based on formulations of ideal calibration. Through theoretical and practical analysis, we highlight ten common limitations in previous approaches. To address these limitations, we propose a probabilistic learning framework for calibration called $h$h-calibration, which theoretically constructs an equivalent learning formulation for canonical calibration with boundedness. On this basis, we design a simple yet effective post-hoc calibration algorithm. Our method not only overcomes the ten identified limitations but also achieves markedly better performance than traditional methods, as validated by extensive experiments. We further analyze, both theoretically and experimentally, the relationship and advantages of our learning objective compared to traditional proper scoring rule. In summary, our probabilistic framework derives an approximately equivalent differentiable objective for learning error-bounded calibrated probabilities, elucidating the correspondence and convergence properties of computational statistics with respect to theoretical bounds in canonical calibration. The theoretical effectiveness is verified on standard post-hoc calibration benchmarks by achieving state-of-the-art performance. This research offers valuable reference for learning reliable likelihood in related fields. Wenjian Huang 0001, Guiping Cao, Jiahao Xia 0001, Jingkun Chen, Hao Wang 0230, Jianguo Zhang 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2025 | Cross-DINO: Cross the Deep MLP and Transformer for Small Object DetectionabstractSmall Object Detection (SOD) poses significant challenges due to limited information and the model's low class prediction score. While Transformer-based detectors have shown promising performance, their potential for SOD remains largely unexplored. In typical DETR-like frameworks, the CNN backbone network, specialized in aggregating local information, struggles to capture the necessary contextual information for SOD. The multiple attention layers in the Transformer Encoder face difficulties in effectively attending to small objects and can also lead to blurring of features. Furthermore, the model's lower class prediction score of small objects compared to large objects further increases the difficulty of SOD. To address these challenges, we introduce a novel approach calledCross-DINO. This approach incorporates the deep MLP network to aggregate initial feature representations with both short and long range information for SOD. Then, a new Cross Coding Twice Module (CCTM) is applied to integrate these initial representations to the Transformer Encoder feature, enhancing the details of small objects. Additionally, we introduce a new kind of soft label named Category-Size (CS), integrating the Category and Size of objects. By treating CS as new ground truth, we propose a new loss function called Boost Loss to improve the class prediction score of the model. Extensive experimental results on COCO, WiderPerson, VisDrone, AI-TOD, and SODA-D datasets demonstrate that Cross-DINO efficiently improves the performance of DETR-like models on SOD. Specifically, our model achieves36.4%AP$_{S}$on COCO for SOD with only 45M parameters, outperforming the DINO by+4.4%AP$_{S}$(36.4% vs. 32.0%) with fewer parameters and FLOPs, under 12 epochs training setting. Guiping Cao, Wenjian Huang 0001, Xiangyuan Lan, Jianguo Zhang 0001, Dongmei Jiang, Yaowei Wang 0001 |
IEEE Trans. Multim. | 1 |
| 2024 | MLP-DINO: Category Modeling and Query Graphing with Deep MLP for Object Detection
Guiping Cao, Wenjian Huang 0001, Xiangyuan Lan, Jianguo Zhang 0001, Dongmei Jiang, Yaowei Wang 0001 |
IJCAI | 1 |
| 2023 | Strip-MLP: Efficient Token Interaction for Vision MLPabstractToken interaction operation is one of the core modules in MLP-based models to exchange and aggregate information between different spatial locations. However, the power of token interaction on the spatial dimension is highly dependent on the spatial resolution of the feature maps, which limits the model’s expressive ability, especially in deep layers where the feature are down-sampled to a small spatial size. To address this issue, we present a novel method called Strip-MLP to enrich the token interaction power in three ways. Firstly, we introduce a new MLP paradigm called Strip MLP layer that allows the token to interact with other tokens in a cross-strip manner, enabling the tokens in a row (or column) to contribute to the information aggregations in adjacent but different strips of rows (or columns). Secondly, a Cascade Group Strip Mixing Module (CGSMM) is proposed to overcome the performance degradation caused by small spatial feature size. The module allows tokens to interact more effectively in the manners of within-patch and cross-patch, which is independent to the feature spatial size. Finally, based on the Strip MLP layer, we propose a novel Local Strip Mixing Module (LSMM) to boost the token interaction power in the local region. Extensive experiments demonstrate that Strip-MLP significantly improves the performance of MLP-based models on small datasets and obtains comparable or even better results on ImageNet. In particular, Strip-MLP models achieve higher average Top-1 accuracy than existing MLP-based models by +2.44% on Caltech-101 and +2.16% on CIFAR-100. The source codes will be available at https://github.com/Med-Process/Strip_MLP. Guiping Cao, Shengda Luo, Wenjian Huang 0001, Xiangyuan Lan, Dongmei Jiang, Yaowei Wang 0001, Jianguo Zhang 0001 |
ICCV | 1 |
| 2020 | Deep Reinforcement Active Learning for Medical Image Classification
Yuguang Yan, Yubing Zhang, Guiping Cao, Ming Yang 0039, Michael Kwok-Po Ng |
MICCAI (1) | 4 |