Jiaming Li 0010

dblp:33/5934-10 · DBLP profile ↗
← Back
9ranked-venue papers
5as first author
8since 2021 · last 2026
0000-0002-1717-9209ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 4 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Image recognition and object detection · 58% Video understanding and tracking · 12% 3D vision · 11%

Topics — the 14 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Computer vision › Image recognition and object detection
object detection
3.342026
Toward Efficient Semi-Supervised Object Detection With Detection Transformer · IEEE Trans. Pattern Anal. Mach. Intell. 2026
OffsetNet: Towards Efficient Multiple Object Tracking, Detection, and Segmentation · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Learning Background Prompts to Discover Implicit Knowledge for Open Vocabulary Object Detection · CVPR 2024
Computer vision › Image recognition and object detection › object detection
semi-supervised object detection
1.722026
Toward Efficient Semi-Supervised Object Detection With Detection Transformer · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Gradient-based Sampling for Class Imbalanced Semi-supervised Object Detection · ICCV 2023
Computer vision › Image recognition and object detection › object detection
open-vocabulary object detection
1.622025
GUIDED: Granular Understanding via Identification, Detection, and Discrimination for Fine-Grained Open-Vocabulary Object Detection · NeurIPS 2025
Learning Background Prompts to Discover Implicit Knowledge for Open Vocabulary Object Detection · CVPR 2024
Computer vision › Image recognition and object detection › object detection
detection transformer
1.012026
Toward Efficient Semi-Supervised Object Detection With Detection Transformer · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Computer vision › Video understanding and tracking
multi-object tracking
0.912025
OffsetNet: Towards Efficient Multiple Object Tracking, Detection, and Segmentation · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Computer vision › Video understanding and tracking › multi-object tracking
multi-object tracking and segmentation
0.912025
OffsetNet: Towards Efficient Multiple Object Tracking, Detection, and Segmentation · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Computer vision › 3D vision
3d object detection
0.812024
Decoupled Pseudo-Labeling for Semi-Supervised Monocular 3D Object Detection · CVPR 2024
Computer vision › Image recognition and object detection › object detection
knowledge distillation for detection
0.812024
Learning Background Prompts to Discover Implicit Knowledge for Open Vocabulary Object Detection · CVPR 2024
Computer vision › 3D vision › 3d object detection › image-based 3d object detection
monocular 3d object detection
0.812024
Decoupled Pseudo-Labeling for Semi-Supervised Monocular 3D Object Detection · CVPR 2024
Machine learning › Learning paradigms › semi-supervised learning
pseudo-labeling
0.812024
Decoupled Pseudo-Labeling for Semi-Supervised Monocular 3D Object Detection · CVPR 2024
Computer vision › Vision and language
vision-language model
0.812024
Learning Background Prompts to Discover Implicit Knowledge for Open Vocabulary Object Detection · CVPR 2024
Machine learning › Learning paradigms
class imbalance
0.712023
Gradient-based Sampling for Class Imbalanced Semi-supervised Object Detection · ICCV 2023
Computer vision › Segmentation and scene understanding
instance segmentation
0.312025
OffsetNet: Towards Efficient Multiple Object Tracking, Detection, and Segmentation · IEEE Trans. Pattern Anal. Mach. Intell. 2025
Robotics › Autonomous driving
perception
0.212024
Decoupled Pseudo-Labeling for Semi-Supervised Monocular 3D Object Detection · CVPR 2024

Methods — techniques the papers use, named apart from their topics

query consistency regularization · 1.0pseudo-labeling · 1.0hybrid matching · 1.0vision-language model · 0.9region-level attribute discrimination · 0.9offset-based representation · 0.9linear self-attention · 0.9attribute embedding fusion · 0.9knowledge distillation · 0.8background prompt learning · 0.8
YearPublicationVenuePosition
2026 Toward Efficient Semi-Supervised Object Detection With Detection Transformer
abstract
Semi-supervised object detection (SSOD) mitigates the annotation burden in object detection by leveraging unlabeled data, providing a scalable solution for modern perception systems. Concurrently, detection transformers (DETRs) have emerged as a popular end-to-end framework, offering advantages such as non-maximum suppression (NMS)-free inference. However, existing SSOD methods are predominantly designed for conventional detectors, leaving the exploration of DETR-based SSOD largely uncharted. This paper presents a systematic study to bridge this gap. We begin by identifying two principal obstacles in semi-supervised DETR training: (1) the inherent one-to-one assignment mechanism of DETRs is highly sensitive to noisy pseudo-labels, which impedes training efficiency; and (2) the query-based decoder architecture complicates the design of an effective consistency regularization scheme, limiting further performance gains. To address these challenges, we propose Semi-DETR++, a novel framework for efficient SSOD with DETRs. Our approach introduces a stage-wise hybrid matching strategy that enhances robustness to noisy pseudo-labels by synergistically combining one-to-many and one-to-one assignments while preserving NMS-free inference. Furthermore, based on our observation of the unique layer-wise decoding behavior in DETRs, we develop a simple yet effective re-decode query consistency training method to regularize the decoder. Extensive experiments demonstrate that Semi-DETR++ enables more efficient semi-supervised learning across various DETR architectures, outperforming existing methods by significant margins. The proposed components are also flexible and versatile, showing superior generalization by readily extending to semi-supervised segmentation tasks.
Jiaming Li 0010, Xiangru Lin, Wei Zhang 0197, Xiao Tan 0001, Hongbo Gao 0001, Jingdong Wang 0001, Guanbin Li
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 GUIDED: Granular Understanding via Identification, Detection, and Discrimination for Fine-Grained Open-Vocabulary Object Detection
abstract
Fine-grained open-vocabulary object detection (FG-OVD) aims to detect novel object categories described by attribute-rich texts. While existing open-vocabulary detectors show promise at the base-category level, they underperform in fine-grained settings due to the semantic entanglement of subjects and attributes in pretrained vision-language model (VLM) embeddings -- leading to over-representation of attributes, mislocalization, and semantic drift in embedding space. We propose GUIDED, a decomposition framework specifically designed to address the semantic entanglement between subjects and attributes in fine-grained prompts. By separating object localization and fine-grained recognition into distinct pathways, GUIDED aligns each subtask with the module best suited for its respective roles. Specifically, given a fine-grained class name, we first use a language model to extract a coarse-grained subject and its descriptive attributes. Then the detector is guided solely by the subject embedding, ensuring stable localization unaffected by irrelevant or overrepresented attributes. To selectively retain helpful attributes, we introduce an attribute embedding fusion module that incorporates attribute information into detection queries in an attention-based manner. This mitigates over-representation while preserving discriminative power. Finally, a region-level attribute discrimination module compares each detected region against full fine-grained class names using a refined vision-language model with a projection head for improved alignment. Extensive experiments on FG-OVD and 3F-OVD benchmarks show that GUIDED achieves new state-of-the-art results, demonstrating the benefits of disentangled modeling and modular optimization.
Jiaming Li 0010, Zhijia Liang, Weikai Chen 0001, Lin Ma 0002, Guanbin Li
NeurIPS1
2025 OffsetNet: Towards Efficient Multiple Object Tracking, Detection, and Segmentation
abstract
Offset-based representation has emerged as a promising approach for modeling semantic relations between pixels and object motion, demonstrating efficacy across various computer vision tasks. In this paper, we introduce a novel one-stage multi-tasking network tailored to extend the offset-based approach to MOTS. Our proposed framework, named OffsetNet, is designed to concurrently address amodal bounding box detection, instance segmentation, and tracking. It achieves this by formulating these three tasks within a unified pixel-offset-based representation, thereby achieving excellent efficiency and encouraging mutual collaborations. OffsetNet achieves several remarkable properties: first, the encoder is empowered by a novel Memory Enhanced Linear Self-Attention (MELSA) block to efficiently aggregate spatial-temporal features; second, all tasks are decoupled fairly using three lightweight decoders that operate in a one-shot manner; third, a novel cross-frame offsets prediction module is proposed to enhance the robustness of tracking against occlusions. With these merits, OffsetNet achieves 76.83% HOTA on KITTI MOTS benchmark, which is the best result without relying on 3D detection. Furthermore, OffsetNet achieves 74.83% HOTA at 50 FPS on the KITTI MOT benchmark, which is nearly 3.3 times faster than CenterTrack with better performance. We hope our approach will serve as a solid baseline and encourage future research in this field.
Wei Zhang 0114, Jiaming Li 0010, Xiao Tan 0001, Yifeng Shi, Zhenhua Huang 0001, Guanbin Li
IEEE Trans. Pattern Anal. Mach. Intell.2
2025 Enhancing out-of-distribution detection via diversified multi-prototype contrastive learning
Yulong Jia, Jiaming Li 0010, Ganlong Zhao, Shuangyin Liu, Weijun Sun, Liang Lin 0004, Guanbin Li
Pattern Recognit.2
2024 Learning Background Prompts to Discover Implicit Knowledge for Open Vocabulary Object Detection
abstract
Open vocabulary object detection (OVD) aims at seeking an optimal object detector capable of recognizing objects from both base and novel categories. Recent advances leverage knowledge distillation to transfer insightful knowledge from pre-trained large-scale vision-language models to the task of object detection, significantly generalizing the powerful capabilities of the detector to identify more unknown object categories. However, these methods face significant challenges in background interpretation and model overfitting and thus often result in the loss of crucial back-ground knowledge, giving rise to sub-optimal inference performance of the detector. To mitigate these issues, we present a novel OVD framework termed LBP to propose learning background prompts to harness explored implicit background knowledge, thus enhancing the detection performance w.r.t. base and novel categories. Specifically, we devise three modules: Background Category-specific Prompt, Background Object Discovery, and Inference Probability Rectification, to empower the detector to discover, represent, and leverage implicit object knowledge explored from background proposals. Evaluation on two benchmark datasets, OV-COCO and OV-LVIS, demonstrates the superiority of our proposed method over existing state-of-the-art approaches in handling the OVD tasks.
Jiaming Li 0010, Jichang Li, Ge Li 0002, Si Liu 0001, Liang Lin 0004, Guanbin Li
CVPR1
2024 Decoupled Pseudo-Labeling for Semi-Supervised Monocular 3D Object Detection
abstract
We delve into pseudo-labeling for semi-supervised monocular 3D object detection (SSM30D) and discover two primary issues: a misalignment between the prediction quality of 3D and 2D attributes and the tendency of depth supervision derived from pseudo-labels to be noisy, leading to significant optimization conflicts with other re-liable forms of supervision. To tackle these issues, we introduce a novel decoupled pseudo-labeling (DPL) approach for SSM30D. Our approach features a Decoupled Pseudo-label Generation (DPG) module, designed to efficiently generate pseudo-labels by separately processing 2D and 3D attributes. This module incorporates a unique homography-based method for identifying dependable pseudo-labels in Bird's Eye View (BEV) space, specifically for 3D attributes. Additionally, we present a Depth Gradient Projection (DGP) module to mitigate optimization conflicts caused by noisy depth supervision of pseudo-labels, effectively decoupling the depth gradient and re-moving conflicting gradients. This dual decoupling strat-egy-at both the pseudo-label generation and gradient lev-els-significantly improves the utilization of pseudo-labels in SSM30D. Our comprehensive experiments on the KITTI benchmark demonstrate the superiority of our method over existing approaches.
Jiaming Li 0010, Xiangru Lin, Wei Zhang 0197, Xiao Tan 0001, Junyu Han, Errui Ding, Jingdong Wang 0001, Guanbin Li
CVPR2
2023 Gradient-based Sampling for Class Imbalanced Semi-supervised Object Detection
abstract
Current semi-supervised object detection (SSOD) algorithms typically assume class balanced datasets (PASCAL VOC etc.) or slightly class imbalanced datasets (MS-COCO, etc). This assumption can be easily violated since real world datasets can be extremely class imbalanced in nature, thus making the performance of semi-supervised object detectors far from satisfactory. Besides, the research for this problem in SSOD is severely under-explored. To bridge this research gap, we comprehensively study the class imbalance problem for SSOD under more challenging scenarios, thus forming the first experimental setting for class imbalanced SSOD (CI-SSOD). Moreover, we propose a simple yet effective gradient-based sampling framework that tackles the class imbalance problem from the perspective of two types of confirmation biases. To tackle confirmation bias towards majority classes, the gradient-based reweighting and gradient-based thresholding modules leverage the gradients from each class to fully balance the influence of the majority and minority classes. To tackle the confirmation bias from incorrect pseudo labels of minority classes, the class-rebalancing sampling module resamples unlabeled data following the guidance of the gradient-based reweighting module. Experiments on three proposed sub-tasks, namely MS-COCO, MS-COCO → Object365 and LVIS, suggest that our method outperforms current class imbalanced object detectors by clear margins, serving as a baseline for future research in CI-SSOD. Code will be available at https://github.com/nightkeepers/CI-SSOD.
Jiaming Li 0010, Xiangru Lin, Wei Zhang 0197, Xiao Tan 0001, Junyu Han, Errui Ding, Jingdong Wang 0001, Guanbin Li
ICCV1
2022 Gradient-Rebalanced Uncertainty Minimization for Cross-Site Adaptation of Medical Image Segmentation
Jiaming Li 0010, Chaowei Fang, Guanbin Li
PRCV (2)1
2018 Detecting Text in the Wild with Deep Character Embedding Network
Jiaming Li 0010, Chengquan Zhang, Yipeng Sun, Junyu Han, Errui Ding
ACCV (4)1