EDBT 2026 Demo / reviewers in the wild / expert
Hongyang Li 0003
dblp:95/8433-3
· DBLP profile ↗
10ranked-venue papers
5as first author
10since 2021 · last 2026
0000-0003-3231-0655ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 9 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 9 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SegDINO3D: 3D Instance Segmentation Empowered by Both Image-Level and Object-Level 2D FeaturesabstractIn this paper, we present SegDINO3D, a novel Transformer encoder-decoder framework for 3D instance segmentation. As 3D training data is generally not as sufficient as 2D training images, SegDINO3D is designed to fully leverage 2D representation from a pre-trained 2D detection model, including both image-level and object-level features, for improving 3D representation. SegDINO3D takes both a point cloud and its associated 2D images as input. In the encoder stage, it first enriches each 3D point by retrieving 2D image features from its corresponding image views and then leverages a 3D encoder for 3D context fusion. In the decoder stage, it formulates 3D object queries as 3D anchor boxes and performs cross-attention from 3D queries to 2D object queries obtained from 2D images using the 2D detection model. These 2D object queries serve as a compact object-level representation of 2D images, effectively avoiding the challenge of keeping thousands of image feature maps in the memory while faithfully preserving the knowledge of the pre-trained 2D model. The introducing of 3D box queries also enables the model to modulate cross-attention using the predicted boxes for more precise querying. SegDINO3D achieves the state-of-the-art performance on the ScanNetV2 and ScanNet200 3D instance segmentation benchmarks. Notably, on the challenging ScanNet200 dataset, SegDINO3D significantly outperforms prior methods by +8.7 and +6.8 mAP on the validation and hidden test sets, respectively, demonstrating its superiority. Jinyuan Qu, Hongyang Li 0003, Xingyu Chen 0002, Shilong Liu 0004, Yukai Shi, Tianhe Ren, Ruitao Jing, Lei Zhang 0001 |
AAAI | 2 |
| 2025 | TAPTR3D: Decoupled 3D Point Tracking Boosts 2D and Further Enhances 3D Tracking Accuracy
Hongyang Li 0003, Jinyuan Qu, Zhaoyang Zeng, Lei Zhang 0001 |
IEEE Signal Process. Lett. | 1 |
| 2024 | Visual in-Context PromptingabstractIn-context prompting in large language models (LLMs) has become a prevalent approach to improve zero-shot capabilities, but this idea is less explored in the vision domain. Existing visual prompting methods focus on referring segmentation to segment the most relevant object, falling short of addressing many generic vision tasks like open-set segmentation and detection. In this paper, we introduce a universal visual in-context prompting framework for both tasks, as shown in Fig. 1. In particular, we build on top of an encoder-decoder architecture, and develop a versatile prompt encoder to support a variety of prompts like strokes, boxes, and points. We further enhance it to take an arbitrary number of reference image segments as the context. Our extensive explorations show that the proposed visual in-context prompting elicits extraordinary referring and generic segmentation capabilities to refer and detect, yielding competitive performance to close-set in-domain datasets and showing promising results on many open-set segmentation datasets. By joint training on COCO and SA-1B, DINOv achieves 57.7 PQ on COCO and 23.2 PQ on ADE20K. Code will be available at https://github.com/UX-Decoder/DINOv Feng Li 0040, Hao Zhang 0097, Tianhe Ren, Shilong Liu 0004, Xueyan Zou, Huaizhe Xu, Hongyang Li 0003, Chunyuan Li, Lei Zhang 0001, Jianfeng Gao 0001 |
CVPR | 8 |
| 2024 | TAPTR: Tracking Any Point with Transformers as Detection
Hongyang Li 0003, Hao Zhang 0097, Shilong Liu 0004, Zhaoyang Zeng, Tianhe Ren, Feng Li 0040, Lei Zhang 0001 |
ECCV (16) | 1 |
| 2024 | LLaVA-Grounding: Grounded Visual Chat with Large Multimodal Models
Hao Zhang 0097, Hongyang Li 0003, Feng Li 0040, Tianhe Ren, Xueyan Zou, Shilong Liu 0004, Shijia Huang, Jianfeng Gao 0001, Leizhang, Chunyuan Li, Jainwei Yang |
ECCV (43) | 2 |
| 2024 | TAPTRv2: Attention-based Position Update Improves Tracking Any PointabstractIn this paper, we present TAPTRv2, a Transformer-based approach built upon TAPTR for solving the Tracking Any Point (TAP) task. TAPTR borrows designs from DEtection TRansformer (DETR) and formulates each tracking point as a point query, making it possible to leverage well-studied operations in DETR-like algorithms. TAPTRv2 improves TAPTR by addressing a critical issue regarding its reliance on cost-volume, which contaminates the point query’s content feature and negatively impacts both visibility prediction and cost-volume computation. In TAPTRv2, we propose a novel attention-based position update (APU) operation and use key-aware deformable attention to realize. For each query, this operation uses key-aware attention weights to combine their corresponding deformable sampling positions to predict a new query position. This design is based on the observation that local attention is essentially the same as cost-volume, both of which are computed by dot-production between a query and its surrounding features. By introducing this new operation, TAPTRv2 not only removes the extra burden of cost-volume computation, but also leads to a substantial performance improvement. TAPTRv2 surpasses TAPTR and achieves state-of-the-art performance on many challenging datasets, demonstrating the effectiveness of our approach. Hongyang Li 0003, Hao Zhang 0097, Shilong Liu 0004, Zhaoyang Zeng, Feng Li 0040, Tianhe Ren, Lei Zhang 0006 |
NeurIPS | 1 |
| 2023 | Lite DETR : An Interleaved Multi-Scale Encoder for Efficient DETRabstractRecent DEtection TRansformer-based (DETR) models have obtained remarkable performance. Its success cannot be achieved without the re-introduction of multi-scale feature fusion in the encoder. However, the excessively increased tokens in multi-scale features, especially for about 75% of low-level features, are quite computationally inefficient, which hinders real applications of DETR models. In this paper, we present Lite DETR, a simple yet efficient end-to-end object detection framework that can effectively reduce the GFLOPs of the detection head by 60% while keeping 99% of the original performance. Specifically, we design an efficient encoder block to update high-level features (corresponding to small-resolution feature maps) and low-level features (corresponding to large-resolution feature maps) in an interleaved way. In addition, to better fuse cross-scale features, we develop a key-aware deformable attention to predict more reliable attention weights. Comprehensive experiments validate the effectiveness and efficiency of the proposed Lite DETR, and the efficient encoder strategy can generalize well across existing DETR-based models. The code will be available in https://github.com/IDEA-Research/Lite-DETR. Feng Li 0040, Ailing Zeng, Shilong Liu 0004, Hao Zhang 0097, Hongyang Li 0003, Lei Zhang 0001, Lionel M. Ni |
CVPR | 5 |
| 2023 | DFA3D: 3D Deformable Attention For 2D-to-3D Feature LiftingabstractIn this paper, we propose a new operator, called 3D DeFormable Attention (DFA3D), for 2D-to-3D feature lifting, which transforms multi-view 2D image features into a unified 3D space for 3D object detection. Existing feature lifting approaches, such as Lift-Splat-based and 2D attention-based, either use estimated depth to get pseudo LiDAR features and then splat them to a 3D space, which is a one-pass operation without feature refinement, or ignore depth and lift features by 2D attention mechanisms, which achieve finer semantics while suffering from a depth ambiguity problem. In contrast, our DFA3D-based method first leverages the estimated depth to expand each view’s 2D feature map to 3D and then utilizes DFA3D to aggregate features from the expanded 3D feature maps. With the help of DFA3D, the depth ambiguity problem can be effectively alleviated from the root, and the lifted features can be progressively refined layer by layer, thanks to the Transformerlike architecture. In addition, we propose a mathematically equivalent implementation of DFA3D which can significantly improve its memory efficiency and computational speed. We integrate DFA3D into several methods that use 2D attention-based feature lifting with only a few modifications in code and evaluate on the nuScenes dataset. The experiment results show a consistent improvement of +1.41% mAP on average, and up to +15.1% mAP improvement when high-quality depth information is available, demonstrating the superiority, applicability, and huge potential of DFA3D. The code is available at https://github.com/IDEAResearch/3D-deformable-attention.git. Hongyang Li 0003, Hao Zhang 0097, Zhaoyang Zeng, Shilong Liu 0004, Feng Li 0040, Tianhe Ren, Lei Zhang 0001 |
ICCV | 1 |
| 2023 | Detection Transformer with Stable MatchingabstractThis paper is concerned with the matching stability problem across different decoder layers in DEtection TRansformers (DETR). We point out that the unstable matching in DETR is caused by a multi-optimization path problem, which is highlighted by the one-to-one matching design in DETR. To address this problem, we show that the most important design is to use and only use positional metrics (like IOU) to supervise classification scores of positive examples. Under the principle, we propose two simple yet effective modifications by integrating positional metrics to DETR’s classification loss and matching cost, named position-supervised loss and position-modulated cost. We verify our methods on several DETR variants. Our methods show consistent improvements over baselines. By integrating our methods with DINO, we achieve 50.4 and 51.5 AP on the COCO detection benchmark using ResNet-50 backbones under 1× (12 epochs) and 2× (24 epochs) training settings, achieving a new record under the same setting. We achieve 63.8 AP on COCO detection test-dev with a Swin-Large backbone. Our code will be made available at https://github.com/IDEA-Research/Stable-DINO. Shilong Liu 0004, Tianhe Ren, Zhaoyang Zeng, Hao Zhang 0097, Feng Li 0040, Hongyang Li 0003, Jun Huang 0007, Hang Su 0006, Jun Zhu 0001, Lei Zhang 0001 |
ICCV | 7 |
| 2022 | DCL-Net: Deep Correspondence Learning Network for 6D Pose Estimation
Hongyang Li 0003, Jiehong Lin, Kui Jia |
ECCV (9) | 1 |