EDBT 2026 Demo / reviewers in the wild / expert
Huazhang Hu
dblp:317/7041
· DBLP profile ↗
3ranked-venue papers
1as first author
3since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Video understanding and tracking · 78% Vision and language · 22% |
Topics — the 5 heaviest of 5, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Computer vision › Video understanding and tracking
video representation learning |
0.7 | 1 | 2023 | Weakly Supervised Video Representation Learning with Unaligned Text for Sequential Videos · CVPR 2023 |
Computer vision › Vision and language › video-text retrieval
video-text matching |
0.7 | 1 | 2023 | Weakly Supervised Video Representation Learning with Unaligned Text for Sequential Videos · CVPR 2023 |
Computer vision › Video understanding and tracking › human action analysis › action understanding
action counting |
0.6 | 1 | 2022 | TransRAC: Encoding Multi-scale Temporal Correlation with Transformers for Repetitive Action Counting · CVPR 2022 |
Computer vision › Video understanding and tracking › human action analysis › action understanding › action counting
repetitive action counting |
0.6 | 1 | 2022 | TransRAC: Encoding Multi-scale Temporal Correlation with Transformers for Repetitive Action Counting · CVPR 2022 |
Computer vision › Video understanding and tracking › human action analysis › action understanding
temporal action analysis |
0.6 | 1 | 2022 | TransRAC: Encoding Multi-scale Temporal Correlation with Transformers for Repetitive Action Counting · CVPR 2022 |
Methods — techniques the papers use, named apart from their topics
transformer · 1.2pseudo-labeling · 0.7contrastive learning · 0.7CLIP · 0.7density map regression · 0.6
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | CQ-DINO: Mitigating Gradient Dilution via Category Queries for Vast Vocabulary Object DetectionabstractWith the exponential growth of data, traditional object detection methods are increasingly struggling to handle vast vocabulary object detection tasks effectively. We analyze two key limitations of classification-based detectors: positive gradient dilution, where rare positive categories receive insufficient learning signals, and hard negative gradient dilution, where discriminative gradients are overwhelmed by numerous easy negatives. To address these challenges, we propose CQ-DINO, a category query-based object detection framework that reformulates classification as a contrastive task between object queries and learnable category queries. Our method introduces image-guided query selection, which reduces the negative space by adaptively retrieving top-K relevant categories per image via cross-attention, thereby rebalancing gradient distributions and facilitating implicit hard example mining.
Furthermore, CQ-DINO flexibly integrates explicit hierarchical category relationships in structured datasets (e.g., V3Det) or learns implicit category correlations via self-attention in generic datasets (e.g., COCO). Experiments demonstrate that CQ-DINO achieves superior performance on the challenging V3Det benchmark (surpassing previous methods by 2.1% AP) while maintaining competitiveness in COCO. Our work provides a scalable solution for real-world detection systems requiring wide category coverage. Huazhang Hu, Yidong Ma, Xu Tang 0007, Yao Hu 0002, Yongchao Xu |
NeurIPS | 2 |
| 2023 | Weakly Supervised Video Representation Learning with Unaligned Text for Sequential VideosabstractSequential video understanding, as an emerging video understanding task, has driven lots of researchers' attention because of its goal-oriented nature. This paper studies weakly supervised sequential video understanding where the accurate time-stamp level text-video alignment is not provided. We solve this task by borrowing ideas from CLIP. Specifically, we use a transformer to aggregate frame-level features for video representation and use a pre-trained text encoder to encode the texts corresponding to each action and the whole video, respectively. To model the correspondence between text and video, we propose a multiple granularity loss, where the video-paragraph contrastive loss enforces matching between the whole video and the complete script, and a fine-grained frame-sentence contrastive loss enforces the matching between each action and its description. As the frame-sentence correspondence is not available, we propose to use the fact that video actions happen sequentially in the temporal domain to generate pseudo frame-sentence correspondence and supervise the network training with the pseudo labels. Extensive experiments on video sequence verification and text-to-video matching show that our method outperforms baselines by a large margin, which validates the effectiveness of our proposed approach. Code is available at https://github.com/svip-lab/WeakSVR. Sixun Dong, Huazhang Hu, Dongze Lian, Weixin Luo, Yicheng Qian, Shenghua Gao |
CVPR | 2 |
| 2022 | TransRAC: Encoding Multi-scale Temporal Correlation with Transformers for Repetitive Action CountingabstractCounting repetitive actions are widely seen in human activities such as physical exercise. Existing methods focus on performing repetitive action counting in short videos, which is tough for dealing with longer videos in more realistic scenarios. In the data-driven era, the degradation of such generalization capability is mainly attributed to the lack of long video datasets. To complement this margin, we introduce a new large-scale repetitive action counting dataset covering a wide variety of video lengths, along with more realistic situations where action interruption or action inconsistencies occur in the video. Besides, we also provide a fine-grained annotation of the action cycles instead of just counting annotation along with a numerical value. Such a dataset contains 1,451 videos with about 20,000 annotations, which is more challenging. For repetitive action counting towards more realistic scenarios, we further propose encoding multi-scale temporal correlation with transformers that can take into account both performance and efficiency. Furthermore, with the help of fine-grained annotation of action cycles, we propose a density map regression-based method to predict the action period, which yields better performance with sufficient interpretability. Our proposed method outperforms state-of-the-art methods on all datasets and also achieves better performance on the unseen dataset without fine-tuning. The dataset and code are available11https://svip-lab.github.io/dataset/RepCount_dataset.html. Huazhang Hu, Sixun Dong, Yiqun Zhao, Dongze Lian, Shenghua Gao |
CVPR | 1 |