Zhenbang Sun

dblp:119/1328 · DBLP profile ↗
← Back
9ranked-venue papers
3as first author
5since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 3 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
7 papers
Video understanding and tracking · 32% Efficient and distributed learning · 25% Representation and self-supervised learning · 18%
Databases, data mining, and information retrieval
3 papers
Spatial and temporal data management · 84% Information retrieval · 16%
Computer graphics and multimedia
2 papers
Multimedia analysis and retrieval · 100%

Topics — the 19 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
inference acceleration
1.012026
OmniSparse: Training-Aware Fine-Grained Sparse Attention for Long-Video MLLMs · AAAI 2026
Machine learning › Efficient and distributed learning › KV cache management
KV cache compression
1.012026
OmniSparse: Training-Aware Fine-Grained Sparse Attention for Long-Video MLLMs · AAAI 2026
Machine learning › Deep learning architectures and training › attention mechanism
sparse attention
1.012026
OmniSparse: Training-Aware Fine-Grained Sparse Attention for Long-Video MLLMs · AAAI 2026
Computer vision › Video understanding and tracking
frame selection
0.912025
Frame-Voyager: Learning to Query Frames for Video Large Language Models · ICLR 2025
Computer vision › Video understanding and tracking
video question answering
0.912025
Frame-Voyager: Learning to Query Frames for Video Large Language Models · ICLR 2025
Machine learning › Representation and self-supervised learning
contrastive learning
0.612022
HCSC: Hierarchical Contrastive Selective Coding · CVPR 2022
Machine learning › Representation and self-supervised learning › representation learning › visual representation learning
image representation learning
0.612022
HCSC: Hierarchical Contrastive Selective Coding · CVPR 2022
Machine learning › Transfer learning and domain adaptation › domain adaptation
unsupervised domain adaptation
0.512021
Cross-category Video Highlight Detection via Set-based Learning · ICCV 2021
Computer vision › Video understanding and tracking › video summarization
video highlight detection
0.512021
Cross-category Video Highlight Detection via Set-based Learning · ICCV 2021
Spatial and temporal data management
trajectory data management
0.412020
Trajectory Similarity Learning with Auxiliary Supervision and Optimal Matching · IJCAI 2020
Spatial and temporal data management › trajectory data management
trajectory similarity
0.412020
Trajectory Similarity Learning with Auxiliary Supervision and Optimal Matching · IJCAI 2020
Multimedia analysis and retrieval › object recognition
sketch recognition
0.322012
Sketch2Tag: automatic hand-drawn sketch recognition · ACM Multimedia 2012
Query-adaptive shape topic mining for hand-drawn sketch recognition · ACM Multimedia 2012
Computer vision › Vision and language › video-language model
video large language model
0.312025
Frame-Voyager: Learning to Query Frames for Video Large Language Models · ICLR 2025
Computer vision › Segmentation and scene understanding › image segmentation
sketch segmentation
0.112012
Free Hand-Drawn Sketch Segmentation · ECCV (1) 2012
Natural language and speech › Information extraction and text analysis
topic model
0.112012
Query-adaptive shape topic mining for hand-drawn sketch recognition · ACM Multimedia 2012
Machine learning › Representation and self-supervised learning › representation learning
embedding learning
0.112020
Trajectory Similarity Learning with Auxiliary Supervision and Optimal Matching · IJCAI 2020
Machine learning › Representation and self-supervised learning › representation learning › embedding learning › temporal embedding
trajectory embedding
0.112020
Trajectory Similarity Learning with Auxiliary Supervision and Optimal Matching · IJCAI 2020
Information retrieval
image retrieval
0.122012
Sketch2Tag: automatic hand-drawn sketch recognition · ACM Multimedia 2012
Query-adaptive shape topic mining for hand-drawn sketch recognition · ACM Multimedia 2012
Information retrieval › image retrieval
sketch-based image retrieval
0.122012
Sketch2Tag: automatic hand-drawn sketch recognition · ACM Multimedia 2012
Query-adaptive shape topic mining for hand-drawn sketch recognition · ACM Multimedia 2012

Methods — techniques the papers use, named apart from their topics

query selection · 1.0dynamic token budget allocation · 1.0KV selection · 1.0KV cache slimming · 1.0ranking · 0.9query-based learning · 0.9prototypical learning · 0.6contrastive learning · 0.6knowledge distillation · 0.5dual-learner framework · 0.5triplet loss · 0.4optimal matching · 0.4metric learning · 0.4real-time recognition · 0.3query-adaptive topic modeling · 0.3generative model · 0.3clipart knowledge base · 0.3
YearPublicationVenuePosition
2026 OmniSparse: Training-Aware Fine-Grained Sparse Attention for Long-Video MLLMs
abstract
Existing sparse attention methods primarily target inference-time acceleration by selecting critical tokens under predefined sparsity patterns. However, they often fail to bridge the training–inference gap and lack the capacity for fine-grained token selection across multiple dimensions—such as queries, key-values (KV), and heads—leading to suboptimal performance and acceleration gains. In this paper, we introduce OmniSparse, a training-aware fine-grained sparse attention of long-video MLLMs, which is applied in both training and inference with dynamic token budget allocation. Specifically, OmniSparse contains three adaptive and complementary mechanisms: (1) query selection as lazy-active classification, aiming to retain active queries that capture broader semantic similarity, while discarding most of lazy ones that focus on limited local context and exhibit high functional redundancy with their neighbors, (2) KV selection with head-level dynamic budget allocation, where a shared budget is determined based on the flattest head and applied uniformly across all heads to ensure attention recall after selection, and (3) KV cache slimming to alleviate head-level redundancy, which selectively fetches visual KV cache according to the head-level decoding query pattern. Experimental results demonstrate that OmniSparse can achieve comparable performance with full attention, achieving 2.7x speedup during prefill and 2.4x memory reduction for decoding.
Feng Chen 0047, Yefei He, Shaoxuan He, Yuanyu He, Jing Liu 0048, Lequan Lin, Akide Liu, Zhenbang Sun, Bohan Zhuang, Qi Wu 0001
AAAI10
2025 Frame-Voyager: Learning to Query Frames for Video Large Language Models
abstract
Video Large Language Models (Video-LLMs) have made remarkable progress in video understanding tasks. However, they are constrained by the maximum length of input tokens, making it impractical to input entire videos. Existing frame selection approaches, such as uniform frame sampling and text-frame retrieval, fail to account for the information density variations in the videos or the complex instructions in the tasks, leading to sub-optimal performance. In this paper, we propose Frame-Voyager that learns to query informative frame combinations, based on the given textual queries in the task. To train Frame-Voyager, we introduce a new data collection and labeling pipeline, by ranking frame combinations using a pre-trained Video-LLM. Given a video of M frames, we traverse its T-frame combinations, feed them into a Video-LLM, and rank them based on Video-LLM's prediction losses. Using this ranking as supervision, we train Frame-Voyager to query the frame combinations with lower losses. In experiments, we evaluate Frame-Voyager on four Video Question Answering benchmarks by plugging it into two different Video-LLMs. The experimental results demonstrate that Frame-Voyager achieves impressive results in all settings, highlighting its potential as a plug-and-play solution for Video-LLMs.
Sicheng Yu, Chengkai Jin, Zhongrong Zuo, Zhenbang Sun, Bingni Zhang, Qianru Sun
ICLR8
2025 ACT as Human: Multimodal Large Language Model Data Annotation with Critical Thinking
abstract
Supervised learning relies on high-quality labeled data, but obtaining such data through human annotation is both expensive and time-consuming. Recent work explores using large language models (LLMs) for annotation, but LLM-generated labels still fall short of human-level quality. To address this problem, we propose the Annotation with Critical Thinking (ACT) data pipeline, where LLMs serve not only as annotators but also as judges to critically identify potential errors. Human effort is then directed towards reviewing only the most "suspicious" cases, significantly improving the human annotation efficiency. Our major contributions are as follows: (1) ACT is applicable to a wide range of domains, including natural language processing (NLP), computer vision (CV), and multimodal understanding, by leveraging multimodal-LLMs (MLLMs). (2) Through empirical studies, we derive 7 insights on how to enhance annotation quality while efficiently reducing the human cost, and then translate these findings into user-friendly guidelines. (3) We theoretically analyze how to modify the loss function so that models trained on ACT data achieve similar performance to those trained on fully human-annotated data. Our experiments show that the performance gap can be reduced to less than 2% on most benchmark datasets while saving up to 90% of human costs.
Lequan Lin, Dai Shi, Andi Han, Feng Chen 0008, Qiuzheng Chen, Zhenbang Sun, Junbin Gao
NeurIPS9
2022 HCSC: Hierarchical Contrastive Selective Coding
abstract
Hierarchical semantic structures naturally exist in an image dataset, in which several semantically relevant image clusters can be further integrated into a larger cluster with coarser-grained semantics. Capturing such structures with image representations can greatly benefit the semantic understanding on various downstream tasks. Existing contrastive representation learning methods lack such an important model capability. In addition, the negative pairs used in these methods are not guaranteed to be semantically distinct, which could further hamper the structural correctness of learned image representations. To tackle these limitations, we propose a novel contrastive learning framework called Hierarchical Contrastive Selective Coding (HCSC). In this framework, a set of hierarchical prototypes are constructed and also dynamically updated to represent the hierarchical semantic structures underlying the data in the latent space. To make image representations better fit such semantic structures, we employ and further improve conventional instance-wise and prototypical contrastive learning via an elaborate pair selection scheme. This scheme seeks to select more diverse positive pairs with similar semantics and more precise negative pairs with truly distinct semantics. On extensive downstream tasks, we verify the state-of-the-art performance of HCSC and also the effectiveness of major model components. We are continually building a comprehensive model zoo (see supplementary material). Our source code and model weights are available at https://github.com/gyfastas/HCSC.
Yuanfan Guo, Bingbing Ni, Zhenbang Sun, Yi Xu 0001
CVPR6
2021 Cross-category Video Highlight Detection via Set-based Learning
abstract
Autonomous highlight detection is crucial for enhancing the efficiency of video browsing on social media platforms. To attain this goal in a data-driven way, one may often face the situation where highlight annotations are not available on the target video category used in practice, while the supervision on another video category (named as source video category) is achievable. In such a situation, one can derive an effective highlight detector on target video category by transferring the highlight knowledge acquired from source video category to the target one. We call this problem cross-category video highlight detection, which has been rarely studied in previous works. For tackling such practical problem, we propose a Dual-Learner-based Video Highlight Detection (DL-VHD) framework. Under this framework, we first design a Set-based Learning module (SL-module) to improve the conventional pair-based learning by assessing the highlight extent of a video segment under a broader context. Based on such learning manner, we introduce two different learners to acquire the basic distinction of target category videos and the characteristics of highlight moments on source video category, respectively. These two types of highlight knowledge are further consolidated via knowledge distillation. Extensive experiments on three benchmark datasets demonstrate the superiority of the proposed SL-module, and the DL-VHD method outperforms five typical Unsupervised Domain Adaptation (UDA) algorithms on various cross-category highlight detection tasks. Our code is available at https://github.com/ChrisAllenMing/Cross_Category_Video_Highlight.
Bingbing Ni, Riheng Zhu, Zhenbang Sun, Changhu Wang
ICCV5
2020 Trajectory Similarity Learning with Auxiliary Supervision and Optimal Matching
abstract
Trajectory similarity computation is a core problem in the field of trajectory data queries. However, the high time complexity of calculating the trajectory similarity has always been a bottleneck in real-world applications. Learning-based methods can map trajectories into a uniform embedding space to calculate the similarity of two trajectories with embeddings in constant time. In this paper, we propose a novel trajectory representation learning framework Traj2SimVec that performs scalable and robust trajectory similarity computation. We use a simple and fast trajectory simplification and indexing approach to obtain triplet training samples efficiently. We make the framework more robust via taking full use of the sub-trajectory similarity information as auxiliary supervision. Furthermore, the framework supports the point matching query by modeling the optimal matching relationship of trajectory points under different distance metrics. The comprehensive experiments on real-world datasets demonstrate that our model substantially outperforms all existing approaches.
Qize Jiang, Baihua Zheng, Zhenbang Sun, Weiwei Sun 0008, Changhu Wang
IJCAI5
2012 Free Hand-Drawn Sketch Segmentation
Zhenbang Sun, Changhu Wang, Liqing Zhang 0001, Lei Zhang 0001
ECCV (1)1
2012 Query-adaptive shape topic mining for hand-drawn sketch recognition
abstract
In this work, we study the problem of hand-drawn sketch recognition. Due to large intra-class variations presented in hand-drawn sketches, most of existing work was limited to a particular domain or limited pre-defined classes. Different from existing work, we target at developing a general sketch recognition system, to recognize any semantically meaningful object that a child can recognize. To increase the recognition coverage, a web-scale clipart image collection is leveraged as the knowledge base of the recognition system. To alleviate the problems of intra-class shape variation and inter-class shape ambiguity in this unconstrained situation, a query-adaptive shape topic model is proposed to mine object topics and shape topics related to the sketch, in which, multiple layers of information such as sketch, object, shape, image, and semantic labels are modeled in a generative process. Besides sketch recognition, the proposed topic model can also be used for related applications such as sketch tagging, image tagging, and sketch-based image search. Extensive experiments on different applications show the effectiveness of the proposed topic model and the recognition system.
Zhenbang Sun, Changhu Wang, Liqing Zhang 0001, Lei Zhang 0001
ACM Multimedia1
2012 Sketch2Tag: automatic hand-drawn sketch recognition
abstract
In this work, we introduce the Sketch2Tag system for hand-drawn sketch recognition. Due to large variations presented in hand-drawn sketches, most of existing work was limited to a particular domain or limited predefined classes. Different from existing work, Sketch2Tag is a general sketch recognition system, towards recognizing any semantically meaningful object that a child can recognize. This system enables a user to draw a sketch on the query panel, and then provides real-time recognition results. To increase the recognition coverage, a web-scale clipart image collection is leveraged as the knowledge base of the recognition system. Better understanding a user's drawing will be of great value to a variety of applications, such as, improving the sketch-based image search by combining the recognition results as textual queries.
Zhenbang Sun, Changhu Wang, Liqing Zhang 0001, Lei Zhang 0001
ACM Multimedia1