Qingbin Liu

dblp:137/6023 · DBLP profile ↗
← Back
17ranked-venue papers
4as first author
17since 2021 · last 2026
0000-0002-2687-9514ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 13 · 2 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 3 since 2021
YearPublicationVenuePosition
2026 Editing the Moving World: Model Editing for Video LLMs
abstract
Qian Zhang, Xinye Li, Xiaokai Wu, Junhao Xu, Zhanyue Qin, Qingbin Liu, Junxian Cai, Xi Chen, Bolin Zhang, Zhiying Tu, Dianhui Chu, Xiaoyan Yu, Dianbo Sui. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xinye Li 0001, Xiaokai Wu, Zhanyue Qin, Qingbin Liu, Junxian Cai, Xi Chen 0003, Zhiying Tu, Dianbo Sui
ACL (1)6
2025 TC-LLaVA: Rethinking the Transfer of LLava from Image to Video Understanding with Temporal Considerations
abstract
Multimodal Large Language Models (MLLMs) have significantly improved performance across various image-language applications. Recently, there has been a growing interest in adapting image pre-trained MLLMs for video-related tasks. However, most efforts concentrate on enhancing the vision encoder and projector components, while the core part, Large Language Models (LLMs), remains comparatively under-explored. In this paper, we propose two strategies to enhance the model's capability in video understanding tasks by improving inter-layer attention computation in LLMs. Specifically, the first approach focuses on the enhancement of Rotary Position Embedding (RoPE) with Temporal-Aware Dual RoPE, which introduces temporal position information to strengthen the MLLM's temporal modeling capabilities while preserving the relative position relationships of both visual and text tokens. The second approach involves enhancing the Attention Mask with the Frame-wise Block Causal Attention Mask, a simple yet effective method that broadens visual token interactions within and across video frames while maintaining the causal inference mechanism. Based on these proposed methods, we adapt LLaVA for video understanding tasks, naming it Temporal-Considered LLaVA (TC-LLaVA). Our TC-LLaVA achieves new state-of-the-art performance across various video understanding benchmarks with only supervised fine-tuning (SFT) on video-related datasets.
Jiangtao Xie, Qingbin Liu, Kevin Zhao, Hui Xiong 0001
AAAI5
2025 VTG-LLM: Integrating Timestamp Knowledge into Video LLMs for Enhanced Video Temporal Grounding
abstract
Video Temporal Grounding (VTG) strives to accurately pinpoint event timestamps in a specific video using linguistic queries, significantly impacting downstream tasks like video browsing and editing. Unlike traditional task-specific models, Video Large Language Models (video LLMs) can handle multiple tasks concurrently in a zero-shot manner. Consequently, exploring the application of video LLMs for VTG tasks has become a burgeoning research area. However, despite considerable advancements in video content understanding, video LLMs often struggle to accurately pinpoint timestamps within videos, limiting their effectiveness in VTG tasks. To address this, we introduce VTG-LLM, a model designed to enhance video LLMs' timestamp localization abilities. Our approach includes: (1) effectively integrating timestamp knowledge into visual tokens; (2) incorporating absolute-time tokens to manage timestamp knowledge without concept shifts; and (3) introducing a lightweight, high-performance, slot-based token compression technique designed to accommodate the demands of a large number of frames to be sampled for VTG tasks. Additionally, we present VTG-IT-120K, a collection of publicly available VTG datasets that we have re-annotated to improve upon low-quality annotations. Our comprehensive experiments demonstrate the superior performance of VTG-LLM in comparison to other video LLM methods across a variety of VTG tasks.
Yongxin Guo 0001, Dingxin Cheng, Xiaoying Tang 0002, Dianbo Sui, Qingbin Liu, Xi Chen 0003, Kevin Zhao
AAAI7
2025 HFF-Tracker: A Hierarchical Fine-grained Fusion Tracker for Referring Multi-Object Tracking
abstract
Referring Multi-Object Tracking (RMOT) aims to track multiple objects based on a provided language expression. Although prior studies have sought to accomplish this by integrating an textual module into the multi-object tracker, these methods combine text and image features in a basic way, neglecting the importance of text features. In this study, we propose a Hierarchical Fine-grained text-image Fusion tracker, named HFF-Tracker, which can perform fine-grained fusion of pixel-level visual features and text features across various semantic levels. Specifically, we have devised a Hierarchical Multi-Modal Fusion (HMMF) module to merge text and image features at an early stage in a hierarchical and detailed manner. The Text-Guided Decoder (TGD) is designed to provide the query with prior semantic information during the decoding process. Additionally, we have crafted a Text-Guided Prediction Head (TGPH) that utilizes text information to enhance the performance of the prediction head. Furthermore, we have implemented an adaptive Look-Back training strategy to maximize the utilization of valuable labeled data. Extensive experiments on the Refer-KITTI dataset and the Refer-KITTI-V2 dataset demonstrate that our proposed HFF-Tracker outperforms other state-of-the-art methods with remarkable margins.
Zeyong Zhao, Yanchao Hao, Qingbin Liu, Dianbo Sui, Shizhu He, Xi Chen 0003
AAAI4
2025 VRoPE: Rotary Position Embedding for Video Large Language Models
abstract
Rotary Position Embedding (RoPE) has shown strong performance in text-based Large Language Models (LLMs), but extending it to video remains a challenge due to the intricate spatiotemporal structure of video frames.Existing adaptations, such as RoPE-3D, attempt to encode spatial and temporal dimensions separately but suffer from two major limitations: positional bias in attention distribution and disruptions in video-text transitions.To overcome these issues, we propose Video Rotary Position Embedding (VRoPE), a novel positional encoding method tailored for Video-LLMs.Specifically, we introduce a more balanced encoding strategy that mitigates attention biases, ensuring a more uniform distribution of spatial focus.Additionally, our approach restructures positional indices to ensure a smooth transition between video and text tokens.Extensive experiments on different models demonstrate that VRoPE consistently outperforms previous RoPE variants, achieving significant improvements in video understanding, temporal reasoning, and retrieval tasks.Code is available at https://github.com/johncaged/VRoPE.
Longteng Guo, Yepeng Tang, Tongtian Yue, Junxian Cai, Qingbin Liu, Jing Liu 0001
EMNLP7
2025 M2Edit: Locate and Edit Multi-Granularity Knowledge in Multimodal Large Language Model
abstract
Multimodal knowledge editing is an important method for modifying outdated or incorrect knowledge in Multimodal Large Language Models (MLLMs).However, existing datasets for multimodal knowledge editing lack multi-granularity knowledge.In this paper, we present a more realistic dataset called M2Edit, which includes three distinct types of knowledge: entity, relation, and action.Additionally, existing knowledge editing methods for MLLMs lack the ability to handle multigranularity knowledge and generalize to multimodal data.To address these limitations, we propose the multimodal knowledge editing method MLE.This approach identifies key knowledge layers within different components and collaboratively edits the various components of MLLMs.As a result, we observe significant improvements in visual generality performance, ranging from 4.8% to 10.8%, and achieve the best overall performance on knowledge data of different granularities.
Yubo Chen 0001, Qingbin Liu, Dianbo Sui, Xi Chen 0003, Kang Liu 0001, Jun Zhao 0001
EMNLP4
2025 TRACE: Temporal Grounding Video LLM via Causal Event Modeling
abstract
Video Temporal Grounding (VTG) is a crucial capability for video understanding models and plays a vital role in downstream tasks such as video browsing and editing. To effectively handle various tasks simultaneously and enable zero-shot prediction, there is a growing trend in employing video LLMs for VTG tasks. However, current video LLM-based methods rely exclusively on natural language generation, lacking the ability to model the clear structure inherent in videos, which restricts their effectiveness in tackling VTG tasks. To address this issue, this paper first formally introduces causal event modeling framework, which represents video LLM outputs as sequences of events, and predict the current event using previous events, video inputs, and textural instructions. Each event consists of three components: timestamps, salient scores, and textual captions. We then propose a novel task-interleaved video LLM called TRACE to effectively implement the causal event modeling framework in practice. The TRACE process visual frames, timestamps, salient scores, and text as distinct tasks, employing various encoders and decoding heads for each. Task tokens are arranged in an interleaved sequence according to the causal event modeling framework's formulation. Extensive experiments on various VTG tasks and datasets demonstrate the superior performance of TRACE compared to state-of-the-art video LLMs. Our model and code are avaliable at \url{https://github.com/gyxxyg/TRACE}.
Yongxin Guo 0001, Qingbin Liu, Xi Chen 0003, Xiaoying Tang 0002
ICLR4
2025 Enhancing Long Video Understanding via Hierarchical Event-Based Memory
abstract
Recently, integrating visual foundation models into large language models (LLMs) to form video understanding systems has attracted widespread attention. Most of the existing models compress diverse semantic information within the whole video and feed it into LLMs for content comprehension. While this method excels in short video understanding, it may result in a blend of multiple event information in long videos due to coarse compression, which causes information redundancy. Consequently, the semantics of key events might be obscured within the vast information that hinders the model’s understanding capabilities. To address this issue, we propose a Hierarchical Event-based Memory-enhanced LLM (HEM-LLM) for better understanding of long videos. Firstly, we design a novel adaptive sequence segmentation scheme to divide multiple events within long videos. In this way, we can perform individual memory modeling for each event to establish intra-event contextual connections, thereby reducing information redundancy. Secondly, while modeling current event, we compress and inject the information of the previous event to enhance the long-term inter-event dependencies in videos. Finally, we perform extensive experiments on various video understanding tasks and the results show that our model achieves state-of-the-art performances.
Dingxin Cheng, Yongxin Guo 0001, Bin Jiang 0011, Qingbin Liu, Xi Chen 0003
ICME6
2024 Editing Language Model-Based Knowledge Graph Embeddings
abstract
Recently decades have witnessed the empirical success of framing Knowledge Graph (KG) embeddings via language models. However, language model-based KG embeddings are usually deployed as static artifacts, making them difficult to modify post-deployment without re-training after deployment. To address this issue, we propose a new task of editing language model-based KG embeddings in this paper. This task is designed to facilitate rapid, data-efficient updates to KG embeddings without compromising the performance of other aspects. We build four new datasets: E-FB15k237, A-FB15k237, E-WN18RR, and A-WN18RR, and evaluate several knowledge editing baselines demonstrating the limited ability of previous models to handle the proposed challenging task. We further propose a simple yet strong baseline dubbed KGEditor, which utilizes additional parametric layers of the hypernetwork to edit/add facts. Our comprehensive experimental results reveal that KGEditor excels in updating specific facts without impacting the overall performance, even when faced with limited training resources. Code and datasets will be available at https://github.com/AnonymousForPapers/DeltaKG.
Siyuan Cheng 0008, Ningyu Zhang 0001, Bozhong Tian, Xi Chen 0003, Qingbin Liu, Huajun Chen
AAAI5
2024 Analyzing Chain-of-thought Prompting in Black-Box Large Language Models via Estimated V-information
abstract
Chain-of-Thought (CoT) prompting combined with large language models (LLM) has shown great potential in improving performance on challenging reasoning tasks. While understanding why CoT prompting is effective is crucial for the application and improvement of CoT prompting, few studies have addressed this issue. Besides, almost no prior work has conducted theoretical analysis on CoT prompting in the context of black-box models. In this paper, we approach the analysis of CoT prompting in black-box LLMs from an information-theoretic perspective. Specifically, we propose a new metric, EPVI (Estimated Pointwise V-Information), which extends the concept of pointwise V-information to black-box models, quantifying the label-relevant new information introduced by CoT prompting beyond the pre-existing information in the input. Based on this, we conduct a series of experiments at both the task and instance levels to analyze CoT prompting, demonstrating that the effectiveness of CoT prompting can be attributed to its capacity to influence the difficulty of model inference by augmenting or reducing the model-usable information. Furthermore, we show that selecting high-quality demonstrations of CoT reasoning based on EPVI can improve the downstream performance of reasoning tasks.
Zecheng Wang, Chunshan Li, Zhao Yang 0004, Qingbin Liu, Yanchao Hao, Xi Chen 0003, Dianbo Sui
LREC/COLING4
2024 Pruning via Merging: Compressing LLMs via Manifold Alignment Based Layer Merging
abstract
Deyuan Liu, Zhanyue Qin, Hairu Wang, Zhao Yang, Zecheng Wang, Fangying Rong, Qingbin Liu, Yanchao Hao, Bo Li, Xi Chen, Cunhang Fan, Zhao Lv, Dianhui Chu, Zhiying Tu, Dianbo Sui. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Deyuan Liu, Zhanyue Qin, Hairu Wang 0002, Zhao Yang 0004, Zecheng Wang, Fangying Rong, Qingbin Liu, Yanchao Hao, Xi Chen 0003, Cunhang Fan, Zhao Lv, Zhiying Tu, Dianbo Sui
EMNLP7
2023 Can We Edit Multimodal Large Language Models?
abstract
In this paper, we focus on editing Multimodal Large Language Models (MLLMs).Compared to editing single-modal LLMs, multimodal model editing is more challenging, which demands a higher level of scrutiny and careful consideration in the editing process.To facilitate research in this area, we construct a new benchmark, dubbed MMEdit, for editing multimodal LLMs and establishing a suite of innovative metrics for evaluation.We conduct comprehensive experiments involving various model editing baselines and analyze the impact of editing different components for multimodal LLMs.Empirically, we notice that previous baselines can implement editing multimodal LLMs to some extent, but the effect is still barely satisfactory, indicating the potential difficulty of this task.We hope that our work can provide the NLP community with insights1.
Siyuan Cheng 0008, Bozhong Tian, Qingbin Liu, Xi Chen 0003, Yongheng Wang, Huajun Chen, Ningyu Zhang 0001
EMNLP3
2023 Unsupervised Domain Adaptation on Sentence Matching Through Self-Supervision
Guirong Bai, Qingbin Liu, Shizhu He, Kang Liu 0001, Jun Zhao 0001
J. Comput. Sci. Technol.2
2023 Unsupervised Dialogue State Tracking for End-to-End Task-Oriented Dialogue with a Multi-Span Prediction Network
Qingbin Liu, Shizhu He, Cao Liu, Kang Liu 0001, Jun Zhao 0001
J. Comput. Sci. Technol.1
2021 Domain-Lifelong Learning for Dialogue State Tracking via Knowledge Preservation Networks
abstract
Dialogue state tracking (DST), which estimates user goals given a dialogue context, is an essential component of task-oriented dialogue systems.Conventional DST models are usually trained offline, which requires a fixed dataset prepared in advance.This paradigm is often impractical in real-world applications since online dialogue systems usually involve continually emerging new data and domains.Therefore, this paper explores Domain-Lifelong Learning for Dialogue State Tracking (DLL-DST), which aims to continually train a DST model on new data to learn incessantly emerging new domains while avoiding catastrophically forgetting old learned domains.To this end, we propose a novel domainlifelong learning method, called Knowledge Preservation Networks (KPN), which consists of multi-prototype enhanced retrospection and multi-strategy knowledge distillation, to solve the problems of expression diversity and combinatorial explosion in the DLL-DST task.Experimental results show that KPN effectively alleviates catastrophic forgetting and outperforms previous state-of-the-art lifelong learning methods by 4.25% and 8.27% of whole joint goal accuracy on the MultiWOZ benchmark and the SGD benchmark, respectively.
Qingbin Liu, Cao Liu, Jiansong Chen, Fan Yang 0087, Shizhu He, Kang Liu 0001, Jun Zhao 0001
EMNLP (1)1
2021 A Unified Shared-Private Network with Denoising for Dialogue State Tracking
Qingbin Liu, Shizhu He, Kang Liu 0001, Shengping Liu, Jun Zhao 0001
J. Comput. Sci. Technol.1
2021 Heterogeneous Relational Graph Neural Networks with Adaptive Objective for End-to-End Task-Oriented Dialogue
Qingbin Liu, Guirong Bai, Shizhu He, Cao Liu, Kang Liu 0001, Jun Zhao 0001
Knowl. Based Syst.1