VLDB 2026 Research / reviewers in the wild / expert
Zijia Zhao
dblp:296/3659
· DBLP profile ↗
17ranked-venue papers
6as first author
17since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 3 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Systems, architecture and hardware · 3 · 2 first-author · 3 since 2021Computer networks · 2 · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | M³-VQA: A Benchmark for Multimodal, Multi-Entity, Multi-Hop Visual Question AnsweringabstractJiatong Ma, Longteng Guo, Yuchen Liu, Zijia Zhao, Dongze Hao, Xuanxu Lin, Jing Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Jiatong Ma, Longteng Guo, Zijia Zhao, Dongze Hao, Xuanxu Lin, Jing Liu 0001 |
ACL (1) | 4 |
| 2026 | Dynamic priority-based area partitioning, trajectory planning, and task scheduling in computing-while-flying UAV networks
Zijia Zhao, Wenhan Zhan, Geyong Min, Xu Jiang 0004, Liang Zhao 0004, Hualong Huang |
Future Gener. Comput. Syst. | 1 |
| 2026 | AITQE: An Adaptive Image-Text Quality Enhancer for Scalable MLLM PretrainingabstractMultimodal large language models (MLLMs) have made significant strides by integrating visual and textual modalities. A critical factor in training MLLMs is the quality of image-text pairs within multimodal pretraining datasets. However, in the process of high-quality data curation, filter-based paradigms often discard a substantial portion of high-quality images due to inadequate semantic alignment between images and texts, leading to inefficiency in data utilization and scalability. In this paper, we propose the Adaptive Image-Text Quality Enhancer (AITQE), a model that dynamically assesses and enhances the quality of image-text pairs. AITQE employs a text rewriting mechanism for low-quality pairs and incorporates a negative sample learning strategy to improve evaluative capabilities by integrating deliberately generated low-quality samples during training. Unlike prior approaches that significantly alter text distributions, our method minimally adjusts text to preserve data volume while enhancing quality. Experimental results demonstrate that AITQE surpasses existing methods on various benchmarks, effectively leveraging raw data and scaling with increasing data volumes. Codes and model are available at https://github.com/hanhuang22/AITQE. Yuqi Huo, Zijia Zhao, Haoyu Lu, Bingning Wang, Qiang Liu 0006, Weipeng Chen, Liang Wang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | EdgeSD: Efficient Speculative Decoding With Vision-Decoding Disaggregation for MLLM Inference in Edge-Cloud NetworksabstractThe deployment of multimodal large language models (MLLMs) in edge-cloud networks faces critical challenges, including computational resource heterogeneity, memory bottlenecks, and bandwidth constraints. To address these issues, we propose EdgeSD, a novel framework that accelerates MLLM inference by integrating speculative decoding (SD) with edge-cloud collaboration. First, EdgeSD decouples the vision encoding and decoding processes of the draft MLLM across heterogeneous edge servers (ESs). This disaggregation architecture overcomes single-node memory constraints, enabling optimized resource utilization and high-resolution input processing. Second, to resolve the communication bottleneck and computational burden inherent in this distributed architecture, EdgeSD integrates a bandwidth-aware dynamic image token merging (ITM) method. Unlike general pruning techniques, this EdgeSD-specific ITM method focuses on minimizing inter-ES transmission latency for vision-decoding disaggregation while maintaining draft quality. Third, to optimize SD efficiency on consumer-grade ESs, EdgeSD employs an adaptive and scalable token tree structure solved using a parallel delta-stepping algorithm. This structure maximizes the number of accepted tokens under strict edge latency constraints. Extensive experiments on six multimodal datasets and five benchmarks with various MLLM pairs demonstrate that EdgeSD achieves substantial acceleration and throughput gains in edge-cloud collaboration scenarios using a lightweight draft MLLM, achieving 3.04-5.12x speedup compared to baseline methods. Hualong Huang, Wenhan Zhan, Hancong Duan, Kai Peng 0002, Geyong Min, Zijia Zhao, Zitian Zhao, Yalan Ye |
IEEE Trans. Mob. Comput. | 6 |
| 2026 | A Collaborative Caching and Offloading Approach for Vehicular Edge ComputingabstractVehicular Edge Computing (VEC) leverages promising technologies, namely the vehicle-to-vehicle (V2V) computation offloading approach and edge service caching, to address latency-sensitive tasks. The V2V offloading method efficiently harnesses idle resources from neighboring vehicles. Edge service caching facilitates the offloading task through pre-caching pertinent service data. However, formulating an efficient caching mechanism to support V2V offloading poses significant challenges, given the dynamic vehicle environment, varying computational resources, and limited caching resources of Roadside Units (RSUs). This paper introduces a collaborative caching and offloading (CACO) scheme. First, to mitigate resource wastage caused by inter-vehicle communication interruptions, we employ Generative Adversarial Network (GAN) for trajectory prediction. This process generates a relationship matrix, predicting the stability of inter-vehicle link connections to assist in V2V offloading decisions. Second, to circumvent redundant uploads and computations for recurring offloading tasks, we analyze the popularity of historical offloading tasks using the Page-Hinkley test (PHT) technique, caching frequently offloaded tasks to reduce the processing latency of offloading tasks. Subsequently, a matching scheme for caching and offloading contents is devised. Finally, the Deep Reinforcement Learning (DRL) algorithm is employed to train the offloading strategy. Results from extensive experiments substantiate that CACO attains superior performance in both system computational latency and offloading success rate. Zijia Zhao, Liang Zhao 0004, Lexi Xu, Na Lin 0001, Zhiyuan Tan 0001 |
IEEE Trans. Sustain. Comput. | 1 |
| 2025 | Efficient Motion-Aware Video MLLMabstractMost current video MLLMs rely on uniform frame sampling and image-level encoders, resulting in inefficient data processing and limited motion awareness. To address these challenges, we introduce EMA, an Efficient Motion-Aware video MLLM that utilizes compressed video structures as inputs. We propose a motion-aware GOP (Group of Pictures) encoder that fuses spatial and motion information within a GOP unit in the compressed video stream, generating compact, informative visual tokens. By integrating fewer but denser RGB frames with more but sparser motion vectors in this native slow-fast input architecture, our approach reduces redundancy and enhances motion representation. Additionally, we introduce MotionBench, a benchmark for evaluating motion understanding across four motion types: linear, curved, rotational, and contact-based. Experimental results show that EMA achieves state-of-the-art performance on both MotionBench and popular video question answering benchmarks, while reducing inference costs. Moreover, EMA demonstrates strong scalability, as evidenced by its competitive performance on long video understanding benchmarks. Zijia Zhao, Yuqi Huo, Tongtian Yue, Longteng Guo, Haoyu Lu, Bingning Wang, Weipeng Chen, Jing Liu 0001 |
CVPR | 1 |
| 2025 | Exploring the Design Space of Visual Context Representation in Video MLLMsabstractVideo Multimodal Large Language Models~(MLLMs) have shown remarkable capability of understanding the video semantics on various downstream tasks. Despite the advancements, there is still a lack of systematic research on visual context representation, which refers to the scheme to select frames from a video and further select the tokens from a frame. In this paper, we explore the design space for visual context representation, and aim to improve the performance of video MLLMs by finding more effective representation schemes. Firstly, we formulate the task of visual context representation as a constrained optimization problem, and model the language modeling loss as a function of the number of frames and the number of embeddings (or tokens) per frame, given the maximum visual context window size. Then, we explore the scaling effects in frame selection and token selection respectively, and fit the corresponding function curve by conducting extensive empirical experiments. We examine the effectiveness of typical selection strategies and present empirical findings to determine the two factors. Furthermore, we study the joint effect of frame selection and token selection, and derive the optimal formula for determining the two factors. We demonstrate that the derived optimal settings show alignment with the best-performed results of empirical experiments. The data and code are available at: https://github.com/RUCAIBox/Opt-Visor. Yifan Du 0002, Yuqi Huo, Kun Zhou 0002, Zijia Zhao, Haoyu Lu, Wayne Xin Zhao, Bingning Wang, Weipeng Chen, Ji-Rong Wen |
ICLR | 4 |
| 2025 | Needle In A Video Haystack: A Scalable Synthetic Evaluator for Video MLLMsabstractVideo understanding is a crucial next step for multimodal large language models (MLLMs).
Various benchmarks are introduced for better evaluating the MLLMs.
Nevertheless, current video benchmarks are still inefficient for evaluating video models during iterative development due to the high cost of constructing datasets and the difficulty in isolating specific skills.
In this paper, we propose VideoNIAH (Video Needle in A Haystack), a benchmark construction framework through synthetic video generation.
VideoNIAH decouples video content from their query-responses by inserting unrelated visual 'needles' into original videos.
The framework automates the generation of query-response pairs using predefined rules, minimizing manual labor. The queries focus on specific aspects of video understanding, enabling more skill-specific evaluations. The separation between video content and the queries also allow for increased video variety and evaluations across different lengths.
Utilizing VideoNIAH, we compile a video benchmark, VNBench, which includes tasks such as retrieval, ordering, and counting to evaluate three key aspects of video understanding: temporal perception, chronological ordering, and spatio-temporal coherence. We conduct a comprehensive evaluation of both proprietary and open-source models, uncovering significant differences in their video understanding capabilities across various tasks. Additionally, we perform an in-depth analysis of the test results and model configurations. Based on these findings, we provide some advice for improving video MLLM training, offering valuable insights to guide future research and model development. Zijia Zhao, Haoyu Lu, Yuqi Huo, Yifan Du 0002, Tongtian Yue, Longteng Guo, Bingning Wang, Weipeng Chen, Jing Liu 0001 |
ICLR | 1 |
| 2025 | ChatSearch: A dataset and a generative retrieval model for general conversational image retrieval
Zijia Zhao, Longteng Guo, Tongtian Yue, Erdong Hu, Shuai Shao 0005, Zehuan Yuan, Jing Liu 0001 |
Pattern Recognit. | 1 |
| 2025 | Dynamic Caching Dependency-Aware Task Offloading in Mobile Edge ComputingabstractMobile Edge Computing (MEC) is a distributed computing paradigm that provides computing capabilities at the periphery of mobile cellular networks. This architecture empowers Mobile Users (MUs) to offload computation-intensive applications to large-scale computing nodes near the edge side, reducing application latency for MUs. The resource allocation and task offloading in MEC has been widely studied. However, the burgeoning complexity inherent to modern applications, often represented as Directed Acyclic Graphs (DAGs) comprising a multitude of subtasks with interdependencies, poses huge challenges for application offloading and resource allocation. Meanwhile, previous work has neglected the impact of edge caching on the offloading execution of dependent tasks. Therefore, this paper introduces a novel dynamiccaching dependency-aware taskoffloading (CachOf) scheme. First, to effectively enhance the rationality of cache and computing resource allocation, we develop a subtask priority computation scheme based on DAG dependencies. This scheme includes the execution sequence priority of subtasks on a single MU and the offloading sequence priority of subtasks from multiple MUs. Second, a dynamic caching scheme, designed to cater to dependent tasks, is proposed. This caching approach can not only assist offloading decisions, but also contribute to load balancing by harmonizing caching resources among edge servers. Finally, based on the task prioritization results and caching results, this paper presents a Deep Reinforcement Learning (DRL)-based offloading scheme to judiciously allocate resources and improve the execution efficiency of applications. Extensive simulation experiments demonstrate that CachOf outperforms other baseline schemes, achieving improved execution efficiency for applications. Liang Zhao 0004, Zijia Zhao, Ammar Hawbani, Zhi Liu 0002, Zhiyuan Tan 0001, Keping Yu |
IEEE Trans. Computers | 2 |
| 2024 | OneDiff: A Generalist Model for Image Difference Captioning
Erdong Hu, Longteng Guo, Tongtian Yue, Zijia Zhao, Shuning Xue, Jing Liu 0001 |
ACCV (3) | 4 |
| 2024 | SC- Tune: Unleashing Self-Consistent Referential Comprehension in Large Vision Language ModelsabstractRecent trends in Large Vision Language Models (LVLMs) research have been increasingly focusing on ad-vancing beyond general image understanding towards more nuanced, object-level referential comprehension. In this paper, we present and delve into the self-consistency ca-pability of LVLMs, a crucial aspect that reflects the mod-els' ability to both generate informative captions for spe-cific objects and subsequently utilize these captions to ac-curately re-identify the objects in a closed-loop process. This capability significantly mirrors the precision and reli-ability of fine- grained visual-language understanding. Our findings reveal that the self-consistency level of existing LVLMs falls short of expectations, posing limitations on their practical applicability and potential. To address this gap, we introduce a novel fine-tuning paradigm named Self-Consistency Tuning (SC-Tune). It features the syn-ergistic learning of a cyclic describer-locator system. This paradigm is not only data-efficient but also exhibits gener-alizability across multiple LVLMs. Through extensive ex-periments, we demonstrate that SC- Tune significantly ele-vates performance across a spectrum of object-level vision-language benchmarks and maintains competitive or im-proved performance on image-level vision-language bench-marks. Both our model and code will be publicly available at https://github.com/ivattyue/SC-Tune. Tongtian Yue, Jie Cheng 0009, Longteng Guo, Xingyuan Dai, Zijia Zhao, Xingjian He, Gang Xiong 0001, Jing Liu 0001 |
CVPR | 5 |
| 2024 | Collaborative Training of Tiny-Large Vision Language ModelsabstractRecently, large vision language models (LVLMs) have advanced AI by integrating visual and linguistic data for tasks like visual conversation, image captioning, and visual question answering. Current LVLM research either scales up model size for performance or reduces parameters for limited computational resources. We believe both large and tiny models have unique strengths and that collaborative training yields better results than independent training. We propose Collaborative Training of Tiny-Large Vision Language Models (CTVLMs), a framework connecting large and tiny models via a projection layer and leveraging a synergistic training strategy. Our framework improves training efficiency by strengthening the interconnection between large and tiny models. Using the parameter efficiency of tiny models, we effectively align image-text features, then apply knowledge distillation to help large models better align cross-modal information. During fine-tuning, the large model's extensive knowledge enhances tiny model's performance. This collaborative approach allows models to adapt to various computational resources and outperforms existing methods in vision-language tasks. Shichen Lu, Longteng Guo, Wenxuan Wang 0002, Zijia Zhao, Tongtian Yue, Jing Liu 0001, Si Liu 0001 |
ACM Multimedia | 4 |
| 2023 | VAST: A Vision-Audio-Subtitle-Text Omni-Modality Foundation Model and DatasetabstractVision and text have been fully explored in contemporary video-text foundational models, while other modalities such as audio and subtitles in videos have not received sufficient attention. In this paper, we resort to establish connections between multi-modality video tracks, including Vision, Audio, and Subtitle, and Text by exploring an automatically generated large-scale omni-modality video caption dataset called VAST-27M. Specifically, we first collect 27 million open-domain video clips and separately train a vision and an audio captioner to generate vision and audio captions. Then, we employ an off-the-shelf Large Language Model (LLM) to integrate the generated captions, together with subtitles and instructional prompts into omni-modality captions. Based on the proposed VAST-27M dataset, we train an omni-modality video-text foundational model named VAST, which can perceive and process vision, audio, and subtitle modalities from video, and better support various tasks including vision-text, audio-text, and multi-modal video-text tasks (retrieval, captioning and QA). Extensive experiments have been conducted to demonstrate the effectiveness of our proposed VAST-27M corpus and VAST foundation model. VAST achieves 22 new state-of-the-art results on various cross-modality benchmarks. Handong Li, Qunbo Wang, Zijia Zhao, Mingzhen Sun, Jing Liu 0001 |
NeurIPS | 4 |
| 2023 | MAMO: Fine-Grained Vision-Language Representations Learning with Masked Multimodal ModelingabstractMultimodal representation learning has shown promising improvements on various vision-language tasks (e.g., image-text retrieval, visual question answering, etc) and has significantly advanced the development of multimedia information systems. Most existing methods excel at building global-level alignment between vision and language while lacking effective fine-grained image-text interaction. In this paper, we propose a jointly masked multimodal modeling method to learn fine-grained multimodal representations. Our method performs joint masking on image-text input and integrates both implicit and explicit targets for the masked signals to recover. The implicit target provides a unified and debiased objective for vision and language, where the model predicts latent multimodal representations of the unmasked input. The explicit target further enriches the multimodal representations by recovering high-level and semantically meaningful information: momentum visual features of image patches and concepts of word tokens. Through such a masked modeling process, our model not only learns fine-grained multimodal interaction, but also avoids the semantic gap between high-level representations and low-or mid-level prediction targets (e.g., image pixels, discrete vision tokens), thus producing semantically rich multimodal representations that perform well on both zero-shot and fine-tuned settings. Our pre-trained model (named MAMO) achieves state-of-the-art performance on various downstream vision-language tasks, including image-text retrieval, visual question answering, visual reasoning, and weakly-supervised visual grounding. Zijia Zhao, Longteng Guo, Xingjian He, Shuai Shao 0005, Zehuan Yuan, Jing Liu 0001 |
SIGIR | 1 |
| 2023 | A Digital Twin-Assisted Intelligent Partial Offloading Approach for Vehicular Edge ComputingabstractVehicle Edge Computing (VEC) is a promising paradigm that exposes Mobile Edge Computing (MEC) to road scenarios. In VEC, task offloading can enable vehicles to offload the computing tasks to nearby Roadside Units (RSUs) that deploy computing capabilities. However, the highly dynamic network topology, strict low-delay constraints, and massive data of tasks of VEC pose significant challenges for implementing efficient offloading. Digital Twin-based VEC is emerging as a promising solution that enables real-time monitoring of the state of the VEC network through mapping and interaction between the physical and virtual worlds, thus assisting in making sound offload decisions in the physical world. Thus, this paper proposes an intelligent partial offloading scheme, namely, Digital Twin-Assisted Intelligent Partial Offloading (IGNITE). First, to find the optimal offloading space in advance, we combine the improved clustering algorithm with the Digital Twin (DT) technique, in which unreasonable decisions can be avoided by reducing the size of the decision space. Second, to reduce the overall cost of the system, Deep Reinforcement Learning (DRL) algorithm is employed to train the offloading strategy, allowing for automatic optimization of computational delay and vehicle service price. To improve the efficiency of cooperation between digital and physical spaces, a feedback mechanism is established. It can adjust the parameters of the clustering algorithm based on the final offloading results in this clustering. To the best of our knowledge, this is the first study on DT-assisted vehicle offloading that proposes a feedback mechanism, forming a complete closed loop as prediction-offloading-feedback. Extensive experiments demonstrate that IGNITE has significant advantages in terms of total system computational cost, total computational delay, and offloading success rate compared with its counterparts. Liang Zhao 0004, Zijia Zhao, Enchao Zhang, Ammar Hawbani, Ahmed Yassin Al-Dubai, Zhiyuan Tan 0001, Amir Hussain 0001 |
IEEE J. Sel. Areas Commun. | 2 |
| 2021 | MM21 Pre-training for Video Understanding Challenge: Video Captioning with Pretraining TechniquesabstractThe quality of video representation directly decides the performance of video related tasks, for both understanding and generation. In this paper, we propose single-modality pretrained feature fusion technique which is composed of reasonable multi-view feature extraction method and designed multi-modality feature fusion strategy. We conduct comprehensive ablation studies on MSR-VTT dataset to demonstrate the effectiveness of proposed method and it surpasses the state-of-the-art methods on both MSR-VTT and VATEX datasets. We further propose the multi-modality pretrained model finetuning technique and dataset augmentation scheme to improve the model's generalization capability. Based on these two proposed pretraining techniques and dataset augmentation scheme, we win the first place in the video captioning track of the MM21 pretraining for video understanding challenge. Dongze Hao, Jiawei Liu 0001, Zijia Zhao, Longteng Guo, Jing Liu 0001 |
ACM Multimedia | 6 |