VLDB 2026 Research / reviewers in the wild / expert
Ziyi Cao
dblp:201/3273
· DBLP profile ↗
9ranked-venue papers
7as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 5 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 first-author · 6 since 2021Systems, architecture and hardware · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | VBF++: Variational Bayesian Fusion with Context-Aware Priors and Recommendation-Guided Adversarial Refinement for Multimodal Video RecommendationabstractMultimodal video recommendation systems face fundamental challenges in determining optimal fusion strategies across diverse content types and user preferences. Existing methods suffer from two critical limitations: (1) their fusion strategies are guided by context-agnostic priors that ignore the semantic structure of content, assuming the same simple distribution (typically a standard multivariate Gaussian prior) governs optimal fusion for all video types, and (2) their optimization objectives, particularly the Evidence Lower Bound (ELBO), are misaligned with the final recommendation goal, optimizing for feature reconstruction rather than ranking performance. To address these fundamental issues, this work proposes VBF++, a novel framework that introduces context-aware structured priors and recommendation-guided adversarial refinement. First, the method designs context-aware priors that learn cluster-specific distributions based on video semantic categories, replacing uninformative priors with structured, content-aware prior distributions. Second, it introduces a Recommendation-Guided Adversarial Refinement (RAR) paradigm that explicitly steers the learning process towards generating recommendation-optimal fusion strategies, resolving the objective misalignment inherent in variational learning. Enhanced with domain-adaptive meta-learning, extensive experiments on three real-world datasets demonstrate consistent improvements of 4.7-8.3 percent in Precision@10 over state-of-the-art methods. Analysis reveals that learned fusion strategies exhibit semantically meaningful patterns, prioritizing visual features for action content, acoustic information for music videos, and textual descriptions for documentary material. Ziyi Cao, Rui Liu 0007, Yong Chen 0008 |
AAAI | 1 |
| 2026 | Sparse Attention Across Multiple-Context KV CacheabstractLarge language models face significant cost challenges in long-sequence inference. To address this, reusing historical Key-Value (KV) Cache for improved inference efficiency has become a mainstream approach. Recent advances further enhance throughput by sparse attention mechanisms to select the most relevant KV Cache, thereby reducing sequence length. However, such techniques are limited to single-context scenarios, where historical KV Cache is computed sequentially with causal-attention dependencies. In retrieval-augmented generation (RAG) scenarios, where retrieved documents as context are unknown beforehand, each document’s KV Cache is computed and stored independently (termed multiple-context KV Cache), lacking cross-attention between contexts. This renders existing methods ineffective. Although prior work partially recomputes multiple-context KV Cache to mitigate accuracy loss from missing cross-attention, it requires retaining all KV Cache throughout, failing to reduce memory overhead. This paper presents SamKV, the first exploration of attention sparsification for multiple-context KV Cache. Specifically, SamKV takes into account the complementary information of other contexts when sparsifying one context, and then locally recomputes the sparsified information. Experiments demonstrate that our method compresses sequence length to 15% without accuracy degradation compared with full-recomputation baselines, significantly boosting throughput in multi-context RAG scenarios. Ziyi Cao, Qingyi Si, Bingquan Liu |
AAAI | 1 |
| 2026 | Dynamic Channel Collaboration Framework for Panoramic Image Enhancement: A Neurobiologically-Inspired ApproachabstractPanoramic images are critical for immersive VR/AR and 6DoF yet degraded by compression artifacts, projection distortion, and uneven sampling, with existing hybrid CNN-Transformer models struggling to reconcile fine details and structural consistency in panoramas; to address this, we propose Dynamic Channel Collaboration (DCC-Former) for panoramic enhancement, inspired by primate vision's hierarchical processing and three strategies: strengthening local feature representation via reparameterization and gating, enhancing global context with adaptive self-attention, and enabling cross-scale aggregation through cascaded multi-scale fusion, aligned with biological vision's ventral-dorsal stream division and fovea-periphery resource allocation to balance detail preservation, global consistency, and computational efficiency-extensive experiments on benchmark datasets demonstrate DCC-Former outperforms SOTA in restoration quality and inference efficiency, providing a practical-efficient paradigm for high-resolution panoramic enhancement. Ziyi Cao, Hongkui Wang, Haibing Yin, Tiansong Li, Jiyong Zhang 0001, Xiaofeng Huang, Xia Wang 0006, Ruiyang Fu |
DCC | 1 |
| 2025 | A Dual Contrastive Learning Framework for Enhanced Multimodal Conversational Emotion RecognitionabstractMultimodal Emotion Recognition in Conversations (MERC) identifies utterance emotions by integrating both contextual and multimodal information from dialogue videos. Existing methods struggle to capture emotion shifts due to label replication and fail to preserve positive independent modality contributions during fusion. To address these issues, we propose a Dual Contrastive Learning Framework (DCLF) that enhances current MERC models without additional data. Specifically, to mitigate label replication effects, we construct context-aware contrastive pairs. Additionally, we assign pseudo-labels to distinguish modality-specific contributions. DCLF works alongside basic models to introduce semantic constraints at the utterance, context, and modality levels. Our experiments on two MERC benchmark datasets demonstrate performance gains of 4.67%-4.98% on IEMOCAP and 5.52%-5.89% on MELD, outperforming state-of-the-art approaches. Perturbation tests further validate DCLF’s ability to reduce label dependence. Additionally, DCLF incorporates emotion-sensitive independent modality features and multimodal fusion representations into final decisions, unlocking the potential contributions of individual modalities. Yunhe Xie, Chengjie Sun, Ziyi Cao, Bingquan Liu, Zhenzhou Ji, Yuanchao Liu, Lili Shan |
COLING | 3 |
| 2025 | PIN: A Prompt-based Implicit Sentiment Analysis Network for ChineseabstractFor sentiment analysis (SA) issue, most current SA models focus on Explicit Sentiment Analysis (ESA), with less attention to Implicit Sentiment Analysis (ISA). Recent ISA models cannot fully consider prior knowledge in Pre-trained Language Model (PLM) for this knowledge intensive task, even introducing complex external knowledge bases. This also leads to limited performance with additional computational resources. To address these problems, we propose a Prompt-based Implicit sentiment analysis Network (PIN) for Chinese ISA, where a topic recognition module is introduced to identify the topic of the review. Then, the topic is embedded in the soft template to predict sentiment based on prompt learning, which can effectively activate PLM knowledge with low computational resources. Experiments conducted on three public datasets demonstrate the effectiveness of our model, as compared with state-of-the-art methods. Meanwhile, we also provide our identified topic as a supplement to the above datasets, forming three new datasets. Kun Bu, Yuanchao Liu, Ziyi Cao |
ICASSP | 4 |
| 2025 | Re3MHQA: Retrieve, Remove, and Return facts in multi-hop QA
Ziyi Cao, Yunhe Xie, Bingquan Liu, Kun Bu |
Expert Syst. Appl. | 1 |
| 2024 | The Attempt on Combining Three Talents by KD with Enhanced Boundary in Co-Salient Object Detection
Ziyi Cao, Shengye Yan |
BMVC | 1 |
| 2024 | TPARN: A Network for Enhancing Synthetic Video Quality After 3D-HEVC Encodingabstract3D-High Efficiency Video Coding (3D-HEVC), as an extension of HEVC in the realm of three-dimensional video, has brought significant coding performance improvements. However, traditional 3D video coding has faced many challenges such as compression distortion in texture and depth videos, as well as non-occlusion issues in Depth Image Based Rendering (DIBR) synthesis, which directly affected the visual quality of synthesized views. A Two-Stream Pyramid Attention Residual Network (TPARN) is proposed to achieve the quality enhancement of synthesized views. First of all, the Global Residual Attention (GRA) module and the Local Pyramid Attention (LPA) module are designed to extract global context information and intricate local texture details, which achieve a comprehensive scene understanding and preserve essential details across different scales. In addition, the Pyramid Attention Module (PAM) and skip connections are utilized to extract multiscale features, promoting seamless interaction among features. Experimental results demonstrate that the proposed method effectively reduces distortion caused by view synthesis, outperforming the latest methods in terms of performance. Ziyi Cao, Tiansong Li, Shaoguo Cui, Kejun Wu, Longwei Zhong, Hongkui Wang, Li Yu 0003 |
ISCAS | 1 |
| 2023 | RPA: Reasoning Path Augmentation in Iterative Retrieving for Multi-Hop QAabstractMulti-hop questions are associated with a series of justifications, and one needs to obtain the answers by following the reasoning path (RP) that orders the justifications adequately. So reasoning path retrieval becomes a critical preliminary stage for multi-hop Question Answering (QA). Within the RP, two fundamental challenges emerge for better performance: (i) what the order of the justifications in the RP should be, and (ii) what if the wrong justification has been in the path. In this paper, we propose Reasoning Path Augmentation (RPA), which uses reasoning path reordering and augmentation to handle the above two challenges, respectively. Reasoning path reordering restructures the reasoning by targeting the easier justification first but difficult one later, in which the difficulty is determined by the overlap between query and justifications since the higher overlap means more lexical relevance and easier searchable. Reasoning path augmentation automatically generates artificial RPs, in which the distracted justifications are inserted to aid the model recover from the wrong justification. We build RPA with a naive pre-trained model and evaluate RPA on the QASC and MultiRC datasets. The evaluation results demonstrate that RPA outperforms previously published reasoning path retrieval methods, showing the effectiveness of the proposed methods. Moreover, we present detailed experiments on how the orders of justifications and the percent of augmented paths affect the question- answering performance, revealing the importance of polishing RPs and the necessity of augmentation. Ziyi Cao, Bingquan Liu, Shaobo Li 0004 |
AAAI | 1 |