VLDB 2026 Research / reviewers in the wild / expert
Yiyun Zhou
dblp:220/9698
· DBLP profile ↗
11ranked-venue papers
4as first author
9since 2021 · last 2026
0009-0001-5801-8540ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Computer networks · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | AccKV: Towards Efficient Audio-Video LLMs Inference via Adaptive-Focusing and Cross-Calibration KV Cache OptimizationabstractRecent advancements in Audio-Video Large Language Models (AV-LLMs) have enhanced their capabilities in tasks like audio-visual question answering and multimodal dialog systems. Video and audio introduce an extended temporal dimension, resulting in a larger key-value (KV) cache compared to static image embedding. A naive optimization strategy is to selectively focus on and retain KV caches of audio or video based on task. However, in the experiment, we observed that the attention of AV-LLMs to various modalities in the high layers is not strictly dependent on the task. In higher layers, the attention of AV-LLMs shifts more towards the video modality. In addition, we also found that directly integrating temporal KV of audio and spatial-temporal KV of video may lead to information confusion and significant performance degradation of AV-LLMs. If audio and video are processed indiscriminately, it may also lead to excessive compression or reservation of a certain modality, thereby disrupting the alignment between modalities. To address these challenges, we propose AccKV, an Adaptive-Focusing and Cross-Calibration KV cache optimization framework designed specifically for efficient AV-LLMs inference. Our method is based on layer adaptive focusing technology, selectively focusing on key modalities according to the characteristics of different layers, and enhances the recognition of heavy hitter tokens through attention redistribution. In addition, we propose a Cross-Calibration technique that first integrates inefficient KV caches within the audio and video modalities, and then aligns low-priority modality with high-priority modality to selectively evict KV cache of low-priority modality. The experimental results show that AccKV can significantly improve the computational efficiency of AV-LLMs while maintaining accuracy. Zhonghua Jiang 0006, Kunxi Li, Keting Yin, Yiyun Zhou, Zhaode Wang, Chengfei Lv, Shengyu Zhang 0001 |
AAAI | 5 |
| 2026 | Collaborative Representation Learning for Alignment of Tactile, Language, and Vision ModalitiesabstractTactile sensing offers rich and complementary information to vision and language, enabling robots to perceive fine-grained object properties. However, existing tactile sensors lack standardization, leading to redundant features that hinder cross-sensor generalization. Moreover, existing methods fail to fully integrate the intermediate communication among tactile, language, and vision modalities. To address this, we propose TLV-CoRe, a CLIP-based Tactile-Language-Vision Collaborative Representation learning method. TLV-CoRe introduces a Sensor-Aware Modulator to unify tactile features across different sensors and employs tactile-irrelevant decoupled learning to disentangle irrelevant tactile features. Additionally, a Unified Bridging Adapter is introduced to enhance tri-modal interaction within the shared representation space. To fairly evaluate the effectiveness of tactile models, we further propose the RSS evaluation framework, focusing on Robustness, Synergy, and Stability across different methods. Experimental results demonstrate that TLV-CoRe significantly improves sensor-agnostic representation learning and cross-modal alignment, offering a new direction for multimodal tactile representation. Yiyun Zhou, Mingjing Xu, Jingwei Shi, Quanjiang Li |
AAAI | 1 |
| 2025 | Contrastive Cross-Course Knowledge Tracing via Concept Graph Guided Knowledge TransferabstractKnowledge tracing (KT) aims to predict learners' future performance based on historical learning interactions. However, existing KT models predominantly focus on data from a single course, limiting their ability to capture a comprehensive understanding of learners' knowledge states. In this paper, we propose TransKT, a contrastive cross-course knowledge tracing method that leverages concept graph guided knowledge transfer to model the relationships between learning behaviors across different courses, thereby enhancing knowledge state estimation. Specifically, TransKT constructs a cross-course concept graph by leveraging zero-shot Large Language Model (LLM) prompts to establish implicit links between related concepts across different courses. This graph serves as the foundation for knowledge transfer, enabling the model to integrate and enhance the semantic features of learners' interactions across courses. Furthermore, TransKT includes an LLM-to-LM pipeline for incorporating summarized semantic features, which significantly improves the performance of Graph Convolutional Networks (GCNs) used for knowledge transfer. Additionally, TransKT employs a contrastive objective that aligns single-course and cross-course knowledge states, thereby refining the model's ability to provide a more robust and accurate representation of learners' overall knowledge states. Our code and datasets are available at https://github.com/DQYZHWK/TransKT/. Wenkang Han, Liya Hu, Zhenlong Dai, Yiyun Zhou, Mengze Li 0001, Chang Yao 0001, Jingyuan Chen 0003 |
IJCAI | 5 |
| 2025 | Cuff-KT: Tackling Learners' Real-time Learning Pattern Adjustment via Tuning-Free Knowledge State Guided Model UpdatingabstractKnowledge Tracing (KT) is a core component of Intelligent Tutoring Systems, modeling learners' knowledge state to predict future performance and provide personalized learning support. Traditional KT models assume that learners' learning abilities remain relatively stable over short periods or change in predictable ways based on prior performance. However, in reality, learners' abilities change irregularly due to factors like cognitive fatigue, motivation, and external stress--a task introduced, which we refer to as Real-time Learning Pattern Adjustment (RLPA). Existing KT models, when faced with RLPA, lack sufficient adaptability, because they fail to timely account for the dynamic nature of different learners' evolving learning patterns. Current strategies for enhancing adaptability rely on retraining, which leads to significant overfitting and high time overhead issues. To address this, we propose Cuff-KT, comprising a controller and a generator. The controller assigns value scores to learners, while the generator generates personalized parameters for selected learners. Cuff-KT controllably adapts to data changes fast and flexibly without fine-tuning. Experiments on five datasets from different subjects demonstrate that Cuff-KT significantly improves the performance of five KT models with different structures under intra- and inter-learner shifts, with an average relative increase in AUC of 10% and 4%, respectively, at a negligible time cost, effectively tackling RLPA task. Our code and datasets are fully available at https://github.com/zyy-2001/Cuff-KT. Yiyun Zhou, Zheqi Lv, Shengyu Zhang 0001, Jingyuan Chen 0003 |
KDD (2) | 1 |
| 2025 | Show and Polish: Reference-Guided Identity Preservation in Face Video RestorationabstractFace Video Restoration (FVR) aims to reconstruct high-quality face videos from degraded input. Traditional methods struggle to preserve fine-grained, identity-specific features when degradation is severe, often producing average-looking faces that lack individual characteristics. To address these challenges, we introduce IP-FVR, a novel method that leverages a high-quality reference face image as a visual prompt to provide identity conditioning during the denoising process. IP-FVR incorporates semantically rich identity information from the reference image using decoupled cross-attention mechanisms, ensuring detailed and identity consistent results. For intra-clip identity drift (within 24 frames), we introduce an identity-preserving feedback learning method that combines cosine similarity-based reward signals with suffix-weighted temporal aggregation. This approach effectively minimizes drift within sequences of frames. For inter-clip identity drift, we develop an exponential blending strategy that aligns identities across clips by iteratively blending frames from previous clips during the denoising process. This method ensures consistent identity representation across different clips. Additionally, we enhance the restoration process with a multi-stream negative prompt, guiding the model's attention to relevant facial attributes and minimizing the generation of low-quality or incorrect features. Extensive experiments on both synthetic and real-world datasets demonstrate that IP-FVR outperforms existing methods in both quality and identity preservation, showcasing its substantial potential for practical applications in face video restoration. Our code and datasets are available at https://ip-fvr.github.io/. Wenkang Han, Yiyun Zhou, Shulei Wang, Chang Yao 0001, Jingyuan Chen 0003 |
ACM Multimedia | 3 |
| 2025 | Theory-Driven Label-Specific Representation for Incomplete Multi-View Multi-Label LearningabstractMulti-view multi-label learning typically suffers from dual data incompleteness
due to limitations in feature storage and annotation costs. The interplay of hetero
geneous features, numerous labels, and missing information significantly degrades
model performance. To tackle the complex yet highly practical challenges, we
propose a Theory-Driven Label-Specific Representation (TDLSR) framework.
Through constructing the view-specific sample topology and prototype association
graph, we develop the proximity-aware imputation mechanism, while deriving class
representatives that capture the label correlation semantics. To obtain semantically
distinct view representations, we introduce principles of information shift, inter
action and orthogonality, which promotes the disentanglement of representation
information, and mitigates message distortion and redundancy. Besides, label
semantic-guided feature learning is employed to identify the discriminative shared
and specific representations and refine the label preference across views. Moreover,
we theoretically investigate the characteristics of representation learning and the
generalization performance. Finally, extensive experiments on public datasets and
real-world applications validate the effectiveness of TDLSR. Quanjiang Li, Tingjin Luo, Yiyun Zhou, Chenping Hou |
NeurIPS | 6 |
| 2025 | Revisiting Applicable and Comprehensive Knowledge Tracing in Large-Scale Data
Yiyun Zhou, Wenkang Han |
ECML/PKDD (7) | 1 |
| 2025 | Disentangled Knowledge Tracing for Alleviating Cognitive BiasabstractIn the realm of Intelligent Tutoring System (ITS), the accurate assessment of students' knowledge states through Knowledge Tracing (KT) is crucial for personalized learning. However, due to data bias, i.e., the unbalanced distribution of question groups ( e.g., concepts), conventional KT models are plagued by cognitive bias, which tends to result in cognitive underload for overperformers and cognitive overload for underperformers. More seriously, this bias is amplified with the exercise recommendations by ITS. After delving into the causal relations in the KT models, we identify the main cause as the confounder effect of students' historical correct rate distribution over question groups on the student representation and prediction score. Towards this end, we propose a Disentangled Knowledge Tracing (DisKT) model, which separately models students' familiar and unfamiliar abilities based on causal effects and eliminates the impact of the confounder in student representation within the model. Additionally, to shield the contradictory psychology ( e.g., guessing and mistaking) in the students' biased data, DisKT introduces a contradiction attention mechanism. Furthermore, DisKT enhances the interpretability of the model predictions by integrating a variant of Item Response Theory. Experimental results on 11 benchmarks and 3 synthesized datasets with different bias strengths demonstrate that DisKT significantly alleviates cognitive bias and outperforms 16 baselines in evaluation accuracy. Yiyun Zhou, Zheqi Lv, Shengyu Zhang 0001, Jingyuan Chen 0003 |
WWW | 1 |
| 2024 | Trust in ESG reporting: The intelligent Veri-Green solution for incentivized verificationabstractIn today's corporate environment, ESG (Environmental, Social, and Governance) reports crucially reflect an organization's commitment to sustainability, environmental preservation, and social responsibility. As corporations share these detailed reports, the responsibility to validate and assure adherence to respected ESG benchmarks critically lies with third-party assurance organizations. However, the essential verification process often encounters challenges related to authenticity, credibility, and fairness, underscoring the need for a new solution. The selection of verifiers is a crucial aspect of this process, as their expertise and impartiality directly impact the validity and trustworthiness of the verification. Consequently, “Veri-Green,” an innovative blockchain-based incentive mechanism, has been introduced to elevate the ESG data verification process. Considering potential risks in verification systems, such as reputational damage due to oversight or inadvertent approval of inaccurate data, and data security risks involving the management of sensitive organizational information, the verifier selection process needs to be thoroughly thought out and designed. Through the utilization of advanced machine learning algorithms, potential verification candidates are precisely identified, followed by the deployment of the Vickrey Clarke Groves (VCG) auction mechanism. This approach ensures the strategic selection of verifiers and cultivates an ecosystem marked by truthfulness, rationality, and computational efficiency throughout the ESG data verification process. In this framework, verifiers are not only encouraged but also properly incentivized, developing a more transparent, and equitable verification process, thereby driving the ESG agenda towards a future defined by genuine, impactful corporate responsibility and sustainability. Zhiguo Ma, Yiyun Zhou, Melissa Fan |
Blockchain Res. Appl. | 3 |
| 2018 | LSTM Recurrent Neural Networks for Influenza Trends Prediction
Yiyun Zhou, Yan Wang 0062 |
ISBRA | 3 |
| 2018 | Understanding Data Breach: A Visualization Aspect
Yan Wang 0062, Yiyun Zhou |
WASA | 4 |