Yongmei Zhou

dblp:96/5435 · DBLP profile ↗
← Back
3ranked-venue papers in the field
0as first author
3since 2021 · last 2025
0000-0003-2661-3078ORCID · corroborated

Domains — venue-derived; a paper can count in several

Big Data, Cloud & Distributed Data Systems · 2Information Retrieval & Web Search · 1
YearPublicationVenuePosition
2025 LLM-Driven Effective Knowledge Tracing by Integrating Dual-Channel Difficulty
Jiahui Cen, Jianghao Lin, Dong Zhou 0001, Weixuan Zhong, Aimin Yang 0002, Yongmei Zhou
IEEE Big Data7
2025 Central-Guided Convolutional Dual Attention for Document-Level Event Argument Extraction
Chengdong Lin, Jianghao Lin, Dong Zhou 0001, Yongmei Zhou, Aimin Yang 0002
IEEE Big Data4
2025 DomainDiff: Unified Two-Stage Optimization for Text-Video Retrieval
abstract
The primary challenge in text-video retrieval lies in achieving cross-modal semantic alignment, particularly the discrepancy between the conciseness of textual descriptions, which often fail to fully encapsulate the breadth of video content, and the redundancy in video data, which introduces noise and masks important semantic features. Current methods align text and video by mapping them into a shared feature space. Despite notable advancements, the inherent differences in modality-specific representations create a bottleneck for fixed-point embedding techniques, making models highly sensitive to dataset distribution and hindering their generalization ability. In this paper, we present DomainDiff, a framework that enhances the embedding space through a two-stage process. In the first stage, stochastic domain modeling, we semantically expand text embeddings to explore potential regions aligned with video content. Simultaneously, we filter video segments to reduce redundancy and highlight key frames. In the second stage, the dynamic agent attention diffusion network, we leverage the generative properties of diffusion models to optimize the embedding space by viewing it from a joint probability distribution perspective. An agent attention mechanism dynamically integrates text and video features, ensuring accurate cross-modal alignment. Experimental results demonstrate that DomainDiff significantly improves retrieval performance across five benchmark datasets, with R@1 improvements ranging from 3% to 7.4%. Moreover, DomainDiff outperforms existing methods in handling long videos and complex textual descriptions, showcasing superior semantic robustness and generalization across varying distributions.
Chenxu Wang 0019, Dong Zhou 0001, Jianghao Lin, Yongmei Zhou, Aimin Yang 0002
ICMR4