EDBT 2026 Demo / reviewers in the wild / expert
Yifei Cao
dblp:276/8120
· DBLP profile ↗
13ranked-venue papers
2as first author
13since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 7 · 1 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | EAGLE: Episodic Appearance- and Geometry-aware Memory for Unified 2D-3D Visual Query Localization in Egocentric VisionabstractEgocentric visual query localization is vital for embodied AI and VR/AR, yet remains challenging due to camera motion, viewpoint changes, and appearance variations. We present EAGLE, a novel framework that leverages episodic appearance- and geometry-aware memory to achieve unified 2D-3D visual query localization in egocentric vision. Inspired by avian memory consolidation, EAGLE synergistically integrates segmentation guided by an appearance-aware meta-learning memory (AMM), with tracking driven by a geometry-aware localization memory (GLM). This memory consolidation mechanism, through structured appearance and geometry memory banks, stores high-confidence retrieval samples, effectively supporting both long- and short-term modeling of target appearance variations. This enables precise contour delineation with robust spatial discrimination, leading to significantly improved retrieval accuracy. Furthermore, by integrating the VQL-2D output with a visual geometry grounded Transformer (VGGT), we achieve a efficient unification of 2D and 3D tasks, enabling rapid and accurate back-projection into 3D space. Our method achieves state-of-the-art performance on the Ego4D-VQ benchmark. Yifei Cao, Yu Liu 0035, Guolong Wang 0001, Kai Wang 0057, Xianjie Zhang, Jizhe Yu, Xun Tu 0001 |
AAAI | 1 |
| 2026 | What Makes a Good Speech Tokenizer for LLM-Centric Speech Generation? A Systematic StudyabstractSpeech-language models (SLMs) offer a promising path toward unifying speech and text understanding and generation. However, challenges remain in achieving effective cross-modal alignment and high-quality speech generation. In this work, we systematically investigate the role of speech tokenizer designs in LLM-centric SLMs, augmented by speech heads and speaker modeling. We compare coupled, semi-decoupled, and fully decoupled speech tokenizers under a fair SLM framework and find that decoupled tokenization significantly improves alignment and synthesis quality. To address the information density mismatch between speech and text, we introduce multi-token prediction (MTP) into SLMs, enabling each hidden state to decode multiple speech tokens. This leads to up to 12× faster decoding and a substantial drop in word error rate (from 6.07 to 3.01). Furthermore, we propose a speaker-aware generation paradigm and introduce RoleTriviaQA, a large-scale role-playing knowledge QA benchmark with diverse speaker identities. Experiments demonstrate that our methods enhance both knowledge understanding and speaker consistency. Xiaoran Fan, Yangfan Gao, Jingfei Xiong, Hang Yan 0001, Yifei Cao, Zhihao Zhang 0002, Zhiheng Xi, Yuhao Zhou 0005, Senjie Jin, Changhao Jiang, Junjie Ye 0005, Ming Zhang 0030, Zhenhua Han, Yunke Zhang, Demei Yan, Shaokang Dong, Tao Gui |
AAAI | 6 |
| 2026 | Beyond Scaling: Measuring and Predicting the Upper Bound of Knowledge Retention in Language Model Pre-TrainingabstractChanghao Jiang, Ming Zhang, Yifei Cao, Junjie Ye, Xiaoran Fan, Shihan Dou, Zhiheng Xi, Jiajun Sun, Yi Dong, Yujiong Shen, Jingqi Tong, Baoyu Fan, Tao Gui, Qi Zhang, Xuanjing Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Changhao Jiang, Ming Zhang 0030, Yifei Cao, Junjie Ye 0005, Xiaoran Fan, Shihan Dou, Zhiheng Xi, Yujiong Shen, Jingqi Tong, Baoyu Fan, Tao Gui, Qi Zhang 0001, Xuanjing Huang 0001 |
ACL (1) | 3 |
| 2026 | BiOVQL: Brain-inspired One-stage Egocentric Visual Query LocalizationabstractVisual query localization (VQL) is pivotal for constructing episodic memory from egocentric videos. However, current methods often rely on computationally intensive attention mechanisms and rigid one-shot regression, which inherently struggle to model uncertainty and exhibit limited adaptability to degraded query appearances. This contrasts sharply with the human brain’s selective encoding and iterative hypothesis verification processes for episodic memory. Inspired by the human brain’s ability to selectively filter irrelevant information and reconstruct vague memory fragments through generative inference, we propose BiOVQL, a brain-inspired one-stage VQL framework. First, inspired by the hippocampus’s selective retention mechanism, we propose the Hippocampus-like Query-guided Spatio-Temporal Compression (HQSTC) module. By leveraging a selective state space model, visual queries are treated as neuromodulators, dynamically gating the video stream to maintain a compact latent state. This guides the model to consistently focus on query-relevant visual cues, enabling query-conditioned feature compression and efficient spatio-temporal memory encoding. Second, inspired by the prefrontal cortex’s re-localization mechanisms, we propose the Prefrontal-like Generative Refinement Localization (PGRL) module. We leverage a diffusion model to reconstruct the localization process as iterative denoising from noise to certainty, which aligns well with the human visual system’s coarse-to-fine perceptual reasoning. This enhances the model’s robustness in handling spatial ambiguities and achieving precise spatio-temporal retrieval. We conducted extensive experiments on the Ego4D-VQ benchmark, demonstrating that BiOVQL achieves state-of-the-art performance with comparable computational efficiency, thus offering an efficient and brain-inspired paradigm for VQL. Yifei Cao, Guolong Wang 0001, Mingliang Hou, Jizhe Yu, Xianjie Zhang, Xiya Bu, Zhizhen Li, Yu Liu 0035 |
ICMR | 1 |
| 2026 | TrackNetV6: A Unified Framework for Lightweight and Robust Fast-Moving Tiny Ball TrackingabstractAlthough vision-based tiny ball tracking has advanced in specific sports, existing methods remain heavily coupled to domain-specific distributions, severely constraining cross-domain generalization. Concurrently, lightweight designs sacrifice representational capacity, while high-performance models incur prohibitive computational costs. To address these challenges, we propose TrackNetV6, a unified fast-moving tiny ball tracking framework that reconciles efficiency with accuracy. Central to our framework is a novel and compact decoding paradigm rooted in the Linear Multistep Method (LMM), designed to supersede conventional single-step feature fusion. This paradigm orchestrates two core components: a Cross-Scale Semantic Consensus Predictor (CSCP) that distills multi-scale features into semantic-correlation location priors, and a Prior-guided Context Corrector (PCC) that injects these priors into current-scale mappings for stable refinement. By iteratively alternating between these components, the model progressively strengthens feature representation for precise tracking. Furthermore, we introduce a Direction-aware Dynamic Fusion (DDF) module as the bottleneck layer, which explicitly models the direction-sensitive feature relationships of the fast-moving ball by synergizing the dynamic interaction between wavelet-based high-frequency cues and deep semantics. Extensive experiments across badminton, table tennis, and tennis benchmarks demonstrate that TrackNetV6 not only sets a new state-of-the-art (SOTA) but also delivers exceptional real-time inference at 183 FPS. Code will be available at https://github.com/Gi-gigi/TrackNetV6. Jizhe Yu, Xiya Bu, Yu Liu 0035, Kaiping Xu, Yifei Cao, Zhizhen Li |
ICMR | 5 |
| 2026 | SUAD: semantic understanding and attribute discrimination for visual grounding
Xiya Bu, Jizhe Yu, Yu Liu 0035, Kaiping Xu, Guolong Wang 0001, Yifei Cao |
Expert Syst. Appl. | 6 |
| 2026 | Two birds with one stone: Query-dependent moment retrieval in muted video or audio via inter-token interactions
Guolong Wang 0001, Xun Tu 0001, Sutian Hou, Yifei Cao, Yu Liu 0035 |
Inf. Sci. | 4 |
| 2025 | Graph neural networks adversarial attacks based on node gradient and importance score
Yifei Sun 0004, Jiale Ju, Shi Cheng 0002, Yifei Cao, Wenya Shi |
Inf. Sci. | 5 |
| 2024 | Global-Guided Weighted Enhancement for Salient Object Detection
Jizhe Yu, Yu Liu 0035, Hongkui Wei, Kaiping Xu, Yifei Cao, Jiangquan Li |
ICANN (2) | 5 |
| 2024 | Towards Highly Effective Moving Tiny Ball Tracking via Vision Transformer
Jizhe Yu, Yu Liu 0035, Hongkui Wei, Kaiping Xu, Yifei Cao, Jiangquan Li |
ICIC (3) | 5 |
| 2024 | Learning Modality-Complementary and Eliminating-Redundancy Representations with Multi-Task Learning for Multimodal Sentiment AnalysisabstractA crucial issue in multimodal language processing is representation learning. Previous works joint training the multimodal and unimodal tasks to learn the consistency and difference of modality representations. However, due to the lack of cross-modal interaction, the extraction of complementary features between modalities is not sufficient. Moreover, during multimodal fusion, the generated multimodal embeddings may be redundant, and unimodal representations also contain noise information, which negatively influence the final sentiment prediction. To this end, we construct a Modality-Complementary and Eliminating-Redundancy multi-task learning model (MCER), and additionally add a cross-modal task to learn complementary features between two modal pairs through gated transformer. Then use two label generation modules to learn modality-specific and modality-complementary representations. Additionally, we introduce the multimodal information bottleneck (MIB) in both multimodal and unimodal tasks to filter out noise information in unimodal representations as well as learn powerful and sufficient multimodal embeddings that is free of redundancy. Last, we conduct extensive experiments on two popular sentiment analysis benchmarks, MOSI and MOSEI. Experimental results demonstrate that our model significantly outperforms the current strong baselines. Xiaowei Zhao 0003, Xinyu Miao, Xiujuan Xu, Yu Liu 0035, Yifei Cao |
IJCNN | 5 |
| 2024 | GCN-SA: a hybrid recommendation model based on graph convolutional network with embedding splicing layer
Yifei Sun 0004, Shi Cheng 0002, Yifei Cao, Wenya Shi, Jiale Ju, Jihui Yin, Qiaosen Yan, Xinqi Yang, Ziang Wang 0002 |
Neural Comput. Appl. | 4 |
| 2022 | Using Segmentation With Multi-Scale Selective Kernel for Visual Object TrackingabstractGeneric visual object tracking is challenging due to various difficulties, e.g. scale variations and deformations. To solve those problems, we propose a novel multi-scale selective kernel module for tracking, which contains small-scale and large-scale branches to model the target at different scales and attention mechanism to capture the more effective appearance information of the target. In our module, we cascade multiple small-scale convolutional blocks as an equivalent large-scale branch to extract large-scale features of the target effectively. Besides, we present a hybrid strategy for feature selection to extract significant information from features of different scales. Based on the current excellent segmentation tracking framework, we propose a novel tracking network that leverages our module at multiple places in the up-sample phase to construct a more accurate and robust appearance model. Extensive experimental results show that our tracker outperforms other state-of-the-art trackers on multiple challenging benchmarks including VOT2018, TrackingNet, DAVIS-2017, and YouTube-VOS-2018 while achieves real-time tracking. Yifei Cao, Shunli Zhang 0005, Beibei Lin, Sicong Zhao |
IEEE Signal Process. Lett. | 2 |