EDBT 2026 Demo / reviewers in the wild / expert
Jizhe Yu
dblp:380/7416
· DBLP profile ↗
5ranked-venue papers in the field
3as first author
5since 2021 · last 2026
0009-0008-8365-2317ORCID · verified
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 5 (3 first)
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | BiOVQL: Brain-inspired One-stage Egocentric Visual Query LocalizationabstractVisual query localization (VQL) is pivotal for constructing episodic memory from egocentric videos. However, current methods often rely on computationally intensive attention mechanisms and rigid one-shot regression, which inherently struggle to model uncertainty and exhibit limited adaptability to degraded query appearances. This contrasts sharply with the human brain’s selective encoding and iterative hypothesis verification processes for episodic memory. Inspired by the human brain’s ability to selectively filter irrelevant information and reconstruct vague memory fragments through generative inference, we propose BiOVQL, a brain-inspired one-stage VQL framework. First, inspired by the hippocampus’s selective retention mechanism, we propose the Hippocampus-like Query-guided Spatio-Temporal Compression (HQSTC) module. By leveraging a selective state space model, visual queries are treated as neuromodulators, dynamically gating the video stream to maintain a compact latent state. This guides the model to consistently focus on query-relevant visual cues, enabling query-conditioned feature compression and efficient spatio-temporal memory encoding. Second, inspired by the prefrontal cortex’s re-localization mechanisms, we propose the Prefrontal-like Generative Refinement Localization (PGRL) module. We leverage a diffusion model to reconstruct the localization process as iterative denoising from noise to certainty, which aligns well with the human visual system’s coarse-to-fine perceptual reasoning. This enhances the model’s robustness in handling spatial ambiguities and achieving precise spatio-temporal retrieval. We conducted extensive experiments on the Ego4D-VQ benchmark, demonstrating that BiOVQL achieves state-of-the-art performance with comparable computational efficiency, thus offering an efficient and brain-inspired paradigm for VQL. Yifei Cao, Guolong Wang 0001, Mingliang Hou, Jizhe Yu, Xianjie Zhang, Xiya Bu, Zhizhen Li, Yu Liu 0035 |
ICMR | 4 |
| 2026 | TrackNetV6: A Unified Framework for Lightweight and Robust Fast-Moving Tiny Ball TrackingabstractAlthough vision-based tiny ball tracking has advanced in specific sports, existing methods remain heavily coupled to domain-specific distributions, severely constraining cross-domain generalization. Concurrently, lightweight designs sacrifice representational capacity, while high-performance models incur prohibitive computational costs. To address these challenges, we propose TrackNetV6, a unified fast-moving tiny ball tracking framework that reconciles efficiency with accuracy. Central to our framework is a novel and compact decoding paradigm rooted in the Linear Multistep Method (LMM), designed to supersede conventional single-step feature fusion. This paradigm orchestrates two core components: a Cross-Scale Semantic Consensus Predictor (CSCP) that distills multi-scale features into semantic-correlation location priors, and a Prior-guided Context Corrector (PCC) that injects these priors into current-scale mappings for stable refinement. By iteratively alternating between these components, the model progressively strengthens feature representation for precise tracking. Furthermore, we introduce a Direction-aware Dynamic Fusion (DDF) module as the bottleneck layer, which explicitly models the direction-sensitive feature relationships of the fast-moving ball by synergizing the dynamic interaction between wavelet-based high-frequency cues and deep semantics. Extensive experiments across badminton, table tennis, and tennis benchmarks demonstrate that TrackNetV6 not only sets a new state-of-the-art (SOTA) but also delivers exceptional real-time inference at 183 FPS. Code will be available at https://github.com/Gi-gigi/TrackNetV6. Jizhe Yu, Xiya Bu, Yu Liu 0035, Kaiping Xu, Yifei Cao, Zhizhen Li |
ICMR | 1 |
| 2026 | Locate Core, Refine Path: A Training-Free Closed-Loop Paradigm for Referring Video Object SegmentationabstractReferring video object segmentation (RVOS) aims to dynamically segment a target object across video sequences based on a given natural language expression. Although recent decoupled methods surpass end-to-end models, they suffer from a critical trade-off: relying on fine-tuned large language models for key frame selection incurs high computational costs, whereas employing sparse sampling risks missing optimal reference frames. Furthermore, the subsequent mask propagation lacks self-correction mechanisms, leading to irreversible cumulative errors over time. In this work, we propose LCRP, a training-free framework driven by two novel components. The Hybrid Candidate Sampling (HCS) module significantly improves frame selection by integrating lightweight semantic parsing with CLIP-based filtering. Meanwhile, the Multi-Anchor Refinement (MAR) module utilizes anchor frames as dynamic checkpoints to detect temporal drift and execute localized re-propagation. Extensive experiments indicate that our proposed method achieves state-of-the-art results. Code is available at https://github.com/Crystal535/Locate-Core-Refine-Path. Jizhe Yu, Hao Zhang 0218, Xiya Bu, Yuhang Duan, Xiaoshuai Wu, Yu Liu 0035 |
ICMR | 1 |
| 2025 | Visual Grounding with Feature Enhancement and Language-Aware Attribute Guidance
Xiya Bu, Jizhe Yu, Yu Liu 0035, Kaiping Xu |
ICMR | 2 |
| 2025 | PAP-SAM: Global-Local Prior Adaptive Perception SAM for Co-Salient Object Detection
Jizhe Yu, Xiya Bu, Yu Liu 0035, Kaiping Xu |
ICMR | 1 |