VLDB 2026 Research / reviewers in the wild / expert
Yuyu Liu
dblp:35/490
· DBLP profile ↗
14ranked-venue papers
6as first author
9since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 6 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Computer networks · 2 · 1 first-author · 2 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SCOPE: Spatio-Temporal Collaborative Caching and Proactive Transfer in LEO Satellite Networks
Yuyu Liu, Qian Wu 0001, Zeqi Lai, Hewu Li, Yuanjie Li, Jun Liu 0063 |
APNet | 1 |
| 2026 | COPELEO: Enhancing Low Earth Orbit Satellite CDN through Collaborative Caching
Yuyu Liu, Qian Wu 0001, Zeqi Lai, Hewu Li, Yuanjie Li, Jun Liu 0063 |
IWCMC | 1 |
| 2026 | From Abstract Events to Grounded Cues: Cue-Guided Vision-Language Anomaly DetectionabstractVideo anomaly detection (VAD) must be reliable under large appearance variation and weak supervision, yet provide explanations grounded in human-interpretable evidence. Vision-Language Models (VLMs) often suffer from brittle direct visual-text alignment which is unstable across domains, while deep models are accurate but have limited interpretability. We address this by introducing event-related but more concrete cues as intermediate representations, making the mapping from frames to evidence more stable than directly mapping frames to event labels. Building on this idea, we propose a two-stage cooperative framework: a VLM discovers cues and generates cue-guided pseudo frame-level labels, and a Symbolic Learning Model (SLM) learns from them to produce segment-level cue presence estimates and anomaly scores via cross-modal matching. In inference, cue estimates, SLM scores and video segments are combined by VLM to output final decisions with cue-based rationales. Experiments show competitive performance. Code is available athttps://github.com/AllenYLJiang/Atoms-to-Events-Categorical-Evidence-Composition-for-Video-Anomaly-Detection. Yalong Jiang, Lian Huai, Yuyu Liu, Xingqun Jiang |
IEEE Signal Process. Lett. | 3 |
| 2025 | Spache: Accelerating Ubiquitous Web Browsing via Schedule-Driven Space CachingabstractIn this paper, we perform a systematic study to explore a pivotal problem facing the web community: is current distributed web cache ready for future satellite Internet? First, through a worldwide performance measurement based on the RIPE Atlas platform and Starlink, the largest low-earth orbit (LEO) satellite network (LSN) today, we identify that the uneven deployment of current distributed cache servers, inter-ISP meandering routes and the last-mile congestion on LEO links jointly prevent existing terrestrial web cache from providing low-latency web access for users in emerging LSNs. Second, we propose Spache, a novel web caching system which addresses the limitations of existing ground-only cache by exploiting a bold idea: integrating web cache into LEO satellites to achieve ubiquitous and low-latency web services. Specifically, Spache leverages a key feature of LSNs called communication schedule to efficiently prefetch web contents on satellites, and adopts a schedule-driven partitioning strategy to avoid cache pollution involved by LEO mobility. Finally, we implement a prototype of Spache, and evaluate it based on real-world HTTP traces and data-driven LSN simulation. Extensive evaluations demonstrate that as compared to existing distributed caching solutions, Spache can improve cache hit ratio by 19.8% on average, reduce latency by up to 17.7%, and maintains consistently low web browsing latency for global LSN users. Qi Zhang 0102, Qian Wu 0001, Zeqi Lai, Hewu Li, Yuyu Liu, Yuanjie Li, Jun Liu 0063 |
WWW | 6 |
| 2024 | Decoupled DETR for Few-Shot Object Detection
Zeyu Shangguan, Lian Huai, Yuyu Liu, Xingqun Jiang |
ACCV (8) | 4 |
| 2024 | Valid Information Guidance Network for Compressed Video Quality EnhancementabstractRestoring high-quality videos from the compressed ones is a crucial research topic in video coding. Most existing methods generally take the raw video as the ground truth to guide the reconstruction. We find that the compressed frames contain less texture details than the raw frames, which we called valid information. As shown in Figure 1 , we propose a Valid Information Guidance (VIG) scheme to recover the raw spatio-temporal distribution via making full use of valid information. Specifically, we propose a Truth Guidance Distillation (TGD) strategy to learn to model the spatio-temporal correspondence from the rich valid information contained in the raw frames. We replace ninety percent of the compressed blocks with the raw blocks for pre-training. Furthermore, we propose an efficient Compressed Redundancy Filtering (CRF) network to extract valid information by filtering the compressed frames. Guannan Chen, Zhanfu An, Yuyu Liu |
DCC | 5 |
| 2024 | Multimodal Visual-Semantic Representations Learning for Scene Text RecognitionabstractScene Text Recognition (STR), the critical step in OCR systems, has attracted much attention in computer vision. Recent research on modeling textual semantics with Language Model (LM) has witnessed remarkable progress. However, LM only optimizes the joint probability of the estimated characters generated from the Vision Model (VM) in a single language modality, ignoring the visual-semantic relations in different modalities. Thus, LM-based methods can hardly generalize well to some challenging conditions, in which the text has weak or multiple semantics, arbitrary shape, and so on. To migrate the above issue, in this paper, we propose Multimodal Visual-Semantic Representations Learning for Text Recognition Network (MVSTRN) to reason and combine the multimodal visual-semantic information for accurate Scene Text Recognition. Specifically, our MVSTRN builds a bridge between vision and language through its unified architecture and has the ability to reason visual semantics by guiding the network to reconstruct the original image from the latent text representation, breaking the structural gap between vision and language. Finally, the tailored multimodal Fusion (MMF) module is motivated to combine the multimodal visual and textual semantics from VM and LM to make the final predictions. Extensive experiments demonstrate our MVSTRN achieves state-of-the-art performance on several benchmarks. Xinjian Gao, Ye Pang, Yuyu Liu, Maokun Han, Jun Yu 0001, Wei Wang 0496, Yuanxu Chen |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2022 | DBCAN: Dual-Branch Cross-Attention Network for Scene Text RecognitionabstractScene text recognition, especially irregular text recognition, is a challenging task due to the large variance in text appearance. Although some existing methods have achieved state-of-the-art performance with the attention-based encoder-decoder framework, they always perform poorly on some challenging text such as severely curved, blurred, and incomplete-semantic text. To address these issues, we propose a Dual-Branch Cross-Attention Network (DBCAN). Different from the previous methods heavily relying on semantic information, DBCAN can enhance the position clues and learn semantic relations with two separate branches and fuse them by a tailored Cross-Attention Module (CAM). Furthermore, a Convolution-Based 2D Positional Embedding (CBPE) is introduced to describe the 2D spatial dependencies of characters. Extensive experiments demonstrate our DBCAN is more accurate and robust than the previous methods and achieves state-of-the-art performance on several benchmarks, particularly CUTE (93.4%). Our code is made publicly available at https://github.com/GaoXinJian-USTC/DBCAN. Xinjian Gao, Ye Pang, Yuyu Liu, Jun Yu 0001, Maokun Han, Wei Wang 0496 |
ICME | 3 |
| 2021 | Radar Object Detection Using Data Merging, Enhancement and FusionabstractCompared to visible images, radar images are generally considered to be an active and robust solution, even in adverse driving situations, for object detection. However, the accuracy of radar object detection (ROD) is always poor. Owing to taking full advantage of data merging, enhancement and fusion, this paper proposes an effective ROD system with only radar images as the input. First, an aggregation module is designed to merge the data from all chirps in the same frame. Then, various gaussian noises with different parameters are employed to increase data diversity and reduce over-fitting based on the analysis of training data. Moreover, due to the process of inference with default parameters is not accurate enough, some hyperparameters are changed to increase the accuracy performance. Finally, a combination strategy is adopted to benefit from multi-model fusion. ROD2021 Challenge is supported by ACM ICMR 2021, and our team (ustc-nelslip) ranked 2nd in the test stage of this challenge. Diverse evaluations also verify the superiority of the proposed system. Jun Yu 0001, Xinlong Hao, Xinjian Gao, Yuyu Liu, Peng Chang 0002, Fang Gao 0001, Feng Shuang 0002 |
ICMR | 5 |
| 2010 | Recovery of audio-to-video synchronization through analysis of cross-modality correlation
Yuyu Liu, Yoichi Sato 0001 |
Pattern Recognit. Lett. | 1 |
| 2009 | Visual localization of non-stationary sound sourcesabstractSound source can be visually localized by analyzing the correlation between audio and visual data. To correctly analyze this correlation, the sound source is required to be stationary in a scene to date. We introduce a technique that localizes the non-stationary sound sources to overcome this limitation. The problem is formulated as finding the optimal visual trajectories that best represent the movement of the sound source over the pixels in a spatio-temporal volume. Using a beam search, we search these optimal visual trajectories by maximizing the correlation between the newly introduced audiovisual features of inconsistency. An incremental correlation evaluation with mutual information is developed here, which significantly reduces the computational cost. The correlations computed along the optimal trajectories are finally incorporated into a segmentation technique to localize a sound source region in the first visual frame of the current time window. Experimental results demonstrate the effectiveness of our method. Yuyu Liu, Yoichi Sato 0001 |
ACM Multimedia | 1 |
| 2008 | Recovering audio-to-video synchronization by audiovisual correlation analysisabstractAudio-to-video synchronization (AV-sync) may drift and is difficult to recover without dedicated human effort. In this work, we develop an interactive method to recover the drifted AV-sync by audiovisual correlation analysis. Given a video segment, a user specifies a rough time span during which a person is speaking. Our system first detects a speaker region using face detection. It then does a two-stage search to find the optimum AV-drift that can maximize the average audiovisual correlation inside the speaker region. The correlation is evaluated using quadratic mutual information with kernel density estimation. AV-sync is finally recovered by the detected optimum AV-drift. Experimental results demonstrate the effectiveness of our method. Yuyu Liu, Yoichi Sato 0001 |
ICPR | 1 |
| 2006 | Sigma-delta based clock recovery using on-chip PLL in FPGAabstractA clock and data recovery (CDR) circuit is proposed based on the sigma-delta quantization. The phase of the new CDR circuit is adjusted by a sigma-delta modulated reference clock that increases the stability of the system and can easily interface with PLL cores embedded in FPGAs. The approximate linear model of the proposed CDR is analyzed for SONET/SDH applications to evaluate its performance. The measurement shows that the jitter tolerance meets the ITU-T requirement with a high margin of 0.3UI. The commercial equipment has been developed using a single FPGA chip based on the SDM-CDR Ning Ge 0001, Yuyu Liu, Huazhong Yang, Hui Wang 0004 |
FPT | 2 |
| 2005 | Eye-contact visual communication with virtual view synthesisabstractIn this paper, we propose a new visual communication system where eye contact is possible by using a virtual image. The virtual image is obtained by view synthesis with stereo matching from two real camera views. We developed a region based dynamic programming (DP) approach with improved matching cost, occlusion cost and vertical smoothness constraint. We also proposed a fast view interpolation method. To achieve real time performance, we developed a hardware system. Furthermore, to avoid the reordering problem in the foreground region, a view change approach with disorder detection is adopted. Experimental results demonstrate the validity of our improved DP matching algorithm and eye-contact visual communication system. Yuyu Liu, Keisuke Yamaoka, Akira Nakamura, Yoshiaki Iwai, Ken-ichiro Ooi, Weiguo Wu, Takayuki Yoshigahara |
CCNC | 1 |