VLDB 2026 Research / reviewers in the wild / expert
Xinyi Wang 0011
dblp:14/7249-11
· DBLP profile ↗
6ranked-venue papers
4as first author
6since 2021 · last 2026
0000-0003-3079-086XORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FGSVQA: Frequency-Guided Short-Form Video Quality Assessment
Xinyi Wang 0011, Angeliki V. Katsenou, Junxiao Shen, David Bull 0001 |
QoMEX | 1 |
| 2026 | CAMP-VQA: Caption-Embedded Multimodal Perception for No-Reference Quality Assessment of Compressed VideoabstractThe prevalence of user-generated content (UGC) on platforms such as YouTube and TikTok has rendered no-reference (NR) perceptual video quality assessment (VQA) vital for optimizing video delivery. Nonetheless, the characteristics of non-professional acquisition and the subsequent transcoding of UGC video on sharing platforms present significant challenges for NR-VQA. Although NR-VQA models attempt to infer mean opinion scores (MOS), their modeling of subjective scores for compressed content remains limited due to the absence of fine-grained perceptual annotations of artifact types. To address these challenges, we propose CAMP-VQA, a novel NR-VQA framework that exploits the semantic understanding capabilities of large vision-language models. Our approach introduces a quality-aware prompting mechanism that integrates video metadata (e.g., resolution, frame rate, bitrate) with key fragments extracted from inter-frame variations to guide the BLIP-2 pretraining approach in generating fine-grained quality captions. A unified architecture has been designed to model perceptual quality across three dimensions: semantic alignment, temporal characteristics, and spatial characteristics. These multimodal features are extracted and fused, then regressed to video quality scores. Extensive experiments on a wide variety of UGC datasets demonstrate that our model consistently outperforms existing NR-VQA methods, achieving improved accuracy without the need for costly manual fine-grained annotations. Our method achieves the best performance in terms of average rank and linear correlation (SRCC: 0.928, PLCC: 0.938) compared to state-of-the-art methods. The source code and trained models, along with a user-friendly demo, are available at: https://github.com/xinyiW915/CAMP-VQA. Xinyi Wang 0011, Angeliki V. Katsenou, Junxiao Shen, David Bull 0001 |
WACV | 1 |
| 2025 | DIVA-VQA: Detecting Inter-Frame Variations in UGC Video QualityabstractThe rapid growth of user-generated (video) content (UGC) has driven increased demand for research on no-reference (NR) perceptual video quality assessment (VQA). NR-VQA is a key component for large-scale video quality monitoring in social media and streaming applications where a pristine reference is not available. This paper proposes a novel NR-VQA model based on spatio-temporal fragmentation driven by inter-frame variations. By leveraging these inter-frame differences, the model progressively analyses quality-sensitive regions at multiple levels: frames, patches, and fragmented frames. It integrates frames, fragmented residuals, and fragmented frames aligned with residuals to effectively capture global and local information. The model extracts both 2D and 3D features in order to characterize these spatio-temporal variations. Experiments conducted on five UGC datasets and against state-of-the-art models ranked our proposed method among the top 2 in terms of average rank correlation (DIVA-VQA-L: 0.898 and DIVA-VQA-B: 0.886). The improved performance is offered at a low runtime complexity, with DIVA-VQA-B ranked top and DIVA-VQA-L third on average compared to the fastest existing NR-VQA method. Code and models are publicly available at: https://github.com/xinyiW915/DIVA-VQA. Xinyi Wang 0011, Angeliki V. Katsenou, David Bull 0001 |
ICIP | 1 |
| 2025 | Guiding WaveMamba with Frequency Maps for Image Debanding
Xinyi Wang 0011, Smaranda Tasmoc, Nantheera Anantrasirichai, Angeliki V. Katsenou |
PCS | 1 |
| 2024 | Rate-Quality or Energy-Quality Pareto Fronts for Adaptive Video Streaming?abstractAdaptive video streaming is a key enabler for optimising the delivery of offline encoded video content. The research focus to date has been on optimisation, based solely on rate-quality curves. This paper adds an additional dimension, the energy expenditure, and explores construction of bitrate ladders based on decoding energy-quality curves rather than the conventional rate-quality curves. Pareto fronts are extracted from the rate-quality and energy-quality spaces to select optimal points. Bitrate ladders are constructed from these points using conventional rate-based rules together with a novel qualitybased approach. Evaluation on a subset of YouTube-UGC videos encoded with x. 265 shows that the energy-quality ladders reduce energy requirements by $28-31 \%$ on average at the cost of slightly higher bitrates. The results indicate that optimising based on energy-quality curves rather than rate-quality curves and using quality levels to create the rungs could potentially improve energy efficiency for a comparable quality of experience. Angeliki V. Katsenou, Xinyi Wang 0011, Daniel Schien, David Bull 0001 |
ICIP | 2 |
| 2024 | Comparative Study of Hardware and Software Power Measurements in Video CompressionabstractThe environmental impact of video streaming services has been discussed as part of the strategies towards sustainable information and communication technologies. A first step towards that is the energy profiling and assessment of energy consumption of existing video technologies. This paper presents a comprehensive study of power measurement techniques for video encoding and decoding that is comparing the use of hardware and software power meters. An experimental methodology to ensure reliability of measurements is introduced. Key findings demonstrate the high correlation of hardware and software based energy measurements for the case of two video codecs across different spatial and temporal resolutions at a lower computational overhead. Angeliki V. Katsenou, Xinyi Wang 0011, Daniel Schien, David Bull 0001 |
PCS | 2 |