Yibo Shi

dblp:325/5297 · DBLP profile ↗
← Back
8ranked-venue papers
2as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 1 first-author · 5 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Computer graphics and multimedia
4 papers
Image and video coding · 96% Image and video processing · 4%
Human-computer interaction and pervasive computing
1 paper
Human-AI interaction · 100%
Network and information security
1 paper
Privacy and data protection · 100%
Artificial intelligence
1 paper
Efficient and distributed learning · 100%

Topics — the 7 heaviest of 9, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Image and video coding › video compression
learned video compression
2.032024
Neural Rate Control for Learned Video Compression · ICLR 2024
High Visual-Fidelity Learned Video Compression · ACM Multimedia 2023
AlphaVC: High-Performance and Efficient Learned Video Compression · ECCV (19) 2022
Human-AI interaction
LLM-based agents
1.012026
Mind the Third Eye! Benchmarking Privacy Awareness in MLLM-powered Smartphone Agents · AAAI 2026
Privacy and data protection › privacy perceptions
privacy awareness
1.012026
Mind the Third Eye! Benchmarking Privacy Awareness in MLLM-powered Smartphone Agents · AAAI 2026
Image and video coding
rate control
0.812024
Neural Rate Control for Learned Video Compression · ICLR 2024
Image and video coding
video compression
0.812024
Neural Rate Control for Learned Video Compression · ICLR 2024
Image and video coding › image compression
learned image compression
0.612022
Content-Oriented Learned Image Compression · ECCV (19) 2022
Image and video processing
video restoration
0.212023
High Visual-Fidelity Learned Video Compression · ACM Multimedia 2023

Methods — techniques the papers use, named apart from their topics

benchmark construction · 2.0learned video compression · 1.1entropy coding · 1.1rate-parameter mapping · 0.8neural rate allocation · 0.8periodic compensation loss · 0.7confidence-based feature reconstruction · 0.7
YearPublicationVenuePosition
2026 Mind the Third Eye! Benchmarking Privacy Awareness in MLLM-powered Smartphone Agents
Zhixin Lin, Jungang Li, Shidong Pan, Yibo Shi, Yue Yao 0001, Dongliang Xu
AAAI4
2025 RTV-Bench: Benchmarking MLLM Continuous Perception, Understanding and Reasoning through Real-Time Video
abstract
Multimodal Large Language Models (MLLMs) increasingly excel at perception,understanding, and reasoning. However, current benchmarks inadequately evaluate their ability to perform these tasks continuously in dynamic, real-world environments. To bridge this gap, we introduce RT V-Bench, a fine-grained benchmark for MLLM real-time video analysis. RTV-Bench includes three key principles: (1) Multi-Timestamp Question Answering (MTQA), where answers evolve with scene changes; (2) Hierarchical Question Structure, combining basic and advanced queries; and (3) Multi-dimensional Evaluation, assessing the ability of continuous perception, understanding, and reasoning. RTV-Bench contains 552 diverse videos (167.2 hours) and 4,631 high-quality QA pairs. We evaluated leading MLLMs, including proprietary (GPT-4o, Gemini 2.0), open-source offline (Qwen2.5-VL, VideoLLaMA3), and open-source real-time (VITA-1.5, InternLM-XComposer2.5-OmniLive) models. Experiment results show open-source real-time models largely outperform offline ones but still trail top proprietary models. Our analysis also reveals that larger model size or higher frame sampling rates do not significantly boost RTV-Bench performance, sometimes causing slight decreases.This underscores the need for better model architectures optimized for video stream processing and long sequences to advance real-time video analysis with MLLMs.
Shuhang Xun, Sicheng Tao, Jungang Li, Yibo Shi, Zhixin Lin, Zhanhui Zhu, Hanqian Li, Linghao Zhang, Shikang Wang, Hanbo Zhang, Xuming Hu
NeurIPS4
2025 VCIP 2025 Ultra Low-Bitrate Video Compression Challenge
abstract
This report presents the VCIP 2025 Grand Challenge on Ultra Low-Bitrate Video Compression, which aims to promote research progress in perceptually optimized and computationally efficient video compression under extreme bandwidth constraints. The challenge focuses on scenarios such as emergency communication, remote monitoring, and low-power transmission, where conventional codecs like HEVC and AV1 struggle to maintain acceptable perceptual quality. Two benchmark tracks are introduced to assess the trade-off between compression ratio, visual quality, and complexity: Track 1 (50 kbps) targets extreme low-bitrate conditions, while Track 2 (200 kbps) allows moderately higher bitrates under real-time constraints. Participants were required to satisfy strict limits on encoding and decoding complexity, evaluated in kMac per pixel and per-frame runtime. The evaluation combined both objective metrics (PSNR, SSIM, VMAF, LPIPS) and subjective human perception scoring to comprehensively assess quality and efficiency. The results reveal two complementary trends for the future of video compression: (1) efficiency-oriented codec architecture design enabling low-latency deployment on edge devices, and (2) perceptual enhancement through post-decoding restoration guided by temporal and semantic priors. Together, these directions signal a paradigm shift from traditional rate–distortion optimization toward a broader rate–perception–complexity trade-off. The insights gained from this challenge are expected to inspire future standards and generative compression models for perceptually-driven, adaptive, and bandwidth-efficient video communication.
Guo Lu, Jing Wang 0194, Yunuo Chen 0002, Chuqin Zhou, Yibo Shi
VCIP5
2025 Determination of barrier surface in Target-Attacker-Defender game with capture radius for superior pursuer
Yibo Shi, Chaoli Wang 0002
Neurocomputing1
2024 Neural Rate Control for Learned Video Compression
abstract
The learning-based video compression method has made significant progress in recent years, exhibiting promising compression performance compared with traditional video codecs. However, prior works have primarily focused on advanced compression architectures while neglecting the rate control technique. Rate control can precisely control the coding bitrate with optimal compression performance, which is a critical technique in practical deployment. To address this issue, we present a fully neural network-based rate control system for learned video compression methods. Our system accurately encodes videos at a given bitrate while enhancing the rate-distortion performance. Specifically, we first design a rate allocation model to assign optimal bitrates to each frame based on their varying spatial and temporal characteristics. Then, we propose a deep learning-based rate implementation network to perform the rate-parameter mapping, precisely predicting coding parameters for a given rate. Our proposed rate control system can be easily integrated into existing learning-based video compression methods. The extensive experimental results show that the proposed method achieves accurate rate control on several baseline methods while also improving overall rate-distortion performance.
Guo Lu, Yunuo Chen 0002, Shen Wang 0013, Yibo Shi, Jing Wang 0194, Li Song 0001
ICLR5
2023 High Visual-Fidelity Learned Video Compression
abstract
With the growing demand for video applications, many advanced learned video compression methods have been developed, outperforming traditional methods in terms of objective quality metrics such as PSNR. Existing methods primarily focus on objective quality but tend to overlook perceptual quality. Directly incorporating perceptual loss into a learned video compression framework is non-trivial and raises several perceptual quality issues that need to be addressed. In this paper, we investigated these issues in learned video compression and propose a novel High Visual-Fidelity Learned Video Compression framework (HVFVC). Specifically, we design a novel confidence-based feature reconstruction method to address the issue of poor reconstruction in newly-emerged regions, which significantly improves the visual quality of the reconstruction. Furthermore, we present a periodic compensation loss to mitigate the checkerboard artifacts related to deconvolution operation and optimization. Extensive experiments have shown that the proposed HVFVC achieves excellent perceptual quality, outperforming the latest VVC standard with only 50% required bitrate.
Meng Li 0050, Yibo Shi, Jing Wang 0194, Yunqi Huang
ACM Multimedia2
2022 Content-Oriented Learned Image Compression
Meng Li 0050, Shangyin Gao, Yihui Feng, Yibo Shi, Jing Wang 0194
ECCV (19)4
2022 AlphaVC: High-Performance and Efficient Learned Video Compression
Yibo Shi, Yunying Ge, Jing Wang 0194, Jue Mao
ECCV (19)1