VLDB 2026 Research / reviewers in the wild / expert
Ying Chen 0011
dblp:21/5521-11
· DBLP profile ↗
70ranked-venue papers
11as first author
35since 2021 · last 2026
0000-0002-1620-9904ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 46 · 10 first-author · 20 since 2021Artificial intelligence and machine learning · 16 · 14 since 2021Systems, architecture and hardware · 9 · 1 first-author · 2 since 2021Computer networks · 6 · 6 since 2021Databases, data management, data science and information retrieval · 4 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Rethinking Bitrate Adaptation in WebRTC-Based Low-Latency Live Streaming SystemsabstractBitrate adaptation (or ABR) plays a crucial role in shaping the user QoE of low-latency live streaming (LLLS) applications. However, the unique characteristics of modern WebRTC-based LLLS systems render traditional ABR paradigms inefficient or even ineffective in real-world scenarios. This motivates us to revisit the bitrate adaptation problem within this complex yet realistic context and propose Salmon, an innovative bitrate adaptation framework designed for WebRTC-based LLLS applications. Salmon pragmatically addresses challenges arising from new QoE objectives, two-stream handovers, user interaction behaviors, and application-specific signal semantics. Deployed on a leading e-commerce LLLS platform, Salmon demonstrates significant performance gains over state-of-the-art algorithms. Notably, in low-bandwidth conditions, Salmon reduces startup delay by 17.7%, stall by 11%, frame jumps by 41.2%, and switching rate by 25.9×. Shibo Wang 0002, Chengxuan Yuan, Zhehao Zhong, Yiding Yu, Zeke Wang, Cuijun Qu, Ying Chen 0011 |
NOSSDAV | 9 |
| 2026 | Self-distilled learning of adaptive interval 3D lookup tables on real-time image enhancement
Ruikai Zhou, Canqian Yang, Meiguang Jin, Xu Jia 0012, Ying Chen 0011, Yi Xu 0001 |
Pattern Recognit. | 6 |
| 2026 | Toward Robust Low-Latency Live Streaming: Measurement, Prediction, and Rate Adaptation Under UncertaintyabstractLow latency live streaming (LLLS) leverages chunked transfer encoding (CTE) to substantially reduce end-to-end latency. However, this paradigm introduces a cascade of challenges for adaptive bitrate (ABR) algorithms: (1) the sending idle periods between chunks in CTE render bandwidth measurement difficult and prone to error; (2) bandwidth prediction in LLLS is an irregular time series forecasting with uncertain future segment size, leading to a circular prediction dependency; (3) stochastic uncertainty within LLLS, such as fluctuating idle time, leads to imprecise buffer evolution and ABR degradation. In this paper, we tackle the issues and present AAR, a novel LLLS framework that comprises 3 key modules: (1) accurate bandwidth measurement that leverages a server-side Flag to identify burst transmission and isolate chunks. We further propose to fuse our two learning and heuristic-based algorithms via confidence estimation; (2) bandwidth prediction via conditional normalizing flow to simultaneously learn joint variable distributions. We further propose a bitrate-aware transformer to capture the intrinsic circular relationships as backbone flow condition; (3) an LLLS tailored ABR with a novel and robust objective to maximize the minimum Quality of Experience (QoE) under uncertainty. We propose two theorems to derive the min solution via download time bounds, and we maximize the QoE via Model Predictive Controller (MPC) with LLLS tailored state evolution. Extensive experiments on real-world network traces demonstrate that AAR significantly outperforms baselines with absolute error reduction by 11%-83% for measurement and up to 17% for prediction. We also improve QoE by up to 102% across all tested network conditions. Jiahui Chen 0009, Yiding Yu, Ying Chen 0011, Tianchi Huang, Lifeng Sun |
IEEE Trans. Mob. Comput. | 4 |
| 2026 | KSIQA: A Knowledge-Sharing Model for No-Reference Image Quality AssessmentabstractNo-reference image quality assessment (NR-IQA) aims to quantitatively measure human perception of visual quality without comparing a distorted image to a reference. Despite recent advances, existing NR-IQR approaches often demonstrate insufficient ability to capture perceptual cues in the absence of a reference, limiting their generalisability across diverse and complex real-world image degradations. These limitations hinder their ability to match the reliability of full-reference IQA (FR-IQA) counterparts. A key challenge, therefore, is to enable NR-IQA models to emulate the reference-aware reasoning exhibited by humans and FR-IQA methods. To address this challenge, we propose a novel NR-IQA model based on a knowledge-sharing (KS) strategy to simulate this capability and predict image quality more effectively. Specifically, we designate an FR-IQA model as the teacher and an NR-IQA model as the student. Unlike conventional knowledge distillation (KD), our proposed architecture enables the NR-IQA student and FR-IQA teacher to share a decoder rather than being independent models. Furthermore, the student model contains a Mental Imagery Generation (MIG) module to learn mental imagery as the reference. To fully exploit local and global information, we adopt a vision transformer (ViT) branch and a convolutional neural network branch for feature extraction (FE). Finally, a quality-aware regressor (QAR) combined with deep ordinal regression is constructed to infer the quality score. Experiments show that our proposed NR-IQA model, KSIQA, has class-leading performance against current no-reference (NR) techniques across widespread benchmark datasets. Huasheng Wang, Hongchen Tan, Jianxun Lou, Xiaochang Liu, Wei Zhou 0021, Ying Chen 0011, Roger M. Whitaker, Walter Colombo, Hantao Liu |
IEEE Trans. Neural Networks Learn. Syst. | 7 |
| 2026 | Understanding and Taming the Inflated Latency in Mobile Cloud RenderingabstractLow-latency cloud rendering enables mobile users to experience high-quality, real-time 3D graphics but achieving low Motion-to-Photon (MTP) latency while maintaining smooth playback is a significant challenge. Our real-world measurement study identifies Receive-to-Composition (R2C) latency, caused by ineffective jitter buffer management, as the primary factor contributing to increased MTP latency. To address this, we introduce JitBright, a client-side optimization strategy that dynamically reduces MTP latency through adaptive jitter buffer management. By adjusting buffer levels based on smoothing playback probability and implementing proactive keyframe requests to mitigate frame dependency, JitBright minimizes both active and passive waiting times. Our large-scale evaluation, conducted over 591,000 sessions across diverse network conditions (WiFi, 4G, 5G) and device types, demonstrates significant improvements in user experience. JitBright reduces median R2C latency by up to 87.5%, increases the proportion of sessions meeting strict MTP latency requirements by 6%–27%, and decreases the video freeze rate from 2.4%–2.8% to 0.4%–1.0%. Yuankang Zhao, Qinghua Wu 0004, Gerui Lv, Furong Yang, Jiuhai Zhang, Yanmei Liu, Zhenyu Li 0001, Ying Chen 0011, Gaogang Xie |
ACM Trans. Multim. Comput. Commun. Appl. | 10 |
| 2025 | Advancing Myopia To Holism: Fully Contrastive Language-Image Pre-trainingabstractIn rapidly evolving field of vision-language models (VLMs), contrastive language-image pre-training (CLIP) has made significant strides, becoming foundation for various downstream tasks. However, relying on one-to-one (image, text) contrastive paradigm to learn alignment from large-scale messy web data, CLIP faces a serious myopic dilemma, resulting in biases towards monotonous short texts and shallow visual expressivity. To overcome these issues, this paper advances CLIP into one novel holistic paradigm, by updating both diverse data and alignment optimization. To obtain colorful data with low cost, we use image-to-text captioning to generate multi-texts for each image, from multiple perspectives, granularities, and hierarchies. Two gadgets are proposed to encourage textual diversity. To match such (image, multi-texts) pairs, we modify the CLIP image encoder into multi-branch, and propose multi-to-multi contrastive optimization for image-text part-to-part matching. As a result, diverse visual embeddings are learned for each image, bringing good interpretability and generalization. Extensive experiments and ablations across over ten benchmarks indicate that our holistic CLIP significantly outperforms existing myopic CLIP, including image-text retrieval, open-vocabulary classification, and dense visual tasks. Project page is available to further promote the prosperity of VLMs: https://voide1220.github.io/Holism/. Haicheng Wang, Chen Ju, Weixiong Lin, Shuai Xiao 0002, Mingshuai Yao, Jinsong Lan, Ying Chen 0011, Qingwen Liu 0002 |
CVPR | 10 |
| 2025 | FPEM: Face Prior Enhanced Facial Attractiveness Prediction for Live Videos with Face Retouching
Hongjiu Yu, Ying Chen 0011, Kai Li 0012, Xiongkuo Min, Huiyu Duan, Guangtao Zhai, Xu Liu 0006 |
ICCV | 4 |
| 2025 | Enhanced Bandwidth Measurement and Robust Rate Adaptation for Low-Latency Live Streaming
Jiahui Chen 0009, Yiding Yu, Ying Chen 0011, Tianchi Huang, Lifeng Sun |
INFOCOM | 4 |
| 2025 | CLIP-DQA: Blindly Evaluating Dehazed Images from Global and Local Perspectives Using CLIPabstractBlind dehazed image quality assessment (BDQA), which aims to accurately predict the visual quality of dehazed images without any reference information, is essential for the evaluation, comparison, and optimization of image dehazing algorithms. Existing learning-based BDQA methods have achieved remarkable success, while the small scale of DQA datasets limits their performance. To address this issue, in this paper, we propose to adapt Contrastive Language-Image Pre-Training (CLIP), pre-trained on large-scale image-text pairs, to the BDQA task. Specifically, inspired by the fact that the human visual system understands images based on hierarchical features, we take global and local information of the dehazed image as the input of CLIP. To accurately map the input hierarchical information of dehazed images into the quality score, we tune both the vision branch and language branch of CLIP with prompt learning. Experimental results on two authentic DQA datasets demonstrate that our proposed approach, named CLIP-DQA, achieves more accurate quality predictions over existing BDQA methods. The code is available at https://github.com/JunFu1995/CLIP-DQA. Yirui Zeng, Jun Fu 0007, Hadi Amirpour, Huasheng Wang, Guanghui Yue 0001, Hantao Liu, Ying Chen 0011, Wei Zhou 0021 |
ISCAS | 7 |
| 2025 | MARC: Motion-Aware Rate Control for Mobile E-commerce Cloud Rendering
Yuankang Zhao, Furong Yang, Gerui Lv, Qinghua Wu 0004, Yanmei Liu, Jiuhai Zhang, Yutang Peng, Ying Chen 0011, Zhenyu Li 0001, Gaogang Xie |
USENIX ATC | 10 |
| 2025 | End-to-End Optimized Image Compression With Deep Gaussian Process RegressionabstractEnd-to-end optimization via deep neural networks has facilitated lossy image compression. Existing neural network-based entropy models for end-to-end optimized image compression are limited by parameterized Gaussian distributions with deterministic mean and variance and cannot achieve accurate rate estimation for bottleneck representation with varying statistics. In this paper, we propose a novel entropy model based on deep Gaussian process regression (DGPR) to address this problem. Specifically, the proposed entropy model leverages autoregressive DGPR to flexibly predict the channel-wise posterior distributions of high-dimensional bottleneck representation for entropy coding. Consequently, we develop a well-established bit-rate estimation scheme via posterior inference of DGPR using the learned probabilistic distribution. Furthermore, scalable training is achieved via tensor train decomposition and Monte Carlo sampling to enable tractable variational inference of DGPR. To our best knowledge, this paper is the first attempt to develop the learnable probabilistic model for flexible parameter estimation in entropy modeling. Experimental results show that the proposed model outperforms conventional image compression methods (e.g., JPEG2000 and BPG) as well as recent end-to-end optimized methods on the Kodak and Tecnick datasets in terms of rate-distortion performance. Maida Cao, Wenrui Dai, Junni Zou, Ying Chen 0011, Hongkai Xiong |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Joint Luminance-Chrominance Learning for Image DebandingabstractBanding is a visually annoying artifact that frequently occurs along the chain of video acquisition, production, distribution, and display, showing a significant need for improvement in many fields. Thus far, efforts on banding removal are mainly knowledge-driven or merely learning on RGB space, which is either limited by domain knowledge or lacks the consideration for banding in chrominance channels. In this work, we propose a unified deep neural network that explicitly disentangles the luminance and chrominance channels, and simultaneously recovers intensity gradients and color discontinuity from detection-free measurement in an end-to-end manner. Our debanding model is comprised of a luminance restoration network (LR-Net) and a chrominance restoration network (CR-Net). Each of them follows an encoder-decoder architecture, where a cascade of residual blocks is employed to exploit hierarchical non-local features in spatial dimensions for more powerful feature representation. Moreover, we investigate the characteristics of banding artifacts and apply specific loss functions to guide the debanding in different channels, thus boosting the restoration performance. Both qualitative and quantitative experiments show that our model significantly surpasses the existing method in terms of all 7 metrics. Ultimately, our network trained on simulated data exhibits good adaptiveness under various compression scenarios, which further demonstrates the effectiveness of the proposed model. Zijian Chen 0001, Wei Sun 0029, Jun Jia, Ru Huang 0002, Fangfang Lu, Ying Chen 0011, Xiongkuo Min, Guangtao Zhai, Wenjun Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Adaptive Spatiotemporal Graph Transformer Network for Action Quality AssessmentabstractLong video action quality assessment (AQA) aims to evaluate the performance of long-term actions depicted in a video and produce an overall assessment for action quality. A video of long-term actions often contains more complicated temporal and spatial information than that of short-term actions. However, existing approaches that segment a video into individual clips for independent analysis potentially disrupt the narrative flow and diminish contextual details within and across clips, impeding comprehensive video understanding. To address this challenge, we propose an adaptive spatiotemporal graph transformer network (ASGTN) that combines multiple graph structures and transformer attention mechanisms to capture both local and global contextual information within and across clips in a long video. Specifically, the adaptive spatiotemporal graph (ASG) combines a spatial graph branch, designed to enrich the local nuanced spatiotemporal relations within an individual clip, and a temporal graph branch, tailored to dynamically learn the semantic context across different clips. Furthermore, a transformer encoder is integrated to amplify the global dependencies across clips in the entire video. This structure is designed to preserve narrative coherence and maintain essential contextual details in video-level features. Finally, we employ a level-focused decoder to predict the action quality score distribution. Experiments demonstrate that our model achieves state-of-the-art results on popular AQA datasets. Our code is available athttps://github.com/jiangliu5/ASGTN_AQA. Huasheng Wang, Wei Zhou 0021, Katarzyna Stawarz, Padraig Corcoran, Ying Chen 0011, Hantao Liu |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | A Deep Transformer-Based Fast CU Partition Approach for Inter-Mode VVCabstractThe latest versatile video coding (VVC) standard proposed by the Joint Video Exploration Team (JVET) has significantly improved coding efficiency compared to that of its predecessor, while introducing an extremely higher computational complexity by $6\sim 26$ times. The quad-tree plus multi-type tree (QTMT)-based coding unit (CU) partition accounts for most of the encoding time in VVC encoding. This paper proposes a data-driven fast CU partition approach based on an efficient Transformer model to accelerate VVC inter-coding. First, we establish a large-scale database for inter-mode VVC, comprising diverse CU partition patterns from more than 800 raw video sequences across various resolutions and contents. Next, we propose a deep neural network model with a Transformer-based temporal topology for predicting the CU partition, named as TCP-Net, which is adaptive to the group of pictures (GOP) hierarchy in VVC. Then, we design a two-stage structured output for TCP-Net, reflecting both the locations of CU edges and the split modes of all possible CUs. Accordingly, we develop a dual-supervised optimization mechanism to train the TCP-Net model with improved accuracy. The experimental results have verified that our approach can reduce the encoding time by $46.89\sim 55.91$ % with negligible rate-distortion (RD) degradation, outperforming other state-of-the-art approaches. Tianyi Li 0004, Mai Xu, Ying Chen 0011, Kai Li 0012 |
IEEE Trans. Image Process. | 4 |
| 2025 | Evaluating Point Cloud From Moving Camera Videos: A No-Reference MetricabstractPoint cloud is one of the most widely used digital representation formats for three-dimensional (3D) contents, the visual quality of which may suffer from noise and geometric shift distortions during the production procedure as well as compression and downsampling distortions during the transmission process. To tackle the challenge of point cloud quality assessment (PCQA), many PCQA methods have been proposed to evaluate the visual quality levels of point clouds by assessing the rendered static 2D projections. Although such projectionbased PCQA methods achieve competitive performance with the assistance of mature image quality assessment (IQA) methods, they neglect that the 3D model is also perceived in a dynamic viewing manner, where the viewpoint is continually changed according to the feedback of the rendering device. Therefore, in this paper, we evaluate the point clouds from moving camera videos and explore the way of dealing with PCQA tasks via using video quality assessment (VQA) methods. First, we generate the captured videos by rotating the camera around the point clouds through several circular pathways. Then we extract both spatial and temporal quality-aware features from the selected key frames and the video clips through using trainable 2D-CNN and pretrained 3D-CNN models respectively. Finally, the visual quality of point clouds is represented by the video quality values. The experimental results reveal that the proposed method is effective for predicting the visual quality levels of the point clouds and even competitive with full-reference (FR) PCQA methods. The ablation studies further verify the rationality of the proposed framework and confirm the contributions made by the qualityaware features extracted via the dynamic viewing manner. The code is available athttps://github.com/zzc-1998/VQA_PC. Wei Sun 0029, Yucheng Zhu, Xiongkuo Min, Wei Wu 0002, Ying Chen 0011, Guangtao Zhai |
IEEE Trans. Multim. | 6 |
| 2025 | Chest X-Ray Visual Saliency Modeling: Eye-Tracking Dataset and Saliency Prediction ModelabstractRadiologists' eye movements during medical image interpretation reflect their perceptual-cognitive processes of diagnostic decisions. The eye movement data can be modeled to represent clinically relevant regions in a medical image and potentially integrated into an artificial intelligence (AI) system for automatic diagnosis in medical imaging. In this article, we first conduct a large-scale eye-tracking study involving 13 radiologists interpreting 191 chest X-ray (CXR) images, establishing a best-of-its-kind CXR visual saliency benchmark. We then perform analysis to quantify the reliability and clinical relevance of saliency maps (SMs) generated for CXR images. We develop CXR image saliency prediction method (CXRSalNet), a novel saliency prediction model that leverages radiologists' gaze information to optimize the use of unlabeled CXR images, enhancing training and mitigating data scarcity. We also demonstrate the application of our CXR saliency model in enhancing the performance of AI-powered diagnostic imaging systems. Jianxun Lou, Huasheng Wang, Xinbo Wu, John Cho Hui Ng, Kaveri A. Thakoor, Padraig Corcoran, Ying Chen 0011, Hantao Liu |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2025 | A Bioinspired Deep Learning Framework for Saliency-Based Image Quality AssessmentabstractAdvancements in deep learning have led to significant progress in no-reference (NR) image quality assessment (NR-IQA) for evaluating the perceived quality of digital images without relying on a reference. However, existing NR-IQA models remain suboptimal in handling complex and diverse natural images. Visual saliency constitutes a critical element for enhancing the reliability of NR-IQA, but the optimal use of saliency in deep learning-based NR-IQA has not heretofore been significantly explored. In this article, we present a novel method for integrating saliency in NR-IQA, which is motivated by the saliency-based visual search mechanism that different parts of the visual input are visited by the focus of attention (FOA) in the order of decreasing saliency. By dividing saliency into the high and low levels of FOA, we build a bioinspired deep neural network-BioSIQNet-based on a multitask learning (MTL) framework. The network architecture consists of two saliency-specific tasks and one primary image quality assessment (IQA) task. The low and high saliency (HS) are separately encoded and integrated into the early and deeper layers of the IQA network, respectively, analogous to the hierarchical processing in the visual cortex of the brain that allocates low attentional resources to process the simple patterns and high resources to learn intricate representations. We demonstrate that leveraging the synergy between visual attention and image quality perception and joint learning of these interconnected visual tasks can enhance the overall learning capabilities of the primary IQA model. Experiments validate the effectiveness of our proposed BioSIQNet for NR-IQA. Huasheng Wang, Yueran Ma, Hongchen Tan, Xiaochang Liu, Ying Chen 0011, Hantao Liu |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2024 | Enhancing Quality of Compressed Images by Mitigating Enhancement Bias Towards Compression DomainabstractExisting quality enhancement methods for compressed images focus on aligning the enhancement domain with the raw domain to yield realistic images. However, these methods exhibit a pervasive enhancement bias towards the compression domain, inadvertently regarding it as more realistic than the raw domain. This bias makes enhanced images closely resemble their compressed counterparts, thus degrading their perceptual quality. In this paper, we propose a simple yet effective method to mitigate this bias and enhance the quality of compressed images. Our method employs a conditional discriminator with the compressed image as a key condition, and then incorporates a domain-divergence regularization to actively distance the enhancement domain from the compression domain. Through this dual strategy, our method enables the discrimination against the compression domain, and brings the enhancement domain closer to the raw domain. Comprehensive quality evaluations confirm the superiority of our method over other state-of-the-art methods without incurring inference overheads. Qunliang Xing, Mai Xu, Shengxi Li, Xin Deng 0002, Meisong Zheng, Huaida Liu, Ying Chen 0011 |
CVPR | 7 |
| 2024 | Chorus: Coordinating Mobile Multipath Scheduling and Adaptive Video StreamingabstractIncreasing bandwidth demands of mobile video streaming pose a challenge in optimizing the Quality of Experience (QoE) for better user engagement. Multipath transmission promises to extend network capacity by utilizing multiple wireless links simultaneously. Previous studies mainly tune the packet scheduler in multipath transmission, expecting higher QoE by accelerating transmission. However, since Adaptive BitRate (ABR) algorithms overlook the impact of multipath scheduling on throughput prediction, multipath adaptive streaming can even experience lower QoE than single-path. This paper proposes Chorus, a cross-layer framework that coordinates multipath scheduling with adaptive streaming to optimize QoE jointly. Chorus establishes two-way feedback control loops between the server and the client. Furthermore, Chorus introduces Coarse-grained Decisions, which assist appropriate bitrate selection by considering the scheduling decision in throughput prediction, and Finegrained Corrections, which meet the predicted throughput by QoE-oriented multipath scheduling. Extensive emulation and real-world mobile Internet evaluations show that Chorus outperforms the state-of-the-art MPQUIC scheduler, improving average QoE by 23.5% and 65.7%, respectively. Gerui Lv, Qinghua Wu 0004, Yanmei Liu, Zhenyu Li 0001, Qingyue Tan, Furong Yang, Ying Chen 0011, Gaogang Xie |
MobiCom | 10 |
| 2024 | JitBright: towards Low-Latency Mobile Cloud Rendering through Jitter Buffer OptimizationabstractLow-latency cloud rendering services use high-performance servers to provide mobile device users with exquisite graphics and convenient access experiences. Due to the complexity of the system and the diversity of impacting factors, identifying system bottlenecks has become a significant challenge. To demystify system performance, we build an online cloud rendering system to measure the latency distribution of its key components. Yuankang Zhao, Qinghua Wu 0004, Gerui Lv, Furong Yang, Jiuhai Zhang, Yanmei Liu, Zhenyu Li 0001, Ying Chen 0011, Gaogang Xie |
NOSSDAV | 9 |
| 2023 | MD-VQA: Multi-Dimensional Quality Assessment for UGC Live VideosabstractUser-generated content (UGC) live videos are often bothered by various distortions during capture procedures and thus exhibit diverse visual qualities. Such source videos are further compressed and transcoded by media server providers before being distributed to end-users. Because of the flourishing of UGC live videos, effective video quality assessment (VQA) tools are needed to monitor and perceptually optimize live streaming videos in the distributing process. In this paper, we address UGC Live VQA problems by constructing a first-of-a-kind subjective UGC Live VQA database and developing an effective evaluation tool. Concretely, 418 source UGC videos are collected in real live streaming scenarios and 3,762 compressed ones at different bit rates are generated for the subsequent subjective VQA experiments. Based on the built database, we develop a Multi-12imensional VQA (MD-VQA) evaluator to measure the visual quality of UGC live videos from semantic, distortion, and motion aspects respectively. Extensive experimental results show that MD-VQA achieves state-of-the-art performance on both our UGC Live VQA database and existing compressed UGC VQA databases. Wei Wu 0002, Wei Sun 0029, Danyang Tu, Wei Lu 0021, Xiongkuo Min, Ying Chen 0011, Guangtao Zhai |
CVPR | 7 |
| 2023 | Lightweight Network Towards Real-Time Image Denoising On Mobile DevicesabstractDeep convolutional neural networks have achieved great progress in image denoising tasks. However, their complicated architectures and heavy computational cost hinder their deployments on mobile devices. Some recent efforts in designing lightweight denoising networks focus on reducing either FLOPs (floating-point operations) or the number of parameters. However, these metrics are not directly correlated with the on-device latency. In this paper, we identify the real bottlenecks that affect the CNN-based models’ runtime performance on mobile devices: memory access cost and NPU-incompatible operations, and build the model based on these. To further improve the denoising performance, the mobile-friendly attention module MFA and the model reparameterization module RepConv are proposed, which enjoy both low latency and excellent denoising performance. To this end, we propose a mobile-friendly denoising network, namely MFDNet. The experiments show that MFDNet achieves state-of-the-art performance on real-world denoising benchmarks SIDD and DND under real-time latency on mobile devices. The code and pre-trained models will be released. Zhuoqun Liu, Meiguang Jin, Ying Chen 0011, Huaida Liu, Canqian Yang, Hongkai Xiong |
ICIP | 3 |
| 2023 | Hierarchical Feature Fusion Transformer for No-Reference Image Quality AssessmentabstractRecently, increasing interest has been drawn in Transformer-based models for No-reference Image Quality Assessment (NR-IQA), especially for the hybrid approach. The hybrid approach tend to apply Transformer to aggregate quality information from feature maps extracted by Convolutional Neural Networks (CNN). However, existing methods cannot fully utilize the information of hierarchical features extracted by the deep neural network, resulting in the limited performance of image quality evaluation. In this work, we propose a novel Hierarchical Feature Fusion Transformer for NR-IQA (HiFFTiq), which is able to effectively exploit complementary strengths of features extracted by different layers. Further, we propose a new Uniform Partition Pooling (UPP) which can reduce the resolution of input features via uniform partitions and can well retain the quality-related information compared to the traditional pooling method Sliding Window Pooling (SWP). The results of experiment demonstrate that HiFFTiq leads to improvements of performance over the state-of-the-art methods on three large scale NR-IQA datasets. Zesheng Wang 0004, Wei Wu 0002, Wei Sun 0029, Ying Chen 0011, Kai Li 0012, Guangtao Zhai |
ICIP | 5 |
| 2023 | PACC: Perception Aware Congestion Control for Real-time CommunicationabstractDue to the network fluctuations, congestion control is indispensable to guarantee the quality of experience (QoE) for Real-Time Communication (RTC) users. This component adjusts the sending rate of media data, which determines the video encoding bitrate. However, existing control schemes either only focus on network numerical indicators or fail to adapt to various network environments. Logically, we propose PACC (Perception Aware Congestion Control) for RTC in this paper. Leveraging the convolutional neural network (CNN), we develop a quality sensor to infer the video quality increasing rate. Assisted with the variation trend analysis for user perception, PACC tunes the bitrate towards the direction of better QoE. Extensive tracedriven experiments demonstrate the effectiveness of PACC, which outperforms the existing landmark schemes by 8.2% to 32.4% and 6.8% to 18.0% in terms of transport and application layer QoE metrics, respectively. Bingcong Lu, Li Song 0001, Rong Xie 0004, Yanmei Liu, Ying Chen 0011 |
ICME | 6 |
| 2023 | A Deep Learning-Based Multidimensional Aesthetic Quality Assessment Method for Mobile Game ImagesabstractMobile games have played an increasingly significant role in people's leisure lives in recent years, thanks to the fast expansion of the gaming industry and the widespread use of mobile devices. The aesthetic quality of game pictures is a very important factor that attracts users' interest. However, evaluating the aesthetic quality of mobile game pictures is difficult since the painting styles of games vary greatly and the evaluation criteria are also diversified. In this article, we propose a multitask deep learning-based method, which is able to predict the aesthetic quality of mobile game images in multiple dimensions. The proposed model consists of two modules, a feature extraction module and a quality regression module. We extract quality-aware features from intermediate layers of the deep convolution neural network and then incorporate them into the final feature representation in the feature extraction module, allowing the model to fully use visual information from low to high levels. The quality regression module uses fully connected layers to map quality-aware features into quality scores across multiple dimensions. The multidimensional aesthetic quality scores are trained using a multitask learning approach, in which quality-aware features are shared across multiple dimensional quality prediction tasks. Finally, several key factors which help the proposed model perform better are analyzed. The experimental results indicate that our proposed method not only achieves the greatest performance on mobile game images, but also is applicable to natural scene images. Tao Wang 0078, Wei Sun 0029, Wei Wu 0002, Ying Chen 0011, Xiongkuo Min, Wei Lu 0021, Guangtao Zhai |
IEEE Trans. Games | 4 |
| 2023 | Deep Multi-Task Learning Based Fast Intra-Mode Decision for Versatile Video CodingabstractThe latest Versatile Video Coding (VVC) standard has significantly coding efficiency improvement compared with its ancestor High Efficiency Video Coding (HEVC) standard, but at the expense of over-high complexity. As measured by the VVC test model (VTM), the intra-mode comparison and selection in the rate-distortion optimization (RDO) search consume most of the encoding time. In this paper, we propose a deep multi-task learning based fast intra-mode decision approach via adaptively pruning off most redundant modes. First, we create a large-scale intra-mode database for VVC, including both normal angular modes and the newly introduced tools, i.e., intra sub-partition (ISP) and matrix-based intra prediction (MIP). Next, we propose a multi-task intra-mode decision network (MID-Net) model to effectively predict the most probable angular modes and whether to skip ISP and MIP modes. Then, a fast intra-coding workflow is designed accordingly, involving rough mode decision (RMD) acceleration and candidate mode list (CML) pruning. For the workflow output, the learning-oriented probability and the statistics-oriented probability are synthesized together to further improve the prediction accuracy, ensuring that only unnecessary intra-modes are skipped. Finally, experimental results show that our approach can significantly reduce 40.48% of encoding time of VVC intra-coding with negligible rate-distortion degradation, outperforming other state-of-the-art approaches. Tianyi Li 0004, Ying Chen 0011, Kaijin Wei, Mai Xu, Honggang Qi |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2022 | A Fine-grained Interpretability Evaluation Benchmark for Neural NLPabstractLijie Wang, Yaozong Shen, Shuyuan Peng, Shuai Zhang, Xinyan Xiao, Hao Liu, Hongxuan Tang, Ying Chen, Hua Wu, Haifeng Wang. Proceedings of the 26th Conference on Computational Natural Language Learning (CoNLL). 2022. Yaozong Shen, Shuyuan Peng, Xinyan Xiao, Hao Liu 0026, Hongxuan Tang, Ying Chen 0011, Hua Wu 0003, Haifeng Wang 0001 |
CoNLL | 8 |
| 2022 | AdaInt: Learning Adaptive Intervals for 3D Lookup Tables on Real-time Image EnhancementabstractThe 3D Lookup Table (3D LUT) is a highly-efficient tool for real-time image enhancement tasks, which models a non-linear 3D color transform by sparsely sampling it into a discretized 3D lattice. Previous works have made efforts to learn image-adaptive output color values of LUTs for flexible enhancement but neglect the importance of sampling strategy. They adopt a sub-optimal uniform sampling point allocation, limiting the expressiveness of the learned LUTs since the (tri-)linear interpolation between uniform sampling points in the LUT transform might fail to model local non-linearities of the color transform. Focusing on this problem, we present AdaInt (Adaptive Intervals Learning), a novel mechanism to achieve a more flexible sampling point allocation by adaptively learning the non-uniform sampling intervals in the 3D color space. In this way, a 3D LUT can increase its capability by conducting dense sampling in color ranges requiring highly non-linear transforms and sparse sampling for near-linear transforms. The proposed AdaInt could be implemented as a compact and efficient plug-and-play module for a 3D LUT-based method. To enable the end-to-end learning of AdaInt, we design a novel differentiable operator called AiLUT-Transform (Adaptive Interval LUT Transform) to locate input colors in the non-uniform 3D LUT and provide gradients to the sampling intervals. Experiments demonstrate that methods equipped with AdaInt can achieve state-of-the-art performance on two public benchmark datasets with a negligible overhead increase. Our source code is available at https://github.com/ImCharlesY/AdaInt. Canqian Yang, Meiguang Jin, Xu Jia 0012, Yi Xu 0001, Ying Chen 0011 |
CVPR | 5 |
| 2022 | Entropy Modeling via Gaussian Process Regression for Learned Image CompressionabstractExisting entropy models in learned image compression are cumbersome to generate fixed mean and variance for estimating Gaussian distributions for latent representation. In this paper, we propose a novel entropy model based on Gaussian process regression (GPR) that flexibly predicts the mean of Gaussians with posterior distributions characterized by covari-ance functions spanned in the high-dimensional feature space. Furthermore, we develop the rate-distortion optimization based on the proposed entropy model by approximating the bitrates with the evidence lower bound (ELBO) derived via variational inference for GPR. The proposed model can be seamlessly integrated into existing end-to-end optimized frame-works by substituting the masked convolution based autoregressive models. Experimental results demonstrate that the proposed model outperforms conventional image compression methods such as JPEG2000 and BPG, as well as recent learning based methods on the Kodak dataset in terms of rate-distortion performance. Maida Cao, Wenrui Dai, Junni Zou, Ying Chen 0011, Hongkai Xiong |
DCC | 6 |
| 2022 | SepLUT: Separable Image-Adaptive Lookup Tables for Real-Time Image Enhancement
Canqian Yang, Meiguang Jin, Yi Xu 0001, Rui Zhang 0052, Ying Chen 0011, Huaida Liu |
ECCV (18) | 5 |
| 2022 | DuReader-Retrieval: A Large-scale Chinese Benchmark for Passage Retrieval from Web Search EngineabstractIn this paper, we present DuReader retrieval , a large-scale Chinese dataset for passage retrieval.DuReader retrieval contains more than 90K queries and over 8M unique passages from a commercial search engine.To alleviate the shortcomings of other datasets and ensure the quality of our benchmark, we (1) reduce the false negatives in development and test sets by manually annotating results pooled from multiple retrievers, and (2) remove the training queries that are semantically similar to the development and testing queries.Additionally, we provide two outof-domain testing sets for cross-domain evaluation, as well as a set of human translated queries for for cross-lingual retrieval evaluation.The experiments demonstrate that DuReader retrieval is challenging and a number of problems remain unsolved, such as the salient phrase mismatch and the syntactic mismatch between queries and paragraphs.These experiments also show that dense retrievers do not generalize well across domains, and cross-lingual retrieval is essentially challenging.DuReader Yifu Qiu, Yingqi Qu, Ying Chen 0011, Qiaoqiao She, Jing Liu 0022, Hua Wu 0003, Haifeng Wang 0001 |
EMNLP | 4 |
| 2022 | DuQM: A Chinese Dataset of Linguistically Perturbed Natural Questions for Evaluating the Robustness of Question Matching ModelsabstractIn this paper, we focus on the robustness evaluation of Chinese Question Matching (QM) models.Most of the previous work on analyzing robustness issues focus on just one or a few types of artificial adversarial examples.Instead, we argue that a comprehensive evaluation should be conducted on natural texts, which takes into account the fine-grained linguistic capabilities of QM models.For this purpose, we create a Chinese dataset namely DuQM which contains natural questions with linguistic perturbations to evaluate the robustness of QM models.DuQM contains 3 categories and 13 subcategories with 32 linguistic perturbations.The extensive experiments demonstrate that DuQM has a better ability to distinguish different models.Importantly, the detailed breakdown of evaluation by the linguistic phenomenon in DuQM helps us easily diagnose the strength and weakness of different models.Additionally, our experiment results show that the effect of artificial adversarial examples does not work on natural texts.Our baseline codes and a leaderboard are now publicly available.1 Hongyu Zhu 0002, Jing Yan 0004, Jing Liu 0022, Yu Hong 0001, Ying Chen 0011, Hua Wu 0003, Haifeng Wang 0001 |
EMNLP | 6 |
| 2022 | Disparity-Aware Reference Frame Generation Network for Multiview Video CodingabstractMultiview video coding (MVC) aims to compress the multiview video through the elimination of video redundancies, where the quality of the reference frame directly affects the compression efficiency. In this paper, we propose a deep virtual reference frame generation method based on a disparity-aware reference frame generation network (DAG-Net) to transform the disparity relationship between different viewpoints and generate a more reliable reference frame. The proposed DAG-Net consists of a multi-level receptive field module, a disparity-aware alignment module, and a fusion reconstruction module. First, a multi-level receptive field module is designed to enlarge the receptive field, and extract the multi-scale deep features of the temporal and inter-view reference frames. Then, a disparity-aware alignment module is proposed to learn the disparity relationship, and perform disparity shift on the inter-view reference frame to align it with the temporal reference frame. Finally, a fusion reconstruction module is utilized to fuse the complementary information and generate a more reliable virtual reference frame. Experiments demonstrate that the proposed reference frame generation method achieves superior performance for multiview video coding. Jianjun Lei 0001, Zongqian Zhang, Zhaoqing Pan, Dong Liu 0002, Xiangrui Liu, Ying Chen 0011, Nam Ling |
IEEE Trans. Image Process. | 6 |
| 2021 | Deep Stereoscopic Image Super-Resolution via Interaction ModuleabstractDeep learning-based methods have achieved remarkable performance in single image super-resolution. However, these methods cannot be effectively applied in stereoscopic image super-resolution without considering the characteristics of stereoscopic images. In this article, an interaction module-based stereoscopic image super-resolution network (IMSSRnet) is proposed to effectively utilize the correlation information in stereoscopic images. The key insight of the network lies with how to explore the complementary information of one view to help the reconstruction of another view. Thus, an interaction module is designed to acquire the enhanced features by utilizing complementary information between different views. Specifically, the interaction module is composed of a series of interaction units with a residual structure. In addition, the single image features of left and right views are obtained by a spatial feature extraction module, which can be realized by any existing single image super-resolution models. In order to obtain high-quality stereoscopic images, a gradient loss is introduced to preserve the texture details in a view, and a disparity loss is developed to constrain the disparity relationship between different views. Experimental results demonstrate that the proposed method achieves a promising performance and outperforms the state-of-the-art methods. Jianjun Lei 0001, Zhe Zhang 0041, Xiaoting Fan, Bolan Yang, Ying Chen 0011, Qingming Huang |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2021 | DeepQTMT: A Deep Learning Approach for Fast QTMT-Based CU Partition of Intra-Mode VVCabstractVersatile Video Coding (VVC), as the latest standard, significantly improves the coding efficiency over its predecessor standard High Efficiency Video Coding (HEVC), but at the expense of sharply increased complexity. In VVC, the quad-tree plus multi-type tree (QTMT) structure of the coding unit (CU) partition accounts for over 97% of the encoding time, due to the brute-force search for recursive rate-distortion (RD) optimization. Instead of the brute-force QTMT search, this paper proposes a deep learning approach to predict the QTMT-based CU partition, for drastically accelerating the encoding process of intra-mode VVC. First, we establish a large-scale database containing sufficient CU partition patterns with diverse video content, which can facilitate the data-driven VVC complexity reduction. Next, we propose a multi-stage exit CNN (MSE-CNN) model with an early-exit mechanism to determine the CU partition, in accord with the flexible QTMT structure at multiple stages. Then, we design an adaptive loss function for training the MSE-CNN model, synthesizing both the uncertain number of split modes and the target on minimized RD cost. Finally, a multi-threshold decision scheme is developed, achieving a desirable trade-off between complexity and RD performance. The experimental results demonstrate that our approach can reduce the encoding time of VVC by 44.65%~66.88% with a negligible Bjøntegaard delta bit-rate (BD-BR) of 1.322%~3.188%, significantly outperforming other state-of-the-art approaches. Tianyi Li 0004, Mai Xu, Runzhi Tang, Ying Chen 0011, Qunliang Xing |
IEEE Trans. Image Process. | 4 |
| 2020 | Deep Virtual Reference Frame Generation For Multiview Video CodingabstractMultiview video has a large amount of data which brings great challenges to both the storage and transmission. Thus, it is essential to increase the compression efficiency of multiview video coding. In this paper, a deep virtual reference frame generation method is proposed to improve the performance of multiview video coding. Specifically, a parallax-guided generation network (PGG-Net) is designed to transform the parallax relation between different viewpoints and generate a high-quality virtual reference frame. In the network, a multilevel receptive field module is designed to enlarge the receptive field and extract the multi-scale deep features. After that, a parallax attention fusion module is used to transform the parallax and merge the features. The proposed method is integrated into the platform of 3D-HEVC and the generated virtual reference frame is inserted into the reference picture list as an additional reference. Experimental results show that the proposed method achieves 5.31% average BD-rate reduction compared to the 3D-HEVC. Jianjun Lei 0001, Zongqian Zhang, Dong Liu 0002, Ying Chen 0011, Nam Ling |
ICIP | 4 |
| 2019 | An Overview of the 2019 Language and Intelligence Challenge
Quan Wang 0002, Wenquan Wu, Yabing Shi, Wei He 0014, Ying Chen 0011, Yajuan Lyu, Hua Wu 0003 |
NLPCC (2) | 8 |
| 2018 | Multi-Turn Response Selection for Chatbots with Deep Attention Matching NetworkabstractXiangyang Zhou, Lu Li, Daxiang Dong, Yi Liu, Ying Chen, Wayne Xin Zhao, Dianhai Yu, Hua Wu. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018. Xiangyang Zhou, Daxiang Dong, Ying Chen 0011, Wayne Xin Zhao, Dianhai Yu, Hua Wu 0003 |
ACL (1) | 5 |
| 2016 | Overview of the Multiview and 3D Extensions of High Efficiency Video CodingabstractThe High Efficiency Video Coding (HEVC) standard has recently been extended to support efficient representation of multiview video and depth-based 3D video formats. The multiview extension, MV-HEVC, allows efficient coding of multiple camera views and associated auxiliary pictures, and can be implemented by reusing single-layer decoders without changing the block-level processing modules since block-level syntax and decoding processes remain unchanged. Bit rate savings compared with HEVC simulcast are achieved by enabling the use of inter-view references in motion-compensated prediction. The more advanced 3D video extension, 3D-HEVC, targets a coded representation consisting of multiple views and associated depth maps, as required for generating additional intermediate views in advanced 3D displays. Additional bit rate reduction compared with MV-HEVC is achieved by specifying new block-level video coding tools, which explicitly exploit statistical dependencies between video texture and depth and specifically adapt to the properties of depth maps. The technical concepts and features of both extensions are presented in this paper. Gerhard Tech, Ying Chen 0011, Karsten Müller 0001, Jens-Rainer Ohm, Anthony Vetro, Ye-Kui Wang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2016 | Multiview and 3D Video Compression Using Neighboring Block Based Disparity VectorsabstractCompression of the statistical redundancy among different viewpoints, i.e., inter-view redundancy, is a fundamental and critical problem in multiview and three-dimensional (3D) video coding. To exploit the inter-view redundancy, disparity vectors are required to identify pixels of the same objects within two different views; in this way, the enhancement coding tools can be efficiently employed as new modes in block-based video codecs to achieve higher compression efficiency. Although disparity can be converted from depth, it is not possible in multiview video coding since depth information is not considered. Even when depth information is coded, it breaks the so-called multiview compatibility wherein texture views can be decoded without depth information. To resolve this problem, in this paper, a neighboring block-based disparity vector derivation (NBDV) method is proposed. The basic concept of NBDV is to derive a disparity vector (DV) of a current block by utilizing the motion information of spatially and temporally neighboring blocks predicted from another view. Through extensive experiments and analysis, it is shown that the proposed NBDV method achieves efficient DV derivation in the state-of-art video codecs, and it keeps the multiview compatibility with a relatively lower complexity. The proposed method has become an essential part of the 3D video standard extensions of H.264/AVC and HEVC. Ying Chen 0011, Xin Zhao 0003, Li Zhang 0006, Je-Won Kang |
IEEE Trans. Multim. | 1 |
| 2015 | Intra Block Copy for HEVC Screen Content CodingabstractSummary form only given. Screen content videos increasingly gain the popularity due to the rapid advances in cloud and multimedia technologies, which in turn requires highly efficient screen content compression. A recent standard, namely SCC is under development in JCT-VC, Joint Collaborative Team on Video Coding between ISO/IEC and ITU-T. In SCC, the most efficient new coding tool is Intra block copy (Intra BC). In this paper, we describe the Intra BC that has been proposed by the authors and adopted in the SCC standard and reference software for coding of screen content. Different from the conventional Intra prediction method where the prediction signal is derived from the spatially neighboring samples, the Intra BC mode greatly improves the prediction efficiency by fully exploiting the redundancy of repetitive patterns which typically appear in screen content. Experimental results suggest that the Intra BC mode can improve the coding efficiency significantly for typical screen content video sequences with 43.2% bit rate reduction on average. Joel Sole, Ying Chen 0011, Vadim Seregin, Marta Karczewicz |
DCC | 3 |
| 2015 | HEVC-Compatible Extensions for Advanced Coding of 3D and Multiview VideoabstractThis article provides an overview of standardized extensions of HEVC for the advanced coding of 3D and multiview video. In those extensions, new coding tools that better exploit the inter-view redundancy of the multiview texture videos have been developed. Additionally, dedicated tools for the improved coding of depth have been extensively studied and incorporated into the standard. In this paper, the performance of these extensions is assessed, and experimental results demonstrate notable gains in coding efficiency. Anthony Vetro, Ying Chen 0011, Karsten Müller 0001 |
DCC | 2 |
| 2015 | Texture based sub-PU motion inheritance for depth codingabstractThe 3D extension of High Efficiency Video Coding (HEVC) has been developed as a standard to code multiview data with depth information. In addition to the existing coding tools in HEVC, in the 3D extension, depth pictures can be coded with supplemental techniques typically designed to utilize the special characteristic of depth pictures. To better utilize the correlation between texture and depth pictures, in this paper, we present an enhanced motion prediction method for depth coding using the motion from the associated texture picture. Such a method inherits the motion information for sub-blocks of depth prediction units. The proposed method provides in average 3.2% bit rate reduction for a 3D video system, wherein the quality is evaluated by the synthesized views. Ying Chen 0011, Xin Zhao 0003 |
ICIP | 1 |
| 2014 | Unified wedgelet genenration for depth coding in 3D-HEVCabstractIn 3D-HEVC, bi-partition based modes, i.e., depth modeling modes (DMM), are applied for depth intra coding. With DMM, a depth prediction unit (PU) is partitioned into two parts using a Wedgelet pattern selected from a predefined large Wedgelet set, and the Wedgelet set is generated during both encoder and decoder initialization for each block size ranging from 4 ×4 to 32×32. The generation of Wedgelet sets involves relatively high complexity in terms of both storage and computation, especially for large block sizes. To simplify the Wedgelet generation, in this paper, we propose a unified Wedgelet generation method which derives large Wedgelet patterns by extending the primitive 4×4 Wedgelet pattern along its partition boundary line, and the storage of large Wedgelet patterns can be saved. Experimental results show almost no coding efficiency degradation using the proposed method, and the complexity of Wedgelet generation process is largely reduced in a unified manner. Xin Zhao 0003, Li Zhang 0006, Ying Chen 0011 |
ICIP | 3 |
| 2014 | Inter-view motion vector prediction for depth codingabstractThis paper presents an inter-view motion prediction technique for efficient compression of motion vectors of the depth views in 3D-HEVC. 3D-HEVC is an extension of HEVC standard for coding the multi-view video plus depth content, known as MVD. In MVD format, the motion characteristics of the adjacent views in the depth video are highly correlated. In this paper, we take benefit of this correlation and propose inter-view motion prediction technique, where motion information of the dependent depth views is predicted from the already coded motion information in a reference depth view. In addition, a novel method for deriving disparity vectors based on neighboring pixels is proposed in order to establish correspondences between the blocks in different depth views. Experimental results show that proposed inter-view motion prediction method provides an average bit-rate savings of 1.5% for the synthesized views when the motion information of the depth views is predicted without using the texture information. Vijayaraghavan Thirumalai, Li Zhang 0006, Ying Chen 0011 |
ICME | 3 |
| 2014 | Low complexity Neighboring Block based Disparity Vector Derivation in 3D-HEVCabstract3D-HEVC incorporates advanced inter-view prediction techniques based on more accurately derived disparity vector to better exploit the correlation between objects in different views. The efficient disparity vector derivation method, namely, Neighboring Block based Disparity Vector Derivation (NBDV) provides disparity without accessing any depth information. The NBDV has been developed as a part of the 3D-HEVC in Joint Collaborative Team on 3D Video Coding (JCT-3V) for several meeting cycles, and adopted as a common coding tool used for all the inter-view prediction techniques for high efficient coding of texture views. This paper presents a low complexity NBDV, which is the state-of-the-art disparity vector derivation method in 3D-HEVC. Je-Won Kang, Ying Chen 0011, Li Zhang 0006, Marta Karczewicz |
ISCAS | 2 |
| 2014 | Low-complexity advanced residual prediction design in 3D-HEVCabstractAdvanced residual prediction (ARP) is an efficient tool for 3D video coding by exploiting the residual correlation between views. In ARP, the residual predictor could be efficiently produced by aligning the motion information at the current view for motion compensation in the reference view. On the other hand, such on-the-fly residual predictor derivation process increases the complexity significantly due to the increased motion compensation steps. In this paper, a low-complexity ARP scheme is proposed. Experimental results demonstrate that the proposed scheme significantly reduces the decoding complexity of the original design in terms of both memory access and computational complexity while keeping comparable coding performance. Li Zhang 0006, Ying Chen 0011, Xiang Li 0003, Shanhua Xue |
ISCAS | 2 |
| 2014 | Inter-view motion prediction in 3D-HEVCabstractThis paper presents a novel inter-view motion prediction technique used in 3D-HEVC which provides efficient compression for motion vectors. The 3D extension of HEVC (High Efficiency Video Coding) namely 3D-HEVC is under development in JCT-3V for coding multi-view video and depth data. Inter-view motion prediction takes benefit of the inter-view correlation between views by inferring the motion information of a view from the already coded motion information in another view as well as the local disparity between two views. Experimental results show that inter-view motion prediction in 3D-HEVC provides an average bit-rate savings of 15% for the dependent views. Li Zhang 0006, Ying Chen 0011, Vijayaraghavan Thirumalai, Jian-Liang Lin, Yi-Wen Chen, Jicheng An, Shawmin Lei, Laurent Guillo, Thomas Guionnet, Christine Guillemot |
ISCAS | 2 |
| 2014 | Derived disparity vector based NBDV for 3D-AVCabstractIn the 3D video extension of H.264/AVC, namely 3D-AVC, Neighboring Based Disparity Vector (NBDV) derivation has been proposed to support multiview/stereo compatibility, therefore texture views can be decoded independently to depth views. NBDV generates a disparity vector for the current macroblock (MB) using the motion information of neighboring blocks, especially those coded with motion vectors pointing to inter-view reference pictures. In 3D-AVC, NBDV has been utilized to access minimum number of spatial and temporal neighboring blocks, therefore there is a high probability that NBDV does not derive an efficient disparity vector. This paper introduces a derived disparity vector scheme, wherein only one disparity vector derived from NBDV is maintained for the whole slice and it is used as the disparity vector of the current MB if NBDV does not derive one from neighboring blocks. Simulation results show that the proposed method provides 3.6% bit rate reduction for multiview coding. Xin Zhao 0003, Ying Chen 0011, Li Zhang 0006 |
VCIP | 2 |
| 2014 | Overview of the MVC + D 3D video coding standard
Ying Chen 0011, Miska M. Hannuksela, Teruhiko Suzuki, Shinobu Hattori |
J. Vis. Commun. Image Represent. | 1 |
| 2014 | Motion Hooks for the Multiview Extension of HEVCabstractMV-HEVC refers to the multiview extension of High Efficiency Video Coding (HEVC). At the time of writing, MV-HEVC was being developed by the Joint Collaborative Team on 3D Video Coding Extension Development (JCT-3V) of International Organization for Standardization (ISO)/International Electrotechnical Commission (IEC) Moving Picture Experts Group and ITU-T VCEG. Before HEVC itself was technically finalized in January 2013, the development of MV-HEVC had already started and it was decided that MV-HEVC would only contain high-level syntax changes compared with HEVC, i.e., no changes to block-level processes, to enable the reuse of the first-generation HEVC decoder hardware as is for constructing an MV-HEVC decoder with only firmware changes corresponding to the high-level syntax part of the codec. Consequently, any block-level process that is not necessary for HEVC itself but on the other hand is useful for MV-HEVC can only be enabled through so-called hooks. Motion hooks refer to techniques that do not have a significant impact on the HEVC single-view version 1 codec and can mainly improve MV-HEVC. This paper presents techniques for efficient MV-HEVC coding by introducing hooks into the HEVC design to accommodate inter-view prediction in MV-HEVC. These hooks relate to motion prediction, hence named motion hooks. Some of the motion hooks developed by the authors have been adopted into HEVC during its finalization. Simulation results show that the proposed motion hooks provide on average 4% of bitrate reduction for the views coded with inter-view prediction. Ying Chen 0011, Li Zhang 0006, Vadim Seregin, Ye-Kui Wang |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2013 | Scalable Video Coding Extension for HEVCabstractThis paper describes a scalable video codec that was submitted as a response to the joint call for proposals issued by ISO/IEC MPEG and ITU-T VCEG on HEVC scalable extension. The proposed codec uses a multi-loop decoding structure. Several inter-layer texture prediction methods are employed to remove the inter-layer redundancy. Inter-layer prediction is also used when coding enhancement layer syntax elements such as motion parameter and intra prediction mode, to further reduce bit overhead. Additionally, alternative transforms as well as adaptive coefficients scanning are used to code the prediction residues more efficiently. Experimental results are presented to demonstrate the effectiveness of the proposed scheme. When compared to HEVC single-layer coding, the additional rate overhead for the proposed scalable extension is 1.2% to 6.4% to achieve two layers of SNR and spatial scalability. Jianle Chen, Krishnakanth Rapaka, Xiang Li 0003, Vadim Seregin, Marta Karczewicz, Geert Van der Auwera, Joel Sole, Xianglin Wang, Chengjie Tu, Ying Chen 0011, Rajan L. Joshi |
DCC | 11 |
| 2013 | Advanced residual predction in 3D-HEVabstractInter-view residual prediction (IVRP) is employed t o efficiently code non-base texture views by exploiting the correlation of residues between two views in 3D video coding extension of HEVC(3D-HEVC). To further improve the performance of texture coding, advanced residual prediction (ARP) is proposed in this paper. In IVRP, residues in a non-base view are predicted from decoded residues in base view. In contrast, ARP makes the residual predictor for the non-base view block based on newly generated base-view residues by applying motion vector at the non-base view to the baseview. Moreover, an adaptive weighting factor is applied to reduce prediction error. Comprehensive simulations show that up to 6.0% luma BD-rate reduction was obtained over IVRP in the reference software of 3D-HEVC. Xiang Li 0003, Li Zhang 0006, Ying Chen 0011 |
ICIP | 3 |
| 2013 | Texture mode dependent depth coding in 3D-HEVCabstractIn 3D-HEVC, depth modeling modes (DMM) are applied for efficient intra depth coding. With DMM modes, a prediction unit is partitioned into two parts, and each part is predicted by a single value. In one DMM mode, the partition pattern is implicitly derived at the decoder by searching all pre-defined Wedgelet patterns on a Co-located Texture Luma Block (CTLB), which increases the decoding complexity drastically. To simplify the design of this DMM mode, in this paper, we propose to utilize the Intra Prediction Mode (IPM) of CTLB to largely skip unnecessary Wedgelet searches at the decoder. To achieve this goal, for each IPM, only a limited number of Wedgelet patterns are selected as candidates in this DMM mode. Experimental results demonstrate that, with almost no coding performance degradation, the proposed method significantly reduces the decoding complexity by skipping 90% of Wedgelet searches. The proposed method has been adopted by 3D-HEVC. Xin Zhao 0003, Ying Chen 0011, Li Zhang 0006, Marta Karczewicz |
ICIP | 2 |
| 2013 | Disparity vector based advanced inter-view prediction in 3D-HEVCabstractCoding multiview video content often benefits significantly from disparity compensation. More advanced inter-view predictions, to predict e.g., motion vectors and residues among views require disparity vector to identify blocks in different views corresponding to the same objects. Block-level disparity vectors in typically approaches are either transmitted thus with additional overhead, or predicted with sophisticated algorithms. This paper describes a new method to derive disparity vectors, based on which advanced inter-view predictions can be better supported. The proposed method derives a disparity vector of a block only from spatial and temporal neighboring blocks. It was implemented and adopted into the 3D-HEVC codec, which is currently developed by Joint Collaborative Team on 3D Video coding development (JCT-3V). Experimental results show that with the proposed method, advanced inter-view prediction in 3D-HEVC provides in average 5.2% bitrate saving of the overall bitrate, for multiview video coding. Li Zhang 0006, Ying Chen 0011, Marta Karczewicz |
ISCAS | 2 |
| 2013 | Neighboring block based disparity vector derivation for 3D-AVCabstract3D-AVC, being developed under Joint Collaborative Team on 3D Video Coding (JCT-3V), significantly outperforms the Multiview Video Coding plus Depth (MVC+D) which has no new macroblock level coding tools compared to Multiview video coding extension of H.264/AVC (MVC). However, for multiview compatible configuration, i.e., when texture views are decoded without accessing depth information, the performance of the current 3D-AVC is only marginally better than MVC+D. The problem is caused by the lack of disparity vectors which can be obtained only from the coded depth views in 3D-AVC. In this paper, a disparity vector derivation method is proposed by using the motion information of neighboring blocks and applied along with existing coding tools in 3D-AVC. The proposed method improves 3D-AVC in the multiview compatible mode substantially, resulting in about 20% bitrate reduction for texture coding. When enabling the so-called view synthesis prediction to further refine the disparity vectors, the performance of the proposed method is 31% better than MVC+D and even better than 3D-AVC under the best performing 3D-AVC configuration. Li Zhang 0006, Je-Won Kang, Xin Zhao 0003, Ying Chen 0011, Rajan L. Joshi |
VCIP | 4 |
| 2012 | Adaptive Depth edge sharpening for 3D video depth codingabstractIn 3D video systems with Multiview Video plus Depth (MVD) representation, intermediate views can be rendered from transmitted texture views and corresponding depth maps by techniques such as Depth Image Based Rendering (DIBR). Recent standardization activities in MPEG include the development for such MVD based 3DV codecs. One codec which is H.264/AVC based is called 3DV-ATM. Because of compression, reconstructed depth maps often have certain distortions, such as blurry depth edges, which can result in noticeable artifacts in the rendered views. In this paper, a method of adaptive depth edge sharpening is proposed for 3D video coding, based on the 3DV-ATM. The proposed adaptive depth edge filtering and smoothing along depth edge techniques adaptively sharpen blurry edges of the reconstructed depth frames caused by compression. Compared to the anchor software 3DV-ATM, the proposed method achieves about 7.2% bitrate reduction on average for the rendered view PSNRs versus overall bitrates with comparable runtime complexity. Rong Zhang 0014, Ying Chen 0011, Marta Karczewicz |
VCIP | 2 |
| 2012 | Overview of HEVC High-Level Syntax and Reference Picture ManagementabstractThe increasing proportion of video traffic in telecommunication networks puts an emphasis on efficient video compression technology. High Efficiency Video Coding (HEVC) is the forthcoming video coding standard that provides substantial bit rate reductions compared to its predecessors. In the HEVC standardization process, technologies such as picture partitioning, reference picture management, and parameter sets are categorized as “high-level syntax.” The design of the high-level syntax impacts the interface to systems and error resilience, and provides new functionalities. This paper presents an overview of the HEVC high-level syntax, including network abstraction layer unit headers, parameter sets, picture partitioning schemes, reference picture management, and supplemental enhancement information messages. Rickard Sjöberg, Ying Chen 0011, Akira Fujibayashi, Miska M. Hannuksela, Jonatan Samuelsson, Thiow Keng Tan, Ye-Kui Wang, Stephan Wenger |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2011 | MVC based scalable codec enhancing frame-compatible stereoscopic videoabstractExisting 3D video solutions take advantages of the 2D video infrastructure, by coding two views of the stereoscopic video content in a frame-compatible manner, such as side-by-side or top-bottom. Content providers and service providers are using the existing 2D video authorizing tools and delivery infrastructure to enhance the 2D video experience into 3D with relatively small additional cost. However, such a frame-compatible solution typically sacrifices the quality of each view, by providing a half-resolution representation. In the future, it is expected that more and more devices are capable of rendering full-resolution 1080p stereoscopic video content. Co-existing of services based on half-resolution stereoscopic video, full-resolution stereoscopic video, as well as full-resolution 2D video is expected in the marketplace. In this paper, a scalable codec is proposed to support the decoding and rendering of frame-compatible stereo, full-resolution 2D and full-resolution stereo video representations. Compared to codecs providing similar functionalities, significant coding gain can be achieved. Ying Chen 0011, Rong Zhang 0014, Marta Karczewicz |
ICME | 1 |
| 2010 | Depth-level-adaptive view synthesis for 3D videoabstractIn the multiview video plus depth (MVD) representation for 3D video, a depth map sequence is coded for each view. In the decoding end, a view synthesis algorithm is used to generate virtual views from depth map sequences. Many of the known view synthesis algorithms introduce rendering artifacts especially at object boundaries. In this paper, a depth-level-adaptive view synthesis algorithm is presented to reduce the amount of artifacts and to improve the quality of the synthesized images. The proposed algorithm introduces awareness of the depth level so that no pixel value in the synthesized image is derived from pixels of more than one depth level. Improvements on objective quality of the synthesized views were achieved in five out of eight test cases, while the subjective quality of the proposed method was similar to or better than that of the view synthesis method used by Moving Picture Experts Group (MPEG). Ying Chen 0011, Weixing Wan, Miska M. Hannuksela, Houqiang Li, Moncef Gabbouj |
ICME | 1 |
| 2009 | Spatial transcoding from Scalable Video Coding to H.264/AVCabstractScalable Video Coding (SVC) is backwards compatible to H.264/AVC in the sense that the base layer sub-bitstream is decodable by an H.264/AVC decoder. However, there are applications wherein it is desirable for an H.264/AVC decoder to obtain a higher resolution video representation than the base layer within SVC. In order to fulfill the needs of such application scenarios, transcoding of SVC enhancement layers to H.264/AVC is required. This paper presents a transcoding scheme that is capable of transcoding a spatial scalable SVC bitstreams to H.264/AVC bitstreams that provide high resolution than the H.264/AVC compliant base layer. To reduce the complexity at the transcoder, a fast mode decision (MD) process is proposed, wherein the original SVC macroblock coding modes and motion information are reused as much as possible. Experimental results show that proposed scheme performs elegantly compared with full-decoding-and-encoding transcoding with low computational complexity. Ye-Kui Wang, Ying Chen 0011, Houqiang Li |
ICME | 3 |
| 2009 | Regionally Adaptive Filtering for Asymmetric Stereoscopic Video CodingabstractIn asymmetric stereoscopic video coding, one view can be coded in a lower resolution of the other. In this scenario, stereoscopic video can be compressed with only moderately increased bandwidth and complexity compared to 2D monoview video coding. The subjective quality degradation of this scenario can be negligible compared to coding two views with original resolution. The low-resolution view can be predicted from the high-resolution view to achieve higher coding efficiency. In this paper, a regionally adaptive filtering algorithm is proposed to generate a predictor of a macroblock (MB) or MB partition of the low-resolution view from the high-resolution view. Different filters are applied for different picture regions. Disparity motion matching and clustering are applied in the encoder for generation of regionally adaptive filters. Simulation results show that the proposed algorithm results in up to 27% bit-rate saving compared with methods without adaptive filtering. Ying Chen 0011, Ye-Kui Wang, Moncef Gabbouj, Miska M. Hannuksela |
ISCAS | 1 |
| 2009 | Joint Texture and Depth Map Video Coding based on the Scalable Extension of H.264/AVCabstractDepth-Image-Based Rendering (DIBR) is widely used for view synthesis in 3D video applications. Compared with traditional 2D video applications, both the texture video and its associated depth map are required for transmission in a communication system that supports DIBR. To efficiently utilize limited bandwidth, coding algorithms, e.g. the Advanced Video Coding (H.264/AVC) standard, can be adopted to compress the depth map using the 4:0:0 chroma sampling format. However, when the correlation between texture video and depth map is exploited, the compression efficiency may be improved compared with encoding them independently using H.264/AVC. A new encoder algorithm which employs Scalable Video Coding (SVC), the scalable extension of H.264/AVC, to compress the texture video and its associated depth map is proposed in this paper. Experimental results show that the proposed algorithm can provide up to 0.97 dB gain for the coded depth maps, compared with the simulcast scheme, wherein texture video and depth map are coded independently by H.264/AVC. Siping Tao, Ying Chen 0011, Miska M. Hannuksela, Ye-Kui Wang, Moncef Gabbouj, Houqiang Li |
ISCAS | 2 |
| 2009 | Coding techniques in Multiview Video Coding and Joint Multiview Video ModelabstractSince early 2006, Joint Video Team has been devoting on the development of Multiview Video Coding (MVC) standard as an extension of H.264/AVC. This MVC standard has been finalized in 2008. During the standardization of MVC, there was also a project namely Joint Multiview Video Model (JMVM), which focused on the advanced coding tools that are potentially useful. Those coding tools adopted into JMVM, including illumination compensation and motion skip, have not been added into MVC specification. In this paper, coding techniques in MVC as well as the tools in JMVM are described and discussed, focusing on the coding efficiency. Ying Chen 0011, Miska M. Hannuksela, Antti Hallapuro, Moncef Gabbouj, Houqiang Li |
PCS | 1 |
| 2009 | Efficient hierarchical inter picture coding for H.264/AVC baseline profileabstractBi-predictive (B) slices are not supported in the Baseline profile of the Advanced Video Coding (H.264/AVC) standard, which results in a decreased coding efficiency compared with other profiles supporting B slices. However, many application standards, such as the mobile multimedia services specified by the Third Generation Partnership Project (3GPP), use only the Baseline profile for H.264/AVC. Therefore, it is worth investigating H.264/AVC coding when only intra (I) and inter (P) slices are supported. In this paper, a content-adaptive Quantization Parameter (QP) cascading scheme for the hierarchical P coding method compatible with Baseline profile of H.264/AVC is proposed. The proposed method is based on a picture-level QP optimization. The proposed method has a significantly better rate-distortion performance than the traditional IPPP coding structure and outperforms hierarchical P coding methods using fixed delta QP settings between temporal levels noticeably with up to 0.53 dB gain in average luminance Peak Signal-to-Noise Ratio (PSNR). Weixing Wan, Ying Chen 0011, Ye-Kui Wang, Miska M. Hannuksela, Houqiang Li, Moncef Gabbouj |
PCS | 2 |
| 2009 | Error Resilient Coding and Error Concealment in Scalable Video CodingabstractScalable video coding (SVC), which is the scalable extension of the H.264/AVC standard, was developed by the Joint Video Team (JVT) of ISO/IEC MPEG (Moving Picture Experts Group) and ITU-T VCEG (Video Coding Experts Group). SVC is designed to provide adaptation capability for heterogeneous network structures and different receiving devices with the help of temporal, spatial, and quality scalabilities. It is challenging to achieve graceful quality degradation in an error-prone environment, since channel errors can drastically deteriorate the quality of the video. Error resilient coding and error concealment techniques have been introduced into SVC to reduce the quality degradation impact of transmission errors. Some of the techniques are inherited from or applicable also to H.264/AVC, while some of them take advantage of the SVC coding structure and coding tools. In this paper, the error resilient coding and error concealment tools in SVC are first reviewed. Then, several important tools such as loss-aware rate-distortion optimized macroblock mode decision algorithm and error concealment methods in SVC are discussed and experimental results are provided to show the benefits from them. The results demonstrate that PSNR gains can be achieved for the conventional inter prediction (IPPP) coding structure or the hierarchical bi-predictive (B) picture coding structure with large group of pictures size, for all the tested sequences and under various combinations of packet loss rates, compared with the basic joint scalable video model (JSVM) design applying no error resilient tools at the encoder and only picture copy error concealment method at the decoder. Ying Chen 0011, Ye-Kui Wang, Houqiang Li, Miska M. Hannuksela, Moncef Gabbouj |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2008 | Picture-level adaptive filter for asymmetric stereoscopic videoabstractIn asymmetric stereoscopic video coding, one view is coded in a quarter of the resolution of the other and the low- resolution view is predicted from the high-resolution view. This way, stereoscopic video effect could be achieved with only moderately increased bandwidth and complexity. Inter-view prediction tools for generating the predictor of a maroblock (MB) or MB partition in the low-resolution view from the high-resolution view play a vital role for coding efficiency in asymmetric video coding. In this paper, we propose a method that applies an adaptive filter to generate picture-level adaptive inter-view predictors for MBs or MB partitions. At the encoder, a low complexity preprocessing module is built to find out the filters. Simulation results show that the proposed method provides a bit-rate saving of 26% at maximum and 5% on average. Ying Chen 0011, Ye-Kui Wang, Miska M. Hannuksela, Moncef Gabbouj |
ICIP | 1 |
| 2008 | Low-complexity asymmetric multiview video codingabstractMultiview video coding (MVC) is currently under development by the Joint Video Team (JVT) as an extension to Advanced Video Coding (H264/AVC). Based on the suppression theory in binocular vision, the fidelity of one of the two views of a stereoscopic display can be reduced without noticeable degradation of subjective quality. Thus, in MVC, a subset of views can be coded with lower spatial resolution at negligible cost to subjective quality. Due to different resolutions, a downsampling process is required in an MVC decoder in order to enable motion compensation (MC) between views. In this paper, a low-complexity MC algorithm is proposed for MVC to enable inter-view prediction between pictures with different resolutions. It requires lower memory consumption and lower computational complexity compared with the conventional downsampled inter-view prediction, while providing comparable efficiency, as shown by the simulation results. Ying Chen 0011, Shujie Liu 0001, Ye-Kui Wang, Miska M. Hannuksela, Houqiang Li, Moncef Gabbouj |
ICME | 1 |
| 2008 | Single-loop decoding for multiview video codingabstractMultiview video coding (MVC) is currently being standardized by the Joint Video Team as an extension of H264/AVC. When an MVC bitstream is decoded, some views (named target views) are to be displayed; some other views (named dependent views) may not be displayed but are needed for inter-view prediction of the target views. The original MVC design requires pictures of the dependent views to be fully decoded and stored. This entails both high decoding complexity and high memory consumption for the pictures in the views which are not intended for display, particularly when the number of dependent views is large. In this paper, a single-loop decoding (SLD) scheme is introduced to address these disadvantages. SLD requires only partial decoding of pictures in dependent views and thus significantly reduces decoding complexity and memory consumption. The proposed method is based on the so-called motion skip, wherein inter-view motion and coding mode prediction is exploited. Experimental results show that compared to coding schemes that require comparable complexity, significant compression gain can be achieved. For example, 25% bit-rate saving on average can be obtained compared to simulcast. Simulation results also show that the proposed SLD scheme provides a substantial reduction of complexity and memory size, at the expense of only a minor compression efficiency loss, compared with multiple-loop decoding MVC schemes. Ying Chen 0011, Ye-Kui Wang, Miska M. Hannuksela, Moncef Gabbouj |
ICME | 1 |
| 2008 | Frame loss error concealment for multiview video codingabstractThe Multiview Video Coding (MVC) standard is currently under development by the Joint Video Team as an extension of the Advanced Video Coding (H.264/AVC) standard. An MVC encoder compresses more than one viewpoint of a scene captured by different cameras. Redundancies between views can be used for inter-view prediction in encoding as well as error concealment in decoding. In this paper, a new algorithm utilizing motion information of pictures from other views to conceal a lost picture is proposed. The algorithm first derives motion information for a lost picture based on motion fields of pictures in adjacent views. Then, traditional motion compensation is invoked within the view containing the lost picture to derive a concealed frame. Experimental results show that the proposed algorithm can improve video quality with a negligible computational complexity overhead compared to simple temporal error concealment algorithms. Shujie Liu 0001, Ying Chen 0011, Ye-Kui Wang, Moncef Gabbouj, Miska M. Hannuksela, Houqiang Li |
ISCAS | 2 |