EDBT 2026 Demo / reviewers in the wild / expert
Hao Chen 0036
dblp:175/3324-36
· DBLP profile ↗
21ranked-venue papers
3as first author
17since 2021 · last 2026
0000-0002-1179-8199ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 16 · 2 first-author · 13 since 2021Computer networks · 4 · 4 since 2021Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 first-authorDatabases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Reinforced Rate Control for Neural Video Compression via Inter-Frame Rate-Distortion AwarenessabstractNeural video compression (NVC) has demonstrated superior compression efficiency, yet effective rate control remains a significant challenge due to complex temporal dependencies. Existing rate control schemes typically leverage frame content to capture distortion interactions, overlooking inter-frame rate dependencies arising from shifts in per-frame coding parameters. This often leads to suboptimal bitrate allocation and cascading parameter decisions. To address this, we propose a reinforcement‑learning (RL)‑based rate control framework that formulates the task as a frame‑by‑frame sequential decision process. At each frame, an RL agent observes a spatiotemporal state and selects coding parameters to optimize a long‑term reward that reflects rate‑distortion (R-D) performance and bitrate adherence. Unlike prior methods, our approach jointly determines bitrate allocation and coding configuration in a single step, independent of group‑of‑pictures (GOP) structure. Extensive experiments across diverse NVC architectures show that our method reduces the average relative bitrate error to 1.20 percent and achieves up to 13.45 percent bitrate savings at typical GOP sizes, outperforming existing approaches. In addition, our framework demonstrates improved robustness to content variation and bandwidth fluctuations with lower encoding/decoding overhead, making it highly suitable for practical deployment. Wuyang Cong, Junqi Shi, Lizhong Wang, Weijing Shi, Ming Lu 0003, Hao Chen 0036, Zhan Ma 0001 |
AAAI | 6 |
| 2026 | DynA: Breaking Fixed-Increase Rigidity in Google Congestion ControlabstractWebRTC’s Google Congestion Control (GCC) suffers from severe bandwidth underutilization during bandwidth recovery, owing to its rigid fixed-step rate increment strategy. To tackle this issue, we propose DynA, a dual-policy learning-augmented congestion control mechanism that dynamically optimizes continuous bitrate step sizes via reinforcement learning. By preserving GCC’s operational structure while introducing decoupled continuous-action policies with phased curriculum training, DynA enables more stable and fine-grained bitrate adaptation than conventional learning-based approaches. Preliminary evaluation results show that DynA improves bandwidth utilization by over 10% compared with vanilla GCC. By balancing aggressive bandwidth utilization with controlled queue growth, DynA achieves the best overall trade-off between throughput and overshoot among all evaluated baselines. Xuehao Wu, Bowei Xu, Hao Chen 0036, Zhan Ma 0001 |
APNet | 3 |
| 2026 | COOPASY: Cooperative Asynchronous Rate-Redundancy Control for Video Conferencing over High-RTT Weak Networks
Guixing Yang, Bowei Xu, Hao Chen 0036, Zhan Ma 0001 |
APNet | 4 |
| 2025 | Towards Loss-Resilient Image Coding for Unstable Satellite NetworksabstractGeostationary Earth Orbit (GEO) satellite communication demonstrates significant advantages in emergency short burst data services. However, unstable satellite networks, particularly those with frequent packet loss, present a severe challenge to accurate image transmission. To address it, we propose a loss-resilient image coding approach that leverages end-to-end optimization in learned image compression (LIC). Our method builds on the channel-wise progressive coding framework, incorporating Spatial-Channel Rearrangement (SCR) on the encoder side and Mask Conditional Aggregation (MCA) on the decoder side to improve reconstruction quality with unpredictable errors. By integrating the Gilbert-Elliot model into the training process, we enhance the model's ability to generalize in real-world network conditions. Extensive evaluations show that our approach outperforms traditional and deep learning-based methods in terms of compression performance and stability under diverse packet loss, offering robust and efficient progressive transmission even in challenging environments. Hongwei Sha, Muchen Dong, Quanyou Luo, Ming Lu 0003, Hao Chen 0036, Zhan Ma 0001 |
AAAI | 5 |
| 2025 | Adaptive Rate Control for Deep Video Compression with Rate-Distortion PredictionabstractDeep video compression has made significant progress in recent years, achieving rate-distortion performance that surpasses that of traditional video compression methods. However, rate control schemes tailored for deep video compression have not been well studied. In this paper, we propose a neural network-based$\lambda$-domain rate control scheme for deep video compression, which determines the coding parameter$\lambda$for each to-be-coded frame based on the rate-distortion-$\lambda\ (\mathrm{R}-\mathrm{D}-\lambda)$relationships directly learned from uncompressed frames, achieving high rate control accuracy efficiently without the need for pre-encoding. Moreover, this content-aware scheme is able to mitigate inter-frame quality fluctuations and adapt to abrupt changes in video content. Specifically, we introduce two neural network-based predictors to estimate the relationship between bitrate and$\lambda$, as well as the relationship between distortion and$\lambda$for each frame. Then we determine the coding parameter$\lambda$for each frame to achieve the target bitrate. Experimental results demonstrate that our approach achieves high rate control accuracy at the mini-GOP level with low time overhead and mitigates inter-frame quality fluctuations across video content of varying resolutions. Bowen Gu, Hao Chen 0036, Ming Lu 0003, Zhan Ma 0001 |
DCC | 2 |
| 2024 | Modeling the non-uniform retinal perception for viewport-dependent streaming of immersive video
Peiyao Guo, Xu Zhang 0006, Hao Chen 0036, Zhan Ma 0001 |
Multim. Syst. | 4 |
| 2024 | Improving Adaptive Real-Time Video Communication via Cross-Layer OptimizationabstractEffective Adaptive Bitrate (ABR) algorithm or policy is of paramount importance for Real-Time Video Communication (RTVC) amid this pandemic to pursue uncompromised quality of experience (QoE). Existing ABR methods mainly separate the network bandwidth estimation and video encoder control, and fine-tune video bitrate towards estimated bandwidth, assuming the maximization of bandwidth utilization yields the optimal QoE. However, the QoE of an RTVC system is jointly determined by the quality of the compressed video, fluency of video playback, and interaction delay. Solely maximizing the bandwidth utilization without comprehensively considering compound impacts incurred by both transport and video application layers, does not assure a satisfactory QoE. The decoupling of the transport and application layer further exacerbates the user experience due to codec-transport incoordination. This work, therefore, proposes the Palette, a reinforcement learning-based ABR scheme that unifies the processing of transport and video application layers to directly maximize the QoE formulated as the weighted function of video quality, stalling rate, and delay. To this aim, a cross-layer optimization is proposed to derive the fine-grained compression factor of the upcoming frame(s) using cross-layer observations like network conditions, video encoding parameters, and video content complexity. As a result, Palette manages to resolve the codec-transport incoordination and to best catch up with the network fluctuation. Compared with state-of-the-art schemes in real-world tests, Palette not only reduces 3.1%-46.3% of the stalling rate, 20.2%-50.8% of the delay but also improves 0.2%-7.2% of the video quality with comparable bandwidth consumption, under a variety of application scenarios. Yueheng Li, Hao Chen 0036, Bowei Xu, Zhan Ma 0001 |
IEEE Trans. Multim. | 2 |
| 2023 | Bamboo: Boosting Training Efficiency for Real-Time Video Streaming via Online Grouped Federated Transfer LearningabstractNo abstract available. Qianyuan Zheng, Hao Chen 0036, Zhan Ma 0001 |
APNet | 2 |
| 2023 | Anableps: Adapting Bitrate for Real-Time Communication Using VBR-encoded VideoabstractContent providers increasingly replace traditional constant bitrate with variable bitrate (VBR) encoding in real-time video communication systems for better video quality. However, VBR encoding often leads to large and frequent bitrate fluctuation, inevitably deteriorating the efficiency of existing adaptive bitrate (ABR) methods. To tackle it, we propose the Anableps to consider the network dynamics and VBR-encoding-induced video bitrate fluctuations jointly for deploying the best ABR policy. With this aim, Anableps uses sender-side information from the past to predict the video bitrate range of upcoming frames. Such bitrate range is then combined with the receiver-side observations to set the proper bitrate target for video encoding using a reinforcement-learning-based ABR model. As revealed by extensive experiments on a real-world trace-driven testbed, our Anableps outperforms the de facto GCC with significant improvement of quality of experience, e.g., 1.88× video quality, 57% less bitrate consumption, 85% less stalling, and 74% shorter interaction delay. Hao Chen 0036, Xun Cao, Zhan Ma 0001 |
ICME | 2 |
| 2023 | Mamba: Bringing Multi-Dimensional ABR to WebRTCabstractContemporary real-time video communication systems, such as WebRTC, use an adaptive bitrate (ABR) algorithm to assure high-quality and low-delay services, e.g., promptly adjusting video bitrate according to the instantaneous network bandwidth. However, target bitrate decisions in the network and bitrate control in the codec are typically incoordinated and simply ignoring the effect of inappropriate resolution and frame rate settings also leads to compromised results in bitrate control, thus devastatingly deteriorating the quality of experience (QoE). To tackle these challenges, Mamba, an end-to-end multi-dimensional ABR algorithm is proposed, which utilizes multi-agent reinforcement learning (MARL) to maximize the user's QoE by adaptively and collaboratively adjusting encoding factors including the quantization parameters (QP), resolution, and frame rate based on observed states such as network conditions and video complexity information in a video conferencing system. We also introduce curriculum learning to improve the training efficiency of MARL. Both the in-lab and real-world evaluation results demonstrate the remarkable efficacy of Mamba. Yueheng Li, Hao Chen 0036, Zhan Ma 0001 |
ACM Multimedia | 3 |
| 2023 | Karma: Adaptive Video Streaming via Causal Sequence ModelingabstractOptimal adaptive bitrate (ABR) decision depends on a comprehensive characterization of state transitions that involve interrelated modalities over time including environmental observations, returns, and actions. However, state-of-the-art learning-based ABR algorithms solely rely on past observations to decide the next action. This paradigm tends to cause a chain of deviations from optimal action when encountering unfamiliar observations, which consequently undermines the model generalization. Bowei Xu, Hao Chen 0036, Zhan Ma 0001 |
ACM Multimedia | 2 |
| 2023 | Quality-of-Experience Assessment for Ultra-Low Latency Live Streaming VideosabstractThe quality-of-experience (QoE) assessment is vitally important for optimizing video streaming systems. Different from traditional streaming video, ultra-low latency live streaming (ULL-LS) video is susceptible to fluctuations in compression quality, playback smoothness, and latency in the time-varying network environment, which makes QoE evaluation extremely complicated. However, most existing streaming QoE researches focus on the effect of video compression quality and rebuffering events, without considering the joint impact of compression quality, playback smoothness, and latency on the user QoE, resulting in inaccurate quality assessment for ULL-LS videos. In this paper, we establish the first continuous-time subjective evaluation database for ULL-LS videos (ULL-LSVD), aiming to study the comprehensive influence of various compression and transmission distortions on the user's QoE. Additionally, we propose an objective assessment model for ULL-LS videos based on correlation analysis. Compared with other well-known objective metrics, our model shows superiority in the QoE evaluation of ULL-LS videos. We believe that our database and analysis will promote the QoE modeling research for ULL-LS videos, and we will make it public soon. Mingliu Sun, Hao Chen 0036, Zhan Ma 0001 |
MMSP | 2 |
| 2023 | Cloud Game Video Coding Based On Human Eye Fixation PointabstractCloud Gaming enables users to run high-quality games on thin clients with limited graphics processing and data computing capabilities. Under the running mode of cloud games, all games are run on the server side, and the rendered game picture is compressed and sent to the user through the network. On the client side, the user's gaming device doesn't need any high-end processors or graphics cards, only basic video extraction capabilities. However, cloud gaming requires a high bandwidth connection to present a good-quality game picture to the user. At present, the main problem is limited bandwidth during transmission, which leads to poor quality of the game image received by users and poor user experience, which has become an important problem hindering the popularity of cloud games. To solve this problem, we observed that when playing a game, the user's eyes are not always focused on the entire picture, and due to the characteristics of visual perception, the user pays more attention to the area around the eye fixation point. In this paper, for the first time, we add information about user interactions with devices to the network to more accurately predict user fixation points. Using visual perception features, more bit rates are assigned to ROI regions near the user's fixation point. Our results show that our network is able to predict more accurate user fixation points, and that our approach can significantly improve the subjective quality of the area near the user's focus of the game frame at the same bitrate, with the VMAF scores is 2 to 3 points higher on average compared to H.264, the most commonly used standard encoder in cloud games today. Geng Wei, Tianjing Zhang, Ming Lu 0003, Hao Chen 0036 |
MMSP | 5 |
| 2023 | Improving ABR Performance for Short Video Streaming Using Multi-Agent Reinforcement Learning with Expert GuidanceabstractIn the realm of short video streaming, popular adaptive bitrate (ABR) algorithms developed for classical long video applications suffer from catastrophic failures because they are tuned to solely adapt bitrates. Instead, short video adaptive bitrate (SABR) algorithms have to properly determine which video at which bitrate level together for content prefetching, without sacrificing the users' quality of experience (QoE) and yielding noticeable bandwidth wastage jointly. Unfortunately, existing SABR methods are inevitably entangled with slow convergence and poor generalization. Thus, in this paper, we propose Incendio, a novel SABR framework that applies Multi-Agent Reinforcement Learning (MARL) with Expert Guidance to separate the decision of video ID and video bitrate in respective buffer management and bitrate adaptation agents to maximize the system-level utilized score modeled as a compound function of QoE and bandwidth wastage metrics. To train Incendio, it is first initialized by imitating the hand-crafted expert rules and then fine-tuned through the use of MARL. Results from extensive experiments indicate that Incendio outperforms the current state-of-the-art SABR algorithm with a 53.2% improvement measured by the utility score while maintaining low training complexity and inference time. Yueheng Li, Qianyuan Zheng, Hao Chen 0036, Zhan Ma 0001 |
NOSSDAV | 4 |
| 2022 | Robust Ultralow Bitrate Video Conferencing with Second Order Motion CoherencyabstractThe emergence of unsupervised deep image animation (DIA) has promised unprecedented prospects of video conferencing applications across ultralow-bandwidth networks. Existing DIA approaches rely on the First Order Motion (FOM) Model to combine compressed sparse motion features (SMFs) from motion driving frames and the appearance feature extracted from the source image for high-quality video generation. This work improves the FOM model by introducing the Second Order Motion (SOM) Coherency for better synthesis, with which we not only best assure the temporal smoothness that is not considered in FOM, but also enable the effective compensation of packet loss often encountered in real-life scenarios. Extensive experiments show that our method outperforms the HEVC-based conferencing with$\approx \mathbf{70}\%$BD-rate gains with network bandwidth$< \mathbf{10}$kbps, and surpasses state-of-art solution proposed by Konuko et al. with about 30 absolute percentage points improvement in BD-rate. Zhehao Chen, Ming Lu 0003, Hao Chen 0036, Zhan Ma 0001 |
MMSP | 3 |
| 2021 | Learned Resolution Scaling Powered Gaming-as-a-Service at ScaleabstractBuilt on the explosive advancement of cloud and telecommunication technologies, Gaming-as-a-Service (GaaS) or cloud gaming system is expected to revolutionize the traditional multi-billion video game market in the near future. This wave is analogous to the rise of live-video-streaming-based-Netflix to replace conventional DVD rental business for movies and TVs. In practice, a successful GaaS platform need to operate in a transparent mode without requiring substantial efforts from both content providers and end users, and offer the pristine quality of experience (QoE) at an affordable cost. Our analysis suggests that GaaS provisioning cost can be reduced significantly by enforcing the game video rendering and streaming at a lower resolution (so as to increase the user concurrency in the cloud and reduce the streaming bandwidth over the network). However, streaming video at a lower resolution may deteriorate the QoE. To maintain the client QoE at the level using the default-native resolution for streaming or even enhance it, we introduce the learned resolution scaling (LRS), which leverages the computational capabilities at clients/edges to restore/improve the reconstructed image/video quality via stacked deep neural networks (DNN). We integrate this LRS into a commercialized GaaS platform - AnyGame, to study its efficiency and complexity quantitatively. Extensive real-life experiments have shown that LRS-powered AnyGame offers the state-of-the-art performance, and the lower operational cost, paving the road for a potential success of GaaS over the Internet. Additionally, we dive into proposed LRS via ablation studies to further demonstrate its consistent performance, including the discussions on trade-off between efficiency and complexity, alternative training sets, etc. Hao Chen 0036, Ming Lu 0003, Zhan Ma 0001, Xu Zhang 0006, Yiling Xu, Qiu Shen, Wenjun Zhang 0001 |
IEEE Trans. Multim. | 1 |
| 2021 | Predicting the Perceptual Quality of Point Cloud: A 3D-to-2D Projection-Based ExplorationabstractPoint cloud is emerged as a promising media format to represent realistic 3D objects or scenes in applications, such as virtual reality, teleportation, etc. How to accurately quantify the subjective point cloud quality for application-driven optimization, however, is still a challenging and open problem. In this paper, we attempt to tackle this problem in a systematic means. First, we produce a fairly large point cloud dataset where ten popular point clouds are augmented with seven types of impairments (e.g., compression, photometry/color noise, geometry noise, scaling) at six different distortion levels, and organize a formal subjective assessment with tens of subjects to collect mean opinion scores (MOS) for all 420 processed point cloud samples (PPCS). We then try to develop an objective metric that can accurately estimate the subjective quality. Towards this goal, we choose to project the 3D point cloud onto six perpendicular image planes of a cube for the color texture image and corresponding depth image, and aggregate image-based global (e.g., Jensen-Shannon (JS) divergence) and local features (e.g., edge, depth, pixel-wise similarity, complexity) among all projected planes for a final objective index. Model parameters are fixed constants after performing the regression using a small and independent dataset previously published. The proposed metric has demonstrated the state-of-the-art performance for predicting the subjective point cloud quality compared with multiple full-reference and no-reference models, e.g., the weighted peak signal-to-noise ratio (PSNR), structural similarity (SSIM), feature similarity (FSIM) and natural image quality evaluator (NIQE). The dataset is made publicly accessible athttp://smt.sjtu.edu.cnorhttp://vision.nju.edu.cnfor all interested audiences. Qi Yang 0003, Hao Chen 0036, Zhan Ma 0001, Yiling Xu, Rongjun Tang, Jun Sun 0005 |
IEEE Trans. Multim. | 2 |
| 2020 | Efficient Mobile Video Streaming via Context-Aware RaptorQ-Based Unequal Error ProtectionabstractMobile video streaming systems typically apply the forward error correction (FEC) at the application layer to cope with packet-level transmission errors, which complements the bit-level correction mechanisms at the physical layer. However, most existing works fail to exploit the block-level dependencies in both intra and interframe coding modes of a single-layer compressed video, and thus are less efficient for the prevailing H.264/AVC and/or H.265/HEVC compatible single-layer video application. To this end, we propose a low-complexity FEC, i.e., context-aware RaptorQ (CA-RQ) with unequal error protection (UEP), to improve the error recovery performance of the singlelayer mobile video streaming, through incorporating the blocklevel dependencies in the compressed video data. We use a packet-level video transmission distortion model that considers the dependencies in both spatial and temporal domains, to quantify the importance of video packets within a group of pictures (GoP). The compressed video packets are categorized and grouped into several classes according to their importance to construct the CA-RQ code with the UEP property. We provide a theoretical analysis on redundancy allocation bounds to demonstrate the superior performance of proposed CA-RQ over the standard RaptorQ code. In the meantime, extensive simulations have shown that our scheme not only offers much better subjective visual quality with less than 50% additional redundant symbols as compared to the Macroblock-Based UEP (MB-UEP) scheme, but also outperforms the MB-UEP and classical equal error protection (EEP)-based schemes, by a 0.45%'5.71% and 0.94%'6.78% margin, respectively, in reconstructed quality evaluated using the structural similarity (SSIM) index, across a reasonable range of redundancy proportions. Hao Chen 0036, Xu Zhang 0006, Yiling Xu, Zhan Ma 0001, Wenjun Zhang 0001 |
IEEE Trans. Multim. | 1 |
| 2019 | T-Gaming: A Cost-Efficient Cloud Gaming System at ScaleabstractCloud gaming (CG) system could pursue both high-quality gaming experience via intensive computing, and ultimate convenience anywhere at anytime through any energy-constrained mobile devices. Despite the abundance of efforts devoted, state-of-the-art CG systems still suffer from multiple key limitations: expensive deployment cost, high bandwidth consumption and unsatisfied quality of experience (QoE). As a result, existing works are not widely adopted in reality. This paper proposes a Transparent Gaming framework called T-Gaming that allows users to play any popular high-end desktop/console games on-the-fly over the Internet. T-Gaming utilizes the off-the-shelf consumer GPUs without resorting to the expensive proprietary GPU virtualization (vGPU) technology to reduce the deployment cost. Moreover, it enables prioritized video encoding based on the human visual feature to reduce the bandwidth consumption without noticeable visual quality degradation. Last but not least, T-Gaming adopts adaptive real-time streaming based on deep reinforcement learning (RL) to improve user's QoE. To evaluate the performance of T-Gaming, we implement and test a prototype system in the real world. Compared with the existing cloud gaming systems, T-Gaming not only reduces the expense per user by 75 percent hardware cost reduction and 14.3 percent network cost reduction, but also improves the normalized average QoE by 3.6-27.9 percent. Hao Chen 0036, Xu Zhang 0006, Yiling Xu, Ju Ren 0001, Jingtao Fan, Zhan Ma 0001, Wenjun Zhang 0001 |
IEEE Trans. Parallel Distributed Syst. | 1 |
| 2016 | A new AL-FEC coding scheme with limited feedbackabstractFor the next generation mobile video broadcasting, especially in-band solutions that serves the mobile devices, a limited feedback scheme via cellular channel polling is feasible to give accurate real-time information on the broadcast receivers' channel erasure rate, and decoding buffer status. In this work, we propose an AL-FEC coding degree scheme based on this feedback, to achieve a better decode efficiency and save the code redundancy. Simulation results demonstrate the effectiveness of this solution, and open up new opportunities in the next generation broadcasting system design. Wei Huang 0012, Hao Chen 0036, Yiling Xu, Zhu Li 0001, Wenjun Zhang 0001 |
MMSP | 2 |
| 2016 | Single-input-multiple-ouput transcoding for video streamingabstractIn this work, a single input multiple output (SIMO) transcoding architecture is proposed. SIMO will benefit the mobile edge computing (such as HTTP Live Streaming requiring multiple copies of the video streams at different quality levels) without resorting to the legacy transcoding that video stream is completed decoded and encoded multiple times without exploring the compressed information. Leveraging the information encoded in the existing video streams, we could reduce the search candidates when transcoding the high quality bitstream to other versions with reduced quality level. As the first step, we have demonstrated the SIMO idea with bit rate shaping (i.e., bit rate transcoding) only scenario. It has shown more than 2x complexity reduction without quality loss using the common test conditions. Hao Zhang 0032, Hao Chen 0036, Yiling Xu, Zhan Ma 0001 |
MMSP | 5 |