VLDB 2026 Research / reviewers in the wild / expert
Nuowen Kan
dblp:226/2477
· DBLP profile ↗
25ranked-venue papers
7as first author
21since 2021 · last 2026
0000-0002-6028-1284ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 19 · 5 first-author · 15 since 2021Artificial intelligence and machine learning · 5 · 5 since 2021Computer networks · 2 · 2 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Error-Resilient Learned Video Compression with Channel Importance-Aware Redundancy AllocationabstractReal-time video communication is critically hampered by packet loss. Building on the neural video codec (NVC) and exploiting the statistics of its latent representations, we propose a channel importance-aware redundancy framework that eliminates insignificant channels and reallocates high redundancy to protect important channels for loss-resilient video coding. We first develop a channel-wise importance evaluation method based on the in-distribution (IND) region length criterion [1] for latent representations of prediction residuals and motion vectors produced by NVCs. Consequently, we derive a heuristic that adapts the number of protected channels to the packet loss rate (PLR). More channels are retained at low PLRs to improve video quality, while only a few critical channels are kept and assigned higher redundancy for enhanced robustness at high PLRs. Experimental results show that the proposed method provides more graceful quality degradation and maintains significant performance advantages, particularly under high PLRs, as illustrated in Figure 1. Nuowen Kan, Wenrui Dai, Junni Zou, Hongkai Xiong |
DCC | 2 |
| 2026 | MERINA+: Improving Generalization for Neural Video Adaptation via Information-Theoretic Meta-Reinforcement LearningabstractAdaptive bitrate (ABR) streaming is a popular technique used to improve the quality of experience (QoE) for users who watch videos online, which, for example, can provide a smoother video playback by dynamically adjusting the requested video quality with associated bitrate according to the constrained yet diverse network conditions. Recently, learning-based ABR algorithms have achieved a notable performance gain with lower inference overhead than the conventional heuristic or model-based baselines. However, their performance may degrade significantly in an unseen network environment with time-varying and heterogeneous throughput dynamics. For a better generalization, in this paper, we propose a meta-reinforcement learning (meta-RL)-based neural ABR algorithm that is able to quickly adapt its policy to these unseen throughput dynamics. Specifically, we propose a model-free system framework comprising an inference network and a policy network. The inference network infers distribution of the latent representation for underlying dynamics based on the recent throughout context, while the policy network is trained to quickly adapt to the changing throughout dynamics with the sampled latent representation. To effectively learn the inference network and meta-policy on mixed dynamics of the practical ABR scenarios, we further design a variational information bottleneck theory-based loss function for training the inference and policy networks, whose objective is to strike a trade-off between brevity of the latent representation and expressiveness of the meta-policy. We also derive a theoretically necessary condition for the bitrate versions that yield higher long-term QoE, based on which a dynamic action pruning strategy is further developed for practical implementation. This pruning strategy can not only prevent unsafe policy outputs in midst of unseen throughput dynamics, but may also reduce the computational complexity of model-based ABR algorithms. Finally, the meta-training and meta-adaptation procedures of our proposed algorithm are implemented across a range of throughput dynamics. The empirical evaluations on various datasets containing real-world network traces verify that our algorithm surpasses the state-of-the-art ABR algorithms, particularly in terms of the average chunk QoE and fast adaptation across out-of-distribution throughput traces. Nuowen Kan, Yuankun Jiang, Wenrui Dai, Junni Zou, Hongkai Xiong, Laura Toni |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2026 | Point Cloud Attribute Compression With Geometry-Aware Lifting-Based Multiscale NetworksabstractPoint cloud attribute compression is challenged by fitting the attribute signals living on irregular geometric structures. Existing methods cannot achieve compact multiscale representation for high-fidelity reconstruction using the handcrafted transforms or deep learning-based techniques. In this paper, we propose a novel geometry-aware lifting-based multiscale network via spatial-channel lifting scheme for point cloud attribute compression. The proposed network cascades geometry-aware spatial lifting to reduce spatial redundancy by adaptively capturing irregular geometric structures and progressive channel lifting to progressively reduce channel-wise redundancy in multiscale representation. Furthermore, we design the split, predict, and update operations for geometry-aware spatial lifting to fully exploit the geometry information representing irregular structures. We develop geometry-aware adaptive split to equally split input points with significance scores indicating their dependencies, and propose geometry-aware cross-attention filtering for the predict and update operations for decorrelation based on geometry information. To our best knowledge, this paper achieves the first lifting-based learned transform for point cloud compression that enjoys reversibility guarantees of multiscale representation to enhance rate-distortion performance. Experimental results show that the proposed framework achieves state-of-the-art performance on extensive point cloud datasets, and outperforms latest MPEG G-PCC standard and most recent deep learning based methods. Xin Li 0165, Wenrui Dai, Nuowen Kan, Junni Zou, Hongkai Xiong |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | Stabilizing and Accelerating Autofocus with Expert Trajectory Regularized Deep Reinforcement LearningabstractAutofocus is a crucial component of modern digital cameras. While recent learning-based methods achieve state-of-the-art in focus prediction accuracy, they unfortunately ignore the potential focus hunting phenomenon of back-and-forth lens movement in the multi-step focusing procedure. To address this, in this paper, we propose an expert regularized deep reinforcement learning (DRL)-based approach for autofocus, which can utilize the sequential information of lens movement trajectory to both enhance the multi-step in-focus prediction accuracy and reduce the chance of focus hunting. Our method generally follows an actor-critic framework. To accelerate the DRL’s training with a higher sample efficiency, we initialize the policy with a pre-trained single-step prediction network, where the network is further improved by modifying the output of absolute in-focus position distribution to the relative lens movement distribution to establish a better mapping between input images and lens movement. To further stabilize DRL’s training with a lower occurrence of focus hunting in the resulting lens movement trajectory, we generate some offline trajectories based on prior knowledge to avoid focus hunting, which are then leveraged as an offline dataset of expert trajectories to regularize the actor network’s training. Empirical evaluations show that our method outperforms those learning-based methods on public benchmarks, with higher single- and multi-step prediction accuracy, and a significant reduction of focus hunting rate. Shouhang Zhu, Yuankun Jiang, Nuowen Kan, Wenrui Dai, Junni Zou, Hongkai Xiong |
CVPR | 5 |
| 2025 | On Disentangled Training for Nonlinear Transform in Learned Image CompressionabstractLearned image compression (LIC) has demonstrated superior rate-distortion (R-D) performance compared to traditional codecs, but is challenged by training inefficiency that could incur more than two weeks to train a state-of-the-art model from scratch. Existing LIC methods overlook the slow convergence caused by compacting energy in learning nonlinear transforms. In this paper, we first reveal that such energy compaction consists of two components, \emph{i.e.}, feature decorrelation and uneven energy modulation. On such basis, we propose a linear auxiliary transform (AuxT) to disentangle energy compaction in training nonlinear transforms. The proposed AuxT obtains coarse approximation to achieve efficient energy compaction such that distribution fitting with the nonlinear transforms can be simplified to fine details. We then develop wavelet-based linear shortcuts (WLSs) for AuxT that leverages wavelet-based downsampling and orthogonal linear projection for feature decorrelation and subband-aware scaling for uneven energy modulation. AuxT is lightweight and plug-and-play to be integrated into diverse LIC models to address the slow convergence issue. Experimental results demonstrate that the proposed approach can accelerate training of LIC models by 2 times and simultaneously achieves an average 1\% BD-rate reduction. To our best knowledge, this is one of the first successful attempt that can significantly improve the convergence of LIC with comparable or superior rate-distortion performance. Wenrui Dai, Maida Cao, Nuowen Kan, Junni Zou, Hongkai Xiong |
ICLR | 5 |
| 2025 | A Generalizable and Expressive Meta-Diffusion Policy for RTC Bandwidth PredictionabstractIn real-time communication (RTC) systems, a congestion control (CC) is indispensable to guarantee the end user’s quality of experience (QoE), the key of which is to predict the bottleneck link capacity and thus set the target bitrate for the sender. Existing deep reinforcement learning (DRL)-based bandwidth prediction methods, though improving the bandwidth prediction performance based on the pre-collected offline dataset, are usually designed for a single RTC scenario with poor generalization ability to other scenarios. What’s worse, the policy learned by these DRL methods can hardly capture the complex multi-modal distribution of underlying optimal policy. To address this, in this paper, we propose an offline DRL-based meta-diffusion policy, where we pre-train an encoder to extract the effective feature embedding of network dynamics, and further learn an expressive meta-policy based on the diffusion model to predict the bandwidth conditioned on the observed network states and extracted network dynamics features. Experiments demonstrate the enhanced accuracy and generalization of our meta-diffusion policy over existing landmark schemes, with a performance gain of 7.84% and 27.12%, in terms of the bandwidth prediction error rate and overestimate rate, respectively. Nuowen Kan, Wenrui Dai, Junni Zou, Hongkai Xiong |
ICME | 2 |
| 2025 | Noise Conditional Variational Score DistillationabstractWe propose Noise Conditional Variational Score Distillation (NCVSD), a novel method for distilling pretrained diffusion models into generative denoisers. We achieve this by revealing that the unconditional score function implicitly characterizes the score function of denoising posterior distributions. By integrating this insight into the Variational Score Distillation (VSD) framework, we enable scalable learning of generative denoisers capable of approximating samples from the denoising posterior distribution across a wide range of noise levels. The proposed generative denoisers exhibit desirable properties that allow fast generation while preserve the benefit of iterative refinement: (1) fast one-step generation through sampling from pure Gaussian noise at high noise levels; (2) improved sample quality by scaling the test-time compute with multi-step sampling; and (3) zero-shot probabilistic inference for flexible and controllable sampling. We evaluate NCVSD through extensive experiments, including class-conditional image generation and inverse problem solving. By scaling the test-time compute, our method outperforms teacher diffusion models and is on par with consistency models of larger sizes. Additionally, with significantly fewer NFEs than diffusion-based methods, we achieve record-breaking LPIPS on inverse problems. Xinyu Peng, Nuowen Kan, Wenrui Dai, Junni Zou, Hongkai Xiong |
ICML | 5 |
| 2025 | 3DGabSplat: 3D Gabor Splatting for Frequency-adaptive Radiance Field Rendering
Junyu Zhou 0001, Wenrui Dai, Junni Zou, Nuowen Kan, Hongkai Xiong |
ACM Multimedia | 6 |
| 2025 | Learning to Optimize Low-Latency Live Streaming from Expertise: An Offline Meta-Reinforcement Learning ApproachabstractIn low-latency live streaming (LLLS), the adaptive bitrate algorithm plays a critical role in optimizing encoding bitrates to meet the strict end-to-end latency requirement, which is mostly several seconds or less. Existing methods usually rely on a fixed preset latency target, limiting their generalization capabilities across the heterogeneous scenarios. To address this, we propose an offline meta-reinforcement learning-based bitrate decision framework for LLLS, which incorporates the expertise of multiple state-of-the-art LLLS algorithms from their collected experience and constructs a universal bitrate selection policy via the offline training. Specifically, we establish an RL-based policy that adaptively adjusts the network throughput measurements and selects the encoding bitrates for LLLS based on the network conditions. The policy’s performance is enhanced by improving the robustness of short-term throughput estimations. To leverage the expertise of current LLLS algorithms, we collect their decision-making trajectories, and directly train the policy on these offline trajectories with implicit Q-learning. In addition, a meta-RL paradigm is further adopted to learn a universal policy that performs uniformly well across the heterogeneous network conditions and varying target latencies. Experimental results demonstrate that, compared to the other baselines, the proposed method achieves at least a 5.8% performance gain in terms of the overall quality of experience (QoE), across the diverse target-latency scenarios in the real-world throughput traces. Yuhui Du, Nuowen Kan, Junni Zou, Wenrui Dai, Qingli Li, Hongkai Xiong |
VCIP | 2 |
| 2025 | Efficient Conditional Entropy Coding for Learned Progressive Image and Video CompressionabstractProgressive coding adapts to reliable image and video transmission over unstable network with fluctuating bandwidth with truncatable bitstreams produced by layer-wise conditional entropy coding. However, recent learned progressive image and video compression methods solely leverage table lookup or real-time computation for conditional entropy coding, and suffer from dramatically increased storage and computational overheads due to the expansion of codewords. To address this problem, in this paper, we propose an efficient conditional entropy coding scheme for learned progressive coding that incorporates the fast lookup-table and low-memory real-time computation based on statistical distribution of the codewords. Specifically, we build a smaller lookup table that stores only the symbols with high occurrence frequency and achieve real-time computation of those symbols with low occurrence frequency. Furthermore, we seamlessly integrate the proposed scheme into existing learned image and video compression frameworks for enhanced coding efficiency. Experimental results show that the proposed scheme accelerates both encoding and decoding processes for learned progressive image and video compression. Remarkably, we for the first time successfully extend the proposed scheme to learned video compression to demonstrate its feasibility. Haolin Luan, Shuoyu Ma, Wenrui Dai, Nuowen Kan, Junni Zou, Hongkai Xiong |
VCIP | 5 |
| 2025 | Learnable Non-Uniform Quantization With Sampling-Based Optimization for Variable-Rate Learned Image CompressionabstractVariable-rate coding is challenging but indispensable for learned image compression (LIC) that is in nature characterized by nonlinear transform coding (NTC). Existing methods for variable-rate LIC are restricted by the non-smooth quantization process with zero gradients almost everywhere, and consequently, suffer from training-test gap and degraded rate-distortion (R-D) performance. To address this problem, in this paper, we propose sampling-based optimization for training NTC models along with non-uniform quantizers. Different from gradient-based optimization, the proposed sampling-based optimization first randomly samples the parameters from Gaussian distributions with progressively reduced variance and then selects the optimal parameters with a R-D indicator. On the basis of sampling-based optimization, we develop a learnable non-uniform dead-zone quantizer by adaptively refining the quantization steps for variable-rate coding with nonlinear transforms. Furthermore, we incorporate the learnable dead-zone quantizer to achieve a variable-rate LIC model with enhanced R-D performance and design rate and distortion control algorithms to adapt to dynamic network conditions. Experimental results show that the proposed method achieves state-of-the-art R-D performance in variable-rate image compression. It obtains an average 8.82% BD-rate reduction compared to latest versatile video compression (VVC) standard, and simultaneously achieves precise rate and distortion control with an average variation of 0.0087 bpp in bit-rates and 0.1265 dB in distortion on the Kodak dataset. Wenrui Dai, Nuowen Kan, Junni Zou, Hongkai Xiong |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | Task-Adapted Learnable Embedded Quantization for Scalable Human-Machine Image CompressionabstractImage compression for both human and machine vision has become prevailing to accommodate to rising demands for machine-machine and human-machine communications. Scalable human-machine image compression is recently emerging as an efficient alternative to simultaneously achieve high accuracy for machine vision in the base layer and obtain high-fidelity reconstruction for human vision in the enhancement layer. However, existing methods achieve scalable coding with heuristic mechanisms, which cannot fully exploit the inter-layer correlations and evidently sacrifice rate-distortion performance. In this paper, we propose task-adapted learnable embedded quantization to address this problem in an analytically optimized fashion. We first reveal the relationship between the latent representations for machine and human vision and demonstrate that optimal representation for machine vision can be approximated with post-training optimization on the learned representation for human vision. On such basis, we propose task-adapted learnable embedded quantization that leverages learnable step predictor to adaptively determine the optimal quantization step for diverse machine vision tasks such that inter-layer correlations between representations for human and machine vision are sufficiently exploited using embedded quantization. Furthermore, we develop a human-machine scalable coding framework by incorporating the proposed embedded quantization into pre-trained learned image compression models. Experimental results demonstrate that the proposed framework achieves state-of-the-art performance on machine vision tasks like object detection, instance segmentation, and panoptic segmentation with negligible loss in rate-distortion performance for human vision. Shuoyu Ma, Wenrui Dai, Nuowen Kan, Fan Cheng 0002, Junni Zou, Hongkai Xiong |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | GDPlan: Generative Network Planning via Graph Diffusion ModelabstractNetwork planning is crucial to facilitate network service under limited network operation costs. However, adapting the network topology (i.e., connections and capacities for physical and IP links) to time-varying and stochastic network states is challenging due to high computational overhead of problem solving. Existing deep learning based methods suffer from prohibitive computational complexity due to simultaneously solving IP link capacities and allocated traffics or iteratively exploring search space with excessive samples via heuristic rules. In this paper, we propose GDPlan, the first generative framework that leverages conditional graph diffusion model to address this challenge. To achieve high solving efficiency, GDPlan decouples network planning into two separate stages, i.e., generation of diverse high-quality solutions to IP link capacity and refinement of the solutions under feasibility constraints. Specifically, we develop generation-oriented graph signal modeling that reformulates solving IP link capacities as generation of graph topology conditioned on the graph signals related to traffic demands, physical links and network costs. Consequently, energy-guided controllable graph generation is achieved to capture significant structural patterns of graph topology with score-based graph diffusion model and produce diverse solutions with a guarantee of feasibility under specific objectives or constraints. GDPlan successfully achieves graph generation based solution to network planning with an order of magnitude higher solving speed than the commercial solver Gurobi. Experimental results demonstrate that, compared with Gurobi, the proposed GDPlan obtains lowest 3.7% average gap with only about 7.5% average running time. Nuowen Kan, Sa Yan, Junni Zou, Wenrui Dai, Xing Gao 0005, Hongkai Xiong |
IEEE Trans. Netw. | 1 |
| 2024 | Task-Oriented Multi-Bitstream Optimization for Image Compression and Transmission via Optimal TransportabstractImage compression for machine vision exhibits various rate-accuracy performance across different downstream tasks and content types. An efficient utilization of constrained network resource for achieving an optimal overall task performance has thus recently attracted a growing attention. In this paper, we propose Tombo, a task-oriented image compression and transmission framework that efficiently identifies the optimal encoding bitrate and routing scheme for multiple image bitstreams delivered simultaneously for different downstream tasks. Specifically, we study the characteristics of image rate-accuracy performance for different machine vision tasks, and formulate the task-oriented joint bitrate and routing optimization problem for multi-bitstreams as a multi-commodity network flow problem with the time-expanded network modeling. To ensure consistency between the encoding bitrate and routing optimization, we also propose an augmented network that incorporates the encoding bitrate variables into the routing variables. To improve computational efficiency, we further convert the original optimization problem to a multi-marginal optimal transport problem, and adopt a Sinkhorn iteration-based algorithm to quickly obtain the near-optimal solution. Finally, we adapt Tombo to efficiently deal with the dynamic network scenario where link capacities may fluctuate over time. Empirical evaluations on three typical machine vision tasks and four real-world network topologies demonstrate that Tombo achieves a comparable performance to the optimal one solved by the off-the-shelf solver Gurobi, with a 5x ~ 114× speedup. Sa Yan, Nuowen Kan, Wenrui Dai, Junni Zou, Hongkai Xiong |
ACM Multimedia | 2 |
| 2024 | Improving Generalization in Federated Learning with Model-Data Mutual Information Regularization: A Posterior Inference ApproachabstractMost of existing federated learning (FL) formulation is treated as a point-estimate of models, inherently prone to overfitting on scarce client-side data with overconfident decisions. Though Bayesian inference can alleviate this issue, a direct posterior inference at clients may result in biased local posterior estimates due to data heterogeneity, leading to a sub-optimal global posterior. From an information-theoretic perspective, we propose FedMDMI, a federated posterior inference framework based on model-data mutual information (MI). Specifically, a global model-data MI term is introduced as regularization to enforce the global model to learn essential information from the heterogeneous local data, alleviating the bias caused by data heterogeneity and hence enhancing generalization. To make this global MI tractable, we decompose it into local MI terms at the clients, converting the global objective with MI regularization into several locally optimizable objectives based on local data. For these local objectives, we further show that the optimal local posterior is a Gibbs posterior, which can be efficiently sampled with stochastic gradient Langevin dynamics methods. Finally, at the server, we approximate sampling from the global Gibbs posterior by simply averaging samples from the local posteriors. Theoretical analysis provides a generalization bound for FL w.r.t. the model-data MI, which, at different levels of regularization, represents a federated version of the bias-variance trade-off. Experimental results demonstrate a better generalization behavior with better calibrated uncertainty estimates of FedMDMI. Nuowen Kan, Wenrui Dai, Junni Zou, Hongkai Xiong |
NeurIPS | 3 |
| 2024 | Successor Feature-Based Transfer Reinforcement Learning for Video Rate Adaptation With Heterogeneous QoE PreferencesabstractIn adaptive video streaming, the design of an adaptive bitrate (ABR) strategy is critical for the quality-of-experience (QoE) perceived by users. Though current learning-based ABR algorithms achieve state-of-the-art performance for users with a given QoE metric setting for training, they may unfortunately suffer the poor generalization issue for other users with different QoE preferences. Besides, how to quantitatively characterize the distinct QoE preference for a user has also not been extensively studied yet. In this paper, we propose STEER, a successor feature-based transfer reinforcement learning framework for fast learning the ABR strategies on heterogeneous QoE preferences. Specifically, we first develop a QoE preference analysis scheme to infer the personal QoE preference of a single user based on the user's actual viewing history. We then formulate the personalized QoE maximization problem as a reinforcement learning (RL) task, which optimizes the ABR strategy to maximize the overall QoE perceived by the user. Further, we model the QoE maximization problem for multiple users with heterogeneous QoE preferences as a multi-task RL problem, with each task distinguished by the user-distinct QoE preference. To efficiently address this problem, the proposed STEER solves for each RL-based ABR task by learning its optimal successor feature (SF) function, which can be exploited as shared knowledge across tasks to facilitate the transfer between tasks. With SF functions, STEER can quickly evaluate the optimal policies of previously learned tasks on a new task, and further use the generalized policy improvement operation to obtain a jumpstart policy. Both theoretically and empirically, we show that this jumpstart policy is a good initial policy with a performance guarantee for better generalization in the new task, and can also lead to a faster convergence to the optimal policy of the new task. Kexin Tang, Nuowen Kan, Yuankun Jiang, Wenrui Dai, Junni Zou, Hongkai Xiong |
IEEE Trans. Multim. | 2 |
| 2023 | Doubly Robust Augmented Transfer for Meta-Reinforcement LearningabstractMeta-reinforcement learning (Meta-RL), though enabling a fast adaptation to learn new skills by exploiting the common structure shared among different tasks, suffers performance degradation in the sparse-reward setting. Current hindsight-based sample transfer approaches can alleviate this issue by transferring relabeled trajectories from other tasks to a new task so as to provide informative experience for the target reward function, but are unfortunately constrained with the unrealistic assumption that tasks differ only in reward functions. In this paper, we propose a doubly robust augmented transfer (DRaT) approach, aiming at addressing the more general sparse reward meta-RL scenario with both dynamics mismatches and varying reward functions across tasks. Specifically, we design a doubly robust augmented estimator for efficient value-function evaluation, which tackles dynamics mismatches with the optimal importance weight of transition distributions achieved by minimizing the theoretically derived upper bound of mean squared error (MSE) between the estimated values of transferred samples and their true values in the target task. Due to its intractability, we then propose an interval-based approximation to this optimal importance weight, which is guaranteed to cover the optimum with a constrained and sample-independent upper bound on the MSE approximation error. Based on our theoretical findings, we finally develop a DRaT algorithm for transferring informative samples across tasks during the training of meta-RL. We implement DRaT on an off-policy meta-RL baseline, and empirically show that it significantly outperforms other hindsight-based approaches on various sparse-reward MuJoCo locomotion tasks with varying dynamics and reward functions. Yuankun Jiang, Nuowen Kan, Wenrui Dai, Junni Zou, Hongkai Xiong |
NeurIPS | 2 |
| 2022 | Improving Generalization for Neural Adaptive Video Streaming via Meta Reinforcement LearningabstractIn this paper, we present a meta reinforcement learning (Meta-RL)-based neural adaptive bitrate streaming (ABR) algorithm that is able to rapidly adapt its control policy to the changing network throughput dynamics. Specifically, to allow rapid adaptation, we discuss the necessity of detaching the inference of throughput dynamics with the universal control mechanism that is in essence shared by all potential throughput dynamics for neural ABR algorithms. To meta-learn the ABR policy, we then build up a model-free system framework, composed of a probabilistic latent encoder that infers the underlying dynamics from the recent throughput context, and a policy network that is conditioned on latent variable and learns to quickly adapt to new environments. Additionally, to address the difficulties caused by training the policy on mixed dynamics, on-policy RL (or imitation learning) algorithms are suggested for policy training, with a mutual information-based regularization to make the latent variable more informative about the policy. Finally, we implement our algorithm's meta-training and meta-adaptation procedures under a variety of throughput dynamics. Empirical evaluations on different QoE metrics and multiple datasets containing real-world network traces demonstrate that our algorithm outperforms state-of-the-art ABR algorithms, in terms of the performance on the average chunk QoE, consistency and fast adaptation across a wide range of throughput patterns. Nuowen Kan, Yuankun Jiang, Wenrui Dai, Junni Zou, Hongkai Xiong |
ACM Multimedia | 1 |
| 2022 | RAPT360: Reinforcement Learning-Based Rate Adaptation for 360-Degree Video Streaming With Adaptive Prediction and TilingabstractTile-based rate adaption can improve the quality of experience (QoE) for adaptive 360-degree video streaming under constrained network conditions, which, however, is a challenging problem due to the requirements of accurate prediction for users’ viewports and optimal bitrate allocation for tiles. In this paper, we propose a strategy that deploys reinforcement learning-based Rate Adaptation with adaptive Prediction and Tiling for 360-degree video streaming, named RAPT360, to address these challenges. Specifically, to improve the accuracy of the state-of-the-art viewport prediction approaches, we fit the time-varying Laplace distribution-based probability density function of the prediction error for different prediction lengths. On the basis of that, we develop a viewport identification method to determine the viewport area of a user depending on the buffer occupancy, where the obtained viewport can cover the real viewport with any given probability confidence level. We then propose a viewport-aware adaptive tiling scheme to improve the bandwidth efficiency, where three types of tile granularities are allocated according to the shape and position of the 2-D projection of that viewport. By establishing an adaptive streaming model and QoE metric specific to 360-degree videos, we finally formulate the rate adaptation problem for tile-based 360-degree video streaming as a non-linear discrete optimization problem that targets at maximizing the long-term user QoE under a bandwidth-constrained network. To efficiently solve this problem, we model the rate adaptation logic as a Markov decision process (MDP) and employ the deep reinforcement learning (DRL)-based algorithm to dynamically learn the optimal bitrate allocation of tiles. Extensive experimental results show that RAPT360 achieves a performance gain of at least 1.47 dB on average chunk QoE, including a video quality improvement of at least 1.33 dB, in comparison to the existing strategies for tile-based adaptive 360-degree video streaming. Nuowen Kan, Junni Zou, Wenrui Dai, Hongkai Xiong |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2021 | Uncertainty-aware robust adaptive video streaming with bayesian neural network and model predictive controlabstractIn this paper, we propose BayesMPC, an uncertainty-aware robust adaptive bitrate (ABR) algorithm on the basis of Bayesian neural network (BNN) and model predictive control (MPC). Specifically, to improve the capacity of learning transition probability of the network throughput, we adopt a BNN-based predictor that is able to predict the statistical distribution of future throughput from the past throughput by not only considering the aleatoric uncertainty (e.g., noise), but also capturing the epistemic uncertainty incurred by lack of adequate training samples. We further show that by using the negative log-likelihood loss function to train this BNN-based throughput predictor, the generalization error can be minimized with the guarantee of PAC-Bayesian theorem. Rather than a point estimate, the learnt uncertainty can contribute to a confidence region for the future throughput, the lower bound of which then leads to an uncertainty-aware robust MPC strategy to maximize the worst-case user quality-of-experience (QoE) w.r.t. this confidence region. Finally, experimental results on three real-world network trace datasets validate the efficiency of both the proposed BNN-based predictor and uncertainty-aware robust MPC strategy, and demonstrate the superior performance compared to other baselines, in terms of both the overall QoE performance and generalization across all ranges of heterogeneous network and user conditions. Nuowen Kan, Caiyi Yang, Wenrui Dai, Junni Zou, Hongkai Xiong |
NOSSDAV | 1 |
| 2021 | Multi-User Adaptive Video Delivery Over Wireless Networks: A Physical Layer Resource-Aware Deep Reinforcement Learning ApproachabstractIn this paper, we investigate the adaptive video delivery for multiple users over time-varying and mutually interfering multi-cell wireless networks. The key research challenge is to jointly design the physical-layer resource allocation scheme and application-layer rate adaptation logic, such that the users' long-term fair quality of experience (QoE) can be maximized. Due to the timescale mismatch between these two layers and the asynchrony of user requests, however, it is difficult to directly model the cross-layer stochastic control problem by using a reinforcement learning framework. To address this difficulty, we propose a novel two-level decision framework where an optimization-based beamforming scheme (performed at the base stations) and a deep reinforcement learning (DRL)-based rate adaptation scheme (performed at the user terminals) are, respectively, developed, such that a highly complex long-term multi-user QoE fairness problem is decomposed into some relatively simple problems and solved effectively. Our strategy represents a significant departure from the existing schemes with consideration of either a short-term multi-user QoE maximization or a long-term single-user point-to-point QoE maximization. Extensive simulations demonstrate that the proposed cross-layer design is effective and promising. Kexin Tang, Nuowen Kan, Junni Zou, Xiao Fu 0001, Mingyi Hong 0001, Hongkai Xiong |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Deep Reinforcement Learning-based Rate Adaptation for Adaptive 360-Degree Video StreamingabstractIn this paper, we propose a deep reinforcement learning (DRL)-based rate adaptation algorithm for adaptive 360-degree video streaming, which is able to maximize the quality of experience of viewers by adapting the transmitted video quality to the time-varying network conditions. Specifically, to reduce the possible switching latency of the field of view (FoV), we design a new QoE metric by introducing a penalty term for the large buffer occupancy. A scalable FoV method is further proposed to alleviate the combinatorial explosion of the action space in the DRL formulation. Then, we model the rate adaptation logic as a Markov decision process and employ the DRL-based algorithm to dynamically learn the optimal video transmission rate. Simulation results show the superior performance of the proposed algorithm compared to the existing algorithms. Nuowen Kan, Junni Zou, Kexin Tang, Hongkai Xiong |
ICASSP | 1 |
| 2019 | A Server-Side Optimized Hybrid Multicast-Unicast Strategy for Multi-User Adaptive 360-Degree Video StreamingabstractThe head-mounted display (HMD) for 360-degree videos cannot be shared by multiple users, which significantly increases the bandwidth consumption when multiple users request the same video content simultaneously. This paper proposes a server-side hybrid multicast-unicast strategy for multi-user adaptive 360-degree video streaming, aiming to ensure the overall quality of experience (QoE) of the users in a bandwidth-constrained environment. A framework is established to realise this strategy, through which the delivery mode (multicast or unicast) and the corresponding bitrate for each tile can be jointly adapted to the dynamics of users' throughputs and field of views (FoVs). To maximize the overall QoE, a user clustering method is proposed to classify the users into several multicast clusters. Based on these clusters, we then formulate the joint delivery mode selection and rate adaptation problem as a non-linear integer programming problem which can be solved using the steepest ascent gradient algorithm. Finally, experiments are carried out to verify the performance of the proposed strategy. Nuowen Kan, Chengming Liu, Junni Zou, Hongkai Xiong |
ICIP | 1 |
| 2019 | Multiuser Video Streaming Rate Adaptation: A Physical Layer Resource-Aware Deep Reinforcement Learning ApproachabstractIn this paper, we propose a cross-layer decision framework for multiuser adaptive video delivery over time-varying and mutually interfering wireless cellular network. The key idea is to synthetically design the physical-layer optimization-based beamforming scheme (performed at the base stations) and the application-layer deep reinforcement learning (DRL)-based rate adaptation scheme (performed at the user terminals), so that a very complex multi-user overall fair long-term quality of experience (QoE) maximization problem can be decomposed to two layers and solved effectively. Extensive simulations show that the proposed cross-layer design is effective and promising. Kexin Tang, Nuowen Kan, Junni Zou, Xiao Fu 0001, Mingyi Hong 0001, Hongkai Xiong |
VCIP | 2 |
| 2018 | Server-Side Rate Adaptation for Multi-User 360-Degree Video StreamingabstractHow to balance the tradeoff between the user experience and bandwidth utilization emerges a critical challenge for multi-user 360-degree video adaptive streaming. This paper studies the server-side rate adaptation strategy for multiple users which are competing for the server bandwidth capacity. A tile visibility probability model is established, by which the tiles are classified into predicted, marginal and invisible types. A fine-grained rate adaptation problem is formulated as a nonlinear integer programming (NIP) problem, which aims at maximizing the video quality and navigation smoothness for multiple users. Thereafter, a steepest ascent algorithm with feasible starting point is developed to solve the proposed NIP problem in polynomial time. Finally, simulation results verify the performance of the proposed rate adaptation strategy. Chengming Liu, Nuowen Kan, Junni Zou, Qin Yang 0002, Hongkai Xiong |
ICIP | 2 |