Jian Qian

dblp:232/2479 · DBLP profile ↗
← Back
29ranked-venue papers
5as first author
23since 2021 · last 2026
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 19 · 4 first-author · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 8 · 2 first-author · 5 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Self-Normalized Martingales and Uniform Regret Bounds for Linear Regression
abstract
Self-normalized martingale inequalities lie at the heart of confidence ellipsoids for online least squares and, more broadly, many bandit and reinforcement-learning results. Yet existing vector and scalar results typically rely on bounded covariates and an explicit regularization matrix, producing bounds that are \emph{not scale-invariant}: although the self-normalized quantity is scale-invariant by definition, its standard upper bounds are not. We characterize when scale-invariant upper bounds on self-normalized martingales are possible. Without further assumptions, we prove that nontrivial scale-invariant bounds exist only in dimension $d=1$; moreover, in $d=1$ we obtain $O(\log T)$ scale-invariant self-normalized bounds without any assumptions on the covariates. In contrast, for $d>1$ we show that no nontrivial scale-invariant bound can hold in full generality. We then connect this dichotomy to \emph{doubly-uniform} regret in online linear regression (i.e., regret bounds that are simultaneously independent of the covariate scale and the comparator norm) and use it to resolve the open question of Gaillard, Gerchinovitz, Huard, and Stoltz, \emph{“Uniform regret bounds over $\mathbb{R}^d$ for the sequential linear regression problem with the square loss”} (ALT 2019): in $d=1$ we give an explicit algorithm with $O(\log T)$ doubly-uniform regret, whereas for $d>1$ sublinear doubly-uniform regret is impossible. Finally, under a natural \emph{smoothness} condition (bounded Radon–Nikodym derivatives of the conditional covariate laws with respect to a fixed base measure), we recover sublinear regret for $d>1$ without bounded covariates and derive a self-normalized concentration inequality free of the usual regularization penalties, yielding arguably a first natural scale-invariant bound for adaptive, non-i.i.d. vector martingales.
Jian Qian, Alexander Rakhlin, Nikita Zhivotovskiy
COLT2
2026 Refined Risk Bounds for Unbounded Losses via Transductive Priors
abstract
We revisit the sequential variants of linear regression with the squared loss, classification problems with hinge loss, and logistic regression, all characterized by unbounded losses in the setup where no assumptions are made on the magnitude of design vectors and the norm of the optimal vector of parameters. The key distinction from existing results lies in our assumption that the set of design vectors is known in advance (though their order is not), a setup sometimes referred to as transductive online learning. While this assumption might seem similar to fixed design regression or denoising, we demonstrate that the sequential nature of our algorithms allows us to convert our bounds into statistical ones with random design without making any additional assumptions about the distribution of the design vectors-an impossibility for standard denoising results. Our key tools are based on the exponential weights algorithm with carefully chosen transductive (design-dependent) priors, which exploit the full horizon of the design vectors, as well as additional aggregation tools that address the possibly unbounded norm of the vector of the optimal solution. Our classification regret bounds have a feature that is only attributed to bounded losses in the literature: they depend solely on the dimension of the parameter space and on the number of rounds, independent of the design vectors or the norm of the optimal solution. For linear regression with squared loss, we further extend our analysis to the sparse case, providing sparsity regret bounds that depend additionally only on the magnitude of the response variables. We argue that these improved bounds are specific to the transductive setting and unattainable in the worst-case sequential setup. Our algorithms, in several cases, have polynomial-time approximations and reduce to sampling with respect to log-concave measures instead of aggregating over hard-to-construct epsilon-covers of classes.
Jian Qian, Alexander Rakhlin, Nikita Zhivotovskiy
J. Mach. Learn. Res.1
2026 Wind power forecasting using ensemble learning based on coupled multiscale decomposition and regression enhancement
Zilong Ma, Tiefeng Zhang, Jian Qian
Knowl. Based Syst.4
2025 Bridging Multiple Worlds: Multi-marginal Optimal Transport for Causal Partial-identification Problem
abstract
Under the prevalent potential outcome model in causal inference, each unit is associated with multiple potential outcomes but at most one of which is observed, leading to many causal quantities being only partially identified. The inherent missing data issue echoes the multi-marginal optimal transport (MOT) problem, where marginal distributions are known, but how the marginals couple to form the joint distribution is unavailable. In this paper, we cast the causal partial identification problem in the framework of MOT with $K$ margins and $d$-dimensional outcomes and obtain the exact partial identified set. In order to estimate the partial identified set via MOT, statistically, we establish a convergence rate of the plug-in MOT estimator for the $\ell_2$ cost function stemming from the variance minimization problem and prove it is minimax optimal for arbitrary $K$ and $d \le 4$. We also extend the convergence result to general quadratic objective functions. Numerically, we demonstrate the efficacy of our method over synthetic datasets and several real-world datasets where our proposal consistently outperforms the baseline by a significant margin (over 70%). In addition, we provide efficient off-the-shelf implementations of MOT with general objective functions.
Zijun Gao, Shu Ge, Jian Qian
AISTATS3
2025 PillarHist: A Quantization-aware Pillar Feature Encoder based on Height-aware Histogram
abstract
Real-time and high-performance 3D object detection plays a critical role in autonomous driving and robotics. Recent pillar-based 3D object detectors have gained significant attention due to their compact representation and low computational overhead, making them suitable for onboard deployment and quantization. However, existing pillar-based detectors still suffer from information loss along height dimension and large numerical distribution difference during pillar feature encoding (PFE), which severely limits their performance and quantization potential. To address above issue, we first unveil the importance of different input information during PFE and identify the height dimension as a key factor in enhancing 3D detection performance. Motivated by this observation, we propose a heightaware pillar feature encoder, called PillarHist. Specifically, PillarHist statistics the discrete distribution of points at different heights within one pillar with the information entropy guidance. This simple yet effective design greatly preserves the information along the height dimension while significantly reducing the computation overhead of the PFE. Meanwhile, PillarHist also constrains the arithmetic distribution of PFE input to a stable range, making it quantization-friendly. Notably, PillarHist operates exclusively within the PFE stage to enhance performance, enabling seamless integration into existing pillar-based methods without introducing complex operations. Extensive experiments show the effectiveness of PillarHist in terms of both efficiency and performance.
Sifan Zhou, Zhihang Yuan, Xing Hu 0010, Jian Qian
CVPR5
2025 The Photoacoustic Quality-Enhancement Neural Network Processor with the Scalable and End-to-End Architecture by Improving the Sparsity Level
abstract
Recent advancements have marked significant progress in photoacoustic imaging as an effective method for acquiring deep bio-tissue visuals in modern medical clinical therapy and the efficacy of U-Net and its variants has been established for imaging quality enhancement in this field. Unlike common computer vision datasets such as ImageNet [1] and PASCAL VOC [2], biomedical images exhibit highly structured patterns, low spatial resolution, and single-channel modality, as shown in Fig. 1. Additionally, the U-Net parameters trained for medical super-resolution tasks demonstrate a high sparsity ratio, making them suitable for implementation on edge-computing platforms. Therefore, developing an energy-efficient photoacoustic imaging setup in this area is a natural progression. However, this development is constrained by the current neural network architectures, which are built around a U-Net backbone. The multi-stage feature extractor, skip connection integration across different blocks, and the encoder-decoder backbone design pose significant challenges to cutting-edge computational hardware platforms. In this study, a scalable, sparsity-supported neural network accelerator architecture for bio-tissue imaging quality enhancement is proposed to meet the stringent requirements of latency and energy efficiency, as depicted in Fig. 2. This architecture achieves desired performance improvements by exploring the sparsity possibilities in neural network during the training process and implementing an end-to-end pixel-first hardware design to minimize data movement and support sparsity computation. Compared with the state-of-the-art related works, this optimized architecture has achieved minimum on-chip storage overhead and the fastest frame for the application of photoacoustic imaging quality enhancement. The scalable architecture has also been implemented on a Xilinx XCZU9EG FPGA and attains a performance of PSNR@ 24 dB and a frame rate of 164 fps at a working frequency of 250 MHz.
Zhengyuan Zhang 0002, Caijie Liang, Boyi Dong, Yange Wang, Zhongzhiguang Lu, Xiangjun Yin, Shenglong Zhuo, Yifan Wu 0009, Yingjie Cao, Tianyang Zhou, Jian Qian, Patrick Chiang 0001, Lei Qiu 0002, Yuanjin Zheng
ISCAS14
2025 Evolution of Information in Interactive Decision Making: A Case Study for Multi-Armed Bandits
abstract
We study the evolution of information in interactive decision making through the lens of a stochastic multi-armed bandit problem. Focusing on a fundamental example where a unique optimal arm outperforms the rest by a fixed margin, we characterize the optimal success probability and mutual information over time. Our findings reveal distinct growth phases in mutual information---initially linear, transitioning to quadratic, and finally returning to linear---highlighting curious behavioral differences between interactive and non-interactive environments. In particular, we show that optimal success probability and mutual information can be decoupled, where achieving optimal learning does not necessarily require maximizing information gain. These findings shed new light on the intricate interplay between information and learning in interactive decision making.
Yuzhou Gu, Yanjun Han, Jian Qian
NeurIPS3
2025 PINR: A physics-integrated neural representation for dynamic fluid scenes
Sifan Zhou, Xiaobo Lu, Jian Qian
Neurocomputing5
2024 Sub-SA: Strengthen In-Context Learning via Submodular Selective Annotation
abstract
In-context learning (ICL) leverages in-context examples as prompts for the predictions of Large Language Models (LLMs). These prompts play a crucial role in achieving strong performance. However, the selection of suitable prompts from a large pool of labeled examples often entails significant annotation costs. To address this challenge, we propose Sub-SA (Submodular Selective Annotation), a submodule-based selective annotation method. The aim of Sub-SA is to reduce annotation costs while improving the quality of in-context examples and minimizing the time consumption of the selection process. In Sub-SA, we design a submodular function that facilitates effective subset selection for annotation and demonstrates the characteristics of monotonically and submodularity from the theoretical perspective. Specifically, we propose RPR (Reward and Penalty Regularization) to better balance the diversity and representativeness of the unlabeled dataset attributed to a reward term and a penalty term, respectively. Consequently, the selection for annotations can be effectively addressed with a simple yet effective greedy search algorithm based on the submodular function. Finally, we apply the similarity prompt retrieval to get the examples for ICL. Compared to existing selective annotation approaches, Sub-SA offers two main advantages. (1.) Sub-SA operates in an end-to-end, unsupervised manner, and significantly reduces the time consumption of the selection process (from hours-level to millisecond-level). (2.) Sub-SA enables a better balance between data diversity and representativeness and obtains state-of-the-art performance. Meanwhile, the theoretical support guarantees their reliability and scalability in practical scenarios. Extensive experiments conducted on diverse models and datasets demonstrate the superiority of Sub-SA over previous methods, achieving millisecond(ms)-level time selection and remarkable performance gains. The efficiency and effectiveness of Sub-SA make it highly suitable for real-world ICL scenarios. Our codes are available at https://github.com/JamesQian11/SubSA
Jian Qian, Sifan Zhou, Ruizhi Hun, Patrick Chiang 0001
ECAI1
2024 TCEKG: A Temporal and Causal Event Knowledge Graph for Power Distribution Network Fault Diagnosis
Feilong Liao, Jianye Huang 0002, Qichuan Liu, Xinjie Peng, Xinxin Wu, Jian Qian
ICIC (12)7
2024 The Non-linear F-Design and Applications to Interactive Learning
abstract
We propose a generalization of the classical G-optimal design concept to non-linear function classes. The criterion, termed F -design, coincides with G-design in the linear case. We compute the value of the optimal design, termed the F-condition number, for several non-linear function classes. We further provide algorithms to construct designs with a bounded F -condition number. Finally, we employ the F-design in a variety of interactive machine learning tasks, where the design is naturally useful for data collection or exploration. We show that in four diverse settings of confidence band construction, contextual bandits, model-free reinforcement learning, and active learning, F-design can be combined with existing approaches in a black-box manner to yield state-of-the-art results in known problem settings as well as to generalize to novel ones.
Alekh Agarwal, Jian Qian, Alexander Rakhlin, Tong Zhang 0001
ICML2
2024 Assouad, Fano, and Le Cam with Interaction: A Unifying Lower Bound Framework and Characterization for Bandit Learnability
abstract
We develop a unifying framework for information-theoretic lower bound in statistical estimation and interactive decision making. Classical lower bound techniques---such as Fano's method, Le Cam's method, and Assouad's lemma---are central to the study of minimax risk in statistical estimation, yet are insufficient to provide tight lower bounds for \emph{interactive decision making} algorithms that collect data interactively (e.g., algorithms for bandits and reinforcement learning). Recent work of Foster et al. provides minimax lower bounds for interactive decision making using seemingly different analysis techniques from the classical methods. These results---which are proven using a complexity measure known as the \emph{Decision-Estimation Coefficient} (DEC)---capture difficulties unique to interactive learning, yet do not recover the tightest known lower bounds for passive estimation. We propose a unified view of these distinct methodologies through a new lower bound approach called \emph{interactive Fano method}. As an application, we introduce a novel complexity measure, the \emph{Fractional Covering Number}, which facilitates the new lower bounds for interactive decision making that extend the DEC methodology by incorporating the complexity of estimation. Using the fractional covering number, we (i) provide a unified characterization of learnability for \emph{any} stochastic bandit problem, (ii) close the remaining gap between the upper and lower bounds in Foster et al. (up to polynomial factors) for any interactive decision making problem in which the underlying model class is convex.
Dylan J. Foster, Yanjun Han, Jian Qian, Alexander Rakhlin, Yunbei Xu
NeurIPS4
2024 Online Estimation via Offline Estimation: An Information-Theoretic Framework
abstract
The classical theory of statistical estimation aims to estimate a parameter of interest under data generated from a fixed design (''offline estimation''), while the contemporary theory of online learning provides algorithms for estimation under adaptively chosen covariates (''online estimation''). Motivated by connections between estimation and interactive decision making, we ask: is it possible to convert offline estimation algorithms into online estimation algorithms in a black-box fashion? We investigate this question from an information-theoretic perspective by introducing a new framework, Oracle-Efficient Online Estimation (OEOE), where the learner can only interact with the data stream indirectly through a sequence of offline estimators produced by a black-box algorithm operating on the stream. Our main results settle the statistical and computational complexity of online estimation in this framework. $\bullet$ Statistical complexity. We show that information-theoretically, there exist algorithms that achieve near-optimal online estimation error via black-box offline estimation oracles, and give a nearly-tight characterization for minimax rates in the OEOE framework. $\bullet$ Computational complexity. We show that the guarantees above cannot be achieved in a computationally efficient fashion in general, but give a refined characterization for the special case of conditional density estimation: computationally efficient online estimation via black-box offline estimation is possible whenever it is possible via unrestricted algorithms. Finally, we apply our results to give offline oracle-efficient algorithms for interactive decision making.
Dylan J. Foster, Yanjun Han, Jian Qian, Alexander Rakhlin
NeurIPS3
2024 How Does Variance Shape the Regret in Contextual Bandits?
abstract
We consider realizable contextual bandits with general function approximation, investigating how small reward variance can lead to better-than-minimax regret bounds. Unlike in minimax regret bounds, we show that the eluder dimension $d_{\text{elu}}$$-$a measure of the complexity of the function class$-$plays a crucial role in variance-dependent bounds. We consider two types of adversary: (1) Weak adversary: The adversary sets the reward variance before observing the learner's action. In this setting, we prove that a regret of $\Omega( \sqrt{ \min (A, d_{\text{elu}}) \Lambda } + d_{\text{elu}} )$ is unavoidable when $d_{\text{elu}} \leq \sqrt{A T}$, where $A$ is the number of actions, $T$ is the total number of rounds, and $\Lambda$ is the total variance over $T$ rounds. For the $A\leq d_{\text{elu}}$ regime, we derive a nearly matching upper bound $\tilde{O}( \sqrt{ A\Lambda } + d_{\text{elu} } )$ for the special case where the variance is revealed at the beginning of each round. (2) Strong adversary: The adversary sets the reward variance after observing the learner's action. We show that a regret of $\Omega( \sqrt{ d_{\text{elu}} \Lambda } + d_{\text{elu}} )$ is unavoidable when $\sqrt{ d_{\text{elu}} \Lambda } + d_{\text{elu}} \leq \sqrt{A T}$. In this setting, we provide an upper bound of order $\tilde{O}( d_{\text{elu}}\sqrt{ \Lambda } + d_{\text{elu}} )$. Furthermore, we examine the setting where the function class additionally provides distributional information of the reward, as studied by Wang et al. (2024). We demonstrate that the regret bound $\tilde{O}(\sqrt{d_{\text{elu}} \Lambda} + d_{\text{elu}})$ established in their work is unimprovable when $\sqrt{d_{\text{elu}} \Lambda} + d_{\text{elu}}\leq \sqrt{AT}$. However, with a slightly different definition of the total variance and with the assumption that the reward follows a Gaussian distribution, one can achieve a regret of $\tilde{O}(\sqrt{A\Lambda} + d_{\text{elu}})$.
Zeyu Jia, Jian Qian, Alexander Rakhlin, Chen-Yu Wei
NeurIPS2
2024 Offline Oracle-Efficient Learning for Contextual MDPs via Layerwise Exploration-Exploitation Tradeoff
abstract
Motivated by the recent discovery of a statistical and computational reduction from contextual bandits to offline regression \citep{simchi2020bypassing}, we address the general (stochastic) Contextual Markov Decision Process (CMDP) problem with horizon $H$ (as known as CMDP with $H$ layers). In this paper, we introduce a reduction from CMDPs to offline density estimation under the realizability assumption, i.e., a model class $\mathcal{M}$ containing the true underlying CMDP is provided in advance. We develop an efficient, statistically near-optimal algorithm requiring only $O(H \log T)$ calls to an offline density estimation algorithm (or oracle) across all $T$ rounds. This number can be further reduced to $O(H \log \log T)$ if $T$ is known in advance. Our results mark the first efficient and near-optimal reduction from CMDPs to offline density estimation without imposing any structural assumptions on the model class. A notable feature of our algorithm is the design of a layerwise exploration-exploitation tradeoff tailored to address the layerwise structure of CMDPs. Additionally, our algorithm is versatile and applicable to pure exploration tasks in reward-free reinforcement learning.
Jian Qian, Haichen Hu, David Simchi-Levi
NeurIPS1
2023 A Novel Approach to Analyzing Defects: Enhancing Knowledge Graph Embedding Models for Main Electrical Equipment
Jianye Huang 0002, Jian Qian, Longqiang Yi, Jinhu Li, Jiangsheng Huang
ICIC (5)3
2023 Improve Knowledge Graph Completion for Diagnosing Defects in Main Electrical Equipment
Jianye Huang 0002, Jian Qian, Yuyou Weng, Guoqing Lin
ICIC (5)2
2023 Model-Free Reinforcement Learning with the Decision-Estimation Coefficient
abstract
We consider the problem of interactive decision making, encompassing structured bandits and reinforcement learning with general function approximation. Recently, Foster et al. (2021) introduced the Decision-Estimation Coefficient, a measure of statistical complexity that lower bounds the optimal regret for interactive decision making, as well as a meta-algorithm, Estimation-to-Decisions, which achieves upper bounds in terms of the same quantity. Estimation-to-Decisions is a reduction, which lifts algorithms for (supervised) online estimation into algorithms for decision making. In this paper, we show that by combining Estimation-to-Decisions with a specialized form of "optimistic" estimation introduced by Zhang (2022), it is possible to obtain guarantees that improve upon those of Foster et al. (2021) by accommodating more lenient notions of estimation error. We use this approach to derive regret bounds for model-free reinforcement learning with value function approximation, and give structural results showing when it can and cannot help more generally.
Dylan J. Foster, Noah Golowich, Jian Qian, Alexander Rakhlin, Ayush Sekhari
NeurIPS3
2023 Convex and Non-convex Optimization Under Generalized Smoothness
abstract
Classical analysis of convex and non-convex optimization methods often requires the Lipschitz continuity of the gradient, which limits the analysis to functions bounded by quadratics. Recent work relaxed this requirement to a non-uniform smoothness condition with the Hessian norm bounded by an affine function of the gradient norm, and proved convergence in the non-convex setting via gradient clipping, assuming bounded noise. In this paper, we further generalize this non-uniform smoothness condition and develop a simple, yet powerful analysis technique that bounds the gradients along the trajectory, thereby leading to stronger results for both convex and non-convex optimization problems. In particular, we obtain the classical convergence rates for (stochastic) gradient descent and Nesterov's accelerated gradient method in the convex and/or non-convex setting under this general smoothness condition. The new analysis approach does not require gradient clipping and allows heavy-tailed noise with bounded variance in the stochastic setting.
Haochuan Li, Jian Qian, Alexander Rakhlin, Ali Jadbabaie
NeurIPS2
2022 Learning-Based Fast Depth Inter Coding for 3D-HEVC via XGBoost
abstract
The 3D extension of High Efficiency Video Coding (3D-HEVC) achieves excellent per-formance for 3D video coding while possessing significant computational complexity. To accelerate the time-consuming coding process of the depth map, a fast algorithm via XG-Boost is proposed in this paper. Specifically, a total of 14 specialized XGBoost models are used for different block sizes and viewpoint types to achieve early coding unit partition de-termination (ECP) and early prediction unit mode selection (EPM) to avoid executing the exhaustive traversal coding process. To promote the prediction accuracy of XGBoost mod-els, multi-domain correlations, including spatiotemporal, inter-view, and inter-component correlations are utilized and plenty of features are selected for model training. Evaluated on HTM-16.0 under random access configuration, the proposed ECP strategy can obtain 51.2% total encoding time saving with a 0.18% BDBR increase and the ECP+EPM can overall achieve 60.8% total encoding time saving with a 0.59% BDBR increase. The source code of our method is available at https://github.com/Joeyrr/_EPM.git.
Zixiang Zhang, Li Yu 0003, Jian Qian, Hongkui Wang
DCC3
2021 Robust learning under clean-label attack
abstract
We study the problem of robust learning under clean-label data-poisoning attacks, where the attacker injects (an arbitrary set of) \emph{correctly-labeled} examples to the training set to fool the algorithm into making mistakes on \emph{specific} test instances at test time. The learning goal is to minimize the attackable rate (the probability mass of attackable test instances), which is more difficult than optimal PAC learning. As we show, any robust algorithm with diminishing attackable rate can achieve the optimal dependence on $\epsilon$ in its PAC sample complexity, i.e., $O(1/\epsilon)$. On the other hand, the attackable rate might be large even for some optimal PAC learners, e.g., SVM for linear classifiers. Furthermore, we show that the class of linear hypotheses is not robustly learnable when the data distribution has zero margin and is robustly learnable in the case of positive margin but requires sample complexity exponential in the dimension. For a general hypothesis class with bounded VC dimension, if the attacker is limited to add at most $t=O(1/\epsilon)$ poison examples, the optimal robust learning sample complexity grows linearly with $t$.
Avrim Blum, Steve Hanneke, Jian Qian, Han Shao 0001
COLT3
2021 Bi-Prediction Enhancement with Deep Frame Prediction Network for Versatile Video Coding
abstract
Bi-prediction is a fundamental module of inter prediction in the blocked-based hybrid video coding framework. Block-based motion estimation(ME) and motion compensation(MC) with simple models are adopted in bi-prediction process. Unfortunately, this MEMC-based scheme can't guarantee the prediction performance when it comes to video with complicated motions. In this paper, a novel inter prediction scheme based on deep frame prediction network (DFP-net) is proposed to enhance bi-prediction accuracy especially in complicate scenes. Specifically, the proposed DFP-net is composed of multi-scale motion alignment, fusion of temporal and spatial correlation and frame synthesis module. The DFP-net can precisely extract and fuse motion features in various scales and completely exploit temporal and spatial correlation to generate the prediction frame in a data-driven manner. Moreover, the DFP-net is integrated into VTM-6.2 to provide an additional prediction frame for biprediction. Since the prediction generated by DFP-net is more similar with to-be-coded frame in the sense of temporal distance and texture, it can be added to reference list to improve the diversity of references. In this manner, the proposed bi-prediction scheme has surpassed VTM-6.2 on average 1.8% BD-rate saving.
Hao Tao, Jian Qian, Li Yu 0003, Hongkui Wang
DCC2
2021 Distortion-based Neural Network for Compression Artifacts Reduction in VVC
abstract
To reduce compression artifacts in video coding, prior knowledge from the codec is often utilized by deep learning based enhancement methods. However, most existing algorithms only take limited kinds of prior knowledge into consideration and directly feed it into neural networks, resulting in finite information obtained by the enhancement module. In this paper, a distortion-based neural network is proposed to better take advantage of features from the codec and further facilitate the quality of decoded videos. Firstly, a variety of prior knowledge is unitedly exploited to estimate the compression distortion. Secondly, an estimation module is also designed for the information obtained from the codec to improve the precision of estimation. With the accurately estimated distortion level, the enhancement module in the proposed method could reduce the corresponding artifacts by flexibly controlling the filtering strength. Moreover, the proposed method is integrated into VTM-7.1, achieving on average 5.92% BD-rate saving for Y component under All Intra configuration.
Jian Qian, Hongkui Wang, Li Yu 0003
VCIP1
2020 Spatial-Temporal Fusion Convolutional Neural Network for Compressed Video Enhancement in HEVC
abstract
Convolutional neural network has witnessed remarkable progress in compressed video quality enhancement in high efficiency video coding (HEVC) standard. But most existing methods focus on single frame quality enhancement where copious temporal and spatial information is neglected. In this paper, we propose a spatial-temporal fusion convolutional neural network (STEF-CNN) to employ spatial and temporal information to improve the performance of in-loop filter in HEVC. Specifically, the STEF-CNN adopts a pre-denoising network which in advance processes the compressed videos frame by frame. The pre-denoising operation alleviates the impact of noise and blocking artifacts. Then the denoised frames are sent to temporal-spatial fusion module which picks out valuable temporal and spatial information. The fused frames are eventually fed to quality enhancement network which is based on residual learning and dense network. The STEF-CNN is capable of capturing abundant information from consecutive neighboring frames. Extensive experimental results demonstrate the effectiveness of the proposed method. The STEF-CNN achieves 11.53% BD-BR reduction in all-intra (AI) configuration and 10.20% BD-BR reduction in random-access (RA) configuration.
Jian Qian, Li Yu 0003, Hongkui Wang, Hao Tao, Shengju Yu
DCC2
2020 A Compact Deep Neural Network for Single Image Super-Resolution
Jian Qian, Li Yu 0003, Shengju Yu, Hao Tao
MMM (2)2
2020 Towards Minimax Optimal Reinforcement Learning in Factored Markov Decision Processes
abstract
We study minimax optimal reinforcement learning in episodic factored Markov decision processes (FMDPs), which are MDPs with conditionally independent transition components. Assuming the factorization is known, we propose two model-based algorithms. The first one achieves minimax optimal regret guarantees for a rich class of factored structures, while the second one enjoys better computational complexity with a slightly worse regret. A key new ingredient of our algorithms is the design of a bonus term to guide exploration. We complement our algorithms by presenting several structure dependent lower bounds on regret for FMDPs that reveal the difficulty hiding in the intricacy of the structures.
Jian Qian, Suvrit Sra
NeurIPS2
2019 Exploration Bonus for Regret Minimization in Discrete and Continuous Average Reward MDPs
abstract
The exploration bonus is an effective approach to manage the exploration-exploitation trade-off in Markov Decision Processes (MDPs). While it has been analyzed in infinite-horizon discounted and finite-horizon problems, we focus on designing and analysing the exploration bonus in the more challenging infinite-horizon undiscounted setting. We first introduce SCAL+, a variant of SCAL (Fruit et al. 2018), that uses a suitable exploration bonus to solve any discrete unknown weakly-communicating MDP for which an upper bound $c$ on the span of the optimal bias function is known. We prove that SCAL+ enjoys the same regret guarantees as SCAL, which relies on the less efficient extended value iteration approach. Furthermore, we leverage the flexibility provided by the exploration bonus scheme to generalize SCAL+ to smooth MDPs with continuous state space and discrete actions. We show that the resulting algorithm (SCCAL+) achieves the same regret bound as UCCRL (Ortner and Ryabko, 2012) while being the first implementable algorithm for this setting.
Jian Qian, Ronan Fruit, Matteo Pirotta, Alessandro Lazaric
NeurIPS1
2019 Importance Resampling for Off-policy Prediction
abstract
Importance sampling (IS) is a common reweighting strategy for off-policy prediction in reinforcement learning. While it is consistent and unbiased, it can result in high variance updates to the weights for the value function. In this work, we explore a resampling strategy as an alternative to reweighting. We propose Importance Resampling (IR) for off-policy prediction, which resamples experience from a replay buffer and applies standard on-policy updates. The approach avoids using importance sampling ratios in the update, instead correcting the distribution before the update. We characterize the bias and consistency of IR, particularly compared to Weighted IS (WIS). We demonstrate in several microworlds that IR has improved sample efficiency and lower variance updates, as compared to IS and several variance-reduced IS strategies, including variants of WIS and V-trace which clips IS ratios. We also provide a demonstration showing IR improves over IS for learning a value function from images in a racing car simulator.
Matthew Schlegel, Wesley Chung, Daniel Graves, Jian Qian, Martha White
NeurIPS4
2019 Dense Inception Attention Neural Network for In-Loop Filter
abstract
Recently, deep learning technology has made significant progresses in high efficiency video coding (HEVC), especially in in-loop filter. In this paper, we propose a dense inception attention network (DIA_Net) to delve into image information and model capacity. The DIA_Net contains multiple inception blocks which have different size kernels so as to dig out various scales information. Meanwhile, attention mechanism including spatial attention and channel attention is utilized to fully exploit feature information. Further we adopt a dense residual structure to deepen the network. We attach DIA_Net to the end of in-loop filter part in HEVC as a post-processor and apply it to luma components. The experimental results demonstrate the proposed DIA_Net has remarkable improvement over the standard HEVC. With all-intra(AI) and random access(RA) configurations, It achieves 8.2% bd-rate reduction in AI configuration and 5.6% bd-rate reduction in RA configuration.
Jian Qian, Li Yu 0003, Hongkui Wang, Xing Zeng, Ning Wang 0109
PCS2