Li-Heng Chen

dblp:168/3493 · DBLP profile ↗
← Back
14ranked-venue papers
9as first author
9since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 11 · 7 first-author · 8 since 2021Artificial intelligence and machine learning · 4 · 3 first-author · 2 since 2021
YearPublicationVenuePosition
2025 GCRayDiffusion: Pose-Free Surface Reconstruction via Geometric Consistent Ray Diffusion
abstract
Accurate surface reconstruction from unposed images is crucial for efficient 3D object or scene creation. However, it remains challenging, particularly for the joint camera pose estimation. Previous approaches have achieved impressive pose-free surface reconstruction results in dense-view settings, but could easily fail for sparse-view scenarios without sufficient visual overlap. In this paper, we propose a new technique for pose-free surface reconstruction, which follows triplane-based signed distance field (SDF) learning but regularizes the learning by explicit points sampled from ray-based diffusion of camera pose estimation. Our key contribution is a novel Geometric Consistent Ray Diffusion model (GCRayDiffusion), where we represent camera poses as neural bundle rays and regress the distribution of noisy rays via a diffusion model. More importantly, we further condition the denoising process of RGRayDiffusion using the triplane-based SDF of the entire scene, which provides effective 3D consistent regularization to achieve multi-view consistent camera pose estimation. Finally, we incorporate RGRayDiffusion into the triplane-based SDF learning by introducing on-surface geometric regularization from the sampling points of the neural bundle rays, which leads to highly accurate pose-free surface reconstruction results even for sparse-view inputs. Extensive evaluations on public datasets show that our GCRayDiffusion achieves more accurate camera pose estimation than previous approaches, with geometrically more consistent surface reconstruction results, especially given sparse-view inputs.
Li-Heng Chen, Zixin Zou, Tianjiao Jing, Yan-Pei Cao 0001, Shi-Sheng Huang, Hongbo Fu 0001, Hua Huang 0001
ICCV1
2025 Estimating the resize parameter in end-to-end learned image compression
Li-Heng Chen, Christos G. Bampis, Zhi Li 0001, Lukas Krasula, Alan C. Bovik
Signal Process. Image Commun.1
2025 Subjective and Objective Quality Assessment of Banding Artifacts on Compressed Videos
abstract
Although there have been notable advancements in video compression technologies in recent years, banding artifacts remain a serious issue affecting the quality of compressed videos, particularly on smooth regions of high-definition videos. Noticeable banding artifacts can severely impact the perceptual quality of videos viewed on a high-end HDTV or high-resolution screen. Hence, there is a pressing need for a systematic investigation of the banding video quality assessment problem for advanced video codecs. Given that the existing publicly available datasets for studying banding artifacts are limited to still picture data only, which cannot account for temporal banding dynamics, we have created a first-of-a-kind open video dataset, dubbed LIVE-YT-Banding, which consists of 160 videos generated by four different compression parameters using the AV1 video codec. A total of 7,200 subjective opinions are collected from a cohort of 45 human subjects. To demonstrate the value of this new resources, we tested and compared a variety of models that detect banding occurrences, and measure their impact on perceived quality. Among these, we introduce an effective and efficient new no-reference (NR) video quality evaluator which we call CBAND. CBAND leverages the properties of the learned statistics of natural images expressed in the embeddings of deep neural networks. Our experimental results show that the perceptual banding prediction performance of CBAND significantly exceeds that of previous state-of-the-art models, and is also orders of magnitude faster. Moreover, CBAND can be employed as a differentiable loss function to optimize video debanding models. The LIVE-YT-Banding database, code, and pre-trained model are all publically available at https://github.com/uniqzheng/CBAND.
Qi Zheng 0004, Li-Heng Chen, Chenlong He, Neil Birkbeck, Yilin Wang 0001, Balu Adsumilli, Alan C. Bovik, Yibo Fan, Zhengzhong Tu
IEEE Trans. Image Process.2
2024 Learned fractional downsampling network for adaptive video streaming
Li-Heng Chen, Christos G. Bampis, Zhi Li 0001, Joel Sole, Chao Chen 0006, Alan C. Bovik
Signal Process. Image Commun.1
2024 Fuzzy-Based Adaptive Reliable Motion Control of a Piezoelectric Nanopositioning System
abstract
The nonlinearity and cross-axis coupling of piezo-driven multiple-degree-of-freedom (multi-DOF) nanopositioning systems impose challenges to achieving precise and reliable motion control. This paper develops a new adaptive reliable control approach for a 2-DOF piezoelectric nanopositioning system utilizing a fuzzy back-stepping strategy. First, a virtual tracking model is constructed to address the stabilization problem via a system transformation method. Then, a fuzzy logical system (FLS) model is introduced to mitigate the effects of hysteresis and unmodeled high-order nonlinearity. To obtain high-precision motion tracking with high reliability, the approximation of the unknown piezoelectric actuator's efficiency factor is injected into the reliable controller of nanopositioning systems. Furthermore, an adaptive mechanism based on tracking errors is designed to adjust control parameters automatically to improve the robustness to unknown perturbations. Simulation and practical experiment examples are presented to show the effectiveness and potential of the developed fuzzy reliable nanopositioning control method over existing control approaches.
Li-Heng Chen, Qingsong Xu 0002
IEEE Trans. Fuzzy Syst.1
2021 Regression or classification? New methods to evaluate no-reference picture and video quality models
abstract
Video and image quality assessment has long been projected as a regression problem, which requires predicting a continuous quality score given an input stimulus. However, recent efforts have shown that accurate quality score regression on real-world user-generated content (UGC) is a very challenging task. To make the problem more tractable, we propose two new methods - binary, and ordinal classification - as alternatives to evaluate and compare no-reference quality models at coarser levels. Moreover, the proposed new tasks convey more practical meaning on perceptually optimized UGC transcoding, or for preprocessing on media processing platforms. We conduct a comprehensive benchmark experiment of popular no-reference quality models on recent in-the-wild picture and video quality datasets, providing reliable baselines for both evaluation methods to support further studies. We hope this work promotes coarse-grained perceptual modeling and its applications to efficient UGC processing.
Zhengzhong Tu, Chia-Ju Chen, Li-Heng Chen, Yilin Wang 0001, Neil Birkbeck, Balu Adsumilli, Alan C. Bovik
ICASSP3
2021 A Progressive Architecture for Learned Fractional Downsampling
abstract
In many image and video processing applications, the ability to resize by a fractional factor, such as from 1080p to 720p, is essential. However, conventional CNN layers can only be used to alter the resolution of their inputs with integer scale factors. In this paper, we propose a downsampling network architecture that progressively reconstructs residuals at different scales. In particular, the aforementioned problem is solved by combining an upsampling sub-network and a downsampling subnetwork, both with integer scale factor. As an application, we apply the proposed downsampling network to an adaptive bitrate video streaming scenario. We extensively evaluate with different video codecs and upsampling algorithms to show the generality of our model. Our experimental results show that improvements in coding efficiency over the conventional Lanczos downsampling and state-of-the-art methods are attained, measured in different perceptual video quality models on large-resolution test videos.
Li-Heng Chen, Christos G. Bampis, Zhi Li 0001, Joel Sole, Alan C. Bovik
PCS1
2021 ProxIQA: A Proxy Approach to Perceptual Optimization of Learned Image Compression
abstract
(p = 1,2) norms has largely dominated the measurement of loss in neural networks due to their simplicity and analytical properties. However, when used to assess the loss of visual information, these simple norms are not very consistent with human perception. Here, we describe a different "proximal" approach to optimize image analysis networks against quantitative perceptual models. Specifically, we construct a proxy network, broadly termed ProxIQA, which mimics the perceptual model while serving as a loss layer of the network. We experimentally demonstrate how this optimization framework can be applied to train an end-to-end optimized image compression network. By building on top of an existing deep image compression model, we are able to demonstrate a bitrate reduction of as much as 31% over MSE optimization, given a specified perceptual quality (VMAF) level.
Li-Heng Chen, Christos G. Bampis, Zhi Li 0001, Andrey Norkin, Alan C. Bovik
IEEE Trans. Image Process.1
2021 Perceptual Video Quality Prediction Emphasizing Chroma Distortions
abstract
Measuring the quality of digital videos viewed by human observers has become a common practice in numerous multimedia applications, such as adaptive video streaming, quality monitoring, and other digital TV applications. Here we explore a significant, yet relatively unexplored problem: measuring perceptual quality on videos arising from both luma and chroma distortions from compression. Toward investigating this problem, it is important to understand the kinds of chroma distortions that arise, how they relate to luma compression distortions, and how they can affect perceived quality. We designed and carried out a subjective experiment to measure subjective video quality on both luma and chroma distortions, introduced both in isolation as well as together. Specifically, the new subjective dataset comprises a total of 210 videos afflicted by distortions caused by varying levels of luma quantization commingled with different amounts of chroma quantization. The subjective scores were evaluated by 34 subjects in a controlled environmental setting. Using the newly collected subjective data, we were able to demonstrate important shortcomings of existing video quality models, especially in regards to chroma distortions. Further, we designed an objective video quality model which builds on existing video quality algorithms, by considering the fidelity of chroma channels in a principled way. We also found that this quality analysis implies that there is room for reducing bitrate consumption in modern video codecs by creatively increasing the compression factor on chroma channels. We believe that this work will both encourage further research in this direction, as well as advance progress on the ultimate goal of jointly optimizing luma and chroma compression in modern video encoders.
Li-Heng Chen, Christos G. Bampis, Zhi Li 0001, Joel Sole, Alan C. Bovik
IEEE Trans. Image Process.1
2020 A Comparative Evaluation Of Temporal Pooling Methods For Blind Video Quality Assessment
abstract
Many objective video quality assessment (VQA) algorithms include a key step of temporal pooling of frame-level quality scores. However, less attention has been paid to studying the relative efficiencies of different pooling methods on noreference (blind) VQA. Here we conduct a large-scale comparative evaluation to assess the capabilities and limitations of multiple temporal pooling strategies on blind VQA of usergenerated videos. The study yields insights and general guidance regarding the application and selection of temporal pooling models. In addition, we also propose an ensemble pooling model built on top of high-performing temporal pooling models. Our experimental results demonstrate the relative efficacies of the evaluated temporal pooling models, using several popular VQA algorithms evaluated on two recent largescale natural video quality databases. Conclusively, we also provide an empirical recipe for applying temporal pooling of frame-based quality predictions.
Zhengzhong Tu, Chia-Ju Chen, Li-Heng Chen, Neil Birkbeck, Balu Adsumilli, Alan C. Bovik
ICIP3
2020 Learning to Distort Images Using Generative Adversarial Networks
abstract
Modeling image and video distortions is an important, but difficult problem of great consequence to numerous and diverse image processing and computer vision applications. While many statistical models have been proposed to synthesize different types of image noise, real-world distortions are far more difficult to emulate. Toward advancing progress on this interesting problem, we consider distortion generation as an image-to-image transformation problem, and solve it via a data-driven approach. Specifically, we use a conditional generative adversarial network (cGAN) which we train to learn four kinds of realistic distortions. We experimentally demonstrate that the learned model can produce the perceptual characteristics of several types of distortion.
Li-Heng Chen, Christos G. Bampis, Zhi Li 0001, Alan C. Bovik
IEEE Signal Process. Lett.1
2018 Adaptive Fuzzy Sliding Mode Control for Network-Based Nonlinear Systems With Actuator Failures
abstract
This paper investigates the robust control problem of nonlinear systems with unknown time-varying actuator faults over digital communication networks. In the study, a new adaptive sliding mode control (SMC) scheme is developed for the investigated nonlinear systems, where the unknown nonlinearity is approximated via the adaptive fuzzy mechanism in the presence of signal quantization. The proposed control law can compensate the time-varying faults and the quantization errors completely by injecting the quantizer parameter into the controller gain. Moreover, in this design, the adaptive fuzzy law is constructed via the quantized state information instead of using the exact original information. Under the proposed quantized SMC scheme, the closed-loop control systems are guaranteed to be asymptotically stable, and the reachability of the sliding surface can be ensured strictly. Finally, two examples are presented to show the effectiveness of the method.
Li-Heng Chen, Ming Liu 0014, Xianlin Huang, Shasha Fu, Jianbin Qiu
IEEE Trans. Fuzzy Syst.1
2018 Adaptive Fuzzy Observer Design for a Class of Switched Nonlinear Systems With Actuator and Sensor Faults
abstract
In this paper, an adaptive fault estimation approach is proposed for a class of switched nonlinear systems. The considered system is assumed to possess unknown nonlinearities, unmeasured states, and simultaneous sensor and actuator faults. Fuzzy logic systems are applied to approximate the unknown nonlinear terms. Two new adaptive fuzzy observers are designed where the sensor and actuator faults can be estimated separately. On the basis of average dwell time approach and Lyapunov stability theory, the resulting error system is proven to be bounded stable with designed parameters. Finally, a simulation example is presented to illustrate the effectiveness of the designed observers.
Shasha Fu, Jianbin Qiu, Li-Heng Chen, Shaoshuai Mou
IEEE Trans. Fuzzy Syst.3
2015 Image Quality Assessment Using Human Visual DOG Model Fused With Random Forest
abstract
Objective image quality assessment (IQA) plays an important role in the development of multimedia applications. Prediction of IQA metric should be consistent with human perception. The release of the newest IQA database (TID2013) challenges most of the widely used quality metrics (e.g., peak-to-noise-ratio and structure similarity index). We propose a new methodology to build the metric model using a regression approach. The new IQA score is set to be the nonlinear combination of features extracted from several difference of Gaussian (DOG) frequency bands, which mimics the human visual system (HVS). Experimental results show that the random forest regression model trained by the proposed DOG feature is highly correspondent to the HVS and is also robust when tested by available databases.
Soo-Chang Pei, Li-Heng Chen
IEEE Trans. Image Process.2