Chenlong He

dblp:190/2766 · DBLP profile ↗
← Back
10ranked-venue papers
3as first author
8since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Systems, architecture and hardware · 5 · 2 first-author · 5 since 2021Artificial intelligence and machine learning · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 HLC: A High-Quality Lightweight Mezzanine Codec Featuring High-Throughput Palette
Chenlong He, Leilei Huang, Wei Li 0257, Hanyang Cui, Zhijian Hao, Xiaoyang Zeng, Yibo Fan
ISCAS1
2026 SENTRY-VQA: A Region-Aware Framework for Surveillance Video Quality Assessment
Chenlong He, Leilei Huang, Minge Jing, Yibo Fan
ISCAS3
2025 A High-Precision and Low-Cost Approximate Transform Accelerator for Video Coding
abstract
The introduction of multiple transform types in the Versatile Video Coding (VVC) standard has yielded notable encoding gains but also imposed considerable computational burdens. Existing transform circuits of different types are typically implemented separately due to their independence, leading to substantial hardware overhead. To address this, we explore the relationship between Discrete Cosine Transform Type-2 (DCT2) and Discrete Sine Transform Type-7 (DST7) matrices and reveal a prominent diagonal aggregation phenomenon in their transfer matrix. Based on this insight, the least-squares method is applied to optimize the transfer matrix sparsity, achieving a high-precision, low-cost approximate conversion from DCT2 to DST7. Furthermore, we optimize DCT2 computation by proposing an elaborate matrix decomposition approach that allows a lightweight shift-adder unit to efficiently generate all required product terms across varying sizes. Leveraging these algorithmic optimizations, we implement a highly reusable and area-efficient approximate transform accelerator that supports sizes from 4 to 32 points and accommodates three types in VVC. Experimental results demonstrate that the proposed accelerator achieves over 44% reduction in circuit resource consumption with negligible BD-BR performance loss of just $\mathbf{0. 5 3 \%}$, maintaining processing capabilities up to $8 K \text{@} 57 \mathrm{fps}$.
Zhijian Hao, Chenlong He, Qi Zheng 0004, Shushi Chen, Jinchang Xu, Yue Hao 0001, Xiaohua Ma 0001
DAC3
2025 A Novel Transform Accelerator With Fast Kernel Selection and Efficient Transform Circuit
abstract
The introduction of multiple transform types into the Versatile Video Coding (VVC) standard has yielded notable encoding gains but also resulted in substantial computational burdens, posing two critical challenges for hardware implementation: fast kernel selection and efficient transform computation design. Existing studies typically address these challenges in isolation, lacking a holistic solution for VVC transform coding. In this paper, we presents a groundbreaking transform accelerator that unifies transform kernel selection and multiple transform circuit within a single framework. In terms of algorithms, driven by mechanistic analysis, we propose a decision tree-based kernel selection algorithm that ensures both high decision accuracy and computational efficiency. Additionally, we design a transfer matrix-based approximation algorithm for Discrete Sine Transform Type-7 and a matrix decomposition-based improved computation for Discrete Cosine Transform Type-2, significantly reducing the computational complexity. On the hardware front, we implement a high-precision and area-efficient transform accelerator, which integrates highly pipelined kernel selection and transform computation architectures. With multiple reuse and parallelism strategies, the accelerator demonstrates substantial resource efficiency advantages. Experimental results reveal that the proposed accelerator achieves a circuit resource reduction of over 44% with a slight performance degradation, while maintaining processing capabilities up to 8K@57 fps. To the best of our knowledge, this is the first comprehensive hardware solution for VVC transform coding that jointly addresses the challenges of kernel selection and transform circuit design.
Zhijian Hao, Chenlong He, Qi Zheng 0004, Jinchang Xu, Peijun Ma, Xiaohua Ma 0001, Yue Hao 0001
IEEE Trans. Circuits Syst. I Regul. Pap.2
2025 An 8K@120fps Advanced Entropy Coding Hardware Design for AVS3
abstract
The third generation audio video coding standard (AVS3) is the latest video coding standard developed by the China AVS working group. The advanced entropy coding (AEC) tool in AVS3 has critical bin-to-bin data dependencies leading to difficulties in parallelization. The use of a 16384-entries lookup table (LUT) in the AEC context update algorithm poses challenges in balancing area and performance. To address these issues, we propose a high-performance, area-efficient hardware design. Firstly, we introduce a novel multicycle-path parallel architecture and optimize area cost through hardware reuse. Next, we construct a context modeling processing unit to replace the large LUT, significantly reducing area. Finally, we propose a new LUT-free dual-context modeling processing unit, effectively resolving critical paths introduced by parallel context conflicts. As a result, our design processes 2.63 bins per cycle. The synthesis results based on the GlobalFoundries’ 28nm process indicate that its maximum frequency is 990MHz, with an overall throughput of 2604 Mbin/s. Compared to state-of-the-art designs, our design leads in performance by 24.7% while reducing area by 28%.
Wei Li 0257, Leilei Huang, Chenlong He, Minge Jing, Yibo Fan
IEEE Trans. Circuits Syst. Video Technol.3
2025 Subjective and Objective Quality Assessment of Banding Artifacts on Compressed Videos
abstract
Although there have been notable advancements in video compression technologies in recent years, banding artifacts remain a serious issue affecting the quality of compressed videos, particularly on smooth regions of high-definition videos. Noticeable banding artifacts can severely impact the perceptual quality of videos viewed on a high-end HDTV or high-resolution screen. Hence, there is a pressing need for a systematic investigation of the banding video quality assessment problem for advanced video codecs. Given that the existing publicly available datasets for studying banding artifacts are limited to still picture data only, which cannot account for temporal banding dynamics, we have created a first-of-a-kind open video dataset, dubbed LIVE-YT-Banding, which consists of 160 videos generated by four different compression parameters using the AV1 video codec. A total of 7,200 subjective opinions are collected from a cohort of 45 human subjects. To demonstrate the value of this new resources, we tested and compared a variety of models that detect banding occurrences, and measure their impact on perceived quality. Among these, we introduce an effective and efficient new no-reference (NR) video quality evaluator which we call CBAND. CBAND leverages the properties of the learned statistics of natural images expressed in the embeddings of deep neural networks. Our experimental results show that the perceptual banding prediction performance of CBAND significantly exceeds that of previous state-of-the-art models, and is also orders of magnitude faster. Moreover, CBAND can be employed as a differentiable loss function to optimize video debanding models. The LIVE-YT-Banding database, code, and pre-trained model are all publically available at https://github.com/uniqzheng/CBAND.
Qi Zheng 0004, Li-Heng Chen, Chenlong He, Neil Birkbeck, Yilin Wang 0001, Balu Adsumilli, Alan C. Bovik, Yibo Fan, Zhengzhong Tu
IEEE Trans. Image Process.3
2024 CTU-Level Adaptive Quantization Method Joint with GOP based Temporal Filter for Video Coding
abstract
Both Versatile Video Coding (VVC) and High Efficiency Video Coding (HEVC) introduce Group of Pictures (GOP) based temporal filter (GBTF) as a pre-filter to improve compression performance. While numerous efforts have been made to optimize GBTF, there is a limited amount of research that explicitly addresses why GBTF could improve compression performance. Additionally, most optimizations have focused on the design of the filter itself, rather than on how to better integrate it with other encoding tools. In this paper, we analyze the reasons behind the superior compression performance of GBTF. Subsequently, we introduce a Coding Tree Unit (CTU)-level adaptive quantization parameter allocation method joint with GBTF to further enhance compression performance for video coding. The experimental results demonstrate that, for VVC, our method provides Bjontegaard delta bit rate (BD-BR) savings of 2.0% for Peak Signal-to-Noise Ratio (PSNR) and 4.0% for Structural Similarity index (SSIM). Furthermore, for HEVC, our method provides BD-BR savings of 3.5% for PSNR and 7.6% for SSIM.
Chenlong He, Xiaoxiang Chen, Zhijian Hao, Chao Liu 0027, Xiaoyang Zeng, Yibo Fan
ISCAS1
2023 Probabilistic model with evolutionary optimization for cognitive diagnosis
abstract
Cognitive Diagnostic Models (CDMs) aim to analyze students' cognitive levels of each knowledge component (KC) by mining educational data. Existing CDMs can be mainly divided into two categories, i.e., traditional probability-based and neural-network-based. Most probabilistic models have the advantages of simplicity and good interpretability, but suffer from slow training time in the case of a large number of KCs. Neural-network-based methods are widely considered to be superior to probabilistic models due to their good performance. However, neural network methods are less interpretable than probabilistic models, thus limiting their usefulness in practice. Because most existing probabilistic models are optimized iteratively based on single-point-based search methods, they may be easily trapped in local optimum due to the influence of the initial points. And evolutionary algorithms (EAs) have good global search ability. Therefore, an interesting question is whether a simple probabilistic model based on evolutionary optimization can rival neural-network in limited optimization time. Thus, a hybrid EA with a customized local search is proposed. Experimental results on three real-world datasets show that our method outperforms the compared 7 models (including 2 state-of-the-art neural-network-based models); and the running time of our method is significantly less than the compared probabilistic models.
Chenyang Bu, Zhiyong Cao, Chenlong He, Yuhong Zhang 0002
GECCO3
2016 Enhance continuous estimation of distribution algorithm by variance enlargement and reflecting sampling
abstract
Estimation of distribution algorithm (EDA) is a kind of typical model-based evolutionary algorithm (EA). Although possessing competitive advantages in theoretical analysis, current EDAs may encounter premature convergence due to the rapid shrinkage of the search range and the relatively low sampling efficiency. Focusing on continuous EDAs with Gaussian models, this paper proposes a novel probability density estimator which can adaptively enlarge the variances and thus endow EDA with flexible search behavior. For the estimated probability density, a reflecting sampling strategy which can further improve the search efficiency is put forward. With these two algorithmic strategies, a new EDA variant named EDAver is developed. Experimental results on a set of benchmark problems demonstrate that EDAver outperforms conventional EDAs and can produce superior solutions in comparison with some state-of-the-art EAs.
Chenlong He, Dexing Zhong, Yongsheng Liang 0002
CEC2
2016 Collective motion of self-propelled particles without collision and fragmentation
abstract
In this paper, we propose a novel model to generate collective motion of self-propelled particles without collision and fragmentation by means of adjusting the absolute velocity instead of the constant speed used in the original Vicsek model. Interactions among particles are represented by the r-limited Delaunay graph to guarantee the locality of the model. The centroid of the Voronoi cell is set as a destination to scatter particles. The formation of the group is controlled by the surface tension generated by particles on the boundary. Without noise and periodic boundary, given a collision-free and cohesive configuration initially, the abundant types of collective motion will emerge from the local behaviors. Numerical simulations demonstrate that our model can produce rich behaviors such as crystallization, rotation, flocking and cluster with different combinations of two coefficients adjusting the amplitude of centering force and surface tension, respectively. Collisions among particles and fragmentations of the group do not appear.
Chenlong He
SMC1