Chuang Zhu

dblp:02/8503 · DBLP profile ↗
← Back
41ranked-venue papers
7as first author
22since 2021 · last 2026
0000-0001-5155-7069ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 27 · 2 first-author · 10 since 2021Artificial intelligence and machine learning · 18 · 3 first-author · 15 since 2021Systems, architecture and hardware · 2 · 1 first-authorDatabases, data management, data science and information retrieval · 2 · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 PEOD: A Pixel-Aligned Event-RGB Benchmark for Object Detection Under Challenging Conditions
abstract
Robust object detection for challenging scenarios increasingly relies on event cameras, yet existing Event-RGB datasets remain constrained by sparse coverage of extreme conditions and low spatial resolution (≤ 640 × 480), which prevents comprehensive evaluation of detectors under challenging scenarios. To address these limitations, we propose PEOD, the first large-scale, pixel-aligned and hign-resolution (1280 × 720) Event-RGB dataset for object detection under challenge conditions. PEOD contains 130+ spatiotemporal-aligned sequences and 340k manual bounding boxes, with 57% of data captured under low-light, overexposure, and high-speed motion. Furthermore, we benchmark 14 methods across three input configurations (Event-based, RGB-based, and Event-RGB fusion) on PEOD. On the full test set and normal subset, fusion-based models achieve the excellent performance. However, in illumination challenge subset, the top event-based model outperforms all fusion models, while fusion models still outperform their RGB-based counterparts, indicating limits of existing fusion methods when the frame modality is severely degraded. PEOD establishes a realistic, high-quality benchmark for multimodal perception and will be publicly released later to facilitate future research.
Luoping Cui, Endian Lin, Donghong Jiang, Chuang Zhu
AAAI7
2026 MoDe: Multi-modal discriminative priors for prompt tuning
Chuang Zhu, Maoyuan Shao, Zekuan Yu
Neurocomputing2
2025 Precise Diffusion Inversion: Towards Novel Samples and Few-Step Models
abstract
The diffusion inversion problem seeks to recover the latent generative trajectory of a diffusion model given a real image. Faithful inversion is critical for ensuring consistency in diffusion-based image editing. Prior works formulate this task as a fixed-point problem and solve it using numerical methods. However, achieving both accuracy and efficiency remains challenging, especially for few-step models and novel samples. In this paper, we propose ***PreciseInv***, a general-purpose test-time optimization framework that enables fast and faithful inversion in as few as two inference steps. Unlike root-finding methods, we reformulate inversion as a learning problem and introduce a dynamic programming-inspired strategy to recursively estimate a parameterized sequence of noise embeddings. This design leverages the smoothness of the diffusion latent space for accurate gradient-based optimization and ensures memory efficiency via recursive subproblem construction. We further provide a theoretical analysis of ***PreciseInv***'s convergence and derive a provable upper bound on its reconstruction error. Extensive experiments on COCO 2017, DarkFace, and a stylized cartoon dataset show that ***PreciseInv*** achieves state-of-the-art performance in both reconstruction quality and inference speed. Improvements are especially notable for few-step models and under distribution shifts. Moreover, precise inversion yields substantial gains in editing consistency for text-driven image manipulation tasks. Code is available at: https://github.com/panda7777777/PreciseInv
Jing Zuo, Luoping Cui, Chuang Zhu, Yonggang Qi
NeurIPS3
2025 Hard-Aware Instance Adaptive Self-Training for Unsupervised Cross-Domain Semantic Segmentation
abstract
The divergence between labeled training data and unlabeled testing data is a significant challenge for recent deep learning models. Unsupervised domain adaptation (UDA) attempts to solve such problem. Recent works show that self-training is a powerful approach to UDA. However, existing methods have difficulty in balancing the scalability and performance. In this paper, we propose a hard-aware instance adaptive self-training framework for UDA on the task of semantic segmentation. To effectively improve the quality and diversity of pseudo-labels, we develop a novel pseudo-label generation strategy with an instance adaptive selector. We further enrich the hard class pseudo-labels with inter-image information through a skillfully designed hard-aware pseudo-label augmentation. Besides, we propose the region-adaptive regularization to smooth the pseudo-label region and sharpen the non-pseudo-label region. For the non-pseudo-label region, consistency constraint is also constructed to introduce stronger supervision signals during model optimization. Our method is so concise and efficient that it is easy to be generalized to other UDA methods. Experiments on GTA5 $\rightarrow$→ Cityscapes, SYNTHIA $\rightarrow$→ Cityscapes, and Cityscapes $\rightarrow$→ Oxford RobotCar demonstrate the superior performance of our approach compared with the state-of-the-art methods.
Chuang Zhu, Kebin Liu 0002, Wenqi Tang, Ke Mei, Jiaqi Zou, Tiejun Huang 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2025 Source-Free Semantic Regularization Learning for Semi-Supervised Domain Adaptation
abstract
Semi-supervised domain adaptation (SSDA) has been extensively researched due to its ability to improve classification performance and generalization ability of models by using a small amount of labeled data on the target domain. However, existing methods cannot effectively adapt to the target domain due to difficulty in fully learning rich and complex target semantic information and relationships. In this paper, we propose a novel SSDA learning framework called semantic regularization learning (SERL), which captures the target semantic information from multiple perspectives of regularization learning to achieve adaptive fine-tuning of the source pre-trained model on the target domain. SERL includes three robust semantic regularization techniques. Firstly, semantic probability contrastive regularization (SPCR) helps the model learn more discriminative feature representations from a probabilistic perspective, using semantic information on the target domain to understand the similarities and differences between samples. Additionally, adaptive weights in SPCR can help the model learn the semantic distribution correctly through the probabilities of different samples. To further comprehensively understand the target semantic distribution, we introduce hard-sample mixup regularization (HMR), which uses easy samples as guidance to mine the latent target knowledge contained in hard samples, thereby learning more complete and complex target semantic knowledge. Finally, target prediction regularization (TPR) regularizes the target predictions of the model by maximizing the correlation between the current prediction and the past learned objective, thereby mitigating the misleading of semantic information caused by erroneous pseudo-labels. Extensive experiments on three benchmark datasets demonstrate that our SERL method achieves state-of-the-art performance.
Chuang Zhu, Ruiying Ren, Tiejun Huang 0001
IEEE Trans. Multim.2
2024 Unsupervised Domain Adaptive Semantic Segmentation Based on Clip-Guided Prototypical Contrastive Learning
abstract
Domain adaptive semantic segmentation aims to improve the model performance by bridging the gap existing between source and target domains. Recent works show that prototypical contrastive learning is a powerful approach. However, the prototypes can become unstable when there are significant variations in visual characteristics (e.g., color, scale and shape) among different images. Additionally, the prototypes generated from the source domain are highly correlated with domain information, which limits further gains in domain alignment. To address these issues, we propose a new method based on CLIP-guided Prototypical Contrastive Learning (CLIP-ProCL). Our approach simultaneously combines the rich text knowledge and image knowledge of CLIP to perform domain alignment. Towards the former, we obtain robust and domain-agnostic prototypes through the utilization of text prompts. Towards the latter, we leverage the image priors of CLIP to further guide the features learned by the segmentation network closer to the CLIP space. Experiments on the benchmark tasks GTA5 $\rightarrow$ Cityscapes and SYNTHIA $\rightarrow$ Cityscapes demonstrate that our approach outperforms the state-of-the-art methods. Our code is available at https://github.com/bupt-ai-cz/CLIP-ProCL.
Kebin Liu 0002, Chuang Zhu
ICIP2
2024 Learning from Noisy Labels for Long-Tailed Data via Optimal Transport
Chuang Zhu
ICONIP (7)2
2024 Learning robust correlation with foundation model for weakly-supervised few-shot segmentation
Chuang Zhu, Kebin Liu 0002, Ruiying Ren
Knowl. Based Syst.2
2023 TorchAudio 2.1: Advancing Speech Recognition, Self-Supervised Learning, and Audio Processing Components for Pytorch
abstract
TorchAudio is an open-source audio and speech processing library built for PyTorch. It aims to accelerate the research and development of audio and speech technologies by providing well-designed, easy-to-use, and performant PyTorch components. Its contributors routinely engage with users to understand their needs and fulfill them by developing impactful features. Here, we survey TorchAudio’s development principles and contents and highlight key features we include in its latest version (2.1): self-supervised learning pre-trained pipelines and training recipes, high-performance CTC decoders, speech recognition models and training recipes, advanced media I/O capabilities, and tools for performing forced alignment, multi-channel speech enhancement, and reference-less speech assessment. For a selection of these features, through empirical studies, we demonstrate their efficacy and show that they achieve competitive or state-of-the-art performance.
Jeff Hwang, Moto Hira, Caroline Chen, Xiaohui Zhang 0007, Zhaoheng Ni, Guangzhi Sun, Pingchuan Ma 0001, Ruizhe Huang, Vineel Pratap, Yuekai Zhang, Anurag Kumar 0003, Chin-Yun Yu, Chuang Zhu, Chunxi Liu, Jacob Kahn, Mirco Ravanelli, Shinji Watanabe 0001, Yangyang Shi, Yumeng Tao
ASRU13
2023 RestNet: Boosting Cross-Domain Few-Shot Segmentation with Residual Transformation Network
Chuang Zhu
BMVC2
2023 Highly Efficient SNNs for High-speed Object Detection
Nemin Qiu, Yuan Li 0014, Chuang Zhu
BMVC4
2023 WUDA: Unsupervised Domain Adaptation Based on Weak Source Domain Labels
abstract
Unsupervised domain adaptation (UDA) for semantic segmentation addresses the cross-domain problem with fine source domain labels. However, the acquisition of semantic labels is often time-consuming, many scenarios only have weak labels (e.g. bounding boxes). When weak supervision and cross-domain problems coexist, this paper defines a new task: unsupervised domain adaptation based on weak source domain labels (WUDA). To explore solutions for WUDA, this paper proposes two intuitive frameworks and conducts comparative experiments. We observe that the two frameworks behave differently when the datasets change. Therefore, we construct datasets with a wide range of domain shifts and conduct extended experiments to analyze the impact of domain shift changes on the two frameworks. In addition, to measure domain shift, we apply the metric representation shift to urban landscape image segmentation for the first time. The source code and constructed datasets can be obtained from this link: https://github.com/bupt-ai-cz/WUDA.
Chuang Zhu, Yuan Li 0014, Wenqi Tang
ICASSP2
2023 A Self-Training Framework Based on Multi-Scale Attention Fusion for Weakly Supervised Semantic Segmentation
abstract
Weakly supervised semantic segmentation (WSSS) based on image-level labels is challenging since it is hard to obtain complete semantic regions. To address this issue, we propose a self-training method that utilizes fused multi-scale class-aware attention maps. Our observation is that attention maps of different scales contain rich complementary information, especially for large and small objects. Therefore, we collect information from attention maps of different scales and obtain multi-scale attention maps. We then apply denoising and reactivation strategies to enhance the potential regions and reduce noisy areas. Finally, we use the refined attention maps to retrain the network. Experiments showthat our method enables the model to extract rich semantic information from multi-scale images and achieves 72.4% mIou scores on both the PASCAL VOC 2012 validation and test sets. The code is available at https://bupt-ai-cz.github.io/SMAF.
Chuang Zhu
ICME2
2023 Semi-supervised Domain Adaptation via Prototype-based Multi-level Learning
abstract
In semi-supervised domain adaptation (SSDA), a few labeled target samples of each class help the model to transfer knowledge representation from the fully labeled source domain to the target domain. Many existing methods ignore the benefits of making full use of the labeled target samples from multi-level. To make better use of this additional data, we propose a novel Prototype-based Multi-level Learning (ProML) framework to better tap the potential of labeled target samples. To achieve intra-domain adaptation, we first introduce a pseudo-label aggregation based on the intra-domain optimal transport to help the model align the feature distribution of unlabeled target samples and the prototype. At the inter-domain level, we propose a cross-domain alignment loss to help the model use the target prototype for cross-domain knowledge transfer. We further propose a dual consistency based on prototype similarity and linear classifier to promote discriminative learning of compact target feature representation at the batch level. Extensive experiments on three datasets, including DomainNet, VisDA2017, and Office-Home, demonstrate that our proposed method achieves state-of-the-art performance in SSDA. Our code is available at https://github.com/bupt-ai-cz/ProML.
Chuang Zhu
IJCAI2
2023 Sample Prior Guided Robust Model Learning to Suppress Noisy Labels
Chuang Zhu
ECML/PKDD (2)2
2023 Negative Prototypes Guided Contrastive Learning for Weakly Supervised Object Detection
Chuang Zhu
ECML/PKDD (2)2
2022 Dual-branch network via pseudo-label training for thyroid nodule detection in ultrasound image
Ruoning Song, Chuang Zhu, Long Zhang 0020, Yihao Luo, Jun Liu 0014, Jie Yang 0023
Appl. Intell.2
2022 Construct informative triplet with two-stage hard-sample generation
Chuang Zhu, Huihui Dong, Zekuan Yu, Shangshang Zhang
Neurocomputing1
2022 DigestPath: A benchmark dataset with challenge review for the pathological detection and segmentation of digestive-system
Qian Da, Zhongyu Li 0002, Yanfei Zuo, Chenbin Zhang, Jingxin Liu 0005, Wen Chen 0001, Jiahui Li 0005, Dou Xu, Hongmei Yi, Zhe Wang 0043, Li Zhang 0040, Xianying He, Xiaofan Zhang 0002, Ke Mei, Chuang Zhu, Weizeng Lu, LinLin Shen, Jun Shi 0006, Jun Li 0106, Sreehari S, Ganapathy Krishnamurthi, Jiangcheng Yang, Tiancheng Lin 0001, Qingyu Song 0004, Xuechen Liu 0004, Simon Graham, Raja Muhammad Saad Bashir, Canqian Yang, Shaofei Qin, Xinmei Tian 0001, Jie Zhao 0014, Dimitris N. Metaxas, Hongsheng Li 0001, Chaofu Wang, Shaoting Zhang 0001
Medical Image Anal.19
2022 Hard Sample Aware Noise Robust Learning for Histopathology Image Classification
abstract
Deep learning-based histopathology image classification is a key technique to help physicians in improving the accuracy and promptness of cancer diagnosis. However, the noisy labels are often inevitable in the complex manual annotation process, and thus mislead the training of the classification model. In this work, we introduce a novel hard sample aware noise robust learning method for histopathology image classification. To distinguish the informative hard samples from the harmful noisy ones, we build an easy/hard/noisy (EHN) detection model by using the sample training history. Then we integrate the EHN into a self-training architecture to lower the noise rate through gradually label correction. With the obtained almost clean dataset, we further propose a noise suppressing and hard enhancing (NSHE) scheme to train the noise robust model. Compared with the previous works, our method can save more clean samples and can be directly applied to the real-world noisy dataset scenario without using a clean subset. Experimental results demonstrate that the proposed scheme outperforms the current state-of-the-art methods in both the synthetic and real-world noisy datasets. The source code and data are available at https://github.com/bupt-ai-cz/HSA-NRL/.
Chuang Zhu, Ying Wang 0043, Mulan Jin
IEEE Trans. Medical Imaging1
2021 SFCN: Symmetric feature comparison network for detecting ischemic stroke lesions on CT images
abstract
Abstract Ischemic stroke is the most common stroke and the leading cause of disability and death in the world. Computed tomography (CT) is a popular and economical diagnostic device for the stroke, However the ischemic stroke lesions are not evident on CT images and the diagnostic result relies on the visual observation of neurologists, which may vary from doctor to doctor. To facilitate the treatment, a computer‐aided detection algorithm on CT images is proposed to help clinician for the ischemic stroke screening. In order to obtain accurate lesion annotation on CT images, novel automatic algorithms are developed to achieve image pairing, calibration, and registration. Then, a new framework with the symmetric feature extraction and comparison is proposed to identify and locate the ischemic stroke lesion. Experimental results show that this method achieves 75% of DICE in the detection of ischemic stroke lesions, which is higher than other methods by 4%. Its competitive results compared with seven latest methods is shown in terms of extensive qualitative and quantitative evaluation. This method can accurately detect the lesion in the CT images through the comparison of symmetric regional features, which has contributed to the clinical diagnosis of ischemic stroke.
Long Zhang 0020, Chuang Zhu, Yuewei Wu, Yang Yang 0007, Yihao Luo, Ruoning Song, Jie Yang 0023
IET Image Process.2
2021 Multi-level colonoscopy malignant tissue detection with adversarial CAC-UNet
Chuang Zhu, Ke Mei, Yihao Luo, Jun Liu 0014, Ying Wang 0043, Mulan Jin
Neurocomputing1
2020 Instance Adaptive Self-training for Unsupervised Domain Adaptation
Ke Mei, Chuang Zhu, Jiaqi Zou, Shanghang Zhang
ECCV (26)2
2020 Cross-Stained Segmentation from Renal Biopsy Images Using Multi-Level Adversarial Learning
abstract
Segmentation from renal pathological images is a key step in automatic analyzing the renal histological characteristics. However, the performance of models varies significantly in different types of stained datasets due to the appearance variations. In this paper, we design a robust and flexible model for cross-stained segmentation. It is a novel multi-level deep adversarial network architecture that consists of three sub-networks: (i) a segmentation network; (ii) a pair of multi-level mirrored discriminators for guiding the segmentation network to extract domain-invariant features; (iii) a shape discriminator that is utilized to further identify the output of the segmentation network and the ground truth. Experimental results on glomeruli segmentation from renal biopsy images indicate that our network is able to improve segmentation performance on target type of stained images and use unlabeled data to achieve similar accuracy to labeled data. In addition, this method can be easily applied to other tasks.
Ke Mei, Chuang Zhu, Jun Liu 0014, Yuanyuan Qiao 0002
ICASSP2
2020 Attention Based Multi-Instance Thyroid Cytopathological Diagnosis with Multi-Scale Feature Fusion
abstract
In recent years, deep learning has been popular in combining with cytopathology diagnosis. Using the whole slide images (WSI) scanned by electronic scanners at clinics, researchers have developed many algorithms to classify the slide (benign or malignant). However, the key area that support the diagnosis result can be relatively small in a thyroid WSI, and only the global label can be acquired, which make the direct use of the strongly supervised learning framework infeasible. What's more, because the clinical diagnosis of the thyroid cells requires the use of visual features in different scales, a generic feature extraction way may not achieve good performance. In this paper, we propose a weakly supervised multi-instance learning framework based on attention mechanism with multi-scale feature fusion (MSF) using convolutional neural network (CNN) for thyroid cytopathological diagnosis. We take each WSI as a bag, each bag contains multiple instances which are the different regions of the WSI, our framework is trained to learn the key area automatically and make the classification. We also propose a feature fusion structure, merge the low-level features into the final feature map and add an instance-level attention module in it, which improves the classification accuracy. Our model is trained and tested on the collected clinical data, reaches the accuracy of 93.2 %, which outperforms the other existing methods. We also tested our model on a public histopathology dataset and achieves better result than the state-of-the-art deep multi-instance method.
Shuhao Qiu, Yao Guo 0004, Chuang Zhu
ICPR3
2019 Breast Cancer Image Classification on WSI with Spatial Correlations
abstract
As common cancer, breast cancer kills thousands of women every year. It’s significant to provide doctors computer-aided diagnosis (CAD) to ease their workload as well as improve detection quality. Patch-level CNNs are usually used to classify the breast tissue slice, and the CNNs classify each patch independently ignoring the spatial correlations, resulting in wrong isolated label map. However, the probability distribution of cancer type is related to their adjacent patches. In this paper, we propose a framework integrating CNN and filter algorithm aimed at extracting spatial information and improving the performance of the classification. The network was trained on a breast cancer dataset provided by ICIAR18. For 4-class classification, compared to CNN methods without using spatial correlations, the proposed method achieved about 10% improvement on accuracy over the validation dataset and get smoother probability maps. Our experiments also show that larger kernel size gets better performance. The code is available at https://github.com/dong100136/Breast-Cancer-Image-Classification-On-WSI-With-Spatial-Correlations.
Jiandong Ye, Yihao Luo, Chuang Zhu, Yue Zhang 0016
ICASSP3
2017 Low-light image enhancement using CNN and bright channel prior
abstract
In this paper, we propose a joint framework to enhance images under low-light conditions. First, a convolutional neural network (CNN) based architecture is proposed to denoise low-light images. Then, based on atmosphere scattering model, we introduce a low-light model to enhance image contrast. In our low-light model, we propose a simple but effective image prior, bright channel prior, to estimate the transmission parameter; besides, an effective filter is designed to adaptively estimate environment light in different image areas. Experimental results demonstrate that our method achieves superior performance over other methods.
Chuang Zhu, Jiawen Song, Huizhu Jia
ICIP2
2017 LLCNN: A convolutional neural network for low-light image enhancement
abstract
In this paper, we propose a CNN based method to perform low-light image enhancement. We design a special module to utilize multiscale feature maps, which can avoid gradient vanishing problem as well. In order to preserve image textures as much as possible, we use SSIM loss to train our model. The contrast of low-light images can be adaptively enhanced using our method. Results demonstrate that our CNN based method outperforms other contrast enhancement methods.
Chuang Zhu, Guoqing Xiang, Yuan Li 0014, Huizhu Jia
VCIP2
2016 Smart query expansion scheme for CDVS based on illumination and key features
abstract
Given a query image, retrieving images depicting the same object in a large scale database is becoming an urgent and challenging task. Recently, Compact Description for Visual Search (CDVS) is drafted by the ISO/IEC Moving Pictures Experts Group (MPEG) to support image retrieval applications, and it has been published as an international standard. Unfortunately, with regard to applications with hugely mutative illumination, perspective and noisy background, CDVS suffers from an inevitable performance loss. In this paper, firstly we introduce the query expansion to address performance loss caused by the scene complexity in CDVS. Secondly, a query expansion instance selection method based on illumination is proposed, which achieves better performance. Thirdly, we adopt a key feature matching score based weighted strategy in basic query expansion to improve retrieval performance. We evaluate our proposed methods on the Oxford (5K images) dataset and a reality traffic vehicle dataset (12K images), and the result shows that the proposed methods boost mean average precision (MAP) by 7% ∼ 10% in Oxford dataset and 7% ∼17% in vehicle dataset.
Chuang Zhu, Huizhu Jia, Ling-Yu Duan, Jiawen Song, Wen Gao 0001
ICPR2
2015 An adaptive inter CU depth decision algorithm for HEVC
abstract
The emerging High-Efficiency Video Coding (HEVC) standard has introduced a number of new coding tools, such as a quad-tree based coding unit (CU). The quadtree-structured coding unit achieves significant coding efficiency improvements compared to H264/AVC. However, the complexity of CU depth decision associated with Rate-Distortion (R-D) cost computation dramatically increased. In order to alleviate the computational burden in HEVC inter coding, a fast CU depth decision algorithm is proposed in this paper. Firstly, zero CU detection method for HEVC is proposed as early termination algorithm. Secondly, the CU depth pruning strategies are adaptively determined according to standard deviation of statistic spatiotemporal depth information. Finally, when the neighbors are not available or have a very weak correlation, edge gradient of current coding tree unit (CTU) is considered as main factor for CU depth pruning method. Experimental results demonstrate that, compared with the original HM16.0 implementation, the proposed algorithm achieves about 40.5% encoding time saving with ignorable coding performance degradation.
Jie Liu 0035, Huizhu Jia, Guoqing Xiang, Xiaofeng Huang, Binbin Cai, Chuang Zhu, Don Xie
VCIP6
2014 Low-delay window-based rate control scheme for video quality optimization in video encoder
abstract
The consistent video quality and encoding latency due to buffering are two important aspects in designing rate control scheme for the application of real-time video coding system. To well balance these two contrary objectives, we firstly analyze the constraint of buffer latency and the definition of a “consistent” video quality. Then a window-based rate control scheme is proposed with one window for controlling the rate and latency, while the other window for optimizing video quality. By applying low complexity frame level ratedistortion model in the testing sequences, our proposed method shows excellent performance in balancing the encoder buffer latency and optimized video quality. Besides, this one-pass rate control scheme is highly practical for the real-time video coding application.
Yuan Li 0014, Huizhu Jia, Chuang Zhu, Meng Li 0016, Wen Gao 0001
ICASSP3
2014 Inter-dependent rate-distortion modeling for video coding and its application to rate control
abstract
Rate control scheme using independent rate-distortion (R-D) model at the minimum coding unit (macroblock) level has been widely discussed in the literature where R-D optimization is performed without consideration of the inter-dependencies of different coding units. In this paper, we extend these techniques to the more general situations - rate control for inter-dependent video coding units. The interdependent distortion-quantization (D-Q) model and rate-quantization (R-Q) model are formulated separately based on the analysis of the relationship between the spatial-domain residual and the transform-domain residual. Then a window-based rate control scheme with frame bit allocation and video quality optimization is proposed, which uses the approximated R-D model to reduce the computational complexity. Simulation results demonstrate that the proposed algorithm shows excellent peak signal-to-noise ratio (PSNR) performance under the bit rate constraint. This one-pass rate control scheme is highly practical for the realtime video coding application.
Yuan Li 0014, Huizhu Jia, Pan Ma, Chuang Zhu, Wen Gao 0001
ICME4
2014 A resolution-adaptive interpolation filter for video codec
abstract
The fraction-pel interpolation filter varies in the video coding standards such as H.264/AVC, AVS and HEVC. Since fractional-pel motion compensation plays an important role in the video encoder, the interpolation of fractional-pel pixels can be refined and designed better to enhance the coding efficiency. In this paper, we firstly propose the generation algorithm of interpolation filter coefficients, and four different tap filters, namely 4tap, 6 tap, 8 tap and 10tap, are tested. A resolution-adaptive interpolation filter for different resolution videos is then introduced based on this algorithm to achieve the maximum bitrate saving. In the proposed scheme, 4 tap filter is applied for the UHD (2560×1600 and above) videos, 6 tap filter and 10 tap filter are performed in the videos whose resolution ranging from 720P (1280×720) to 1080P (1920×1080) and the videos with the resolution below 720P, respectively. When 4 tap filter and 6 tap filter are used in high-definition video, the coding efficiency can increase and the computational complexity will reduce greatly, which is actually beneficial to make hardware optimization more effectively especially SIMD (Single Instruction Multiple Data) and VLSI design. Experiments show that the average BD-rate gains on luma Y, chroma U and V are 1.4%, 0.7% and 0.7% for LP-Main configuration, when conducted in HEVC reference software HM11.0. The coding efficiency gains are significant for some video sequences and can reach up to 6.1%.
Ronggang Wang, Yuan Li 0014, Chuang Zhu, Huizhu Jia, Wen Gao 0001
ISCAS4
2014 Multi-level low-complexity coefficient discarding scheme for video encoder
abstract
Rate-Distortion (R-D) optimization technique plays an important role in video coding. R-D sense discarding (thresholding) technique can make great improvement on the coding efficiency. This work first proposes a multi-level coefficient discarding scheme, which is composed of coefficient-level (CL), block-level and macroblock-level discarding. In CL, coefficient-level R-D cost function is formulated and then CL discarding scheme is developed. At last, an effective implementation method is proposed to reduce the complexity of the proposed scheme. The experimental results show that our proposed multi-level discarding scheme can improve the coding performance of video encoder by 0.15db in average.
Chuang Zhu, Huizhu Jia, Jie Liu 0035, Xianghu Ji, Wen Gao 0001
ISCAS1
2014 Fast algorithm of coding unit depth decision for HEVC intra coding
abstract
The emerging high efficiency video coding standard (HEVC) achieves significantly better coding efficiency than all existing video coding standards. The quad tree structured coding unit (CU) is adopted in HEVC to improve the compression efficiency, but this causes a very high computational complexity because it exhausts all the combinations of the prediction unit (PU) and transform unit (TU) in every CU attempt. In order to alleviate the computational burden in HEVC intra coding, a fast CU depth decision algorithm is proposed in this paper. The CU texture complexity and the correlation between the current CU and neighbouring CUs are adaptively taken into consideration for the decision of the CU split and the CU depth search range. Experimental results show that the proposed scheme provides 39.3% encoder time savings on average compared to the default encoding scheme in HM-RExt-13.0 with only 0.6% BDBR penalty in coding performance.
Xiaofeng Huang, Huizhu Jia, Kaijin Wei, Jie Liu 0035, Chuang Zhu, Zhengguang Lv, Don Xie
VCIP5
2014 Window-based rate control for video quality optimization with a novel INTER-dependent rate-distortion model
Yuan Li 0014, Huizhu Jia, Chuang Zhu, Mingyuan Yang, Wen Gao 0001
Signal Process. Image Commun.3
2013 On a Highly Efficient RDO-Based Mode Decision Pipeline Design for AVS
abstract
Rate distortion optimization (RDO) is the best known mode decision method, while the high implementation complexity limits its applications and almost no real-time hardware encoder is truly full-featured RDO based. In this paper, first, a full-featured RDO-based mode decision (MD) algorithm is presented, which makes more modes enter RDO process. Second, the throughput of RDO-based MD pipeline is thoroughly analyzed and modeled. Third, a highly efficient adaptive block-level pipelining architecture of RDO-based MD for AVS video encoder is proposed which can achieve the highest throughput to alleviate the RDO burden. Our design is described in high-level Verilog/VHDL hardware description language and implemented under SMIC 0.18-$\mu$m CMOS technology with 232 K logic gates and 85 Kb SRAMs. The implementation results validate our architectural design and the proposed architecture can support real time processing of 1080P@30 fps. The coding efficiency of our adopted method far outperforms (0.57 dB PSNR gain in average) the traditional low-complexity MD (LCMD) methods and the throughput of our designed pipeline is increased by 11.3%, 19% and 17% for I, P and B frames, respectively, compared with the existed RDO-based architecture.
Chuang Zhu, Huizhu Jia, Shanghang Zhang, Xiaofeng Huang, Wen Gao 0001
IEEE Trans. Multim.1
2011 A hardware-efficient architecture for multi-resolution motion estimation using fully reconfigurable processing element array
abstract
Integer motion estimation (IME) for block-based video coding presents a significant challenge in external memory bandwidth, data latency, and circuit area with the increase of coding complexity and video resolution. To conquer these problems, this paper proposes a hardware-efficient VLSI architecture for multi-resolution motion estimation algorithm (MMEA) based on fully reconfigurable processing element (PE) array. On-chip storage and PE array are carefully designed to support parallel computation and hardware resource sharing. In addition, low data latency is obtained by arranging internal logics in parallel according to the data dependency. As a result, our design can support real time processing of 1080P@30fps with 2 reference frames and a search range of 256×192 and it is implemented under SMIC 0.18-µm CMOS technology with 920K logic gates and 192 KB SRAMs. Compared with previous work, our design can achieve the best performance-price rate benefiting from the proposed re-configurable PE array.
Xianghu Ji, Chuang Zhu, Huizhu Jia, Hai Bing Yin
ICME2
2011 A highly efficient pipeline architecture of RDO-based mode decision design for AVS HD video encoder
abstract
Like H.264, AVS video coding standard also uses macroblock (MB) based motion compensation (MC) and mode decision (MD). Rate distortion optimization (RDO) is the best known mode decision method, but with a high computational complexity that limits its applications. In our paper, firstly an MD algorithm based on RDO is given, which makes more mode candidates enter into RDO mode decision with little hardware resource increment. We further analyze the pipeline structure in detail, and implement a block-level 5-stage hardware pipeline. It can support the real time RDO mode decision processing of 1080P@30fps, and the coding efficiency is about 0.5db higher than the traditional SAD method. Our design is described in high-level Verilog/VHDL hardware description language and implemented under SMIC 0.18-µm CMOS technology with 215K logic gates and 80 KB SRAMs.
Chuang Zhu, Yuan Li 0014, Huizhu Jia, Hai Bing Yin
ICME1
2011 Adaptive integer-precision Lagrange multiplier selection for high performance AVS video coding
abstract
In AVS and H.264/AVC, Lagrangian Rate distortion (RD) optimization techniques are widely adopted for coding mode selection and displacement vector estimation. The optimal Lagrange multipliers in these two cases are both floating-point values. If RD optimized video encoder is implemented on computation-constrained fixed-point platform such as FPGA and ASIC, fixed-point Lagrange multiplier selection is an important problem to trade-off the RD performance and computation complexity. This work focuses on fixed-point Lagrange multiplier selection for RD mode decision. Adaptive scaling matrix is used to trade-off complexity and RD performance. Also, intensive simulation results and analysis on precision, hardware cost, and RD performance are given. The proposed approach is also well-suited for RD optimized motion estimation for computation-constrained video coding.
Hai Bing Yin, Bingqian Zhou, Chuang Zhu, Huizhu Jia
VCIP3
2010 Algorithm analysis and architecture design for rate distortion optimized mode decision in high definition AVS video encoder
Hai Bing Yin, Honggang Qi, Huizhu Jia, Chuang Zhu
Signal Process. Image Commun.4