VLDB 2026 Research / reviewers in the wild / expert
Huizhu Jia
dblp:56/8496
· DBLP profile ↗
75ranked-venue papers
0as first author
21since 2021 · last 2025
0000-0002-2778-3768ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 62 · 18 since 2021Artificial intelligence and machine learning · 10 · 5 since 2021Systems, architecture and hardware · 9 · 1 since 2021Databases, data management, data science and information retrieval · 4 · 3 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DesignEdit: Unify Spatial-Aware Image Editing via Training-free Inpainting with a Multi-Layered Latent Diffusion FrameworkabstractSpatial-aware image editing focuses on modifying the position and size of elements within a given image. However, previous works still struggle with maintaining background harmony in the original editing areas, as well as preserving the initial identity of the edited elements, making it difficult to achieve complex multi-object editing in a single pass. In this paper, we aim to perform flexible spatial editing in a simple yet straightforward manner. We propose to inpaint the background first and develop a two-stage multi-layered latent diffusion framework to edit each element independently. Specifically, we design a key-masking self-attention scheme alongside artifact suppression to achieve background inpainting within the denoising process, leveraging the powerful generative capabilities of the Latent Diffusion Model, Stable Diffusion XL-1.0. The latent decomposition and fusion framework is capable of unifying various spatial-aware operations, including removal, resizing, relocation, flipping, addition, camera panning, zooming out, occlusion-aware editing, and cross-image editing. Experiments demonstrate the superior inpainting quality for object removal, along with enhanced versatility and higher precision in spatial-aware editing achieved by our method. Yueru Jia, Aosong Cheng, Yuhui Yuan, Chuke Wang, Ji Li 0006, Huizhu Jia, Shanghang Zhang |
AAAI | 6 |
| 2025 | Three-Stage Progressive Pre-Analysis Framework for VMAF Controllable Image CodingabstractTo achieve controllable subjective quality in image coding, this paper proposes a Video Multi-method Assessment Fusion (VMAF)-oriented image coding pre-analysis algorithm, enabling the adaptive derivation of quantization parameters corresponding to a specified quality target. First, a$Q-\mathcal{V}$model is constructed to describe the relationship between encoding quantization and VMAF distortion. Then, a Three-stage Progressive Control (TPC) algorithm, shown in Fig. 1(a), is designed to adapt quantization parameters using the discovered$Q-\mathcal{V}$model. The first two stages, based on lightweight feature extraction, iteratively fit the distortion metrics intrinsically calculated by VMAF to predict the VMAF value for a given sample under specified distortion conditions. The final stage fits the$Q-\mathcal{V}$model parameters using multi-point VMAF distortion data and outputs the corresponding quantization step for encoder control. A two-pass refinement algorithm, depicted in Fig. 1(b), further adjusts the quantization parameters based on the first encoding pass, improving quality control accuracy and framework robustness. Experiments on four datasets show that the quality control error remains below 1.293% for various VMAF targets, and the two-pass refinement reduces it further to 0.710%, outperforming existing methods. Guoqing Xiang, Wenzhao Li, Mingyuan Yang, Fan Yang 0053, Shanghang Zhang, Huizhu Jia |
DCC | 8 |
| 2025 | Hardware-friendly rate estimation algorithm and architecture design for AVS3
Yunyao Yan, Guoqing Xiang, Jie Chen 0001, Xiaofeng Huang, Peng Zhang 0007, Huizhu Jia |
Multim. Tools Appl. | 6 |
| 2025 | Scene-Adaptive Unsupervised Crowd Counting for Video SurveillanceabstractIn recent years, significant advancements in deep learning have expanded its application in a variety of computer vision tasks. However, the performance of these models heavily depends on the quality of the training data. While existing crowd counting methods yield satisfactory results on labeled datasets, they often face serious domain adaptation issues when applied to unlabeled data, the latter being more common in real-world scenarios. To mitigate this issue, we present a novel Scene-adaptive Unsupervised Crowd Counting (SUCC) framework aimed at enhancing the domain adaptability of counting models. This framework integrates a bi-branch attention network (BBA-Net) that leverages human prior knowledge to generate highly accurate density and anchor maps, which are essential for producing intermediate domain data as pseudo labels. Our SUCC framework eliminates the need for laborious manual annotation within the new data domain. Instead, it continually performs adaptive intermediate domain generation and model fine-tuning, establishing a beneficial feedback loop. Comprehensive experiments on multiple video crowd counting datasets show that our SUCC framework significantly improves domain generalizability. Furthermore, it exhibits satisfactory model stability and algorithm interpretability, attributes that are vital for the practical deployment of counting applications. The open-source code and model weights can be found on Github. Rui Ma 0032, Chenxuan Li 0003, Huizhu Jia |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2024 | VLUReID: Exploiting Vision-Language Knowledge for Unsupervised Person Re-IdentificationabstractThe superior performances of pre-trained vision-language models on various downstream tasks demonstrate the effectiveness of integrating cross-modal vision-language knowledge into visual tasks. However, this knowledge is hardly used for visual-based person re-identification (re-ID) because the datasets lack textual descriptions. Existing efforts require manual annotations for training, which can be time-consuming. We propose VLUReID, a framework that improves visual-based person re-ID using vision-language knowledge without requiring manual annotations from datasets. Specifically, the Vision-to-Text Association (VTA) module uses designed textual prompts to prompt the vision-language model in generating pseudo-semantic labels for visual inputs. Subsequently, within the Dual-Branch Asymmetric Training (DBAT) module, we propose an asymmetric training strategy to extract cross-modal knowledge from pseudo-semantic labels and integrate it into the person re-ID model. The experimental results on two widely-used benchmarks for unsupervised video-based person re-ID demonstrate the effectiveness of our framework. Ray Zhang 0002, Fan Yang 0053, Yuan Li 0014, Huizhu Jia, Shanghang Zhang |
ICME | 5 |
| 2024 | Joint Frame-Level and Block-Level Rate-Perception Optimized Preprocessing for Video Coding
Huajie Tan, Guoqing Xiang, Huizhu Jia |
MMAsia | 4 |
| 2024 | Two-Stage Perceptual Quality Oriented Rate Control Algorithm for HEVCabstractAs a practical technique in mainstream video coding applications, rate control dominates important to ensure compression quality with limited bitrates constraints. However, most rate control methods mainly focus on objective quality while ignoring the perceptual quality improvement for human eyes. In this paper, we propose a two-stage rate control algorithm to optimize the perceptual quality at the frame encoding stage and the coding tree unit (CTU) encoding stage for high efficiency video coding (HEVC), respectively. Firstly, for the frame encoding stage, with inter-frame distortion dependency consideration, a frame-level rate control method is presented by adjusting the frame-level Lagrange multiplier adaptively with a preprocessing method. Secondly, for the CTU encoding stage, we propose a saliency-based CTU-level perceptual quality rate control algorithm, which employs CTU-level saliency weight to adjust the perceptual rate-distortion (R-D) model. We conduct the CTU-level rate control by an optimized Lagrange multiplier and quantization parameter (QP) to achieve perceptual quality optimization. Extensive experimental results reveal that, compared with state-of-the-art rate control methods on HEVC, our algorithm achieves significant perceptual coding performance with improved subjective visual quality. Yunyao Yan, Guoqing Xiang, Huizhu Jia, Jie Chen 0001, Xiaofeng Huang |
ACM Trans. Multim. Comput. Commun. Appl. | 3 |
| 2023 | Multi-Agent Automated Machine LearningabstractIn this paper, we propose multi-agent automated machine learning (MA2ML) with the aim to effectively handle joint optimization of modules in automated machine learning (AutoML). MA2ML takes each machine learning module, such as data augmentation (AUG), neural architecture search (NAS), or hyper-parameters (HPO), as an agent and the final performance as the reward, to formulate a multi-agent reinforcement learning problem. MA2ML explicitly assigns credit to each agent according to its marginal contribution to enhance cooperation among modules, and incorporates off-policy learning to improve search efficiency. Theoretically, MA2ML guarantees monotonic improvement of joint optimization. Extensive experiments show that MA2ML yields the state-of-the-art top-1 accuracy on ImageNet under constraints of computational cost, e.g., 79.7%/80.5% with FLOPs fewer than 600M/800M. Exten\sive ablation studies verify the benefits of credit assignment and off-policy learning of MA2ML. Zhaozhi Wang, Kefan Su, Jian Zhang 0018, Huizhu Jia, Qixiang Ye, Zongqing Lu 0002 |
CVPR | 4 |
| 2023 | A Hardware-friendly CTU-level IME Algorithm for VVCabstractThe new coding tools improved the performance for H.266/VVC but also brought challenges for hardware integer motion estimation (IME). First, the data dependency in deriving a predicted motion vector (PMV) is more severe. Second, the overhead of IME is increased by the complex partition mechanism. The challenges are tougher for IME in coding tree unit (CTU) level pipelined encoder. In this paper, we propose a hardware-friendly CTU-level IME algorithm with three innovative designs. First, a PMV prediction is proposed to derive PMVs in advance. Second, all divided blocks are categorized into either binary/quadra tree (BTQT) or ternary tree (TT) blocks. The motion vectors (MVs) of BTQT blocks are estimated with a multi-resolution search. The MVs of TT blocks are inferred from the estimated MVs with an inference algorithm. The proposed algorithm suffers $ 1.20\%$ degradation but reduced the complexity by $ 80\%$ compared to the reference software. Xizhong Zhu, Guoqing Xiang, Xiaofeng Huang, Yunyao Yan, Huizhu Jia |
DCC | 5 |
| 2023 | An Efficient Real-Time Hardware Architecture for Deblocking Filter in AVS3abstractTo achieve higher video compression efficiency to cope with the demand for ultra high definition video applications, the AVS3 standard has been proposed recently. As a block partition-based coding standard, AVS3 suffers from the blocking artifact problem especially at low bitrates, which can be alleviated by deblocking filter. This paper presents an efficient hardware architecture for the deblocking filter for AVS3 with high-throughput. First, a fast and hardware-friendly algorithm is proposed to localize the blocking artifact boundaries. Then, a buffer organization method and a data caching strategy are proposed to solve the problem of data dependency among coding units. Based on the proposed optimized algorithm and caching strategy, a four-stage pipelined deblocking filter module architecture is designed. The experimental results show that the proposed architecture can achieve 4K@120fps video processing at 100MHz which is sufficient for real-time application. Xiaofeng Huang, Guoqing Xiang, Xizhong Zhu, Jiaojiao Yang, Peng Zhang 0007, Huizhu Jia |
ICME | 7 |
| 2023 | A Hardware-efficient Unified Motion Estimation for Video CodingabstractMotion estimation (ME) is one of the most critical tools in video coding and consumes the majority of the encoding complexity. Three types of ME are utilized in the latest video coding standards, namely integer, fractional, and affine MEs. They are implemented as three searches for the integer motion vector (IMV), fractional motion vector (FMV), and control point motion vectors (CPMVs). Many algorithms were proposed to reduce the complexity for them individually, but the overall overhead of three searches is still challenging for hardware implementations. Therefore, we propose a hardware-efficient Unified Motion Estimation (UME) to derive three types of MVs with only one search. An IME with sub-block refinement is performed to collect extra motion information while searching for the IMV. The FMV and CPMVs are then derived from the collected information using a mixed error surface and an overdetermined system. Compared to the default ME algorithms in VVC, the time cost for ME is reduced by 41.63% with a coding loss of only 0.87% under LDB configuration. For hardware implementations, the minimum required resources and corresponding latency are significantly reduced by 75.35% and 69.17%, respectively. Xizhong Zhu, Guoqing Xiang, Peng Zhang 0007, Huizhu Jia |
ACM Multimedia | 4 |
| 2023 | Frame-Recurrent Video Crowd CountingabstractSince video data contains temporal information, video crowd counting demonstrates more potential than single-frame crowd counting for scenarios requiring high accuracy. However, learning robust relationships among frames efficiently and cheaply is very challenging. Existing methods for video crowd counting lack explicit temporal correlation modeling and robustness, and they are complex. In this paper, we propose the Frame-Recurrent Video Crowd Counting (FRVCC) framework to solve these issues. Specifically, we design a frame-recurrent manner to recursively relate the density maps in the temporal dimension, which efficiently explores long-term inter-frame knowledge and ensures the continuity of feature map responses. FRVCC consists of three plug-in modules: an optical flow estimation module, a single-frame counting module, and a density map fusion module. For the fusion module, we propose the ResTrans network to robustly learn complementary features between visual-based and correlation-based feature maps through residual strategy and vision transformer. To constrain the output distribution to be consistent with the ground truth distribution, we introduce an adversarial loss to rectify the training process. Additionally, we release a large-scale synthetic video crowd-counting dataset, CrowdXV, to evaluate the proposed method and further improve its performance. We have conducted extensive experiments on several video-counting datasets. The results demonstrate that FRVCC achieves state-of-the-art performance and, concurrently, high generalization, high flexibility, and less complexity. Shanghang Zhang, Rui Ma 0032, Huizhu Jia |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2022 | Enhancing and Dissecting Crowd Counting by Synthetic DataabstractIn this article, we propose a simulated crowd counting dataset CrowdX, which has a large scale, accurate labeling, parameterized realization, and high fidelity. The experimental results of using this dataset as data enhancement show that the performance of the proposed streamlined and efficient benchmark network ESA-Net can be improved by 8.4%. The other two classic heterogeneous architectures MCNN and CSRNet pre-trained on CrowdX also show significant performance improvements. Considering many influencing factors determine performance, such as background, camera angle, human density, and resolution. Although these factors are important, there is still a lack of research on how they affect crowd counting. Thanks to the CrowdX dataset with rich annotation information, we conduct a large number of data-driven comparative experiments to analyze these factors. Our research provides a reference for a deeper understanding of the crowd counting problem and puts forward some useful suggestions in the actual deployment of the algorithm. Chengyang Li 0001, Yuheng Lu, Yuan Li 0014, Huizhu Jia |
ICASSP | 6 |
| 2022 | Efficient Algorithm and Hardware Architecture for Rate Estimation in Mode Decision of AVS3abstractTowards enabling advanced video coding for emerging ap-plications, the AVS3 standard has been developed recently, achieving twice the coding efficiency of the AVS2 stan-dard through complex coding tools including advanced rate-distortion optimization (RDO) to select the best mode. The bit-rates are produced with the Advanced-Entropy-Coding (AEC) in the RDO process of AVS3. However, AEC dom-inates the time complexity of RDO and among the steps, con-text updating and interval subdivision are performed recur-sively, which is not conducive to real-time application, espe-cially for the hardware implementation. Thus this paper pro-poses an adaptive rate estimation algorithm with a piece-wise linear function that is very friendly to hardware implemen-tation to accelerate the rate estimation process in the RDO for AVS3 practical applications. The proposed architecture can meet the requirement of 4K@120fps ultra-high-definition videos at 200 MHz, whereas the BD-Rate increases only by 0.67% under the All-Intra (AI) configuration. Yunyao Yan, Guoqing Xiang, Huizhu Jia, Yuan Li 0014, Peng Zhang 0007, Jie Chen 0001 |
ICME | 4 |
| 2022 | Towards Blind Watermarking: Combining Invertible and Non-invertible MechanismsabstractBlind watermarking provides powerful evidence for copyright protection, image authentication, and tampering identification.However, it remains a challenge to design a watermarking model with high imperceptibility and robustness against strong noise attacks. To resolve this issue, we present a framework Combining the Invertible and Non-invertible (CIN) mechanisms. The CIN is composed of the invertible part to achieve high imperceptibility and the non-invertible part to strengthen the robustness against strong noise attacks. For the invertible part, we develop a diffusion and extraction module (DEM) and a fusion and split module (FSM) to embed and extract watermarks symmetrically in an invertible way. For the non-invertible part, we introduce a non-invertible attention-based module (NIAM) and the noise-specific selection module (NSM) to solve the asymmetric extraction under a strong noise attack. Extensive experiments demonstrate that our framework outperforms the current state-of-the-art methods of imperceptibility and robustness significantly. Our framework can achieve an average of 99.99% accuracy and 67.66 dB PSNR under noise-free conditions, while 96.64% and 39.28 dB combined strong noise attacks. The code will be available in https://github.com/RM1110/CIN. Rui Ma 0032, Mengxi Guo, Fan Yang 0053, Yuan Li 0014, Huizhu Jia |
ACM Multimedia | 6 |
| 2021 | Hierarchically and Cooperatively Learning Traffic Signal ControlabstractDeep reinforcement learning (RL) has been applied to traffic signal control recently and demonstrated superior performance to conventional control methods. However, there are still several challenges we have to address before fully applying deep RL to traffic signal control. Firstly, the objective of traffic signal control is to optimize average travel time, which is a delayed reward in a long time horizon in the context of RL. However, existing work simplifies the optimization by using queue length, waiting time, delay, etc., as immediate reward and presumes these short-term targets are always aligned with the objective. Nevertheless, these targets may deviate from the objective in different road networks with various traffic patterns. Secondly, it remains unsolved how to cooperatively control traffic signals to directly optimize average travel time. To address these challenges, we propose a hierarchical and cooperative reinforcement learning method-HiLight. HiLight enables each agent to learn a high-level policy that optimizes the objective locally by selecting among the sub-policies that respectively optimize short-term targets. Moreover, the high-level policy additionally considers the objective in the neighborhood with adaptive weighting to encourage agents to cooperate on the objective in the road network. Empirically, we demonstrate that HiLight outperforms state-of-the-art RL methods for traffic signal control in real road networks with real traffic. Bingyu Xu, Yaowei Wang 0001, Zhaozhi Wang, Huizhu Jia, Zongqing Lu 0002 |
AAAI | 4 |
| 2021 | Hardware-Friendly Coding Unit Decision Scheme for HEVCabstractQuad-tree based coding unit partition in High-Efficiency Video Coding (HEVC) achieved significant coding efficiency improvements, but also brought increasing computational complexity. Especially, design challenges like data dependence, large area cost, and imbalance of processing time of each coding tree unit (CTU), make it hard to achieve a real-time structure for real-time hardware encoder for all CTU sizes. To solve these problems, we proposed a hardware-friendly fast CU decision scheme with multi-stage algorithms for HEVC hardware encoder, aiming at the most complex modules: IME (Integer Motion Estimation), FME/IP (Fractional Motion Estimation/ Intra Prediction), and MD (Mode Decision). Firstly, in IME stage, a zero block detection method based fast CU and PU decision algorithm was presented. Secondly, we presented an estimated RDO (Rate-Distortion Optimization) based algorithm in the Hadamard domain for the early CU decision further in FME/IP stage. Finally, under the condition of hardware computing time limitation of several CU sizes, we proposed a computation time constraint CU fast decision algorithm for MD stage. Experiments demonstrated that, compared with the original HM13.0 implementation, the proposed scheme achieved about 53.9% encoding time saving with merely 2.3% coding performance degradation. What's more, significant area cost and data dependency have been alleviated, which will be more hardware-friendly for HEVC encoder design. Ju Huang, Xiaofeng Huang, Guoqing Xiang, Yuan Li 0014, Huizhu Jia |
ISCAS | 5 |
| 2021 | Deep Human-Interaction and Association by Graph-Based Learning for Multiple Object Tracking in the Wild
Cong Ma 0006, Fan Yang 0053, Yuan Li 0014, Huizhu Jia, Wen Gao 0001 |
Int. J. Comput. Vis. | 4 |
| 2021 | Deep Trajectory Post-Processing and Position Projection for Single & Multiple Camera Multiple Object Tracking
Cong Ma 0006, Fan Yang 0053, Yuan Li 0014, Huizhu Jia, Wen Gao 0001 |
Int. J. Comput. Vis. | 4 |
| 2021 | Digital Retina: A Way to Make the City Brain More Efficient by Visual CodingabstractThe ubiquitous camera networks in the city brain system grow at a rapid pace, creating massive amounts of images and videos at a range of spatial-temporal scales and thereby forming the “biggest” big data. However, the sensing system often lags behind the construction of the fast-growing city brain system, in the sense that such exponentially growing data far exceed today’s sensing capabilities. Therefore, critical issues arise regarding how to better leverage the existing city brain system and significantly improve the city-scale performance in intelligent applications. To tackle the unprecedented challenges, we articulate a vision towards a novel visual computing framework, termed asdigital retina, which aligns high-efficiency sensing models with the emerging Visual Coding for Machine (VCM) paradigm. In particular, digital retina may consist of video coding, feature coding, model coding, as well as their joint optimization. The digital retina is biologically-inspired, rooted on the widely accepted view that the retina encodes the visual information for human perception, and extracts features by the brain downstream areas to disentangle the visual objects. Within the digital retina framework, three streams, i.e., video stream, feature stream, and model stream, work collaboratively over the end-edge-cloud platform. In particular, the compressed video stream serves for human vision, the compact feature stream targets for machine vision, and the model stream incrementally updates deep learning models to improve the performance of human/machine vision tasks. We have developed a prototype to demonstrate the technical advantages of digital retina, and extensive experiments have been conducted to validate that it is able to effectively support the video big data analysis and retrieval in the intelligent city system. In particular, up to$7000\times $compression ratio could be realized for visual data compression while maintaining competitive performance with pristine signal in a series of visual analysis tasks. Wen Gao 0001, Siwei Ma 0001, Ling-Yu Duan, Yonghong Tian 0001, Peiyin Xing, Yaowei Wang 0001, Shanshe Wang, Huizhu Jia, Tiejun Huang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 8 |
| 2021 | Part-aware Progressive Unsupervised Domain Adaptation for Person Re-IdentificationabstractUnsupervised domain adaptation (UDA) aims to mitigate the domain shift that occurs when transferring knowledge from a labeled source domain to an unlabeled target domain. While it has been studied for application in unsupervised person re-identification (ReID), the relations of feature distribution across the source and target domains remain underexplored, as they either ignore the local relations or omit the in-depth consideration of negative transfer when two domains do not share identical label spaces. In light of the above, this paper presents an innovative part-aware progressive adaptation network (PPAN) that exploits global and local relations for UDA-based ReID across domains. A multi-branch network is developed that explicitly learns discriminative feature representation from both whole-body images and body-part images under the supervision of a labeled source domain. Within each network branch, an independent UDA constraint is designed that aligns the global and local feature distributions from a labeled source domain with those of an unlabeled target domain. In addition, a novel progressive adaptation strategy (PAS) is designed that effectively alleviates the negative influence of outlier source identities. The proposed unsupervised ReID model is evaluated on five widely used datasets (Market-1501, DukeMTMC-reID, CUHK03, VIPeR and PRID), and experimental results demonstrate its superior robustness and effectiveness relative to state-of-the-art approaches. Fan Yang 0053, Shijian Lu, Huizhu Jia, Don Xie, Zongqiao Yu, Feiyue Huang, Wen Gao 0001 |
IEEE Trans. Multim. | 4 |
| 2020 | FFA-Net: Feature Fusion Attention Network for Single Image DehazingabstractIn this paper, we propose an end-to-end feature fusion at-tention network (FFA-Net) to directly restore the haze-free image. The FFA-Net architecture consists of three key components:1) A novel Feature Attention (FA) module combines Channel Attention with Pixel Attention mechanism, considering that different channel-wise features contain totally different weighted information and haze distribution is uneven on the different image pixels. FA treats different features and pixels unequally, which provides additional flexibility in dealing with different types of information, expanding the representational ability of CNNs. 2) A basic block structure consists of Local Residual Learning and Feature Attention, Local Residual Learning allowing the less important information such as thin haze region or low-frequency to be bypassed through multiple local residual connections, let main network architecture focus on more effective information. 3) An Attention-based different levels Feature Fusion (FFA) structure, the feature weights are adaptively learned from the Feature Attention (FA) module, giving more weight to important features. This structure can also retain the information of shallow layers and pass it into deep layers.The experimental results demonstrate that our proposed FFA-Net surpasses previous state-of-the-art single image dehazing methods by a very large margin both quantitatively and qualitatively, boosting the best published PSNR metric from 30.23 dB to 36.39 dB on the SOTS indoor test dataset. Code has been made available at GitHub. Xu Qin, Zhilin Wang, Yuanchao Bai, Huizhu Jia |
AAAI | 5 |
| 2020 | BBA-NET: A Bi-Branch Attention Network For Crowd CountingabstractIn the field of crowd counting, the current mainstream CNNbased regression methods simply extract the density information of pedestrians without finding the position of each person. This makes the output of the network often found to contain incorrect responses, which may erroneously estimate the total number and not conducive to the interpretation of the algorithm. To this end, we propose a Bi-Branch Attention Network (BBA-NET) for crowd counting, which has three innovation points. i) A two-branch architecture is used to estimate the density information and location information separately. ii) Attention mechanism is used to facilitate feature extraction, which can reduce false responses. iii) A new density map generation method combining geometric adaptation and Voronoi split is introduced. Our method can integrate the pedestrian’s head and body information to enhance the feature expression ability of the density map. Extensive experiments performed on two public datasets show that our method achieves a lower crowd counting error compared to other state-of-the-art methods. Chengyang Li 0001, Fan Yang 0053, Cong Ma 0006, Yuan Li 0014, Huizhu Jia |
ICASSP | 7 |
| 2020 | Fusion Target Attention Mask Generation Network For Video SegmentationabstractVideo segmentation aims to segment target objects in a video sequence, which remains a challenge due to the motion and deformation of objects. In this paper, we propose a novel attention-driven hybrid encoder-decoder network that generates object segmentation by fully leveraging spatial and temporal information. Firstly, a multi-branch network is designed to learn feature representation from object appearance, location and motion. Secondly, a target attention module is proposed to further exploit context information from learned representation. In addition, a novel edge loss is designed which constraints the model to generate salient edge features and accurate segmentation. The proposed model has been evaluated over two widely used public benchmarks, and experiments demonstrate its superior robustness and effectiveness as compared with the state of the arts. Yunyi Li, Fangping Chen, Fan Yang 0053, Yuan Li 0014, Huizhu Jia |
ICIP | 5 |
| 2020 | Optical Flow-Guided Mask Generation Network for Video SegmentationabstractThe purpose of video segmentation is to segment foreground objects from a video sequence. In this paper, we propose a CNN based method for the semi-supervised video object segmentation, where a hybrid encoder-decoder network is designed to generate pixel-wise foreground object segmentation in use of both spatial and temporal information. In order to minimize cumulative error of the network as much as possible, we develop a two-stage training scheme: alternate training and back-propagation-through-time training. Then the performances of our method and other state-of-the-art ones are compared on two annotated video segmentation databases. Furthermore, we also run an extensive ablation study to test the effects of different components from our method. Yunyi Li, Fangping Chen, Fan Yang 0053, Cong Ma 0006, Yuan Li 0014, Huizhu Jia |
ISCAS | 6 |
| 2020 | A Novel Quality Enhanced Low Complexity Rate Control Algorithm for HEVCabstractRate control (RC) is a key technology in video coding which is mainly responsible for adapting the compressed video quality as much as possible under limited bandwidth. Typical RC methods consist of initial quantization parameters of I-frames decision and control parameters updating procedures. However, on the one hand, I-frames quantization parameter (QP) is decided by the sum of absolute transformed difference (SATD) computation, which is time consuming for low delay applications. On the other hand, the RC parameters are only updated ignoring the distortion characteristics for inter frame, which cannot achieve optimal rate distortion (RD) performance. Therefore, a novel RC method which considers distortion characteristics for model parameters updating for quality enhancement is proposed in this paper, including a low complexity based I-frame QP decision strategy of low delay applications. For which estimated distortion characteristics and the previous QPs are employed, respectively. Through the experimental results, with more accurate RC precision and negligible I-frame QP computation time, the proposed rate control scheme improves the performance gain of 2.6% bitrate savings for the whole test sequences on average in HM 16.9. ChungWen Ku, Guoqing Xiang, Huizhu Jia, Yuehua Cui, Yuan Li 0014 |
VCIP | 4 |
| 2020 | An adaptive spatio-temporal perception aware quantization algorithm for AVS2
Yunyao Yan, Guoqing Xiang, Yuan Li 0014, Huizhu Jia |
J. Vis. Commun. Image Represent. | 5 |
| 2020 | Single-Image Blind Deblurring Using Multi-Scale Latent Structure PriorabstractBlind image deblurring is a challenging problem in computer vision, which aims to restore both the blur kernel and the latent sharp image from only a blurry observation. Inspired by the prevalent self-example prior in image super-resolution, in this paper, we observe that a coarse enough image down-sampled from a blurry observation is approximately a low-resolution version of the latent sharp image. We prove this phenomenon theoretically and define the coarse enough image as a latent structure prior of the unknown sharp image. Starting from this prior, we propose to restore sharp images from the coarsest scale to the finest scale on a blurry image pyramid and progressively update the prior image using the newly restored sharp image. These coarse-to-fine priors are referred to as multi-scale latent structures (MSLSs). Leveraging the MSLS prior, our algorithm comprises two phases: 1) we first preliminarily restore sharp images in the coarse scales and 2) we then apply a refinement process in the finest scale to obtain the final deblurred image. In each scale, to achieve lower computational complexity, we alternately perform a sharp image reconstruction with fast local self-example matching, an accelerated kernel estimation with error compensation, and a fast non-blind image deblurring, instead of computing any computationally expensive non-convex priors. We further extend the proposed algorithm to solve more challenging non-uniform blind image deblurring problem. The extensive experiments demonstrate that our algorithm achieves the competitive results against the state-of-the-art methods with much faster running speed. Yuanchao Bai, Huizhu Jia, Ming Jiang 0001, Xianming Liu 0005, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2019 | Deep Association: End-to-end Graph-Based Learning for Multiple Object Tracking with Conv-Graph Neural NetworkabstractMultiple Object Tracking (MOT) has a wide range of applications in surveillance retrieval and autonomous driving. The majority of existing methods focus on extracting features by deep learning and hand-crafted optimizing bipartite graph or network flow. In this paper, we proposed an efficient end-to-end model, Deep Association Network (DAN), to learn the graph-based training data, which are constructed by spatial-temporal interaction of objects. DAN combines Convolutional Neural Network (CNN), Motion Encoder (ME) and Graph Neural Network (GNN). The CNNs and Motion Encoders extract appearance features from bounding box images and motion features from positions respectively, and then the GNN optimizes graph structure to associate the same object among frames together. In addition, we presented a novel end-to-end training strategy for Deep Association Network. Our experimental results demonstrate the effectiveness of DAN up to the state-of-the-art methods without extra-dataset on MOT16 and DukeMTMCT. Cong Ma 0006, Yuan Li 0014, Fan Yang 0053, Ziwei Zhang 0003, Yueqing Zhuang, Huizhu Jia |
ICMR | 6 |
| 2019 | Attention driven person re-identification
Fan Yang 0053, Shijian Lu, Huizhu Jia, Wen Gao 0001 |
Pattern Recognit. | 4 |
| 2019 | A Novel Joint Rate Allocation Scheme of Multiple StreamsabstractEncoding multiple videos in parallel and transmitting them as one joint stream over a limited bandwidth have become a popular strategy for broadcasting, which brings an opportunity to allocate different bitrate for each sequence to meet different demands. In this paper, considering visual experience for human beings, we propose a joint rate allocation scheme aims to reach an equal visual quality among all sequences by minimizing the distortion variance of all the sequences (denoted as minVAR problems). Existing methods assigned bits directly in proportion to their complexity measures and we named them as complexity based allocation scheme (CAS) methods. CAS methods rely on the accuracy of the complexity measures which can hardly be improved under limited computing resources. Also complexities may not be directly related to the distortions. To address these problems, we present a novel joint rate-distortion (R-D) based allocation scheme (RDAS) in this paper. Our proposed scheme can fit for different R-D models and in our method we model the R-D relationship with a hyperbolic function (RDAS-H). We also derive a closed-form solution of RDAS-H by a proposed joint R-D relationship. We integrated the RDAS-H method in high efficiency video coding reference software HM16.0. Experimental results demonstrate that our RDAS-H saves 75.29% variance on average over the related CAS-based method, where we apply both low delay and random access configurations with four different overall bandwidths for all classes recommended by the Joint Collaborative Team on Video Coding. Besides, RDAS-H also saves 36.62% variance on average over our previous method. The proposed RDAS-H method improves the performance significantly while requiring negligible computational cost. Hongfei Fan, Lin Ding 0002, Huizhu Jia |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2018 | Dense Relation Network: Learning Consistent and Context-Aware Representation for Semantic Image SegmentationabstractSemantic image segmentation, which aims at assigning pixel-wise category, is one of challenging image understanding problems. Global context plays an important role on local pixel-wise category assignment. To make the best of global context, in this paper, we propose dense relation network (DRN) and context-restricted loss (CRL) to aggregate global and local information. DRN uses Recurrent Neural Network (RNN) with different skip lengths in spatial directions to get context-aware representations while CRL helps aggregate them to learn consistency. Compared with previous methods, our proposed method takes full advantage of hierarchical contextual representations to produce high-quality results. Extensive experiments demonstrate that our method achieves significant state-of-the-art performances on Cityscapes and Pascal Context benchmarks, with mean-IoU of 82.8% and 49.0% respectively. Yueqing Zhuang, Fan Yang 0053, Cong Ma 0006, Ziwei Zhang 0003, Yuan Li 0014, Huizhu Jia, Wen Gao 0001 |
ICIP | 7 |
| 2018 | Trajectory Factory: Tracklet Cleaving and Re-Connection by Deep Siamese Bi-GRU for Multiple Object TrackingabstractMulti-Object Tracking (MOT) is a challenging task in the complex scene such as surveillance and autonomous driving. In this paper, we propose a novel tracklet processing method to cleave and re-connect tracklets on crowd or longterm occlusion by Siamese Bi-Gated Recurrent Unit (GRU). The tracklet generation utilizes object features extracted by CNN and RNN to create the high-confidence tracklet candidates in sparse scenario. Due to mis-tracking in the generation process, the tracklets from different objects are split into several sub-tracklets by a bidirectional GRU. After that, a Siamese GRU based tracklet re-connection method is applied to link the sub-tracklets which belong to the same object to form a whole trajectory. In addition, we extract the track-let images from existing MOT datasets and propose a novel dataset to train our networks. The proposed dataset contains more than 95160 pedestrian images. It has 793 different persons in it. On average, there are 120 images for each person with positions and sizes. Experimental results demonstrate the advantages of our model over the state-of-the-art methods on MOTI6. Cong Ma 0006, Changshui Yang, Fan Yang 0053, Yueqing Zhuang, Ziwei Zhang 0003, Huizhu Jia |
ICME | 6 |
| 2018 | RelationNet: Learning Deep-Aligned Representation for Semantic Image SegmentationabstractSemantic image segmentation, which assigns labels in pixel level, plays a central role in image understanding. Recent approaches have attempted to harness the capabilities of deep learning. However, one central problem of these methods is that deep convolutional neural network gives little consideration to the correlation among pixels. To handle this issue, in this paper, we propose a novel deep neural network named RelationNet, which utilizes CNN and RNN to aggregate context information. Besides, a spatial correlation loss is applied to train RelationNet to align features of spatial pixels belonging to same category. Importantly, since it is expensive to obtain pixel-wise annotations, we exploit a new training method to combine the coarsely and finely labeled data. Experiments show the detailed improvements of each proposal. Experimental results demonstrate the effectiveness of our proposed method to the problem of semantic image segmentation, which obtains state-of-the-art performance on the Cityscapes benchmark and Pascal Context dataset. Yueqing Zhuang, Fan Yang 0053, Cong Ma 0006, Ziwei Zhang 0003, Huizhu Jia |
ICPR | 6 |
| 2018 | Learning a collaborative multiscale dictionary based on robust empirical mode decomposition
Rui Chen 0006, Huizhu Jia, Wen Gao 0001 |
Neurocomputing | 2 |
| 2018 | A perceptually temporal adaptive quantization algorithm for HEVC
Guoqing Xiang, Huizhu Jia, Mingyuan Yang, Xinfeng Zhang 0001, Xiaofeng Huang, Jie Liu 0035 |
J. Vis. Commun. Image Represent. | 2 |
| 2018 | A novel adaptive quantization method for video coding
Guoqing Xiang, Huizhu Jia, Mingyuan Yang, Yuan Li 0014 |
Multim. Tools Appl. | 2 |
| 2017 | A cache-based bandwidth optimized motion compensation architecture for video decoderabstractIn video decoder applications, motion compensation (MC) is bandwidth consuming because of the non-regular memory access. Especially with the popularity of UHD video and the development of new coding standard (HEVC), external memory bandwidth becomes a crucial bottleneck. In this paper, we propose an area efficiency cache-based bandwidth optimization strategy to minimize the memory bandwidth. First a four-way parallel cache architecture is described. Then partially replacement strategy is proposed to further reduce memory bandwidth and power consumption. At last a column based storage scheme is provided to reduce the precharge/active frequency. We realize this idea using high level synthesis, which allow multiple iterations with quick turnaround time for micro architecture changes, and the results show that the averagely bandwidth reduction is up to 79.9% with moderate resource utilization, which outperforms the state-of-the-art works. Meng Li 0016, Huizhu Jia, Jason Cong, Wen Gao 0001 |
ICASSP | 2 |
| 2017 | Correlation preserving on graphs for image denoisingabstractIn this paper, we propose a novel dictionary-driven image denoising method based on correlation preserving on graphs. To overcome the drawbacks of the instable and unreliable correlations among a set of learned basis vectors, two effective regularized strategies are employed in our coding process. Specifically, a graph-based regularizer is built to preserve the global similarity, which can adaptively capture both geometric structures and discriminative features of textured patches. In particular, edge weights in the graph are obtained by seeking a nonnegative low-rank construction. Besides, a locality constraint is designed to automatically preserve not only spatial neighborhood information but also internal consistency present in noisy patches while learning an overcomplete dictionary. Experimental results show that our method achieves state-of-the-art denoising results in terms of both PSNR and subjective visual quality. Rui Chen 0006, Huizhu Jia, Wen Gao 0001 |
ICIP | 2 |
| 2017 | Low-light image enhancement using CNN and bright channel priorabstractIn this paper, we propose a joint framework to enhance images under low-light conditions. First, a convolutional neural network (CNN) based architecture is proposed to denoise low-light images. Then, based on atmosphere scattering model, we introduce a low-light model to enhance image contrast. In our low-light model, we propose a simple but effective image prior, bright channel prior, to estimate the transmission parameter; besides, an effective filter is designed to adaptively estimate environment light in different image areas. Experimental results demonstrate that our method achieves superior performance over other methods. Chuang Zhu, Jiawen Song, Huizhu Jia |
ICIP | 5 |
| 2017 | Fast rate distortion optimized quantization method for HEVCabstractRate-Distortion Optimized Quantization (RDOQ) brings significant improvement of coding performance in High Efficiency Video Coding (HEVC). However, it results in high computational complexity to determine the best quantization levels for each transform coefficient when applying rate and distortion optimization operation. In this paper, we firstly proposed an efficient way to skip the all zero coding units. Then we further propose a fast RDOQ scheme by skipping the optimization step of All Quantized Zero Blocks except DC Coefficient (AQZB-DC). Moreover, a rate difference model is established to select optimal quantization level for AQZB-DC blocks. Experiments are carried out on reference software HM 16.0 and the results show that the proposed method achieves 42.00% and 39.63% quantization time saving under Random Access (RA) and Low Delay (LD) configuration on average, while the BD-Rate loss is only 0.03% and 0.06%, respectively. Meng Wang 0017, Hongfei Fan, Shanshe Wang, Shengfu Dong, Guoqing Xiang, Huizhu Jia |
ISCAS | 8 |
| 2017 | LLCNN: A convolutional neural network for low-light image enhancementabstractIn this paper, we propose a CNN based method to perform low-light image enhancement. We design a special module to utilize multiscale feature maps, which can avoid gradient vanishing problem as well. In order to preserve image textures as much as possible, we use SSIM loss to train our model. The contrast of low-light images can be adaptively enhanced using our method. Results demonstrate that our CNN based method outperforms other contrast enhancement methods. Chuang Zhu, Guoqing Xiang, Yuan Li 0014, Huizhu Jia |
VCIP | 5 |
| 2016 | AnalogCast: Full linear coding and pseudo analog transmission for satellite remote-sensing imagesabstractIn this paper, we propose a novel image coding and transmission scheme called AnalogCast, which is a pseudo analog coding system for transmitting satellite remote-sensing images to large number of receivers. AnalogCast follows the idea originally developed for SoftCast [1-3] but with two special techniques, scale factor estimate based on L-shaped chunk division and scale factor fix curve-fitting model, so that there is no need to transmit side information as is required for SoftCast [1-3]. This improvement can get better mobility, less bandwidth and computational budget than SoftCast [1-3]. For robustness, AnalogCast adopts a mode of GOP (Group of Pictures) interweave structure to smooth image quality and improve perceptual quality when packet loss happens. Experimental results show that the proposed method can work well in low SNR condition in -7dB to 4dB where SoftCast system and JPEG2000 system fall into "cliff effect". And in high SNR condition, our proposed method works even better than SoftCast system at SNR 1.5~3dB and a gain of 1~10dB over JPEG2000 with forward error correction. For robustness, the AnalogCast works better than SoftCast at SNR 8dB when the burst packet loss rate is higher than 35%. And the scheme gets better perceptual quality under packet loss. Xiange Wen, Huizhu Jia, Wen Gao 0001 |
ICASSP | 3 |
| 2016 | Structure preserving single image super-resolutionabstractIn this paper, we present a novel structure preserving method for single image super-resolution to well construct edge structures and small detail structures. In our approach, the sharp edges are recovered via a novel edge preserving interpolation technique based on a well estimated gradient field and the edge preserving method, which incorporate the local and non-local structure information. The gradient of interpolated high-resolution(HR) image is then regarded as an edge preserving constraint to reconstruct the detail structures. Experimental results demonstrate that the new approach can reconstruct faithfully the HR images with sharp edges and texture structures, and annoying artifacts (blurring, jaggies, ringing, etc.) are greatly suppressed. It outperforms the state-of-the-art approaches, based on subjective and objective evaluations. Fan Yang 0053, Don Xie, Huizhu Jia, Rui Chen 0006, Guoqing Xiang, Wen Gao 0001 |
ICIP | 3 |
| 2016 | Smart query expansion scheme for CDVS based on illumination and key featuresabstractGiven a query image, retrieving images depicting the same object in a large scale database is becoming an urgent and challenging task. Recently, Compact Description for Visual Search (CDVS) is drafted by the ISO/IEC Moving Pictures Experts Group (MPEG) to support image retrieval applications, and it has been published as an international standard. Unfortunately, with regard to applications with hugely mutative illumination, perspective and noisy background, CDVS suffers from an inevitable performance loss. In this paper, firstly we introduce the query expansion to address performance loss caused by the scene complexity in CDVS. Secondly, a query expansion instance selection method based on illumination is proposed, which achieves better performance. Thirdly, we adopt a key feature matching score based weighted strategy in basic query expansion to improve retrieval performance. We evaluate our proposed methods on the Oxford (5K images) dataset and a reality traffic vehicle dataset (12K images), and the result shows that the proposed methods boost mean average precision (MAP) by 7% ∼ 10% in Oxford dataset and 7% ∼17% in vehicle dataset. Chuang Zhu, Huizhu Jia, Ling-Yu Duan, Jiawen Song, Wen Gao 0001 |
ICPR | 3 |
| 2016 | Adaptive perceptual preprocessing for video codingabstractThe quantization of block DCT coefficients is too coarse, and will result in much non-uniform artifacts, therefore, blocking and ring artifacts are usually visible in reconstructed video frames especially when the bitrate is not sufficient. In this paper, we present an adaptive perceptual preprocessing (APP) method to reduce the probability of these artifacts generated in video encoding. The APP algorithm employs a novel filter based on just noticeable distortion filter (JNDF) and adaptive bilateral filter (ABF), and they can be adaptively chosen by block characteristics and quantization parameters. Experimental results demonstrate our proposed algorithm can significantly improve the subjective quality of reconstructed image. Guoqing Xiang, Huizhu Jia, Jie Liu 0035, Binbin Cai, Yuan Li 0014 |
ISCAS | 2 |
| 2016 | Hardware-oriented adaptive multi-resolution motion estimation algorithm and its VLSI architectureabstractIn this paper, we propose a hardware architecture of an adaptive multi-resolution motion estimation algorithm (AMMEA) for high definition video encoder to reduce hardware cost. The texture-based search strategies are based on temporal stationarity and spatial homogeneity with Sobel edge operator. The proposed algorithm makes motion estimation more concise. We also propose Sobel edge operator hardware architecture. The four-pixel SAD unit which is the basic processing element (PE) in our proposed architecture is used for SAD calculation and Sobel edge operator computation. The hardware architecture achieves very high data utilization and data throughout. Using our proposed AMMEA with regular data flow, simulation results show that the proposed architecture can significantly reduce the hardware cost with a negligible PSNR loss of 0.03dB compared with the full-search. The design is implemented with SMIC 0.18μm CMOS technology and costs 950K gates count, and it supports the real-time encoding of 1080P@30fps with two reference frames under a clock frequency of 150MHz. Guoqing Xiang, Huizhu Jia, Jie Liu 0035, Yuan Li 0014 |
ISCAS | 2 |
| 2016 | A structure-preserving image restoration method with high-level ensemble constraintsabstractIn this paper, we present a new image restoration framework based on two high-level regularizations that can predict and preserve the better informative structures in the image. The sparse representation of a blurred image is first obtained to globally encode the salient structures by applying a group of coupled framelet filters. Then a physical meaning regularizer is derived to estimate the point spread function based on the frequency response characteristics of the image. Moreover, based on the operator of structure tensor, a novel nonlocal total variation as the regularizer is established to measure the image variation and non-local self-similarity. Finally, these two high-level regularizers are integrated into an objective function to constrain the ill-posedness. Compared with the state-of-the-art restoration methods, our algorithm can not only suppress strong noises effectively but also recover the sharp structures from the severe and complex blurred images. Rui Chen 0006, Huizhu Jia, Wen Gao 0001 |
VCIP | 2 |
| 2016 | Rate control for consistent video quality with inter-dependent distortion model for HEVCabstractConsistent video quality is important for video coding applications, which is also a popular optimization target for rate control. In this paper, a rate control scheme is proposed to reduce the fluctuation of video quality. First, we set up the optimization formulation of distortion and derive the distortion model by analyzing quadtree-based coding unit (CU) structure in high efficiency video coding (HEVC). Then the frame bit allocation algorithm is proposed by considering the inter-dependency. After the frame basic quantization parameter (QP) is obtained, the solution of optimization formulation finally regulates QP to maintain the consistent video quality. Experimental results show that the proposed rate control scheme can reduce the fluctuation of video quality up to 92.9% than benchmark, where the average reduction is 74.1%. Yuan Li 0014, Huizhu Jia, Tiejun Huang 0001 |
VCIP | 2 |
| 2016 | Hybrid Zero Block Detection for High Efficiency Video CodingabstractIn this paper we propose an efficient hybrid zero block early detection method for high efficiency video coding (HEVC). Our method detects both genuine zero blocks (GZBs) and pseudo zero blocks (PZBs). For GZB detection, we use a two sum of absolute difference bounds and a one sum of absolute transformed difference threshold to decrease the GZB detection complexity. A fast rate-distortion estimation algorithm for HEVC is proposed to improve the PZB detection rate. Experimental results on the HM platform show that the proposed method saves about 50% of the rate-distortion optimization (RTO) time, with negligible Bjøntegaard delta bit rate loss. Our method is faster than other state-of-the-art ZB detection methods for HEVC by 10%-30%. Hongfei Fan, Ronggang Wang, Lin Ding 0002, Huizhu Jia, Wen Gao 0001 |
IEEE Trans. Multim. | 5 |
| 2015 | A fast super-resolution method based on sparsity propertiesabstractSuper-resolution enhancement is a kind of promising approach to enhance the spatial resolution of images. To super-resolve a satisfying result, regularization term design and blur kernel estimation are two important aspects which need to be carefully considered. In this paper, we propose a robust regularized super-resolution reconstruction approach based on two sparsity properties to deal with these two aspects. Firstly, we design a sparse reweighted TV L1 prior to restrict the first derivative of the upsampled image. Then, noticing that only deblurring sparse high gradient areas can sharpen the super-resolution result, we design an over-deblurring control method to decrease the artifacts caused by inaccurate blur kernel estimation. We also design a fast optimization algorithm to solve our model. The experimental results show that the proposed approach achieves a remarkable performance both in visual quality and run time. Yuanchao Bai, Huizhu Jia, Rui Chen 0006, Ming Jiang 0001, Wen Gao 0001 |
VCIP | 2 |
| 2015 | A HVS-guided approach for real-time image interpolationabstractIn this paper, we propose a novel interpolation algorithm for adapting the human visual system (HVS) and applying in real-time image upscaling. The defined statistical features are first computed in a local window of the low-resolution (LR) counterpart. Then the most correlated neighbours of a missing pixel in high-resolution (HR) image are adaptively selected based on local structural analysis for the prominent edges and fine textures. Finally, the unknown pixel values are estimated through a designed directional clustering model (DCM) which incorporates HVS information into the weighted coefficients. The extensive experimental results show that the proposed image interpolation method can accurately reconstruct the structures of HR image in term of arbitrary magnification factors and effectively suppress the jaggy/ringing artifacts with low computation complexity. Rui Chen 0006, Huizhu Jia, Wen Gao 0001 |
VCIP | 2 |
| 2015 | An adaptive inter CU depth decision algorithm for HEVCabstractThe emerging High-Efficiency Video Coding (HEVC) standard has introduced a number of new coding tools, such as a quad-tree based coding unit (CU). The quadtree-structured coding unit achieves significant coding efficiency improvements compared to H264/AVC. However, the complexity of CU depth decision associated with Rate-Distortion (R-D) cost computation dramatically increased. In order to alleviate the computational burden in HEVC inter coding, a fast CU depth decision algorithm is proposed in this paper. Firstly, zero CU detection method for HEVC is proposed as early termination algorithm. Secondly, the CU depth pruning strategies are adaptively determined according to standard deviation of statistic spatiotemporal depth information. Finally, when the neighbors are not available or have a very weak correlation, edge gradient of current coding tree unit (CTU) is considered as main factor for CU depth pruning method. Experimental results demonstrate that, compared with the original HM16.0 implementation, the proposed algorithm achieves about 40.5% encoding time saving with ignorable coding performance degradation. Jie Liu 0035, Huizhu Jia, Guoqing Xiang, Xiaofeng Huang, Binbin Cai, Chuang Zhu, Don Xie |
VCIP | 2 |
| 2015 | Robust image/video super-resolution displayabstractThis paper describes a new method to reconstruct high-resolution video sequences from several observed low-resolution images based on an adaptive Mumford-shah model which is extended by using nonlocal information and low-rank representation. In our regularization framework, joint image restoration and motion estimation are first implemented and then detailed information can be recovered by incorporating new model as a prior term. Rui Chen 0006, Huizhu Jia, Wen Gao 0001 |
VRST | 2 |
| 2015 | Computation-constrained dynamic search range control for real-time video encoder
Xianghu Ji, Huizhu Jia, Jie Liu 0035, Wen Gao 0001 |
Signal Process. Image Commun. | 2 |
| 2014 | Low-delay window-based rate control scheme for video quality optimization in video encoderabstractThe consistent video quality and encoding latency due to buffering are two important aspects in designing rate control scheme for the application of real-time video coding system. To well balance these two contrary objectives, we firstly analyze the constraint of buffer latency and the definition of a “consistent” video quality. Then a window-based rate control scheme is proposed with one window for controlling the rate and latency, while the other window for optimizing video quality. By applying low complexity frame level ratedistortion model in the testing sequences, our proposed method shows excellent performance in balancing the encoder buffer latency and optimized video quality. Besides, this one-pass rate control scheme is highly practical for the real-time video coding application. Yuan Li 0014, Huizhu Jia, Chuang Zhu, Meng Li 0016, Wen Gao 0001 |
ICASSP | 2 |
| 2014 | Inter-dependent rate-distortion modeling for video coding and its application to rate controlabstractRate control scheme using independent rate-distortion (R-D) model at the minimum coding unit (macroblock) level has been widely discussed in the literature where R-D optimization is performed without consideration of the inter-dependencies of different coding units. In this paper, we extend these techniques to the more general situations - rate control for inter-dependent video coding units. The interdependent distortion-quantization (D-Q) model and rate-quantization (R-Q) model are formulated separately based on the analysis of the relationship between the spatial-domain residual and the transform-domain residual. Then a window-based rate control scheme with frame bit allocation and video quality optimization is proposed, which uses the approximated R-D model to reduce the computational complexity. Simulation results demonstrate that the proposed algorithm shows excellent peak signal-to-noise ratio (PSNR) performance under the bit rate constraint. This one-pass rate control scheme is highly practical for the realtime video coding application. Yuan Li 0014, Huizhu Jia, Pan Ma, Chuang Zhu, Wen Gao 0001 |
ICME | 2 |
| 2014 | A resolution-adaptive interpolation filter for video codecabstractThe fraction-pel interpolation filter varies in the video coding standards such as H.264/AVC, AVS and HEVC. Since fractional-pel motion compensation plays an important role in the video encoder, the interpolation of fractional-pel pixels can be refined and designed better to enhance the coding efficiency. In this paper, we firstly propose the generation algorithm of interpolation filter coefficients, and four different tap filters, namely 4tap, 6 tap, 8 tap and 10tap, are tested. A resolution-adaptive interpolation filter for different resolution videos is then introduced based on this algorithm to achieve the maximum bitrate saving. In the proposed scheme, 4 tap filter is applied for the UHD (2560×1600 and above) videos, 6 tap filter and 10 tap filter are performed in the videos whose resolution ranging from 720P (1280×720) to 1080P (1920×1080) and the videos with the resolution below 720P, respectively. When 4 tap filter and 6 tap filter are used in high-definition video, the coding efficiency can increase and the computational complexity will reduce greatly, which is actually beneficial to make hardware optimization more effectively especially SIMD (Single Instruction Multiple Data) and VLSI design. Experiments show that the average BD-rate gains on luma Y, chroma U and V are 1.4%, 0.7% and 0.7% for LP-Main configuration, when conducted in HEVC reference software HM11.0. The coding efficiency gains are significant for some video sequences and can reach up to 6.1%. Ronggang Wang, Yuan Li 0014, Chuang Zhu, Huizhu Jia, Wen Gao 0001 |
ISCAS | 5 |
| 2014 | Multi-level low-complexity coefficient discarding scheme for video encoderabstractRate-Distortion (R-D) optimization technique plays an important role in video coding. R-D sense discarding (thresholding) technique can make great improvement on the coding efficiency. This work first proposes a multi-level coefficient discarding scheme, which is composed of coefficient-level (CL), block-level and macroblock-level discarding. In CL, coefficient-level R-D cost function is formulated and then CL discarding scheme is developed. At last, an effective implementation method is proposed to reduce the complexity of the proposed scheme. The experimental results show that our proposed multi-level discarding scheme can improve the coding performance of video encoder by 0.15db in average. Chuang Zhu, Huizhu Jia, Jie Liu 0035, Xianghu Ji, Wen Gao 0001 |
ISCAS | 2 |
| 2014 | Fast algorithm of coding unit depth decision for HEVC intra codingabstractThe emerging high efficiency video coding standard (HEVC) achieves significantly better coding efficiency than all existing video coding standards. The quad tree structured coding unit (CU) is adopted in HEVC to improve the compression efficiency, but this causes a very high computational complexity because it exhausts all the combinations of the prediction unit (PU) and transform unit (TU) in every CU attempt. In order to alleviate the computational burden in HEVC intra coding, a fast CU depth decision algorithm is proposed in this paper. The CU texture complexity and the correlation between the current CU and neighbouring CUs are adaptively taken into consideration for the decision of the CU split and the CU depth search range. Experimental results show that the proposed scheme provides 39.3% encoder time savings on average compared to the default encoding scheme in HM-RExt-13.0 with only 0.6% BDBR penalty in coding performance. Xiaofeng Huang, Huizhu Jia, Kaijin Wei, Jie Liu 0035, Chuang Zhu, Zhengguang Lv, Don Xie |
VCIP | 2 |
| 2014 | Layer-based image completion by poisson surface reconstructionabstractImage completion has been widely used to repair damaged regions of a given digital image in a visually plausible way. However, it is difficult to infer appropriate information, meanwhile keep globally coherent just from the origin image when its critical parts are missing. To address this problem, we propose a novel layer-divided image completion scheme, which contains two major steps. First, we extract foregrounds of both target image and source image, and then we apply a guided Poisson surface reconstruction technique to complete the target foreground according to parameters obtained from optimal-matching calculation. Second, to fill the remaining damaged part, a related exemplar-based image completion algorithm is further devised. Several experiments and comparisons show the effectiveness and robustness of our proposed algorithm. Hengjin Liu, Huizhu Jia, Yuanchao Bai, Wen Gao 0001 |
VCIP | 2 |
| 2014 | Window-based rate control for video quality optimization with a novel INTER-dependent rate-distortion model
Yuan Li 0014, Huizhu Jia, Chuang Zhu, Mingyuan Yang, Wen Gao 0001 |
Signal Process. Image Commun. | 2 |
| 2013 | A high-throughput low-latency arithmetic encoder design for HDTVabstractIn this paper, we propose a high-throughput low-latency arithmetic encoder (AE) design suitable for high definition (HD) real-time applications employing advanced video coding standards such as H.264/AVC or AVS and using a macroblock (MB) level pipeline. First, in order to derive the performance requirement on the AE, a buffer model in connected with which it is designed is thoroughly analyzed. Then, using joint algorithm-architecture optimization and multi-bin processing techniques, we introduce a novel binary arithmetic coder (BAC) architecture with throughput of 2∼4 bins per cycle sufficient for real-time encoding. Furthermore, a hybrid context memory scheme is presented to meet the throughput requirement on the BAC. Simulation result shows that our design can support 1080p at 60 fps for AVS HDTV real-time coding with a bin rate up to 107K per MB line. Synthesized with the TSMC 0.13μm technology, the AE can run at 200MHz and costs 47.3K gates. By operating at 130MHz, the design is also verified in an AVS HD encoder on a Xilinx Virtex-6 FPGA prototype board for 1080p at 30 fps. Yuan Li 0014, Shanghang Zhang, Huizhu Jia, Wen Gao 0001 |
ISCAS | 3 |
| 2013 | On a Highly Efficient RDO-Based Mode Decision Pipeline Design for AVSabstractRate distortion optimization (RDO) is the best known mode decision method, while the high implementation complexity limits its applications and almost no real-time hardware encoder is truly full-featured RDO based. In this paper, first, a full-featured RDO-based mode decision (MD) algorithm is presented, which makes more modes enter RDO process. Second, the throughput of RDO-based MD pipeline is thoroughly analyzed and modeled. Third, a highly efficient adaptive block-level pipelining architecture of RDO-based MD for AVS video encoder is proposed which can achieve the highest throughput to alleviate the RDO burden. Our design is described in high-level Verilog/VHDL hardware description language and implemented under SMIC 0.18-$\mu$m CMOS technology with 232 K logic gates and 85 Kb SRAMs. The implementation results validate our architectural design and the proposed architecture can support real time processing of 1080P@30 fps. The coding efficiency of our adopted method far outperforms (0.57 dB PSNR gain in average) the traditional low-complexity MD (LCMD) methods and the throughput of our designed pipeline is increased by 11.3%, 19% and 17% for I, P and B frames, respectively, compared with the existed RDO-based architecture. Chuang Zhu, Huizhu Jia, Shanghang Zhang, Xiaofeng Huang, Wen Gao 0001 |
IEEE Trans. Multim. | 2 |
| 2012 | An Optimized Hardware Video Encoder for AVS with Level C+ Data Reuse Scheme for Motion EstimationabstractIn a hardware video encoder, Level C+ data reuse for motion estimation can reuse two-dimensional overlapped search window (SW) and thus is a good choice to trade off the memory bandwidth with the on-chip buffer size. However, the irregular zigzag coding order brings some other troubles to the encoder implementation. This paper mainly focuses on the special considerations for a Level C+ zigzag encoder. First we present a guideline about how to select the Level C+ zigzag HFmVn scan for the adopted encoder pipeline. Second, according to the guideline, zigzag HF5V3 coding order is applied into our Level C+ encoder in which a new function is added to alter zigzag bit-stream into standard raster order and exact motion vector predictor (MVP) can be used for most macro blocks (MBs) except some corner MBs to increase the coding performance. Third, zigzag-aware scheduling for prefetching the SW is proposed so that the pipeline will never be disturbed by this irregular coding order and can smoothly run MB by MB. In addition, balancing the bandwidth into each MB processing period can improve the bandwidth utilization. With these techniques, a real-time high-definition (HD) 1080P AVS encoder is successfully implemented on FPGA verification board with search range [-128, 128]×[-96, 96] and two reference frames at an operating frequency of 160 MHz. Kaijin Wei, Rongwei Zhou, Shanghang Zhang, Huizhu Jia, Don Xie, Wen Gao 0001 |
ICME | 4 |
| 2012 | A high speed and efficient architecture of VLD for AVS HD video decoderabstractIn this paper, we present a high speed and efficient architecture of Variable Length Decoder for AVS video standard targeted for all-hardware implementation. Besides the regular operations of decoding Fixed Length Code, unsigned or signed k-th Exp-Golomb Code and 2D Variable Length Code, the proposed design provides other functions such as de-stuffing or pre-fetching the Bitstream. It can perform decoding syntax elements of sequence, frame, slice, and macro block. The complete architecture has been described in Verilog HDL, simulated with Modelsim SE 6.3c simulator, implemented using FPGA of Xinlinx Vertex 5 VLX330. Without any strict constraint, the design can achieve a working frequency at 190 MHz after synthesis with Synplify_pro 9.4, and the critical path is less than 6.5ns after place & route. The throughput of the design is 1 codeword per clock. In all, the architecture fully meets the demands of AVS HD decoder. It can support real-time decoding for 1080P @ 30 frame/s or 1080i @ 60 field/s videos. Inevitably, the cost for such a high speed design is consuming more hardware resources. Report of Place & Route shows about 9.1K LUTs (4% of the total LUTs in FPGA chip) are consumed by our design. Although the VLD architecture was originally designed for AVS video standard, the idea of the design can be easily adapted to other video standards. Zhenqiang Yang, Huizhu Jia, Don Xie |
PCS | 3 |
| 2012 | A flexible and high-performance hardware video encoder architectureabstractThis paper presents a new video encoder architecture for H.264 and AVS, which adopts a novel macroblock (MB) encoding order. As a replacement of Level C+ zigzag coding order, the so-called Level C+ slash scan coding order with NOP insertion is used as MB scheduling to remove MB-level data dependency of the pipeline so that the left MB's coded results such as motion vector (MV) and reconstructed pixels can be obtained early in motion estimation (ME) stages. As a result, by sharing the reconstruction (REC) loop, sequential intra prediction (INTRA) can be split into multiple pipeline stages to explore more block-level parallelization and rate distortion optimization (RDO) based mode decision is apt to implement. The exact MV predictors (MVP) obtained in motion estimation can not only improve coding performance but also make pre-skip ME algorithm able to be applied into this architecture for low power applications. Since the proposed scheme is attributed to Level C+ data reuse, the bandwidth is decreased greatly. A real-time high-definition (HD) 1080P AVS encoder implementation on FPGA verification board with search range [-128, 128]×[-96, 96] and two reference frames at an operating frequency of 160 MHz validates the efficiency of proposed architecture. Kaijin Wei, Shanghang Zhang, Huizhu Jia, Don Xie, Wen Gao 0001 |
PCS | 3 |
| 2012 | A comparison of fractional-pel interpolation filters in HEVC and H.264/AVCabstractThe fractional-pel interpolation filter adopted in H.264/AVC improves motion compensation greatly. Recently, a new DCT-based fractional-pel interpolation filter is adopted in the oncoming standard HEVC. We are interested in the differences between these two types of fractional-pel interpolation filters. In this paper we describe the derivations of fractional-pel interpolation filters in HEVC and H.264/AVC in detail, and compare them on properties of frequency responses. We find that the half-pel interpolation filters in HEVC and H.264/AVC are very similar, but the low-pass properties of quarter-pel interpolation filters in HEVC are much better than those in H.264/AVC. Experimental results validate this phenomenon, the fractional-pel interpolation in H.264/AVC tends to increase BD-rates by more than 10% compared with that in HEVC, and this performance loss mainly comes from quarter-pel interpolation filters. On the other hand, the complexity of fractional-pel interpolation filtering in HEVC is greatly increased than that in H.264/AVC. Ronggang Wang, Huizhu Jia, Wen Gao 0001 |
VCIP | 4 |
| 2012 | An efficient foreground-based surveillance video coding scheme in low bit-rate compressionabstractMany works have been done in the area of surveillance video compression, while problems still exist. The block-based schemes have blocking artifacts in the edge of foreground, while the object-based coding schemes have excessive bit consumption for coding the object shape. A novel foreground-based (FG-based) coding scheme is presented in this paper to solve these two problems and can gain better video quality at low bit-rate. The improvement comes from: 1) obtaining a foreground frame (FG-frame) by segmentation, in which proper constant value 128 is adopted to represent the luminance and chrominance value of background pixel and thus the residue error is reduced; 2) FG-based motion estimation (ME) and motion compensation (MC), which are more accurate for the foreground prediction and reduce the residue error of edge block in the foreground; 3) a new coding mode (BG-mode) is designed to better code the background when it is falsely segmented as foreground in FG-frames; 4) FG-based rate distortion optimized (RDO) mode decision (MD) is proposed to emphasize the foreground by calculating the distortion in the foreground domain; 5) avoiding shape coding by recovering the shape mask from the reconstructed foreground (REC-FG) frame and the constant background value 128. Our scheme is implemented with AVS encoder platform and the experiment results show the efficiency of the proposed scheme. Shanghang Zhang, Kaijin Wei, Huizhu Jia, Wen Gao 0001 |
VCIP | 3 |
| 2011 | A hardware-efficient architecture for multi-resolution motion estimation using fully reconfigurable processing element arrayabstractInteger motion estimation (IME) for block-based video coding presents a significant challenge in external memory bandwidth, data latency, and circuit area with the increase of coding complexity and video resolution. To conquer these problems, this paper proposes a hardware-efficient VLSI architecture for multi-resolution motion estimation algorithm (MMEA) based on fully reconfigurable processing element (PE) array. On-chip storage and PE array are carefully designed to support parallel computation and hardware resource sharing. In addition, low data latency is obtained by arranging internal logics in parallel according to the data dependency. As a result, our design can support real time processing of 1080P@30fps with 2 reference frames and a search range of 256×192 and it is implemented under SMIC 0.18-µm CMOS technology with 920K logic gates and 192 KB SRAMs. Compared with previous work, our design can achieve the best performance-price rate benefiting from the proposed re-configurable PE array. Xianghu Ji, Chuang Zhu, Huizhu Jia, Hai Bing Yin |
ICME | 3 |
| 2011 | A highly efficient pipeline architecture of RDO-based mode decision design for AVS HD video encoderabstractLike H.264, AVS video coding standard also uses macroblock (MB) based motion compensation (MC) and mode decision (MD). Rate distortion optimization (RDO) is the best known mode decision method, but with a high computational complexity that limits its applications. In our paper, firstly an MD algorithm based on RDO is given, which makes more mode candidates enter into RDO mode decision with little hardware resource increment. We further analyze the pipeline structure in detail, and implement a block-level 5-stage hardware pipeline. It can support the real time RDO mode decision processing of 1080P@30fps, and the coding efficiency is about 0.5db higher than the traditional SAD method. Our design is described in high-level Verilog/VHDL hardware description language and implemented under SMIC 0.18-µm CMOS technology with 215K logic gates and 80 KB SRAMs. Chuang Zhu, Yuan Li 0014, Huizhu Jia, Hai Bing Yin |
ICME | 3 |
| 2011 | Adaptive integer-precision Lagrange multiplier selection for high performance AVS video codingabstractIn AVS and H.264/AVC, Lagrangian Rate distortion (RD) optimization techniques are widely adopted for coding mode selection and displacement vector estimation. The optimal Lagrange multipliers in these two cases are both floating-point values. If RD optimized video encoder is implemented on computation-constrained fixed-point platform such as FPGA and ASIC, fixed-point Lagrange multiplier selection is an important problem to trade-off the RD performance and computation complexity. This work focuses on fixed-point Lagrange multiplier selection for RD mode decision. Adaptive scaling matrix is used to trade-off complexity and RD performance. Also, intensive simulation results and analysis on precision, hardware cost, and RD performance are given. The proposed approach is also well-suited for RD optimized motion estimation for computation-constrained video coding. Hai Bing Yin, Bingqian Zhou, Chuang Zhu, Huizhu Jia |
VCIP | 4 |
| 2010 | Efficient macroblock pipeline structure in high definition AVS video encoder VLSI architectureabstractIn traditional four-stage pipeline structures for H.264 video encoder hardware implementation, rate distortion optimization (RDO) based mode decision was turned off, and dual-port or ping-pang on-chip search window SRAM was used to achieve data reuse between the integer and fractional pixel motion estimation. To support RDO based mode decision for efficient high definition AVS video coding implementation, we propose an improved four-stage MB pipeline structure. Also on-chip buffer structure is optimized to achieve the balance between circuit consumption and coding performance. The Jizhun profile AVS video encoder is successfully mapped into hardware implementation with the proposed pipeline structure with small performance degradation. Hai Bing Yin, Honggang Qi, Huizhu Jia, Don Xie, Wen Gao 0001 |
ISCAS | 3 |
| 2010 | Algorithm analysis and architecture design for rate distortion optimized mode decision in high definition AVS video encoder
Hai Bing Yin, Honggang Qi, Huizhu Jia, Chuang Zhu |
Signal Process. Image Commun. | 3 |
| 2010 | A Hardware-Efficient Multi-Resolution Block Matching Algorithm and its VLSI Architecture for High Definition MPEG-Like Video EncodersabstractHigh throughput, heavy bandwidth requirement, huge on-chip memory consumption, and complex data flow control are major challenges in high definition integer motion estimation hardware implementation. This paper proposes an efficient very large scale integration architecture for integer multi-resolution motion estimation based on optimized algorithm. There are three major contributions in this paper. First, this paper proposes a hardware friendly multi-resolution motion estimation algorithm well-suited for high definition video encoder. Second, parallel processing element (PE) array structure is proposed to implement three-level hierarchical motion estimation, only 256PEs are enough for one reference frame real-time high definition motion estimation by efficient PE reuse. Third, efficient on-chip reference pixel buffer sharing mechanism between integer and fractional motion estimation is proposed with almost 50% SRAM saving and memory bandwidth reduction. The proposed multi-resolution motion estimation algorithm reached a good balance between complexity and performance with rate distortion optimized variable block size motion estimation support. Also, we have achieved moderate logic circuit and on-chip SRAM consumption. The proposed architecture is well-suited for all MPEG-like video coding standards such as H.264, audio video coding standard, and VC-1. Hai Bing Yin, Huizhu Jia, Honggang Qi, Xianghu Ji, Wen Gao 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |