EDBT 2026 Demo / reviewers in the wild / expert
Yuan Li 0014
dblp:86/6196-14
· DBLP profile ↗
34ranked-venue papers
5as first author
12since 2021 · last 2024
0000-0002-8479-3049ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 4 first-author · 8 since 2021Artificial intelligence and machine learning · 6 · 5 since 2021Systems, architecture and hardware · 6 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2024 | DMR: Decomposed Multi-Modality Representations for Frames and Events Fusion in Visual Reinforcement LearningabstractWe explore visual reinforcement learning (RL) using two complementary visual modalities: frame-based RGB cam-era and event-based Dynamic Vision Sensor (DVS). Ex-isting multi-modality visual RL methods often encounter challenges in effectively extracting task-relevant information from multiple modalities while suppressing the in-creased noise, only using indirect reward signals instead of pixel-level supervision. To tackle this, we propose a Decomposed Multi-Modality Representation (DMR) framework for visual RL. It explicitly decomposes the inputs into three distinct components: combined task-relevant features (co-features), RGB-specific noise, and DVS-specific noise. The co-features represent the full information from both modalities that is relevant to the RL task; the two noise components, each constrained by a data reconstruction loss to avoid information leak, are contrasted with the co-features to maximize their difference. Extensive experiments demonstrate that, by explicitly separating the different types of information, our approach achieves substan-tially improved policy performance compared to state-of-the-art approaches. Haoran Xu 0004, Peixi Peng, Guang Tan, Yuan Li 0014, Xinhai Xu, Yonghong Tian 0001 |
CVPR | 4 |
| 2024 | Enhanced Blind Watermarking Against Black-Box Noise: Leveraging CIN FrameworkabstractBlind watermarking is a technology for image copyright protection and digital fingerprinting. However, the introduction of non-differentiable noise makes it challenging to be trained end-to-end for black-box scenes. The phased training technique is used for coping with black-box noise, but it limits the performance since the encoder and decoder cannot be end-to-end optimized. This work proposes a blind watermarking framework CIN+ based on the CIN to address black-box noise. Combining the structural characteristics of an Invertible Neural Network (INN) with the two-stage strategy allows joint updates of the encoder and decoder when encountering non-differentiable noise. We utilize Noise and Gradient Propagation Gate (NGPG) modules to perform a batch-level optimization akin to the two-stage approach, allowing encoder parameters to remain unlocked, thereby enhancing the model’s ability to resist blackbox attacks. Additionally, a Pre-Extraction Module (PEM) is introduced to simplify the complexity and usability of CIN. Our experimental results reveal that CIN+ achieves a new state-of-the-art performance. Rui Ma 0032, Mengxi Guo, Peidong Jia, Chenxuan Li 0003, Yuan Li 0014, Shanghang Zhang |
ICME | 6 |
| 2024 | VLUReID: Exploiting Vision-Language Knowledge for Unsupervised Person Re-IdentificationabstractThe superior performances of pre-trained vision-language models on various downstream tasks demonstrate the effectiveness of integrating cross-modal vision-language knowledge into visual tasks. However, this knowledge is hardly used for visual-based person re-identification (re-ID) because the datasets lack textual descriptions. Existing efforts require manual annotations for training, which can be time-consuming. We propose VLUReID, a framework that improves visual-based person re-ID using vision-language knowledge without requiring manual annotations from datasets. Specifically, the Vision-to-Text Association (VTA) module uses designed textual prompts to prompt the vision-language model in generating pseudo-semantic labels for visual inputs. Subsequently, within the Dual-Branch Asymmetric Training (DBAT) module, we propose an asymmetric training strategy to extract cross-modal knowledge from pseudo-semantic labels and integrate it into the person re-ID model. The experimental results on two widely-used benchmarks for unsupervised video-based person re-ID demonstrate the effectiveness of our framework. Ray Zhang 0002, Fan Yang 0053, Yuan Li 0014, Huizhu Jia, Shanghang Zhang |
ICME | 4 |
| 2023 | Highly Efficient SNNs for High-speed Object Detection
Nemin Qiu, Yuan Li 0014, Chuang Zhu |
BMVC | 3 |
| 2023 | WUDA: Unsupervised Domain Adaptation Based on Weak Source Domain LabelsabstractUnsupervised domain adaptation (UDA) for semantic segmentation addresses the cross-domain problem with fine source domain labels. However, the acquisition of semantic labels is often time-consuming, many scenarios only have weak labels (e.g. bounding boxes). When weak supervision and cross-domain problems coexist, this paper defines a new task: unsupervised domain adaptation based on weak source domain labels (WUDA). To explore solutions for WUDA, this paper proposes two intuitive frameworks and conducts comparative experiments. We observe that the two frameworks behave differently when the datasets change. Therefore, we construct datasets with a wide range of domain shifts and conduct extended experiments to analyze the impact of domain shift changes on the two frameworks. In addition, to measure domain shift, we apply the metric representation shift to urban landscape image segmentation for the first time. The source code and constructed datasets can be obtained from this link: https://github.com/bupt-ai-cz/WUDA. Chuang Zhu, Yuan Li 0014, Wenqi Tang |
ICASSP | 3 |
| 2022 | Enhancing and Dissecting Crowd Counting by Synthetic DataabstractIn this article, we propose a simulated crowd counting dataset CrowdX, which has a large scale, accurate labeling, parameterized realization, and high fidelity. The experimental results of using this dataset as data enhancement show that the performance of the proposed streamlined and efficient benchmark network ESA-Net can be improved by 8.4%. The other two classic heterogeneous architectures MCNN and CSRNet pre-trained on CrowdX also show significant performance improvements. Considering many influencing factors determine performance, such as background, camera angle, human density, and resolution. Although these factors are important, there is still a lack of research on how they affect crowd counting. Thanks to the CrowdX dataset with rich annotation information, we conduct a large number of data-driven comparative experiments to analyze these factors. Our research provides a reference for a deeper understanding of the crowd counting problem and puts forward some useful suggestions in the actual deployment of the algorithm. Chengyang Li 0001, Yuheng Lu, Yuan Li 0014, Huizhu Jia |
ICASSP | 5 |
| 2022 | Efficient Algorithm and Hardware Architecture for Rate Estimation in Mode Decision of AVS3abstractTowards enabling advanced video coding for emerging ap-plications, the AVS3 standard has been developed recently, achieving twice the coding efficiency of the AVS2 stan-dard through complex coding tools including advanced rate-distortion optimization (RDO) to select the best mode. The bit-rates are produced with the Advanced-Entropy-Coding (AEC) in the RDO process of AVS3. However, AEC dom-inates the time complexity of RDO and among the steps, con-text updating and interval subdivision are performed recur-sively, which is not conducive to real-time application, espe-cially for the hardware implementation. Thus this paper pro-poses an adaptive rate estimation algorithm with a piece-wise linear function that is very friendly to hardware implemen-tation to accelerate the rate estimation process in the RDO for AVS3 practical applications. The proposed architecture can meet the requirement of 4K@120fps ultra-high-definition videos at 200 MHz, whereas the BD-Rate increases only by 0.67% under the All-Intra (AI) configuration. Yunyao Yan, Guoqing Xiang, Huizhu Jia, Yuan Li 0014, Peng Zhang 0007, Jie Chen 0001 |
ICME | 5 |
| 2022 | Towards Blind Watermarking: Combining Invertible and Non-invertible MechanismsabstractBlind watermarking provides powerful evidence for copyright protection, image authentication, and tampering identification.However, it remains a challenge to design a watermarking model with high imperceptibility and robustness against strong noise attacks. To resolve this issue, we present a framework Combining the Invertible and Non-invertible (CIN) mechanisms. The CIN is composed of the invertible part to achieve high imperceptibility and the non-invertible part to strengthen the robustness against strong noise attacks. For the invertible part, we develop a diffusion and extraction module (DEM) and a fusion and split module (FSM) to embed and extract watermarks symmetrically in an invertible way. For the non-invertible part, we introduce a non-invertible attention-based module (NIAM) and the noise-specific selection module (NSM) to solve the asymmetric extraction under a strong noise attack. Extensive experiments demonstrate that our framework outperforms the current state-of-the-art methods of imperceptibility and robustness significantly. Our framework can achieve an average of 99.99% accuracy and 67.66 dB PSNR under noise-free conditions, while 96.64% and 39.28 dB combined strong noise attacks. The code will be available in https://github.com/RM1110/CIN. Rui Ma 0032, Mengxi Guo, Fan Yang 0053, Yuan Li 0014, Huizhu Jia |
ACM Multimedia | 5 |
| 2022 | On the Correlation Among Edge, Pose and ParsingabstractSemantic parsing, edge detection, and pose estimation of human are three closely-related tasks. They present human characteristics from three complementary aspects. Compared to learning them individually, solving these tasks jointly can explore the interaction of their contextual cues. However, prior works usually study the fusion of two of them, e.g., parsing and pose, parsing and edge. In this paper, we explore how pixel-level semantics, human boundaries and joint locations can be effectively learned in a unified model. Specifically, we propose an end-to-end trainable Human Task Correlation Machine (HTCorrM) to implement the three tasks. It is asymmetric in that it supports a main task using the other two as auxiliary tasks. We also introduce a Heterogeneous Non-Local module (HNL) to discover the correlations of the three heterogeneous domains. HNL fully explores the global dependency among tasks between any two positions in the feature map. Experimental results on human parsing, pose estimation and body edge detection demonstrate that HTCorrM achieves competitive performance. We show that when designated as the main task, the accuracy of each of the three tasks is improved. Importantly, comparative studies confirm the advantages of our proposed feature correlation strategy over feature concatenation or post processing. Ziwei Zhang 0003, Chi Su, Liang Zheng 0001, Yuan Li 0014 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2021 | Hardware-Friendly Coding Unit Decision Scheme for HEVCabstractQuad-tree based coding unit partition in High-Efficiency Video Coding (HEVC) achieved significant coding efficiency improvements, but also brought increasing computational complexity. Especially, design challenges like data dependence, large area cost, and imbalance of processing time of each coding tree unit (CTU), make it hard to achieve a real-time structure for real-time hardware encoder for all CTU sizes. To solve these problems, we proposed a hardware-friendly fast CU decision scheme with multi-stage algorithms for HEVC hardware encoder, aiming at the most complex modules: IME (Integer Motion Estimation), FME/IP (Fractional Motion Estimation/ Intra Prediction), and MD (Mode Decision). Firstly, in IME stage, a zero block detection method based fast CU and PU decision algorithm was presented. Secondly, we presented an estimated RDO (Rate-Distortion Optimization) based algorithm in the Hadamard domain for the early CU decision further in FME/IP stage. Finally, under the condition of hardware computing time limitation of several CU sizes, we proposed a computation time constraint CU fast decision algorithm for MD stage. Experiments demonstrated that, compared with the original HM13.0 implementation, the proposed scheme achieved about 53.9% encoding time saving with merely 2.3% coding performance degradation. What's more, significant area cost and data dependency have been alleviated, which will be more hardware-friendly for HEVC encoder design. Ju Huang, Xiaofeng Huang, Guoqing Xiang, Yuan Li 0014, Huizhu Jia |
ISCAS | 4 |
| 2021 | Deep Human-Interaction and Association by Graph-Based Learning for Multiple Object Tracking in the Wild
Cong Ma 0006, Fan Yang 0053, Yuan Li 0014, Huizhu Jia, Wen Gao 0001 |
Int. J. Comput. Vis. | 3 |
| 2021 | Deep Trajectory Post-Processing and Position Projection for Single & Multiple Camera Multiple Object Tracking
Cong Ma 0006, Fan Yang 0053, Yuan Li 0014, Huizhu Jia, Wen Gao 0001 |
Int. J. Comput. Vis. | 3 |
| 2020 | BBA-NET: A Bi-Branch Attention Network For Crowd CountingabstractIn the field of crowd counting, the current mainstream CNNbased regression methods simply extract the density information of pedestrians without finding the position of each person. This makes the output of the network often found to contain incorrect responses, which may erroneously estimate the total number and not conducive to the interpretation of the algorithm. To this end, we propose a Bi-Branch Attention Network (BBA-NET) for crowd counting, which has three innovation points. i) A two-branch architecture is used to estimate the density information and location information separately. ii) Attention mechanism is used to facilitate feature extraction, which can reduce false responses. iii) A new density map generation method combining geometric adaptation and Voronoi split is introduced. Our method can integrate the pedestrian’s head and body information to enhance the feature expression ability of the density map. Extensive experiments performed on two public datasets show that our method achieves a lower crowd counting error compared to other state-of-the-art methods. Chengyang Li 0001, Fan Yang 0053, Cong Ma 0006, Yuan Li 0014, Huizhu Jia |
ICASSP | 6 |
| 2020 | Fusion Target Attention Mask Generation Network For Video SegmentationabstractVideo segmentation aims to segment target objects in a video sequence, which remains a challenge due to the motion and deformation of objects. In this paper, we propose a novel attention-driven hybrid encoder-decoder network that generates object segmentation by fully leveraging spatial and temporal information. Firstly, a multi-branch network is designed to learn feature representation from object appearance, location and motion. Secondly, a target attention module is proposed to further exploit context information from learned representation. In addition, a novel edge loss is designed which constraints the model to generate salient edge features and accurate segmentation. The proposed model has been evaluated over two widely used public benchmarks, and experiments demonstrate its superior robustness and effectiveness as compared with the state of the arts. Yunyi Li, Fangping Chen, Fan Yang 0053, Yuan Li 0014, Huizhu Jia |
ICIP | 4 |
| 2020 | Spatiotemporal Perception Aware Quantization Algorithm For Video CodingabstractAdaptive quantization (AQ) proves to be an effective tool to improve coding performance. In this paper, we propose an adaptively spatiotemporal perception aware quantization algorithm to increase subjective coding performance. First, the perceptual complexity models are conducted with spatial and temporal characteristics to measure the spatiotemporally perceptual redundancies, respectively. With the help of the models, the adaptively spatial and temporal quantization parameter (QP) offsets are then calculated for each coding tree unit (CTU), respectively. Finally, the perceptually optimal Lagrange multiplier of each CTU is determined with the spatial-temporal QP offset. Experimental results show that the proposed algorithm reduces 8.6% BD-Rate with SSIM (Structural Similarity Index Metric) in average over the AVS2 (the second generation of Audio Video Coding Standard) reference software RD17.0 in Low-Delay P (LDP) configurations. The subjective assessment proves the proposed algorithm can significantly reduce the bit rates with the same subjective quality. Yunyao Yan, Guoqing Xiang, Yuan Li 0014, Wei Yan 0020, Yungang Bao |
ICME | 3 |
| 2020 | High-Level Task-Driven Single Image Deraining: Segmentation in Rainy Days
Mengxi Guo, Mingtao Chen, Cong Ma 0006, Yuan Li 0014 |
ICONIP (1) | 4 |
| 2020 | Optical Flow-Guided Mask Generation Network for Video SegmentationabstractThe purpose of video segmentation is to segment foreground objects from a video sequence. In this paper, we propose a CNN based method for the semi-supervised video object segmentation, where a hybrid encoder-decoder network is designed to generate pixel-wise foreground object segmentation in use of both spatial and temporal information. In order to minimize cumulative error of the network as much as possible, we develop a two-stage training scheme: alternate training and back-propagation-through-time training. Then the performances of our method and other state-of-the-art ones are compared on two annotated video segmentation databases. Furthermore, we also run an extensive ablation study to test the effects of different components from our method. Yunyi Li, Fangping Chen, Fan Yang 0053, Cong Ma 0006, Yuan Li 0014, Huizhu Jia |
ISCAS | 5 |
| 2020 | A Novel Quality Enhanced Low Complexity Rate Control Algorithm for HEVCabstractRate control (RC) is a key technology in video coding which is mainly responsible for adapting the compressed video quality as much as possible under limited bandwidth. Typical RC methods consist of initial quantization parameters of I-frames decision and control parameters updating procedures. However, on the one hand, I-frames quantization parameter (QP) is decided by the sum of absolute transformed difference (SATD) computation, which is time consuming for low delay applications. On the other hand, the RC parameters are only updated ignoring the distortion characteristics for inter frame, which cannot achieve optimal rate distortion (RD) performance. Therefore, a novel RC method which considers distortion characteristics for model parameters updating for quality enhancement is proposed in this paper, including a low complexity based I-frame QP decision strategy of low delay applications. For which estimated distortion characteristics and the previous QPs are employed, respectively. Through the experimental results, with more accurate RC precision and negligible I-frame QP computation time, the proposed rate control scheme improves the performance gain of 2.6% bitrate savings for the whole test sequences on average in HM 16.9. ChungWen Ku, Guoqing Xiang, Huizhu Jia, Yuehua Cui, Yuan Li 0014 |
VCIP | 6 |
| 2020 | An adaptive spatio-temporal perception aware quantization algorithm for AVS2
Yunyao Yan, Guoqing Xiang, Yuan Li 0014, Huizhu Jia |
J. Vis. Commun. Image Represent. | 3 |
| 2019 | Deep Association: End-to-end Graph-Based Learning for Multiple Object Tracking with Conv-Graph Neural NetworkabstractMultiple Object Tracking (MOT) has a wide range of applications in surveillance retrieval and autonomous driving. The majority of existing methods focus on extracting features by deep learning and hand-crafted optimizing bipartite graph or network flow. In this paper, we proposed an efficient end-to-end model, Deep Association Network (DAN), to learn the graph-based training data, which are constructed by spatial-temporal interaction of objects. DAN combines Convolutional Neural Network (CNN), Motion Encoder (ME) and Graph Neural Network (GNN). The CNNs and Motion Encoders extract appearance features from bounding box images and motion features from positions respectively, and then the GNN optimizes graph structure to associate the same object among frames together. In addition, we presented a novel end-to-end training strategy for Deep Association Network. Our experimental results demonstrate the effectiveness of DAN up to the state-of-the-art methods without extra-dataset on MOT16 and DukeMTMCT. Cong Ma 0006, Yuan Li 0014, Fan Yang 0053, Ziwei Zhang 0003, Yueqing Zhuang, Huizhu Jia |
ICMR | 2 |
| 2019 | Bit Allocation based on Visual Saliency in HEVCabstractAs one of the important part in the HEVC reference software, R-lambda model adopts mean absolute difference (MAD) for the coding unit tree (CTU) level bit allocation. However, this optimum method may neglect some important characteristics of human visual system (HVS). In this paper, we propose a novel bit allocation algorithm to process some salient visual information priority. Firstly, an improved video saliency detection algorithm is proposed, which induces temporal correlation into a 2D visual attention model. Secondly, the visual saliency based CTU level bit allocation algorithm is presented by allocating bits for CTUs with their saliency weights. What's more, with considerations of the temporal quality consistence among Saliency Areas (SAs), a window based weight smoothing model is proposed to achieve better subjective quality. Finally, several experiments are performed on the HEVC reference software, HM16.9, under the low delay P configuration, and the experimental results show that the average BD-Rate of the entire test sequences and of the SAs reduce 1.7% and 6.2%, respectively. The proposed algorithm can also improve subjective quality remarkably. ChungWen Ku, Guoqing Xiang, Wei Yan 0020, Yuan Li 0014 |
VCIP | 5 |
| 2019 | Robust estimation for image noise based on eigenvalue distributions of large sample covariance matrices
Rui Chen 0006, Changshui Yang, Yuan Li 0014, Tiejun Huang 0001 |
J. Vis. Commun. Image Represent. | 4 |
| 2018 | Dense Relation Network: Learning Consistent and Context-Aware Representation for Semantic Image SegmentationabstractSemantic image segmentation, which aims at assigning pixel-wise category, is one of challenging image understanding problems. Global context plays an important role on local pixel-wise category assignment. To make the best of global context, in this paper, we propose dense relation network (DRN) and context-restricted loss (CRL) to aggregate global and local information. DRN uses Recurrent Neural Network (RNN) with different skip lengths in spatial directions to get context-aware representations while CRL helps aggregate them to learn consistency. Compared with previous methods, our proposed method takes full advantage of hierarchical contextual representations to produce high-quality results. Extensive experiments demonstrate that our method achieves significant state-of-the-art performances on Cityscapes and Pascal Context benchmarks, with mean-IoU of 82.8% and 49.0% respectively. Yueqing Zhuang, Fan Yang 0053, Cong Ma 0006, Ziwei Zhang 0003, Yuan Li 0014, Huizhu Jia, Wen Gao 0001 |
ICIP | 6 |
| 2018 | A novel adaptive quantization method for video coding
Guoqing Xiang, Huizhu Jia, Mingyuan Yang, Yuan Li 0014 |
Multim. Tools Appl. | 4 |
| 2017 | LLCNN: A convolutional neural network for low-light image enhancementabstractIn this paper, we propose a CNN based method to perform low-light image enhancement. We design a special module to utilize multiscale feature maps, which can avoid gradient vanishing problem as well. In order to preserve image textures as much as possible, we use SSIM loss to train our model. The contrast of low-light images can be adaptively enhanced using our method. Results demonstrate that our CNN based method outperforms other contrast enhancement methods. Chuang Zhu, Guoqing Xiang, Yuan Li 0014, Huizhu Jia |
VCIP | 4 |
| 2016 | Adaptive perceptual preprocessing for video codingabstractThe quantization of block DCT coefficients is too coarse, and will result in much non-uniform artifacts, therefore, blocking and ring artifacts are usually visible in reconstructed video frames especially when the bitrate is not sufficient. In this paper, we present an adaptive perceptual preprocessing (APP) method to reduce the probability of these artifacts generated in video encoding. The APP algorithm employs a novel filter based on just noticeable distortion filter (JNDF) and adaptive bilateral filter (ABF), and they can be adaptively chosen by block characteristics and quantization parameters. Experimental results demonstrate our proposed algorithm can significantly improve the subjective quality of reconstructed image. Guoqing Xiang, Huizhu Jia, Jie Liu 0035, Binbin Cai, Yuan Li 0014 |
ISCAS | 5 |
| 2016 | Hardware-oriented adaptive multi-resolution motion estimation algorithm and its VLSI architectureabstractIn this paper, we propose a hardware architecture of an adaptive multi-resolution motion estimation algorithm (AMMEA) for high definition video encoder to reduce hardware cost. The texture-based search strategies are based on temporal stationarity and spatial homogeneity with Sobel edge operator. The proposed algorithm makes motion estimation more concise. We also propose Sobel edge operator hardware architecture. The four-pixel SAD unit which is the basic processing element (PE) in our proposed architecture is used for SAD calculation and Sobel edge operator computation. The hardware architecture achieves very high data utilization and data throughout. Using our proposed AMMEA with regular data flow, simulation results show that the proposed architecture can significantly reduce the hardware cost with a negligible PSNR loss of 0.03dB compared with the full-search. The design is implemented with SMIC 0.18μm CMOS technology and costs 950K gates count, and it supports the real-time encoding of 1080P@30fps with two reference frames under a clock frequency of 150MHz. Guoqing Xiang, Huizhu Jia, Jie Liu 0035, Yuan Li 0014 |
ISCAS | 4 |
| 2016 | Rate control for consistent video quality with inter-dependent distortion model for HEVCabstractConsistent video quality is important for video coding applications, which is also a popular optimization target for rate control. In this paper, a rate control scheme is proposed to reduce the fluctuation of video quality. First, we set up the optimization formulation of distortion and derive the distortion model by analyzing quadtree-based coding unit (CU) structure in high efficiency video coding (HEVC). Then the frame bit allocation algorithm is proposed by considering the inter-dependency. After the frame basic quantization parameter (QP) is obtained, the solution of optimization formulation finally regulates QP to maintain the consistent video quality. Experimental results show that the proposed rate control scheme can reduce the fluctuation of video quality up to 92.9% than benchmark, where the average reduction is 74.1%. Yuan Li 0014, Huizhu Jia, Tiejun Huang 0001 |
VCIP | 1 |
| 2014 | Low-delay window-based rate control scheme for video quality optimization in video encoderabstractThe consistent video quality and encoding latency due to buffering are two important aspects in designing rate control scheme for the application of real-time video coding system. To well balance these two contrary objectives, we firstly analyze the constraint of buffer latency and the definition of a “consistent” video quality. Then a window-based rate control scheme is proposed with one window for controlling the rate and latency, while the other window for optimizing video quality. By applying low complexity frame level ratedistortion model in the testing sequences, our proposed method shows excellent performance in balancing the encoder buffer latency and optimized video quality. Besides, this one-pass rate control scheme is highly practical for the real-time video coding application. Yuan Li 0014, Huizhu Jia, Chuang Zhu, Meng Li 0016, Wen Gao 0001 |
ICASSP | 1 |
| 2014 | Inter-dependent rate-distortion modeling for video coding and its application to rate controlabstractRate control scheme using independent rate-distortion (R-D) model at the minimum coding unit (macroblock) level has been widely discussed in the literature where R-D optimization is performed without consideration of the inter-dependencies of different coding units. In this paper, we extend these techniques to the more general situations - rate control for inter-dependent video coding units. The interdependent distortion-quantization (D-Q) model and rate-quantization (R-Q) model are formulated separately based on the analysis of the relationship between the spatial-domain residual and the transform-domain residual. Then a window-based rate control scheme with frame bit allocation and video quality optimization is proposed, which uses the approximated R-D model to reduce the computational complexity. Simulation results demonstrate that the proposed algorithm shows excellent peak signal-to-noise ratio (PSNR) performance under the bit rate constraint. This one-pass rate control scheme is highly practical for the realtime video coding application. Yuan Li 0014, Huizhu Jia, Pan Ma, Chuang Zhu, Wen Gao 0001 |
ICME | 1 |
| 2014 | A resolution-adaptive interpolation filter for video codecabstractThe fraction-pel interpolation filter varies in the video coding standards such as H.264/AVC, AVS and HEVC. Since fractional-pel motion compensation plays an important role in the video encoder, the interpolation of fractional-pel pixels can be refined and designed better to enhance the coding efficiency. In this paper, we firstly propose the generation algorithm of interpolation filter coefficients, and four different tap filters, namely 4tap, 6 tap, 8 tap and 10tap, are tested. A resolution-adaptive interpolation filter for different resolution videos is then introduced based on this algorithm to achieve the maximum bitrate saving. In the proposed scheme, 4 tap filter is applied for the UHD (2560×1600 and above) videos, 6 tap filter and 10 tap filter are performed in the videos whose resolution ranging from 720P (1280×720) to 1080P (1920×1080) and the videos with the resolution below 720P, respectively. When 4 tap filter and 6 tap filter are used in high-definition video, the coding efficiency can increase and the computational complexity will reduce greatly, which is actually beneficial to make hardware optimization more effectively especially SIMD (Single Instruction Multiple Data) and VLSI design. Experiments show that the average BD-rate gains on luma Y, chroma U and V are 1.4%, 0.7% and 0.7% for LP-Main configuration, when conducted in HEVC reference software HM11.0. The coding efficiency gains are significant for some video sequences and can reach up to 6.1%. Ronggang Wang, Yuan Li 0014, Chuang Zhu, Huizhu Jia, Wen Gao 0001 |
ISCAS | 3 |
| 2014 | Window-based rate control for video quality optimization with a novel INTER-dependent rate-distortion model
Yuan Li 0014, Huizhu Jia, Chuang Zhu, Mingyuan Yang, Wen Gao 0001 |
Signal Process. Image Commun. | 1 |
| 2013 | A high-throughput low-latency arithmetic encoder design for HDTVabstractIn this paper, we propose a high-throughput low-latency arithmetic encoder (AE) design suitable for high definition (HD) real-time applications employing advanced video coding standards such as H.264/AVC or AVS and using a macroblock (MB) level pipeline. First, in order to derive the performance requirement on the AE, a buffer model in connected with which it is designed is thoroughly analyzed. Then, using joint algorithm-architecture optimization and multi-bin processing techniques, we introduce a novel binary arithmetic coder (BAC) architecture with throughput of 2∼4 bins per cycle sufficient for real-time encoding. Furthermore, a hybrid context memory scheme is presented to meet the throughput requirement on the BAC. Simulation result shows that our design can support 1080p at 60 fps for AVS HDTV real-time coding with a bin rate up to 107K per MB line. Synthesized with the TSMC 0.13μm technology, the AE can run at 200MHz and costs 47.3K gates. By operating at 130MHz, the design is also verified in an AVS HD encoder on a Xilinx Virtex-6 FPGA prototype board for 1080p at 30 fps. Yuan Li 0014, Shanghang Zhang, Huizhu Jia, Wen Gao 0001 |
ISCAS | 1 |
| 2011 | A highly efficient pipeline architecture of RDO-based mode decision design for AVS HD video encoderabstractLike H.264, AVS video coding standard also uses macroblock (MB) based motion compensation (MC) and mode decision (MD). Rate distortion optimization (RDO) is the best known mode decision method, but with a high computational complexity that limits its applications. In our paper, firstly an MD algorithm based on RDO is given, which makes more mode candidates enter into RDO mode decision with little hardware resource increment. We further analyze the pipeline structure in detail, and implement a block-level 5-stage hardware pipeline. It can support the real time RDO mode decision processing of 1080P@30fps, and the coding efficiency is about 0.5db higher than the traditional SAD method. Our design is described in high-level Verilog/VHDL hardware description language and implemented under SMIC 0.18-µm CMOS technology with 215K logic gates and 80 KB SRAMs. Chuang Zhu, Yuan Li 0014, Huizhu Jia, Hai Bing Yin |
ICME | 2 |