EDBT 2026 Demo / reviewers in the wild / expert
Wei Zhou 0020
dblp:69/5011-20
· DBLP profile ↗
37ranked-venue papers
4as first author
17since 2021 · last 2026
0000-0001-9715-6957ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 24 · 3 first-author · 11 since 2021Artificial intelligence and machine learning · 14 · 1 first-author · 10 since 2021Systems, architecture and hardware · 9 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Score-Based Model for Low-Rank Tensor RecoveryabstractLow-rank tensor decompositions (TDs) provide an effective framework for multiway data analysis. Traditional TD methods rely on predefined structural assumptions, such as CP or Tucker decompositions. From a probabilistic perspective, these methods effectively model the relationships between latent factors and the low-rank tensor using Dirac delta distributions. However, tensor low-rank decomposition is inherently non-unique, leading to a multimodal distribution over possible solutions. Critically, such prior knowledge is rarely available in practical scenarios, particularly regarding the optimal rank structure and contraction rules. To address this issue, we propose a score-based model that eliminates the need for predefined structural or distributional assumptions, enabling the learning of compatibility between tensors and latent factors. Specifically, a neural network is designed to learn the energy function, which is optimized via score matching to capture the gradient of the joint log-probability of tensor entries and latent factors. Our method allows for modeling structures and distributions beyond the Dirac delta assumption. Moreover, integrating the block coordinate descent (BCD) algorithm with the proposed smooth regularization enables the model to perform both tensor completion and denoising. Experimental results demonstrate significant performance improvements across various tensor types, including sparse and continuous-time tensors, as well as visual data. Zhengyun Cheng, Guanwen Zhang, Yi Xu 0008, Wei Zhou 0020, Xiangyang Ji |
AAAI | 5 |
| 2026 | AMA-ViT: Acoustic-Mechanism-Aware Vision Transformer for Underwater Target Recognition
Zhangjie Cai, Ruiting Sun, Zhenhong Liao, Guanwen Zhang, Wei Zhou 0020 |
ICPR (9) | 5 |
| 2026 | SLIQ-ViT: Sensitivity-aware log-uniform integer quantization for efficient vision transformers
Ruiting Sun, Honglu He, Guanwen Zhang, Wei Zhou 0020 |
Expert Syst. Appl. | 4 |
| 2026 | Low-rank tensor recovery via variational schatten-p quasi-norm and Jacobian regularization
Zhengyun Cheng, Guanwen Zhang, Yi Xu 0008, Xiangyang Ji, Wei Zhou 0020 |
Neurocomputing | 6 |
| 2025 | KPDepth-VO: Self-Supervised Learning of Scale-Consistent Visual Odometry and Depth With Keypoint Features From Monocular VideoabstractMonocular visual odometry (VO) is crucial for the application of various autonomous systems. However, the inherent scale ambiguity issue in monocular methods greatly limits their performance in pose estimation. In this paper, we propose a hybrid monocular VO system named KPDepth-VO, which solves camera pose from monocular video based on sparse keypoints. To estimate the scale-consistent relative pose, we present a novel photometric-sensitive depth uncertainty model that accounts for the depth uncertainty introduced by limitations in the photometric error constraint. We also introduce an uncertainty-aware scale recovery strategy that incorporates depth uncertainty for reliable scale alignment. Additionally, we propose a novel difference attention mechanism to construct a point filter that effectively filters out less distinctive points, ensuring high-quality matches for more accurate and efficient pose estimation in the proposed system. Experimental results on the KITTI dataset and Oxford Robotcar dataset demonstrate that our system can predict scale-consistent trajectories from monocular videos and achieve state-of-the-art performance among similar methods. Meanwhile, the depth network within our system achieves competitive depth estimation performance on KITTI depth benchmark. Guanwen Zhang, Zhengyun Cheng, Wei Zhou 0020 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Parallel Heterogeneous Networks With Adaptive Routing for Online Video Lane DetectionabstractLane detection plays a critical role in the field of autonomous driving. Most previous lane detection methods focus exclusively on analyzing individual images and overlook the inter-frame dynamics, while car-mounted cameras capture continuous streams amenable to leveraging intra-video context. In challenging scenes with blur, occlusion, or illumination variations, considering visible lanes from previous frames can aid current frame interpretation. To tackle above challenges, we introduce a parallel heterogeneous framework, called PHNet, for video lane detection. First, a novel router automatically analyzes multi-level features to determine the inference route of candidates with different visual cues. Then, a cross-frame attention mechanism leverages relationships between current candidates and positive embeddings from past frame to aggregate contextual cues. Unlike existing offline video lane detection methods, which face limitations in handling long video sequences and real-time video streams due to computational constraints, our proposed method operates as an online framework capable of processing clips of any length. Our method demonstrates robust lane detection through temporally association modeling and efficient online inference. The extensive experiments on public benchmark show that PHNet has superior performance versus state-of-the-art video lane detection methods. Zhengyun Cheng, Guanwen Zhang, Wei Zhou 0020 |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | A Multi-Scale Feature Fusion Network for Chip Surface Defect DetectionabstractChip surface defect detection plays a crucial role in semiconductor manufacturing and the electronics industry. However, chip surface defects have relatively tiny defect areas and complex defect features make the chip surface defect detection task low in accuracy. In this paper, we propose a multi-scale feature fusion network based on encoder-decoder architecture to solve this challenge. Particularly, we propose a Residual Feature Fusion Attention Module (RFFAM) and an Intersection over Union Loss for Defect with Auxiliary Bounding Box (DA-IoU), and we incorporate a Transformer Prediction Head (TPH) for the network. Additionally, in order to solve the problem of the shortage of chip surface defect datasets, this paper proposes a chip surface defect dataset containing four defect categories. The experimental results show that the method not only improves the detection accuracy but also maintains a small number of parameters, satisfying the engineering requirements of chip surface defect detection. Haoang Ren, Mengke Tian, Guanwen Zhang, Wei Zhou 0020 |
ICIP | 4 |
| 2024 | Video-Based Semi-automatic Drivable Area Segmentation
Zhengyun Cheng, Guanwen Zhang, Wei Zhou 0020 |
ICPR (17) | 4 |
| 2024 | MSDNet: A Multi-scale Dense Network for Chip Surface Defect Segmentation
Guanwen Zhang, Wei Zhou 0020 |
ICPR (25) | 4 |
| 2024 | A Data Augmentation Approach for Well Log Interpretation
Yaobin Wang, Guanwen Zhang, Wei Zhou 0020 |
ICPR (25) | 6 |
| 2024 | Self-Supervised Learning of Monocular Visual Odometry and Depth with Uncertainty-Aware Scale ConsistencyabstractThe inherent scale ambiguity issue greatly limits the performance of monocular visual odometry. In recent years, a variety of methods have been proposed for self-supervised learning of ego-motion and depth estimation, incorporating specifically designed scale-consistency constraints that utilize estimated depth as a reference. However, these existing methods neglect the influence of the depth uncertainty introduced by the dominant photometric loss, which leads to unreliable depth estimation in difficult regions and detrimentally affects scale alignment. To solve these problems, we introduces a feature-based visual odometry learning system with an effective scale recovery strategy in this paper. Additionally, we propose a learning method to estimate the photometric-sensitive depth uncertainty for guiding the scale recovery. The proposed method is evaluated on KITTI odometry, and the experimental results demonstrate that our system can predict scale-consistent trajectories from monocular videos and achieves state-of-the-art performance. Moreover, the proposed method achieves competitive performance on KITTI depth estimation. Guanwen Zhang, Wei Zhou 0020 |
ICRA | 3 |
| 2024 | MonoBooster: Semi-Dense Skip Connection With Cross-Level Attention for Boosting Self-Supervised Monocular Depth Estimation
Guanwen Zhang, Zhengyun Cheng, Wei Zhou 0020 |
IEEE Signal Process. Lett. | 4 |
| 2023 | COP: A Combinational Optimization Power Budgeting Method for Manycore Systems in Dark SiliconabstractDark silicon is a phenomenon of under-utilization in today's manycore systems due to power and thermal limitations. In order to improve the performance of dark silicon systems, it is necessary to adopt dynamic power constraints for different core mapping decisions. However, existing power budgeting methods are generally over pessimistic, e.g., Thermal Safe Power (TSP), or over optimistic, e.g., Greedy based Dynamic Power (GDP). This paper proposes a practical power budgeting method, called Combinational Optimization Power (COP). Different from existing methods, which ignore some actual factors, such as communication overhead and lifetime reliability, COP formulates the power budgeting problem as a thermal-constrained combinational optimization power problem. For the steady-state case, COP achieves the target fusion of optimized temperature and communication energy consumption by applying task priority ranking and task-to-core mapping. For the transient case, COP uses the rainflow counting algorithm to construct the reliability framework based on the thermal cycling failure mechanism, and then establishes a linear time-invariant transient temperature model to obtain the core mapping selection and the corresponding dynamic power budget. Experimental results demonstrate that COP is capable of providing an optimized core mapping decision, which can maximize power budget while ensuring the system performance. Xin Li 0042, Yaqi Ju, Rongyao Wang, Wei Zhou 0020 |
IEEE Trans. Computers | 6 |
| 2022 | DILane: Dynamic Instance-Aware Network for Lane Detection
Zhengyun Cheng, Guanwen Zhang, Wei Zhou 0020 |
ACCV (2) | 4 |
| 2022 | Rethinking Low-Level Features for Interest Point Detection and Description
Guanwen Zhang, Zhengyun Cheng, Wei Zhou 0020 |
ACCV (2) | 4 |
| 2021 | Combinational Optimization Power (COP) - A Practical Power Budgeting Method for Many-coresabstractIn order to improve the performance of many-core systems, it is necessary to adopt dynamic power constraints for different core mapping decisions. However, existing power budgeting methods are generally over pessimistic, e.g. Thermal Safe Power (TSP), or over optimistic, e.g. Greedy based Dynamic Power (GDP). This paper proposes a practical power budgeting method, called Combinational Optimization Power (COP). Different from existing methods, which ignore some actual factors, such as communication overhead and lifetime reliability, COP formulates the power budgeting problem as a thermal-constrained combinational optimization power problem. For the steady-state case, COP achieves the target fusion of optimized temperature and communication energy consumption by applying task priority ranking and task-to-core mapping. For the transient case, COP uses the rainflow counting algorithm to construct the reliability framework based on the thermal cycling failure mechanism, and then establishes a linear time-invariant transient thermal model to obtain the core mapping selection and the corresponding dynamic power budget. Experimental results show that COP is capable of providing an optimized core mapping decision, which can maximize power budget while ensuring the system performance. Xin Li 0042, Yaqi Ju, Huajie Lin, Wei Zhou 0020 |
IECON | 5 |
| 2021 | A deep learning method for video-based action recognitionabstractAbstract In this paper, a deep learning method for video‐based action recognition is proposed. On the one hand, boundary compensation on the basis of a deep neural network is performed to achieve action proposal. Boundary compensation considering non‐maximum suppression according to sliding window priority is applied to remove redundant windows. To accurately detect boundaries, a boundary compensation network is established with multiple networks to process different numbers of segments. On the other hand, action recognition based on the resultant action proposals is performed. To further utilise boundary compensation, three methods are introduced for key frame selection. Optical flow and RGB features are combined via a channel fusion to realise feature representation. A two‐stream network with a spatiotemporal structure is adopted for action recognition. The proposed method is evaluated on three public datasets. The experimental results demonstrate that the proposed method achieves a superior performance to that of state‐of‐the‐art methods. Guanwen Zhang, Yukun Rao, Wei Zhou 0020, Xiangyang Ji |
IET Image Process. | 4 |
| 2020 | Recurrent Deep Attention Network for Person Re-IdentificationabstractPerson re-identification (re-id) is an important task in video surveillance. It is challenging due to the appearance of person varying a wide range across non-overlapping camera views. Recent years, attention-based models are introduced to learn discriminative representation. In this paper, we consider the attention selection in a natural way as like human moving attention on different parts of the visual field for person re-id. In concrete, we propose a Recurrent Deep Attention Network (RDAN) with an attention selection mechanism based on reinforcement learning. The proposed RDAN aims to progressively observe the identity-sensitive regions to build up the representation of individuals. Extensive experiments on three person reid benchmarks Market-1501, DukeMTMC-reID, and CUHK03-NP demonstrate the proposed method can achieve competitive performance. Xianfei Duan, Guanwen Zhang, Wei Zhou 0020 |
ICPR | 5 |
| 2020 | TSMSAN: A Three-Stream Multi-Scale Attentive Network for Video Saliency DetectionabstractVideo saliency detection is an important low-level task that has been used in a large range of high-level applications. In this paper, we proposed a three-stream multi-scale attentive network (TSMSAN) for saliency detection in dynamic scenes. TSMSAN integrates motion vector (MV) representation, static saliency map, and RGB information in multi-scales together into one framework on the basis of Fully Convolutional Network (FCN) and spatial attention mechanism. On the one hand, the respective motion features, spatial features, as well as the scene features can provide abundant information for video saliency detection. On the other hand, spatial attention mechanism can combine features with multi-scales to focus on key information in dynamic scenes. In this manner, the proposed TSMSAN can encode the spatiotemporal features of the dynamic scene comprehensively. We evaluate the proposed approach on two public dynamic saliency datasets. The experimental results demonstrate TSMSAN is able to achieve the state-of-the-art performance as well as the excellent generalization ability. Furthermore, the proposed TSMSAN can provide more convincing video saliency information, in line with human perception. Guanwen Zhang, Jiaming Yan, Wei Zhou 0020 |
ICPR | 4 |
| 2020 | Deep Reinforcement Learning for Autonomous Driving by Transferring Visual FeaturesabstractDeep reinforcement learning (DRL) has achieved great success in processing vision-based driving tasks. However, the end-to-end training manner makes DRL agents suffer from overfitting training scenes. The agents easily fail to generalize to unseen environments. In this paper, we propose a deep reinforcement learning for autonomous driving by transferring visual features. We formulate the DRL training as a perception and control module and introduce adversarial training mechanism for autonomous driving. The perception module is able to extract invariant features between different domains through adversarial training. While the DRL agent can then be trained on the basis of low dimensional states. In this manner, the proposed approach enables trained agents to adapt to unseen environments by learning robust features invariant across various scenes. We evaluate the proposed approach by transferring visual features between different simulators. The experimental results demonstrate the driving policy trained in the source domain can be directly applied in the target domain, and achieve great efficient and effective performance for autonomous driving. Guanwen Zhang, Wei Zhou 0020 |
ICPR | 4 |
| 2020 | Frame level rate control algorithm based on GOP level quality dependency for low-delay hierarchical video coding
Wei Zhou 0020, Henglu Wei, Xin Zhou 0001, Zhemin Duan |
Signal Process. Image Commun. | 2 |
| 2020 | Accurate On-Chip Temperature Sensing for Multicore Processors Using Embedded Thermal SensorsabstractThermal issues seriously restrict the quality, reliability, and lifetime of semiconductor chips. To prevent thermal runaway situations, modern processors deploy numerous on-die thermal sensors to collect temperature information, which is then used to guide dynamic thermal management (DTM) mechanisms. Accurate on-chip temperature information is critical for DTM as temperature overestimation will degrade the performance by an unnecessary invocation of thermal control mechanisms, and underestimation will result in the reliability issues of the systems. In this article, two effective techniques are proposed to achieve an accurate on-chip temperature sensing. The first technique is a synergistic calibration method for forecasting the actual temperatures of noisy thermal sensors. Second, we propose a full thermal characterization technique based on convolutional neural networks (CNNs) to accurately recover the entire thermal maps by using a limited number of thermal sensors. In a realistic scenario, these two techniques can be used in combination to provide a more accurate thermal monitoring. By utilizing the sophisticated infrared imaging setup, the effectiveness of the proposed techniques is validated on a real 45-nm AMD quad-core chip. The simulation results show that the methods achieve significant improvements compared with existing techniques in the literature. The successful implementation of the proposed methods will significantly improve the efficiency of DTM. Xin Li 0042, Wei Zhou 0020, Zhemin Duan |
IEEE Trans. Very Large Scale Integr. Syst. | 3 |
| 2019 | High-Resolution Thermal Maps Extraction of Multi-Core Processors Based on Convolutional Neural NetworksabstractThermal issues are a major concern in high-end computing systems as they severely constrain the performance and shorten the lifetime of integrated circuits. Using embedded on-die thermal sensors, state-of-the-art processors widely employ dynamic thermal management (DTM) mechanisms to prevent thermal runaway situations in multi-core architectures. Full thermal characterization is particularly useful for fine-grained thermal management techniques, examples of which include per-core workload scheduling, and voltage and frequency scaling. In this paper, a new direction for full thermal reconstruction of multi-core processors is proposed based on convolutional neural networks (CNNs) to precisely recover the overall thermal maps from a small number of sensors. The effectiveness of our method is verified on a real AMD quad-core processor. Experimental results indicate that the proposed method is capable of handling high-resolution thermal extraction, and achieves significant improvements compared to available techniques in the literature. The main contribution of this work is to demonstrate the ability of deep learning approaches for full thermal reconstruction of semiconductor chips. The success of our proposed method will assist DTM to achieve a more accurate thermal monitoring. Xin Li 0042, Xingtao Ou, Wei Zhou 0020, Zhemin Duan |
IECON | 5 |
| 2019 | All zero block detection for HEVC based on the quantization level of the maximum transform coefficient
Henglu Wei, Wei Zhou 0020, Xin Zhou 0001, Zhemin Duan |
Multim. Tools Appl. | 2 |
| 2019 | High-Performance FPGA-Based CNN Accelerator With Block-Floating-Point ArithmeticabstractConvolutional neural networks (CNNs) are widely used and have achieved great success in computer vision and speech processing applications. However, deploying the large-scale CNN model in the embedded system is subject to the constraints of computation and memory. An optimized block-floating-point (BFP) arithmetic is adopted in our accelerator for efficient inference of deep neural networks in this paper. The feature maps and model parameters are represented in 16-bit and 8-bit formats, respectively, in the off-chip memory, which can reduce memory and off-chip bandwidth requirements by 50% and 75% compared to the 32-bit FP counterpart. The proposed 8-bit BFP arithmetic with optimized rounding and shifting-operation-based quantization schemes improves the energy and hardware efficiency by three times. One CNN model can be deployed in our accelerator without retraining at the cost of an accuracy loss of not more than 0.12%. The proposed reconfigurable accelerator with three parallelism dimensions, ping-pong off-chip DDR3 memory access, and an optimized on-chip buffer group is implemented on the Xilinx VC709 evaluation board. Our accelerator achieves a performance of 760.83 GOP/s and 82.88 GOP/s/W under a 200-MHz working frequency, significantly outperforming previous accelerators. Xiaocong Lian, Zhenyu Liu 0001, Zhourui Song, Jiwu Dai, Wei Zhou 0020, Xiangyang Ji |
IEEE Trans. Very Large Scale Integr. Syst. | 5 |
| 2018 | Synergistic Calibration of Noisy Thermal Sensors Using Smoothing Filter-Based Kalman PredictorabstractEmbedded thermal sensors are very susceptible to a variety of noise sources, including environmental uncertainty and process variation. This causes the discrepancies between actual temperatures and those observed by on-chip thermal sensors, which seriously affect the efficiency of dynamic thermal management (DTM). In this paper, a smoothing filter-based Kalman prediction technique is proposed to estimate the accurate temperatures of noisy sensors. On this basis, a multi-sensor synergistic calibration algorithm is proposed to improve the simultaneous prediction accuracy of multiple sensors. Moreover, an infrared imaging-based temperature measurement technique is also proposed to capture the thermal traces of an AMD quad-core processor in real-time. The acquired real temperature data are used to evaluate our prediction performance. Simulation shows that the synergistic calibration scheme can achieve an average reduction of the root-mean-square error (RMSE) by 75.9% compared with assuming the thermal sensor readings to be ideal. Additionally, the average false alarm rate (FAR) of the corrected sensor temperature readings can be reduced by 21.6%. These results clearly demonstrate that if our approach is used to perform the temperature estimation, the response mechanisms of DTM can be triggered to adjust the voltages, frequencies, and cooling fan speeds at more appropriate times. Xin Li 0042, Xingtao Ou, Henglu Wei, Wei Zhou 0020 |
ISCAS | 4 |
| 2018 | Parallel Content-Aware Adaptive Quantization-Oriented Lossy Frame Memory Recompression for HEVCabstractSince the development of ultrahigh-definition video, the huge bandwidth and power requirements of external memory have hindered the development of video encoder applications. Power constraints have become a particularly serious problem for portable video codec systems. With high-rate configurations [quantization parameter (QP) ≤ 22 in HEVC test model (HM) reference software], the compression performance of the existing lossless compression algorithms noticeably degrades, because the reference frames are becoming rich of textures. On the other hand, the mathematical analysis of this paper revealed that more quantization noises can be endured by the texture-rich area. Therefore, we develop an adaptive quantization-oriented parallel lossy frame memory recompression algorithm. The contributions of this paper include the following. First, a content-aware adaptive quantization method is devised to achieve a stable high compression ratio that does not deteriorate for highly quality texture-rich pictures. When QP$\in $ [12, 22], a data reduction ratio improvement of up to 14% is obtained compared with the best lossless algorithm. Furthermore, it can reduce the quality loss by 0.49-3.36 dB in terms of Bjøntegaard delta peak signal-to-noise rate (BD-PSNR) compared with the fixed length quantization method. Second, to solve the low throughput problem caused by the pixel-grain prediction method, a parallel directional prediction scheme is developed. It can double or quadruple the throughput with a prediction accuracy loss of only 1.7% or 3.3%, respectively. Using the above-mentioned methods, bandwidth and memory requirements are reduced up to 70.6% and 41.0%, respectively, with a corresponding savings of 59.3% in the dynamic power consumption of the off-chip dynamic random access memory, while the BD-PSNR is -0.04 dB, or, equivalently, Bjøntegaard delta bit rate (BD-BR) is 1.27%. Using TSMC 65-nm CMOS technology, the proposed frame memory compressor and decompressor can achieve the throughputs of up to 2.89 and 2.26 Gpixels/s, respectively. It is applicable to a Super Hi-Vision(8K)@68-frames/s real-time encoding with a Level D reference data reuse scheme. Xiaocong Lian, Zhenyu Liu 0001, Wei Zhou 0020, Zhemin Duan |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2016 | Visual saliency based perceptual video coding in HEVCabstractPerceptual video coding has the potential to provide the same visual quality at a lower bit-rate, compared with the traditional objective quality based scheme. Visual saliency represents the probability of human attention over frames, and it is used for allocating coding bits or controlling visual quality. In this paper, a HEVC compliant perceptual video coding scheme is proposed based on visual saliency. At first visual saliency map is attained to indicate the distribution of saliency. Then refined distortion allocating method is performed in CU level with adaptive QP which is adjusted by the average visual saliency. Besides, a fast CU mode decision algorithm suitable for perceptual video coding in HEVC is proposed to accelerate the encoder. In the fast algorithm, the average saliency is used to estimate texture complexity and movements in videos. Experimental results show that up to 22.52% bit-rate and 43.48% encoding time can be saved by our methods with negligible perceptual quality loss. Henglu Wei, Xin Zhou 0001, Wei Zhou 0020, Chang Yan, Zhemin Duan, Nana Shan |
ISCAS | 3 |
| 2016 | Perceptual CU Size Decision and Fast Prediction Mode Decision Algorithm for HEVC Intra CodingabstractIntra coding has been significantly improved in HEVC over H.264/AVC with quad-tree based coding unit (CU) structure from size 64×64 to 8×8 and more prediction modes. However, these techniques cause a dramatic increase in computational complexity. In this paper, a novel intra coding algorithm is proposed consists of perceptual CU size decision algorithm and fast intra prediction mode decision algorithm. Firstly, based on the visual saliency detection, an adaptive and perceptual CU size decision method is proposed to alleviate intra encoding complexity. Furthermore, a fast intra prediction mode decision algorithm with step halving rough mode decision method is presented to selectively check the potential modes and effectively reduce the complexity of computation. Experimental results show that our proposed method reduces the computational complexity of the current HM to about 54.18% in encoding time with only 0.36% increases in BD rate and reasonable peak signal-to-noise ratio losses. Xin Zhou 0001, Guangming Shi, Wei Zhou 0020 |
ISM | 3 |
| 2016 | Lossless Frame Memory Compression Using Pixel-Grain Prediction and Dynamic Order Entropy CodingabstractPower constraints constitute a critical design issue for the portable video codec system, in which the external dynamic random access memory (DRAM) accounts for more than half of the overall system power requirements. With the ultrahigh-definition video specifications, the power consumed by accessing reference frames in the external DRAM has become the bottleneck for the portable video encoding system design. To relieve the dynamic power stresses introduced by the DRAM, a lossless compression algorithm is devised to reduce the external traffic and the memory requirements of reference frames. First, pixel-granularity directional prediction is adopted to decrease the prediction residual energy by 54.1% over the previous horizontal prediction. Second, the dynamickth-order unary/Exp-Golomb rice coding is applied to accommodate the large-valued prediction residues. With the aforementioned techniques, an average data traffic reduction of 68.5% for the off-chip reference frames is obtained, which consequently reduces the dynamic power requirements of the DRAM by 42.3%. Based on the high data reduction ratio of the proposed compression algorithm, a partition group table-based storage space reduction scheme is provided to improve the utilization of row buffers in the DRAM. Consequently, an additional 14.5% of the DRAM dynamic power can be saved by reducing the number of row buffer activations. In total, a 56.8% decrease in the dynamic power requirements of the external reference frame access can be obtained using our strategies. With TSMC 65-nm CMOS logic technology, our algorithm was implemented in a parallel VLSI architecture based on a compressor and decompressor at the cost of 36.5k and 34.7k, respectively, in terms of gate count. The throughputs of the proposed compressor and decompressor are 1.54 and 0.78 Gpixels/s, which are suitable for quad full high definition (4K) @ 94 frames/s real-time encoding with the level-D reference data reuse scheme. Xiaocong Lian, Zhenyu Liu 0001, Wei Zhou 0020, Zhemin Duan |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2016 | A High-Throughput and Multi-Parallel VLSI Architecture for HEVC Deblocking FilterabstractThis paper presents a high-throughput and multi-parallel VLSI hardware architecture for the deblocking filter in the HEVC video coding standard. First, an implementation-friendly and fast boundary judgment method is proposed to avoid using the original recursion loop approach. Then a dedicated parallel VLSI architecture composed of four parallel filtering cores is presented based on the proposed boundary judgment method. With the parallel luma/chroma filtering and parallel vertical/horizontal edges filtering order, the proposed VLSI architecture can process filtering operations for one largest coding unit (LCU) with less filtering cycles than other conventional approaches. Furthermore, filtering efficiency is improved due to a novel ping-pang buffer architecture and the on-chip single-port SRAM with dedicated data arrangement in the memory modules. Experimental results demonstrate that the proposed deblocking filter architecture improves the performance by 28-89% at the expense of the slightly increased gate count compared to the previously known architecture in HEVC. The proposed architecture can reach a high operating clock frequency of 278 MHz with TSMC 90 nm library and meet the real time requirement of the deblocking filter for 8 K × 4 K video format at 123 frame/s. Wei Zhou 0020, Jingzhi Zhang, Xin Zhou 0001, Zhenyu Liu 0001, Xiaoxiang Liu |
IEEE Trans. Multim. | 1 |
| 2015 | An efficient interpolation filter VLSI architecture for HEVCabstractFirstly, an implementation-friendly interpolation filter algorithm is proposed in this paper. It can save 19.6% processing time on average with negligible coding quality degradation. Then based on the proposed algorithm, an optimized interpolation filter VLSI architecture, composed of the reused data path of interpolation, efficient memory organization and the pipeline interpolation filter engine is presented to reduce the implement hardware area. The resulting design can achieve 240 MHz with only 37.2K gate count and support real-time interpolation filter operation of 3840×2160@47fps video application by using 90nm CMOS technology. Wei Zhou 0020, Xin Zhou 0001, Xiaocong Lian |
ICASSP | 1 |
| 2015 | Pixel-grain prediction and K-order UEG-rice entropy coding oriented lossless frame memory compression for motion estimation in HEVCabstractWith the resolution of video sequences increasing, the memory bandwidth and space requirement has become one bottleneck of a video coding system. Most of previous frame memory compression works merely focused on reducing the averaged bandwidth demand. In this paper, a lossless frame memory compression algorithm for motion estimation in HEVC and the corresponding hardware architecture are described. The compression algorithm is composed of the 64×4 partition compression algorithm and the partition group table based storage scheme. Experimental results demonstrate that, on average, the proposed algorithm reduces 70.1% memory bandwidth and 41% memory space, while saving 59.0% dynamic energy. With TSMC 90nm CMOS logic technology, our algorithm were implemented in a parallel VLSI architecture of compressor and decompressor at the cost of 48.1K and 44.2K respectively in gate-count. The achieved peak throughput is 1.51Gpixel/s, which is sufficient to handle the SHV(8K)@30fps real-time encoding with the level-D reference data reuse scheme. Xiaocong Lian, Zhenyu Liu 0001, Wei Zhou 0020, Zhemin Duan |
ICIP | 3 |
| 2015 | An efficient all zero block detection algorithm based on frequency characteristics of DCT in HEVCabstractLike the previous video coding standard, DCT and quantization are also adopted in HEVC. Compared with H.264/AVC, HEVC employs larger transform blocks, which makes the all zero block detection algorithm designed for H.264/AVC inefficient for HEVC. An efficient all zero block detection algorithm aimed at HEVC, especially for 16×16 and 32×32 transform blocks, is proposed in this paper based on frequency characteristics of DCT. By analysing the distribution of residual energy in frequency domain, Hadamard transform is used to evaluate only a part of DCT coefficients which usually consume most of energy. To make the proposed algorithm efficient in complex sequences, a sum of absolute transformed difference based method is used to evaluate the maximum value of the rest of the coefficients. Experimental results show that more than 90% all zero blocks for 4×4, 8×8 and 16×16 transform blocks and about 80% for 32×32 transform blocks can be detected by the proposed algorithm. In addition, about 50% computational complexity in DCT/quantization can be reduced with negligible loss of video quality and compression efficiency. Henglu Wei, Wei Zhou 0020, Xin Zhou 0001, Zhemin Duan |
VCIP | 2 |
| 2015 | A high-throughput deblocking filter VLSI architecture for HEVCabstractThis paper presents a novel VLSI hardware architecture for the real-time high-throughput implementation of the HEVC deblocking filtering. Based on the proposed implementation-friendly boundary judgment method, a dedicated multi-parallel architecture composed of four parallel filtering cores, parallel luma/chroma filtering and parallel vertical/horizontal edges filtering is presented. Experimental results demonstrate that the proposed architecture can greatly improve the performance at the expense of the slightly increased hardware cost compared to the previously known architecture in HEVC. The proposed architecture can also meet the real-time requirement of the deblocking filter for 8K×4K video format at 123fps under 278MHz clock rate. Wei Zhou 0020, Jingzhi Zhang, Xin Zhou 0001, Tongqing Liu |
VCIP | 1 |
| 2014 | A Fast CABAC Algorithm for Transform Coefficients in HEVC
Nana Shan, Wei Zhou 0020, Zhemin Duan |
ICA3PP (1) | 2 |
| 2007 | Efficient Motion Estimation Scheme for H.264 Based on BP Neural Network
Wei Zhou 0020, Haoshan Shi, Zhemin Duan, Xin Zhou 0001 |
ISNN (3) | 1 |