EDBT 2026 Demo / reviewers in the wild / expert
Heming Sun
dblp:119/0191
· DBLP profile ↗
80ranked-venue papers
13as first author
54since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 56 · 10 first-author · 35 since 2021Systems, architecture and hardware · 22 · 3 first-author · 18 since 2021Artificial intelligence and machine learning · 8 · 6 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LLMdoctor: Token-Level Flow-Guided Preference Optimization for Efficient Test-Time Alignment of Large Language Models
Tiesunlong Shen, Rui Mao 0010, Jin Wang 0008, Heming Sun, Jian Zhang 0087, Xuejie Zhang 0002, Erik Cambria |
AAAI | 4 |
| 2026 | Post-Training Quantization for Overfitted Neural Image Coding
Yile Tu, Heming Sun |
ISCAS | 2 |
| 2026 | Variable-Rate Learned Image Compression Using Parameter Efficient Fine-Tuning and Trainable Quantization Step-Size
Ran Wang 0015, Heming Sun, Jiro Katto |
ISCAS | 3 |
| 2026 | FPGA-Based Low-Power Signed Approximate Multipliers for Diverse Error-Resilient ApplicationsabstractThe Booth algorithm is widely used for efficient signed multiplication due to its ability to reduce partial products. A higher radix Booth multiplier generates fewer partial products, while it also increases hardware complexity in the generator, diminishing the advantage of fewer accumulators. Previous optimizations of generators and accumulators were designed for application-specific integrated circuits (ASICs), but their performance gains cannot be comparably translated to field-programmable gate arrays (FPGAs) due to differences in architecture. This article proposes FPGA-friendly approximate Booth multipliers that combine approximate hybrid-radix partial product generation with resource-efficient accumulation techniques. Initially, to improve generation efficiency, an look-up table (LUT)-reused exact radix-8 generator is introduced through logical partitioning to integrate two types of partial products into a single LUT. In addition, approximate adjacent-compensation radix-8 and radix-16 generators are developed based on the Booth encoding bit-repetition principle. Later, to speed up partial product accumulation, an overlap-parallel accumulation scheme and various accumulators are proposed, reducing compression steps and enhancing resource utilization. Last, performance-configurable hybrid radix-8/-16 approximate Booth multipliers are designed to meet the needs of different error-resilient applications. The most hardware-efficient configuration of the proposed 16-bit multiplier reduces power–delay product (PDP) and LUT consumption by 38.31% and 35.66%, respectively, compared with the exact multiplier. Furthermore, the proposed designs offer a better balance between accuracy and hardware complexity than existing approximate multipliers. The practicality of these multipliers is demonstrated in both joint photographic experts group (JPEG) image compression and finite impulse response (FIR) filtering applications. An open-source library of the proposed multipliers is available athttps://github.com/YnuGuoLab/FPGA_Signed_Approx_Multo support further research. Xuetao Li, Heming Sun, Haroon Waris, Weiqiang Liu 0001 |
IEEE Trans. Very Large Scale Integr. Syst. | 4 |
| 2025 | Learned Image Codec on FPGA: Algorithm, Architecture and System DesignabstractThis paper describes our design for learned image codec (LIC) on FPGA, from the aspects of algorithm, architecture and system. For the algorithm, we build the neural network on the hyperprior structure. Besides, we present a quantization aware training scheme specifically adapted to LIC. For the architecture, we propose a fine-grained pipeline architecture. Channel parallelism constraint and neural network search are proposed to improve the DSP utilization and efficiency, respectively. For the system, we make a CPU-FPGA heterogeneous coding system in which a system-level pipeline is proposed to maximize the throughput. A 720P@30FPS demo and a cross-platform demo are provided in the websites.12 Heming Sun, Jing Wang 0181, Silu Liu, Shinji Kimura, Masahiro Fujita 0004 |
ASP-DAC | 1 |
| 2025 | FLAVC: Learned Video Compression with Feature Level AttentionabstractLearned Video Compression (LVC) aims to reduce redundancy in sequential data through deep learning approaches. Recent advances have significantly boosted LVC performance by shifting compression operations to the feature domain, often combining Motion Estimation and Motion Compensation modules (MEMC) with CNN-based context extraction. However, reliance on motions and convolution-driven context models limits generalizability and global perception. To address these issues, we propose a Feature-level Attention (FLA) module within a Transformer-based framework that explicitly perceives full-frame, thus bypassing confined motion signatures. FLA accomplishes global perception by converting high-level local patch embeddings into one-dimensional batch-wise vectors and replacing traditional attention weights to a global context matrix. Additionally, a dense overlapping patcher (DP) is introduced to retain local features before embedding projection. Furthermore, a Transformer-CNN mixed encoder is applied to alleviate the spatial feature bottleneck without increasing latent size. Experiments demonstrate excellent generalizability with universally efficient redundancy reduction in different scenarios. Extensive tests on four video compression datasets show that our method achieves state-of-the-art Rate-Distortion performance compared to existing LVC methods and traditional codecs. A down-scaled version of our model reduced computation overhead by a great margin while maintained strong performance. The code is available at https://github.com/Z-CV-code/FLAVC. Heming Sun, Jiro Katto |
CVPR | 2 |
| 2025 | Assessing the Reusability of Cloud-Received Feature Streams on Advanced NetworksabstractThe goal of this paper is to raise awareness of challenges and opportunities in the Collaborative Intelligence (CI) field and promote research on related standards. We begin by identifying a key challenge in CI applications, i.e., is it still possible for feature streams received in the cloud to be reused in the future by more advanced multitasking networks to achieve effective task accuracy? We then propose a framework to explore the generalization ability of cloud-received feature streams on more advanced networks from a coarse-grained to a fine-grained manner. We design a series of adapters of varying complexity to further explore the potential of feature streams for task network adaptation. Experiments show that sharing feature streams across multiple task networks could achieve an average of nearly 80% bitrate saving compared to Versatile Video Coding (VVC), which demonstrates the reuse potential of cloud-received feature streams. In addition, we make theoretical inferences about the adaptation range of shared feature streams, especially for those networks with high precision. Jiawang Liu, Hualong Yu, Heming Sun, Lu Yu 0003 |
ISCAS | 4 |
| 2025 | A Multi-Grid Implicit Neural Representation for Multi-View Videos
Qingyue Ling, Zhengxue Cheng, Donghui Feng 0003, Shen Wang 0013, Guo Lu, Heming Sun, Jiro Katto, Li Song 0001 |
PCS | 7 |
| 2025 | CGICM: CLIP-Guided Semantic Frequency Adaptation in Image Compression for MachinesabstractIn recent years, deep learning-based image compression techniques have advanced rapidly, surpassing traditional methods in terms of rate-distortion performance. However, in machine-oriented image compression, preserving high-level semantic information is of greater importance. Most existing methods employ only image-level prompts to guide frequency domain processing, leading to suboptimal preservation of semantic information for downstream machine vision tasks. To address this limitation, we propose a CLIP-guided semantic frequency domain adaptation module that extracts frequency features by applying both the fast Fourier transform and the wavelet transform. Guided by text-based semantics, the module further enhances the frequency components relevant to the target task, thereby improving machine perception performance. The proposed adapter is designed to be plug-and-play with existing learned image compression (LIC) models without requiring retraining of the full model. Experimental results demonstrate that our method outperforms state-of-the-art approaches in multiple machine vision tasks. Feng Liang 0001, Heming Sun, Jiro Katto |
VCIP | 3 |
| 2025 | Storage-and-Memory-Efficient Learned Image Compression With Quality-Aware Hyperprior PruningabstractABSTRACT Learned image compression (LIC) has become more and more important in recent years. The hyperprior‐module‐based LIC models, which use hyperprior module to predict the distribution of image features and improve entropy coder performance, have achieved remarkable rate‐distortion (RD) performance. However, the storage and memory costs of these LIC models are too high, resulting in higher difficulty to be applied to various devices, especially portable or edge devices. The storage and memory cost are directly linked to the parameter number. As a preliminary experiment, we manually assigned half channels for the hyperprior module in LIC models, reducing about 30% parameters in the model. The pruned models still kept similar RD performance to the original ones. This reveals that the hyperprior module in LIC models is highly redundant. In the meanwhile, LIC models with different reconstruction qualities require different amounts of parameters for the hyperprior module. Based on these phenomena, we propose a quality‐aware hyperprior pruning method that efficiently reduces the storage and memory cost of the hyperprior module and various context models. It consists of two parts. The first part is the pruning method itself, called enhanced ResRep on hyper path (ERHP). The second part is a quality‐aware threshold searching method, called pruning threshold searching (PTS), which prunes the hyperprior module based on the reconstruction qualities of LIC models. The experiments on various LIC models show that our methods reduce large volumes of storage cost (up to 74.6%) and memory cost (up to 41.5%), while keeping the performance the same before pruning. Ao Luo, Diego Fujii, Keisuke Nonaka, Heming Sun, Jiro Katto |
IET Image Process. | 4 |
| 2025 | MDLPCC: Misalignment-aware dynamic LiDAR point cloud compressionabstractLiDAR point cloud plays an important role in various real-world areas. It is usually generated as sequences by LiDAR on moving vehicles. Regarding the large data size of LiDAR point clouds, Dynamic Point Cloud Compression (DPCC) methods are developed to reduce transmission and storage data costs. However, most existing DPCC methods neglect the intrinsic misalignment in LiDAR point cloud sequences, limiting the rate–distortion (RD) performance. This paper proposes a Misalignment-aware Dynamic LiDAR Point Cloud Compression method (MDLPCC), which alleviates the misalignment problem in both macroscope and microscope. MDLPCC exploits a global transformation (GlobTrans) method to eliminate the macroscopic misalignment problem, which is the obvious gap between two continuous point cloud frames. MDLPCC also uses a spatial–temporal mixed structure to alleviate the microscopic misalignment, which still exists in the detailed parts of two point clouds after GlobTrans. The experiments on our MDLPCC show superior performance over existing point cloud compression methods. Ao Luo, Linxin Song, Keisuke Nonaka, Jinming Liu 0001, Kyohei Unno, Kohei Matsuzaki, Heming Sun, Jiro Katto |
J. Vis. Commun. Image Represent. | 7 |
| 2025 | Single model learned image compression utilizing multiple scaling factorsabstractImage compression is a critical task in multimedia. However, all learned-based single rate compression methods face challenges, such as prolonged training time due to the need for a dedicated model per bitrate and increased memory usage. Some variable rate methods require extra input, conditional networks, or still involve training multiple models. In this paper, we propose a unified approach using scaling factors to enable variable rate compression within a single model. The scaling factors consist of multi-gain units and quantization step size. The multi-gain units reduce redundancy in encoder and decoder representations, while the quantization step size controls quantization error. We also observe unevenness among slices in the Channel-Wise entropy model, and propose channel-wise quantization compensation by assigning specific step sizes to each slice. Our method supports continuous rate adaptation without retraining. Extensive experiments on CNN-based, Transformer-based, and CNN-Transformer mixed models demonstrate superior performance across a wide range of bitrates. • We enable variable-rate compression with a single model via multiple scaling factors. • We introduce channel-wise quantization compensation with slice-specific step sizes. • Our approach supports continuous rate adaptation without additional parameters. • Weighted training of larger Lagrange multipliers improves performance at all rates. • We validate our methods on CNN, Transformer, and hybrid CNN-Transformer models. Ran Wang 0015, Heming Sun, Jiro Katto |
J. Vis. Commun. Image Represent. | 3 |
| 2025 | Q-LIC: Quantizing Learned Image Compression With Channel SplittingabstractLearned image compression (LIC) has reached a comparable coding gain with traditional hand-crafted methods such as VVC intra. However, the large network complexity prohibits the usage of LIC on resource-limited embedded systems. Network quantization is an efficient way to reduce the network burden. This paper presents a quantized LIC (QLIC) by channel splitting. First, we explore that the influence of quantization error to the reconstruction error is different for various channels. Second, we split the channels whose quantization has larger influence to the reconstruction error. After the splitting, the dynamic range of channels is reduced so that the quantization error can be reduced. Finally, we prune several channels to keep the number of overall channels as origin. By using the proposal, in the case of 8-bit quantization for weight and activation of both main and hyper path, we can reduce the BD-rate by 0.61%-4.74% compared with the previous QLIC. Besides, we can reach better coding gain compared with the state-of-the-art network quantization method when quantizing MS-SSIM models. Moreover, our proposal can be combined with other network quantization methods to further improve the coding gain. The moderate coding loss caused by the quantization validates the feasibility of the hardware implementation for QLIC in the future. Heming Sun, Lu Yu 0003, Jiro Katto |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2024 | SCP: Spherical-Coordinate-Based Learned Point Cloud CompressionabstractIn recent years, the task of learned point cloud compression has gained prominence. An important type of point cloud, LiDAR point cloud, is generated by spinning LiDAR on vehicles. This process results in numerous circular shapes and azimuthal angle invariance features within the point clouds. However, these two features have been largely overlooked by previous methodologies. In this paper, we introduce a model-agnostic method called Spherical-Coordinate-based learned Point cloud compression (SCP), designed to fully leverage the features of circular shapes and azimuthal angle invariance. Additionally, we propose a multi-level Octree for SCP to mitigate the reconstruction error for distant areas within the Spherical-coordinate-based Octree. SCP exhibits excellent universality, making it applicable to various learned point cloud compression techniques. Experimental results demonstrate that SCP surpasses previous state-of-the-art methods by up to 29.14% in point-to-point PSNR BD-Rate. Ao Luo, Linxin Song, Keisuke Nonaka, Kyohei Unno, Heming Sun, Masayuki Goto, Jiro Katto |
AAAI | 5 |
| 2024 | High-Efficiency FPGA - Based Approximate Multipliers with LUT Sharing and Carry SwitchingabstractApproximate multiplier saves energy and improves hardware performance for error-tolerant computation-intensive applications. This work proposes hardware-efficient FPGA-based approximate multipliers with look-up table (LUT) sharing and carry switching. Sharing two LUTs with the same inputs enables to fully utilize the available LUT resources. To mitigate the accuracy loss incurred from this approach, the truncated carry is partially reserved by switching it to the adjacent calculation. In addition, we create a library of 8×8 approximate multipliers to provide various multiplication choices. The proposed design can provide enhancements of up to 38.75% in power, 17.29% in latency, and 28.17% in area compared to the Xilinx exact multiplier. Our proposed designs are open-source at https://github.com/YnuGuoLab/DATE_FPGA_Approx_Mul and assist in further reproducing and development. Qilin Zhou, Xiu Chen, Heming Sun |
DATE | 4 |
| 2024 | Real-Time Video Prediction With Fast Video Interpolation Model and Prediction TrainingabstractTransmission latency significantly affects users’ quality of experience in real-time interaction and actuation. As latency is principally inevitable, video prediction can be utilized to mitigate the latency and ultimately enable zero-latency transmission. However, most of the existing video prediction methods are computationally expensive and impractical for real-time applications. In this work, we therefore propose real-time video prediction towards the zero-latency interaction over networks, called IFRVP (Intermediate Feature Refinement Video Prediction). Firstly, we propose three training methods for video prediction that extend frame interpolation models, where we utilize a simple convolution-only frame interpolation network based on IFRNet. Secondly, we introduce ELAN-based residual blocks into the prediction models to improve both inference speed and accuracy. Our evaluations show that our proposed models perform efficiently and achieve the best trade-off between prediction accuracy and computational speed among the existing video prediction methods. A demonstration movie is also provided at http://bit.ly/IFRVPDemo. Shota Hirose, Kazuki Kotoyori, Kasidis Arunruangsirilert, Fangzheng Lin, Heming Sun, Jiro Katto |
ICIP | 5 |
| 2024 | Privacy-preserving with Flexible Autoencoder for Video Coding for MachinesabstractThe dataset for Video Coding for Machines (VCM) contains sensitive information that requires privacy preservation to address vulnerabilities. Achieving a balance to protect this sensitive data while maintaining VCM performance is crucial. We introduce an autoencoder integrated with a deep learning network that utilizes the ResNet architecture. This design blurs private details while preserving the contours, offering a high-dimensional representation that upholds privacy and VCM performance. The division position between the encoder and decoder is critical, influencing the equilibrium between compression efficacy and machine task performance. We craft a flexible, position-adjustable setting for the autoencoder to optimize this, facilitating a harmonious trade-off between bitrate and mAP across various deep-learning networks. This adaptation demonstrates superior performance relative to existing models. With FasterRCNN, our methods achieve 62.3 of mAP and 5681.29 of bitrate, and their versatility is further validated using YoloV5 and SSD. Aorui Gou, Heming Sun, Xiaoyang Zeng, Yibo Fan |
ISCAS | 2 |
| 2024 | Power-Efficient and Small-Area Approximate Multiplier Design with FPGA-Based CompressorsabstractApproximate computing has become an emerging technique to reduce power consumption. In numerous applications, multiplication is a crucial operation, designing it for approximation is effective in optimizing system performance. In this paper, we propose low-power FPGA-based multipliers by employing the novel compressor designs. Given that the compressor is the primary unit in the multiplier, we introduce novel exact and approximate compressors with low-complexity circuits to parallelly accumulate the elements. To flexibly configure the proposed compressors, a compressor-based once-through structure is proposed to 8×8 multipliers. Two variants of the approximate multipliers are provided with different accuracy-hardware trade-offs. Compared with the exact multiplier, the proposed approximate multiplier reduces power by 57.90%, area by 33.80%, and delay by 24.78%. With a similar accuracy loss, the proposed designs save more hardware resources than others. In addition, the effectiveness of approximate multipliers is assessed in image sharpening. Yi Guo 0010, Xiu Chen, Qilin Zhou, Heming Sun |
ISCAS | 4 |
| 2024 | Accelerating Learnt Video Codecs with Gradient Decay and Layer-Wise DistillationabstractIn recent years, end-to-end learnt video codecs have demonstrated their potential to compete with conventional coding algorithms in term of compression efficiency. However, most learning-based video compression models are associated with high computational complexity and latency, in particular at the decoder side, which limits their deployment in practical applications. In this paper, we present a novel model-agnostic pruning scheme based on gradient decay and adaptive layer-wise distillation. Gradient decay enhances parameter exploration during sparsification whilst preventing runaway sparsity and is superior to the standard Straight-Through Estimation. The adaptive layer-wise distillation regulates the sparse training in various stages based on the distortion of intermediate features. This stage-wise design efficiently updates parameters with minimal computational overhead. The proposed approach has been applied to three popular end-to-end learnt video codecs, FVC, DCVC, and DCVC-HEM. Results confirm that our method yields up to 65% reduction in MACs and 2× speedup with less than 0.3dB drop in BD-PSNR. Supporting code and supplementary material can be downloaded from: https://jasminepp.github.io/lightweighltdvc/. Tianhao Peng 0004, Ge Gao 0005, Heming Sun, Fan Zhang 0017, David Bull 0001 |
PCS | 3 |
| 2024 | Lightweight Stochastic Video Prediction via Hybrid WarpingabstractAccurate video prediction by deep neural networks, especially for dynamic regions, is a challenging task in computer vision for critical applications such as autonomous driving, remote working, and telemedicine. Due to inherent uncertainties, existing prediction models often struggle with the complexity of motion dynamics and occlusions. In this paper, we propose a novel stochastic long-term video prediction model that focuses on dynamic regions by employing a hybrid warping strategy. By integrating frames generated through forward and backward warpings, our approach effectively compensates for the weaknesses of each technique, improving the prediction accuracy and realism of moving regions in videos while also addressing uncertainty by making stochastic predictions that account for various motions. Furthermore, considering real-time predictions, we introduce a MobileNet-based lightweight architecture into our model. Our model, called SVPHW, achieves state-of-the-art performance on two benchmark datasets. Kazuki Kotoyori, Shota Hirose, Heming Sun, Jiro Katto |
VCIP | 3 |
| 2024 | Tell Codec What Worth Compressing: Semantically Disentangled Image Coding for Machine with LMMsabstractWe present a new image compression paradigm to achieve "intelligently coding for machine" by cleverly leveraging the common sense of Large Multimodal Models (LMMs). We are motivated by the evidence that large language/multimodal models are powerful general-purpose semantics predictors for understanding the real world. Different from traditional image compression typically optimized for human eyes, the image coding for machines (ICM) framework we focus on requires the compressed bitstream to more comply with different downstream intelligent analysis tasks. To this end, we employ LMM to${\text{tell codec what to compress}}$: 1) first utilize the powerful semantic understanding capability of LMMs w.r.t object grounding, identification, and importance ranking via prompts, to disentangle image content before compression, 2) and then based on these semantic priors we accordingly encode and transmit objects of the image in order with a structured bitstream. In this way, diverse vision benchmarks including image classification, object detection, instance segmentation, etc., can be well supported with such a semantically structured bitstream. We dub our method "SDComp" for "Semantically Disentangled Compression", and compare it with state-of-the-art codecs on a wide variety of different vision tasks. SDComp codec leads to more flexible reconstruction results, promised decoded visual quality, and a more generic/satisfactory intelligent task-supporting ability. Jinming Liu 0001, Yuntao Wei, Junyan Lin, Shengyang Zhao, Heming Sun, Zhibo Chen 0001, Wenjun Zeng 0001, Xin Jin 0014 |
VCIP | 5 |
| 2024 | LMM-driven Semantic Image-Text Coding for Ultra Low-bitrate Learned Image CompressionabstractSupported by powerful generative models, low-bitrate learned image compression (LIC) models utilizing perceptual metrics have become feasible. Some of the most advanced models achieve high compression rates and superior perceptual quality by using image captions as sub-information. This paper demonstrates that using a large multi-modal model (LMM), it is possible to generate captions and compress them within a single model. We also propose a novel semantic-perceptual-oriented fine-tuning method applicable to any LIC network, resulting in a 41.58% improvement in LPIPS BD-rate compared to existing methods. Our implementation and pre-trained weights are available at https://github.com/tokkiwa/ImageTextCoding. Shimon Murai, Heming Sun, Jiro Katto |
VCIP | 2 |
| 2024 | Variable Bitrate Models For Learned Image Compression with Multi-gain units and Weighted Probability AssignmentabstractWith the advancement of deep learning techniques, learned image compression (LIC) has surpassed traditional compression methods. However, these methods typically require training separate models to achieve optimal rate-distortion performance, leading to increased time and resource consumption. To tackle this challenge, we propose leveraging multi-gain and inverse multi-gain unit pairs to enable variable rate adaptation within a single model. Nevertheless, experiments have shown that rate-distortion performance may degrade at certain bitrates. Therefore, we introduce weighted probability assignment, where different selection probabilities are assigned during training based on lambda values, to increase the model’s training frequency under specific bitrate conditions. To validate our approach, extensive experiments were conducted on Transformer-based and CNN-based models. The experimental results validate the efficiency of our proposed method. Ran Wang 0015, Heming Sun, Jiro Katto |
VCIP | 3 |
| 2024 | Hardware-Efficient Multipliers With FPGA-Based Approximation for Error-Resilient ApplicationsabstractApproximate multipliers enable hardware savings for error-resilient computation-intensive applications. Most existing approximate multipliers have been on ASIC-based circuits. They might not achieve comparable performance gains when used for FPGA-based accelerators. In this paper, we propose hardware-efficient FPGA-based accurate and approximate$4\boldsymbol {\times }4$multipliers with novel methodologies of look-up table (LUT) sharing and carry switching. The LUT resources can be fully utilized by sharing two LUTs with the same inputs. To compensate for the accuracy loss, the truncated carry is partially reserved by switching it to the adjacent calculation. For higher-order multipliers, three approximate adders are proposed to sum the result of the multipliers with arbitrary size. 140 types of$8\boldsymbol {\times }8$multipliers are constructed by combining the proposed$4\boldsymbol {\times }4$multipliers and adders, providing various multiplication choices for different demands. The proposed approximate$8\boldsymbol {\times }8$multiplier can achieve up to 38.75%, 17.29%, and 28.17% improvements in power, latency, and area over the Xilinx exact multiplier, respectively. Moreover, the proposed accurate and approximate$8\boldsymbol {\times }8$multipliers with different adders are extended to$16\boldsymbol {\times }16$multipliers. As evidenced by the performance of the$16\boldsymbol {\times }16$multipliers, our methodology demonstrates the capability to design higher-order multipliers flexibly. Compared with previous works under a similar accuracy loss, the proposed multiplier achieves more hardware savings. Furthermore, the approximate multipliers are assessed on the application of image processing to validate the practical applicability. We create a library of the proposed multipliers which is open-source athttps://github.com/YnuGuoLab/Approx_Mul_FPGA/and assist in further reproducing and development. Qilin Zhou, Xiu Chen, Heming Sun |
IEEE Trans. Circuits Syst. I Regul. Pap. | 4 |
| 2023 | Learned Image Compression with Mixed Transformer-CNN ArchitecturesabstractLearned image compression (LIC) methods have exhibited promising progress and superior rate-distortion performance compared with classical image compression standards. Most existing LIC methods are Convolutional Neural Networks-based (CNN-based) or Transformer-based, which have different advantages. Exploiting both advantages is a point worth exploring, which has two challenges: 1) how to effectively fuse the two methods? 2) how to achieve higher performance with a suitable complexity? In this paper, we propose an efficient parallel Transformer-CNN Mixture (TCM) block with a controllable complexity to incorporate the local modeling ability of CNN and the non-local modeling ability of transformers to improve the overall architecture of image compression models. Besides, inspired by the recent progress of entropy estimation models and attention modules, we propose a channel-wise entropy model with parameter-efficient swin-transformer-based attention (SWAtten) modules by using channel squeezing. Experimental results demonstrate our proposed method achieves state-of-the-art rate-distortion performances on three different resolution datasets (i.e., Kodak, Tecnick, CLIC Professional Validation) compared to existing LIC methods. The code is at https://github.com/jmliu206/LIC_TCM. Jinming Liu 0001, Heming Sun, Jiro Katto |
CVPR | 2 |
| 2023 | Multistage Spatial Context Models for Learned Image CompressionabstractRecent state-of-the-art Learned Image Compression methods feature spatial context models, achieving great rate-distortion improvements over hyperprior methods. However, the autoregressive context model requires serial decoding, limiting run-time performance. The Checkerboard context model allows parallel decoding at a cost of reduced RD performance. We present a series of multistage spatial context models allowing both fast decoding and better RD performance. We split the latent space into square patches and decode serially within each patch while different patches are decoded in parallel. The proposed method features a comparable decoding speed to Checkerboard while reaching the RD performance of Autoregressive and even also outperforming Autoregressive. Inside each patch, the decoding order must be carefully decided as a bad order negatively impacts performance; therefore, we also propose a decoding order optimization algorithm. Fangzheng Lin, Heming Sun, Jinming Liu 0001, Jiro Katto |
ICASSP | 2 |
| 2023 | Recoil: Parallel rANS Decoding with Decoder-Adaptive ScalabilityabstractEntropy coding is essential to data compression, image and video coding, etc. The Range variant of Asymmetric Numeral Systems (rANS) is a modern entropy coder, featuring superior speed and compression rate. As rANS is not designed for parallel execution, the conventional approach to parallel rANS partitions the input symbol sequence and encodes partitions with independent codecs, and more partitions bring extra overhead. This approach is found in state-of-the-art implementations such as DietGPU. It is unsuitable for content-delivery applications, as the parallelism is wasted if the decoder cannot decode all the partitions in parallel, but all the overhead is still transferred. Fangzheng Lin, Kasidis Arunruangsirilert, Heming Sun, Jiro Katto |
ICPP | 3 |
| 2023 | Fast VVC Intra Encoding for Video Coding for MachinesabstractTraditional video coding technologies compress and reconstruct the video frames, which focus on human perception. However, video coding for machines (VCM) uses the feature stream to bridge the correlation between human perception and machine intelligence for vision tasks. We extract the features for the CU with different shapes with part of resnet architecture for VCM. However, the feature-based methods use the model to complete the forward process, which is very time-consuming for its complex architecture and parameter size. The CU architecture for the feature extraction further increases the operation times. A fast algorithm based on the Histogram of oriented gradient (H OG) is proposed for the video coding for machines with VVC intra to overcome the time-consuming problems while maintaining the performance for the vision tasks with codec. The correlation of the mode decision with the VCM performance is discussed to motivate the fast intra coding for V CM. Moreover, the VTM and VVenc are used to verify the universality of the proposed method. The proposed methods can speed up the fast encoding for 35.21 % time saving with 0.26 increment for AP50 for the cityscapes dataset compared with the VTM10.0. Aorui Gou, Heming Sun, Xiaoyang Zeng, Yibo Fan |
ISCAS | 2 |
| 2023 | PTS-LIC: Pruning Threshold Searching for Lightweight Learned Image CompressionabstractLearned Image Compression (LIC), which uses neural networks to compress images, has experienced significant growth in recent years. The hyperprior-module-based LIC model has achieved higher performance than classical codecs. However, the LIC models are too heavy (in calculation and parameter amounts) to apply to edge devices. To solve this problem, some former papers focus on structural pruning for LIC models. However, they either cause noticeable performance decrement or neglect the appropriate pruning threshold for each LIC model. These problems keep their pruning results sub-optimal. This paper proposes a Pruning Threshold Searching on the hyperprior module for different-quality LIC models. Our method removes most parameters and calculations while keeping the performance the same as the models before pruning. We removed at least 49.8% of parameters and 28.5% of calculations for the Channel-Wise-Context-Model-based models and 29.1% of parameters for the Cheng-2020 models. Ao Luo, Heming Sun, Jinming Liu 0001, Fangzheng Lin, Jiro Katto |
VCIP | 2 |
| 2023 | A novel fast intra algorithm for VVC based on histogram of oriented gradientabstractThe latest Versatile Video Coding (VVC) standard incorporates a series of effective and complex new intra coding tools, which obtains superior coding efficiency than the High Efficiency Video Coding (HEVC). However, this makes the intra coding more complicated and time-consuming. A fast algorithm for VVC from two aspects of fast mode decision and fast partition decision is proposed in this paper. For the fast mode decision, the relationship between bins with Histogram of Oriented Gradient (HOG) and intra modes is created for the mode selection, decreasing the planar modes for SATD and RDO. Moreover, we analyze the maximum bins to determine the final modes, and we use the modes of left and upper blocks as a reference for the current CU, which can early terminate RDO. Moreover, a two-step fast partition algorithm is proposed based on HOG for fast partition decision, in which two thresholds are investigated to control the uniformity of textures. The proposed fast algorithm is implemented on the VVC test model, and the experimental results show that it can achieve 69.07% time savings with only 2.96% BDBR increases averagely, which outperforms other relatively existing state-of-the-art methods. Moreover, to convince the universality of our algorithm, we further implement our method in Fraunhofer Versatile Video Encoder (VVenc) and Fraunhofer Versatile Video Decoder (VVdec), which have five settings to control the trade-off between encoding quality and efficiency for intra coding. The fast intra mode decision algorithm and fast partition algorithm decrease the complexity of intra coding for both VTM and VVenc, which shows the efficiency and universality of the proposed fast partition and fast mode decision algorithms. Aorui Gou, Heming Sun, Chao Liu 0027, Xiaoyang Zeng, Yibo Fan |
J. Vis. Commun. Image Represent. | 2 |
| 2023 | A Reconfigurable Multiple Transform Selection Architecture for VVCabstractVideo coding plays an important role in the highly information-based world as videos contribute the largest part of network traffic. The latest video coding standard Versatile Video Coding (VVC) introduces a new transform scheme multiple transform selection (MTS), which brings considerable coding gains at the expense of high coding complexity. In this article, we propose a reconfigurable MTS architecture that supports all transform types in VVC with square and rectangular sizes ranging from$4\times $4 to 32$\times32$. Firstly, we explore the features of three types of transform matrices and extract the features that are beneficial to designing a unified architecture. Then, we present an improved calculation scheme for general transforms, where the transform matrix is decomposed into two simpler matrices to increase the similarity and decrease the complexity of matrices involved in three types of transform operations. Thanks to the improved calculated scheme, a unified shift-adder unit (SAU) is designed and highly reused by different types. Moreover, we provide a twirling two-point splicing (T2S) scheme to improve reusability and deal with issues of data mismatch when conducting discrete cosine transform (DCT)-II of different sizes. As a consequence, an architecture with constant throughput of 32 pixels/cycle is implemented and specified in Verilog HDL. The synthesis results indicate that the application specific integrated circuit (ASIC)-based and field-programmable gate array (FPGA)-based hardware architectures achieve significant advantages both in area reduction and power consumption compared to existing methods in the literature. Zhijian Hao, Heming Sun, Guoqing Xiang, Peng Zhang 0007, Xiaoyang Zeng, Yibo Fan |
IEEE Trans. Very Large Scale Integr. Syst. | 2 |
| 2022 | Learned Video Compression With Residual Prediction And Feature-Aided Loop FilterabstractIn this paper, we propose a learned video codec with a residual prediction network (RP-Net) and a feature-aided loop filter (LF-Net). For the RP-Net, we exploit the residual of previous multiple frames to further eliminate the redundancy of the current frame residual. For the LF-Net, the features from residual decoding network and the motion compensation network are used to aid the reconstruction quality. To reduce the complexity, a light ResNet structure is used as the backbone for both RP-Net and LF-Net. Experimental results illustrate that we can save about 10% BD-rate compared with previous learned video compression frameworks. Moreover, we can achieve faster coding speed due to the ResNet backbone. Chao Liu 0027, Heming Sun, Xiaoyang Zeng, Yibo Fan |
ICIP | 2 |
| 2022 | Streaming-Capable High-Performance Architecture of Learned Image Compression CodecsabstractLearned image compression allows achieving state-of-the-art accuracy and compression ratios, but their relatively slow runtime performance limits their usage. While previous attempts on optimizing learned image codecs focused more on the neural model and entropy coding, we present an alternative method to improving the runtime performance of various learned image compression models. We introduce multi-threaded pipelining and an optimized memory model to enable GPU and CPU workloads’ asynchronous execution, fully taking advantage of computational resources. Our architecture alone already produces excellent performance without any change to the neural model itself. We also demonstrate that combining our architecture with previous tweaks to the neural models can further improve runtime performance. We show that our implementations excel in throughput and latency compared to the baseline and demonstrate the performance of our implementations by creating a real-time video streaming encoder-decoder sample application, with the encoder running on an embedded device. Fangzheng Lin, Heming Sun, Jiro Katto |
ICIP | 2 |
| 2022 | Memory-Efficient Learned Image Compression with Pruned Hyperprior ModuleabstractLearned Image Compression (LIC) gradually became more and more famous in these years. The hyperprior-module-based LIC models have achieved remarkable rate-distortion performance. However, the memory cost of these LIC models is too large to actually apply them to various devices, especially to portable or edge devices. The parameter scale is directly linked with memory cost. In our research, we found the hyperprior module is not only highly over-parameterized, but also its latent representation contains redundant information. Therefore, we propose a novel pruning method named ERHP in this paper to efficiently reduce the memory cost of hyperprior module, while improving the network performance. The experiments show our method is effective, reducing at least 22.6% parameters in the whole model while achieving better rate-distortion performance. Ao Luo, Heming Sun, Jinming Liu 0001, Jiro Katto |
ICIP | 2 |
| 2022 | Improving Multiple Machine Vision Tasks in the Compressed DomainabstractThere is a growing number of images that are analyzed by machines rather than just humans. Recently, most machine vision tasks are based on decoded images which require an image compression (encoding/decoding) framework. However, using the decoded images in the pixel-domain has two drawbacks: 1) the complexity is high for the decoder part, 2) the accuracy (e.g., mIoU, mean absolute error, and average precision) of machine vision tasks will be degraded since decoded images only aim to optimize the human perceived quality (e.g., PSNR) so that information required for machine vision tasks will be lost during the decoding process. In this paper, we improve the machine vision tasks in the compressed domain. 1) A gate module is utilized to effectively select some compressed-domain features. 2) Knowledge distillation is introduced to improve the accuracy. 3) A training strategy is explored to support multiple tasks including the image compression. The experimental results show that we can achieve better rate-accuracy/distortion and lower complexity compared with the state-of-the-art pixel-domain work that can take both machine and human vision tasks. Jinming Liu 0001, Heming Sun, Jiro Katto |
ICPR | 2 |
| 2022 | Fast Intra Mode Decision for VVC Based on Histogram of Oriented GradientabstractThe latest Versatile Video Coding (VVC) standard incorporates a series of effective and complex new intra coding tools, which obtains superior coding efficiency than the High Efficiency Video Coding (HEVC). However, this makes the intra coding more complicated and time-consuming. A fast algorithm for VVC is proposed from two aspects of model selection and early terminating to reduce coding complexity in this paper. The relationship between bins with HOG and intra modes is created for the mode selection, decreasing the planar modes for SATD and RDO. Moreover, we analyze the maximum bins to determine the final modes, and we use the modes of left and upper blocks as a reference for the current CU, which can early terminate RDO. The proposed algorithm is implemented on VVC test model, and the experimental results show that it can achieve 36.61% time savings with only 0.94% BDBR increases averagely, which outperforms other relative existing state-of-the-art methods. Aorui Gou, Heming Sun, Jiro Katto, Xiaoyang Zeng, Yibo Fan |
ISCAS | 2 |
| 2022 | An Area-efficient Unified Transform Architecture for VVCabstractThe next-generation video coding standard Versatile Video Coding (VVC) adopts Multiple Transform Selection (MTS) to the transform module, improving coding efficiency at the expense of high computational complexity. Compared to High Efficiency Video Coding (HEVC), VVC supports larger sizes and extends the transform types to Discrete Cosine Transform (DCT)-II, Discrete Sine Transform (DST)-VII, and DCT-VIII. This paper presents an area-efficient unified architecture for VVC. To reduce the area consumption, we propose an optimized calculation scheme for general transformations where the transform matrix is decomposed into two simpler matrices named the Low-value matrix and the Error matrix. Based on the decomposition algorithm, Shift-Addition Units (SAUs)-based circuits are designed to conduct matrix multiplication and can be reused by three types. As a result, this unified architecture is capable of performing all types and sizes in VVC. The synthesis results indicate that this architecture achieves an area reduction of 37.9% $\sim$ 72.2% compared with related works for 32-point transforms. Zhijian Hao, Qi Zheng 0004, Yibo Fan, Guoqing Xiang, Peng Zhang 0007, Heming Sun |
ISCAS | 6 |
| 2022 | A QP-adaptive Mechanism for CNN-based Filter in Video CodingabstractConvolutional neural network (CNN)-based in-loop filtering have been very successful in video coding. For most existing works, however, a specific model was required for each quantization parameter (QP) band. In this paper, we introduce a generic method for helping CNN-filters deal with variable quantization noises. A feasible solution to this problem can be implemented on CNN by introducing a quantization step (Qstep) into the CNN. As the quantization noise changes, the CNN filter’s ability to suppress noise changes accordingly. The (vanilla) convolution layer can be replaced directly by this method in existing CNN filters. Compared with the VVenC anchor, only one CNN filter is used and achieves about 3.6% BD-rate reduction for the luminance component of random-access configuration. Also, about 0.8% BD-rate reduction has been achieved compared with the previous QP-map method. Chao Liu 0027, Heming Sun, Jiro Katto, Xiaoyang Zeng, Yibo Fan |
ISCAS | 2 |
| 2022 | Project-Based Learning: Bridging the Gap Between Algorithm and Architecture in Neural Network CourseabstractNeural network has shown its powerful ability in many research fields in the recent years. By using different network structures, many new algorithms are developed to enhance the accuracy. Along with the algorithm development, corresponding architectures are also proposed for the acceleration. However, pure algorithm may not be hardware friendly. As a result, we need to find an optimal trade-off between algorithmic accuracy and architectural efficiency. To help students build the gap between algorithm and architecture, this paper introduces a project-based learning. The project is called learned image compression, which is composed of three phases: algorithm design, architecture mapping and algorithm-architecture co-optimization. Through the project, the students are expected to develop a neural network with high image compression ratio and hardware performance. Furthermore, these kind of knowledge can be extended to any neural network applications. Heming Sun, Lu Yu 0003 |
ISCAS | 1 |
| 2022 | Optimizing CABAC architecture with prediction based context model prefetchingabstractContext Adaptive Binary Arithmetic Coding (CABAC) is the entropy coding module widely used in recent video coding standards such as HEVC/H.265 and VVC/H.266. CABAC is a well-known throughput bottleneck due to its strong data dependencies. Because the required context model of the current bin often depends on the results of the previous bin, the context model cannot be prefetched early enough, and then costs pipeline stalls. To solve this problem, we propose a prediction-based context model prefetching strategy. If the prediction is correct, pipeline stalls can be eliminated, and the stalling cycles won't get worse with the wrong prediction. Moreover, the data interaction process between CABAC modules and the multi-stage pipeline structure are optimized to maximize the working frequency. The proposed pipeline architecture can reduce pipeline stalls and save up to 45.66% encoding time, the improved results show that it provides more significant gains in All Intra (AI) under low QP test conditions, which is better than the Random Access (RA) and Low Delay (LD) configuration. The highest hardware efficiency (Mbins/s Per k gates) is higher than the existing advanced pipeline architecture. Heming Sun, Jiayao Xu, Jinjia Zhou |
MMSP | 2 |
| 2022 | Semantic Segmentation In Learned Compressed DomainabstractMost machine vision tasks (e.g., semantic segmentation) are based on images encoded and decoded by image compression algorithms (e.g., JPEG). However, these decoded images in the pixel domain introduce distortion, and they are optimized for human perception, making the performance of machine vision tasks suboptimal. In this paper, we propose a method based on the compressed domain to improve segmentation tasks. i) A dynamic and a static channel selection method are proposed to reduce the redundancy of compressed representations that are obtained by encoding. ii) Two different transform modules are explored and analyzed to help the compressed representation be transformed as the features in the segmentation network. The experimental results show that we can save up to 15.8% bitrates compared with a state-of-the-art compressed domain-based work while saving up to about 83.6% bitrates and 44.8% inference time compared with the pixel domain-based method. Jinming Liu 0001, Heming Sun, Jiro Katto |
PCS | 2 |
| 2022 | Learning from the NN-based Compressed Domain with Deep Feature Reconstruction LossabstractTo speedup the image classification process which conventionally takes the reconstructed images as input, compressed domain methods choose to use the compressed images without decompression as input. Correspondingly, there will be a certain decline about the accuracy. Our goal in this paper is to raise the accuracy of compressed domain classification method using compressed images output by the NN-based image compression networks. Firstly, we design a hybrid objective loss function which contains the reconstruction loss of deep feature map. Secondly, one image reconstruction layer is inte-grated into the image classification network for up-sampling the compressed representation. These methods greatly help increase the compressed domain image classification accuracy and need no extra computational complexity. Experimental results on the benchmark ImageNet prove that our design outperforms the latest work ResNet-41 with a large accuracy gain, about 4.49% on the top-1 classification accuracy. Besides, the accuracy lagging behinds the method using reconstructed images is also reduced to 0.47 %. Moreover, our designed classification network has the lowest computational complexity and model complexity. Liuhong Chen, Heming Sun, Xiaoyang Zeng, Yibo Fan |
VCIP | 2 |
| 2022 | On Pre-chewing Compression Degradation for Learned Video CompressionabstractArtificial Intelligence (AI) needs huge amounts of data, and so does Learned Restoration for Video Compression. There are two main problems regarding training data. 1) Preparing training compression degradation using a video codec (e.g., Versatile Video Coding - VVC) costs a considerable resource. Significantly, the more Quantization Parameters (QPs) we compress with, the more coding time and storage are required. 2) The common way of training a newly initialized Restoration Network on pure compression degradation at the beginning is not effective. To solve these problems, we propose a Degradation Network to pre-chew (generalize and learn to synthesize) the real compression degradation, then present a hybrid training scheme that allows a Restoration Network to be trained on unlimited videos without compression. Concretely, we propose a QP-wise Degradation Network to learn how to compress video frames like VVC in real-time and can transform the degradation output between QPs linearly. The real compression degradation is thus pre-chewed as our Degradation Network can synthesize the more generalized degradation for a newly initialized Restoration Network to learn easier. To diversify training video content without compression and avoid overfitting, we design a Training Framework for Semi-Compression Degradation (TF-SCD) to train our model on many fake compressed videos together with real compressed videos. As a result, the Restoration Network can quickly jump to the near-best optimum at the beginning of training, proving our promising scheme of using pre-chewed data for the very first steps of training. In other words, a newly initialized Learned Video Compression can be warmed up efficiently but effectively with our pre-trained Degradation Network. Besides, our proposed TF-SCD can further enhance the restoration performance in a specific range of QPs and provide a better generalization about QPs compared with the common way of training a restoration model. Our work is available at https://minhmanho.github.io/prechewing_degradation. Man M. Ho, Heming Sun, Jinjia Zhou |
VCIP | 2 |
| 2022 | Improving Latent Quantization of Learned Image Compression with Gradient ScalingabstractLearned image compression (LIC) has shown its superior compression ability. Quantization is an inevitable stage to generate quantized latent for the entropy coding. To solve the non-differentiable problem of quantization in the training phase, many differentiable approximated quantization methods have been proposed. However, the derivative of quantized latent to non-quantized latent are set as one in most of the previous methods. As a result, the quantization error between non-quantized and quantized latent is not taken into consideration in the gradient descent. To address this issue, we exploit the gradient scaling method to scale the gradient of non-quantized latent in the back-propagation. The experimental results show that we can outperform the recent LIC quantization methods. Heming Sun, Lu Yu 0003, Jiro Katto |
VCIP | 1 |
| 2022 | Real-time Learned Image Codec on FPGAabstractThis demo paper gives a real-time learned image codec on FPGA. By using Xilinx VCU128, the proposed system reaches 720P@30fps codec, which is 7.76x faster than prior work. Heming Sun, Qingyang Yi, Fangzheng Lin, Lu Yu 0003, Jiro Katto |
VCIP | 1 |
| 2022 | A method of underwater bridge structure damage detection method based on a lightweight deep convolutional networkabstractAbstract The problem of the underwater structure disease of the bridge is increasingly obvious, which has seriously affected the safe operation of the bridge structure, so it is necessary to detect the underwater structure regularly. There are many kinds of bridge underwater structure diseases. This paper targets the bridge underwater structural crack diseases adopts multiple image recognition networks for verification, compares the advantages of different networks, and takes the YOLO‐v4 network as the main body to build a lightweight convolutional neural network.Mobilenetv3 replaced CSPDarkent as the backbone feature extraction network, while the feature layer scale of Mobilenetv3 was modified, and the extracted preliminary feature layer was input into the enhanced feature extraction network for feature fusion. The PANet networks are replaced by the depthwise separable convolution. Using ablation experiments to compare the performance of four algorithm combinations in lightweight networks. At the same time, the disease identification accuracy of each network and the performance of the network are tested in various experimental environments, and the feasibility of the lightweight network is verified in the application of bridge underwater structure damage identification. Heming Sun, Taiyi Song, Qinghang Meng |
IET Image Process. | 2 |
| 2022 | Deep image compression based on multi-scale deformable convolution
Daowen Li, Yingming Li, Heming Sun, Lu Yu 0003 |
J. Vis. Commun. Image Represent. | 3 |
| 2022 | QA-Filter: A QP-Adaptive Convolutional Neural Network Filter for Video CodingabstractConvolutional neural network (CNN)-based filters have achieved great success in video coding. However, in most previous works, individual models were needed for each quantization parameter (QP) band, which is impractical due to limited storage resources. To explore this, our work consists of two parts. First, we propose a frequency and spatial QP-adaptive mechanism (FSQAM), which can be directly applied to the (vanilla) convolution to help any CNN filter handle different quantization noise. From the frequency domain, a FQAM that introduces the quantization step (Qstep) into the convolution is proposed. When the quantization noise increases, the ability of the CNN filter to suppress noise improves. Moreover, SQAM is further designed to compensate for the FQAM from the spatial domain. Second, based on FSQAM, a QP-adaptive CNN filter called QA-Filter that can be used under a wide range of QP is proposed. By factorizing the mixed features to high-frequency and low-frequency parts with the pair of pooling and upsampling operations, the QA-Filter and FQAM can promote each other to obtain better performance. Compared to the H.266/VVC baseline, average 5.25% and 3.84% BD-rate reductions for luma are achieved by QA-Filter with default all-intra (AI) and random-access (RA) configurations, respectively. Additionally, an up to 9.16% BD-rate reduction is achieved on the luma of sequence BasketballDrill. Besides, FSQAM achieves measurably better BD-rate performance compared with the previous QP map method. Chao Liu 0027, Heming Sun, Jiro Katto, Xiaoyang Zeng, Yibo Fan |
IEEE Trans. Image Process. | 2 |
| 2021 | COUGH: A Challenge Dataset and Models for COVID-19 FAQ RetrievalabstractWe present a large, challenging dataset, COUGH, for COVID-19 FAQ retrieval.Similar to a standard FAQ dataset, COUGH consists of three parts: FAQ Bank, Query Bank and Relevance Set.The FAQ Bank contains ∼16K FAQ items scraped from 55 credible websites (e.g., CDC and WHO).For evaluation, we introduce Query Bank and Relevance Set, where the former contains 1,236 human-paraphrased queries while the latter contains ∼32 humanannotated FAQ items for each query.We analyze COUGH by testing different FAQ retrieval models built on top of BM25 and BERT, among which the best model achieves 48.8 under P@5, indicating a great challenge presented by COUGH and encouraging future research for further improvement.Our COUGH dataset is available at https://github. com/sunlab-osu/covid-faq. *Work was done when the first two authors were at OSU. 1 q and a are question and answer fields in an FAQ item.Question1: Should children wear masks?Answer1: In general, children 2 years and older should wear a mask...Appropriate and consistent use of masks...FAQ Bank Question2: Coping with Self-Quarantine Answer2: Remind yourself that difficult emotions are normal during self-quarantine... Query1: Is it possible for human beings to get sick with COVID-19 transmitted to them from animals?Query2: Is it possible to get infected by COVID 19 if I touch food surface packaging?Query Bank Question3: COVID-19是如何在⼈与⼈之间传播的? (How does COVID-19 spread between people?) Answer3: . Xinliang Frederick Zhang, Heming Sun, Xiang Yue, Simon M. Lin, Huan Sun 0001 |
EMNLP (1) | 2 |
| 2021 | Deep Pedestrian Density Estimation For Smart City MonitoringabstractRecently, requirement of city monitoring and maintenance using ICT techniques increases with the help of transportation system. In addition, the spread of COVID-19 has increased the demand for managing pedestrian traffic volume. To contribute to these trends, in this paper, we propose a new pedestrian radar map system in order to estimate pedestrian density on streets and sidewalks. Our system uses e-bikes to collect 360-degree images and visualize pedestrian positions as a radar map. In evaluations, we confirm the accuracies of the radar maps and pedestrian density by using KITTI dataset and by carrying out a field experiment. Kazuki Murayama, Kenji Kanai, Masaru Takeuchi, Heming Sun, Jiro Katto |
ICIP | 4 |
| 2021 | Approximated Reconfigurable Transform Architecture for VVCabstractAs the demand for high-resolution videos grows, the next generation video coding standard Versatile Video Coding introduces many new proposals, including Adaptive Multiple Transforms (AMT), to improve coding efficiency. This paper presents a reconfigurable transform core for the VVC standard where the implementation of 1D DST-VII and DCT-VIII for all transform sizes are enabled. To offer a very low circuit complexity, a simple approximation strategy with a little coding performance loss is proposed. An 8x8 Processing Element (PE) array is employed as the core computational unit, where each PE can be configured dynamically based on the transform type. In addition, the transforms of larger sizes can be realized in the finite PE units with the Partitioned Matrix Multiplication (PMM) scheme. The experimental and synthesis results show that this design can save at least 29.1% area compared with other works in literature with the negligible degradation of video quality and a slight increase in the bit rate. Yixuan Zeng, Heming Sun, Jiro Katto, Yibo Fan |
ISCAS | 2 |
| 2021 | Accelerating Convolutional Neural Network Inference Based on a Reconfigurable Sliced Systolic ArrayabstractConvolutional neural networks (CNNs) have achieved great successes on many computer vision tasks, such as image recognition, video processing, and target detection. In recent years, many hardware designs have been devoted to accelerating CNN inference. In order to further speed up CNN inference and reduce data waste, this work proposed a reconfigurable sliced systolic array: 1) Depending on the number of network nodes in each layer, the slice mode could be dynamically configured to achieve high throughput and resource utilization. 2) To take full advantage of convolution reuse and weight reuse, this work designed a tile-column sliding (TCS) processing dataflow. 3) A four-stage for loop algorithm was employed, which divides the CNN calculation into several parts based on the input nodes and output nodes. The entire CNN inference is carried out using integer-only arithmetic originated from TensorLite. Experimental results prove that these strategies lead to significant improvement in inference performance and energy efficiency. Yixuan Zeng, Heming Sun, Jiro Katto, Yibo Fan |
ISCAS | 2 |
| 2021 | Learned Image Compression with Fixed-point ArithmeticabstractLearned image compression (LIC) has achieved superior coding performance than traditional image compression standards such as HEVC intra in terms of both PSNR and MS-SSIM. However, most LIC frameworks are based on floating-point arithmetic which has two potential problems. First is that using traditional 32-bit floating-point will consume huge memory and computational cost. Second is that the decoding might fail because of the floating-point error coming from different encoding/decoding platforms. To solve the above two problems. 1) We linearly quantize the weight in the main path to 8-bit fixed-point arithmetic, and propose a fine tuning scheme to reduce the coding loss caused by the quantization. Analysis transform and synthesis transform are fine tuned layer by layer. 2) We exploit look-up-table (LUT) for the cumulative distribution function (CDF) to avoid the floating-point error. When the latent node follows non-zero mean Gaussian distribution, to share the CDF LUT for different mean values, we restrict the range of latent node to be within a certain range around mean. As a result, 8-bit weight quantization can achieve negligible coding gain loss compared with 32-bit floating-point anchor. In addition, proposed CDF LUT can ensure the correct coding at various CPU and GPU hardware platforms. Heming Sun, Lu Yu 0003, Jiro Katto |
PCS | 1 |
| 2021 | Learning in Compressed Domain for Faster Machine Vision TasksabstractLearned image compression (LIC) has illustrated good ability for reconstruction quality driven tasks (e.g. PSNR, MS-SSIM) and machine vision tasks such as image understanding. However, most LIC frameworks are based on pixel domain, which requires the decoding process. In this paper, we develop a learned compressed domain framework for machine vision tasks. 1) By sending the compressed latent representation directly to the task network, the decoding computation can be eliminated to reduce the complexity. 2) By sorting the latent channels by entropy, only selective channels will be transmitted to the task network, which can reduce the bitrate. As a result, compared with the traditional pixel domain methods, we can reduce about 1/3 multiply-add operations (MACs) and 1/5 inference time while keeping the same accuracy. Moreover, proposed channel selection can contribute to at most 6.8% bitrate saving. Jinming Liu 0001, Heming Sun, Jiro Katto |
VCIP | 2 |
| 2020 | Small-Area and Low-Power FPGA-Based Multipliers using Approximate Elementary ModulesabstractApproximate multiplier design is an effective technique to improve hardware performance at the cost of accuracy loss. The current approximate multipliers are mostly ASIC-based and are dedicated for one particular application. In contrast, FPGA has been an attractive choice for many applications, because of its high performance, reconfigurability, and fast development. This paper presents a novel methodology for designing approximate multipliers by employing the FPGA-based fabrics. The area and latency are significantly reduced by cutting the carry propagation path in the multiplier. Moreover, we explore higher-order multipliers on architectural space by using our proposed small-size approximate multipliers as elementary modules. For different accuracy requirements, eight configurations for approximate 8 × 8 multiplier are discussed. In terms of mean relative error distance (MRED), the accuracy loss of the proposed 8 × 8 multiplier is low as 0.17%. Compared with the exact multiplier, our proposed design can reduce area by 43.66% and power by 20.36%. The critical path latency reduction is up to 27.66%. The proposed multiplier design has a better accuracy-hardware tradeoff than other designs with com-parable accuracy. Yi Guo 0010, Heming Sun, Shinji Kimura |
ASP-DAC | 2 |
| 2020 | Learned Image Compression With Discretized Gaussian Mixture Likelihoods and Attention ModulesabstractImage compression is a fundamental research field and many well-known compression standards have been developed for many decades. Recently, learned compression methods exhibit a fast development trend with promising results. However, there is still a performance gap between learned compression algorithms and reigning compression standards, especially in terms of widely used PSNR metric. In this paper, we explore the remaining redundancy of recent learned compression algorithms. We have found accurate entropy models for rate estimation largely affect the optimization of network parameters and thus affect the rate-distortion performance. Therefore, in this paper, we propose to use discretized Gaussian Mixture Likelihoods to parameterize the distributions of latent codes, which can achieve a more accurate and flexible entropy model. Besides, we take advantage of recent attention modules and incorporate them into network architecture to enhance the performance. Experimental results demonstrate our proposed method achieves a state-of-the-art performance compared to existing learned compression methods on both Kodak and high-resolution datasets. To our knowledge our approach is the first work to achieve comparable performance with latest compression standard Versatile Video Coding (VVC) regarding PSNR. More importantly, our approach generates more visually pleasant results when optimized by MS-SSIM. Zhengxue Cheng, Heming Sun, Masaru Takeuchi, Jiro Katto |
CVPR | 2 |
| 2020 | Learned Lossless Image Compression with A Hyperprior and Discretized Gaussian Mixture LikelihoodsabstractLossless image compression is an important task in the field of multimedia communication. Traditional image codecs typically support lossless mode, such as WebP, JPEG2000, FLIF. Recently, deep learning based approaches have started to show the potential at this point. HyperPrior is an effective technique proposed for lossy image compression. This paper generalizes the hyperprior from lossy model to lossless compression, and proposes a L2-norm term into the loss function to speed up training procedure. Besides, this paper also investigated different parameterized models for latent codes, and propose to use Gaussian mixture likelihoods to achieve adaptive and flexible context models. Experimental results validate our method can outperform existing deep learning based lossless compression, and outperform the JPEG2000 and WebP for JPG images. Zhengxue Cheng, Heming Sun, Masaru Takeuchi, Jiro Katto |
ICASSP | 2 |
| 2020 | Scalable Learned Image Compression With A Recurrent Neural Networks-Based HyperpriorabstractRecently learned image compression has achieved many great progresses, such as representative hyperprior and its variants based on convolutional neural networks (CNNs). However, CNNs are not fit for scalable coding and multiple models need to be trained separately to achieve variable rates. In this paper, we incorporate differentiable quantization and accurate entropy models into recurrent neural networks (RNNs) architectures to achieve a scalable learned image compression. First, we present an RNN architecture with quantization and entropy coding. To realize the scalable coding, we allocate the bits to multiple layers, by adjusting the layer-wise lambda values in Lagrangian multiplier-based rate-distortion optimization function. Second, we add an RNN-based hyperprior to improve the accuracy of entropy models for multiple-layer residual representations. Experimental results demonstrate that our performance can be comparable with recent CNN-based hyperprior methods on Kodak dataset. Besides, our method is a scalable and flexible coding approach, to achieve multiple rates using one single model, which is very appealing. Rige Su, Zhengxue Cheng, Heming Sun, Jiro Katto |
ICIP | 3 |
| 2020 | End-To-End Learned Image Compression With Fixed Point Weight QuantizationabstractLearned image compression (LIC) has reached the traditional hand-crafted methods such as JPEG2000 and BPG in terms of the coding gain. However, the large model size of the network prohibits the usage of LIC on resource-limited embedded systems. This paper presents a LIC with 8-bit fixed-point weights. First, we quantize the weights in groups and propose a non-linear memory-free codebook. Second, we explore the optimal grouping and quantization scheme. Finally, we develop a novel weight clipping fine tuning scheme. Experimental results illustrate that the coding loss caused by the quantization is small, while around 75% model size can be reduced compared with the 32-bit floating-point anchor. As far as we know, this is the first work to explore and evaluate the LIC fully with fixed-point weights, and our proposed quantized LIC is able to outperform BPG in terms of MS-SSIM. Heming Sun, Zhengxue Cheng, Masaru Takeuchi, Jiro Katto |
ICIP | 1 |
| 2020 | Fully Neural Network Mode Based Intra Prediction of Variable Block SizeabstractIntra prediction is an essential component in the image coding. This paper gives an intra prediction framework completely based on neural network modes (NM). Each NM can be regarded as a regression from the neighboring reference blocks to the current coding block. (1) For variable block size, we utilize different network structures. For small blocks 4×4 and 8×8, fully connected networks are used, while for large blocks 16×16 and 32×32, convolutional neural networks are exploited. (2) For each prediction mode, we develop a specific pre-trained network to boost the regression accuracy. When integrating into HEVC test model, we can save 3.55%, 3.03% and 3.27% BD-rate for Y, U, V components compared with the anchor. As far as we know, this is the first work to explore a fully NM based framework for intra prediction, and we reach a better coding gain with a lower complexity compared with the previous work. Heming Sun, Lu Yu 0003, Jiro Katto |
VCIP | 1 |
| 2020 | A Pipelined 2D Transform Architecture Supporting Mixed Block Sizes for the VVC StandardabstractFor the next-generation video coding standard Versatile Video Coding (VVC), several new contributions have been proposed to improve the coding efficiency, especially in the transformation operations. This paper proposes a unified $32\times 32$ block-based transform architecture for the VVC standard that enables 2D Discrete Sine Transform-VII (DST-VII) and Discrete Cosine Transform-VIII (DCT-VIII) of all sizes. It mainly gives three contributions: 1) The N-Dimensional Reduced Adder Graph (RAG-n) algorithm is adopted to design the minimal adder-oriented computational units. 2) The storage of the asymmetric transform units can be realized in the dual-port SRAM-based transpose memory. 3) The pipelined 2D transformations of mixed block sizes are achieved with the throughput rate of 32 samples per cycle. The synthesis results indicate that this architecture can reduce area by up to 73.1% compared with other state-of-the-art works. Moreover, power saving ranging from 4.9% to 9.9% can be achieved. Regarding the transpose memory, at least 21.9% of the area can be saved by using SRAM. Yibo Fan, Yixuan Zeng, Heming Sun, Jiro Katto, Xiaoyang Zeng |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2020 | Energy Compaction-Based Image Compression Using Convolutional AutoEncoderabstractImage compression has been an important research topic for many decades. Recently, deep learning has achieved great success in many computer vision tasks, and its use in image compression has gradually been increasing. In this paper, we present an energy compaction-based image compression architecture using a convolutional autoencoder (CAE) to achieve high coding efficiency. Our main contributions include three aspects: 1) we propose a CAE architecture for image compression by decomposing it into several down(up)sampling operations; 2) for our CAE architecture, we offer a mathematical analysis on the energy compaction property and we are the first work to propose a normalized coding gain metric in neural networks, which can act as a measurement of compression capability; 3) based on the coding gain metric, we propose an energy compaction-based bit allocation method, which adds a regularizer to the loss function during the training stage to help the CAE maximize the coding gain and achieve high compression efficiency. The experimental results demonstrate our proposed method outperforms BPG (HEVC-intra), in terms of the MS-SSIM quality metric. Additionally, we achieve better performance in comparison with existing bit allocation methods, and provide higher coding efficiency compared with state-of-the-art learning compression methods at high bit rates. Zhengxue Cheng, Heming Sun, Masaru Takeuchi, Jiro Katto |
IEEE Trans. Multim. | 2 |
| 2020 | Enhanced Intra Prediction for Video Coding by Using Multiple Neural NetworksabstractThis paper enhances the intra prediction by using multiple neural network modes (NM). Each NM serves as an end-to-end mapping from the neighboring reference blocks to the current coding block. For the provided NMs, we present two schemes (appending and substitution) to integrate the NMs with the traditional modes (TM) defined in high efficiency video coding (HEVC). For the appending scheme, each NM is corresponding to a certain range of TMs. The categorization of TMs is based on the expected prediction errors. After determining the relevant TMs for each NM, we present a probability-aware mode signaling scheme. The NMs with higher probabilities to be the best mode are signaled with fewer bits. For the substitution scheme, we propose to replace the highest and lowest probable TMs. New most probable mode (MPM) generation method is also employed when substituting the lowest probable TMs. Experimental results demonstrate that using multiple NMs will improve the coding efficiency apparently compared with the single NM. Specifically, proposed appending scheme with seven NMs can save 2.6%, 3.8%, and 3.1% BD-rate for Y, U, and V components compared with using single NM in the state-of-the-art works. Heming Sun, Zhengxue Cheng, Masaru Takeuchi, Jiro Katto |
IEEE Trans. Multim. | 1 |
| 2019 | Learning Image and Video Compression Through Spatial-Temporal Energy CompactionabstractCompression has been an important research topic for many decades, to produce a significant impact on data transmission and storage. Recent advances have shown a great potential of learning based image and video compression. Inspired from related works, in this paper, we present an image compression architecture using a convolutional autoencoder, and then generalize image compression to video compression, by adding an interpolation loop into both encoder and decoder sides. Our basic idea is to realize spatial-temporal energy compaction in learning image and video compression. Thereby, we propose to add a spatial energy compaction-based penalty into loss function, to achieve higher image compression performance. Furthermore, based on temporal energy distribution, we propose to select the number of frames in one interpolation loop, adapting to the motion characteristics of video contents. Experimental results demonstrate that our proposed image compression outperforms the latest image compression standard with MS-SSIM quality metric, and provides higher performance compared with state-of-the-art learning compression methods at high bit rates, which benefits from our spatial energy compaction approach. Meanwhile, our proposed video compression approach with temporal energy compaction can significantly outperform MPEG-4, and is competitive with commonly used H.264. Both our image and video compression can produce more visually pleasant results than traditional standards. Zhengxue Cheng, Heming Sun, Masaru Takeuchi, Jiro Katto |
CVPR | 2 |
| 2019 | Perceptual Quality Study on Deep Learning Based Image CompressionabstractRecently deep learning based image compression has made rapid advances with promising results based on objective quality metrics. However, a rigorous subjective quality evaluation on such compression schemes have rarely been reported. This paper aims at perceptual quality studies on learned compression. First, we build a general learned compression approach, and optimize the model. In total six compression algorithms are considered for this study. Then, we perform subjective quality tests in a controlled environment using high-resolution images. Results demonstrate learned compression optimized by MS-SSIM yields competitive results that approach the efficiency of state-of-the-art compression. The results obtained can provide a useful benchmark for future developments in learned image compression. Zhengxue Cheng, Pinar Akyazi, Heming Sun, Jiro Katto, Touradj Ebrahimi |
ICIP | 3 |
| 2019 | A Gamut-Extension Method Considering Color Information Restoration using Convolutional Neural NetworksabstractRecently, Ultra HDTV (UHDTV) services become popular over satellite and on the internet. On the contrary, there are tremendously huge volume of High Definition Television (HDTV) and Standard Definition Television (SDTV) contents stored in broadcasting companies and storage devices. In this paper, we propose a color space conversion (also known as gamut mapping) method from BT. 709 (used for current HDTV broadcast) to BT. 2020 (used for UHDTV broadcast), which estimates and restores lost color information. It learns an end-to-end conversion method from BT. 709 image to BT. 2020 image with restoring lost color information using Convolutional Neural Network (CNN). By experiments, we confirm that our method can achieve 2.31dB gain against the conventional method on average. Masaru Takeuchi, Yusuke Sakamoto, Ryota Yokoyama, Heming Sun, Yasutaka Matsuo, Jiro Katto |
ICIP | 4 |
| 2019 | Dual Learning-based Video Coding with Inception Dense BlocksabstractIn this paper, a dual learning-based method in intra coding is introduced for PCS Grand Challenge. This method is mainly composed of two parts: intra prediction and reconstruction filtering. They use different network structures, the neural network-based intra prediction uses the full-connected network to predict the block while the neural network-based reconstruction filtering utilizes the convolutional networks. Different with the previous filtering works, we use a network with more powerful feature extraction capabilities in our reconstruction filtering network. And the filtering unit is the block-level so as to achieve a more accurate filtering compensation. To our best knowledge, among all the learning-based methods, this is the first attempt to combine two different networks in one application, and we achieve the state-of-the-art performance for AI configuration on the HEVC Test sequences. The experimental result shows that our method leads to significant BD-rate saving for provided 8 sequences compared to HM-16.20 baseline (average 10.24% and 3.57% bitrate reductions for all-intra and random-access coding, respectively). For HEVC test sequences, our model also achieved a 9.70% BD-rate saving compared to HM-16.20 baseline for all-intra configuration. Chao Liu 0027, Heming Sun, Zhengxue Cheng, Masaru Takeuchi, Jiro Katto, Xiaoyang Zeng, Yibo Fan |
PCS | 2 |
| 2019 | Fast QTMT Partition Decision Algorithm in VVC Intra Coding based on Variance and GradientabstractQuadtree with nested multi-type tree (QTMT) partition structure in Versatile Video Coding (VVC) contributes to superior encoding performance compared to the basic quad-tree (QT) structure in High Efficiency Video Coding (HEVC). However, the improvement of performance leads to an un-avoidable increase of computational complexity. To achieve a balance between coding efficiency and compression quality, we propose a fast intra partition algorithm based on variance and gradient to solve the rectangular partition problem in VVC. First, further splitting of smooth areas is terminated. Then, QT partition is chosen depending on the gradient features extracted by Sobel operator. Finally, one partition from five possible QTMT partitions is directly chosen by computing the variance of variance of sub-CUs. The theoretical basis of our method is that a homogeneous area tends to be predicted with a larger coding unit (CU), and sub-parts of a split CU are prone to have different textures from each other. To our knowledge, this is the first attempt to apply traditional method to accelerating the rectangular partition problem in VVC intra prediction. Experimental results show that the proposed method can save averagely 53.17% encoding time with only 1.62% BDBR increase and 0.09dB BDPSNR loss compared to anchor VTM4.0. Heming Sun, Jiro Katto, Xiaoyang Zeng, Yibo Fan |
VCIP | 2 |
| 2018 | Sparse ternary connect: Convolutional neural networks using ternarized weights with enhanced sparsityabstractConvolutional Neural Networks (CNNs) are indispensable in a wide range of tasks to achieve state-of-the-art results. In this work, we exploit ternary weights in both inference and training of CNNs and further propose Sparse Ternary Connect (STC) where kernel weights in float value are converted to 1, -1 and 0 based on a new conversion rule with the controlled ratio of 0. STC can save hardware resource a lot with small degradation of precision. The experimental evaluation on 2 popular datasets (CIFAR-10 and SVHN) shows that the proposed method can reduce resource utilization (by 28.9% of LUT, 25.3% of FF, 97.5% of DSP and 88.7% of BRAM on Xilinx Kintex-7 FPGA) with less than 0.5% accuracy loss. Canran Jin, Heming Sun, Shinji Kimura |
ASP-DAC | 2 |
| 2018 | Deep Convolutional AutoEncoder-based Lossy Image CompressionabstractImage compression has been investigated as a fundamental research topic for many decades. Recently, deep learning has achieved great success in many computer vision tasks, and is gradually being used in image compression. In this paper, we present a lossy image compression architecture, which utilizes the advantages of convolutional autoencoder (CAE) to achieve a high coding efficiency. First, we design a novel CAE architecture to replace the conventional transforms and train this CAE using a rate-distortion loss function. Second, to generate a more energy-compact representation, we utilize the principal components analysis (PCA) to rotate the feature maps produced by the CAE, and then apply the quantization and entropy coder to generate the codes. Experimental results demonstrate that our method outperforms traditional image coding algorithms, by achieving a 13.7% BD-rate decrement on the Kodak database images compared to JPEG2000. Besides, our method maintains a moderate complexity similar to JPEG2000. Zhengxue Cheng, Heming Sun, Masaru Takeuchi, Jiro Katto |
PCS | 2 |
| 2018 | Design of Power and Area Efficient Lower-Part-OR Approximate MultiplierabstractApproximate computing has been paid attention as a promising technique to decrease power and area for error-tolerant applications by simplifying the internal operations with sacrificing their accuracy. In this paper, a new power and area efficient approximate multiplier is proposed using OR based compressor with no carry propagation for lower bit positions and a carry propagation compressor with inexact half adders and full adders for upper bit positions. The proposal is effective to reduce the critical path delay with almost the same precision with previous methods. Firstly, an inexact half adder and an inexact full adder are proposed and a construction method of 4×4 multiplier is shown. Then, a construction method of 8×8 multiplier is proposed using OR based compressor, approximate 4×4 multipliers and an accurate 4×4 multiplier. The proposed construction method can also be applied to 16×16 multiplier. The accuracy loss of proposed multipliers is evaluated using MATLAB simulation and that of the proposed 8×8 multiplier is low as 0.20%, the effect of which is shown to be negligible by applying to discrete cosine transform (DCT), inverse DCT and convolutional neural networks for image classification. The proposed 8×8 multiplier reduces power and area by 50.78% and 53.19%, respectively, compared with the accurate Wallace tree multiplier when evaluated using SMIC 40nm process. Yi Guo 0010, Heming Sun, Shinji Kimura |
TENCON | 2 |
| 2017 | A low-cost approximate 32-point transform architectureabstractThis paper presents an area-efficient approximate method for 32-point transform which is one of the most area-consuming parts in High Efficiency Video Coding (HEVC) applications. Compared to prior literatures, this work reduces the hardware cost of transform by 1) eliminating all the arithmetic operations of 6 least significant bits (LSB), 2) presenting a low-delay method for generating carry propagation from the remaining 5 LSBs and 3) truncating the most significant bits (MSB) according to the position of component. In the implementation of a 32-point forward transform, the experimental results show that 27% area consumption can be saved and the coding efficiency loss aroused by the approximation is only 0.044% compared with the origin. Heming Sun, Zhengxue Cheng, Amir Masoud Gharehbaghi, Shinji Kimura, Masahiro Fujita 0004 |
ISCAS | 1 |
| 2017 | Fast Algorithm and VLSI Architecture of Rate Distortion Optimization in H.265/HEVCabstractIn H.265/high efficiency video coding (HEVC) encoding, rate distortion optimization (RDO) is an important cost function for mode decision and coding structure decision. Despite being near-optimum in terms of coding efficiency, RDO suffers from a high complexity. To address this problem, this paper presents a fast RDO algorithm and its very large scale implementation (VLSI) for both intra- and inter-frame coding. The proposed algorithm employs a quantization-free framework that significantly reduces the complexity for rate and distortion optimization. Meanwhile, it maintains a low degradation of coding efficiency by taking the syntax element organization and probability model of HEVC into consideration. The algorithm is also designed with hardware architecture in mind to support an efficient VLSI implementation. When implemented in the HEVC test model, the proposed algorithm achieves 62% RDO time reduction with 1.85% coding efficiency loss for the “all-intra” configuration. The hardware implementation achieves 1.6 × higher normalized throughput relative to previous works, and it can support a throughput of 8k@30fps (for four fine-processed modes per prediction unit) with 256 k logic gates when working at 200 MHz. Heming Sun, Dajiang Zhou, Landan Hu, Shinji Kimura, Satoshi Goto |
IEEE Trans. Multim. | 1 |
| 2016 | Power-efficient and slew-aware three dimensional gated clock tree synthesisabstractThis paper presents a three dimensional (3D) gated clock tree synthesis (CTS) approach, which consists of two steps: 1) abstract tree topology generation; and 2) 3D gated and buffered clock routing. 3D Pair Matching (3D-PM) algorithm is proposed to generate the initial tree topology and then the proposed TSV-minimization algorithm is applied to generate TSV-aware tree topology. Based on TSV-aware tree topology, 3D gated and buffered clock tree routing is done using the proposed 3D Gated and Buffered Deferred-Merge Embedding (3D-GB-DME) algorithm. The slew constraint satisfaction is considered and the clock skew is minimized in our approach. Experimental results show that the proposed method achieves 29.11% power reduction compared with the state-of-the-art 2D work. Minghao Lin, Heming Sun, Shinji Kimura |
VLSI-SoC | 2 |
| 2015 | Merge mode based fast inter prediction for HEVCabstractThe latest High Efficiency Video Coding (HEVC/H.265) obtains 50% bit rate reduction than H.264/AVC standard with comparable quality, but at the cost of high computational complexity. Inter prediction accounts for large complexity and merge mode is one of the most important new features introduced in HEVC. To address this issue, this paper utilizes the merge mode to accelerate inter prediction by three fast mode decision methods. 1) A merge candidate decision is proposed to select the best merge mode by Sum of Absolute Transformed Difference (SATD) cost to reduce the merge time. 2) An early merge termination is presented still based on SATD cost with more than 90% accuracy. 3) Based on efficient merge mode, symmetric motion partition (SMP) modes can be disabled for non-8 × 8 code units (CUs). Experimental results demonstrate that our work can achieve 53.1%-54.2% time reduction on average with 1.57%-2.30% BD-rate increment. Besides, our method achieves an improvement of 18%-30% time reduction with 0.89%-2.85% BD-rate increment when combined with other existing approaches. Zhengxue Cheng, Heming Sun, Dajiang Zhou, Shinji Kimura |
VCIP | 2 |
| 2014 | VLSI architecture of HEVC intra prediction for 8K UHDTV applicationsabstractThis paper presents an efficient VLSI architecture of intra prediction for 8K×4K HEVC decoder. It supports all 35 intra prediction modes and prediction sizes ranging from 4×4 to 64×64. This works proposed a Cyclic SRAM Banks based Parallel Reference Sample Fetching (CSB-PRSF), which guarantees enough reference samples for prediction and reduces the number of registers used for storing reference samples. To guarantee high throughput, 16 pixels are predicted by 4×4 Block Based Pipelining, and dependency between neighboring blocks is eliminated by Hybrid Data Forwarding and Block Reordering. This architecture is synthesized using 90nm technology and the maximum working frequency is 469 MHz, with 72.1K gates area. Running at 397MHz, the architecture can support 4320p@120fps HEVC intra decoding, with full modes and full sizes. Jian-Bin Zhou, Dajiang Zhou, Heming Sun, Satoshi Goto |
ICIP | 3 |
| 2014 | Low-Complexity Rate-Distortion Optimization Algorithms for HEVC Intra Prediction
Zhe Sheng, Dajiang Zhou, Heming Sun, Satoshi Goto |
MMM (1) | 3 |
| 2014 | An area-efficient 4/8/16/32-point inverse DCT architecture for UHDTV HEVC decoderabstractThis paper presents a new VLSI architecture for HEVC inverse discrete cosine transform (TDCT). Compared to prior arts, this work reduces hardware cost by: reducing computational logic of 1-D IDCTs with a reordered parallel-in serial-out (RPISO) scheme that shares the inputs of the butterfly structure; and reducing the area of the transpose buffer with a cyclic memory organization that achieves 100% I/O utilization of the SRAMs. In the implementation of a unified 4/8/16/32-point IDCT, the proposed schemes demonstrate 35% and 62% reduction of logic and memory costs, respectively. The IDCT implementation can support real-time decoding of 4K×2K 60fps video with a total hardware cost of 357,250um2on 2-D IDCT and 80,988um2on transpose memory in 90nm process. Heming Sun, Dajiang Zhou, Jiayi Zhu 0001, Shinji Kimura, Satoshi Goto |
VCIP | 1 |
| 2013 | Multi-scale bidirectional local template patterns for real-time human detectionabstractIn this paper, a feature named multi-scale bidirectional local template patterns (MBLTP) is proposed for human detection. As an extension of bidirectional local template patterns (BLTP), MBLTP not only integrates the textural and gradient information according to the four predefined templates but also calculates information for additional feature vectors by adjusting the scale of the training samples. These additional feature vectors contain multi-scale information on the samples, which can make the feature more discriminative than its original form. Experimental results for an INRIA dataset show that the detection rate of our proposed MBLTP feature outperforms those of other features such as the multi-level histogram of orientated gradient (multi-level HOG), multi scale block histogram of template (MB-HOT), and HOG-LBP. Moreover, in order to make our feature meet real-time requirements, an implementation based on a graphic process unit (GPU) is adopted to accelerate the calculation. Jiu Xu, Ning Jiang 0002, Xinwei Xue, Heming Sun, Wenxin Yu 0001, Satoshi Goto |
MMSP | 4 |
| 2012 | A Low-Complexity HEVC Intra Prediction Algorithm Based on Level and Mode FilteringabstractHEVC achieves a better coding efficiency relative to prior standards, but also involves increased complexity. For intra prediction, complexity is especially intensive due to a highly flexible coding unit structure and a large number of prediction modes. This paper presents a low-complexity intra prediction algorithm for HEVC. A fast preprocessing stage based on a simplified cost model is proposed. Based on its results, a level filtering scheme reduces the number of prediction unit levels that requires fine processing from 5 to 2. To supply level filtering decision with appropriate thresholds, a fast training method is also designed. A mode filtering scheme further reduces the maximum number of angular modes to be evaluated from 34 to 9. Complexity reduction from HM 3.0 is over 50% and stable for various sequences, which makes the proposed algorithm suitable for real-time applications. The corresponding bit rate increase is lower than 2.5%. Heming Sun, Dajiang Zhou, Satoshi Goto |
ICME | 1 |