Haisheng Fu

dblp:244/2588 · DBLP profile ↗
← Back
17ranked-venue papers
11as first author
15since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 13 · 11 first-author · 11 since 2021Artificial intelligence and machine learning · 3 · 1 first-author · 3 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Spike-hammer: An efficient spike-driven hybrid architecture for multi-modal emotion recognition with physiological signals
Haisheng Fu, Yuchen Zou, Guohe Zhang, Jie Liang 0001
Neural Networks2
2025 SELIC: Semantic-Enhanced Learned Image Compression via High-Level Textual Guidance
abstract
Learned image compression (LIC) techniques have achieved remarkable progress; however, effectively integrating high-level semantic information remains challenging. In this work, we present a Semantic-Enhanced Learned Image Compression framework, termed SELIC, which leverages high-level textual guidance to improve rate-distortion performance. Specifically, SELIC employs a text encoder to extract rich semantic descriptions from the input image. These textual features are transformed into fixed-dimension tensors and seamlessly fused with the image-derived latent representation. By embedding the SELIC tensor directly into the compression pipeline, our approach enriches the bitstream without requiring additional inputs at the decoder, thereby maintaining fast and efficient decoding. Extensive experiments on benchmark datasets (e.g., Kodak) demonstrate that integrating semantic information substantially enhances compression quality. Our SELIC-guided method outperforms a baseline LIC model without semantic integration by approximately 0.1-0.15 dB across a wide range of bit rates in PSNR and achieves a 4.9% BD-rate improvement over VVC. Moreover, this improvement comes with minimal computational overhead, making the proposed SELIC framework a practical solution for advanced image compression applications.
Haisheng Fu, Jie Liang 0001, Zhenman Fang, Jingning Han
ICME1
2025 Toward Multitask Perception for Remote Sensing Imagery via Compression and Prompt Tuning
abstract
Recently, advancements in satellite technology have greatly increased the availability of high-resolution remote sensing images. Concurrently, learning-based image compression (LIC) has significantly improved the efficiency of transmitting and storing such images. As machine recognition tasks increasingly depend on transmitting visual data across devices, compressed images play a key role in both human and machine perception during downstream tasks. However, most LIC approaches are not optimized for machine recognition tasks. To address this limitation, we propose a remote sensing image compression network called RSIC, which integrates multi-task perception and supports downstream tasks such as object detection. Specifically, we introduce a wavelet-based frequency-spatial block (WFSB) that separates frequency components and processes them using Transformer and CNN blocks to effectively capture frequency-specific features. Within WFSB, the Prompting Swin-Transformer Block (PSTB) extracts spatial information while enabling prompt tuning. Additionally, after primary codec training, instance and task prompts are applied during the encoding and decoding stages, respectively, facilitating machine perception without full fine-tuning. Extensive experimental results show that our model achieves better rate-distortion performance for image compression on the AID test dataset, surpassing the traditional VVC codec and several recent LIC methods. Furthermore, our method demonstrates superior performance in terms of rate-accuracy for machine perception on the NWPU VHR-10 and HRSID remote sensing datasets.
Feng Liang 0001, Haisheng Fu, Jiro Katto
IEEE Geosci. Remote. Sens. Lett.4
2025 S2LIC: Learned image compression with the SwinV2 block, Adaptive Channel-wise and Global-inter attention Context
Haisheng Fu, Shang Wang 0006, Zhenjiao Chen, Feng Liang 0001
Neural Networks2
2025 SC-IMC: Algorithm-Architecture Co-Optimized SRAM-Based In-Memory Computing for Sine/Cosine and Convolutional Acceleration
abstract
Sine/cosine (SC) is widely used in practical engineering applications, such as image compression and motor control. Nevertheless, due to power sensitivity and speed demands, SC acceleration suffers from limitations in traditional von-Neumann architectures. To overcome this challenge, we propose accelerating SC and convolution using a static random access memory (SRAM)-based in-memory computing (IMC) architecture through an algorithm-architecture co-optimization manner. We develop the first SC algorithm that transforms nonlinear operations into the IMC paradigm, enabling IMC array to handle both SC and artificial intelligence (AI) tasks and making the IMC array a reusable module. Our architecture extends computing functions of macro dedicated to convolutional neural networks (CNNs), with less than a 1% area increase. The proposed SC algorithm for FP32 data achieves high accuracy within 1 unit in the least significant place (ulp) error margin compared withCmath library. Moreover, we build an intelligent IMC system that supports various CNNs. Our IMC macro implements 512-kb binary weight storage within 3.0366-mm2area in SMIC 28-nm technology and presents area/energy efficiency of 2160.29–270.04 GOPS/mm2and 513.95–8.03 TOPS/W in CNN mode. The proposed algorithm and architecture facilitate the integration of more nonlinear functions into IMC with minimal area overhead.
Shang Wang 0006, Haisheng Fu, Qifan Gao, Zhenjiao Chen, Feng Liang 0001
IEEE Trans. Very Large Scale Integr. Syst.3
2024 Learned Image Compression with Dual-Branch Encoder and Conditional Information Coding
abstract
Recent advancements in deep learning-based image compression are notable. However, prevalent schemes that employ a serial context-adaptive entropy model to enhance rate-distortion (R-D) performance are markedly slow. Furthermore, the complexities of the encoding and decoding networks are substantially high, rendering them unsuitable for some practical applications. In this paper, we propose two techniques to balance the trade-off between complexity and performance. First, we introduce two branching coding networks to independently learn a low-resolution latent representation and a high-resolution latent representation of the input image, discriminatively representing the global and local information therein. Second, we utilize the high-resolution latent representation as conditional information for the low-resolution latent representation, furnishing it with global information, thus aiding in the reduction of redundancy between low-resolution information. We do not utilize any serial entropy models. Instead, we employ a parallel channel-wise auto-regressive entropy model for encoding and decoding low-resolution and high-resolution latent representations. Experiments demonstrate that our method is approximately twice as fast in both encoding and decoding compared to the parallelizable checkerboard context model, and it also achieves a 1.2% improvement in R-D performance compared to state-of-the-art learned image compression schemes. Our method also outperforms classical image codecs including H.266/VVC-intra (4:4:4) and some recent learned methods in rate-distortion performance, as validated by both PSNR and MS-SSIM metrics on the Kodak dataset.
Haisheng Fu, Feng Liang 0001, Jie Liang 0001, Zhenman Fang, Guohe Zhang, Jingning Han
DCC1
2024 WeConvene: Learned Image Compression with Wavelet-Domain Convolution and Entropy Model
Haisheng Fu, Jie Liang 0001, Zhenman Fang, Jingning Han, Feng Liang 0001, Guohe Zhang
ECCV (50)1
2024 Efficient Learned Image Compression with Selective Kernel Residual Module and Channel-Wise Causal Context Model
abstract
Recently, learning-based image compression approaches have achieved superior performance over classical image compression methods. However, their complexities remain quite high. In this paper, we propose two efficient modules to reduce the complexity. First, we introduce a selective kernel residual module into the core network, which effectively expands the receptive field and captures global information. Second, we present an improved channel-wise causal context model, designed to not only reduce encoding and decoding time but also ensure rate-distortion performance. Experimental results demonstrate that our proposed method achieves better tradeoff than recent leading learned image compression methods, and also outperforms the latest H.266/VVC (4:4:4) in terms of PSNR and MS-SSIM metrics.
Haisheng Fu, Feng Liang 0001, Jie Liang 0001, Zhenman Fang, Guohe Zhang, Jingning Han
ICASSP1
2024 Fast and High-Performance Learned Image Compression With Improved Checkerboard Context Model, Deformable Residual Module, and Knowledge Distillation
abstract
Deep learning-based image compression has made great progresses recently. However, some leading schemes use serial context-adaptive entropy model to improve the rate-distortion (R-D) performance, which is very slow. In addition, the complexities of the encoding and decoding networks are quite high and not suitable for many practical applications. In this paper, we propose four techniques to balance the trade-off between the complexity and performance. We first introduce the deformable residual module to remove more redundancies in the input image, thereby enhancing compression performance. Second, we design an improved checkerboard context model with two separate distribution parameter estimation networks and different probability models, which enables parallel decoding without sacrificing the performance compared to the sequential context-adaptive model. Third, we develop a three-pass knowledge distillation scheme to retrain the decoder and entropy coding, and reduce the complexity of the core decoder network, which transfers both the final and intermediate results of the teacher network to the student network to improve its performance. Fourth, we introduce$L_{1}$regularization to make the numerical values of the latent representation more sparse, and we only encode non-zero channels in the encoding and decoding process to reduce the bit rate. This also reduces the encoding and decoding time. Experiments show that compared to the state-of-the-art learned image coding scheme, our method can be about 20 times faster in encoding and 70-90 times faster in decoding, and our R-D performance is also 2.3% higher. Our method achieves better rate-distortion performance than classical image codecs including H.266/VVC-intra (4:4:4) and some recent learned methods, as measured by both PSNR and MS-SSIM metrics on the Kodak and Tecnick-40 datasets.
Haisheng Fu, Feng Liang 0001, Jie Liang 0001, Zhenman Fang, Guohe Zhang, Jingning Han
IEEE Trans. Image Process.1
2023 ROI-Based Deep Image Compression with Swin Transformers
abstract
Encoding the Region Of Interest (ROI) with better quality than the background has many applications including video conferencing systems, video surveillance and object-oriented vision tasks. In this paper, we propose a ROI-based image compression framework with Swin transformers as main building blocks for the autoencoder network. The binary ROI mask is integrated into different layers of the network to provide spatial information guidance. Based on the ROI mask, we can control the relative importance of the ROI and non-ROI by modifying the corresponding Lagrange multiplier λ for different regions. Experimental results show our model achieves higher ROI PSNR than other methods and modest average PSNR for human evaluation. When tested on models pre-trained with original images, it has superior object detection and instance segmentation performance on the COCO validation dataset.
Jie Liang 0001, Haisheng Fu, Jingning Han
ICASSP3
2023 Learned image compression with generalized octave convolution and cross-resolution parameter estimation
Haisheng Fu, Feng Liang 0001
Signal Process.1
2023 Asymmetric Learned Image Compression With Multi-Scale Residual Block, Importance Scaling, and Post-Quantization Filtering
abstract
Recently, deep learning-based image compression has made significant progresses, and has achieved better rate-distortion (R-D) performance than the latest traditional method, H.266/VVC, in both MS-SSIM metric and the more challenging PSNR metric. However, a major problem is that the complexities of many leading learned schemes are too high. In this paper, we propose an efficient and effective image coding framework, which achieves similar R-D performance with lower complexity than the state of the art. First, we develop an improved multi-scale residual block (MSRB) that can expand the receptive field and capture global information more efficiently, which further reduces the spatial correlation of the latent representations. Second, an importance scaling network is introduced to directly scale the latents to achieve content-adaptive bit allocation without sending side information, which is more flexible than previous importance map methods. Third, we apply a post-quantization filter (PQF) to reduce the quantization error, motivated by the Sample Adaptive Offset (SAO) filter in video coding. Moreover, our experiments show that the performance of the system is less sensitive to the complexity of the decoder. Therefore, we design an asymmetric paradigm, in which the encoder employs three stages of MSRBs to improve the learning capacity, whereas the decoder only uses one stage of MSRB, which reduces the decoder complexity and still yields satisfactory performance. Experimental results show that compared to the state-of-the-art method, the encoding and decoding time of the proposed method are about 17 times faster, and the R-D performance is only reduced by about 1% on both Kodak and Tecnick-40 datasets, which is still better than H.266/VVC(4:4:4) and other leading learning-based methods. Our source code is publicly available athttps://github.com/fengyurenpingsheng.
Haisheng Fu, Feng Liang 0001, Jie Liang 0001, Guohe Zhang, Jingning Han
IEEE Trans. Circuits Syst. Video Technol.1
2023 Learned Image Compression With Gaussian-Laplacian-Logistic Mixture Model and Concatenated Residual Modules
abstract
Recently deep learning-based image compression methods have achieved significant achievements and gradually outperformed traditional approaches including the latest standard Versatile Video Coding (VVC) in both PSNR and MS-SSIM metrics. Two key components of learned image compression are the entropy model of the latent representations and the encoding/decoding network architectures. Various models have been proposed, such as autoregressive, softmax, logistic mixture, Gaussian mixture, and Laplacian. Existing schemes only use one of these models. However, due to the vast diversity of images, it is not optimal to use one model for all images, even different regions within one image. In this paper, we propose a more flexible discretized Gaussian-Laplacian-Logistic mixture model (GLLMM) for the latent representations, which can adapt to different contents in different images and different regions of one image more accurately and efficiently, given the same complexity. Besides, in the encoding/decoding network design part, we propose a concatenated residual blocks (CRB), where multiple residual blocks are serially connected with additional shortcut connections. The CRB can improve the learning ability of the network, which can further improve the compression performance. Experimental results using the Kodak, Tecnick-100 and Tecnick-40 datasets show that the proposed scheme outperforms all the leading learning-based methods and existing compression standards including VVC intra coding (4:4:4 and 4:2:0) in terms of the PSNR and MS-SSIM. The source code is available at https://github.com/fengyurenpingsheng.
Haisheng Fu, Feng Liang 0001, Bing Li 0022, Jie Liang 0001, Guohe Zhang, Dong Liu 0002, Chengjie Tu, Jingning Han
IEEE Trans. Image Process.1
2022 Learned Image Compression with Inception Residual Blocks and Multi-Scale Attention Module
Haisheng Fu, Feng Liang 0001, Jie Liang 0001, Guohe Zhang, Jiangning Han
PCS1
2021 An extended context-based entropy hybrid modeling for image compression
Haisheng Fu, Feng Liang 0001, Qian Zhang 0082, Jie Liang 0001, Chengjie Tu, Guohe Zhang
Signal Process. Image Commun.1
2020 Variable-Rate Multi-Frequency Image Compression using Modulated Generalized Octave Convolution
abstract
In this proposal, we design a learned multi-frequency image compression approach that uses generalized octave convolutions to factorize the latent representations into high-frequency (HF) and low-frequency (LF) components, and the LF components have lower resolution than HF components, which can improve the rate-distortion performance, similar to wavelet transform. Moreover, compared to the original octave convolution, the proposed generalized octave convolution (GoConv) and octave transposed-convolution (GoTConv) with internal activation layers preserve more spatial structure of the information, and enable more effective filtering between the HF and LF components, which further improve the performance. In addition, we develop a variable-rate scheme using the Lagrangian parameter to modulate all the internal feature maps in the autoencoder, which allows the scheme to achieve the large bitrate range of the JPEG AI with only three models. Experiments show that the proposed scheme achieves much better Y MS-SSIM than VVC. In terms of YUV PSNR, our scheme is very similar to HEVC.
Haisheng Fu, Qian Zhang 0082, Shang Wang 0006, Jie Liang 0001, Dong Liu 0002, Feng Liang 0001, Guohe Zhang, Chengjie Tu
MMSP3
2020 Improved hybrid layered image compression using deep learning and traditional codecs
Haisheng Fu, Feng Liang 0001, Nai Bian, Qian Zhang 0082, Jie Liang 0001, Chengjie Tu
Signal Process. Image Commun.1