VLDB 2026 Research / reviewers in the wild / expert
Haichuan Ma
dblp:213/5747
· DBLP profile ↗
14ranked-venue papers
5as first author
10since 2021 · last 2025
0000-0003-1126-9382ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 2 · 2 first-author · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Space-Time Video Super-Resolution With Neural OperatorabstractThis paper addresses the task of space-time video super-resolution (STVSR). Existing methods generally suffer from inaccurate motion estimation and motion compensation (MEMC) problems for large motions. Inspired by recent progress in physics-informed neural networks, we model the challenges of MEMC in STVSR as a mapping between two continuous function spaces. Specifically, our approach transforms independent low-resolution representations in the coarse-grained continuous function space into refined representations with enriched spatiotemporal details in the fine-grained continuous function space. To achieve efficient and accurate MEMC, we design a Galerkin-type attention function to perform frame alignment and temporal interpolation. Due to the linear complexity of the Galerkin-type attention mechanism, our model avoids patch partitioning and offers global receptive fields, enabling precise estimation of large motions. The experimental results show that the proposed method surpasses state-of-the-art techniques in both fixed-size and continuous space-time video super-resolution tasks. Code is publicly available at the URL https://github.com/hahazh/STVSR-NO. Yuantong Zhang, Hanyou Zheng, Daiqin Yang, Zhenzhong Chen 0001, Haichuan Ma, Wenpeng Ding |
IEEE Trans. Image Process. | 5 |
| 2024 | Temporal Wavelet Transform-Based Low-Complexity Perceptual Quality Enhancement of Compressed VideoabstractThe past few years have witnessed a great success in applying deep learning to enhance the perceptual quality of compressed video. These methods usually perform frame-by-frame quality enhancement, incurring high computational complexity. Low-complexity perceptual quality enhancement is addressed in this paper, motivated by the observation of temporal correlations among video frames. We propose to decompose video content into temporal low-frequency and high-frequency components, and to focus the enhancement of the temporal low-frequency component, which may significantly reduce the computational complexity. Specifically, we employ the temporal wavelet transform (TWT) for the temporal frequency analysis, and build a TWT-based multiple-input multiple-output perceptual quality enhancement scheme. First, we use a motion estimation method on the input video to acquire the motion information, and then use TWT to obtain the temporal low- and high-frequency components. Second, we design a deep network to enhance the quality of the temporal low-frequency component. Finally, the temporal high-frequency component and the enhanced temporal low-frequency component are combined by the temporal wavelet inverse transform (TWIT) to generate the enhanced video. Experimental results show that our method achieves comparable perceptual quality to that of the state-of-the-art methods, but reduces the computational complexity to 1/13. Cunhui Dong, Haichuan Ma, Zhuoyuan Li 0001, Li Li 0040, Dong Liu 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | DBVC: An End-to-End 3-D Deep Biomedical Video Coding FrameworkabstractBiomedical videos require tremendous storage space and transmission bandwidth, so efficient coding methods are urgently required. Existing methods can be roughly divided into motion-based methods and wavelet-based methods. Motion-based methods use motion estimation designed for natural videos and independently optimize prediction, transform, and entropy coding modules. Wavelet-based methods treat the more redundant time dimension exactly the same as other spatial dimensions. They are both unable to completely remove the redundant spatial-temporal information in biomedical videos. In this paper, to address these problems, we build an end-to-end framework named DBVC with 3-D motion estimation, MV coding, 3-D motion compensation, and residual coding networks for efficient 3-D biomedical video coding. First, we propose a simple yet efficient 3-D motion estimation network to extract motion information. Specifically, we obtain the region with the most intense motion by a segmentation network and then perform unsupervised motion estimation exclusively on this region. After that, to encode and decode the estimated motion vectors, we apply a 3-D autoencoder-based MV coding network. Moreover, we use a lossless learnable wavelet transform for residual coding, which makes lossless coding possible. To the best of our knowledge, this is the first end-to-end video coding framework that supports both lossy and lossless coding, thus meeting the requirements of 3-D biomedical video coding. Extensive experiments demonstrate that our framework achieves state-of-the-art performance on both 3-D biological videos and 3-D medical videos. Dongmei Xue, Haichuan Ma, Li Li 0040, Dong Liu 0002, Zhiwei Xiong, Houqiang Li |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Learning a Single Convolutional Layer Model for Low Light Image EnhancementabstractLow-light image enhancement (LLIE) aims to improve the illuminance of images due to insufficient light exposure. Recently, various lightweight learning-based LLIE methods have been proposed to handle the challenges of unfavorable prevailing low contrast, low brightness, etc. In this paper, we have streamlined the architecture of the network to the utmost degree. By utilizing the effective structural re-parameterization technique, a single convolutional layer model (SCLM) is proposed that provides global low-light enhancement as the coarsely enhanced results. In addition, we introduce a local adaptation module that learns a set of shared parameters to accomplish local illumination correction to address the issue of varied exposure levels in different image regions. Experimental results demonstrate that the proposed method performs favorably against the state-of-the-art LLIE methods in both objective metrics and subjective visual effects. Additionally, our method has fewer parameters and lower inference complexity compared to other learning-based schemes. Code will be made publicly available at the URL https://gitee.com/zhanghahaxixi/SCLM. Yuantong Zhang, Baoxin Teng, Daiqin Yang, Zhenzhong Chen 0001, Haichuan Ma, Wenpeng Ding |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Rectified Wasserstein Generative Adversarial Networks for Perceptual Image RestorationabstractWasserstein generative adversarial network (WGAN) has attracted great attention due to its solid mathematical background, i.e., to minimize the Wasserstein distance between the generated distribution and the distribution of interest. In WGAN, the Wasserstein distance is quantitatively evaluated by the discriminator, also known as the critic. The vanilla WGAN trained the critic with the simple Lipschitz condition, which was later shown less effective for modeling complex distributions, like the distribution of natural images. We try to improve the WGAN training by introducing pairwise constraint on the critic, oriented to image restoration tasks. In principle, pairwise constraint is to suggest the critic assign a higher rating to the original (real) image than to the restored (generated) image, as long as such a pair of images are available. We show that such pairwise constraint may be implemented by rectifying the gradients in WGAN training, which leads to the proposed rectified Wasserstein generative adversarial network (ReWaGAN). In addition, we build interesting connections between ReWaGAN and the perception-distortion tradeoff. We verify ReWaGAN on two representative image restoration tasks: single image super-resolution (4× and 8×) and compression artifact reduction, where our ReWaGAN not only beats the vanilla WGAN consistently, but also outperforms the state-of-the-art perceptual quality-oriented methods significantly. Our code and models are publicly available at https://github.com/mahaichuan/ReWaGAN. Haichuan Ma, Dong Liu 0002, Feng Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2023 | aiWave: Volumetric Image Compression With 3-D Trained Affine Wavelet-Like TransformabstractVolumetric image compression has become an urgent task to effectively transmit and store images produced in biological research and clinical practice. At present, the most commonly used volumetric image compression methods are based on wavelet transform, such as JP3D. However, JP3D employs an ideal, separable, global, and fixed wavelet basis to convert input images from pixel domain to frequency domain, which seriously limits its performance. In this paper, we first design a 3-D trained wavelet-like transform to enable signal-dependent and non-separable transform. Then, an affine wavelet basis is introduced to capture the various local correlations in different regions of volumetric images. Furthermore, we embed the proposed wavelet-like transform to an end-to-end compression framework called aiWave to enable an adaptive compression scheme for various datasets. Last but not least, we introduce the weight sharing strategies of the affine wavelet-like transform according to the volumetric data characteristics in the axial direction to reduce the number of parameters. The experimental results show that: 1) when cooperating our trained 3-D affine wavelet-like transform with a simple factorized entropy coding module, aiWave performs better than JP3D and is comparable in terms of encoding and decoding complexities; 2) when adding a context module to remove signal redundancy further, aiWave can achieve a much better performance than HEVC. Dongmei Xue, Haichuan Ma, Li Li 0040, Dong Liu 0002, Zhiwei Xiong |
IEEE Trans. Medical Imaging | 2 |
| 2022 | Wavelet-Based Learned Scalable Video CodingabstractScalability is an important requirement for video coding when coded videos stream over dynamic-bandwidth networks. The state-of-the-art scalable video coding schemes adopt layer-based methods upon H.265, represented by the SHVC standard. Compared to layer-based schemes, wavelet-based schemes were suspected less efficient for a long while. We try to improve the compression efficiency of wavelet-based scalable video coding by leveraging the recent progresses of deep learning. First, we propose an entropy coding method, using trained convolutional neural networks (CNNs) for probability estimation, to compress the wavelet-transformed subbands. Second, we design a CNN-based method for inverse temporal wavelet transform. We integrate the two proposed methods into a traditional wavelet-based scalable video coding scheme, named Interframe-EZBC. The two methods together achieve more than 20% bits savings. Then, our scheme outperforms the SHVC reference software by 9.09%, 6.55%, and 8.66% BD-rate reductions in YUV respectively. Cunhui Dong, Haichuan Ma, Dong Liu 0002, John W. Woods |
ISCAS | 2 |
| 2022 | End-to-End Optimized Versatile Image Compression With Wavelet-Like TransformabstractBuilt on deep networks, end-to-end optimized image compression has made impressive progress in the past few years. Previous studies usually adopt a compressive auto-encoder, where the encoder part first converts image into latent features, and then quantizes the features before encoding them into bits. Both the conversion and the quantization incur information loss, resulting in a difficulty to optimally achieve arbitrary compression ratio. We propose iWave++ as a new end-to-end optimized image compression scheme, in which iWave, a trained wavelet-like transform, converts images into coefficients without any information loss. Then the coefficients are optionally quantized and encoded into bits. Different from the previous schemes, iWave++ is versatile: a single model supports both lossless and lossy compression, and also achieves arbitrary compression ratio by simply adjusting the quantization scale. iWave++ also features a carefully designed entropy coding engine to encode the coefficients progressively, and a de-quantization module for lossy compression. Experimental results show that lossy iWave++ achieves state-of-the-art compression efficiency compared with deep network-based methods; on the Kodak dataset, lossy iWave++ leads to 17.34 percent bits saving over BPG; lossless iWave++ achieves comparable or better performance than FLIF. Our code and models are available at https://github.com/mahaichuan/Versatile-Image-Compression. Haichuan Ma, Dong Liu 0002, Ning Yan 0001, Houqiang Li, Feng Wu 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2021 | Recycling Discriminator: Towards Opinion-Unaware Image Quality Assessment Using Wasserstein GANabstractGenerative adversarial networks (GANs) have been extensively used for training networks that perform image generation. After training, the discriminator in GAN was not used anymore. We propose to recycle the trained discriminator for another use: no-reference image quality assessment (NR-IQA). We are motivated by twofold facts. First, in Wasserstein GAN (WGAN), the discriminator is designed to calculate the distance between the distribution of generated images and that of real images; thus, the trained discriminator may encode the distribution of real-world images. Second, NR-IQA often needs to leverage the distribution of real-world images for assessing image quality. We then conjecture that using the trained discriminator for NR-IQA may help get rid of any human-labeled quality opinion scores and lead to a new opinion-unaware (OU) method. To validate our conjecture, we start from a restricted NR-IQA problem, that is IQA for artificially super-resolved images. We train super-resolution (SR) WGAN with two kinds of discriminators: one is to directly evaluate the entire image, and the other is to work on small patches. For the latter kind, we obtain patch-wise quality scores, and then have the flexibility to fuse the scores, e.g., by weighted average. Moreover, we directly extend the trained discriminators for authentically distorted images that have different kinds of distortions. Our experimental results demonstrate that the proposed method is comparable to the state-of-the-art OU NR-IQA methods on SR images and is even better than them on authentically distorted images. Our method provides a better interpretable approach to NR-IQA. Our code and models are available at https://github.com/YunanZhu/RecycleD. Yunan Zhu 0001, Haichuan Ma, Jialun Peng, Dong Liu 0002, Zhiwei Xiong |
ACM Multimedia | 2 |
| 2021 | iWave3D: End-to-end Brain Image Compression with Trainable 3-D Wavelet TransformabstractWith the rapid development of whole brain imaging technology, a large number of brain images have been produced, which puts forward a great demand for efficient brain image compression methods. At present, the most commonly used compression methods are all based on 3-D wavelet transform, such as JP3D. However, traditional 3-D wavelet transforms are designed manually with certain assumptions on the signal, but brain images are not as ideal as assumed. What's more, they are not directly optimized for compression task. In order to solve these problems, we propose a trainable 3-D wavelet transform based on the lifting scheme, in which the predict and update steps are replaced by 3-D convolutional neural networks. Then the proposed transform is embedded into an end-to-end compression scheme called iWave3D, which is trained with a large amount of brain images to directly minimize the rate-distortion loss. Experimental results demonstrate that our method outperforms JP3D significantly by 2.012 dB in terms of average BD-PSNR. Dongmei Xue, Haichuan Ma, Li Li 0040, Dong Liu 0002, Zhiwei Xiong |
VCIP | 2 |
| 2020 | Improving Compression Artifact Reduction via End-to-End Learning of Side InformationabstractWe propose to improve neural network-based compression artifact reduction by transmitting side information for the neural network. The side information consists of artifact descriptors that are obtained by analyzing the original and compressed images in the encoder. In the decoder, the received descriptors are used as additional input to a well-designed conditional post-processing neural network. To reduce the transmission overhead, the entire model is optimized under the rate-distortion constraint via end-to-end learning. Experimental results show that introducing the side information greatly improves the ability of the post-processing neural network, and improves the rate-distortion performance. Haichuan Ma, Dong Liu 0002, Feng Wu 0001 |
VCIP | 1 |
| 2020 | iWave: CNN-Based Wavelet-Like Transform for Image CompressionabstractWavelet transform is a powerful tool for multiresolution time-frequency analysis. It has been widely adopted in many image processing tasks, such as denoising, enhancement, fusion, and especially compression. Wavelets lead to the successful image coding standard JPEG-2000. Traditionally, wavelets were designed from the signal processing theory with certain assumption on the signal, but natural images are not as ideal as assumed by the theory. How to design content-adaptive wavelets for natural images remains a difficulty. Inspired by the recent progress of convolutional neural network (CNN), we propose iWave as a framework for deriving wavelet-like transform that is more suitable for natural image compression. iWave adopts an update-first lifting scheme, where the prediction filter is a trained CNN, to achieve wavelet-like transform. The CNN can be embedded into a deep network that is analogous to an auto-encoder, which is trained end-to-end. The trained wavelet-like transform still possesses the lifting structure, which ensures perfect reconstruction, supports multiresolution analysis, and is more interpretable than the deep networks trained as “black boxes.” We perform experiments to verify the generality as well as the speciality of iWave in comparison with JPEG-2000. When trained with a generic set of natural images and tested on the Kodak dataset, iWave achieves on average 4.4% and up to 14% BD-rate reductions. When trained and tested with a specific kind of textures, iWave provides as high as 27% BD-rate reduction. Haichuan Ma, Dong Liu 0002, Ruiqin Xiong, Feng Wu 0001 |
IEEE Trans. Multim. | 1 |
| 2019 | A CNN-Based Image Compression Scheme Compatible with JPEG-2000abstractWe propose a convolutional neural network (CNN) based image compression scheme that is compatible with JPEG-2000. Specifically, our scheme reuses the existing JPEG-2000 encoders to achieve bitstream, and features two components in addition to JPEG-2000: bitstream re-compression and decoder-side post-processing. First, we propose an advanced arithmetic codec that adopts CNN-based probability estimation to exploit the correlation between wavelet coefficients within and across subbands. Second, we propose a CNN-based post-processing method to improve the quality of reconstructed images. Experimental results show that the proposed two CNN-based components both help improve the compression efficiency by a significant margin. Haichuan Ma, Dong Liu 0002, Ruiqin Xiong, Feng Wu 0001 |
ICIP | 1 |
| 2018 | CNN-Based DCT-Like Transform for Image Compression
Dong Liu 0002, Haichuan Ma, Zhiwei Xiong, Feng Wu 0001 |
MMM (2) | 2 |