Haoming Chen

dblp:09/10698 · DBLP profile ↗
← Back
26ranked-venue papers
10as first author
17since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 8 first-author · 8 since 2021Artificial intelligence and machine learning · 10 · 2 first-author · 10 since 2021Systems, architecture and hardware · 3 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Computer networks · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MS-PPO: Mean Standard Deviation Proximal Policy Optimization for Reliable Parking Space Search in Structured Environments
abstract
This paper investigates the reliable parking space search problem in structured environments, with the objective of minimizing the linear combination of mean and standard deviation (mean-std) parking space search time. While canonical parking space search algorithms usually target the minimal expected search time, we argue that risk-averse users would like to trade expectation with its variance, leading to the reliable parking space search problem, which minimizes the mean-std search time. However, the non-additive nature of standard deviation makes the reliable parking space search problem difficult to solve with canonical search algorithms. To address the challenge, we propose a model-free reinforcement learning algorithm, namely MS-PPO, which simultaneously estimates the mean and standard deviation of the current decision-making policy's search time, and performs policy optimization via clipped mean-std advantage function maximization. MS-PPO is compared with several baseline parking space search algorithms as well as canonical reinforcement learning algorithms in a range of representative parking lot networks, and achieves the best overall performance in terms of the mean-std parking space search time. We also validate the effectiveness of MS-PPO in a real parking garage by deploying it to an autonomous vehicle testbed.
Haoming Chen
AAAI1
2026 A Geometric Perspective on Optimizing Vector Quantized Latent Diffusion Model for Image Restoration
abstract
In this paper, we investigate the limitations of the Vector Quantized Latent Diffusion Model (VQ-LDM) in restoration tasks. We identify a performance gap between the Vector Quantization (VQ) and Diffusion Model components, manifested as a significant discrepancy between the reconstruction quality of ground truth images processed via VQ autoregression and degraded images restored by VQ-LDM. Through experiments, we attribute this gap primarily to the lack of robustness in the mapped points of VQ within the original VQ-LDM framework. To address this issue, we propose a geometric based optimization approach. First, we introduce a simple yet effective method, termed interpolation-based latent initial state optimization, which mitigates the performance gap by replacing the original mapped points with interpolated values, supported by theoretical analysis. Here, the latent initial state refers specifically to the input of the diffusion model. Building upon this, we further propose a Chebyshev center-based latent initial state optimization, an elegant theoretical solution from a geometric perspective, that further enhances restoration performance. Our improvements consistently achieve superior results across nine benchmark datasets.
Chen Hang, Haoming Chen, Xuwei Fang, Weisheng Xie, Xiangxiang Gao, Faming Fang, Guixu Zhang
AAAI2
2026 Video-based human pose estimation via feature decoupling and multi-hypothesis calibration
Runyang Feng, Tze Ho Elden Tse, Haoming Chen, Hyung Jin Chang, Haifeng Zhong, Yixing Gao 0001
Pattern Recognit.3
2026 LSPC-LA: Local Structure Preserving Clustering With Learnable Anchors
abstract
K-means algorithm divides samples into c classes based on their structural characteristics. However, due to the non convex nature of the clustering problem, algorithms are prone to converge to poor local minima. To address the aforementioned issues, we propose the Local Structure Preserving Clustering with Learnable Anchors (LSPC-LA) method. We assume that with a well-designed anchor selection strategy, samples near the same anchor tend to belong to the same cluster, which reveal high confidence Must-Link local structural information for clustering. Based on this observation, we first construct an anchor-based bipartite graph, transforming the sample clustering problem into anchor clustering problem by local structural information, thus reducing the solution space and minimizing the risk of poor local minima. Then we create an anchor guiding matrix to allow anchors to learn the sample structure, improving clustering performance. Subsequently, an alternating iterative algorithm is proposed to optimize the LSPC-LA model. Finally, extensive experiments demonstrate the accuracy of the local structural information and the effectiveness of LSPC-LA.
Haonan Xin, Haoming Chen, Zhezheng Hao, Danyang Wu, Rong Wang 0001, Feiping Nie 0001
IEEE Trans. Knowl. Data Eng.2
2025 Learning semantical dynamics and spatiotemporal collaboration for human pose estimation in video
Runyang Feng, Haoming Chen
Neurocomputing2
2025 Visual comparative analytics of multimodal transportation
abstract
Contemporary urban transportation systems frequently depend on a variety of modes to provide residents with travel services. Understanding a multimodal transportation system is pivotal for devising well-informed planning; however, it is also inherently challenging for traffic analysts and planners. This challenge stems from the necessity of evaluating and contrasting the quality of transportation services across multiple modes. Existing methods are constrained in offering comprehensive insights into the system, primarily due to the inadequacy of multimodal traffic data necessary for fair comparisons and their inability to equip analysts and planners with the means for exploration and reasoned analysis within the urban spatial context. To this end, we first acquire sufficient multimodal trips leveraging well-established navigation platforms that can estimate the routes with the least travel time given an origin and a destination (an OD pair). We also propose TraDyssey, a visual analytics system that enables analysts and planners to evaluate and compare multiple modes by exploring acquired massive multimodal trips. TraDyssey follows a streamlined query-and-explore workflow supported by user-friendly and effective interactive visualizations. Specifically, a revisited difference-aware parallel coordinate plot (PCP) is designed for overall mode comparisons based on multimodal trips. Trip groups can be flexibly queried on the PCP based on differential features across modes. The queried trips are then organized and presented on a geographic map by OD pairs, forming a group-OD-trip hierarchy of visual exploration. Domain experts gained valuable insights into transportation planning through real-world case studies using TraDyssey.
Zikun Deng, Haoming Chen, Qing-Long Lu, Zicheng Su, Tobias Schreck, Jie Bao 0003, Yi Cai 0001
Vis. Informatics2
2024 Building a Strong Pre-Training Baseline for Universal 3D Large-Scale Perception
abstract
An effective pre-training framework with universal 3D representations is extremely desired in perceiving large- scale dynamic scenes. However, establishing such an ideal framework that is both task-generic and label-efficient poses a challenge in unifying the representation of the same primitive across diverse scenes. The current contrastive 3D pre-training methods typically follow a frame-level consistency, which focuses on the 2D-3D relationships in each detached image. Such inconsiderate consistency greatly hampers the promising path of reaching an universal pre-training framework: (1) The cross-scene semantic self-conflict, i.e., the intense collision between primitive segments of the same semantics from different scenes; (2) Lacking a globally unified bond that pushes the cross-scene semantic consistency into 3D representation learning. To address above challenges, we propose a CSC framework that puts a scene-level semantic consistency in the heart, bridging the connection of the similar semantic segments across various scenes. To achieve this goal, we combine the coherent semantic cues provided by the vision foundation model and the knowledge-rich cross-scene prototypes derived from the complementary multi-modality information. These allow us to train a universal 3D pre-training model that facilitates various downstream tasks with less fine-tuning efforts. Empirically, we achieve consistent improvements over SOTA pre-training approaches in semantic segmentation (+1.4% mIoU), object detection (+ 1.0% mAP), and panoptic segmentation (+3.0% PQ) using their task-specific 3D network on nuScenes. Code is released at https://github.com/chenhaomingbob/CSC, hoping to inspire future research.
Haoming Chen, Zhizhong Zhang 0001, Yanyun Qu, Xin Tan 0002, Yuan Xie 0006
CVPR1
2024 Exploring Fixed Point in Image Editing: Theoretical Support and Convergence Optimization
abstract
In image editing, Denoising Diffusion Implicit Models (DDIM) inversion has become a widely adopted method and is extensively used in various image editing approaches. The core concept of DDIM inversion stems from the deterministic sampling technique of DDIM, which allows the DDIM process to be viewed as an Ordinary Differential Equation (ODE) process that is reversible. This enables the prediction of corresponding noise from a reference image, ensuring that the restored image from this noise remains consistent with the reference image. Image editing exploits this property by modifying the cross-attention between text and images to edit specific objects while preserving the remaining regions. However, in the DDIM inversion, using the $t-1$ time step to approximate the noise prediction at time step $t$ introduces errors between the restored image and the reference image. Recent approaches have modeled each step of the DDIM inversion process as finding a fixed-point problem of an implicit function. This approach significantly mitigates the error in the restored image but lacks theoretical support regarding the existence of such fixed points. Therefore, this paper focuses on the study of fixed points in DDIM inversion and provides theoretical support. Based on the obtained theoretical insights, we further optimize the loss function for the convergence of fixed points in the original DDIM inversion, improving the visual quality of the edited image. Finally, we extend the fixed-point based image editing to the application of unsupervised image dehazing, introducing a novel text-based approach for unsupervised dehazing.
Chen Hang, Haoming Chen, Xuwei Fang, Vincent Xie, Faming Fang, Guixu Zhang
NeurIPS3
2024 An Effective Optimization Method for Fuzzy $k$k-Means With Entropy Regularization
abstract
Fuzzy$k$-Means with Entropy Regularization method (ERFKM) is an extension to Fuzzy$k$-Means (FKM) by introducing a maximum entropy term to FKM, whose purpose is trading off fuzziness and compactness. However, ERFKM often converges to a poor local minimum, which affects its performance. In this paper, we propose an effective optimization method to solve this problem, called IRW-ERFKM. First a new equivalent problem for ERFKM is proposed; then we solve it through Iteratively Re-Weighted (IRW) method. Since IRW-ERFKM optimizes the problem with$k\times 1$instead of$d\times k$intermediate variables, the space complexity of IRW-ERFKM is greatly reduced. Extensive experiments on clustering performance and objective function value show IRW-ERFKM can get a better local minimum than ERFKM with fewer iterations. Through time complexity analysis, it verifies IRW-ERFKM and ERFKM have the same linear time complexity. Moreover, IRW-ERFKM has advantages on evaluation metrics compared with other methods. What's more, there are two interesting findings. One is when we use IRW method to solve the equivalent problem of ERFKM with one factor$\mathbf{U}$, it is equivalent to ERFKM. The other is when the inner loop of IRW-ERFKM is executed only once, IRW-ERFKM and ERFKM are equivalent in this case.
Yun Liang 0003, Qiong Huang 0001, Haoming Chen, Feiping Nie 0001
IEEE Trans. Knowl. Data Eng.4
2023 Layer-wise partitioning and merging for efficient and scalable deep learning
abstract
Deep Neural Network (DNN) models are usually trained sequentially from one layer to another, which causes forward, backward and update locking problems, leading to poor performance in terms of training time. The existing parallel strategies to mitigate these problems provide suboptimal runtime performance. In this work, we have proposed a novel layer-wise partitioning and merging, forward and backward pass parallel framework to provide better training performance. The novelty of the proposed work consists of 1) a layer-wise partition and merging model which can minimise communication overhead between devices without the memory cost of existing strategies during the training process; 2) a forward pass and backward pass parallelisation to address the update locking problem and minimise the total training cost. The experimental evaluation on real use cases shows that the proposed method outperforms the state-of-the-art approaches in terms of training speed; and achieves almost linear speedup without compromising the accuracy performance of the non-parallel approach.
Samson B. Akintoye, Liangxiu Han, Huw Lloyd, Darren Dancey, Haoming Chen, Daoqiang Zhang
Future Gener. Comput. Syst.6
2023 2D Human pose estimation: a survey
Haoming Chen, Runyang Feng, Sifan Wu 0001, Fengcheng Zhou, Zhenguang Liu
Multim. Syst.1
2023 A Transferred Daily Activity Recognition Method Based on Sensor Sequences
Jinghuan Guo, Jianxun Ren, Haoming Chen, Shuo Han 0014, Shaoxi Li
Neural Process. Lett.3
2023 CXR-Net: A Multitask Deep Learning Network for Explainable and Accurate Diagnosis of COVID-19 Pneumonia From Chest X-Ray Images
abstract
Accurate and rapid detection of COVID-19 pneumonia is crucial for optimal patient treatment. Chest X-Ray (CXR) is the first-line imaging technique for COVID-19 pneumonia diagnosis as it is fast, cheap and easily accessible. Currently, many deep learning (DL) models have been proposed to detect COVID-19 pneumonia from CXR images. Unfortunately, these deep classifiers lack the transparency in interpreting findings, which may limit their applications in clinical practice. The existing explanation methods produce either too noisy or imprecise results, and hence are unsuitable for diagnostic purposes. In this work, we propose a novel explainable CXR deep neural Network (CXR-Net) for accurate COVID-19 pneumonia detection with an enhanced pixel-level visual explanation using CXR images. An Encoder-Decoder-Encoder architecture is proposed, in which an extra encoder is added after the encoder-decoder structure to ensure the model can be trained on category samples. The method has been evaluated on real world CXR datasets from both public and private sources, including healthy, bacterial pneumonia, viral pneumonia and COVID-19 pneumonia cases. The results demonstrate that the proposed method can achieve a satisfactory accuracy and provide fine-resolution activation maps for visual explanation in the lung disease detection. Compared to current state-of-the-art visual explanation methods, the proposed method can provide more detailed, high-resolution, visual explanation for the classification results. It can be deployed in various computing environments, including cloud, CPU and GPU environments. It has a great potential to be used in clinical practice for COVID-19 pneumonia diagnosis.
Xin Zhang 0033, Liangxiu Han, Tamir Sobeih, Lianghao Han, Nina C. Dempsey-Hibbert, Symeon Lechareas, Ascanio Tridente, Haoming Chen, Stephen White, Daoqiang Zhang
IEEE J. Biomed. Health Informatics8
2022 Temporal Feature Alignment and Mutual Information Maximization for Video-Based Human Pose Estimation
abstract
Multi-frame human pose estimation has long been a compelling and fundamental problem in computer vision. This task is challenging due to fast motion and pose occlusion that frequently occur in videos. State-of-the-art methods strive to incorporate additional visual evidences from neighboring frames (supporting frames) to facilitate the pose estimation of the current frame (key frame). One aspect that has been obviated so far, is the fact that current methods directly aggregate unaligned contexts across frames. The spatial-misalignment between pose features of the current frame and neighboring frames might lead to unsatisfactory results. More importantly, existing approaches build upon the straightforward pose estimation loss, which unfortunately cannot constrain the network to fully leverage useful information from neighboring frames. To tackle these problems, we present a novel hierarchical alignment framework, which leverages coarse-to-fine deformations to progressively update a neighboring frame to align with the current frame at the feature level. We further propose to explicitly supervise the knowledge extraction from neighboring frames, guaranteeing that useful complementary cues are extracted. To achieve this goal, we theoretically analyzed the mutual information between the frames and arrived at a loss that maximizes the task-relevant mutual information. These allow us to rank No.1 in the Multi-frame Person Pose Estimation Challenge on benchmark dataset PoseTrack2017, and obtain state-of-the-art performance on benchmarks Sub-JHMDB and Pose-Track2018. Our code is released at https://github.com/Pose-Group/FAMI-Pose, hoping that it will be useful to the community.
Zhenguang Liu, Runyang Feng, Haoming Chen, Shuang Wu 0002, Yixing Gao 0001, Yunjun Gao, Xiang Wang 0010
CVPR3
2022 Personalized motion kernel learning for human pose estimation
abstract
Estimating human poses from a video is at the foundation of many visual intelligent systems. Various convolutional neural networks have been proposed, achieving state-of-the-art performance on different image datasets. However, most existing approaches are image based, which deliver unreliable estimations on videos since they fail to model temporal consistency across video frames. Recently, another line of work leverages temporal cues for multi-frame person pose estimation, yet still in an instance-unaware fashion, disregarding the specific traits of different instances (persons) or different joints. In this paper, we propose a novel approach to learn specific keypoint motion representations for each person, termed Personalized Motion-Aware Network (PMAN). In the PMAN, we devise three components: (i) an Instance-Sensitive Extractor that adaptively computes the spatial features according to human physical characteristics; (ii) a Keypoint Motion Encoder that separately generates convolution kernels with fine-grained keypoint motion encoding; (iii) a Motion Driven Decoder that parses multi-frame spatial features of the same person to provide precise human pose estimations. Extensive experiments on PoseTrack2017 and PoseTrack2018 datasets demonstrate that our approach greatly improves the performance of multi-frame human pose estimation. It is worth mentioning that our approach surpasses the state-of-the-art method by +1.7 mAP and achieves 82.9 mAP on PoseTrack2017 dataset.
Runyang Feng, Haoming Chen, Roger Zimmermann, Zhenguang Liu, Hengchang Liu
Int. J. Intell. Syst.3
2022 GLPose: Global-Local Representation Learning for Human Pose Estimation
abstract
Multi-frame human pose estimation is at the core of many computer vision tasks. Although state-of-the-art approaches have demonstrated remarkable results for human pose estimation on static images, their performances inevitably come short when being applied to videos. A central issue lies in the visual degeneration of video frames induced by rapid motion and pose occlusion in dynamic environments. This problem, by nature, is insurmountable for a single frame. Therefore, incorporating complementary visual cues from other video frames becomes an intuitive paradigm. Current state-of-the-art methods usually leverage information from adjacent frames, which unfortunately place excessive focus on only the temporally nearby frames. In this paper, we argue that combining global semantically similar information and local temporal visual context will deliver more comprehensive and more robust representations for human pose estimation. Towards this end, we present an effective framework, namely global-local enhanced pose estimation ( GLPose ) network. Our framework consists of a feature processing module that conditionally incorporates global semantic information and local visual context to generate a robust human representation and a feature enhancement module that excavates complementary information from this aggregated representation to enhance keyframe features for precise estimation. We empirically find that the proposed GLpose outperforms existing methods by a large margin and achieves new state-of-the-art results on large benchmark datasets.
Yingying Jiao, Haipeng Chen 0002, Runyang Feng, Haoming Chen, Sifan Wu 0001, Yifang Yin, Zhenguang Liu
ACM Trans. Multim. Comput. Commun. Appl.4
2021 Deep Dual Consecutive Network for Human Pose Estimation
abstract
Multi-frame human pose estimation in complicated situations is challenging. Although state-of-the-art human joints detectors have demonstrated remarkable results for static images, their performances come short when we apply these models to video sequences. Prevalent shortcomings include the failure to handle motion blur, video defocus, or pose occlusions, arising from the inability in capturing the temporal dependency among video frames. On the other hand, directly employing conventional recurrent neural networks incurs empirical difficulties in modeling spatial contexts, especially for dealing with pose occlusions. In this paper, we propose a novel multi-frame human pose estimation framework, leveraging abundant temporal cues between video frames to facilitate keypoint detection. Three modular components are designed in our framework. A Pose Temporal Merger encodes keypoint spatiotemporal context to generate effective searching scopes while a Pose Residual Fusion module computes weighted pose residuals in dual directions. These are then processed via our Pose Correction Network for efficient refining of pose estimations. Our method ranks No.1 in the Multi-frame Person Pose Estimation Challenge on the large-scale benchmark datasets PoseTrack2017 and PoseTrack2018. We have released our code, hoping to inspire future research.
Zhenguang Liu, Haoming Chen, Runyang Feng, Shuang Wu 0002, Shouling Ji, Bailin Yang, Xun Wang 0007
CVPR2
2017 Signal Dependent Transform Based on SVD for HEVC Intracoding
abstract
Transform is used to compact the energy of the blocks into a small number of coefficients and is widely used in recent image/video coding standards. In the latest video coding standard high efficiency video coding (HEVC), a combination of discrete cosine transform (DCT) and discrete sine transform (DST) is adopted to transform the residuals from intra prediction. Since the DCT and DST are the fixed transforms that are derived from the Gauss-Markov model, some of residual blocks may not be compacted well by the DCT/DST. In this paper, we propose a signal dependent transform based on singular value decomposition (SVD) for HEVC intracoding. The proposed transform (SDT-SVD) is derived by performing SVD on the synthetic block and applied to the residual block considering the structural similarity between them. Furthermore, we extend SDT-SVD to template matching prediction (TMP) to further improve the intracoding performance. Experimental results show that the proposed transform on angular intra prediction (AIP) outperforms the latest HEVC reference software with a bit rate reduction of 1.0% on average and it can be up to 2.1%. When the proposed transform is extended to TMP-based intracoding, the overall bit rate reduction is 2.7% on average and can be up to 5.8%.
Tao Zhang 0013, Haoming Chen, Ming-Ting Sun, Debin Zhao, Wen Gao 0001
IEEE Trans. Multim.2
2016 Improving Intra Prediction in High-Efficiency Video Coding
abstract
Intra prediction is an important tool in intra-frame video coding to reduce the spatial redundancy. In current coding standard H.265/high-efficiency video coding (HEVC), a copying-based method based on the boundary (or interpolated boundary) reference pixels is used to predict each pixel in the coding block to remove the spatial redundancy. We find that the conventional copying-based method can be further improved in two cases: 1) the boundary has an inhomogeneous region and 2) the predicted pixel is far away from the boundary that the correlation between the predicted pixel and the reference pixels is relatively weak. This paper performs a theoretical analysis of the optimal weights based on a first-order Gaussian Markov model and the effects when the pixel values deviate from the model and the predicted pixel is far away from the reference pixels. It also proposes a novel intra prediction scheme based on the analysis that smoothing the copying-based prediction can derive a better prediction block. Both the theoretical analysis and the experimental results show the effectiveness of the proposed intra prediction method. An average gain of 2.3% on all intra coding can be achieved with the HEVC reference software.
Haoming Chen, Tao Zhang 0013, Ming-Ting Sun, Ankur Saxena, Madhukar Budagavi
IEEE Trans. Image Process.1
2015 Improvements on Intra Block Copy in natural content video coding
abstract
The Intra Block Copy (IntraBC) is a newly adopted tool in the HEVC extension for the screen content video coding. The IntraBC tool efficiently encodes repeating patterns in a picture. The current IntraBC scheme achieves about 1.0% bit-rate reduction on average and up to 4.3% bitrate reduction on natural content video for a database consisting of 2K, 4K, and 8K sequences. In this paper, we propose to improve the IntraBC with a template matching block vector and a fractional search IntraBC. With these two tools, the gain on natural content video coding can be further improved by 0.5% on average and up to 2.0%.
Haoming Chen, Yu-Sheng Chen, Ming-Ting Sun, Ankur Saxena, Madhukar Budagavi
ISCAS1
2015 Hybrid angular intra/template matching prediction for HEVC intra coding
abstract
In the latest HEVC video coding standard, angular intra prediction (AIP) applies 35 modes including 33 angular modes which can handle blocks with direction information well, and other 2 modes (DC and planar) which are used to predict smooth blocks. However, for blocks with complex texture, these modes may not give good predictions. Some complex blocks can be predicted well by template matching prediction (TMP) which was proposed to predict blocks having similar patterns in the coded regions of the same frame without the cost of large overheads. In this paper, a novel hybrid AIP/TMP is proposed to improve the prediction efficiency. Experimental results show that the proposed method can achieve a coding gain of about 2.5% for high resolution sequences on average compared to the AIP in HEVC. The gain can be up to 4.2%.
Tao Zhang 0013, Haoming Chen, Ming-Ting Sun, Debin Zhao, Wen Gao 0001
VCIP2
2015 Adaptive intra-refresh for low-delay error-resilient video coding
Haoming Chen, Chen Zhao 0002, Ming-Ting Sun, Aaron Drake
J. Vis. Commun. Image Represent.1
2014 Nearest-neighbor intra prediction for screen content video coding
abstract
Screen content video coding is becoming increasingly important in various applications, such as desktop sharing, video conferencing, and remote education. In general, compared to natural camera-captured content, screen content has different characteristics, such as sharp edges. In this paper, we propose a novel intra prediction scheme for screen content video. In the proposed scheme, bilinear interpolation in angular intra prediction in HEVC is selectively replaced by nearest-neighbor (NN) interpolation to preserve the sharp edges in screen content video. We present two different variants of NN interpolation. In the first implicit pixel-based method, both the encoder, and the decoder determine whether to perform NN interpolation based on the prediction pixels. The second method comprises of the encoder performing a Rate-Distortion search at a block-level, and explicitly signaling a flag to the decoder to indicate when to use the NN interpolation. Both the proposed variants provide significant gains over HEVC, and simulation results show that average gains of 3.3% BD-bitrate are achieved for screen content video. The HEVC proposal of this method was accepted in the core experiments, and would be a technology under consideration in the ongoing Screen Content Coding extension of HEVC scheduled to begin in March 2014.
Haoming Chen, Ankur Saxena, Felix C. A. Fernandes
ICIP1
2012 Design of low-complexity, non-separable 2-D transforms based on butterfly structures
abstract
The transform used in most image and video coding standards is the separable 2-D discrete cosine transform (DCT), which has been proven to be a robust approximation of the optimal Karhunen-Loève transform (KLT) for the 1st-order Markov sources with a large correlation coefficient. However, such separable 2-D DCT surely is not the best choice when it is applied on some residual or directional signals. Based on the butterfly architecture for DCT's fast implementation, we present in this paper a novel design of non-separable 2-D transforms that get much closer to the KLT but at the implementation cost no bigger than that of the DCT. The critical issue in our design is how to pair all node-variables in various stages of the butterfly structure. We propose a near-optimal pairing strategy to solve this problem and present some examples to demonstrate its effectiveness.
Haoming Chen, Bing Zeng 0001
ISCAS1
2012 New Transforms Tightly Bounded by DCT and KLT
abstract
It is well known that the discrete cosine transform (DCT) and Karhunen–Loève transform (KLT) are two good representatives in image and video coding: the first can be implemented very efficiently while the second offers the best R-D coding performance. In this work, we attempt to design some new transforms with two goals: i) approaching to the KLT's R-D performance and ii) maintaining the implementation cost no bigger than that of DCT. To this end, we follow a cascade structure of multiple butterflies to develop an iterative algorithm: two out of N nodes are selected at each stage to form a Givens rotation (which is equivalent to a butterfly); and the best rotation angle is then determined by maximizing the resulted coding gain. We give the closed-form solutions for the node-selection as well as the angle-determination, together with some design examples to demonstrate their superiority.
Haoming Chen, Bing Zeng 0001
IEEE Signal Process. Lett.1
2011 Design of non-separable transforms for directional 2-D sources
abstract
Traditionally, a 2-D block-based transform is always implemented through two separate 1-D transforms along each block's vertical and horizontal dimensions. Such a framework is however not highly suitable for a directional 2-D source in which the dominant directional information is neither horizontal nor vertical. On the other hand, the R-D performance upper bound for all block-based transform coding schemes applied on such 2-D directional sources can be obtained by the non-separable Karhunen-Loève transform (KLT) — which is unfortunately very expensive computationally. In this paper, we present a new framework for designing some non-separable transforms that offer an R-D performance closer to that of the KLT, but can be implemented with nearly the same complexity as that of the discrete cosine transform (DCT).
Haoming Chen, Shuyuan Zhu, Bing Zeng 0001
ICIP1