EDBT 2026 Demo / reviewers in the wild / expert
Kai-Kuang Ma
dblp:49/926
· DBLP profile ↗
168ranked-venue papers
11as first author
49since 2021 · last 2026
0000-0003-2932-5709ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 146 · 10 first-author · 40 since 2021Artificial intelligence and machine learning · 19 · 2 first-author · 9 since 2021Computer networks · 6 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2Systems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | MIROS: Elusive Unauthorized AAV Positioning by Multi-View Radar-Vision Cognitive Fusion
Yuben Qu, Kai-Kuang Ma |
INFOCOM | 5 |
| 2026 | Deep bilateral learning for image interpolation
Jiahuan Ji, Kai-Kuang Ma, Baojiang Zhong, Fuhui Zhou, Qihui Wu 0001 |
Knowl. Based Syst. | 2 |
| 2026 | ARPose: Anatomical relation-driven token pruning for human pose estimation
Xiaodi Sun, Baojiang Zhong, Minghao Piao, Kai-Kuang Ma |
Multim. Syst. | 4 |
| 2026 | CMUA++: A cross-model universal active defense framework for combating deepfake
Xiaoyu Ye, Yongtao Wang, Kai-Kuang Ma |
Pattern Recognit. | 5 |
| 2026 | Highway Camera Calibration and Vehicle Speed Estimation Using Multilayered Lane-Line KeypointsabstractCamera calibration enables the automatic estimation of intrinsic and extrinsic camera parameters, uncovering correspondences between 2D images and 3D real-world coordinates. For highway surveillance cameras, existing methods often rely on cumbersome procedures to extract limited priors (e.g., vanishing points or reference points) and provide incomplete estimations (e.g., roll angle). Therefore, we leverage the multilayered lane lines on highways, which offer rich priors such as segment lengths, intervals, and lane widths, to develop a novel camera calibration and vehicle speed estimation method. For camera calibration, our approach performs road instance segmentation and extractsmultilayered lane-line keypoints (MLK)while mitigating environmental interference and dynamic vehicle occlusions. An MLK-based calibration model is constructed and anangle-polling Levenberg-Marquardt algorithmis designed to estimate key parameters, including focal length, three rotation angles, and lane-line distance. For vehicle speed estimation, multi-object tracking (MOT) algorithms are integrated with the calibration model to infer the average speeds of all identified vehicles. We collected real highway video footage from four different camera setups in Chinese highways. Experimental results demonstrate that our method outperforms existing methods across all setups. The impact of key parameters is evaluated to determine the optimal configuration. Lastly, its effectiveness in vehicle speed estimation is assessed based on advanced MOT algorithms. Fan Xu 0005, Xiaoguang Zhai, Chuibin Chen, Kai-Kuang Ma, Qihui Wu 0001, Xiaofei Zhang 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2026 | An Unsupervised Image Dehazing With Scene Geometry Prior for Road Traffic ScenariosabstractDespite advances in single image dehazing, robust dehazing for real-world road traffic scenes remains challenging due to scarce paired data, traffic-specific geometry, and real-time constraints. To address this issue, we propose a novel image prior for road traffic scenes, termed scene geometry prior (SGP), which leverages depth cues derived from vanishing point (VP) to provide geometry-aware guidance and reduce reliance on paired training data. Our SGP comprises two components: a global SGP (G-SGP) that captures the global geometric distribution and a non-local SGP (NL-SGP) that corrects the errors, among obstructions belonging to the same category, in captured global distribution. Building on the proposed prior, we develop a lightweight and unsupervised road traffic image dehazing network (RTDnet). It consists of a main sub-network guided by the G-SGP to reconstruct the haze-free image, alongside two auxiliary sub-networks that leverage the NL-SGP and VP information to respectively estimate transmission map, and atmospheric light. During training, we introduce an atmospheric scattering model (ASM)-driven mutual-boost learning mechanism (ASM-ML), which is rooted in Bayesian theory and effectively integrates the strengths of different priors without mutual interference while distilling ASM-based physical knowledge into each sub-network. By coupling SGP with ASM-ML, RTDnet can be trained without paired traffic data by exploiting traffic-specific geometry, whose accurate guidance reduces the reliance on large model capacity and enables lightweight real-time deployment. Experiments demonstrate that our RTDnet surpasses state-of-the-art competitors in terms of restoration quality, efficiency, and model size. Moreover, its robust dehazing performance benefits downstream tasks operating in hazy conditions. Mingye Ju, Tianyi Lyu, Chunming He, Qingshan Liu 0001, Kai-Kuang Ma |
IEEE Trans. Image Process. | 5 |
| 2026 | Bayesian Learning-Based Spectrum Mapping With UAV Path Dynamic Optimization Under 3-D Unknown EnvironmentsabstractSpectrum mapping (SM) visualizes spectrum information across a geographical area, constructing radio environment maps (REMs), which serve as a foundation for spectrum monitoring, management, and security. Most existing SM schemes rely on spatially distributed sensors or vehicle-mounted equipment, and assume prior environmental knowledge, limiting their applicability in dynamic or unknown 3D environments. In this paper, we propose a Bayesian learning-based three-dimensional (3D) SM framework that enables accurate REM construction through adaptive UAV sampling in complex and unknown environments. First, a mutual-information-driven UAV path planner is designed by integrating an enhanced sampling-based optimization scheme, enabling efficient data collection according to the maximum mutual information criterion and recent sensing data. Second, a semi-deterministic channel dictionary, refined with sampled field data, is established to model the correlation between observed spectrum values and environmental features. Based on this dictionary, a Bayesian learning-based recovery algorithm reconstructs the spectrum distribution at unsampled positions, producing the corresponding 3D REM. Experimental results on open simulated and measured datasets demonstrate that the proposed framework reduces the mean absolute error by over 60% compared with CS-based methods and by 35% with data-driven interpolation. It also improves sampling efficiency by up to 70% for a given recovery accuracy, highlighting the effectiveness in unknown 3D environments. Jie Wang 0165, Qiuming Zhu, Yuanjin Zheng, Zhipeng Lin 0001, Qihui Wu 0001, Kai-Kuang Ma, Qianhao Gao, Yiran Chen 0024 |
IEEE Trans. Wirel. Commun. | 6 |
| 2025 | Perception-Enhanced Network for Accurate Human Pose EstimationabstractHuman pose estimation in computer vision is particularly challenging with images containing multiple individuals. Existing methods often integrate spatial and channel attention by simply adding them up through a cascade or parallel connection. The features extracted in this way could lead to less accurate key point predictions, especially in cases where the limbs of different people are obstructed or tangled with each other. To tackle this crucial issue, we develop a novel network that enhances key point detection by combining the spatial and channel attention in a more effective manner. Specifically, our network features a lightweight perception-enhanced module (PEM) that adaptively fuses spatial and channel features through a Hadamard product, thereby refining the overall feature representation. Moreover, by exploiting the initial feature map as a guide to generate the pixel attention, we further boost key point prediction accuracy. Extensive experimental results show that our developed network can clearly outperform the current state-of-the-art methods. Xiaodi Sun, Baojiang Zhong, Kai-Kuang Ma |
ICASSP | 3 |
| 2025 | Deformable Attention-Based Edge-Aware Network for Single Image Super-ResolutionabstractAccurately reconstructing object edges is a key challenge in single image super-resolution (SISR), as it greatly influences our visual perception of image quality. To address this fundamental issue, we propose a novel SISR approach named the deformable attention-based edge-aware (DAE) network. The DAE network features a deformable attention block that dynamically adjusts attention weights to align with edge structures, thereby improving edge awareness and reconstruction quality. Moreover, our network incorporates a multi-patterns window block that captures fine-grained edge details and enhances information flow across network layers. This combination results in visually superior SISR outputs. Extensive experiments and comparisons with the state-of-the-art methods demonstrate that our DAE network excels in both quantitative metrics and visual quality. Baojiang Zhong, Kai-Kuang Ma |
ICASSP | 3 |
| 2025 | Lightweight Single Image Super-Resolution With High-Continuity AttentionabstractWindow attention has become a popular choice in single image super-resolution (SISR) network design due to its efficient computation. However, its self-attention is restricted to fixed-size windows, leading to a lack of cross-window interaction. To address this, the benchmark SwinIR model adopts a shifted window strategy to capture long-range dependencies. However, we observe that its attention still suffers from discontinuities at window boundaries, resulting in inferior SISR performance. To address this issue, we propose a newscale-dual attention(SDA) module, consisting of three parallel branches that integratewindowattention andpoolingattention via three complementary scales. This enables hierarchical local-global interactions, yielding high-continuity attention maps. To validate the effectiveness of our proposed SDA, we develop a lightweightscale-dual attention network(SDAN) with approximately 878K parameters for SISR. Extensive experiments demonstrate that our SDAN achieves superior performance, outperforming state-of-the-art methods in both accuracy and efficiency. Baojiang Zhong, Kai-Kuang Ma |
IEEE Signal Process. Lett. | 3 |
| 2025 | BAN: A Boundary-Aware Network for Accurate Colorectal Polyp SegmentationabstractColonoscopy images exhibit multi-frequency features, with polyp boundaries residing in a mid-frequency range, which are critical for accurate polyp segmentation. However, current deep learning models tend to prioritize low-frequency features, leading to reduced segmentation performance. To address this challenge, we propose a novelboundary-aware network(BAN) that integrates trainable Gabor filters into the polyp segmentation process through a dedicated module calledGabor-driven feature extraction(GFE). By developing and using atrajectory-directed frequency learningapproach, Gabor filters are trained along adamping sinusoidalpath, dynamically optimizing their frequency parameters within a proper mid-frequency range. This enhances boundary feature representation and significantly improves polyp segmentation accuracy. Extensive experiments demonstrate that our BAN outperforms existing state-of-the-art methods. Nengxiang Zhang, Baojiang Zhong, Minghao Piao, Kai-Kuang Ma |
IEEE Signal Process. Lett. | 4 |
| 2025 | All-Inclusive Image Enhancement for Degraded Images Exhibiting Low-Frequency CorruptionabstractIn this paper, a novel image enhancement method, called the all-inclusive image enhancement (AIIE), is proposed that can effectively enhance the degraded images for improving the visibility of image content. These imageries were acquired under various types of weather conditions such as haze, low-light, underwater, and sandstorm, etc. One commonality shared by this class of noise is that the resulted degradations on visual quality or visibility are caused by low-frequency interference. Existing image enhancement methods lack the ability to deal with all types of degradations from this class, while our proposed AIIE offers a unified treatment for them. To achieve this goal, a statistical property is obtained from the study of the discrete cosine transform (DCT) of 1,000 high- and 1000 low-quality images on their DCT domains. It shows that the normalized DCT coefficients (between 0 and 1) of high-quality images has about 95% fall in the interval [0, 0.2]; for low-quality images, almost all the coefficients are in the same interval. This fundamental property, called the DCT prior (DCT-P), is instrumental to the development of our AIIE algorithm proposed in this paper. Since the proposed DCT-P delineates the attributes of high- and low-quality images clearly, it becomes a highly effective ‘tool’ to convert low-quality images to its enhanced version. Extensive experimental results have clearly validated the superior performance of the AIIE conducted on different types of deteriorated images in terms of visual quality and efficiency as well as significant advantages on computational complexity, which is essential for real-time applications. Mingye Ju, Chunming He, Can Ding 0002, Wenqi Ren, Lin Zhang 0014, Kai-Kuang Ma |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Deep Multi-Modal Ship Detection and Classification NetworkabstractWhile a majority of single-modal ship detectors solely rely on RGB images, a novel multi-modal real-time transformer-based ship detection and classification method, called the MM-ShipNet, is proposed in this paper that integrates the data acquired from three modalities—i.e., RGB camera, radar, and automatic identification system (AIS). First, a bounding box is generated based on the position information from radar and ship’s actual size information from AIS. This physical information are fused and projected onto the camera-acquired RGB image frame. Each bounding box is then possibly weighted depending on the ship size presented on the image. The generated weighted ship masks (WSMs) will be exploited for facilitating ship classification task. In the second stage of MM-ShipNet, multi-modal detection transformer (MM-DETR) introduces an multi-modal cross-scale encoder (MCE) for improving ship detection and classification performance. Our MCE exploits a dual-flow structure to fuse the features extracted from the WSMs and the RGB images under different scales. Since our method is the first work entailing three aforementioned modalities, no such dataset with all modalities can be found in the open source. Thus, we construct a multi-modal ship dataset, termed MMShips, as another contribution. Our MMShips dataset comprises 9,513 camera-acquired real-life maritime RGB images and their aligned ship masks generated from radar and AIS. Experimental results clearly demonstrate that our MM-ShipNet significantly outperforms multiple state-of-the-art single-modal and multi-modal ship detectors. Fan Xu 0005, Chuibin Chen, Zhigao Shang, Kai-Kuang Ma, Qihui Wu 0001, Zebin Lin, Jie Zhan, Yizhou Shi |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Cognitive Contour Detection of Sparse-Structured Objects in the Alpha-Shape Scale SpaceabstractIn this paper, we introduce cognitive contour, a novel image attribute that encapsulates the global shape perceived from sparsely distributed, identical or similar objects-such as drone swarms or flocks of geese-collectively termed sparse-structured objects. Unlike traditional contour analysis that delineates the boundaries of individual objects, cognitive contours reflect a gestalt-inspired perception of the overall structure formed by the ensemble, capturing higher-level visual organization. Detecting cognitive contours is challenging due to the sparsity and multiplicity of constituent elements. To tackle this, we propose a scale-space method that integrates alpha shapes into a scale-space framework. An alpha-shape scale space is constructed for the sparse-structured object, and the optimal scale is adaptively selected to extract cognitively meaningful contours with appropriate structural detail. Extensive experiments validate the effectiveness and robustness of the proposed method, enhancing visual inference and offering flexibility across diverse image-based applications. Code and data are available at: https://github.com/CookiC/Sparse. Yuxiang Shen, Baojiang Zhong, Kai-Kuang Ma |
IEEE Trans. Image Process. | 3 |
| 2025 | APSNR: Artifact Peak Signal-to-Noise Ratio for Image Quality AssessmentabstractAn image processing pipeline typically involves key operations like compression, denoising, and resizing, along with enhancements such as sharpening, histogram equalization, and low-light compensation. Within this pipeline, image artifacts are often introduced, which could severely degrade perceptual quality and mislead downstream vision tasks. Yet, current image quality assessment (IQA) models fail to distinguish between harmful artifacts and beneficial enhancements, as they generally apply a rigid fidelity criterion that penalizes all deviations from the reference image. We address this gap with the artifact peak signal-to-noise ratio (APSNR), a new IQA metric that adopts a selective fidelity criterion-allowing legitimate enhancements while penalizing only spurious artifacts. Specifically, APSNR detects artifacts by identifying pixels that violate an "artifact-free" intensity mapping between the processed and reference images, and then computes PSNR exclusively within the artifact-corrupted regions. Extensive experiments demonstrate that our APSNR consistently correlates with human perception of artifacts while remaining robust to enhancements. This enables a more nuanced evaluation of image processing algorithms and provides a principled tool for benchmarking artifact suppression. Baojiang Zhong, Kai-Kuang Ma |
IEEE Trans. Image Process. | 3 |
| 2024 | Ellipse Detection Based On Structure-Preserving Anisotropic Edge ExtractionabstractExisting methods for ellipse detection popularly adopt the edge-linking strategy—i.e., first combining the elliptical arcs extracted from an edge map into groups and then fitting each group of arcs to an ellipse. However, such methods generally use the Canny operator to extract edges, which tends to corrupt salient elliptical shapes in texture regions, thus preventing the effective detection of ellipses. To overcome the difficulty, we propose a novel ellipse detector based on our developed structure-preserving anisotropic edge extraction (SPAEE) approach, which can remove redundant textures and preserve continuous structural edges, thus improving the performance of ellipse detection. In addition, an adaptive validation strategy is proposed to further enhance the detection quality. To our best knowledge, this is the first attempt to detect ellipses with an anisotropic edge extraction process for preserving structural edges. Experimental results have shown that the mean F-score on five benchmark datasets has increased from 0.61 to 0.67 (about 10% performance gain). Baojiang Zhong, Kai-Kuang Ma |
ICASSP | 4 |
| 2024 | Corner Detection Based on a Rotation-Invariant and Noise-Insensitive Curvature MeasurementabstractCorner detection is extensively applied across various computer vision tasks. Current corner detectors typically assume that the distance between every two nearby pixels is constant. However, this assumption is invalid in real-world scenarios. As a result, the pixel-based curvature measurements designed and used in these corner detectors may suffer instability under rotation transformation and noise interference. To tackle this issue, a novel curvature measurement is proposed in this paper, which exploits the length of the subpixelized chord to estimate the discrete curvatures of digital curves. The proposed curvature model is invariant to rotation transformation and insensitive to image noise interference. Based on this curvature measurement, a new corner detector is further developed. Experimental results demonstrate that our proposed corner detector outperforms existing state-of-the-art methods. Baojiang Zhong, Kai-Kuang Ma |
ICASSP | 3 |
| 2024 | Ellipse Detection Based on Contrast-Guided Arc EnhancementabstractA common practice of existing ellipse detection methods is to produce an ellipse by grouping elliptic arcs. Thus, the quality (i.e., continuity) of elliptic arcs will have a significant impact on the performance of ellipse detection. In this paper, a novel contrast-guided ellipse detection method is proposed for generating elliptic arcs with improved quality. After applying an edge detector to the input image, all the nearby pixels along each edge are evaluated for possibly re-classifying them to the opposite class; that is, non-edge pixels might be re-classified as edge pixels, and vice versa. With the proposed contrast-guided arc enhancement, the produced elliptic arcs tend to have better quality, which is essential to detect ellipses more accurately. To further assure the detected ellipses, a novel ellipse validation process is developed and used to discard those ellipses with low confidence. Extensive experimental results have shown that the proposed method can deliver superior performance in nearly all test cases. Baojiang Zhong, Kai-Kuang Ma |
ICASSP | 3 |
| 2024 | ATU-NET: An Adaptive Transformation-Based U-NET for Medical Image SegmentationabstractBoth the U-Net and its variants, which produce the state-of-the-art performance in the field of medical image segmentation, are founded on an encoder-decoder architecture. However, this architecture generally processes the input image in the spatial domain only, overlooking potential insights that could be gained from other transform domains. For that, an adaptive transformation-based U-Net (ATU-Net) is proposed in this paper. Our ATU-Net is based on a novel network architecture called the adaptive transformation-encoder-decoder (ATU), which adaptively transforms the input image into a more suitable domain by training the transformation kernel for processing to make it easier for the encoder and decoder to extract key features of the image. Extensive experimental results obtained on benchmark datasets have shown that our proposed ATU-Net can deliver superior performance to the existing methods. Qianyu Du, Baojiang Zhong, Kai-Kuang Ma |
ICIP | 3 |
| 2024 | An Image Decomposition-Guided Network for Image InterpolationabstractA novel image decomposition-guided network (IDGN) for image interpolation is proposed in this paper by incorporating the fundamentals of subband image decomposition into the design of our deep-learning network. In our work, a filter bank consisting of a Gaussian filter and a differenceof-Gaussian filter is designed for decomposing the low-resolution input image into multiple subbands of the same resolution without downsampling. These subbands are inherited with different low-frequency and high-frequency information and are ready to be interpolated individually in our developed network. For training our IDGN, the decomposed low-resolution subbands need to be paired up with their corresponding ground-truth high-resolution subbands. Since our human visual system is sensitive to high-frequency signals, a perception-regulated (PR) loss function is proposed to guide our IDGN by putting more emphasis on the high-frequency subbands during the training process. Extensive experimental results have shown that our IDGN can achieve superior performance when compared with a number of state-of-the-art image interpolation methods. Jiahuan Ji, Baojiang Zhong, Kai-Kuang Ma, Fuhui Zhou, Qihui Wu 0001 |
ICIP | 3 |
| 2024 | A Robust Airport Detection Method Based on Environment-Insensitive Saliency Analysis
Hongtao Chen, Baojiang Zhong, Kai-Kuang Ma |
PRICAI (5) | 3 |
| 2024 | Unsupervised video-based action recognition using two-stream generative adversarial network
Wei Lin 0021, Huanqiang Zeng, Jianqing Zhu, Chih-Hsien Hsia, Junhui Hou, Kai-Kuang Ma |
Neural Comput. Appl. | 6 |
| 2024 | A Channel-Wise Multi-Scale Network for Single Image Super-ResolutionabstractExisting multi-scale feature extraction methods extract image features using various convolution window sizes conducted on the spatial dimension of the feature maps. However, such an approach inevitably encounters redundant convolution operations. To address this concern, we propose to extract multiscale features on the channel dimension rather than on the spatial dimension. To demonstrate, a channel-wise multi-scale network (CMSN) is proposed for conducting single image super-resolution (SISR). In our CMSN, a sequence of channel-wise multi-scale blocks (CMSBs) is designed to extract multi-scale features at increasing levels by performing convolutions with different channel numbers (i.e., scales). To fuse the image features generated from different levels in our CMSN, a hybrid attention-aware feature fusion block (HAFFB) is proposed. Extensive experimental results have clearly shown the superiority of our CMSN to that of several state-of-the-art SISR methods on delivering superior high-resolution images, both objectively and subjectively. This reveals the potential of channel-wise, versus spatial-wise, on the effectiveness of multi-scale feature extraction. Jiahuan Ji, Baojiang Zhong, Qihui Wu 0001, Kai-Kuang Ma |
IEEE Signal Process. Lett. | 4 |
| 2024 | Contrast-Guided Line Segment DetectionabstractDue to the effects of quantization error and image noise, detecting ‘meaningful’ line segments from an image with high continuity is a challenging task. To pursue this goal, a novel line segment detector, called thecontrast-guided line segment detector(CGLSD), is proposed in this paper. Our basic idea is to integrate a low-level image attribute, i.e.,edge contrast, into the line segment detection process for improving line continuity. After applying an edge detector to the input image, the edge contrast is exploited to guide the growth of aline-support regionfor each line segment individually. This is achieved by evaluating edge pixels as well as those non-edge pixels that are nearby the edges. As a result, some of the non-edge pixels are re-considered as ‘edge’ pixels and included for establishing the support region. Reversely, certain edge pixels might be treated as ‘non-edge’ pixels instead and excluded from the region. Since each support region is supposed to yield onlyoneline segment, each formed support region needs to have arefinementby removing those edge pixels that do not belong to it. Lastly, the support region is required to pass through avalidationcheck that might lead to a complete discard of the line segment due to its low confidence. Extensive experiments are conducted and compared with multiple state-of-the-arts on two datasets, including the one from us with manually-annotated ground truth. The results have shown that the proposed CGLSD can deliver superior performance in nearly all test cases. Zikai Wang 0007, Baojiang Zhong, Dongxu Han, Kai-Kuang Ma |
IEEE Signal Process. Lett. | 4 |
| 2024 | Anisotropic Scale-Invariant Ellipse DetectionabstractDetecting ellipses poses a challenging low-level task indispensable to many image analysis applications. Existing ellipse detection methods commonly encounter two fundamental issues. First, inferior detection accuracy could be incurred on a small ellipse than that on a large one; this introduces the scale issue. Second, inferior detection accuracy could be yielded along the minor axis than along the major one of the same ellipse; this leads to the anisotropy issue. To address these issues simultaneously, a novel anisotropic scale-invariant (ASI) ellipse detection methodology is proposed. Our basic idea is to perform ellipse detection in a transformed image space referred to as the ellipse normalization (EN) space, in which the desired ellipse from the original image is 'normalized' to the unit circle. With the establishment of the EN-space, an analytical ellipse fitting scheme and a set of distance measures are developed. Theoretical justifications are then derived to prove that both our ellipse fitting scheme and distance measures are invariant to anisotropic scaling, and thus each ellipse can be detected with the same accuracy regardless of its size and ellipticity. By incorporating these components into two recent state-of-the-art algorithms, two ASI ellipse detectors are finally developed and exploited to verify the effectiveness of our proposed methodology. Zikai Wang 0007, Baojiang Zhong, Kai-Kuang Ma |
IEEE Trans. Image Process. | 3 |
| 2024 | Width-Adaptive CNN: Fast CU Partition Prediction for VVC Screen Content CodingabstractScreen content coding (SCC) in Versatile Video Coding (VVC) improves the coding efficiency of screen content videos (SCVs) significantly but results in high computational complexity due to the quad-tree plus multi-type tree (QTMT) structure of the coding unit (CU) partitioning. Therefore, we make the first attempt to reduce the encoding complexity from the perspective of CU partitioning for SCC in VVC. To this end, a fast CU partition prediction method is technically developed for VVC-SCC. First, to solve the problem of lacking sufficient SCC training data, SCVs are collected to establish a database containing CUs of various sizes and corresponding partition labels. Second, to determine the partition decision in advance, a novel WA-CNN model is proposed, which is capable of predicting two large CUs for VVC-SCC by adjusting the feature channels based on the size of input CU blocks. Finally, considering the imbalanced proportion of diverse partition decisions, a loss function with the weight that equalizes the contribution of imbalanced data is formulated to train the proposed WA-CNN model. Experimental results show that the proposed model reduces the SCC intra-encoding time by 35.65%${\sim }$38.31% with an average of 1.84%${\sim }$2.42% BDBR increase. Chao Jiao, Huanqiang Zeng, Jing Chen 0001, Chih-Hsien Hsia, Tianlei Wang, Kai-Kuang Ma |
IEEE Trans. Multim. | 6 |
| 2023 | Learning a Simple Low-Light Image Enhancer from Paired Low-Light InstancesabstractLow-light Image Enhancement (LIE) aims at improving contrast and restoring details for images captured in lowlight conditions. Most of the previous LIE algorithms adjust illumination using a single input image with several handcrafted priors. Those solutions, however, often fail in revealing image details due to the limited information in a single image and the poor adaptability of handcrafted priors. To this end, we propose PairLIE, an unsupervised approach that learns adaptive priors from low-light image pairs. First, the network is expected to generate the same clean images as the two inputs share the same image content. To achieve this, we impose the network with the Retinex theory and make the two reflectance components consistent. Second, to assist the Retinex decomposition, we propose to remove inappropriate features in the raw image with a simple self-supervised mechanism. Extensive experiments on public datasets show that the proposed PairLIE achieves comparable performance against the state-of-the-art approaches with a simpler network and fewer handcrafted priors. Code is available at: https://github.com/zhenqifu/PairLIE. Zhenqi Fu, Xiaotong Tu, Yue Huang 0001, Xinghao Ding, Kai-Kuang Ma |
CVPR | 6 |
| 2023 | A Multi-Scale Cell Segmentation Method for Detecting Hematological DisordersabstractCell segmentation, conducted on a microscopic image that contains blood cells, plays a crucial role for the detection of various hematological disorders. Existing methods often yield inferior performance in the presence of elongated and irregularly-shaped cells, as well as to those adjacent cells with partial overlapping among themselves. To address these issues, a novel multi-scale cell segmentation (MCS) method is proposed in this paper that involves three scales, denoted by coarse, medium, and fine, for demonstrating the effectiveness and efficacy of the proposed multiscale approach. It has been shown in our work that noise and insignificant cell structures can be effectively suppressed at the coarse scale. Consequently, those elongated and irregularly-shaped cells are more accurately identified. Furthermore, our developed coarse-to-fine scale tracing algorithm, performed to trace each identified cell, improves segmentation accuracy. To solve the problem of cell overlapping, an effective cell identification technique is developed for extracting cell contours for each scale. Extensive experimental results obtained from three benchmark datasets have shown that the proposed MCS outperforms the state-of-the-art methods by a large margin. Yating Fang, Baojiang Zhong, Kai-Kuang Ma |
ICIP | 4 |
| 2023 | Learning Multi-Scale Features for Jpeg Image Artifacts RemovalabstractA recently proposed quantization-table convolutional network (QCN) has proven as a state-of-the-art JPEG image artifacts removal method. To reduce computational complexity, the QCN learns image features from the down-sampled version of the input image. Consequently, the performance might be compromised as some salient features that can only be learned from the original input image with full resolution will be lost. To solve this problem, a novel multi-scale feature extraction block (MFEB) is proposed in this paper, which contains a coarse-scale branch and a fine-scale branch for learning salient image features from the down-sampled input image and the original full resolutions, respectively. To avoid introducing too much additional computational complexity due to the fine-scale branch, in our MFEB, this branch uses only one layer, while the coarse-scale branch exploits multiple layers. The feature sets obtained from the two branches are then fused. With the MFEB, a novel multi-scale artifacts removal network (MARN) is then developed to remove JPEG image artifacts. Extensive experiments have clearly shown that our MARN can deliver superior performance to that of a number of state-of-the-art methods. Jiahuan Ji, Baojiang Zhong, Weigang Song, Kai-Kuang Ma |
ICIP | 4 |
| 2023 | On the generalized k-cosine arithmetic-mean curvature for multi-scale corner detection
Baojiang Zhong, Kai-Kuang Ma |
Expert Syst. Appl. | 4 |
| 2023 | Corner Detection via Scale-Space Behavior-Guided Trajectory TracingabstractExistingcurvature scale-space(CSS) methods detect corners by tracing the CSS trajectories from a determined high scale toward the lowest one. For those images with sophisticated details, such approach could often yield unsatisfactory corner detection results; i.e., miss-detectedtruecorners (false negatives) androundcorners (false positives). In this letter, these two fundamental problems are investigated. To tackle them, a novel CSS-based corner detector is proposed by incorporating our mathematically derived scale-space properties of the planar curves and corner points into the developed trajectory tracing algorithm, called thescale-space behavior-guided trajectory tracing(SBTT). In view of lacking a benchmark dataset with the ground truth, another contribution from our work is on the establishment of an augmented test image dataset, containing 147 test images with manually-labelled ground truth and their augmented images up to 62,328 images in total. Based on the ground truth, four commonly-used metrics are exploited to conduct corner detection performance evaluation. The obtained simulation results show that our proposed corner detector yields the highest F-score, when compared with that of nine state-of-the-art methods. Baojiang Zhong, Jianyu Yang 0002, Kai-Kuang Ma |
IEEE Signal Process. Lett. | 4 |
| 2023 | 3D-Gradient Guided Rate Control Model for Screen Content Video CodingabstractCompared with natural videos,screen content videos(SCVs) have particular features, such as fruitful sharper edges, lots of computer-generated graphics and texts, a large amount of flat areas. New tools are adopted toHEVC extensions on Screen Content Coding(HEVC-SCC), the traditional video rate control methods for natural videos are not effective for SCVs. For that, a3D-gradient guided rate control modelfor SCV coding, named 3DG-RC, is proposed to allocate bitrate more efficiently serving for SCVs. By considering the particular spatial-temporal characteristics of SCVs, the spatial and temporal feature extraction scheme is developed by using 3D-gradient filter and performed on the SCV to extract the spatial and temporal features simultaneously for guiding the bit allocation. The spatial-temporal feature similarity between three original reference SCV frames and their reconstructed ones is used to estimate the encoding parameters of the current block and frame. Experimental results demonstrate that compared with the classical and state-of-the-art rate control methods for HEVC-SCC, the proposed 3DG-RC algorithm achieves significant bitrate mismatch reduction and coding efficiency improvement for HEVC-SCC. In specific, the proposed 3DG-RC model outperforms the rate control model in SCM-8.8 with over 41.33% and 37.95% BD-BR savings on average, forlow delay B(LDB) andrandom access(RA) coding structure, respectively. Jing Chen 0001, Huanqiang Zeng, Chih-Hsien Hsia, Tianlei Wang, Kai-Kuang Ma |
IEEE Trans. Multim. | 6 |
| 2023 | Deep Cross-Modal Hashing Based on Semantic Consistent RankingabstractThe amount of multi-modal data available on the Internet is enormous. Cross-modal hash retrieval maps heterogeneous cross-modal data into a single Hamming space to offer fast and flexible retrieval services. However, existing cross-modal methods mainly rely on the feature-level similarity between multi-modal data and ignore the relationship between relative rankings and label-level fine-grained similarity of neighboring instances. To overcome these issues, we propose a novelDeepCross-modalHashing based onSemanticConsistentRanking (DCH-SCR) that comprehensively investigates the intra-modal semantic similarity relationship. Firstly, to the best of our knowledge, it is an early attempt to preserve semantic similarity for cross-modal hashing retrieval by combining label-level and feature-level information. Secondly, the inherent gap between modalities is narrowed by developing a ranking alignment loss function. Thirdly, the compact and efficient hash codes are optimized based on the common semantic space. Finally, we use the gradient to specify the optimization direction and introduce the Normalized Discounted Cumulative Gain (NDCG) to achieve varying optimization strengths for data pairs with different similarities. Extensive experiments on three real-world image-text retrieval datasets demonstrate the superiority of DCH-SCR over several state-of-the-art cross-modal retrieval methods. Huanqiang Zeng, Yifan Shi 0001, Jianqing Zhu, Chih-Hsien Hsia, Kai-Kuang Ma |
IEEE Trans. Multim. | 6 |
| 2022 | CMUA-Watermark: A Cross-Model Universal Adversarial Watermark for Combating DeepfakesabstractMalicious applications of deepfakes (i.e., technologies generating target facial attributes or entire faces from facial images) have posed a huge threat to individuals' reputation and security. To mitigate these threats, recent studies have proposed adversarial watermarks to combat deepfake models, leading them to generate distorted outputs. Despite achieving impressive results, these adversarial watermarks have low image-level and model-level transferability, meaning that they can protect only one facial image from one specific deepfake model. To address these issues, we propose a novel solution that can generate a Cross-Model Universal Adversarial Watermark (CMUA-Watermark), protecting a large number of facial images from multiple deepfake models. Specifically, we begin by proposing a cross-model universal attack pipeline that attacks multiple deepfake models iteratively. Then, we design a two-level perturbation fusion strategy to alleviate the conflict between the adversarial watermarks generated by different facial images and models. Moreover, we address the key problem in cross-model optimization with a heuristic approach to automatically find the suitable attack step sizes for different models, further weakening the model-level conflict. Finally, we introduce a more reasonable and comprehensive evaluation method to fully test the proposed method and compare it with existing ones. Extensive experimental results demonstrate that the proposed CMUA-Watermark can effectively distort the fake facial images generated by multiple deepfake models while achieving a better performance than existing methods. Our code is available at https://github.com/VDIGPKU/CMUA-Watermark. Yongtao Wang, Zhaoyu Chen 0001, Zhi Tang 0001, Jingdong Chen, Weisi Lin, Kai-Kuang Ma |
AAAI | 10 |
| 2022 | Uncertainty Inspired Underwater Image Enhancement
Zhenqi Fu, Yue Huang 0001, Xinghao Ding, Kai-Kuang Ma |
ECCV (18) | 5 |
| 2022 | MA-NET: Multi-Scale Attention-Aware Network for Optical Flow EstimationabstractExisting methods for optical flow estimation can perform well on images with small offsets of large objects. However, they could fail when there are a number of small and fast-moving objects. In particular, current deep learning methods often impose a down-sampling of the input data for a reduction of computational complexity. This practice inevitably incurs a loss of image details and causes small objects be ignored. To solve the problem, a multi-scale attention-aware network (MA-Net) is proposed, in which coarse-scale and fine-scale features are extracted in parallel and then attention is paid to them for producing the optical flow estimation. Extensive experiments conducted on the Sintel and KITTI2015 datasets show that the proposed MA-Net can capture fast-moving small objects with high accuracy and thus deliver superior performance over a number of state-of-the-art methods. Baojiang Zhong, Kai-Kuang Ma |
ICASSP | 3 |
| 2022 | Deep Rank Cross-Modal Hashing with Semantic Consistent for Image-Text RetrievalabstractCross-modal hashing retrieval approaches maps heterogeneous multi-modal data into a common hamming space to achieve efficient and flexible retrieval performance. However, existing cross-modal methods mainly exploit feature-level similarity between multi-modal data, the label-level similarity and relative ranking relationship between adjacent instances have been ignored. To address these problems, we propose a novel Deep Rank Cross-modal Hashing(DRCH) method that fully explores the intra-modal semantic similarity relationship. Firstly, DRCH preserves semantic similarity by combining both label-level and feature-level information. Secondly, the inherent gap between modalities are narrowed by proposing a ranking alignment loss function. Finally, the compact and efficient hash codes are optimized from the common semantic space. Extensive experiments on two real-world image-text retrieval datasets demonstrate the superiority of DRCH compared with several state-of-the-art(SOTA) methods. Huanqiang Zeng, Yifan Shi 0001, Jianqing Zhu, Kai-Kuang Ma |
ICASSP | 5 |
| 2022 | Rangeinet: Fast Lidar Point Cloud Temporal InterpolationabstractDue to the low scan rate of LiDAR sensors, LiDAR point cloud streams usually have a low frame rate, which is far below that of other sensors such as cameras. This could incur frame rate mismatch while conducting multi-sensor data fusion. LiDAR point cloud temporal interpolation aims to synthesize the non-existing intermediate frame between input frames to improve the frame rate of point clouds. However, the existing methods heavily depend on 3D scene flow or 2D flow estimation, which yield huge computational complexity and obstacles in real-time applications. To resolve this issue, we propose a fast and non-flow involved method, which analyzes the LiDAR point cloud by exploiting its corresponding 2D range images (RIs). Specifically, we develop a Siamese context extractor containing asymmetrical convolution kernels to learn the shape context and spatial feature of RIs, and the 3D space-time convolutions are introduced to precisely capture the temporal characteristics. Experimental results have clearly shown that our method is much faster than the state-of-the-art LiDAR point cloud temporal interpolation methods on various datasets, while delivering either comparable or superior frame interpolation performance. Lili Zhao 0001, Xuhu Lin, Wenyi Wang 0005, Kai-Kuang Ma |
ICASSP | 4 |
| 2022 | Anisotropic Edge Detection in Catadioptric ImagesabstractCatadioptric images are produced in omnidirectional vision systems and can be expressed on Riemannian manifolds. The existing edge detectors are operated either in Euclidean space, or on Riemannian manifolds with isotropic image filtering. In this paper, a new type of edge detection is proposed—it is operated on Riemannian manifolds with anisotropic image filtering. For that, an anisotropic image filtering kernel on Riemannian manifolds is derived by solving the anisotropic heat equation embedded with Riemannian metric. With this kernel, a novel anisotropic edge detector is then developed. Compared to an edge detector operated in Euclidean space, our edge detector is more suitable for catadioptric images, since their geometric structure information will be taken into account in the detection process. Compared to existing edge detectors customized for catadioptric images, the new edge detector has a higher efficiency in preserving image edges and thus can produce more true positives. Enzhuang Zheng, Baojiang Zhong, Kai-Kuang Ma |
ICIP | 3 |
| 2022 | Object-Scale Adaptive Optical Flow Estimation Network
Baojiang Zhong, Kai-Kuang Ma |
PRICAI (3) | 3 |
| 2022 | Point Cloud Quality Assessment via 3D Edge Similarity MeasurementabstractIn this letter, a new full-reference metric is presented to assess the perceptual quality of the point clouds (PCs). The human visual system (HVS) always shows a high sensitivity to the three-dimensional (3D) edge features inherent in the PCs. With this motivation, the three-dimensional edge similarity-based model (TDESM) is proposed, which makes the first attempt to apply 3D Difference of Gaussian (3D-DOG) on point cloud quality assessment (PCQA). Specifically, the 3D edge features are captured by convolving the dual-scale 3D-DOG filters with both reference and distorted PCs. The quality scores of distorted PCs are generated by combining the 3D edge similarity measured from different scales. The experiments are conducted on four publicly available PCQA datasets, i.e., Torlig2018, M-PCCD, ICIP2020, and SJTU-PCQA. Compared with multiple state-of-the-art PCQA metrics, our proposed approach is able to be higher consistent with the subjective perception on the PCs. Zian Lu, Hailiang Huang 0002, Huanqiang Zeng, Junhui Hou, Kai-Kuang Ma |
IEEE Signal Process. Lett. | 5 |
| 2022 | A Direction-Decoupled Non-Local Attention Network for Single Image Super-ResolutionabstractThenon-local attentionmechanism has often been exploited in deep learning to capturelong-range dependencies(LRDs) from the same image for enhancing the performance of various image processing methods. However, the initially proposed non-local attention process inevitably yields extremely-high computation complexity, sinceallthe feature points are involved in computing the LRDs. To address this concern, a recently proposedcriss-cross network(CCNet), which has arecurrent criss-cross attention(RCCA) module, is used to compute the LRDs by involving only a small set of feature points for significantly reducing computation. Motivated by the RCCA, a noveldirection-decoupled non-local attention(DNA) module is proposed in this paper that is able to further reduce the computation complexity of RCCA by half approximately. To verify the performance of our new non-local attention module, a DNA network is developed for conducting single image super-resolution (SISR). Extensive experimental results have clearly demonstrated the superiority of using our DNA network for SISR when compared with that of state-of-the-art methods. Zijiang Song, Baojiang Zhong, Jiahuan Ji, Kai-Kuang Ma |
IEEE Signal Process. Lett. | 4 |
| 2022 | Real-Time Scene-Aware LiDAR Point Cloud Compression Using Semantic Prior RepresentationabstractExisting LiDAR point cloud compression (PCC) methods tend to treat compression as afidelityissue, without sufficiently addressing itsmachine perceptionaspect. The latter issue is often encountered by the decoder agents that might aim to conduct scene-understanding related tasks only, such as computing the localization information. For tackling this challenge, a novel LiDAR PCC system is proposed to compress the point cloud geometry, which contains aback channelfor allowing the decoder to initiate such request to the encoder. The key success of our PCC method lies in our proposedsemantic prior representation(SPR) and its lossy encoding algorithm with variable precision to generate the final bitstream; the entire process is fast and achieves real-time performance. Note that our SPR is a compact and effective representation of three-dimensional (3D) input point clouds, and it consists oflabels, predictions, andresiduals. These information can be generated by first exploiting ascene-aware object segmentationto a set of 2D range images (frames) individually, which were generated from the 3D point clouds via a projection process. Based on the generated labels, the pixels associated with those moving objects are considered as noisy information and should be removed for not only saving bit budget on transmission but also, most importantly, improving the accuracy of localization computed at the decoder. Experimental results conducted on the commonly-used test dataset have shown that our proposed system outperforms the MPEG’s G-PCC (TMC13-v14.0) in a large bitrate range. In fact, the performance gap will become even larger when more and/or large moving objects are involved in the input point clouds. Lili Zhao 0001, Kai-Kuang Ma, Zhili Liu, Qian Yin 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | A Spatial and Geometry Feature-Based Quality Assessment Model for the Light Field ImagesabstractThis paper proposes a new full-reference image quality assessment (IQA) model for performing perceptual quality evaluation on light field (LF) images, called the spatial and geometry feature-based model (SGFM). Considering that the LF image describe both spatial and geometry information of the scene, the spatial features are extracted over the sub-aperture images (SAIs) by using contourlet transform and then exploited to reflect the spatial quality degradation of the LF images, while the geometry features are extracted across the adjacent SAIs based on 3D-Gabor filter and then explored to describe the viewing consistency loss of the LF images. These schemes are motivated and designed based on the fact that the human eyes are more interested in the scale, direction, contour from the spatial perspective and viewing angle variations from the geometry perspective. These operations are applied to the reference and distorted LF images independently. The degree of similarity can be computed based on the above-measured quantities for jointly arriving at the final IQA score of the distorted LF image. Experimental results on three commonly-used LF IQA datasets show that the proposed SGFM is more in line with the quality assessment of the LF images perceived by the human visual system (HVS), compared with multiple classical and state-of-the-art IQA models. Hailiang Huang 0002, Huanqiang Zeng, Junhui Hou, Jing Chen 0001, Jianqing Zhu, Kai-Kuang Ma |
IEEE Trans. Image Process. | 6 |
| 2022 | Screen Content Video Quality Assessment Model Using Hybrid Spatiotemporal FeaturesabstractIn this paper, a full-reference video quality assessment (VQA) model is designed for the perceptual quality assessment of the screen content videos (SCVs), called the hybrid spatiotemporal feature-based model (HSFM). The SCVs are of hybrid structure including screen and natural scenes, which are perceived by the human visual system (HVS) with different visual effects. With this consideration, the three dimensional Laplacian of Gaussian (3D-LOG) filter and three dimensional Natural Scene Statistics (3D-NSS) are exploited to extract the screen and natural spatiotemporal features, based on the reference and distorted SCV sequences separately. The similarities of these extracted features are then computed independently, followed by generating the distorted screen and natural quality scores for screen and natural scenes. After that, an adaptive screen and natural quality fusion scheme through the local video activity is developed to combine them for arriving at the final VQA score of the distorted SCV under evaluation. The experimental results on the Screen Content Video Database (SCVD) and Compressed Screen Content Video Quality (CSCVQ) databases have shown that the proposed HSFM is more in line with the perceptual quality assessment of the SCVs perceived by the HVS, compared with a variety of classic and latest IQA/VQA models. Huanqiang Zeng, Hailiang Huang 0002, Junhui Hou, Jiuwen Cao, Yongtao Wang, Kai-Kuang Ma |
IEEE Trans. Image Process. | 6 |
| 2021 | Rpattack: Refined Patch Attack on General Object DetectorsabstractNowadays, general object detectors like YOLO and Faster R-CNN as well as their variants are widely exploited in many applications. Many works have revealed that these detectors are extremely vulnerable to adversarial patch attacks. The perturbed regions generated by previous patch-based attack works on object detectors are very large which are not necessary for attacking and perceptible for human eyes. To generate much less but more efficient perturbation, we propose a novel patch-based method for attacking general object detectors. Firstly, we propose a patch selection and refining scheme to find the pixels which have the greatest importance for attack and remove the inconsequential perturbations gradually. Then, for a stable ensemble attack, we balance the gradients of detectors to avoid over-optimizing one of them during the training phase. Our RPAttack can achieve an amazing missed detection rate of 100% for both Yolo v4 and Faster R-CNN while only modifies 0.32% pixels on VOC 2007 test set. Our code is available at https://github.com/VDIGPKU/RPAttack. Yongtao Wang, Zhaoyu Chen 0001, Zhi Tang 0001, Kai-Kuang Ma |
ICME | 6 |
| 2021 | Cascading Scene and Viewpoint Feature Learning for Pedestrian Gender RecognitionabstractPedestrian gender recognition plays an important role in smart city. To effectively improve the pedestrian gender recognition performance, a new method, called cascading scene and viewpoint feature learning (CSVFL), is proposed in this article. The novelty of the proposed CSVFL lies on the joint consideration of two crucial challenges in pedestrian gender recognition, namely, scene and viewpoint variation. For that, the proposed CSVFL starts with the scene transfer (ST) scheme, followed by the viewpoint adaptation (VA) scheme in a cascading manner. Specifically, the ST scheme exploits the key pedestrian segmentation network to extract the key pedestrian masks for the subsequent key pedestrian transfer generative adversarial network, with the goal of encouraging the input pedestrian image to have the similar style to the target scene while preserving the image details of the key pedestrian as much as possible. Afterward, the obtained scene-transferred pedestrian images are fed to train the deep feature learning network with the VA scheme, in which each neuron will be enabled/disabled for different viewpoints depending on whether it has contribution on the corresponding viewpoint. Extensive experiments conducted on the commonly used pedestrian attribute data sets have demonstrated that the proposed CSVFL approach outperforms multiple recently reported pedestrian gender recognition methods. Huanqiang Zeng, Jianqing Zhu, Jiuwen Cao, Yongtao Wang, Kai-Kuang Ma |
IEEE Internet Things J. | 6 |
| 2021 | Single Image Super-Resolution Using Asynchronous Multi-Scale NetworkabstractAn existing multi-scale residual network (MSRN) has demonstrated its success on conducting the single image super-resolution (SISR) task. The MSRN consists of a number of multi-scale residual blocks (MSRBs), and each MSRB performs convolutions by exploiting two different sizes of windows for conducting multi-scale feature extraction. The smaller window is used to extract image features at a low scale, while the larger one is used for a high scale. To significantly reduce the number of parameters involved in the MSRB, a new feature extraction module, called the asynchronous multi-scale block (AMB), is proposed in this paper. It is based on the fact that the larger window used in the MSRB can be replaced by two smaller windows without affecting the original MSRB's function. Consequently, by replacing each MSRB with our AMB, an asynchronous multi-scale network (AMNet) is then constructed, which can yield a significant reduction on computational complexity. This means that more AMBs can be used in our AMNet to deliver superior SISR performance, while maintaining the same or comparable computational complexity to that of the MSRN. To consolidate all image features generated from all scales, a new fusion scheme, called the adaptive feature fusion block (AFFB), is proposed that weights the extracted features according to their importance for further increasing SISR's performance. Extensive experimental results have clearly shown the superiority of our proposed AMNet when compared with multiple state-of-the-arts. Jiahuan Ji, Baojiang Zhong, Kai-Kuang Ma |
IEEE Signal Process. Lett. | 3 |
| 2021 | A Light Field Image Quality Assessment Model Based on Symmetry and Depth FeaturesabstractThis paper presents a new full-reference image quality assessment (IQA) method for conducting the perceptual quality evaluation of the light field (LF) images, called the symmetry and depth feature-based model (SDFM). Specifically, the radial symmetry transform is first employed on the luminance components of the reference and distorted LF images to extract their symmetry features for capturing the spatial quality of each view of an LF image. Second, the depth feature extraction scheme is designed to explore the geometry information inherited in an LF image for modeling its LF structural consistency across views. The similarity measurements are subsequently conducted on the comparison of their symmetry and depth features separately, which are further combined to achieve the quality score for the distorted LF image. Note that the proposed SDFM that explores the symmetry and depth features is conformable to the human vision system, which identifies the objects by sensing their structures and geometries. Extensive simulation results on the dense light fields dataset have clearly shown that the proposed SDFM outperforms multiple classical and recently developed IQA algorithms on quality evaluation of the LF images. Huanqiang Zeng, Junhui Hou, Jing Chen 0001, Jianqing Zhu, Kai-Kuang Ma |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2020 | Single Image Super-Resolution Via A Progressive Mixture ModelabstractIn this paper, a progressive mixture model (PMM) for single image super-resolution is proposed. Our model consists of an offline training stage and an online reconstruction stage, and both stages are conducted progressively by exploiting a uniform iterative scheme. In the training stage, the training dataset is clustered into finer and finer groups, and a set of mixture models are sequentially learned. In the reconstruction stage, residuals of the low-resolution (LR) input image are estimated by using the trained mixture models at increasing levels, and then they are progressively added to the LR image for producing the high-resolution (HR) image. Extensive experimental simulation results have clearly shown that the proposed model consistently delivers highly accurate and visually pleasant HR images, compared to that of the state-of the-art image super-resolution methods. Run Su, Baojiang Zhong, Jiahuan Ji, Kai-Kuang Ma |
ICIP | 4 |
| 2020 | Learning Local and Global Priors for JPEG Image Artifacts RemovalabstractLossy compression will inevitably introduce image artifacts in the decoded image and degrade the image quality. In recent years, convolutional neural network (CNN) has been exploited for removing compression artifacts with great success. However, most existing CNN-based methods only utilize image's local prior without considering the global prior on the training of their networks. In this letter, a novel CNN, called the local and global priors network (LGPNet), is proposed that simultaneously learns both the local and the global priors for removing compression image artifacts. To achieve this goal, a dual-attention unit (DAU) is developed and incorporated into the well-known U-Net architecture for learning a better local prior. Meanwhile, the global prior is also learned from the entire image via our proposed global prior network. Extensive experimental results have clearly demonstrated that our proposed LGPNet is able to effectively remove image artifacts and greatly improve the image quality of JPEG-compressed images. Yongtao Wang, Haihua Xie, Kai-Kuang Ma |
IEEE Signal Process. Lett. | 4 |
| 2020 | Screen Content Video Quality Assessment: Subjective and Objective StudyabstractIn this paper, we make the first attempt to study the subjective and objective quality assessment for the screen content videos (SCVs). For that, we construct the first large-scale video quality assessment (VQA) database specifically for the SCVs, called the screen content video database (SCVD). This SCVD provides 16 reference SCVs, 800 distorted SCVs, and their corresponding subjective scores, and it is made publicly available for research usage. The distorted SCVs are generated from each reference SCV with 10 distortion types and 5 degradation levels for each distortion type. Each distorted SCV is rated by at least 32 subjects in the subjective test. Furthermore, we propose the first full-reference VQA model for the SCVs, called the spatiotemporal Gabor feature tensor-based model (SGFTM), to objectively evaluate the perceptual quality of the distorted SCVs. This is motivated by the observation that 3D-Gabor filter can well stimulate the visual functions of the human visual system (HVS) on perceiving videos, being more sensitive to the edge and motion information that are often-encountered in the SCVs. Specifically, the proposed SGFTM exploits 3D-Gabor filter to individually extract the spatiotemporal Gabor feature tensors from the reference and distorted SCVs, followed by measuring their similarities and later combining them together through the developed spatiotemporal feature tensor pooling strategy to obtain the final SGFTM score. Experimental results on SCVD have shown that the proposed SGFTM yields a high consistency on the subjective perception of SCV quality and consistently outperforms multiple classical and state-of-the-art image/video quality assessment models. Shan Cheng, Huanqiang Zeng, Jing Chen 0001, Junhui Hou, Jianqing Zhu, Kai-Kuang Ma |
IEEE Trans. Image Process. | 6 |
| 2020 | 3D Point Cloud Attribute Compression Using Geometry-Guided Sparse Representationabstract3D point clouds associated with attributes are considered as a promising paradigm for immersive communication. However, the corresponding compression schemes for this media are still in the infant stage. Moreover, in contrast to conventional image/video compression, it is a more challenging task to compress 3D point cloud data, arising from the irregular structure. In this paper, we propose a novel and effective compression scheme for the attributes of voxelized 3D point clouds. In the first stage, an input voxelized 3D point cloud is divided into blocks of equal size. Then, to deal with the irregular structure of 3D point clouds, a geometry-guided sparse representation (GSR) is proposed to eliminate the redundancy within each block, which is formulated as an ℓ0-norm regularized optimization problem. Also, an inter-block prediction scheme is applied to remove the redundancy between blocks. Finally, by quantitatively analyzing the characteristics of the resulting transform coefficients by GSR, an effective entropy coding strategy that is tailored to our GSR is developed to generate the bitstream. Experimental results over various benchmark datasets show that the proposed compression scheme is able to achieve better rate-distortion performance and visual quality, compared with state-of-the-art methods. Shuai Gu, Junhui Hou, Huanqiang Zeng, Hui Yuan 0001, Kai-Kuang Ma |
IEEE Trans. Image Process. | 5 |
| 2020 | Image Interpolation Using Multi-Scale Attention-Aware Inception NetworkabstractA new multi-scale deep learning (MDL) framework is proposed and exploited for conducting image interpolation in this paper. The core of the framework is a seeding network that needs to be designed for the targeted task. For image interpolation, a novel attention-aware inception network (AIN) is developed as the seeding network; it has two key stages: 1) feature extraction based on the low-resolution input image; and 2) feature-to-image mapping to enlarge image's size or resolution. Note that the designed seeding network, AIN, needs to be trained with a matched training dataset at each scale. For that, multi-scale image patches are generated using our proposed pyramid cut, which outperforms the conventional image pyramid method by completely avoiding aliasing issue. After training, the trained AINs are then combined for processing the input image in the testing stage. Extensive experimental simulation results obtained from seven image datasets (comprising 359 images in total) have clearly shown that the proposed MAIN consistently delivers highly accurate interpolated images. Jiahuan Ji, Baojiang Zhong, Kai-Kuang Ma |
IEEE Trans. Image Process. | 3 |
| 2020 | Learning a Single Model With a Wide Range of Quality Factors for JPEG Image Artifacts RemovalabstractLossy compression brings artifacts into the compressed image and degrades the visual quality. In recent years, many compression artifacts removal methods based on convolutional neural network (CNN) have been developed with great success. However, these methods usually train a model based on one specific value or a small range of quality factors. Obviously, if the test images quality factor does not match to the assumed value range, then degraded performance will be resulted. With this motivation and further consideration of practical usage, a highly robust compression artifacts removal network is proposed in this paper. Our proposed network is a single model approach that can be trained for handling a wide range of quality factors while consistently delivering superior or comparable image artifacts removal performance. To demonstrate, we focus on the JPEG compression with quality factors, ranging from 1 to 60. Note that a turnkey success of our proposed network lies in the novel utilization of the quantization tables as part of the training data. Furthermore, it has two branches in parallel-i.e., the restoration branch and the global branch. The former effectively removes the local artifacts, such as ringing artifacts removal. On the other hand, the latter extracts the global features of the entire image that provides highly instrumental image quality improvement, especially effective on dealing with the global artifacts, such as blocking, color shifting. Extensive experimental results performed on color and grayscale images have clearly demonstrated the effectiveness and efficacy of our proposed single-model approach on the removal of compression artifacts from the decoded image. Yongtao Wang, Haihua Xie, Kai-Kuang Ma |
IEEE Trans. Image Process. | 4 |
| 2020 | Color Image Demosaicing Using Progressive Collaborative RepresentationabstractIn this paper, a progressive collaborative representation (PCR) framework is proposed that is able to incorporate any existing color image demosaicing method for further boosting its demosaicing performance. Our PCR consists of two phases: (i) offline training and (ii) online refinement. In phase (i), multiple training-and-refining stages will be performed. In each stage, a new dictionary will be established through the learning of a large number of feature-patch pairs, extracted from the demosaicked images of the current stage and their corresponding original full-color images. After training, a projection matrix will be generated and exploited to refine the current demosaicked image. The updated image with improved image quality will be used as the input for the next training-and-refining stage and performed the same processing likewise. At the end of phase (i), all the projection matrices generated as above-mentioned will be exploited in phase (ii) to conduct online demosaicked image refinement of the test image. Extensive simulations conducted on two commonly-used test datasets (i.e., the IMAX and Kodak) for evaluating the demosaicing algorithms have clearly demonstrated that our proposed PCR framework is able to constantly boost the performance of any image demosaicing method we experimented, in terms of the objective and subjective performance evaluations. Zhangkai Ni, Kai-Kuang Ma, Huanqiang Zeng, Baojiang Zhong |
IEEE Trans. Image Process. | 2 |
| 2020 | Light Field Image Quality Assessment via the Light Field CoherenceabstractIn this paper, a novel full-referenceimage quality assessment(IQA) method for evaluating the quality of the distortedlight field(LF) image against its reference LF image is proposed, called thelog-Gabor feature-basedlight field coherence (LGF-LFC). Based on the fact that to compare two LF images, it essentially boils down to measure howcoherentof these two LF images, we attempt to measure the degree of their LFcoherence(LFC). To pursue this goal, the salient features from the reference and distorted LF images under comparison need to be extracted. By considering that the Gabor feature has the ability to well characterize thehuman visual system(HVS) perception, and the special characteristics of the LF images, themulti-scale andsingle-scale Gabor feature extraction schemes are developed to extract the multi-scale log-Gabor features from thesub-aperture images(SAIs) and the single-scale log-Gabor feature from theepi-polar images(EPIs), respectively. Note that the former can reflect the image details (via the SAIs), while the latter indicates the viewing consistency (via the EPI’s depth information). The similarity measurements are subsequently conducted on the comparison of their SAIs and that of their EPIs separately, followed by combining them together for arriving at the final score. Extensive simulation results have clearly demonstrated that the proposed LGF-LFC is more consistent with the perception of the HVS on the quality evaluation of the LF images than multiple classical and state-of-the-art IQA methods. Huanqiang Zeng, Junhui Hou, Jing Chen 0001, Kai-Kuang Ma |
IEEE Trans. Image Process. | 5 |
| 2020 | PNEN: Pyramid Non-Local Enhanced NetworksabstractExisting neural networks proposed for low-level image processing tasks are usually implemented by stacking convolution layers with limited kernel size. Every convolution layer merely involves in context information from a small local neighborhood. More contextual features can be explored as more convolution layers are adopted. However it is difficult and costly to take full advantage of long-range dependencies. We propose a novel non-local module, Pyramid Non-local Block, to build up connection between every pixel and all remain pixels. The proposed module is capable of efficiently exploiting pairwise dependencies between different scales of low-level structures. The target is fulfilled through first learning a query feature map with full resolution and a pyramid of reference feature maps with downscaled resolutions. Then correlations with multi-scale reference features are exploited for enhancing pixel-level feature representation. The calculation procedure is economical considering memory consumption and computational cost. Based on the proposed module, we devise a Pyramid Non-local Enhanced Networks for edge-preserving image smoothing which achieves state-of-the-art performance in imitating three classical image smoothing algorithms. Additionally, the pyramid non-local block can be directly incorporated into convolution neural networks for other image restoration tasks. We integrate it into two existing methods for image denoising and single image super-resolution, achieving consistently improved performance. Feida Zhu 0003, Chaowei Fang, Kai-Kuang Ma |
IEEE Trans. Image Process. | 3 |
| 2019 | A Log-Gabor Feature-Based Quality Assessment Model for Screen Content ImagesabstractIn this paper, an image quality assessment (IQA) model for conducting objective evaluations of screen content images (S-CIs) is proposed, called the log-Gabor feature-based model (LGFM). From the standpoint of signal representation, the log-Gabor filters outperform the classical Gabor filters since the outputs of the log-Gabor filters are more consistent with the perception of visual cortex in human visual system (HVS). Furthermore, the following two remarkable characteristics of the log-Gabor filters are highly beneficial to develop a more accurate IQA model; i.e., (i) zero response at the DC, and (ii) stronger response at high frequencies. In our proposed L-GFM, the log-Gabor filters are used to extract features from the luminance of the reference SCIs and that of the distorted SCIs for measuring their degree of similarity. Together with the measurements from the other two chrominance components, the final LGFM score will be arrived at the output of the pooling stage. Extensive simulation results have shown that our proposed LGFM is highly consistent with the human perception, compared to other state-of-the-art IQA models. Kai-Kuang Ma, Huanqiang Zeng |
ICIP | 2 |
| 2019 | Multi-Scale Defense of Adversarial ImagesabstractDeep learning has achieved great success in image classification. However, recent researches show that existing deep learning-based classifiers remain weak for recognizing adversarial images. In this paper, an effective multi-scale defense method is proposed to solve the problem. In our method, an input image is first evolved by Gaussian kernels of different intensities to generate a multi-scale representation of the image. These evolved images are then fed into a classifier trained with a multi-scale strategy to yield multi-scale confidences. Finally, an average confidence is exploited to generate classification result. Furthermore, by monitoring the change of confidence values during the image evolution process, our method is able to achieve an indication of the attacking risk. Experimental results show that our method performs favorably against a number of state-of-the-art methods. Jiahuan Ji, Baojiang Zhong, Kai-Kuang Ma |
ICIP | 3 |
| 2019 | Shape Matching Based on Rectangularized Curvature Scale-Space MapsabstractThe curvature scale-space (CSS) technique is a shape descriptor in the MPEG-7 standard. To match a shape under query, a common approach is to exploit a simplified CSS map, which is produced by representing each arch-shaped contour of the original CSS map by a vertical line segment with the same height. However, we consider this approach is oversimplified, since some very useful information presented in the original CSS map has been totally ignored. To solve the problem, a new CSS map approximation is proposed in this paper, which is a rectangularized bin circumscribing each arch-shaped contour of the CSS map so that both the height and width characteristics are simultaneously captured and preserved. For conducting shape matching, an efficient method is further proposed, which has a significantly more succinct structure than the existing CSS matching technique used in MPEG-7. Simulation results have shown that the proposed method can yield superior performance. Baojiang Zhong, Kai-Kuang Ma |
ICIP | 3 |
| 2019 | Confusion-Aware Convolutional Neural Network for Image Classification
Liguang Yan, Baojiang Zhong, Kai-Kuang Ma |
ICONIP (1) | 3 |
| 2019 | Multi-label learning with multi-label smoothing regularization for vehicle re-identification
Jinhui Hou, Huanqiang Zeng, Jianqing Zhu, Jing Chen 0001, Kai-Kuang Ma |
Neurocomputing | 6 |
| 2019 | Predictor-corrector image interpolation
Baojiang Zhong, Kai-Kuang Ma, Zhifang Lu |
J. Vis. Commun. Image Represent. | 2 |
| 2018 | Selecting Informative Frames for Action Recognition with Partial ObservationsabstractGiven a video clip that contains only one type of action (e.g., golfing), the goal of action recognition is to recognize this action category from a given set of action types. To deliver fast response for practical video applications, existing works have been endevouring on processing the leading frames of the input video. In our view, only the informative key frames extracted from this `partial video' should be used for performing action recognition task. This will not only further speed up action recognition process due to less amount of data to be processed but also achieve higher recognition accuracy owing to more distinctive features presented to the learning network. For that, a novel a two-stage learning network architecture is proposed in this paper that consists of aselection network(S-net) and arecognition network(R-net). The S-net is a relatively-shallow network designed to efficiently identify informative key frames, while the R-net is a deep network to perform the final action recognition. In the S-net, a key frame selection criterion is further proposed for identifying informative key frames. Extensive experiments based on two benchmark datasets, UCF101 and HMDB51, have been conducted and clearly shown that our approach significantly outperforms existing state-of-the-art methods. Yanjun Zhu, Gang Yu 0002, Junsong Yuan 0001, Kai-Kuang Ma |
ICIP | 4 |
| 2018 | A multi-order derivative feature-based quality assessment model for light field image
Huanqiang Zeng, Jing Chen 0001, Jianqing Zhu, Kai-Kuang Ma |
J. Vis. Commun. Image Represent. | 6 |
| 2018 | Screen Content Image Quality Assessment Using Multi-Scale Difference of GaussianabstractIn this paper, a novel image quality assessment (IQA) model for the screen content images (SCIs) is proposed by using multi-scale difference of Gaussian (MDOG). Motivated by the observation that the human visual system (HVS) is sensitive to the edges while the image details can be better explored in different scales, the proposed model exploits MDOG to effectively characterize the edge information of the reference and distorted SCIs at two different scales, respectively. Then, the degree of edge similarity is measured in terms of the smaller-scale edge map. Finally, the edge strength computed based on the larger-scale edge map is used as the weighting factor to generate the final SCI quality score. Experimental results have shown that the proposed IQA model for the SCIs produces high consistency with human perception of the SCI quality and outperforms the state-of-the-art quality models. Ying Fu 0004, Huanqiang Zeng, Lin Ma 0002, Zhangkai Ni, Jianqing Zhu, Kai-Kuang Ma |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2018 | A Gabor Feature-Based Quality Assessment Model for the Screen Content ImagesabstractIn this paper, an accurate and efficient full-reference image quality assessment (IQA) model using the extracted Gabor features, called Gabor feature-based model (GFM), is proposed for conducting objective evaluation of screen content images (SCIs). It is well-known that the Gabor filters are highly consistent with the response of the human visual system (HVS), and the HVS is highly sensitive to the edge information. Based on these facts, the imaginary part of the Gabor filter that has odd symmetry and yields edge detection is exploited to the luminance of the reference and distorted SCI for extracting their Gabor features, respectively. The local similarities of the extracted Gabor features and two chrominance components, recorded in the LMN color space, are then measured independently. Finally, the Gabor-feature pooling strategy is employed to combine these measurements and generate the final evaluation score. Experimental simulation results obtained from two large SCI databases have shown that the proposed GFM model not only yields a higher consistency with the human perception on the assessment of SCIs but also requires a lower computational complexity, compared with that of classical and state-of-the-art IQA models. The source code for the proposed GFM will be available at http://smartviplab.org/pubilcations/GFM.html. Zhangkai Ni, Huanqiang Zeng, Lin Ma 0002, Junhui Hou, Jing Chen 0001, Kai-Kuang Ma |
IEEE Trans. Image Process. | 6 |
| 2018 | Blurriness-Guided Unsharp MaskingabstractIn this paper, a highly-adaptive unsharp masking (UM) method is proposed and called the blurriness-guided UM, or BUM, in short. The proposed BUM exploits the estimated local blurriness as the guidance information to perform pixel-wise enhancement. The consideration of local blurriness is motivated by the fact that enhancing a highly-sharp or a highly-blurred image region is undesirable, since this could easily yield unpleasant image artifacts due to over-enhancement or noise enhancement, respectively. Our proposed BUM algorithm has two powerful adaptations as follows. First, the enhancement strength is adjusted for each pixel on the input image according to the degree of local blurriness measured at the local region of this pixel's location. All such measurements collectively form the blurriness map, from which the scaling matrix can be obtained using our proposed mapping process. Second, we also consider the type of layer-decomposition filter exploited for generating the base layer and the detail layer, since this consideration would effectively help to prevent over-enhancement artifacts. In this paper, the layer-decomposition filter is considered from the viewpoint of edge-preserving type versus non-edge-preserving type. Extensive simulations experimented on various test images have clearly demonstrated that our proposed BUM is able to consistently yield superior enhanced images with better perceptual quality to that of using a fixed enhancement strength or other state-of-the-art adaptive UM methods. Wei Ye 0005, Kai-Kuang Ma |
IEEE Trans. Image Process. | 2 |
| 2017 | Semantic image content filtering via edge-preserving scale-aware filterabstractIn this paper, we highlight a new filtering concept and methodology, called the semantic image content filtering (SICF), which aims to remove insignificant small details from the image while preserving its main structure. Such image content separation is not possible to achieve by using any conventional linear filter as it is essentially designed to perform frequency separation. To realize an effective SICF, a novel image filtering algorithm, called the edge-preserving scale-aware filter (ESF), is proposed in this paper. Our proposed ESF yields a significant improvement over a recently-developed scale-aware filter, called the rolling guidance filter (RGF). The key success of our ESF lies in the developed adaptive relative total variation filter (ARTVF), which replaces the RGF's Gaussian filter for generating a much improved initial guidance image. Extensive simulation results obtained from various test images have clearly demonstrated that the proposed ESF outperforms other state-of-the-art methods on conducting SICF task. That is, the semantically-important large-scale image structure has been better preserved, while the insignificant small details have been removed more effectively. Wei Ye 0005, Kai-Kuang Ma |
ICIP | 2 |
| 2017 | Blurriness-guided unsharp maskingabstractIt has been observed that enhancing a highly-blurred image region could often lead to unpleasant noise amplification. Motivated by this, an adaptive unsharp masking (UM) method is proposed in this paper, which incorporates the estimated local blurriness information into the enhancement process to adaptively determine the scaling factor for each pixel on the detail layer. To achieve this goal, a pixel-wise local blurriness estimation method is developed for generating a pixel-wise blurriness map, followed by individually converting each blurriness measurement on the map to a scaling factor via a mapping process. The proposed method not only avoids noise amplification in blurred regions but also addresses `high-level' considerations, such as photographer's original intention on making background more blurred for creating special aesthetic effect. Extensive simulations conducted on various test images have demonstrated that our approach is able to deliver much superior perceptual quality of enhanced images compared to other state-of-the-art UM methods. Wei Ye 0005, Kai-Kuang Ma |
ICIP | 2 |
| 2017 | Sum-of-gradient based fast intra coding in 3D-HEVC for depth map sequence (SOG-FDIC)
Jing Chen 0001, Huanqiang Zeng, Canhui Cai, Kai-Kuang Ma |
J. Vis. Commun. Image Represent. | 5 |
| 2017 | ESIM: Edge Similarity for Screen Content Image Quality AssessmentabstractIn this paper, an accurate full-reference image quality assessment (IQA) model developed for assessing screen content images (SCIs), called the edge similarity (ESIM), is proposed. It is inspired by the fact that the human visual system (HVS) is highly sensitive to edges that are often encountered in SCIs; therefore, essential edge features are extracted and exploited for conducting IQA for the SCIs. The key novelty of the proposed ESIM lies in the extraction and use of three salient edge features-i.e., edge contrast, edge width, and edge direction. The first two attributes are simultaneously generated from the input SCI based on a parametric edge model, while the last one is derived directly from the input SCI. The extraction of these three features will be performed for the reference SCI and the distorted SCI, individually. The degree of similarity measured for each above-mentioned edge attribute is then computed independently, followed by combining them together using our proposed edge-width pooling strategy to generate the final ESIM score. To conduct the performance evaluation of our proposed ESIM model, a new and the largest SCI database (denoted as SCID) is established in our work and made to the public for download. Our database contains 1800 distorted SCIs that are generated from 40 reference SCIs. For each SCI, nine distortion types are investigated, and five degradation levels are produced for each distortion type. Extensive simulation results have clearly shown that the proposed ESIM model is more consistent with the perception of the HVS on the evaluation of distorted SCIs than the multiple state-of-the-art IQA methods. Zhangkai Ni, Lin Ma 0002, Huanqiang Zeng, Jing Chen 0001, Canhui Cai, Kai-Kuang Ma |
IEEE Trans. Image Process. | 6 |
| 2016 | Screen content image quality assessment using edge modelabstractSince the human visual system (HVS) is highly sensitive to edges, a novel image quality assessment (IQA) metric for assessing screen content images (SCIs) is proposed in this paper. The turnkey novelty lies in the use of an existing parametric edge model to extract two types of salient attributes - namely, edge contrast and edge width, for the distorted SCI under assessment and its original SCI, respectively. The extracted information is subject to conduct similarity measurements on each attribute, independently. The obtained similarity scores are then combined using our proposed edge-width pooling strategy to generate the final IQA score. Hopefully, this score is consistent with the judgment made by the HVS. Experimental results have shown that the proposed IQA metric produces higher consistency with that of the HVS on the evaluation of the image quality of the distorted SCI than that of other state-of-the-art IQA metrics. Zhangkai Ni, Lin Ma 0002, Huanqiang Zeng, Canhui Cai, Kai-Kuang Ma |
ICIP | 5 |
| 2016 | Low complexity depth intra coding in 3D-HEVC based on depth classificationabstractThe latest high efficiency video coding-based three dimensional video coding (3D-HEVC) exploits sophisticated intra prediction scheme to improve the coding performance of the depth video, but incurring heavy computational complexity. To address this problem, a low complexity depth intra coding method is presented for 3D-HEVC based on depth classification. Firstly, a database of depth prediction units (PUs) with three kinds of complexities is collected based on their optimal intra prediction mode. Then, the histogram of oriented gradient (HOG) features are extracted on these established database to train the classifier using support vector machine (SVM). For the current depth PU, the trained classifier is applied to determine its most possible complexity class so as to select the corresponding modes for involving the mode decision process. Experimental results show that the proposed method is able to significantly reduce the computational complexity while keeping almost the same coding performance of depth video and video quality of the synthesized view, compared with the exhaustive mode decision in 3D-HEVC. Huijie Zheng, Jianqing Zhu, Huanqiang Zeng, Jing Chen 0001, Canhui Cai, Kai-Kuang Ma |
VCIP | 6 |
| 2016 | Quad binary pattern and its application in mean-shift tracking
Huanqiang Zeng, Jing Chen 0001, Xiaolin Cui, Canhui Cai, Kai-Kuang Ma |
Neurocomputing | 5 |
| 2016 | Multiple description video coding based on adaptive data reuse
Meng Dong, Huanqiang Zeng, Jing Chen 0001, Canhui Cai, Kai-Kuang Ma |
J. Vis. Commun. Image Represent. | 5 |
| 2016 | Gradient Direction for Screen Content Image Quality AssessmentabstractIn this letter, we make the first attempt to explore the usage of the gradient direction to conduct the perceptual quality assessment of the screen content images (SCIs). Specifically, the proposed approach first extracts the gradient direction based on the local information of the image gradient magnitude, which not only preserves gradient direction consistency in local regions, but also demonstrates sensitivities to the distortions introduced to the SCI. A deviation-based pooling strategy is subsequently utilized to generate the corresponding image quality index. Moreover, we investigate and demonstrate the complementary behaviors of the gradient direction and magnitude for SCI quality assessment. By jointly considering them together, our proposed SCI quality metric outperforms the state-of-the-art quality metrics in terms of correlation with human visual system perception. Zhangkai Ni, Lin Ma 0002, Huanqiang Zeng, Canhui Cai, Kai-Kuang Ma |
IEEE Signal Process. Lett. | 5 |
| 2016 | Convolutional Edge Diffusion for Fast Contrast-guided Image InterpolationabstractA recently introduced image interpolation method, called the contrast-guided interpolation (CGI), has shown superior performance on producing high-quality interpolated image. However, its iterative edge diffusion (IED) process for diffusing continuous-valued directional variation (DV) fields inevitably incurs high computational complexity due to its iterative optimization process. The key objective of this letter lies in how to greatly reduce the computation of this diffusion process while maintaining CGI's superior performance on its interpolated image. The novelty of this letter started with a critical observation as follows. Since each diffused DV field needs to be thresholded for generating a binary contrast-guided decision map (CDM) in the subsequent step, such binarization operation will definitely destroy the fidelity that was preserved previously through the data term of the IED's energy functional. Therefore, the data term is lifted in our approach to yield a new energy functional. It turns out that the diffusion equation derived from this simplified functional is, in fact, the well-known heat equation, from which a highly attractive property of the heat equation can be exploited for conducting diffusion. That is, given a desired amount of diffusion to yield, it can be realized by simply convolving the DV field with a Gaussian kernel once, rather than gradually updating the DV field through iterations. Note that the variance of the Gaussian kernel corresponds to the amount of diffusion desired. As a result, the total computation time is significantly reduced. Extensive simulation results have shown that the proposed CED can generate nearly identical CDMs as those produced by the IED, while only requiring about 1/10 of its computation time. By replacing the IED with the proposed CED in the CGI framework, the total run time of our fast CGI is only 1/4 of the original CGI's on average. Wei Ye 0005, Kai-Kuang Ma |
IEEE Signal Process. Lett. | 2 |
| 2015 | Multiple Description Coding for Multi-view Video
Jing Chen 0001, Canhui Cai, Xiaolan Wang 0006, Huanqiang Zeng, Kai-Kuang Ma |
ACIVS | 5 |
| 2015 | A location-aware scale-space method for salient object detectionabstractMany existing saliency detection methods made an assumption that the salient object is on the center of the image and incorporated such center-biased assumption in the design of their algorithms. Obviously, this is not always proper to set, especially for those imageries acquired by unmanned monitoring system or device (e.g., surveillance camera), in which the salient object could appear in any location within the image. Consequently, the resulted saliency detection performance could be greatly degraded. In this paper, an existing hypercomplex Fourier transform (HFT) based saliency detection algorithm is investigated and modified for improving the saliency detection performance. In details, we remove its prior assumption on `center bias' and exploit a location-aware strategy to identify the optimal saliency map across multiple scales of the image. Extensive simulation results have justified that the proposed location-aware HFT-based approach clearly outperforms existing five state-of-the-art algorithms on saliency detection. Dan Xiang, Baojiang Zhong, Kai-Kuang Ma |
ICIP | 3 |
| 2015 | Feature histogram equalization for feature contrast enhancement
Kai-Kuang Ma |
J. Vis. Commun. Image Represent. | 2 |
| 2015 | SIFT-flow-based color correction for multi-view video
Huanqiang Zeng, Kai-Kuang Ma, Canhui Cai |
Signal Process. Image Commun. | 2 |
| 2015 | Color Image Demosaicing Using Iterative Residual InterpolationabstractA recently developed demosaicing methodology, called residual interpolation (RI), has demonstrated superior performance over the conventional color-component difference interpolation. However, it has been observed that the existing RI-based methods fail to fully exploit the potential of RI strategy on the reconstruction of the most important G channel, as only the R and B channels are restored through the RI strategy. Since any reconstruction error introduced in the G channel will be carried over into the demosaicing process of the other two channels, this makes the restoration of the G channel highly instrumental to the quality of the final demosaiced image. In this paper, a novel iterative RI (IRI) process is developed for reconstructing a highly accurate G channel first; in essence, it can be viewed as an iterative refinement process for the estimation of those missing pixel values on the G channel. The key novelty of the proposed IRI process is that all the three channels will mutually guide each other until a stopping criterion is met. Based on the restored G channel, the mosaiced R and B channels will be, respectively, reconstructed by exploiting the existing RI method without iteration. Extensive simulations conducted on two commonly-used test datasets for demosaicing algorithms have demonstrated that our algorithm has achieved the best performance in most cases, compared with the existing state-of-the-art demosaicing methods on both objective and subjective performance evaluations. Wei Ye 0005, Kai-Kuang Ma |
IEEE Trans. Image Process. | 2 |
| 2014 | Image demosaicing by using iterative residual interpolationabstractA new demosaicing approach has been introduced recently, which is based on conducting interpolation on the generated residual fields rather than on the color-component difference fields as commonly practiced in most demosaicing methods. In view of its attractive performance delivered by such residual interpolation (RI) strategy, a new RI-based demosaicing method is proposed in this paper that has shown much improved performance. The key success of our approach lies in that the RI process is iteratively deployed to all the three channels for generating a more accurately reconstructed G channel, from which the R channel and the B channel can be better reconstructed as well. Extensive simulations conducted on two commonly-used test datasets have clearly demonstrated that our algorithm is superior to the existing state-of-the-art demosaicing methods, both on objective performance evaluation and on subjective perceptual quality. Wei Ye 0005, Kai-Kuang Ma |
ICIP | 2 |
| 2014 | Bipartite graph-based mismatch removal for wide-baseline image matching
Kai-Kuang Ma |
J. Vis. Commun. Image Represent. | 2 |
| 2014 | Common Visual Pattern Discovery via Directed GraphabstractA directed graph (or digraph) approach is proposed in this paper for identifying all the visual objects commonly presented in the two images under comparison. As a model, the directed graph is superior to the undirected graph, since there are two link weights with opposite orientations associated with each link of the graph. However, it inevitably draws two main challenges: 1) how to compute the two link weights for each link and 2) how to extract the subgraph from the digraph. For 1), a novel n-ranking process for computing the generalized median and the Gaussian link-weight mapping function are developed that basically map the established undirected graph to the digraph. To achieve this graph mapping, the proposed process and function are applied to each vertex independently for computing its directed link weight by not only considering the influences inserted from its immediately adjacent neighboring vertices (in terms of their link-weight values), but also offering other desirable merits-i.e., link-weight enhancement and computational complexity reduction. For 2), an evolutionary iterative process for solving the non-cooperative game theory is exploited to handle the non-symmetric weighted adjacency matrix. The abovementioned two stages of processes will be conducted for each assumed scale-change factor, experimented over a range of possible values, one factor at a time. If there is a match on the scale-change factor under experiment, the common visual patterns with the same scale-change factor will be extracted. If more than one pattern are extracted, the proposed topological splitting method is able to further differentiate among them provided that the visual objects are sufficiently far apart from each other. Extensive simulation results have clearly demonstrated the superior performance accomplished by the proposed digraph approach, compared with those of using the undirected graph approach. Kai-Kuang Ma |
IEEE Trans. Image Process. | 2 |
| 2013 | A novel multiple description video coding based on data reuseabstractA novel H.264-based multiple description coding (MDC) framework, called data reuse MDC (DR-MDC), is proposed in this paper. The input video sequence is first down-sampled by a factor of two in both horizontal and vertical directions, respectively, on each frame to generate four sub-sequences, followed by grouping them into two descriptions via the quincunx manner. In each description, one sub-sequence is directly encoded by applying the H.264/AVC encoder, while the other is examined at each macroblock (MB) to determine whether the encoding of the current MB should be conducted or skipped completely based on the following criterion: If the MB is considered locating in a homogeneous or still background region, the encoding process will be skipped. Otherwise, the neighboring prediction algorithm will be used to predict the pixel values of this MB, and the resultant prediction errors will be further encoded and transmitted. Experimental results have shown that the proposed DR-MDC scheme is more error resilient and yields better reconstructed video than the existing state-of-the-art MDC methods. Meng Dong, Canhui Cai, Kai-Kuang Ma |
ICIP | 3 |
| 2013 | Curvature scale-space of open curves: Theory and shape representationabstractThe problem of extending the curvature scale-space (CSS) technique to represent open curves is addressed. Various approaches for dealing with the endpoint problem of open curves are considered, and one is selected which allows us to handle the evolution of the open curves as a special case of the evolution of closed curves. The convergence theory of evolved open curves is established, and the CSS shape representation is investigated. Baojiang Zhong, Kai-Kuang Ma, Jiwen Yang |
ICIP | 2 |
| 2013 | Layered moving-object segmentation for stereoscopic video using motion and depth information
Yibin Chen, Canhui Cai, Kai-Kuang Ma, Xiaolan Wang 0006 |
J. Vis. Commun. Image Represent. | 3 |
| 2013 | Contrast-Guided Image InterpolationabstractIn this paper a contrast-guided image interpolation method is proposed that incorporates contrast information into the image interpolation process. Given the image under interpolation, four binary contrast-guided decision maps (CDMs) are generated and used to guide the interpolation filtering through two sequential stages: 1) the 45(°) and 135(°) CDMs for interpolating the diagonal pixels and 2) the 0(°) and 90(°) CDMs for interpolating the row and column pixels. After applying edge detection to the input image, the generation of a CDM lies in evaluating those nearby non-edge pixels of each detected edge for re-classifying them possibly as edge pixels. This decision is realized by solving two generalized diffusion equations over the computed directional variation (DV) fields using a derived numerical approach to diffuse or spread the contrast boundaries or edges, respectively. The amount of diffusion or spreading is proportional to the amount of local contrast measured at each detected edge. The diffused DV fields are then thresholded for yielding the binary CDMs, respectively. Therefore, the decision bands with variable widths will be created on each CDM. The two CDMs generated in each stage will be exploited as the guidance maps to conduct the interpolation process: for each declared edge pixel on the CDM, a 1-D directional filtering will be applied to estimate its associated to-be-interpolated pixel along the direction as indicated by the respective CDM; otherwise, a 2-D directionless or isotropic filtering will be used instead to estimate the associated missing pixels for each declared non-edge pixel. Extensive simulation results have clearly shown that the proposed contrast-guided image interpolation is superior to other state-of-the-art edge-guided image interpolation methods. In addition, the computational complexity is relatively low when compared with existing methods; hence, it is fairly attractive for real-time image applications. Kai-Kuang Ma |
IEEE Trans. Image Process. | 2 |
| 2012 | Content-adaptive temporal consistency enhancement for depth videoabstractThe video plus depth format, which is composed of the texture video and the depth video, has been widely used for free viewpoint TV. However, the temporal inconsistency is often encountered in the depth video due to the error incurred in the estimation of the depth values. This will inevitably deteriorate the coding efficiency of depth video and the visual quality of synthesized view. To address this problem, a content-adaptive temporal consistency enhancement (CTCE) algorithm for the depth video is proposed in this paper, which consists of two sequential stages: (1) classification of stationary and non-stationary regions based on the texture video, and (2) adaptive temporal consistency filtering on the depth video. The result of the first stage is used to steer the second stage so that the filtering process will be conducted in an adaptive manner. Extensive experimental results have shown that the proposed CTCE algorithm can effectively mitigate the temporal inconsistency in the original depth video and consequently improve the coding efficiency of depth video and the visual quality of synthesized view. Huanqiang Zeng, Kai-Kuang Ma |
ICIP | 2 |
| 2012 | Prediction-Compensated Polyphase Multiple Description Image Coding With Adaptive Redundancy ControlabstractIn this paper, a novel multiple description coding (MDC) system is proposed, consisting of two thrust contributions: 1) a new polyphase MDC scheme, called the prediction-compensated polyphase MDC (PCP-MDC); and 2) an adaptive redundancy control (ARC) scheme for yielding optimal tradeoff between coding efficiency and error resilience. The PCP-MDC partitions each quincunx-downsampled description into two subdescriptions, called the primary subdescription (PS) and dual subdescription (DS). The PS is encoded by the H.264/AVC intra coding, while the DS is subject to prediction coding based on the reconstructed PS. For prediction, the mode-guided directional prediction algorithm is developed to conduct a fast and accurate prediction for the DS. For the second thrust contribution, two fundamental issues regarding the inserted redundancy are addressed in the proposed ARC scheme: quality and quantity. For the quality issue, the residuals of cross prediction are inserted into each description. For the quantity issue, the amount of redundancy bits allocated to each description is determined according to the network condition; for that, the probability of channel failure is incorporated into mathematical formulation on the derivation of optimal redundancy allocation condition. To implement the derived optimal condition, a fast estimation algorithm of the rate-distortion function is then developed, followed by exploiting successive approaching algorithm to identify the optimal partition of the target bitrate between the primary part of the description and its complementary part (i.e., the redundancy part). We also develop a post-processing filter, called the switching Gaussian filter, to remove the granular artifacts that tend to occur in any polyphase MDC approach when two descriptions are received and encoded at low bitrates. Extensive simulation results have consistently shown that the proposed MDC system significantly outperforms other state-of-the-art MDC methods on coding performance with much lower computational complexity. Kai-Kuang Ma, Canhui Cai |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2011 | Common visual pattern discovery via directed graph modelabstractIn this paper, a novel directed graph (or digraph) model-based approach is proposed to discover visual patterns commonly shared by two images. Unlike the conventional undirected graph model with only one weight value on each link, the directed graph model has two link weights, one for each direction of the link. In our work, it takes two phases to compute the link weights. First, the principle of pairwise spatial consistency is exploited to generate the initial link weights. The entire initial weights are then modified to generate the relative link weights by further considering the “relativeness” of neighboring vertices for each vertex using our proposed n-ranking value. Consequently, the resulted relative link weights are more robust to combat various commonly encountered scenarios such as large viewpoint variations and inaccurate feature descriptors. Based on the relative link weights, the strongly-connected subgraph for each scale value under consideration is then extracted from the graph by applying the non-cooperative game theory for handling non-symmetric adjacency matrix issue. All the vertices (i.e., point-to-point feature correspondences) belonging to the subgraph are collectively denoted as one common visual pattern. Preliminary simulation results have reasonably demonstrated the efficacy and robustness of the proposed method on discovering common visual patterns. Kai-Kuang Ma |
ICIP | 2 |
| 2011 | Fast Mode Decision for Multiview Video Coding Using Mode CorrelationabstractExhaustive mode decision has been exploited in multiview video coding for effectively improving the coding efficiency, but at the expense of yielding much higher computational complexity. In this paper, a fast mode decision algorithm, called the mode correlation-based mode decision (MCMD), is proposed to speed up the encoding process by reducing the number of the modes required to be checked. In our approach, all the prediction modes are first categorized into five motion-activity classes, and only one of them will be chosen to identify the optimal mode in a hierarchical manner, as follows. For each macroblock (MB), the proposed MCMD algorithm always begins with checking whether the rate-distortion cost computed at the SKIP mode (i.e., Class 1) is below an adaptive threshold for providing a possible early termination chance. If this early termination condition is not met, one of the remaining four motion-activity classes will be chosen for further mode checking according to the analysis of the predicted motion vector (PMV) of the current MB. The above-mentioned adaptive threshold and PMV are derived by exploiting the mode correlation between the current MB and a set of adjacent MBs (i.e., region of support) in the current view and its neighboring view. Experimental results have shown that compared with exhaustive mode decision, which is a default approach set in the joint multiview video model (JMVM) reference software, the proposed MCMD algorithm achieves a reduction of the computational complexity by 73.39% on average, while incurring only 0.07 dB loss in peak signal-to-noise ratio (PSNR) and 2.22% increment on the total bit rate. Huanqiang Zeng, Kai-Kuang Ma, Canhui Cai |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2011 | Multitemporal Image Change Detection Using Undecimated Discrete Wavelet Transform and Active ContoursabstractIn this paper, an unsupervised change detection method for satellite images is proposed. Owing to its robustness against noise, the undecimated discrete wavelet transform is exploited to obtain a multiresolution representation of the difference image, which is obtained from two satellite images acquired from the same geographical area but at different time instances. A region-based active contour model is then applied to the multiresolution representation of the difference image for segmenting the difference image into the “changed” and “unchanged” regions. The proposed change detection method has been conducted on two types of image data sets, i.e., the synthetic aperture radar images and the optical images. The change detection results are compared with several state-of-the-art techniques. The extensive simulation results clearly show that the proposed change detection method consistently yields superior performance. Turgay Çelik 0001, Kai-Kuang Ma |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2010 | Histogram-offset-based color correction for multi-view video codingabstractIn multi-view video system, variations of different camera setups (for example, camera positions, lighting conditions and camera characteristics) might cause discrepancies on the luminance and chrominance components of different views. From the viewpoint of source compression, this will lead to inaccurate inter-view prediction and lower coding efficiency. In this paper, a histogram-offset-based color correction method is developed for benefitting multi-view video coding. First, disparity estimation is conducted on the rank-transformed domain to identify the maximum matching regions between the reference view and the target view. Within the identified matching regions, the histograms of the reference view and the target view are then calculated, respectively. By using an iterative thresholding approach, a histogram offset is generated and exploited to correct the target view. Experimental results have shown that the proposed color correction method outperforms the histogram matching method on the improvement of coding efficiency. Yibin Chen, Kai-Kuang Ma, Canhui Cai |
ICIP | 2 |
| 2010 | Mode-correlation-based early termination mode decision for multi-view video codingabstractExhaustive mode decision is exploited in multi-view video coding for effectively improving the coding efficiency, but at the expense of yielding higher computational complexity. In this paper, a fast mode decision algorithm, called the mode-correlation-based early termination (MET), is proposed. For each macroblock, the proposed MET algorithm always starts with checking whether the rate-distortion (RD) cost computed at the SKIP mode is below an adaptive threshold for providing a possible early termination chance. This adaptive threshold is calculated by using the mode correlation between the current macroblock and a set of adjacent macroblocks in the current view and its neighboring view. Experimental results have shown that compared with exhaustive mode decision, which is a default approach set in the JMVM reference software, the proposed MET algorithm achieves a reduction of the computational complexity by 65.91% and the total bit rate by 0.98% on average, while incurring only 0.06 dB loss in peak signal-to-noise ratio (PSNR). Huanqiang Zeng, Kai-Kuang Ma, Canhui Cai |
ICIP | 2 |
| 2010 | Edge-contrast-guided image interpolation using directional variation field diffusionabstractIn this paper, a novel interpolation method is proposed, called the edge-contrast-guided interpolation (ECGI), which takes the edge contrast into consideration. Similar to the existing edge-guided methods, for those edge pixels under interpolation, the proposed ECGI will conduct the interpolation along their associated edge's direction. However, the novelty of our proposed ECGI lies in the treatment of non-edge pixels: We additionally view those non-edge pixels as `edge' pixels, if they are locating in the vicinity of true edges. In this case, each `edge' pixel is subject to conduct edge-guided interpolation along the same direction of the true edge that influences it. In our work, we use the edge contrast to determine how far a true edge could yield such influence on its nearby pixels; consequently, the stronger the contrast, the more pixels in its vicinity will be viewed as `edge' pixels. To implement this idea, we apply a variational approach over the directional variation (DV) fields to obtain new diffused DV fields for conducting edge-guided interpolation. Compared with the state-of-the-art interpolation methods, our method is able to deliver superior performance, yielding more clear edges without introducing ringing or other artifacts, while enjoying fairly low computational complexity. Wei Zhe, Kai-Kuang Ma, Canhui Cai |
ICIP | 2 |
| 2010 | Motion activity-based block size decision for multi-view video codingabstractMotion estimation and disparity estimation using variable block sizes have been exploited in multi-view video coding to effectively improve the coding efficiency, but at the expense of yielding higher computational complexity. In this paper, a fast block size decision algorithm, called motion activity-based block size decision (MABSD), is proposed. In our approach, the various motion estimation and disparity estimation block sizes are classified into four classes, and only one of them will be chosen to further identify the optimal block size within that class according to the measured motion activity of the current macroblock. The above-mentioned motion activity can be measured by the maximum city-block distance of a set of motion vectors taken from the adjacent macroblocks in the current view and its neighboring view. Experimental results have shown that compared with exhaustive block size decision, which is a default approach set in the JMVM reference software, the proposed MABSD algorithm achieves a reduction of computational complexity by 42% on average, while incurring only 0.01 dB loss in peak signal-to-noise ratio (PSNR) and 1% increment on the total bit rate. Huanqiang Zeng, Kai-Kuang Ma, Canhui Cai |
PCS | 2 |
| 2010 | Stochastic super-resolution image reconstruction
Jing Tian 0002, Kai-Kuang Ma |
J. Vis. Commun. Image Represent. | 2 |
| 2010 | Error-Resilient H.264/AVC Video Transmission Using Two-Way Decodable Variable Length Data BlockabstractStandard video coders utilize variable length coding (VLC) to obtain more data compression in addition to what lossy coding has achieved at the expense of making the compressed bitstream very vulnerable to channel errors. Even a 1-bit error incurred in the bitstream may cause the follow-up bitstream to be either erroneously decoded or completely undecodable, and this could further result in error propagation. To mitigate this phenomenon, a new VLC coding scheme is proposed in this paper, called the two-way decodable variable length data block (TDVLDB), which allows the compressed bitstream to bebidirectionallydecodable without exploiting data partitioning. The proposed TDVLDB scheme is able to effectively recover more uncorrupted data from the corrupted packets. Furthermore, it is able to correct some, if not all, channel errors of a finite-length burst error. To effectively identify the location of the first actual error incurred within the current slice, abitstream similarity measurement(BSM) algorithm is proposed. Note that the proposed TDVLDB scheme is generic in the sense that it can be exploited in any image or video coding framework as long as it involves the use of VLC and requires error-resilience capability. In this paper, the proposed TDVLDB is incorporated into the H.264/advanced video coding (AVC) coder to evaluate its error-resilience performance in terms of rate-distortion coding efficiency. Compared with the baseline H.264/AVC coding, the TDVLDB-incorporated H.264/AVC-based coding scheme has demonstrated significant objective and subjective video quality improvements when the bitstream is transmitted over error-prone channels. Kai-Kuang Ma |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2010 | Hierarchical Intra Mode Decision for H.264/AVCabstractThe intra mode prediction via exhaustive mode decision exploited in the H.264/advanced video coding effectively improves the coding efficiency, but at the expense of yielding higher computational complexity. In this letter, a fast intra mode decision algorithm, called the hierarchical intra mode decision (HIMD), is proposed to speed up the mode decision process by reducing the number of modes required to be checked for each macroblock. The novelty of the proposed HIMD algorithm lies at the following accounts. 1) An early decision with adaptive thresholding is developed for the mode decision of the luma component. 2) The candidate modes are selected according to their Hadamard distances and prediction directions. 3) Only one of the hierarchical paths will be chosen to compute its least rate-distortion cost. Experimental results have shown that the proposed HIMD algorithm achieves a reduction of 85.75% computational complexity on average, while incurring only 0.164 dB loss in peak signal-to-noise ratio (PSNR) and 2.336% increment on the total bit rate compared with that of exhaustive mode decision, which is a default approach set in the joint model reference software. Huanqiang Zeng, Kai-Kuang Ma, Canhui Cai |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2010 | Unsupervised Change Detection for Satellite Images Using Dual-Tree Complex Wavelet TransformabstractIn this paper, an unsupervised change-detection method for multitemporal satellite images is proposed. The algorithm exploits the inherent multiscale structure of the dual-tree complex wavelet transform (DT-CWT) to individually decompose each input image into one low-pass subband and six directional high-pass subbands at each scale. To avoid illumination variation issue possibly incurred in the low-pass subband, only the DT-CWT coefficient difference resulted from the six high-pass subbands of the two satellite images under comparison is analyzed in order to decide whether each subband pixel intensity has incurred a change. Such a binary decision is based on an unsupervised thresholding derived from a mixture statistical model, with a goal of minimizing the total error probability of change detection. The binary change-detection mask is thus formed for each subband, and all the produced subband masks are merged by using both the intrascale fusion and the interscale fusion to yield the final change-detection mask. For conducting the performance evaluation of change detection, the proposed DT-CWT-based unsupervised change-detection method is exploited for both the noise-free and the noisy images. Extensive simulation results clearly show that the proposed algorithm not only consistently provides more accurate detection of small changes but also demonstrates attractive robustness against noise interference under various noise types and noise levels. Turgay Çelik 0001, Kai-Kuang Ma |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2010 | On the Convergence of Planar Curves Under SmoothingabstractCurve smoothing has two important applications in computer vision and image processing: 1) the curvature scale-space (CSS) technique for shape analysis, and 2) the Gaussian filter for noise suppression. In this paper, we study how planar curves converge as they are smoothed with increasing scales. First, two types of convergence behavior are clarified. The coined term shrinkage refers to the reduction of arc-length of a smoothed planar curve, which describes the convergence of the curve latitudinally; and another coined term collapse refers to the movement of each point to its limiting position, which describes the convergence of the curve longitudinally. A systematic study on the shrinkage and collapse of three categories of curve models is then presented. The corner models helps to reveal how the local structures of planar curves collapse and what the smoothed curves may converge to. The sawtooth models allows us to gain insights regarding how noise is suppressed from noisy planar curves by the Gaussian filter. Our investigation on the closed curves shows that each curve collapses to a point at its center of mass. However, different curves may yield different limiting shapes at the infinity scale. Finally, based upon the derived results the performance of the CSS technique in corner detection and shape representation is analyzed, and a fast implementation method of the Gaussian filter for noise suppression is proposed. Baojiang Zhong, Kai-Kuang Ma |
IEEE Trans. Image Process. | 2 |
| 2009 | Stereoscopic video error concealment for missing frame recovery using disparity-based frame difference projectionabstractAt low bit-rate video communications, packet loss may easily cause whole-frame loss that, in return, leads to annoying frame drop phenomenon. In this paper, a novel error concealment algorithm is specifically developed for stereoscopic video, called the disparity-based frame difference projection (DFDP), to recover the lost frames at the decoder. The proposed DFDP contains three key components: 1) change detection, 2) disparity estimation, and 3) frame difference projection, which exploits both the intra-view frame difference from one view and interview correlation to estimate the lost frame in another view. The change region computed on the correctly received frame will be used to predict the change region between current missing frame and its previous frame through the estimated disparity, which is the summation of the estimated global disparity and the estimated local disparity. Experimental results have shown that the proposed stereoscopic video error concealment method can effectively restore the lost frames at the decoder and deliver attractive performance, in terms of objective measurement (in peak signal-to-noise ratio) and subjective visual quality. Yibin Chen, Canhui Cai, Kai-Kuang Ma |
ICIP | 3 |
| 2009 | Scale-Space Behavior of Planar-Curve CornersabstractThe curvature scale-space (CSS) technique is suitable for extracting curvature features from objects with noisy boundaries. To detect corner points in a multiscale framework, Rattarangsi and Chin investigated the scale-space behavior of planar-curve corners. Unfortunately, their investigation was based on an incorrect assumption, viz., that planar curves have no shrinkage under evolution. In the present paper, this mistake is corrected. First, it is demonstrated that a planar curve may shrink nonuniformly as it evolves across increasing scales. Then, by taking into account the shrinkage effect of evolved curves, the CSS trajectory maps of various corner models are investigated and their properties are summarized. The scale-space trajectory of a corner may either persist, vanish, merge with a neighboring trajectory, or split into several trajectories. The scale-space trajectories of adjacent corners may attract each other when the corners have the same concavity, or repel each other when the corners have opposite concavities. Finally, we present a standard curvature measure for computing the CSS maps of digital curves, with which it is shown that planar-curve corners have consistent scale-space behavior in the digital case as in the continuous case. Baojiang Zhong, Kai-Kuang Ma, Wenhe Liao |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2009 | Fast Mode Decision for H.264/AVC Based on Macroblock Motion ActivityabstractThe intra-mode and inter-mode predictions have been made available in H.264/AVC for effectively improving coding efficiency. However, exhaustively checking for all the prediction modes for identifying the best one (commonly referred to asexhaustivemodedecision) greatly increases computational complexity. In this paper, a fast mode decision algorithm, called themotionactivity-basedmodedecision(MAMD), is proposed to speed up the encoding process by reducing the number of modes required to be checked in a hierarchical manner, and is as follows. For each macroblock, the proposed MAMD algorithm always starts with checking the rate-distortion (RD) cost computed at the SKIP mode for a possible early termination, once the RD cost value is below a predetermined ldquolowrdquo threshold. On the other hand, if the RD cost exceeds another ldquohighrdquo threshold, then this indicates that only the intra modes are worthwhile to be checked. If the computed RD cost falls between the above-mentioned two thresholds, the remaining seven modes, which are classified into three motion activity classes in our work, will be examined, and only one of the three classes will be chosen for further mode checking. The above-mentioned motion activity can be quantitatively measured, which is equal to the maximum city-block length of the motion vector taken from a set of adjacent macroblocks (i.e., region of support, ROS). This measurement is then used to determine the most possible motion-activity class for the current macroblock. Experimental results have shown that, on average, the proposed MAMD algorithm reduces the computational complexity by 62.96%, while incurring only 0.059 dB loss in PSNR (peak signal-to-noise ratio) and 0.19% increment on the total bit rate compared to that of exhaustive mode decision, which is a default approach set in the JM reference software. Huanqiang Zeng, Canhui Cai, Kai-Kuang Ma |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2007 | Super-resolution Imaging Using Grid ComputingabstractThe super-resolution (SR) imaging is to overcome the inherent limitations of the image acquisition systems to produce high-resolution images from their low-resolution counterparts. In our recent work, the Markov chain Monte Carlo (MCMC) technique has been successfully developed and shown as a promising stochastic approach for addressing the SR problem. However, the MCMC SR approach requires substantial amounts of computational resources, for it not only needs to generate a huge number of samples, but also requires an exhaustive search for obtaining an optimal prior image model. To tackle the above computation challenge, Grid computing is introduced for tackling the SR problem in this paper. The computationally- intensive MCMC SR task is broke down into a set of independent and small sub-tasks, which are further distributed and implemented in the grid computing environment. Their respective results are finally assembled to produce a high-resolution image as the final result of the entire MCMC SR task. Experiments are conducted to show that grid computing can effectively accelerating the execution time of the MCMC SR algorithm. Jing Tian 0002, Kai-Kuang Ma |
CCGRID | 2 |
| 2007 | Expected Run-Time Distortion Based Scheduling for Scalable Video Transmission with Hybrid FEC/ARQ Error ControlabstractThe optimal packet scheduling for transmitting scalable media over the packet erasure networks has been extensively studied in the past. In the existing work, only retransmission is used for packet loss recovery. As a result, when the round-trip-time (RTT) of the network increases, the performance of the system degrades fast. In this paper, a scheduling scheme for hybrid FEC/ARQ error control is proposed. In our proposed scheme, the importance of both data packets and FEC packets are evaluated by considering several factors, such as data dependency structure of scalable video, transmission history of other packets, as well as the decoding deadline, and the most important packet is chosen to be sent out. In this way, our proposed scheme is able to achieve more stable playback video quality by taking advantage of both FEC and ARQ, as demonstrated by the experimental results. Tong Gan, Lu Gan 0002, Kai-Kuang Ma |
ICASSP (1) | 3 |
| 2007 | Rotation-invariant and scale-invariant Gabor features for texture image retrieval
Ju Han, Kai-Kuang Ma |
Image Vis. Comput. | 2 |
| 2007 | Undersampled Boundary Pre-/Postfilters for Low Bit-Rate DCT-Based Block CodersabstractIt has been well established that critically sampled boundary pre-/postfiltering operators can improve the coding efficiency and mitigate blocking artifacts in traditional discrete cosine transform-based block coders at low bit rates. In these systems, both the prefilter and the postfilter are square matrices. This paper proposes to use undersampled boundary pre- and postfiltering modules, where the pre-/postfilters are rectangular matrices. Specifically, the prefilter is a "fat" matrix, while the postfilter is a "tall" one. In this way, the size of the prefiltered image is smaller than that of the original input image, which leads to improved compression performance and reduced computational complexities at low bit rates. The design and VLSI-friendly implementation of the undersampled pre-/postfilters are derived. Their relations to lapped transforms and filter banks are also presented. Two design examples are also included to demonstrate the validity of the theory. Furthermore, image coding results indicate that the proposed undersampled pre-/postfiltering systems yield excellent and stable performance in low bit-rate image coding. Lu Gan 0002, Chengjie Tu, Jie Liang 0001, Trac D. Tran, Kai-Kuang Ma |
IEEE Trans. Image Process. | 5 |
| 2007 | Wiener Filter-Based Error Resilient Time-Domain Lapped TransformabstractIn this paper, the design of the error resilient time-domain lapped transform is formulated as a linear minimal mean-squared error problem. The optimal Wiener solution and several simplifications with different tradeoffs between complexity and performance are developed. We also prove the persymmetric structure of these Wiener filters. The existing mean reconstruction method is proven to be a special case of the proposed framework. Our method also includes as a special case the linear interpolation method used in DCT-based systems when there is no pre/postfiltering and when the quantization noise is ignored. The design criteria in our previous results are scrutinized and improved solutions are obtained. Various design examples and multiple description image coding experiments are reported to demonstrate the performance of the proposed method. Jie Liang 0001, Chengjie Tu, Lu Gan 0002, Trac D. Tran, Kai-Kuang Ma |
IEEE Trans. Image Process. | 5 |
| 2006 | H.264-based Multiple Description Video Coder and Its DSP ImplementationabstractIn this paper, a novel H.264-based multiple description coding (MDC) framework, called the prediction-based spatial polyphase transform (PSPT) multiple description video coding, is introduced to enhance the error-resilience ability and reduce the computational complexity of the H.264 codec. The proposed PSPT-MDC down-samples each input frame in both the horizontal and the vertical directions, forming four subframes. Instead of directly coding and transporting all the four subframes, two of them are predicted from the other two subframes. The proposed H.264-based PSPT-MDC video codec is then implemented on the TI TMS320C6416 digital signal processor with some optimizations to demonstrate its feasibility and attractive performance, in terms of the decoded video quality and error-resilience capability. Canhui Cai, Kai-Kuang Ma |
ICIP | 3 |
| 2006 | Reducing video-quality fluctuations for streaming scalable video using unequal error protection, retransmission, and interleavingabstractForward error correction based multiple description (MD-FEC) transcoding for transmitting embedded bitstream over the packet erasure networks has been extensively studied in the past. In the existing work, a single embedded source bitstream, e.g., the bitstream of a group of pictures (GOP) encoded using three-dimensional set partitioning in hierarchical trees is optimally protected unequal error protection (UEP) in the rate-distortion sense. However, most of the previous work on transmitting embedded video using MD-FEC assumed that one GOP is transmitted only once, and did not consider the chance of retransmission. This may lead to noticeable video quality variations due to varying channel conditions. In this paper, a novel window-based packetization scheme is proposed, which combats bursty packet loss by combining the following three techniques: UEP, retransmission, and GOP-level interleaving. In particular, two retransmission mechanisms, namely segment-wise retransmission and byte-wise retransmission, are proposed based on different types of receiver feedback. Moreover, two levels of rate allocations are introduced: intra-GOP rate allocation minimizes the distortion of individual GOP; while inter-GOP rate allocation intends to reduce video quality fluctuations by adaptively allocating bandwidth according to video signal characteristics and client buffer status. In this way, more consistent video quality can be achieved under various packet loss probabilities, as demonstrated by our experimental results. Tong Gan, Lu Gan 0002, Kai-Kuang Ma |
IEEE Trans. Image Process. | 3 |
| 2006 | A switching median filter with boundary discriminative noise detection for extremely corrupted imagesabstractA novel switching median filter incorporating with a powerful impulse noise detection method, called the boundary discriminative noise detection (BDND), is proposed in this paper for effectively denoising extremely corrupted images. To determine whether the current pixel is corrupted, the proposed BDND algorithm first classifies the pixels of a localized window, centering on the current pixel, into three groups--lower intensity impulse noise, uncorrupted pixels, and higher intensity impulse noise. The center pixel will then be considered as "uncorrupted," provided that it belongs to the "uncorrupted" pixel group, or "corrupted." For that, two boundaries that discriminate these three groups require to be accurately determined for yielding a very high noise detection accuracy--in our case, achieving zero miss-detection rate while maintaining a fairly low false-alarm rate, even up to 70% noise corruption. Four noise models are considered for performance evaluation. Extensive simulation results conducted on both monochrome and color images under a wide range (from 10% to 90%) of noise corruption clearly show that our proposed switching median filter substantially outperforms all existing median-based filters, in terms of suppressing impulse noise while preserving image details, and yet, the proposed BDND is algorithmically simple, suitable for real-time implementation and application. Pei-Eng Ng, Kai-Kuang Ma |
IEEE Trans. Image Process. | 2 |
| 2005 | A MCMC approach for Bayesian super-resolution image reconstructionabstractIn this paper, we consider the super-resolution image reconstruction problem. We propose a Markov chain Monte Carlo (MCMC) approach to find the maximum a posterior probability (MAP) estimation of the unknown high-resolution image. Firstly, Gaussian Markov random field (GMRF) is exploited for modeling the prior probability distribution of the unknown high-resolution image. Then, a MCMC technique (in particular, the Gibbs sampler) is introduced to generate samples from the posterior probability distribution to compute the MAP estimation of the unknown high-resolution image, which is obtained as the mean of the samples. Moreover, we derive a bound on the convergence time of the proposed MCMC approach. Finally, the experimental results are presented to verify the superior performance of the proposed approach and the validity of the proposed bound. Jing Tian 0002, Kai-Kuang Ma |
ICIP (1) | 2 |
| 2005 | A new state-space approach for super-resolution image sequence reconstructionabstractSuper-resolution imaging is to overcome the inherent limitations of image acquisition to create high-resolution images from their low-resolution counterparts. In this paper, a novel state-space approach is proposed to incorporate the temporal correlations among the low-resolution observations into the framework of the Kalman filtering. The proposed approach exploits both the temporal correlations information among the high-resolution images and the temporal correlations information among the low-resolution images to improve the quality of the reconstructed high-resolution sequence. Experimental results show that the proposed framework is superior to bi-linear interpolation, bi-cubic spline interpolation and the conventional Kalman filter approach, due to the consideration of the temporal correlations among the low-resolution images. Jing Tian 0002, Kai-Kuang Ma |
ICIP (1) | 2 |
| 2005 | Accurate optical flow estimation in noisy sequences by robust tensor-driven anisotropic diffusionabstractIn this paper, a new tensor-driven anisotropic diffusion filtering method is proposed for achieving accurate optical flow estimation in noisy image sequences. The novelties of our approach are: (1) robust tensor-driven anisotropic diffusion computation, (2) new thresholding criterion for normalization function. By utilizing the decomposed eigenvectors and eigenvalues of the 3D structure tensor, the robust diffusion tensor is computed to steer the anisotropic filtering over the input image sequence. The moving orientations of the local spatio-temporal structures are precisely captured during the denoising process. For achieving more accurate diffusion tensor computation, a new thresholding criterion is developed in the normalization function to threshold the decomposed eigenvalues. As compared with that of existing methods, our experimental results demonstrate much improved accuracy on both motion field classification and optical flow estimation. Kai-Kuang Ma |
ICIP (3) | 2 |
| 2005 | Adaptive irregular pattern search with matching prejudgment for fast block-matching motion estimationabstractIn this paper, a simple and effective fast block matching algorithm (BMA) is proposed, called adaptive irregular pattern search with matching prejudgment (AIPS-MP). In the AIPS-MP, a dynamic search pattern, adaptive irregular pattern (AIP), is constructed for each block based on its spatial and temporal neighboring motion vectors (MVs). The AIP is used to quickly identify the best center point for the follow-up refined search to effectively reduce unnecessary intermediate searches and avoid the search trapping at local-minimum matching error position. The construction of the AIP jointly takes the tradeoff of rate and distortion into consideration and favors the differential coding of MVs. The matching prejudgment (MP) is imbedded in the AIPS and exploits an adaptive threshold determined for each block based on the prediction error of the temporal neighboring block. The simulation results show that the proposed AIPS-MP consistently achieves very close, and sometimes even higher, peak signal-to-noise ratio (PSNR) than the full search, while achieving 167.6-871.5 times of speed-up or computational gain, measured in terms of the average number of search points per MV generation. It also greatly outperforms two well-referenced fast BMAs, the diamond search and the motion vector field adaptive search technique (MVFAST), adopted in the MPEG-4 Verification Model and the Optimization Model, respectively, on the average PSNR and the search speed. Yao Nie, Kai-Kuang Ma |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2005 | Weighted unequal error protection for transmitting scalable object-oriented images over packet-erasure networksabstractIn this paper, we investigate the problem of transmitting embedded encoded object-oriented images over the packet-erasure networks. After giving a review of the existing combined unequal error protection (CUEP) and individual unequal error protection (IUEP) schemes, a novel weighted unequal error protection (WUEP) packetization scheme is proposed, which serves as an alternative to the existing methods. In our proposed framework, the embedded bitstreams of all concerned image objects are packetized into multiple description packet streams before transmission. Two levels of rate allocation are introduced: intraobject rate allocation provides unequal error protection to the embedded bitstream of each object and minimizes its associated mean distortion; interobject rate allocation aims at minimizing the weighted mean distortion by adaptively allocating the rate budget among different objects according to their importance. Furthermore, our proposed packetization scheme ensures independent access and manipulation of individual image object. A detailed comparison between CUEP, IUEP, and WUEP is presented along with the experimental results, so that one can choose the most suitable approach according to the requirements. Tong Gan, Kai-Kuang Ma |
IEEE Trans. Image Process. | 2 |
| 2005 | Dual-plan bandwidth smoothing for layer-encoded videoabstractTraditional bandwidth smoothing techniques can be naturally supported by the renegotiated constant bit rate (RCBR) service model, but renegotiation failure in RCBR may cause buffer underflow and interrupt the playback of video. To address this concern, a novel dual-plan bandwidth smoothing (DBS) scheme is proposed in this paper by taking advantage of the SNR scalability of layer-encoded video. Upon renegotiation failure, the proposed scheme can adaptively discard certain enhancement layers to guarantee continuous video playback at the original frame rate. Experiments are carried out to demonstrate the validity of the proposed scheme. The impacts of renegotiation interval, granularity of enhancement layers, and playback buffer size on resulted video quality are also studied. From the simulation results, it is shown that the performance of the RCBR-based DBS scheme can be improved by 1) reducing the minimum time gap of renegotiation interval; 2) employing multilayer video encoding with finer granularity; and/or 3) increasing the playback buffer size. Tong Gan, Kai-Kuang Ma, Liren Zhang |
IEEE Trans. Multim. | 2 |
| 2004 | Modeling two-windows TCP behavior in differentiated services networks
Jianhua He 0001, Zongkai Yang, Zhen Fan 0004, Zuoyin Tang, Liren Zhang, Kai-Kuang Ma |
Comput. Commun. | 6 |
| 2003 | Sliding-window packetization for forward error correction based multiple description transcodingabstractForward error correction based multiple description (MD-FEC) transcoding, which performs unequal loss protection (ULP) for an embedded bitstream, allows robust video transmission over packet erasure channels. However, most of the existing works focus on rate-distortion optimization of individual encoding unit, e.g., a group of pictures (GOP), and do not examine the problem of rate allocation among different units. Such a transcoding strategy would lead to noticeable video quality variations when the video signal is highly nonstationary, and/or when large transmission rate fluctuations occur. In this paper, a novel window-based packetization scheme is proposed for reducing such quality variation through GOP interleaving. Performance evaluations are conducted by using a 3D-SPIHT embedded video encoder. Tong Gan, Kai-Kuang Ma |
ICASSP (5) | 2 |
| 2003 | On efficient implementation of oversampled linear phase perfect reconstruction filter banksabstractIn this paper, we first present an alternative way of generating oversampled linear phase perfect reconstruction filter banks (OSLPPRFB). We show that this method provides the minimal factorization of a subset of existing OSLPPRFB. The combination of the new structure and the conventional one leads to efficient implementations of a general class of OSLPPRFB. Possible application of the new scheme is discussed. Jie Liang 0001, Lu Gan 0002, Chengjie Tu, Trac D. Tran, Kai-Kuang Ma |
ICASSP (6) | 5 |
| 2003 | Motion field discontinuity classification for tensor-based optical flow estimationabstractA much more accurate classification scheme is proposed for structure tensor-based optical flow estimation to address the difficulties of interpreting motion field discontinuities. The key novelties of this approach are: (1) a scale-adaptive spatio-temporal filter; (2) a weighted structure tensor; (3) confidence measurements. Multiple motions of moving objects are matched by utilizing a spatio-temporal Gaussian filter with adaptive scale selection, which is steered by the condition number. To capture the neighborhood structure of local discontinuities, weighting the structure tensors is attempted. A new normalization function is exploited to facilitate accurate thresholding for confidence measurements. Experimental results demonstrate that these three novelties together effectively contribute much improved performance on motion field discontinuity classification compared with that of existing methods. Kai-Kuang Ma |
ICASSP (3) | 2 |
| 2003 | Oversampled lapped transforms via time-domain pre- and post-processingabstractThis paper introduces a large family of oversampled lapped transforms with symmetric basis functions. These new transforms are implemented by adding time-domain oversampled pre- and post- filters to the DCT and the IDCT, respectively. Structures and parameterizations of the corresponding pre-/postfilters are proposed. Two design examples along with some image coding results are presented to demonstrate the validity of the theory and the potential of the new transforms. Lu Gan 0002, Kai-Kuang Ma |
ICIP (3) | 2 |
| 2003 | Sliding-window packetization for unequal loss protection based multiple description codingabstractUnequal loss protection (ULP) for transmitting embedded bitstream over packet erasure networks has been extensively studied. However, most of the existing works focused on rate-distortion optimization of individual encoding unit, e.g., a group of pictures (GOP), which may lead to noticeable video quality variations due to varying channel conditions. In this paper, a novel sliding-window packetization scheme is proposed, which combat bursty packet loss through GOP interleaving. Two levels of rate allocation are introduced: intra-GOP rate allocation minimizes the distortion of individual GOP, and provides hybrid FEC/retransmission packet loss recovery; while inter-GOP rate allocation intends to reduce video quality fluctuations by adaptively allocating bandwidth according to video signal characteristics. Through this way, more consistent video quality can be achieved under various packet loss probability, as verified by our experimental results. Tong Gan, Kai-Kuang Ma |
ICIP (3) | 2 |
| 2003 | Unequal-arm adaptive rood pattern search for fast block-matching motion estimation in the JVT/H.26LabstractThe adaptive rood pattern search (ARPS) algorithm proposed by Nie and Ma (2002) has shown two to three times of search speed-up improvement over that of diamond search (DS) based on the MPEG-4 verification model encoding platform. In this paper, first we have shown that the distribution of motion vectors bears a rood shape. An improved ARPS algorithm, called ARPS-3 or unequal-arm ARPS, is then proposed and experimented on the JVT/H.26L JM encoding platform. Due to complex modes and multiframe prediction, a new normalized computational cost metrics is also proposed for objectively measuring the computational gain or search speed-up. Experimental results show that ARPS-3 has achieved superb performance on many accounts while maintaining fairly close rate-distortion performance compared with that of the full search. It also outperforms its predecessors, ARPS and ARPS-2 (or equal-arm ARPS). Kai-Kuang Ma, G. Qiu |
ICIP (1) | 1 |
| 2003 | Automatic video object segmentation via 3D structure tensorabstract3D structure tensor is an effective representation of the local motion information of video object (VO) and has been exploited for performing VO segmentation. However, existing 3D structure tensor-based VO segmentation approaches often yield inaccurate objects' boundaries, and high computation is needed for estimating dense motion field. To address these concerns, a new scheme is proposed in this paper by generating the spatial-constrained motion masks without computing dense motion field. For that, scale-adaptive spatio-temporal filtering steered by the condition number is developed to handle multiple motions contributed from different VOs. As rigid, and nonrigid VO motions need to be handled differently on mask generation, rigidity analysis is conducted based on standard deviation of correlation coefficients over a range of successive video frames in order to identify whether each video sequence frame contains rigid or nonrigid motion. Various masks, such as eigenmaps, coherency-measurement maps, and change-detection maps, are produced and fused for generating the final VO motion masks. With boundary refinement by graph-based spatial segmentation, experimental results present accurately segmented moving VOs using different kinds of test sequences. Kai-Kuang Ma |
ICIP (1) | 2 |
| 2003 | On efficient implementation of oversampled linear phase perfect reconstruction filter banksabstractIn this paper, we first present an alternative way of generating over-sampled linear phase perfect reconstruction filter banks (OSLP-PRFB). We show that this method provides the minimal factorization of a subset of existing OSLPPRFB. The combination of the new structure and the conventional one leads to efficient implementations of a general class of OSLPPRFB. Possible application of the new scheme is discussed. Jie Liang 0001, Lu Gan 0002, Chengjie Tu, Trac D. Tran, Kai-Kuang Ma |
ICME | 5 |
| 2003 | Error concealment for video transmission with dual multiscale Markov random field modelingabstractA novel error concealment algorithm based on a stochastic modeling approach is proposed as a post-processing tool at the decoder side for recovering the lost information incurred during the transmission of encoded digital video bitstreams. In our proposed scheme, both the spatial and the temporal contextual features in video signals are separately modeled using the multiscale Markov random field (MMRF). The lost information is then estimated using maximum a posteriori (MAP) probabilistic approach based on the spatial and temporal MMRF models; hence, a unified MMRF-MAP framework. To preserve the high frequency information (in particular, the edges) of the damaged video frames through iterative optimization, a new adaptive potential function is also introduced in this paper. Comparing to the existing MRF-based schemes and other traditional concealment algorithms, the proposed dual MMRF (DMMRF) modeling method offers significant improvement on both objective peak signal-to-noise ratio (PSNR) measurement and subjective visual quality of restored video sequence. Kai-Kuang Ma |
IEEE Trans. Image Process. | 2 |
| 2002 | Unsupervised semantic video objects segmentation over optical-flow fieldabstractAn unsupervised semantic video objects segmentation system is introduced in this paper, which is a region-based non-parametric spatio-temporal approach over optical-flow field. The proposed method overcomes multiple drawbacks inherited in existing supervised pixel-based parametric schemes. The unsupervised mechanism is realized by extracting the phase of the optical-flow field and forming the phase histogram to identify the number of dominant video objects contained within the video frame. Through extensive simulations, dominant video objects are automatically detected and segmented with high accuracy. The segmented VOs have semantic meaning that matches human being's perception; thus, the proposed segmentation system should be very useful to many applications encountered in multimedia, virtual reality and computer vision. Kai-Kuang Ma |
ICARCV | 1 |
| 2002 | Color distance histogram: a novel descriptor for color image segmentationabstractA novel color image descriptor, called color distance histogram (CDH), is proposed in this paper as a fundamental signal feature, readily to be exploited for various color image applications, such as indexing and retrieval, segmentation, and so on. To establish CDH, color image is first represented in CIE L*a*b* color space, followed by using CIE L*a*b*'s uniform color distance metric on computing the distance of each pixel with respect to the reference color. Consequently, CDH accurately reflects the degree of color similarity of individual pixel with respect to the reference color. To demonstrate the use and effectiveness of CDH, it is further extended into a set of CDHs, called dominant color profile (DCP), for color image segmentation. Experimental results clearly indicate that the proposed CDH-based or DCP-based segmentation method yields superior segmentation to other thresholding and clustering methods, in terms of accuracy, robustness, efficiency and computational complexity. Kai-Kuang Ma, Junxian Wang |
ICARCV | 1 |
| 2002 | Adaptive irregular pattern search with zero-motion prejudgement for fast block-matching motion estimationabstractA novel fast block matching algorithm, called adaptive irregular pattern search with zero-motion prejudgement (AIPS-ZMP), is proposed in this paper. The algorithm effectively exploits the spatial and temporal correlation existed among the motion vectors (MVs). In AIPS-ZMP, an adaptive irregular pattern (AIP) is dynamically formed for each block according to its spatial and temporal neighboring MVs to identify the most promising search center so that unnecessary intermediate searches and the risk of being trapped into local-minimum matching-error point are both avoided. The refined search is then performed from this center to find out the target MV. To further speed up the search, zero-motion prejudgement (ZMP) technique is embedded in the AIPS's framework to detect static blocks before invoking any search. Consequently, fairly large amount of computational gain is obtained with little degradation on picture quality. Compared with the fast BMA adopted in MEG-4 Verification Model (VMS) - diamond search (SD), our proposed AIPS-ZMP method improves the average peak signal-to-noise ratio (PSNR) by up to 0.67 dB with computational gain in the range of 2.12/spl sim/5.14 times. Yao Nie, Kai-Kuang Ma |
ICARCV | 2 |
| 2002 | Theory and lattice factorization of oversampled linear-phase perfect reconstruction filter banksabstractThis paper presents the theory and structure of a large family of oversampled linear-phase perfect reconstruction filter banks (OLPPRFBs). For such filter banks, we first derive the necessary existence conditions on the number of symmetric filters and antisymmetric filters. We then develop lattice factorizations of these OLPPRFBs, followed by two design examples to confirm the validity of the theory. Lu Gan 0002, Kai-Kuang Ma |
ICASSP | 2 |
| 2002 | On lattice factorization of symmetric-antisymmetric multifilter banksabstractWe introduce a new structure for symmetric-antisymmetric multiwavelets (SAMWTs) and symmetric-antisymmetric multifilter banks (SAMFBs). First, by exploring the connection between SAMFBs and traditional (scalar) linear phase perfect reconstruction filter banks (LPPRFBs), we show that the implementation and design of an SAMFB can be converted into that of a LPPRFB. Then, based on the lattice factorization for LPPRFBs, we propose a fast, modular, minimal structure for SAMFBs. To demonstrate the effectiveness of the proposed lattice structure, a multiplierless SAMWT design example is presented along with its application in image coding. Lu Gan 0002, Kai-Kuang Ma |
ICIP (1) | 2 |
| 2002 | Accurate optical flow estimation using adaptive scale-space and 3D structure tensorabstractComputing optical flow for image sequences is often an essential step to many image processing and computer vision applications. In this paper, a novel, unified optical flow estimation method is developed for simultaneously tackling the aperture problem and multiple motions, and consequently, yielding more accurate optical flow estimation. By integrating Gaussian scale-space with 3D structure tensor, the estimation difficulty encountered in multiple motions resulting from multiple video objects has been handled reasonably well. The obtained normal flow is then treated separately from the real flow, by further applying the least-squares estimation, with the assist of the automatic scale selection mechanism, to produce the estimated real flow. Our proposed automatic scale selection for spatial scale-space is developed from the viewpoint of numerical stability, and the condition number is exploited for adaptively choosing local scales (window sizes). For performance evaluation, we adopted the angular error as the quantitative measurement and used several benchmark image sequences. Experimental results show that the accuracy of our optical flow estimation method is superior to several leading algorithms. Kai-Kuang Ma |
ICIP (2) | 2 |
| 2002 | Dual-plan bandwidth smoothing for layer-encoded videoabstractTraditional bandwidth smoothing techniques can be supported naturally by the renegotiated constant bit rate (RCBR) service model, but renegotiation failure in RCBR may cause buffer underflow and interrupt the playback of video. To address this concern, a novel dual-plan bandwidth smoothing scheme is proposed by taking advantage of layer-encoded video. Upon renegotiation failure, the proposed scheme can adaptively discard enhancement layers to guarantee continuous video playback at the original frame rate. Experimental results are provided to demonstrate the validity of the proposed scheme and the impact of playback buffer size on the resulting video quality degradation. Tong Gan, Kai-Kuang Ma, Liren Zhang |
ICME (1) | 2 |
| 2002 | Region-based nonparametric optical flow segmentation with pre-clustering and post-clusteringabstractA region-based nonparametric video object segmentation over an optical-flow field is proposed to overcome the drawbacks inherited in pixel-based parametric approaches. The key novelties of this approach are: (1) motion field smoothing; (2) pre-clustering and post-clustering. By utilizing both spatial and temporal information extracted from the input video sequence, the raw optical-flow field is partitioned into homogeneous regions, with each region undergoing a common translational motion. Such an objective can be achieved through iterative spatio-temporal processing until the predetermined error-tolerance threshold is met. To facilitate fuzzy c-means clustering, pre-clustering and post-clustering are proposed. Experimental results demonstrate that they also effectively contribute a much improved performance in video object segmentation. Kai-Kuang Ma |
ICME (2) | 1 |
| 2002 | Using eigencolor normalization for illumination-invariant color object recognition
Zhenyong Lin, Junxian Wang, Kai-Kuang Ma |
Pattern Recognit. | 3 |
| 2002 | On the completeness of the lattice factorization for linear-phase perfect reconstruction filter banksabstractIn this letter, we re-examine the completeness of the lattice factorization for M-channel linear-phase perfect reconstruction filter bank (LPPRFB) with filters of the same length L=KM as discussed by Tran et al. (see IEEE Trans. Signal Processing, vol.48, p.133-47, Jan. 2000). We point out that the assertion of completeness is incorrect. Examples are presented to show that the proposed lattice structure of Tran et al. is not complete when K>2. In addition, we verify that the lattice structure is complete only when K/spl les/2. Lu Gan 0002, Kai-Kuang Ma, Truong Q. Nguyen, Trac D. Tran, Ricardo L. de Queiroz |
IEEE Signal Process. Lett. | 2 |
| 2002 | Fuzzy color histogram and its use in color image retrievalabstractA conventional color histogram (CCH) considers neither the color similarity across different bins nor the color dissimilarity in the same bin. Therefore, it is sensitive to noisy interference such as illumination changes and quantization errors. Furthermore, CCHs large dimension or histogram bins requires large computation on histogram comparison. To address these concerns, this paper presents a new color histogram representation, called fuzzy color histogram (FCH), by considering the color similarity of each pixel's color associated to all the histogram bins through fuzzy-set membership function. A novel and fast approach for computing the membership values based on fuzzy c-means algorithm is introduced. The proposed FCH is further exploited in the application of image indexing and retrieval. Experimental results clearly show that FCH yields better retrieval results than CCH. Such computing methodology is fairly desirable for image retrieval over large image databases. Ju Han, Kai-Kuang Ma |
IEEE Trans. Image Process. | 2 |
| 2002 | Adaptive rood pattern search for fast block-matching motion estimationabstractIn this paper, we propose a novel and simple fast block-matching algorithm (BMA), called adaptive rood pattern search (ARPS), which consists of two sequential search stages: 1) initial search and 2) refined local search. For each macroblock (MB), the initial search is performed only once at the beginning in order to find a good starting point for the follow-up refined local search. By doing so, unnecessary intermediate search and the risk of being trapped into local minimum matching error points could be greatly reduced in long search case. For the initial search stage, an adaptive rood pattern (ARP) is proposed, and the ARP's size is dynamically determined for each MB, based on the available motion vectors (MVs) of the neighboring MBs. In the refined local search stage, a unit-size rood pattern (URP) is exploited repeatedly, and unrestrictedly, until the final MV is found. To further speed up the search, zero-motion prejudgment (ZMP) is incorporated in our method, which is particularly beneficial to those video sequences containing small motion contents. Extensive experiments conducted based on the MPEG-4 Verification Model (VM) encoding platform show that the search speed of our proposed ARPS-ZMP is about two to three times faster than that of the diamond search (DS), and our method even achieves higher peak signal-to-noise ratio (PSNR) particularly for those video sequences containing large and/or complex motion contents. Yao Nie, Kai-Kuang Ma |
IEEE Trans. Image Process. | 2 |
| 2001 | A simplified lattice factorization for linear-phase perfect reconstruction filter bankabstractWe propose a simplified version of lattice factorization for linear-phase perfect reconstruction filter bank (LPPRFB) derived by T.D. Tran et al. (see IEEE Trans. Signal Processing. vol.48, no.1, p.133-47, Jan. 2000). The proposed new lattice structure spans the same class of LPPRFB, while substantially reducing free parameters in nonlinear optimization and saving computation cost in hardware implementation. To further address the importance of our proposed structure, we generalize our factorization to multidimensional LPPRFB (MD-LPPRFB), and show its effectiveness. Lu Gan 0002, Kai-Kuang Ma |
IEEE Signal Process. Lett. | 2 |
| 2001 | Analysis of a full-memory multidestination ARQ protocol over broadcast linksabstractBased on an assumption that a steady state exists in the full-memory multidestination automatic repeat request (ARQ) scheme, we propose a novel analytical method called steady-state function method (SSFM), to evaluate the performance of the scheme with any size of receiver buffer. For a wide range of system parameters, SSFM has higher accuracy on throughput estimation as compared to the conventional analytical methods. Jianhua He 0001, K. R. Subramanian, Liren Zhang, Kai-Kuang Ma |
IEEE Trans. Commun. | 4 |
| 2001 | Noise adaptive soft-switching median filterabstractExisting state-of-the-art switching-based median filters are commonly found to be nonadaptive to noise density variations and prone to misclassifying pixel characteristics at high noise density interference. This reveals the critical need of having a sophisticated switching scheme and an adaptive weighted median filter. We propose a novel switching-based median filter with incorporation of fuzzy-set concept, called the noise adaptive soft-switching median (NASM) filter, to achieve much improved filtering performance in terms of effectiveness in removing impulse noise while preserving signal details and robustness in combating noise density variations. The proposed NASM filter consists of two stages. A soft-switching noise-detection scheme is developed to classify each pixel to be uncorrupted pixel, isolated impulse noise, nonisolated impulse noise or image object's edge pixel. "No filtering" (or identity filter), standard median (SM) filter or our developed fuzzy weighted median (FWM) filter will then be employed according to the respective characteristic type identified. Experimental results show that our NASM filter impressively outperforms other techniques by achieving fairly close performance to that of ideal-switching median filter across a wide range of noise densities, ranging from 10% to 70% How-Lung Eng, Kai-Kuang Ma |
IEEE Trans. Image Process. | 2 |
| 2000 | Noise adaptive soft-switching median filter for image denoisingabstractWe observed that certain fundamental concerns commonly exist in some state-of-the-art switching-based median filters: (i) fixed thresholding for the pre-assumed noise density, (ii) the noise decision accuracy at high density impulse noise, and (iii) the filtering scheme adopted in response to pixel characteristic type identified. In this paper, we propose a novel noise adaptive soft-switching median (NASM) filter to effectively address the above-mentioned issues and achieve much improved filtering performance in terms of efficiency in removing impulse noise and robustness against noise density variations. Experimental results also reveal that the performance of our NASM filter is fairly close to that of ideal-switching median filter. How-Lung Eng, Kai-Kuang Ma |
ICASSP | 2 |
| 2000 | Fuzzy color histogram: an efficient color feature for image indexing and retrievalabstractThe conventional color histogram (CCH) considers neither the color similarity across different bins nor the color dissimilarity in the same bin, thus it is sensitive to noisy interference such as illumination changes. We propose a new concept of color histogram representation, called fuzzy color histogram (FCH), to address the above mentioned issue by considering the color similarity of each pixel's color associated to all the histogram bins through fuzzy-set membership function, individually. A novel and fast approach for computing the membership values based on fuzzy c-means algorithm is introduced. The proposed FCH is further exploited in the application of image indexing and retrieval. Experimental results clearly show that FCH yields better retrieval results than CCH. In addition, in contrast with quadratic histogram distance, our method shifts the computation load from on-line retrieval to off-line indexing. Such computing methodology is fairly desirable for image retrieval over large image databases. Ju Han, Kai-Kuang Ma |
ICASSP | 2 |
| 2000 | Unsupervised Image Object Segmentation over Compressed DomainabstractDirect processing of JPEG images based on DCT coefficients could avoid computationally intensive full decoding and large memory storage. In this paper, we exploit the inherent information extracted from DCT coefficients to achieve unsupervised segmentation of image objects. First, a maximum entropy fuzzy clustering (MEFC) algorithm is proposed to achieve a coarse segmentation based on DCT-DC coefficients. The DCT-AC coefficients are then utilized to refine the segmentation boundary by a maximum a posteriori (MAP) approach. The major challenge of the problem is to achieve satisfactory segmentation simply based on DCT coefficients, which are quantized and coarse information in essence. Experimental results show the promising potential of the proposed algorithm in overcoming these fundamental limitations. How-Lung Eng, Kai-Kuang Ma |
ICIP | 2 |
| 2000 | Novel video signal processor with VLIW-controlled SIMD architecture
Kai-Kuang Ma, Qingdong Yao |
VCIP | 2 |
| 2000 | Fundamental error analysis and geometric interpretation for block truncation coding techniques
Kai-Kuang Ma, Shan Zhu |
Signal Process. Image Commun. | 1 |
| 2000 | A new diamond search algorithm for fast block-matching motion estimationabstractBased on the study of motion vector distribution from several commonly used test image sequences, a new diamond search (DS) algorithm for fast block-matching motion estimation (BMME) is proposed in this paper. Simulation results demonstrate that the proposed DS algorithm greatly outperforms the well-known three-step search (TSS) algorithm. Compared with the new three-step search (NTSS) algorithm, the DS algorithm achieves close performance but requires less computation by up to 22% on average. Experimental results also show that the DS algorithm is better than the four-step search (4SS) and block-based gradient descent search (BBGDS), in terms of mean-square error performance and required number of search points. Shan Zhu, Kai-Kuang Ma |
IEEE Trans. Image Process. | 2 |
| 2000 | Correction to "a new diamond search algorithm for fast block-matching motion estimation"
Shan Zhu, Kai-Kuang Ma |
IEEE Trans. Image Process. | 2 |
| 1999 | Motion Trajectory Extraction Based on Macroblock Motion Vectors for Video IndexingabstractAs video sequences are composed of dynamic video objects (VOs) in nature, VO's motion is an effective feature used to provide an overall description about the content of video sequences. In this paper, we propose a novel video indexing technique that extracts VO's motion trajectory based on macroblock motion vectors (M-Vs) of MPEG encoded bitstreams, without performing full decoding. This is an attractive approach, as it promises large saving of computation and memory storage. The proposed motion trajectory extraction algorithm comprises two phases: (i) unsupervised VO segmentation and (ii) automatic tracking of each segmented VO. Experimental results reveal promising potential of the established algorithm in extracting VOs' motion trajectories from the limited information provided by MVs. How-Lung Eng, Kai-Kuang Ma |
ICIP (3) | 2 |
| 1999 | Optimal Algorithm for Progressive Polygon Approximation of Discrete Planar CurvesabstractThe problem of optimal polygon approximation of a discrete planar curve is addressed in this paper. Towards this end, an optimal algorithm using the progressive polygon approximation approach is proposed for a given acceptable approximation error and initial vertex. The proposed algorithm is optimal because it determines the minimal number of edges for a given approximation error tolerance. The proposed scheme can be extended to the approximation of digital contours wherein the contour points and the polygon vertices are restricted to the integer plane Z/sup 2/. Prabhudev I. Hosur, Kai-Kuang Ma |
ICIP (1) | 2 |
| 1999 | Bidirectional motion tracking for video indexingabstractMotion is recognized as one of the most essential video object (VO) features in indexing video contents. Previously, we directly exploited macroblock motion vectors (MVs) of MPEG bitstream without performing full decoding to extract motion trajectories of VOs. However, the performance of VO segmentation based on MVs is constrained by the limited information extracted from the MV field, since it is quantized and irregular in nature. Therefore, an effective VO tracking scheme would be essential to compensate this limitation. In this paper, we propose a bidirectional motion tracking methodology to (i) firstly validate current segmented VOs by exploring their correlations to the VOs obtained from previous and future frames and (ii) subsequently track each validated VO using discrete Kalman filter. The proposed scheme has been well tested using MPEG-7 test materials and achieves more reliable VO motion trajectory extraction. How-Lung Eng, Kai-Kuang Ma |
MMSP | 2 |
| 1999 | A novel scheme for progressive polygon approximation of shape contoursabstractThis paper presents an efficient algorithm for polygon approximation of shape contours. The proposed algorithm approximates a shape contour by a polygon with minimal number of vertices for given allowable approximation error and initial vertex. Furthermore, it is designed to provide a low computational complexity and simple implementation. The efficacy of the proposed algorithm is demonstrated through experimental results. Prabhudev I. Hosur, Kai-Kuang Ma |
MMSP | 2 |
| 1999 | Colour Image Indexing Using SOM for Region-of-Interest Retrieval
Tao Chen 0044, Lihui Chen 0001, Kai-Kuang Ma |
Pattern Anal. Appl. | 3 |
| 1999 | Tri-state median filter for image denoisingabstractIn this work, a novel nonlinear filter, called tri-state median (TSM) filter, is proposed for preserving image details while effectively suppressing impulse noise. We incorporate the standard median (SM) filter and the center weighted median (CWM) filter into a noise detection framework to determine whether a pixel is corrupted, before applying filtering unconditionally. Extensive simulation results demonstrate that the proposed filter consistently outperforms other median filters by balancing the tradeoff between noise reduction and detail preservation. Tao Chen 0044, Kai-Kuang Ma, Lihui Chen 0001 |
IEEE Trans. Image Process. | 2 |
| 1998 | ROI-Oriented Image Query and Indexing for Content-based RetrievalabstractA new scheme for image query and indexing based on the concepts of region-of-interest (ROI) is proposed for content-based retrieval. Users are allowed to impose an ROI directly over the sample image, and the subsequent query is then focused on the content of selected ROI in order to find those images containing similar regions or objects from the database. To accomplish this objective, a neural network model, called a self-organizing map (SOM), is exploited to adaptively separate each image into several homogeneous regions without using any prior knowledge followed by a distance measure for ROI matching. Experimental results demonstrate that the proposed approach is fairly flexible, efficient and effective with great potential for further development. Tao Chen 0044, Lihui Chen 0001, Kai-Kuang Ma |
ICIP (2) | 3 |
| 1998 | Discrete wavelet frame representations of color texture features for image queryabstractWe propose a wavelet-based multi-channel scheme to extract human-perception relevant color texture features for image indexing and querying. For each spectral band, a two-dimensional discrete wavelet frame (DWF) decomposition is applied first, followed by an enveloping operation performed on the resulting wavelet coefficients. The unichrome features computed from the enveloped coefficients of the individual band as well as the opponent features that provide the spatial correlation between different spectral bands are jointly exploited for accurate image classification. The experimental results are promisingly, showing that the proposed approach is suitable for browsing color texture images. Tao Chen 0044, Kai-Kuang Ma, Lihui Chen 0001 |
MMSP | 2 |
| 1997 | A DCT Embedded Subband AMBTC Image CoderabstractFirst, we propose a DCT-based subband absolute moment block truncation coding (SAMBTC) method and compare its coding performance with that of QMF-based SAMBTC. Second, the characteristics of the human visual system based on the Weber's law model is incorporated into these schemes, independently. The objective is to further reduce the bit rate without incurring more noticeable degradation. By comparison with the JPEG standard, our image coders suffer much less blocking artifacts at low bit rates and obtain superior subjective image quality. Kai-Kuang Ma |
ICIP (2) | 1 |
| 1997 | Put absolute moment block truncation coding in perspectiveabstractThe purpose of this letter is to point out that the Udpikar and Raina (1987) modified block truncation coding (BTC) has been misleading researchers as a different BTC scheme in multiple publications for years. In fact, their algorithm is exactly identical to the second version of the Lema and Mitchell (1984) absolute moment block truncation coding (AMBTC). Mathematical proof is provided. Kai-Kuang Ma |
IEEE Trans. Commun. | 1 |
| 1996 | Modified absolute moment block truncation codingabstractBy merging absolute moment block truncation coding (AMBTC) and block truncation coding (BTC), a new three-moment AMBTC (TAMBTC) is presented. A generalized error analysis is presented to explain why TAMBTC is outperformed by AMBTC. More insights on the aspects of computational issues and preserved moments are also discussed. Kai-Kuang Ma |
ICIP (1) | 1 |
| 1995 | New properties of AMBTC [absolute moment block truncation coding]abstractNew properties of absolute moment block truncation coding (AMBTC) are presented with proof. The main purposes of this work are (1) to provide some fundamental insights into the AMBTC algorithm and (2) to show that AMBTC is a robust choice among the 1-bit quantizers considering both computational complexity and quantization error.> Kai-Kuang Ma, Sarah A. Rajala |
IEEE Signal Process. Lett. | 1 |
| 1994 | Generalized Optimum Dynamic Bit Allocation Scheme for Source CompressionabstractDynamic bit allocation directly impacts the overall coding performance in block transform (e.g., DCT, Hadamard, etc.) coding, subband coding, and wavelet coding. In very low bit rate coding (such as the evolving MPEG-4 standard), additional factors impacting the bit allocation process need to be identified and properly exploited to achieve optimal coding performance. In the paper, a generalized optimum dynamic bit allocation algorithm is presented. The algorithm is based on the Shannon rate-distortion bound and a weighted distortion criterion using a Lagrange multiplier optimization technique. The results look promising for applications in real-time multimedia communications.> Kai-Kuang Ma, Sarah A. Rajala |
ICIP (2) | 1 |
| 1991 | Subband coding of digital images using absolute moment block truncationabstractA combination of subband coding and absolute moment block truncation coding (AMBTC) is presented to effectively eliminate the blocking effect (or grid noise) that the AMBTC severely suffers. To dynamically allocate bits for each subband and efficiently achieve quality images for a given compression ratio, there are two fundamental questions; (1) how to choose subbands for applying AMBTC, and (2) how to choose the appropriate window sizes for these subbands. The concept of intra-subband and inter-subband bit allocation has been proposed to address the former issue, and a variable window size approach is suggested for the latter one. Each solution shows promising results.> Kai-Kuang Ma, Sarah A. Rajala |
ICASSP | 1 |