Susanto Rahardja

dblp:98/3034 · DBLP profile ↗
← Back
207ranked-venue papers
2as first author
51since 2021 · last 2026
0000-0003-0831-6934ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 146 · 23 since 2021Artificial intelligence and machine learning · 21 · 10 since 2021Systems, architecture and hardware · 17 · 1 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 9 since 2021Computer networks · 8 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 4 · 4 since 2021Security and privacy · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 Removing Box-Free Watermarks for Image-to-Image Models via Query-Based Reverse Engineering
abstract
The intellectual property of deep generative networks (GNets) can be protected using a cascaded hiding network (HNet) which embeds watermarks (or marks) into GNet outputs, known as box-free watermarking. Although both GNet and HNet are encapsulated in a black box (called operation network, or ONet), with only the generated and marked outputs from HNet being released to end users and deemed secure, in this paper, we reveal an overlooked vulnerability in such systems. Specifically, we show that the hidden GNet outputs can still be reliably estimated via query-based reverse engineering, leaking the generated and unmarked images, despite the attacker's limited knowledge of the system. Our first attempt is to reverse-engineer an inverse model for HNet under the stringent black-box condition, for which we propose to exploit the query process with specially curated input images. While effective, this method yields unsatisfactory image quality. To improve this, we subsequently propose an alternative method leveraging the equivalent additive property of box-free model watermarking and reverse-engineering a forward surrogate model of HNet, with better image quality preservation. Extensive experimental results on image processing and image generation tasks demonstrate that both attacks achieve impressive watermark removal success rates (100%) while also maintaining excellent image quality (reaching the highest PSNR of 34.69 dB), substantially outperforming existing attacks, highlighting the urgent need for robust defensive strategies to mitigate the identified vulnerability in box-free model watermarking.
Haonan An 0001, Guang Hua 0001, Hangcheng Cao, Zhengru Fang, Guowen Xu, Susanto Rahardja, Yuguang Fang
AAAI6
2026 Towards reliable recognition for plant diseases and weeds by learning soft probability population
Maowen Zhou, Mingle Xu, Erma Rahayu Mohd Faizal Abdullah, Aznul Qalid Md Sabri, Susanto Rahardja
Eng. Appl. Artif. Intell.6
2026 Subsystem-Aware Stackelberg Game for Optimal Anti-Jamming Power Control in Periodically Switched Cyber-Physical Systems
abstract
This paper presents a subsystem-aware Stackelberg game framework for optimal power control in periodically switched cyber-physical systems (PSCPS) operating over wireless networks under jamming attacks. To account for subsystem switching dynamics, a switching Kalman filter is employed for mode-dependent state estimation, with corresponding error covariance matrices derived for each mode. The impact of switching is quantified using a subsystem-importance-based weighting factor, obtained by analyzing trace variations in the covariance matrices. This factor is integrated into the defender’s utility function to enable adaptive power control. First, a subsystem-aware Stackelberg game model without power constraints is formulated, and the equilibrium strategies of the defender (leader) and the attacker (follower) are obtained via the convex optimization. The framework is then extended to the power-constrained scenario, where feasibility issues arise. By analyzing the boundary and extreme points of the solution space, closed-form equilibrium strategies are derived. Simulation results show that the proposed approach enhances the resilience of PSCPS against jamming attacks in wireless communication environments.
Peilin Jia, Jiyun Tian, Jie Lian 0001, Susanto Rahardja
IEEE Internet Things J.5
2026 DiffPixelFormer: Differential Pixel-Aware Transformer for RGB-D Indoor Scene Segmentation
abstract
Indoor semantic segmentation is fundamental to computer vision and robotics, supporting applications such as autonomous navigation, augmented reality, and smart environments. Although RGB-D fusion leverages complementary appearance and geometric cues, existing methods often depend on computationally intensive cross-attention mechanisms and insufficiently model intra- and inter-modal feature relationships, resulting in imprecise feature alignment and limited discriminative representation. To address these challenges, we propose DiffPixelFormer, a differential pixel-aware Transformer for RGB-D indoor scene segmentation that simultaneously enhances intra-modal representations and models inter-modal interactions. At its core, the Intra-Inter Modal Interaction Block (IIMIB) captures intra-modal long-range dependencies via self-attention and models inter-modal interactions with the Differential–Shared Inter-Modal (DSIM) module to disentangle modality-specific and shared cues, enabling fine-grained, pixel-level cross-modal alignment. Furthermore, a dynamic fusion strategy balances modality contributions and fully exploits RGB-D information according to scene characteristics. Extensive experiments on the SUN RGB-D and NYUDv2 benchmarks demonstrate that DiffPixelFormer-L achieves mIoU scores of 54.28% and 59.95%, outperforming DFormer-L by 1.78% and 2.75%, respectively. Moreover, its effectiveness is further validated on the large-scale ScanNetv2 dataset, indicating strong generalization capability. Code is available at https://github.com/gongyan1/DiffPixelFormer.
Jianli Lu, Yongsheng Gao 0002, Jie Zhao 0003, Susanto Rahardja
IEEE Trans. Circuits Syst. Video Technol.6
2026 Decoder Gradient Shields: A Family of Provable and High-Fidelity Methods Against Gradient-Based Box-Free Watermark Removal
abstract
Box-free model watermarking has gained significant attention in deep neural network (DNN) intellectual property protection due to its model-agnostic nature and its ability to flexibly manage high-entropy image outputs from generative models. Typically operating in a black-box manner, it employs an encoder-decoder framework for watermark embedding and extraction. While existing research has focused primarily on the encoders for the robustness to resist various attacks, the decoders have been largely overlooked, leading to attacks against the watermark. In this paper, we identify one such attack against the decoder, where query responses are utilized to obtain backpropagated gradients to train a watermark remover. To address this issue, we propose Decoder Gradient Shields (DGSs), a family of defense mechanisms, including DGS at the output (DGS-O), at the input (DGS-I), and in the layers (DGS-L) of the decoder, with a closed-form solution for DGS-O and provable performance for all DGS. Leveraging the joint design of reorienting and rescaling of the gradients from watermark channel gradient leaking queries, the proposed DGSs effectively prevent the watermark remover from achieving training convergence to the desired low-loss value, while preserving image quality of the decoder output. We demonstrate the effectiveness of our proposed DGSs in diverse application scenarios. Our experimental results on deraining and image generation tasks with the state-of-the-art box-free watermarking show that our DGSs achieve a defense success rate of 100% under all settings.
Haonan An 0001, Guang Hua 0001, Hangcheng Cao, Yihang Tao, Guowen Xu, Susanto Rahardja, Yuguang Fang
IEEE Trans. Dependable Secur. Comput.7
2026 MMM: A Unified Weakly-Supervised Anomaly Detection Framework for Multi-Distributional Data
abstract
Weakly-Supervised Anomaly Detection (WSAD) has garnered increasing research interest in recent years, as it enables superior detection performance while demanding only a small fraction of labeled data. However, existing WSAD methods face two major limitations. From the data aspect, they struggle to detect anomalies between normal clusters or collective anomalies due to overlooking the multi-distribution and complex manifolds of real-world data. From the label aspect, they fall short of detecting unknown anomalies because of the label-insufficiency and anomaly contamination. To address these issues, we propose MMM, a unified WSAD framework for multi-distributional data. The framework consists of three components: a Multi-distribution data modeler captures latent representations of complex data distributions, followed by a Multiform feature extractor that extracts multiple underlying features from the modeler, highlighting the characteristics of potential anomalies. Finally, a Multi-strategy anomaly score estimator converts these features into anomaly scores, with the aid of a novel training approach with three strategies that maximize the utility of both data and labels. Experimental results showed that MMM achieved superior performance and robustness compared to state-of-the-art WSAD methods, while providing interpretable results that facilitate practical anomaly analysis.
Xu Tan 0004, Junqi Chen 0001, Jiawei Yang 0001, Jie Chen 0022, Susanto Rahardja
IEEE Trans. Knowl. Data Eng.5
2026 Test-Time Learning for Outlier Detection
abstract
In this work, the concept of test-time learning is presented, wherein Machine-Learning (ML) models are constructed by involving unlabeled test samples. Based on this concept, we propose an unsupervised method called Local Augment (LA) designed to improve the performance of trained outlier detectors at the prediction stage without altering the trained models or accessing the training data. LA operates under the only assumption that the model should produce similar outputs for similar inputs, implying that the prediction of a given sample can be enhanced by the predictions for its similar samples. Specifically, LA boosts outlier detection performance during prediction by fusing the outlier score of a given sample with the scores of synthetically neighboring samples generated by adding random perturbations to the given sample. This simple method demonstrates an average improvement of +0.04 Area Under the Receiver Operating Characteristic curve (AUROC) across 22 real-world datasets for all 11 tested detectors. Notably, this represents the pioneering work of enhancing ML models during the prediction stage without the need to modify the trained models or access the training dataset. This work opens up new possibilities for addressing existing bottleneck problems in various ML tasks beyond outlier detection in diverse domains.
Jiawei Yang 0001, Jingdong Chen, Susanto Rahardja
IEEE Trans. Knowl. Data Eng.3
2026 Bootstrap Deep Spectral Clustering With Optimal Transport
abstract
Spectral clustering is a leading clustering method. Two of its major shortcomings are the disjoint optimization process and the limited representation capacity. To address these issues, we propose a deep spectral clustering model (named BootSC), which jointly learns all stages of spectral clustering—affinity matrix construction, spectral embedding, and$k$-means clustering—using a single network in an end-to-end manner. BootSC leverages effective and efficient optimal-transport-derived supervision to bootstrap the affinity matrix and the cluster assignment matrix. Moreover, a semantically-consistent orthogonal re-parameterization technique is introduced to orthogonalize spectral embeddings, significantly enhancing the discrimination capability. Experimental results indicate that BootSC achieves state-of-the-art clustering performance. For example, it accomplishes a notable 16% NMI improvement over the runner-up method on the challenging ImageNet-Dogs dataset. Our code is available athttps://github.com/spdj2271/BootSC.
Wengang Guo, Wei Ye 0001, Chunchun Chen, Xin Sun 0003, Christian Böhm 0001, Claudia Plant, Susanto Rahardja
IEEE Trans. Multim.7
2025 Decoder Gradient Shield: Provable and High-Fidelity Prevention of Gradient-Based Box-Free Watermark Removal
abstract
The intellectual property of deep image-to-image models can be protected by the so-called box-free watermarking. It uses an encoder and a decoder, respectively, to embed into and extract from the model’s output images invisible copyright marks. Prior works have improved watermark robustness, focusing on the design of better watermark encoders. In this paper, we reveal an overlooked vulnerability of the unprotected watermark decoder which is jointly trained with the encoder and can be exploited to train a watermark removal network. To defend against such an attack, we propose the decoder gradient shield (DGS) as a protection layer in the decoder API to prevent gradient-based watermark removal with a closed-form solution. The fundamental idea is inspired by the classical adversarial attack, but is utilized for the first time as a defensive mechanism in the box-free model watermarking. We then demonstrate that DGS can reorient and rescale the gradient directions of watermarked queries and stop the watermark remover’s training loss from converging to the level without DGS, while retaining decoder output image quality. Experimental results verify the effectiveness of the proposed method. Code of paper is available at https://github.com/haonanAN309/CVPR-2025-Official-Implementation-Decoder-Gradient-Shield.
Haonan An 0001, Guang Hua 0001, Zhengru Fang, Guowen Xu, Susanto Rahardja, Yuguang Fang
CVPR5
2025 Smoothing Outlier Scores is All You Need to Improve Outlier Detectors (Extended Abstract)
abstract
Existing outlier detectors calculate outlier scores for data objects independently, ignoring the consistency between score similarity and object similarity. As a result, these detectors may produce inconsistent scores for similar objects, leading the scores of some normal objects to exceed some of outlier objects, increasing the possibility of misclassification. To address this issue, we first assume that similar objects should have similar scores. Then, based on this assumption, we propose neighborhood averaging, an outlier score post-processing technique to improve any single outlier detector, which is the first of its kind.
Jiawei Yang 0001, Susanto Rahardja, Pasi Fränti
ICDE2
2025 MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled Embedding
abstract
Audio-driven emotional 3D facial animation aims to generate synchronized lip movements and vivid facial expressions. However, most existing approaches focus on static and predefined emotion labels, limiting their diversity and naturalness. To address these challenges, we propose MEDTalk, a novel framework for fine-grained and dynamic emotional talking head generation. Our approach first disentangles content and emotion embedding spaces from motion sequences using a carefully designed cross-reconstruction process, enabling independent control over lip movements and facial expressions. Beyond conventional audio-driven lip synchronization, we integrate audio and speech text, predicting frame-wise intensity variations and dynamically adjusting static emotion features to generate realistic emotional expressions. Furthermore, to enhance control and personalization, we incorporate multimodal inputs-including text descriptions and reference expression images-to guide the generation of user-specified facial expressions. With MetaHuman as the priority, our generated results can be conveniently integrated into the industrial production pipeline. The code is available at: https://github.com/SJTU-Lucy/MEDTalk.
Chang Liu 0021, Susanto Rahardja, Xiaokang Yang 0001
ACM Multimedia4
2025 A comb concatenation diffusion model for hyperspectral image super-resolution
Yinghao Xu 0003, Hao Wang 0192, Xin Sun 0003, Qianlong Xie, Peng Ren 0001, Fei Zhou 0007, Susanto Rahardja
Eng. Appl. Artif. Intell.8
2025 MSS-PAE: Saving Autoencoder-based Outlier Detection from Unexpected Reconstruction
Xu Tan 0004, Jiawei Yang 0001, Junqi Chen 0001, Sylwan Rahardja, Susanto Rahardja
Pattern Recognit.5
2025 Multi-granularity acoustic information fusion for sound event detection
Han Yin, Jisheng Bai, Mou Wang, Susanto Rahardja, Dongyuan Shi, Woon-Seng Gan
Signal Process.5
2025 Fast tone mapping operator for high dynamic range image using prior information
Xueyu Han, Xin Sun 0003, Susanto Rahardja
Signal Process. Image Commun.3
2025 Real-Valued Discrete Fractional Hadamard Transform: Fast Algorithms and Implementations
abstract
This paper introduces a new real-valued discrete fractional Hadamard transform (RFHT) designed to address the issues of high computational complexity and large storage demands found in the traditional discrete fractional Hadamard transform (FHT) and its modifications. Additionally, a fast algorithm for the RFHT is developed, and the corresponding computational complexity analysis demonstrates that for sizes ranging from$N = 2$to 1024, the proposed fast algorithm can reduce the number of multiplications and additions by up to 50.0% and 83.3%, respectively, compared to state-of-the-art fast algorithms. The comparison results show that the RFHT also has lower execution time and power consumption. Furthermore, due to its real-valued property, the RFHT has been applied in an image watermarking system and implemented on real iOS devices, demonstrating enhanced information security. Lower execution and power consumption, reduced storage and transmission requirements, and superior information protection make the RFHT a superior candidate compared to WHT, FHT, and its modifications.
Zi-Chen Fan, Di Li 0006, Susanto Rahardja
IEEE Trans. Circuits Syst. I Regul. Pap.3
2025 MambaHSISR: Mamba Hyperspectral Image Super-Resolution
abstract
One of the main challenges facing hyperspectral image super-resolution is the complex high dimensional data processing. Mamba leverages its ability to model long-range dependencies of linear complexity to capture the global spatial and spectral information of high-dimensional data while maintaining linear complexity. However, its visual state space equation mainly focuses on the band dimension mapping of the image, while ignoring the modeling of the spatial dimension. To overcome this limitation, we develop a Mamba hyperspectral image super-resolution framework, which comprises three essential components. The first component, i.e., spatial Mamba sub-network, models the spatial dimensions of hyperspectral data. It captures long-range dependencies in the pixel space, thereby integrating global spatial information into the framework. The second component, i.e., spectral Mamba sub-network, serves to capture long-range spectral dependencies. The third component, i.e., reconstruction, generates hyperspectral images with rich spatial and spectral details through pixel interpolation. Our Mamba framework fully develops the potential of the Mamba model in hyperspectral image super-resolution, significantly enhancing the restoration quality and accuracy of hyperspectral images. Extensive experiments on the Houston and QUST-1 datasets show that our framework outperforms state-of-the-art methods in both quantitative metrics and visual quality across diverse scenarios. We release our source code at https://gitee.com/xu_yinghao/MambaHSISR for public evaluations.
Yinghao Xu 0003, Hao Wang 0192, Fei Zhou 0007, Chunbo Luo, Xin Sun 0003, Susanto Rahardja, Peng Ren 0001
IEEE Trans. Geosci. Remote. Sens.6
2025 Hyperspectral Texture Metrology Based on Distance Measures in an Information-Theoretic Framework
abstract
The present work sought to instil metrology in existing hyperspectral texture feature extraction methods. Specifically, we propose distance-based expressions of graylevel cooccurrence matrix (GLCM), local binary pattern (LBP), and Gabor filtering directly computable for hyperspectral images without any pre- or post-processing. At the core of our proposition is Radical of Extended Mean Information for Discrimination (REID), a novel spectral distance with information-theoretic roots. Respecting the physics of spectrum as continuous function of wavelengths, REID is mathematically decomposable into spectral direction and spectral magnitude distances. The resulted feature calculations are fullband (utilizing all wavelengths), yet lightweight and fully interpretable. A similarity measure based on information theory is also justified. Their efficiency is demonstrated in the context of texture classification, content-based image retrieval, and cancer detection in which they consistently outperform existing computations based on dimensionally reduced space using PCA, ICA, and NMF. The propositions could be potentially integrated into machine/deep learning systems towards explainable AI (XAI).
Rui Jian Chu, Jie Chen 0022, Susanto Rahardja
IEEE Trans. Image Process.3
2025 Rethinking Affine Transform for Efficient Image Enhancement: A Color Space Perspective
abstract
In recent years, we have observed significant advancements in learning-based techniques for image-enhancement tasks. However, most of the existing methods are either purely based on image-to-image convolutional neural networks, which cannot handle high-resolution images in real-time, or resort to 3D Lookup Tables, which fall short of local tone adjustments. In this paper, we rethink affine transform through a color space perspective, and then propose AttnBL (Attentional Bilateral Grid Learning), a novel hybrid image enhancement algorithm to process ultra-high-definition images in real-time. Our algorithm consists of two paths, the low-resolution chroma prediction path that aims to learn the chroma coefficients and the full-resolution luma adaptation path that aims to preserve brightness details. Specifically, we propose a carefully designed hierarchical transformer to capture the global information in an efficient way and introduce a feature extraction module to adaptively learn a luma guidance for bilateral upsampling. Our algorithm can process a 4K-resolution image in 20 milliseconds. This efficiency provides a practical solution for high resolution real-time preview. Without bells and whistles, our model outperforms previous state-of-the-art methods on two well-known datasets in image enhancement tasks both quantitatively and qualitatively. Our analysis also provides some interesting findings that may enlighten further studies.
Di Li 0006, Susanto Rahardja
IEEE Trans. Multim.2
2024 A Novel Discrete Fractional Complex Hadamard Transform for Medical Image Encryption
abstract
This paper introduces a new discrete fractional complex Hadamard transform (FCHT) and its generalized form, the multiple-parameter FCHT (MFCHT). The MFCHT is applied to the medical image encryption. Both subjective observations and objective evaluations are conducted to validate the effectiveness of the proposed algorithm, and the simulation results demonstrate that the proposed MFCHT outperforms previous transform-based algorithms, including the discrete fractional Fourier transform and the discrete fractional Hadamard transform, in terms of robust image preservation against blind attacks in the medical image encryption.
Zi-Chen Fan, Di Li 0006, Susanto Rahardja
ICASSP3
2024 Ensemble of Deep Variational Mixture Models for Unsupervised Clustering
abstract
Deep variational mixture models (DVMMs) have demonstrated promising performance in unsupervised clustering for complicated high-dimensional data such as images. However, their prediction accuracy is often unstable and significantly influenced by randomness, particularly during the initialization of parameters. To reduce this uncertainty, we propose an ensemble approach that combines the predictions of multiple base models. Specifically, we introduce two individual ensemble strategies: voting and merging. In the voting strategy, the final label is determined by selecting the predicted class label with the most votes and lowest Shannon entropy. In the merging strategy, the class probability vectors (scaled by the temperature parameter) from different models are combined to predict the final class label. Experimental results on two image datasets demonstrate that these proposed methods yield reliable and superior clustering performance.
Xu Tan 0004, Junqi Chen 0001, Jiawei Yang 0001, Sylwan Rahardja, Mou Wang, Susanto Rahardja
ICIP6
2024 FlexAE: A Self-Conditioned Detector To Prevent Model Overfitting For Unsupervised Video Anomaly Detection
abstract
Unsupervised Video Anomaly Detection (VAD) has garnered significant attention for its ability to exploit unlabeled videos. However, VAD faces two primary challenges arising from the absence of labels: (i) Striking a balance between overfitting and underfitting, and (ii) Optimal parameter tuning. To tackle these challenges, we propose a novel detector named Flexible AutoEncoder (FlexAE). A fitting-parameter is introduced to regulate the model’s fitting capacity, and a novel Negative Learning (NL) mechanism is integrated to mitigate the influence of anomalies during training. For self-conditioning, a novel algorithm is devised to autonomously update the fitting-parameter and the threshold used in NL based on the reconstruction error. Comprehensive experiments on two benchmark datasets, UCF-Crime and ShanghaiTech, demonstrate that our proposed FlexAE outperforms state-of-the-art methods without the need for manual hyperparameter tuning.
Junqi Chen 0001, Xu Tan 0004, Jiawei Yang 0001, Sylwan Rahardja, Susanto Rahardja
ICIP5
2024 Low-Rank Matrix and Tensor Decomposition Using Randomized Two-Sided Subspace Iteration With Application to Video Reconstruction
abstract
The low-rank approximation of big data matrices and tensors plays a pivotal role in many modern applications. Recently, the randomized subspace iteration has shown to be a powerful tool in approximating large matrices. In this paper we present a rank-revealing, two-sided variant of the randomized subspace iteration. Novelty of our work lies in the utilization of the unpivoted QR factorization, rather than the singular value decomposition (SVD), for factorizing the compressed matrix. We provide bounds on the rank-revealingness of our algorithm as well as bounds on the error of the low-rank approximations, in both 2- and Frobenius norm. In addition, we employ the proposed algorithm to efficiently compute the low rank tensor decomposition using the truncated higher-order SVD. We conduct tests on (i) two classes of matrices, and (ii) synthetic data tensor and real dataset to demonstrate the efficacy of the proposed algorithms.
Maboud F. Kaloorazi, Salman Ahmadi-Asl, Susanto Rahardja
ICIP3
2024 Audiolog: LLMs-Powered Long Audio Logging with Hybrid Token-Semantic Contrastive Learning
abstract
Previous studies in automated audio captioning have faced difficulties in accurately capturing the complete temporal details of acoustic scenes and events within long audio sequences. This paper presents AudioLog, a large language models (LLMs)-powered audio logging system with hybrid token-semantic contrastive learning. Specifically, we propose to fine-tune the pre-trained hierarchical token-semantic audio Transformer by incorporating contrastive learning between hybrid acoustic representations. We then leverage LLMs to generate audio logs that summarize textual descriptions of the acoustic environment. Finally, we evaluate the AudioLog system on two datasets with both scene and event annotations. Experiments show that the proposed system achieves exceptional performance in acoustic scene classification and sound event detection, surpassing existing methods in the field. Further analysis of the prompts to LLMs demonstrates that AudioLog can effectively summarize long audio sequences1. To the best of our knowledge, this approach is the first attempt to leverage LLMs for summarizing long audio sequences.
Jisheng Bai, Han Yin, Mou Wang, Dongyuan Shi, Woon-Seng Gan, Susanto Rahardja
ICME7
2024 Unsupervised Image Enhancement via Contrastive Learning
abstract
Recent years have witnessed significant achievements for image enhancement tasks. However, many advanced algorithms are trained in a supervised manner and thus rely on a huge collection of paired data, for which the collection is itself a challenge especially for real-world scenarios. We address this issue by proposing a novel GAN framework designed for unsupervised training. To be specific, our approach introduces a contrastive loss to ensure that the content remains consistent across multiple scales in both input and output representations. In addition, we propose a multi-scale discriminator to strengthen the adversarial learning. Extensive experiments conducted in this paper showed that our algorithm achieved state-of-the-art performance on MIT-Adobe-FiveK dataset both quantitively and qualitatively.
Di Li 0006, Susanto Rahardja
ISCAS2
2024 i6mA-CNN: A Web-based System to Identify DNA N6-Methyladenine Sites in Mouse Genomes
abstract
N6-methyladenine (6mA) is one of the most common epigenetic modifications of DNA sequences found in both eukaryotes and prokaryotes. In prokaryotes, 6mA is closely associated with various biochemical processes such as DNA replication, repair, transcription, and cellular defense. In eukaryotes, the biological roles and behaviors of this methylation type have not been fully understood. Therefore, gaining more knowledge about 6mA sites is important and contributes to uncovering the characteristics and unexplored biological functions of 6mA. In our study, we propose an effective computational method called i6mA-CNN using convolutional neural networks with a fusion of multiple receptive fields. The 6mA sequences of Mus musculus (mice) were retrieved from the MethSMRT database and then refined to create a benchmark dataset. To fairly evaluate the performance of the model, we conducted multiple experiments and compared i6mA-CNN with other methods on the same independent test set. The results indicated that i6mACNN achieved better performance, with a value of 0.98 for both the area under the receiver operating characteristic curve and the area under the precision-recall curve.
Thanh-Hoang Nguyen-Vo, Susanto Rahardja, Binh P. Nguyen
ISCAS2
2024 Identifying Nephrotoxicity of Small Molecules Using Machine Learning
abstract
Nephrotoxicity is a severe condition characterized by kidney damage resulting from exposure to harmful substances such as drugs, diagnostic agents, chemicals, or environmental toxins. The potential for nephrotoxicity in drug molecules remains significant, often leading to severe consequences for patients. Despite existing computational methods for identifying nephrotoxic molecules, these approaches fail to provide stable and reliable performance due to biased modeling (e.g., small sample sizes, imbalanced classes, and data leakage). In this study, we offer a refined dataset for nephrotoxicity prediction tasks. Our dataset was collected from existing studies, rebalanced, and carefully curated to improve the quality of data for Quantitative Structure-Activity Relationship modeling. Additionally, we implemented a series of 32 prediction models using eight machine learning algorithms in combination with three types of molecular representations. The implementation of these machine learning models serves as a preliminary survey of how conventional methods perform on the refined dataset. Our findings indicated that all implemented models achieved satisfactory performance. Our dataset could serve as a valuable resource for developing more advanced prediction methods in the future.
Thanh-Hoang Nguyen-Vo, Linh Bui, Trang T. T. Do 0001, Susanto Rahardja, Binh P. Nguyen
TENCON4
2024 Image tone mapping based on clustering and human visual system models
Xueyu Han, Ishtiaq Rasool Khan, Susanto Rahardja
Signal Process. Image Commun.3
2024 Contrastive learning for deep tone mapping operator
Di Li 0006, Mou Wang, Susanto Rahardja
Signal Process. Image Commun.3
2024 Joint Selective State Space Model and Detrending for Robust Time Series Anomaly Detection
abstract
Deep learning-based sequence models are extensively employed in Time Series Anomaly Detection (TSAD) tasks due to their effective sequential modeling capabilities. However, the ability of TSAD is limited by two key challenges: (i) the ability to model long-range dependency and (ii) the generalization issue in the presence of non-stationary data. To tackle these challenges, an anomaly detector that leverages the selective state space model known for its proficiency in capturing long-term dependencies across various domains is proposed. Additionally, a multi-stage detrending mechanism is introduced to mitigate the prominent trend component in non-stationary data to address the generalization issue. Extensive experiments conducted on real-world public datasets demonstrate that the proposed methods surpass all 12 compared baseline methods.
Junqi Chen 0001, Xu Tan 0004, Sylwan Rahardja, Jiawei Yang 0001, Susanto Rahardja
IEEE Signal Process. Lett.5
2024 Transformer-Based End-to-End Speech Translation With Rotary Position Embedding
abstract
Recently, many Transformer-based models have been applied to end-to-end speech translation because of their capability to model global dependencies. Position embedding is crucial in Transformer models as it facilitates the modeling of dependencies between elements at various positions within the input sequence. Most position embedding methods employed in speech translation such as the absolute and relative position embedding, often encounter challenges in leveraging relative positional information or adding computational burden to the model. In this letter, we introduce a novel approach by incorporating rotary position embedding into Transformer-based speech translation (RoPE-ST). RoPE-ST first adds absolute position information by multiplying the input vector with rotation matrices, and then implements relative position embedding through the dot-product of the self-attention mechanism. The main advantage of the proposed method over the original method is that rotary position embedding combines the benefits of absolute and relative position embedding, which is suited for position embedding in speech translation tasks. We conduct experiments on a multilingual speech translation corpus MuST-C. Results show that RoPE-ST achieves an average improvement of 2.91 BLEU over the method without rotary position embedding in eight translation directions.
Xueqing Li 0003, Shengqiang Li, Xiao-Lei Zhang 0001, Susanto Rahardja
IEEE Signal Process. Lett.4
2024 Multi-Resolution Convolutional Residual Neural Networks for Monaural Speech Dereverberation
abstract
It is known that the reverberant speech in different acoustic environments varies according to reverberation time. However, most deep learning based speech dereverberation methods rely on a single deep model to learn the context information. It may make the deep model biased to only part of the reverberant time durations. In this paper, we propose a multi-resolution framework to address this issue. The framework integrates the dereverberant ability of multiple deep subnetworks with different time resolutions into a unified model by transferring the dereverberant information from high-resolution subnetworks to low-resolution subnetworks. By doing so, the unified model can perform well in both long and short reverberant time. We further propose two implementations of the framework based on advanced convolutional residual neural networks. The first implementation, named multi-resolution UNet, uses our new implementation of UNet based on convolutional blocks as the dereverberation subnetwork. The second implementation, named multi-resolution stacked convolutional blocks, uses our new stacked convolutional blocks as the subnetwork. Experimental results in both simulated and real-world environments show that the proposed algorithms outperform the state-of-the-art dereverberation methods in terms of both the evaluation metrics for speech dereverberation and word error rate (WER) for speech recognition.
Lei Zhao 0031, Shengqiang Li, Xiao-Lei Zhang 0001, Susanto Rahardja
IEEE ACM Trans. Audio Speech Lang. Process.6
2024 Efficient Computation for Discrete Fractional Hadamard Transform
abstract
This paper introduces a new fast algorithm for the discrete fractional Hadamard transform (FHT). The proposed algorithm demonstrates superior computational efficiency. For data lengths ranging from$2 \leq N \leq 1024$, our algorithm achieves a reduction in the number of multiplications by up to 96.53%, 81.82%, 33.33%, and 90% compared to four existing fast algorithms for the FHT. Additionally, we compare the execution times with those of existing fast algorithms, and the results show that the proposed algorithm has better performance. The reduced computational complexity makes the proposed algorithm a potential candidate for calculating the FHT.
Zi-Chen Fan, Di Li 0006, Susanto Rahardja
IEEE Trans. Circuits Syst. I Regul. Pap.3
2024 FOOR: Be Careful for Outlier-Score Outliers When Using Unsupervised Outlier Ensembles
abstract
Outlier detection is a very important tool in analyzing patterns and detecting unexpected events in social systems. However, the process of outlier detection could be fraught with uncertainty, with difficulties in determining the veracity of an object’s outlier score. We propose a framework for outlier-score outlier removal (FOOR). FOOR is a selection method, which aims to remove inaccurate outlier scores prior to data processing by ensemble techniques, to improve the accuracy of all ensembles. FOOR has rigorously tested with 30 real-world datasets and seven state-of-the-art ensembles over 25 different base detectors. Simulated experiments showed that FOOR significantly improves the existing techniques, with an average (AVG) of +0.05 AUC (from 0.81 to 0.86 AUC). Thus, we recommend FOOR as the new standard for outlier-score preprocessing before ensembles.
Jiawei Yang 0001, Sylwan Rahardja, Susanto Rahardja
IEEE Trans. Comput. Soc. Syst.3
2024 Smoothing Outlier Scores Is All You Need to Improve Outlier Detectors
abstract
We hypothesize thatsimilar objects should have similar outlier scores. To the best of our knowledge, all existing outlier detectors calculate the outlier score for each object independently regardless of the outlier scores of the other objects. Therefore, they do not guarantee that similar objects have similar outlier scores. To verify our proposed hypothesis, we propose an outlier score post-processing technique for outlier detectors, called neighborhood averaging (NA) for neighborhood smoothing in outlier score space. It pays attention to objects and their neighbors and guarantees them to have more similar outlier scores than their original scores. Given an object and its outlier score from any outlier detector, NA modifies its outlier score by combining it with its$k$nearest neighbors' scores. We demonstrate the effectivity of NA by using the well-known$k$nearest neighbors ($k$-NN). Experimental results show that NA improves all 10 tested baseline detectors by 13% on average relative to the original results (from 0.70 to 0.79 AUC) evaluated on nine real-world datasets. Moreover, deep-learning-based detectors and even outlier detectors that are already based on$k$-NN are also improved. The experiments also show that in some applications, the choice of detector is no more significant when detectors are jointly used with NA. This may pose a challenge to the generally considered idea that the data model is the most important factor. We open our code on www.outlierNet.com for reproducibility.
Jiawei Yang 0001, Susanto Rahardja, Pasi Fränti
IEEE Trans. Knowl. Data Eng.2
2024 Learning Deep Representations for Photo Retouching
abstract
Photo enhancement is a long-standing and challenging problem in image processing community. Despite having witnessed significant achievements in recent years, many of them are built upon supervised learning theories and thus required expertise in constructing a huge collection of paired data, which is well-known to be a problem as the acquisition of such data in real life can be impractical. We address this issue by proposing a multi-scale GAN framework that can be trained in an unsupervised fashion. Notably, we unify the design principle of the generator and discriminator in our framework so as to maximize the ability to learn deep latent representations. Specifically, rather than maintaining the content consistency through complicated two-way loss, we present a one-way loss that measures the content distance between multi-scale latent representations of inputs and outputs to speed up the training by$\text{1.7}\times$. Furthermore, we redesign the discriminator into a multi-scale-multi-stage manner to strengthen the adversarial learning, where the multiple latent features with varying scales are produced by the main discriminator and these features are then sent to auxiliary discriminators for final recognition. Extensive experiments have been conducted in the well-known MIT-Adobe-fivek and HDR+ datasets, and the results demonstrated that the proposed multi-scale representation learning framework shows outstanding performance in photo enhancement task.
Di Li 0006, Susanto Rahardja
IEEE Trans. Multim.2
2023 High Dynamic Range Image Tone Mapping Based on Layer Decomposition and Image Fusion
abstract
Common displays have limited dynamic range and therefore cannot support direct reproduction of high dynamic range (HDR) images. Generally, this gap can be effectively solved by image tone mapping. One of the typical tone mapping methods is known as layer decomposition based algorithms. They decompose an HDR image into base and detail layers, and reduce the dynamic range by employing gamma functions on the base layer. However, simple gamma curves are not sufficient to produce good local contrasts. This paper proposes a layer decomposition based method by utilizing image fusion techniques to enhance local contrasts. Our method constructs virtual images to stretch dark, normal and bright regions of the base layer respectively and fuses them together for contrast enhancement. Experimental results validate that the proposed method produces enhanced local contrasts and more natural colors in the tone-mapped images.
Xueyu Han, Susanto Rahardja
ICIP2
2023 An Alternative to Bilinear and Nearest-Neighbour Enlarging for Monitor Displays
abstract
This paper proposes a simple and effective algorithm for image enlargement, aims to improve upon the widely-used bi-linear and nearest-neighbour interpolations for monitor displays. The proposed method uses nearest-neighbour interpolation to generate preliminary enlarged images, and selectively modifies only diagonal edges to reduce blocking artifacts. Experiments demonstrate the proposed method produces clear images with fewer artifacts across different enlargement factors and scenes.
Shumin Liu, Jie Chen 0022, Susanto Rahardja
ICIP3
2023 Correction: A lightweight classification of adaptor proteins using transformer networks
Sylwan Rahardja, Mou Wang, Binh P. Nguyen, Pasi Fränti, Susanto Rahardja
BMC Bioinform.5
2023 Pure Number Discrete Fractional Complex Hadamard Transform
abstract
This paper introduces a novel discrete fractional transform termed as pure number discrete fractional complex Hadamard transform (PN-FCHT). The proposed PN-FCHT offers three advantages over the traditional discrete fractional Hadamard transform (FHT). Firstly, the higher-order PN-FCHT matrix exhibits the Self-Kronecker product structure, which allows for the recursive generation from the$2\times 2$core PN-FCHT matrix. Secondly, it possesses two important properties for computation, i.e. pure number property. Lastly, compared to existing state-of-the-art fast FHT algorithms, the PN-FCHT can reduce the transform multiplication computational complexity by up to 80% and this results in a more efficient hardware implementation.
Zi-Chen Fan, Di Li 0006, Susanto Rahardja
IEEE Signal Process. Lett.3
2023 Classification of Interbeat Interval Time-Series Using Attention Entropy
abstract
Classification of interbeat interval time-series which fluctuates in an irregular and complex manner is very challenging. Typically, entropy methods are employed to quantify the complexity of the time-series for classifying. Traditional entropy methods focus on the frequency distribution of all the observations in a time-series. This requires a relatively long time-series with at least a couple of thousands of data points, which limits their usages in practical applications. The methods are also sensitive to the parameter settings. In this paper, we propose a conceptually new approach calledattention entropy, which pays attention only to the key observations. Instead of counting the frequency of all observations, it analyzes the frequency distribution of the intervals between the key observations in a time-series. Attention entropy does not need any parameter to tune, it is robust to the time-series length, and requires only linear time to compute. Experiments show that it outperforms fourteen state-of-the-art entropy methods evaluated by real-world datasets. It achieves average classification accuracy of AUC = 0.71 while the second-best method, multiscale entropy, achieves AUC = 0.62 when classifying four groups of people with a time-series length of 100.
Jiawei Yang 0001, Gulraiz Iqbal Choudhary, Susanto Rahardja, Pasi Fränti
IEEE Trans. Affect. Comput.3
2023 End-to-End Multi-Modal Speech Recognition on an Air and Bone Conducted Speech Corpus
abstract
Automatic speech recognition (ASR) has been significantly improved in the past years. However, most robust ASR systems are based on air-conducted (AC) speech, and their performances in low signal-to-noise-ratio (SNR) conditions are not satisfactory. Bone-conducted (BC) speech is intrinsically insensitive to environmental noise, and therefore can be used as an auxiliary source for improving the performance of an ASR at low SNR. In this paper, we first develop a multi-modal Mandarin corpus, which contains air- and bone-conducted synchronized speech (ABCS). The multi-modal speeches are recorded with a headset equipped with both AC and BC microphones. To our knowledge, it is by far the largest corpus for conducting bone conduction ASR research. Then, we propose a multi-modal conformer ASR system based on a novel multi-modal transducer (MMT). The proposed system extracts semantic embeddings from the AC and BC speech signals by a conformer-based encoder and a transformer-based truncated decoder. The semantic embeddings of the two speech sources are fused dynamically with adaptive weights by the MMT module. Experimental results demonstrate the proposed multi-modal system outperforms single-modal systems with either AC or BC modality and multi-modal baseline system by a large margin at various SNR levels. It also shows the two modalities complement with each other, and our method can effectively utilize the complementary information of different sources.
Mou Wang, Junqi Chen 0001, Xiao-Lei Zhang 0001, Susanto Rahardja
IEEE ACM Trans. Audio Speech Lang. Process.4
2023 Exploiting Symmetries in the Design and Implementation of Polynomial/Trigonometric Beamformers for Uniform Circular Arrays
abstract
The polynomial and trigonometric beamformers for uniform circular arrays (UCAs), which are respectively based on the polynomial and trigonometric interpolation, have been shown to enable dynamic and continuous beam steering over the entire 360° azimuth range through simple online parameter tuning. The ability of beam steering of the polynomial/trigonometric beamformers, however, comes at the cost of higher complexity in the beamformer design and implementation. This paper derives some symmetry properties of the coefficients of the polynomial/trigonometric beamformers for UCAs. By jointly exploiting the symmetries of the beamformer coefficients, a reduced-dimension design approach and a simplified implementation structure for the polynomial/trigonometric beamformers for UCAs are proposed. It is shown that the dimension of the beamformer design optimization problem can be reduced by as much as over 60% using the reduced-dimension design, enabling a fast design process. Moreover, the simplified implementation structure is highly cost-effective. Compared with the original implementation structure, a reduction rate of nearly 75% in the coefficient storage requirement can be achieved. And the number of the multipliers and adders can also be substantially reduced by around 50%.
Congwei Feng, Susanto Rahardja
IEEE Trans. Circuits Syst. I Regul. Pap.4
2022 End-To-End Multi-Modal Speech Recognition with Air and Bone Conducted Speech
abstract
Improving the performance of automatic speech recognition (ASR) in adverse acoustic environments is a long-term tough task. Although many robust ASR systems based on conventional microphones have been developed, their performance with air-conducted (AC) speech is still far from satisfactory in low signal-to-noise-ratio (SNR) environments. Bone-conducted (BC) speech is relatively insensitive to ambient noise, and has a potential of promoting the ASR performance at such low SNR environments as an auxiliary source. In this paper, we propose a conformer-based multi-modal speech recognition system. It uses a conformer encoder and a transformer-based truncated decoder to extract the semantic information from AC and BC channels respectively. The semantic information of the two channels are re-weighted and integrated by a novel multi-modal transducer. Experimental results show the effectiveness of the proposed method. For example, given a 0 dB SNR environment, it yields a character error rate of over 59.0% lower than a noise-robust baseline conducted on AC channel only, and over 12.7% lower than a multi-modal baseline that takes the concatenated features of AC and BC speech as the input.
Junqi Chen 0001, Mou Wang, Xiao-Lei Zhang 0001, Zhiyong Huang 0001, Susanto Rahardja
ICASSP5
2022 A lightweight classification of adaptor proteins using transformer networks
abstract
BACKGROUND: Adaptor proteins play a key role in intercellular signal transduction, and dysfunctional adaptor proteins result in diseases. Understanding its structure is the first step to tackling the associated conditions, spurring ongoing interest in research into adaptor proteins with bioinformatics and computational biology. Our study aims to introduce a small, new, and superior model for protein classification, pushing the boundaries with new machine learning algorithms. RESULTS: We propose a novel transformer based model which includes convolutional block and fully connected layer. We input protein sequences from a database, extract PSSM features, then process it via our deep learning model. The proposed model is efficient and highly compact, achieving state-of-the-art performance in terms of area under the receiver operating characteristic curve, Matthew's Correlation Coefficient and Receiver Operating Characteristics curve. Despite merely 20 hidden nodes translating to approximately 1% of the complexity of previous best known methods, the proposed model is still superior in results and computational efficiency. CONCLUSIONS: The proposed model is the first transformer model used for recognizing adaptor protein, and outperforms all existing methods, having PSSM profiles as inputs that comprises convolutional blocks, transformer and fully connected layers for the use of classifying adaptor proteins.
Sylwan Rahardja, Mou Wang, Binh P. Nguyen, Pasi Fränti, Susanto Rahardja
BMC Bioinform.5
2022 Perceptual Loss-Constrained Adversarial Autoencoder Networks for Hyperspectral Unmixing
abstract
Recently, the use of a deep autoencoder-based method in blind spectral unmixing has attracted great attention as the method can achieve superior performance. However, most autoencoder-based unmixing methods use non-structured reconstruction loss to train networks, leading to the ignorance of band-to-band-dependent characteristics and fine-grained information. To cope with this issue, we propose a general perceptual loss-constrained adversarial autoencoder network for hyperspectral unmixing. Specifically, the adversarial training process is used to update our framework. The discriminate network is found to be efficient in discovering the discrepancy between the reconstructed pixels and their corresponding ground truth. Moreover, the general perceptual loss is combined with the adversarial loss to further improve the consistency of high-level representations. Ablation studies verify the effectiveness of the proposed components of our framework, and experiments with both synthetic and real data illustrate the superiority of our framework when compared with other competing methods.
Min Zhao 0014, Mou Wang, Jie Chen 0022, Susanto Rahardja
IEEE Geosci. Remote. Sens. Lett.4
2022 Sparse random projection isolation forest for outlier detection
Xu Tan 0004, Jiawei Yang 0001, Susanto Rahardja
Pattern Recognit. Lett.3
2022 Sparse Linear Spectral Unmixing of Hyperspectral Images Using Expectation-Propagation
abstract
This article presents a novel Bayesian approach for hyperspectral image unmixing. The observed pixels are modeled by a linear combination of material signatures weighted by their corresponding abundances. A spike-and-slab abundance prior is adopted to promote sparse mixtures and an Ising prior model is used to capture spatial correlation of the mixture support across pixels. We approximate the posterior distribution of the abundances using the expectation-propagation (EP) method. We show that it can significantly reduce the computational complexity of the unmixing stage and meanwhile provide uncertainty measures, compared to expensive Monte Carlo strategies traditionally considered for uncertainty quantification. Moreover, many variational parameters within each EP factor can be updated in a parallel manner, which enables mapping of efficient algorithmic architectures based on graphics processing units (GPUs). Under the same approximate Bayesian framework, we then extend the proposed algorithm to semi-supervised unmixing, whereby the abundances are viewed as latent variables and the expectation-maximization (EM) algorithm is used to refine the endmember matrix. Experimental results on synthetic data and real hyperspectral data illustrate the benefits of the proposed framework over state-of-art linear unmixing methods.
Zeng Li 0001, Yoann Altmann, Jie Chen 0022, Steve McLaughlin 0001, Susanto Rahardja
IEEE Trans. Geosci. Remote. Sens.5
2022 Hyperspectral Unmixing for Additive Nonlinear Models With a 3-D-CNN Autoencoder Network
abstract
Spectral unmixing is an important task in hyperspectral image processing for separating the mixed spectral data pertaining to various materials observed aiming at analyzing the material components in observed pixels. Recently, nonlinear spectral unmixing has received particular attention in hyperspectral image processing, as there are many situations in which the linear mixture model may not be appropriate and could be advantageously replaced by a nonlinear one. Existing nonlinear unmixing approaches are often based on specific assumptions on the nonlinearity and can be less effective when used for scenes with unknown nonlinearity. This article presents an unsupervised nonlinear spectral unmixing method that addresses a general model that consists of a linear mixture part and an additive nonlinear mixture part. The structure of a deep autoencoder network, which has a clear physical interpretation, is specifically designed to achieve this purpose. Moreover, a convolutional neural network (CNN) is used to capture the spectral-spatial priors from hyperspectral data. Extensive experiments with synthetic and real data illustrate the generality and effectiveness of this scheme compared with state-of-the-art methods.
Min Zhao 0014, Mou Wang, Jie Chen 0022, Susanto Rahardja
IEEE Trans. Geosci. Remote. Sens.4
2021 Hyperspectral Shadow Removal via Nonlinear Unmixing
abstract
Removing shadows that are often present in remotely sensed hyperspectral images is important for both enhancing the interpretability of the data and further target analysis. Shadow removal approaches based on spectral unmixing have been proposed in the literature using the linear mixture model. However, objects that produce shadows may also introduce light scattering, and the higher order interactions of photons can cause nonlinearity. This letter integrates the nonlinear hyperperspectral unmixing into the unmixing-based shadow removal, and the effects of applying typical nonlinear algorithms within the approach are investigated. The usefulness of nonlinear unmixing in hyperspectral shadow removal is verified based on the results of applications to both laboratory-created real data and actual airborne data.
Min Zhao 0014, Jie Chen 0022, Susanto Rahardja
IEEE Geosci. Remote. Sens. Lett.3
2021 Mean-shift outlier detection and filtering
abstract
Traditional outlier detection methods create a model for data and then label as outliers for objects that deviate significantly from this model. However, when dat has many outliers, outliers also pollute the model. The model then becomes unreliable, thus rendering most outlier detectors to become ineffective. To solve this problem, we propose a mean-shift outlier detector. This detector employs a mean-shift technique to modify data and cancel the bias caused by the outliers. The mean-shift technique replaces every object by the mean of its k-nearest neighbors which essentially removes the effect of outliers before clustering without the need to know the outliers. In addition, it also detects outliers based on the distance shifted. Our experiments show that the proposed method works well regardless of the number of outliers in the data. This method outperforms all state-of-the-art methods tested, with both real-world numeric datasets as well as generated numeric and string datasets.
Jiawei Yang 0001, Susanto Rahardja, Pasi Fränti
Pattern Recognit.2
2020 Sparse Spectral Unmixing of Hyperspectral Images using Expectation-Propagation
abstract
The aim of spectral unmixing of hyperspectral images is to determine the component materials and their associated abundances from mixed pixels. In this paper, we present sparse linear unmixing via an Expectation-Propagation method based on the classical linear mixing model and a spike-and-slab prior promoting abundance sparsity. The proposed method, which allows approximate uncertainty quantification (UQ), is compared to existing sparse unmixing methods, including Monte Carlo strategies traditionally considered for UQ. Experimental results on synthetic data and real hyperspectral data illustrate the benefits of the proposed algorithm over state-of-art linear unmixing methods.
Zeng Li 0001, Yoann Altmann, Jie Chen 0022, Steve McLaughlin 0001, Susanto Rahardja
VCIP5
2020 A Novel Area-Power Efficient Design for Approximated Small-Point FFT Architecture
abstract
Fast Fourier transform (FFT) is an essential algorithm in digital signal processing and advanced mobile communications. With the continuous development of modern technology, the area-power efficient hardware implementation of FFT has attracted a lot of attention. In this article, a novel design for FFT implementation is proposed. The number of resource-expensive multiplications in our design is decreased by a twiddle factor merging technique that reduces the hardware area. Subsequently, a common subexpression sharing scheme is applied to reuse the hardware resources to further save the hardware area. In addition, a magnitude-response aware approximation algorithm is proposed for applications where the transformation accuracy can be compromised a little bit for lesser hardware area and power dissipation. Logic synthesis shows that the proposed 16-point FFT architecture can save hardware area and power dissipation on application-specific integrated circuit (ASIC) by up to 65.7% and 53.1% compared with recently published designs. Similarly, the proposed 32-point FFT architecture achieves up to 58.8% reduction on hardware area and 60.0% reduction on power dissipation on ASIC.
Xueyu Han, Jiajia Chen 0002, Boyu Qin, Susanto Rahardja
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2020 A New Multi-Focus Image Fusion Algorithm and Its Efficient Implementation
abstract
Wavelet transform using Haar filter is a fast process for decomposing an image into low and high frequency sub-bands, which is an important step in multi-focus image fusion. This transformation alone is usually not sufficient to produce satisfying fused images, and therefore additional tools like focused region decision map and feature extraction are needed to enhance the fusion quality. In this paper, we propose a multi-focus image fusion algorithm that is suitable for efficient hardware implementation. The algorithm speed is significantly improved by utilizing Haar wavelet and computationally simple fusion rules. To further improve the circuit power and delay, we limit the logic operations to additions, subtractions, and multiplications only. Experiments on different benchmarks demonstrate that our proposed algorithm can reduce the fusion time by up to 99.97% compared with the state-of-the-art competing methods without compromising the fusion quality. In addition, the hardware implementation of the proposed algorithm reduces the circuit area and delay by 79.5% and 92.5%, respectively, when compared to the most competitive algorithm in this paper while maintaining similar fusion quality.
ShuMin Liu, Jiajia Chen 0002, Susanto Rahardja
IEEE Trans. Circuits Syst. Video Technol.3
2020 An Optimized Quantization Constraints Set for Image Restoration and its GPU Implementation
abstract
This paper presents a novel optimized quantization constraint set, acting as an add-on to existing DCT-based image restoration algorithms. The constraint set is created based on generalized Gaussian distribution which is more accurate than the commonly used uniform, Gaussian or Laplacian distributions when modeling DCT coefficients. More importantly, the proposed constraint set is optimized for individual input images and thus it is able to enhance image quality significantly in terms of signal-to-noise ratio. Experimental results indicate that the signal-to-noise ratio is improved by at least 6.78% on top of the existing state-of-the-art methods, with a corresponding expense of only 0.38% in processing time. The proposed algorithm has also been implemented in GPU, and the processing speed increases further by 20 times over that of CPU implementation. This makes the algorithm well suited for fast image retrieval in security and quality monitoring system.
ShuMin Liu, Jiajia Chen 0002, Ye Ai, Susanto Rahardja
IEEE Trans. Image Process.4
2019 AUC Optimization for Deep Learning Based Voice Activity Detection
abstract
Voice activity detection (VAD) based on deep neural networks (DNN) has demonstrated good performance in adverse acoustic environments. Current DNN based VAD optimizes a surrogate function, e.g. minimum cross-entropy or minimum squared error, at a given decision threshold. However, VAD usually works on-the-fly with a dynamic decision threshold; and ROC curve is a global evaluation metric of VAD that reflects the performance of VAD at all possible decision thresholds. In this paper, we propose to optimize the area under ROC curve (AUC) by DNN, which can maximize the performance of VAD in terms of the ROC curve. Experimental results show that optimizing AUC by DNN results in higher performance than the common method of optimizing the minimum squared error by DNN.
Zi-Chen Fan, Zhongxin Bai, Xiao-Lei Zhang 0001, Susanto Rahardja, Jingdong Chen
ICASSP4
2019 Residual U-Net for Retinal Vessel Segmentation
abstract
In recent years, the influence of deep learning on retinal vessel segmentation has grown rapidly. Most of the available deep learning based methods use relatively shallow structures. However, due to the limited representative capacity, shallow networks will restrain deep learning models to segment both vessel and non-vessel pixels accurately. In this paper, we propose a residual U-Net for retinal vessel segmentation. Our network has several advantages. First, the network uses a new residual block structure. In the new structure, batch normalization layers are placed before the activation unit to achieve better performance and accelerate the convergence. Also, a dropout layer is utilized in the structure to alleviate over-fitting problems. Second, the depth of the network is increased by adding more residual blocks and strong dropouts which then allow the network to extract features better. Fundus images from the publicly available DRIVE and STARE datasets are used to evaluate the proposed network. Experimental result shows that the proposed modified residual U-Net has better performance than existing state-of-the-art algorithms.
Di Li 0006, Dhimas Arief Dharmawan, Boon Poh Ng, Susanto Rahardja
ICIP4
2019 iProDNA-CapsNet: identifying protein-DNA binding residues using capsule neural networks
abstract
BACKGROUND: Since protein-DNA interactions are highly essential to diverse biological events, accurately positioning the location of the DNA-binding residues is necessary. This biological issue, however, is currently a challenging task in the age of post-genomic where data on protein sequences have expanded very fast. In this study, we propose iProDNA-CapsNet - a new prediction model identifying protein-DNA binding residues using an ensemble of capsule neural networks (CapsNets) on position specific scoring matrix (PSMM) profiles. The use of CapsNets promises an innovative approach to determine the location of DNA-binding residues. In this study, the benchmark datasets introduced by Hu et al. (2017), i.e., PDNA-543 and PDNA-TEST, were used to train and evaluate the model, respectively. To fairly assess the model performance, comparative analysis between iProDNA-CapsNet and existing state-of-the-art methods was done. RESULTS: Under the decision threshold corresponding to false positive rate (FPR) ≈ 5%, the accuracy, sensitivity, precision, and Matthews's correlation coefficient (MCC) of our model is increased by about 2.0%, 2.0%, 14.0%, and 5.0% with respect to TargetDNA (Hu et al., 2017) and 1.0%, 75.0%, 45.0%, and 77.0% with respect to BindN+ (Wang et al., 2010), respectively. With regards to other methods not reporting their threshold settings, iProDNA-CapsNet also shows a significant improvement in performance based on most of the evaluation metrics. Even with different patterns of change among the models, iProDNA-CapsNets remains to be the best model having top performance in most of the metrics, especially MCC which is boosted from about 8.0% to 220.0%. CONCLUSIONS: According to all evaluation metrics under various decision thresholds, iProDNA-CapsNet shows better performance compared to the two current best models (BindN and TargetDNA). Our proposed approach also shows that CapsNet can potentially be used and adopted in other biological applications.
Binh P. Nguyen, Quang H. Nguyen 0001, Giang-Nam Doan-Ngoc, Thanh-Hoang Nguyen-Vo, Susanto Rahardja
BMC Bioinform.5
2019 Nonlinear Unmixing of Hyperspectral Data via Deep Autoencoder Networks
abstract
Nonlinear spectral unmixing is an important and challenging problem in hyperspectral image processing. Classical nonlinear algorithms are usually derived based on specific assumptions on the nonlinearity. In recent years, deep learning shows its advantage in addressing general nonlinear problems. However, existing ways of using deep neural networks for unmixing are limited and restrictive. In this letter, we develop a novel blind hyperspectral unmixing scheme based on a deep autoencoder network. Both encoder and decoder of the network are carefully designed so that we can conveniently extract estimated endmembers and abundances simultaneously from the nonlinearly mixed data. Because an autoencoder is essentially an unsupervised algorithm, this scheme only relies on the current data and, therefore, does not require additional training. Experimental results validate the proposed scheme and show its superior performance over several existing algorithms.
Mou Wang, Min Zhao 0014, Jie Chen 0022, Susanto Rahardja
IEEE Geosci. Remote. Sens. Lett.4
2018 New Hardware and Power Efficient Sporadic Logarithmic Shifters for DSP Applications
abstract
Shifting an input data by variable amounts is commonly found in arithmetic operations, data encoding and bit-indexing. Although some shift amounts along the entire shift range are not required, the typical realization by a full-range logarithm shifter requires full implementation and therefore suffers from complexity and power overhead. In this paper, the notion of sporadic logarithmic shifter (SLS) is introduced for the first time, and a new design methodology is proposed for its optimization. By reusing parts of existing substructure of conventional logarithmic shifter or post-multiplexing the hardwired shifts, contagious subranges of desirable shift amounts are successively realized. Synthesis results on 8-bit and 16-bit SLSs show average application-specified integrated circuit area and power savings of up to 73.24% and 63.90%, respectively, over conventional logarithmic shifters. In addition, by applying the proposed SLSs to the discrete cosine transform (DCT) architecture, at least 47.5% area savings and 2.9% power savings can be achieved over two constant multipliers-based DCT architectures reported in the literature.
Jiajia Chen 0002, Chip-Hong Chang, Juan Zhao 0005, Susanto Rahardja
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2016 Distributed private online learning for social big data computing over data center networks
abstract
With the rapid growth of Internet technologies, cloud computing and social networks have become ubiquitous. An increasing number of people participate in social networks and massive online social data are obtained. In order to exploit knowledge from copious amounts of data obtained and predict social behavior of users, we urge to realize data mining in social networks. Almost all online websites use cloud services to effectively process the large scale of social data, which are gathered from distributed data centers. These data are so large-scale, high-dimension and widely distributed that we propose a distributed sparse online algorithm to handle them. Additionally, privacy-protection is an important point in social networks. We should not compromise the privacy of individuals in networks, while these social data are being learned for data mining. Thus we also consider the privacy problem in this article. Our simulations shows that the appropriate sparsity of data would enhance the performance of our algorithm and the privacy-preserving method does not significantly hurt the performance of the proposed algorithm.
Chencheng Li, Pan Zhou 0001, Yingxue Zhou, Kaigui Bian, Tao Jiang 0002, Susanto Rahardja
ICC6
2014 A new error-mapping scheme for scalable audio coding
abstract
In scalable audio coders, such as the MPEG-4 SLS, error-mapping is used to map quantization errors in the core coder to an error signal before passing through bit-plane coding. In this paper, we propose a new error-mapping scheme that is derived by observing statistical properties of the error signal. Compared with the error-mapping in SLS, the proposed scheme improves coding efficiency as well as computational complexity of the coder. An average improvement of 9 points in MUSHRA score has been achieved by the proposed scheme in subjective listening tests. The proposed error-mapping adds a useful new tool to the existing toolset for constructing next-generation scalable audio coders.
Susanto Rahardja
MMSP2
2013 Perceptually relevant energy function for seam carving
abstract
Seam carving, an image re-targeting method, works by progressively finding and removing connected paths of low energy pixels in an image until a desired image aspect ratio is reached. In this paper, we first cast the problem of minimizing an energy function as that of minimizing a distortion cost. We then leverage on our understanding of image quality metrics/ distortion metrics in proposing a perceptually relevant energy function. Experimental results show that our proposed energy function can generate more desirable resized images in which the original structures of the images are better preserved.
Hui Li Tan, Yih Han Tan, Zhengguo Li, Susanto Rahardja, Chuohao Yeo
ICASSP4
2013 A novel scalable audio coding scheme
abstract
A new scalable audio coding scheme is introduced in this paper. Its core idea is to create one additional scalability dimension during the encoding process for the purpose of generating a plural of scalable sub-bitstreams. Based on the multiple sub-streams, a smart truncator is designed that can truncate these sub-bitstreams with optimal rate-distortion (R-D) tradeoff. Benefited from the flexible R-D trade-off, the proposed new scheme could, within a wide bitrate range, outperform those traditional scalable coding schemes, which usually provides a fixed R-D relationship designed at a specified bitrate. To verify the performance, the proposed scheme is further implemented based on a prior art scalable audio codec. Significant quality improvement is observed from the new codec via a series of subjective listening tests.
Haiyan Shu, Rongshan Yu, Susanto Rahardja
ICASSP5
2013 Time utility function based packet scheduling algorithm for streaming scalable media
abstract
In this paper, a packet scheduling algorithm that is based on a Time-Utility Function (TUF) is proposed. In the proposed system, the scalable media is partitioned into data units of different quality layers, which are then prioritized and transmitted according to their TUF's that capture both their quality contributions to the decoded media and urgencies with respect to their playback schedule. For optimal streaming quality and meanwhile maintaining a reasonable computational complexity, the scheduling of packet transmission is obtained from a low-complexity packet scheduling algorithm based on utility accrual maximization. Experimental results show that the proposed scheduling algorithm achieves near optimal performance when compared to the operational rate distortion bound of the stream source at the capacity of the network.
Rongshan Yu, Haiyan Shu, Susanto Rahardja
ICME3
2013 Intra Coding With Adaptive Partial Reconstruction
abstract
Intra prediction improves coding performance by reducing inter pixel redundancy. However, to accommodate the use of block transforms, not all pixels can be predicted from reconstructed pixels that are located close to themselves. This causes prediction performance to suffer as pixel values further apart are less correlated. This paper presents additional intra coding modes designed with the goal of improving prediction performance. Experimental results show an average gain of about 2% in the key technical area software when the new modes are incorporated in the current 8$\,\times\,$8 prediction modes. Since the new coding modes (8$\,\times\,$8) are designed with transform size smaller than coding block size, the modes can also be useful when the source block is larger than the maximum transform size.
Yih Han Tan, Chuohao Yeo, Zhengguo Li, Susanto Rahardja
IEEE Trans. Circuits Syst. Video Technol.4
2013 A Perceptually Relevant MSE-Based Image Quality Metric
abstract
Image quality metrics (IQMs), such as the mean squared error (MSE) and the structural similarity index (SSIM), are quantitative measures to approximate perceived visual quality. In this paper, through analyzing the relationship between the MSE and the SSIM under an additive noise distortion model, we propose a perceptually relevant MSE-based IQM, MSE-SSIM, which is expressed in terms of the variance of the source image and the MSE between the source and distorted images. Evaluations on three publicly available databases (LIVE, CSIQ, and TID2008) show that the proposed metric, despite requiring less computation, compares favourably in performance to several existing IQMs. In addition, due to its simplicity, MSE-SSIM is amenable for the use in a wide range of image and video tasks that involve solving an optimization problem. As an example, MSE-SSIM is used as the objective function in designing a Wiener filter that aims at optimizing the perceptual visual quality of the output. Experimental results show that the images filtered with a MSE-SSIM-optimal Wiener filter have better visual quality than those filtered with a MSE-optimal Wiener filter.
Hui Li Tan, Zhengguo Li, Yih Han Tan, Susanto Rahardja, Chuohao Yeo
IEEE Trans. Image Process.4
2013 Hybrid Patching for a Sequence of Differently Exposed Images With Moving Objects
abstract
It is very challenging to synthesize a high dynamic range (HDR) image from multiple differently exposed low dynamic range images when there are moving objects in the images. This is due to the fact that the moving objects will cause ghosting artifacts to appear in the synthesized HDR image. To prevent such artifacts, a patching algorithm is required to correct motion regions such that all the moving objects are synchronized in the differently exposed images. In this paper, a new optimization problem is formulated to correct the motion regions of the multiple differently exposed images by considering both spatial and temporal consistencies. The resultant scheme is a hybrid patching scheme composed of a correction method which is an intensity mapping function at pixel level, and a hole-filling method that uses block-level template matching. The proposed patching scheme is not only robust to large intensity changes in these input images, but also at regions that are over- or underexposed. Experimental results show that the proposed method is able to prevent ghosting artifacts from appearing in the final synthesized HDR image.
Jinghong Zheng 0001, Zhengguo Li, Shiqian Wu, Susanto Rahardja
IEEE Trans. Image Process.5
2012 A local intensity adaptive structural similarity index
abstract
Existing structural similarity (SSIM) index comprises of one term on luminance comparison and the other term on contrast and structure comparison. In this paper, the SSIM index is first improved by introducing three weighting factors to the second term such that it is adaptive to local intensities of two images to be compared. The improved SSIM (iSSIM) index is further extended for two images with possibly different exposures. Experimental results show that the proposed indices are more robust to large intensity changes of two images from the same scene and more sensitive to two images from different scenes than the existing SSIM index.
Zhengguo Li, Chuohao Yeo, Yih Han Tan, Susanto Rahardja
ICASSP4
2012 A bilateral filter in gradient domain
abstract
In this paper, a bilateral filter in gradient domain is first proposed. It is then applied to study detail enhancement via multi-light images and noise reduction of differently exposed low dynamic range images. These two applications show that the proposed filter can be applied to extract fine details from a set of images simultaneously and to provide flexibility for noise reduction from selected areas of an image.
Zhengguo Li, Jinghong Zheng 0001, Shiqian Wu, Susanto Rahardja
ICASSP5
2012 An alternating direction method for frame-based image deblurring with balanced regularization
abstract
In this paper, we propose an efficient algorithm for solving a balanced approach in frame-based image deblurring. The balanced approach is usually formulated as a minimization problem involving an ℓ2data-fidelity term, an ℓ1regularizer on sparsity of frame coefficients, and a penalty on distance of sparse frame coefficients to the canonical frame coefficients. The balanced approach bridges synthesis-based and analysis-based approaches. Our algorithm is based on a variable splitting strategy and the classical alternating direction method (ADM). This paper shows how the proposed algorithm can be applied to solve the balanced approach efficiently. More precisely, a regularized version of the Hessian matrix of the ℓ2data-fidelity term is involved, and by exploiting fast tight frame and circular structure of the observation matrix, the matrix can perform efficiently for image deblurring application. Convergence of the proposed algorithm is guaranteed by the existing ADM theory. Numerical simulations illustrate the efficiency of our proposed algorithm in frame-based image deblurring.
Shoulie Xie, Susanto Rahardja
ICASSP2
2012 Noise reduction for differently exposed images
abstract
For scenes under low lighting condition, cameras are usually set to a high sensitivity (ISO) mode to reduce motion blur at the cost of increased image noise. When multiple differently exposed images are used to generate a high dynamic range (HDR) image, the high ISO noise from each low dynamic range (LDR) image can be further amplified by the HDR synthesis algorithm which would result in severely degradation of visual quality. This paper proposes an intensity mapping function based noise reduction method for differently exposed images with high ISO noise. The proposed method does not require any knowledge on either camera response functions or exposure times. In addition, the method is simple yet effective for noise removal from the LDR images without introducing any blurring or other artifacts.
Wei Yao 0001, Zhengguo Li, Susanto Rahardja
ICASSP3
2012 Anti-ghost of differently exposed images with moving objects
abstract
In a typical image synthesis where multiple differently exposed images are captured for processing, it is important to design an anti-ghost algorithm so as to prevent ghosting artifacts from appearing in the final image. An anti-ghost algorithm is usually composed of a detection module and a correction module. In this paper, a new detection module is proposed to detect non-consistent pixels of all input images without predefining any initial reference image. The proposed module is suitable when an interactive mode is desired. In addition, a bidirectional approach is introduced to correct the non-consistent pixels in the correction module. Compared with existing unidirectional correction methods, the proposed bidirectional correction approach uses information from two adjacent images of a detected image to correct its non-consistent pixels. This leads to a quality improvement in the final image.
Zhengguo Li, Shiqian Wu, Shoulie Xie, Susanto Rahardja
ICIP5
2012 Joint rate allocation for statistical multiplexing of SVC
abstract
This paper presents a joint rate allocation scheme for statistical multiplexing of multiprogram video coding in broadcasting systems. The scheme is based on a scalable video coding (SVC) platform that does not require computationally expensive re-encoding or transcoding to adjust the bit-rate of each video program. A piecewise linear model is used to estimate the rate-distortion relationship in SVC enhancement layers. Based on the model, a joint rate allocation scheme is developed to dynamically allocate the available bandwidth by taking into consideration both inter-program fairness and intraprogram smoothness constraints. Experiments have been carried out to compare the performance of existing methods with our proposed scheme. Results demonstrate that the proposed scheme achieves a fine balance in picture quality across all statistical multiplexed programs as well as within each program.
Wei Yao 0001, Lap-Pui Chau, Susanto Rahardja
ICIP3
2012 Object Recognition by Discriminative Combinations of Line Segments, Ellipses, and Appearance Features
abstract
We present a novel contour-based approach that recognizes object classes in real-world scenes using simple and generic shape primitives of line segments and ellipses. Compared to commonly used contour fragment features, these primitives support more efficient representation since their storage requirements are independent of object size. Additionally, these primitives are readily described by their geometrical properties and hence afford very efficient feature comparison. We pair these primitives as shape-tokens and learn discriminative combinations of shape-tokens. Here, we allow each combination to have a variable number of shape-tokens. This, coupled with the generic nature of primitives, enables a variety of class-specific shape structures to be learned. Building on the contour-based method, we propose a new hybrid recognition method that combines shape and appearance features. Each discriminative combination can vary in the number and the types of features, where these two degrees of variability empower the hybrid method with even more flexibility and discriminative potential. We evaluate our methods across a large number of challenging classes, and obtain very competitive results against other methods. These results show the proposed shape primitives are indeed sufficiently powerful to recognize object classes in complex real-world scenes.
Alex Yong Sang Chia, Deepu Rajan, Maylor Karhang Leung, Susanto Rahardja
IEEE Trans. Pattern Anal. Mach. Intell.4
2012 Analysis of Bit-Plane Probability for Generalized Gaussian Distribution and its Application in Audio Coding
abstract
Bit-plane probability of data is useful information for many applications such as entropy coding and rate estimation, particularly for scalable coding system. Typically, the value of bit-plane probability varies according to the distribution of the data represented, and due to inter-plane correlation, it is hard to get an analytic solution of bit-plane probability for the data with generalized Gaussian distribution. In this paper, bit-plane probability is analyzed when Generalized Gaussian distribution is used to model the input data. Based on the study of bit-plane probability for Laplace distribution and the relationship between different bit-planes, an approximated bit-plane probability for Generalized Gaussian distribution is presented. This closed-form expression is of low computational cost. Furthermore, a much practical format is derived with reduced complexity for implementation in the state-of-the-art MPEG-4 scalable to lossless audio coding system. With the same computational cost, the proposed algorithm presents higher compression efficiency than MPEG-4 scalable to lossless audio coding, which considers Laplace distribution only.
Haiyan Shu, Susanto Rahardja
IEEE Trans. Speech Audio Process.3
2012 Mode-Dependent Transforms for Coding Directional Intra Prediction Residuals
abstract
The use of mode-dependent transforms for coding directional intra prediction residuals has been previously shown to provide coding gains, but the transform matrices have to be derived from training. In this paper, we derive a set of separable mode-dependent transforms by using a simple separable, directional, and anisotropic image correlation model. Our analysis shows that only one additional transform, the odd type-3 discrete sine transform (ODST-3), is required for the optimal implementation of mode-dependent transforms. In addition, the four-point ODST-3 also has a structure that can be exploited to reduce the operation count of the transform operation. Experimental results show that in terms of coding efficiency, our proposed approach matches or improves upon the performance of a mode-dependent transforms approach that uses transform matrices obtained through training.
Chuohao Yeo, Yih Han Tan, Zhengguo Li, Susanto Rahardja
IEEE Trans. Circuits Syst. Video Technol.4
2012 Detail-Enhanced Exposure Fusion
abstract
In a typical processing chain of image enhancement, an exposure fusion scheme can be used to synthesize a more detailed low dynamic range (LDR) image directly from a set of differently exposed LDR images, without generation of an intermediate high dynamic range image. In this brief, we introduce a new quadratic optimization-based method to extract fine details from a vector field. The new method extracts fine details from a set of differently exposed LDR images simultaneously. The extracted fine details are then added to an intermediate LDR image which is fused by simply using an existing exposure fusion scheme. With this, the proposed scheme can enhance fine details to produce sharper images.
Zhengguo Li, Jinghong Zheng 0001, Susanto Rahardja
IEEE Trans. Image Process.3
2012 Alternating Direction Method for Balanced Image Restoration
abstract
This paper presents an efficient algorithm for solving a balanced regularization problem in the frame-based image restoration. The balanced regularization is usually formulated as a minimization problem, involving an l(2) data-fidelity term, an l(1) regularizer on sparsity of frame coefficients, and a penalty on distance of sparse frame coefficients to the range of the frame operator. In image restoration, the balanced regularization approach bridges the synthesis-based and analysis-based approaches, and balances the fidelity, sparsity, and smoothness of the solution. Our proposed algorithm for solving the balanced optimal problem is based on a variable splitting strategy and the classical alternating direction method. This paper shows that the proposed algorithm is fast and efficient in solving the standard image restoration with balanced regularization. More precisely, a regularized version of the Hessian matrix of the l(2) data-fidelity term is involved, and by exploiting the related fast tight Parseval frame and the special structures of the observation matrices, the regularized Hessian matrix can perform quite efficiently for the frame-based standard image restoration applications, such as circular deconvolution in image deblurring and missing samples in image inpainting. Numerical simulations illustrate the efficiency of our proposed algorithm in the frame-based image restoration with balanced regularization.
Shoulie Xie, Susanto Rahardja
IEEE Trans. Image Process.2
2011 Structural similarity indices for high dynamic range imaging
abstract
In this paper, a structural similarity index is first proposed for two images with possibly different dynamic ranges and intensities as well as possibly small rotation and translation. The proposed index is then extended by dividing two images into local windows, and the similarity is detected by checking all pairs of local windows. It is shown by experimental results that the proposed indices are more robust to large intensity and dynamic range changes of two images from the same scene than the structural similarity (SSIM) index in.
Zhengguo Li, Susanto Rahardja
ICASSP3
2011 Quadratic optimization based small scale details extraction
abstract
In many image processing problems, it is required to extract small scale details from an image or a set of images. In this paper, we introduce a new framework for extracting small scale details from a single input image or a set of input images. We then show how to apply the framework to address several important problems in the field of image processing including tone mapping of high dynamic range images, de-noising of a non-flash image with a pair of non flash and flash images, as well as details enhancement via multi light images and a single input image. Experimental results show that the proposed framework outperforms existing methods.
Zhengguo Li, Jinghong Zheng 0001, Chuohao Yeo, Susanto Rahardja
ICASSP4
2011 Fast movement detection for high dynamic range imaging
abstract
When a high dynamic range image is synthesized by using a set of differently exposed low dynamic range (LDR) images, it is important to detect moving objects so as to remove ghosting from the final HDR image. A pixel level movement detection scheme was recently proposed in [8]. It included a pixel level similarity index for differently exposed LDR images, an adaptive threshold for the classification of pixels and an approach that utilizes intensity mapping (IMF) function for patching invalid regions. In this paper, we first propose a new adaptive threshold and a new patching approach to improve the scheme in [8]. Then, a sub-sampling method is introduced to simplify the improved movement detection scheme. Experimental results show that the improved movement detection scheme indeed outperforms the scheme in [8]. In addition, the speed is significantly improved by the proposed fast movement detection scheme.
Zhengguo Li, Susanto Rahardja
ICIP3
2011 Intra-prediction with adaptive sub-sampling
abstract
Intra-prediction improves coding performance by reducing inter-pixel redundancy. However, to accommodate the use of block transforms, not all pixels can be predicted from reconstructed pixels that are located close to themselves. This causes prediction performance to suffer as pixel values further apart are less correlated. This paper presents additional intra-prediction modes designed with the goal of improving prediction performance. Experimental results show an average gain of about 2% in KTA when the new modes are incorporated in the current 8×8 prediction modes. Since the new coding modes (8×8) are designed with transform size smaller than coding block size, the modes can also be useful when the prediction unit is larger than the maximum transform size. The use of smaller transform sizes can potentially lead to reduction of decoder complexity and implementation costs.
Yih Han Tan, Chuohao Yeo, Zhengguo Li, Susanto Rahardja
ICIP4
2011 Chroma intra prediction using template matching with reconstructed luma components
abstract
Intra coding in the current H.264/AVC video coding standard achieves high compression efficiency, in part due to the highly effective intra prediction process that exploits spatial directional correlation. However, intra prediction of chroma components in YUV 4:2:0 videos uses a limited set of possible predictions available for coding of luma components. Furthermore, coding of chroma components proceeds somewhat independently of luma components, and ignores any possible correlation between them. In this paper, we show a way of using reconstructed luma pixels to help with intra prediction of chroma pixels. By making use of the reconstructed co-located luma block to perform template matching in the luma plane, we are able to use as predictors the co-located chroma blocks of the matched luma blocks. Simulations results indicate that the proposed approach is able to obtain up to 33% chroma bit-rate reduction and up to 8% overall bit-rate reduction over H.264/AVC.
Chuohao Yeo, Yih Han Tan, Zhengguo Li, Susanto Rahardja
ICIP4
2011 De-ghosting of HDR images with double-credit intensity mapping
abstract
Ghosting artifacts are usually caused by moving object when composing a high dynamic range image from multiple differently exposed conventional images. In this paper, a robust de-ghosting algorithm is proposed based on a double-credit intensity mapping function (IMF) and an adaptive threshold model derived from statistical training. The double-credit IMF is estimated using both pixel intensity distribution and spatial correlation. A statistical threshold model is trained from the image database, and the key parameters are determined on the fly with variance vector calculated during the IMF estimation to adapt to different scenarios. Optimal bidirectional comparison is used for further improves the detection accuracy. The experiments show the effectiveness of the proposed de-ghosting method.
Zhengguo Li, Susanto Rahardja, Pasi Fränti
ICIP3
2011 Mode-dependent fast separable KLT for block-based intra coding
abstract
In this paper, we derive separable KLTs for coding H.264/AVC intra prediction residuals, using a simple image correlation model. Our analysis shows that for some intra prediction modes, we can in fact just use the DCT for performing either the row-wise or column-wise transform. Furthermore, we also compute the KLT that should be used based on the image correlation model, which happens to have sinosuidal terms. The 4×4 transform also has a structure that can be exploited to reduce the operation count of the transform operation. In our simplified implementation of mode-dependent directional transforms (MDDT), we only need to make use of two matrices: the DCT and the derived KLT. Our experimental results show that in terms of coding efficiency, our proposed approach has similar performance when compared with MDDT. More importantly, compared to MDDT, our approach requires no training and has lower computational and storage costs.
Chuohao Yeo, Yih Han Tan, Zhengguo Li, Susanto Rahardja
ISCAS4
2011 Low-complexity priority based packet scheduling for streaming MPEG-4 SLS
abstract
In this paper, we propose a low-complexity priority based packet scheduling algorithm for streaming MPEG-4 Scalable to Lossless (SLS) encoded audio. In the proposed system, the SLS encoded frames are partitioned into data units of different quality layers, which are transmitted according to their quality contribution to the final decoded audio and their urgency relative to the playback progress. Experimental results show that the proposed scheduling algorithm has an even lower compared to traditional greedy algorithm for packet scheduling, while outperforms them by a significant margin in for terms of quality of the streamed audio.
Rongshan Yu, Dajun Wu, Susanto Rahardja
MMSP4
2011 A Split and Merge Based Ellipse Detector With Self-Correcting Capability
abstract
A novel ellipse detector based upon edge following is proposed in this paper. The detector models edge connectivity by line segments and exploits these line segments to construct a set of elliptical-arcs. Disconnected elliptical-arcs which describe the same ellipse are identified and grouped together by incrementally finding optimal pairings of elliptical-arcs. We extract hypothetical ellipses of an image by fitting an ellipse to the elliptical-arcs of each group. Finally, a feedback loop is developed to sieve out low confidence hypothetical ellipses and to regenerate a better set of hypothetical ellipses. In this aspect, the proposed algorithm performs self-correction and homes in on "difficult" ellipses. Detailed evaluation on synthetic images shows that the algorithm outperforms existing methods substantially in terms of recall and precision scores under the scenarios of image cluttering, salt-and-pepper noise and partial occlusion. Additionally, we apply the detector on a set of challenging real-world images. Successful detection of ellipses present in these images is demonstrated. We are not aware of any other work that can detect ellipses from such difficult images. Therefore, this work presents a significant contribution towards ellipse detection.
Alex Yong Sang Chia, Susanto Rahardja, Deepu Rajan, Maylor Karhang Leung
IEEE Trans. Image Process.2
2010 Object recognition by discriminative combinations of line segments and ellipses
abstract
We present a contour based approach to object recognition in real-world images. Contours are represented by generic shape primitives of line segments and ellipses. These primitives offer substantial flexibility to model complex shapes. We pair connected primitives as shape tokens, and learn category specific combinations of shape tokens. We do not restrict combinations to have a fixed number of tokens, but allow each combination to flexibly evolve to best represent a category. This, coupled with the generic nature of primitives, enables a variety of discriminative shape structures of a category to be learned. We compare our approach with related methods and state-of-the-art contour based approaches on two demanding datasets across 17 categories. Highly competitive results are obtained. In particular, on the challenging Weizmann horse dataset, we attain improved image classification and object detection results over the best contour based results published so far.
Alex Yong Sang Chia, Susanto Rahardja, Deepu Rajan, Maylor K. H. Leung
CVPR2
2010 Bit-plane arithmetic coding for Laplacian source
abstract
For bit-plane coding of random source, we establish a relationship between the probability density function of an arbitrary non-negative distribution and the bit-plane symbol probability. Applied to Laplacian source, the relationship gives a simple, closed form of bit-plane symbol probability, which is a superset of the probability assignment rule adopted in the MPEG-4 Audio Scalable Lossless Coding (SLS) international standard. For bit-plane arithmetic coding of Laplacian source, we propose an algorithm and apply it to MPEG-4 SLS. Experimental results show that without changing the encoder/decoder complexity, the proposed algorithm consistently improved the compression ratio over SLS for all the 51 MPEG test sequences. The overall improvement in the compression ratio is 0.11%.
Haiyan Shu, Susanto Rahardja
ICASSP3
2010 Robust generation of high dynamic range images
abstract
A robust scheme is proposed to generate an anti-ghosting high dynamic range (HDR) image from a set of low dynamic range (LDR) images with different exposure times. Three major contributions of this paper are 1) a bi-directional prediction method; 2) an adaptive threshold for the classification of pixels; 3) Bayes estimator based methods for the on-line updating of predicted values and the synthesis of pixels to fill in the regions of moving objects to preserve their dynamic ranges. The proposed scheme is suitable for both static and dynamic scenes.
Zhengguo Li, Shoulie Xie, Shiqian Wu, Susanto Rahardja
ICASSP5
2010 Enhanced scalable to lossless audio coding scheme
abstract
Scalable to lossless (SLS) audio coding is a state-of-art audio coding technique that has been adopted as MPEG scalable audio coding tool. To realize bit-plane refinement, this technique employs bit-plane arithmetic coding for lossless entropy coding, and Laplacian distribution is used to model the input data to realize high compression efficiency. In this paper, bit-plane probability is analyzed when generalized Gaussian distribution is used to model the input data. Based on the result of bit-plane probability for generalized Gaussian distribution, a low cost bit-plane arithmetic coding method is presented. This scheme is implemented in the SLS audio coding platform. With the same computational complexity, the proposed algorithm presents higher compression efficiency than SLS.
Haiyan Shu, Susanto Rahardja
ICASSP4
2010 Non-cooperative optimization of wireless video encoders
abstract
With an empirically-derived complexity-rate-distortion model of a complexity scalable video encoder, we study the distributed optimization of several encoders sharing a wireless network, using the PSNR of the decoded video as the utility function. We will demonstrate how universal frequency reuse and source rate control can help the system degrade gracefully with increasing users, show the value of a simple pricing mechanism on transmission power in improving performance and study the effect of power allocation between transmitter and source encoder.
Yih Han Tan, Wei Siong Lee, Jo Yew Tham, Susanto Rahardja
ICASSP4
2010 Half-quadratic regularization based de-noising for high dynamic range image synthesis
abstract
It is possible to synthesis a high dynamic range (HDR) image by combining differently exposed low dynamic range (LDR) images of the same scene into one single image. For an HDR scene under low light condition, captured LDR images tend to be noisy, and the noise could usually be amplified during the synthesis process, causing severe degradation of the image quality. In order to reduce the noise during the HDR synthesis process, a de-noising scheme based on half-quadratic regularization is proposed for the synthesis of HDR images in this paper. By taking into consideration of two unique HDR image features, the proposed scheme effectively removes the noise, while the edges still being well preserved in the synthesized HDR images.
Wei Yao 0001, Zhengguo Li, Susanto Rahardja, Susu Yao, Jinghong Zheng 0001
ICASSP3
2010 Collaborative image processing algorithm for detail refinement and enhancement via multi-light images
abstract
In this paper, we introduce a new collaborative image processing algorithm to enhance contours and surface details of a scene through combination of multi-light images, which capture the same scene with fixed view point but different lighting conditions. Firstly, a detail layer that contains all contents of the multi-light images is generated through a gradient domain method and a quadratic filter. A new shadow detection algorithm is introduced to remove the artifacts from the detail layer. Subsequently, a base layer is constructed by using one input image. To further increase the visibility of details in dark areas, a simple tool is presented to brighten all dark areas of the base layer. Finally, the detail layer is multiplied to the base layer to synthesize the enhanced image that contains the desirable details. This enhanced image can present an elaborate description of the real scene. The proposed scheme also gives some interactivities that allow users to easily change the appearance of the enhanced image according to their preferences.
Jinghong Zheng 0001, Zhengguo Li, Susanto Rahardja, Susu Yao, Wei Yao 0001
ICASSP3
2010 Movement detection for the synthesis of high dynamic range images
abstract
In this paper, we propose an intensity mapping function (IMF) based scheme to detect moving objects in a set of low dynamic range (LDR) images with different known exposure times. The objective is to remove ghosting artifacts from the eventual high dynamic range (HDR) image. Our contributions include a bidirectional similarity detection method, an adaptive threshold for movement detection, and an IMF based method for the synthesis of pixels to fill in the regions of moving objects. Experimental results show that the proposed scheme outperforms existing schemes.
Zhengguo Li, Susanto Rahardja, Shoulie Xie, Shiqian Wu
ICIP2
2010 A robust and fast anti-ghosting algorithm for high dynamic range imaging
abstract
This paper presents a robust and fast algorithm for automatically generating high dynamic range (HDR) images in presence of camera movement and moving objects. This scheme comprises five modules: 1) image alignment, 2) estimation of camera response function (CRF) in dynamic scenes, 3) moving object detection, 4) progressive image correction, and 5) construction of HDR images. The key advantage of the algorithm is the ability to generate HDR images without ghost artifact. The proposed algorithm is fast as it is a one-shot solution without iterative computation and post-processing or even manual operation. Experimental results demonstrate that the proposed method outperforms the existing commercial products.
Shiqian Wu, Shoulie Xie, Susanto Rahardja, Zhengguo Li
ICIP3
2010 Detecting and composing near-identical HDR images without exposure information
abstract
In high dynamic range (HDR) imaging, two essential problems are to compose HDR image from conventional image set without any prior information about their exposures, and to access the synthesis result. To solve these problems, we first develop an exposure ratio estimation algorithm based on intensity mapping function (IMF). Then, we introduce an HDR image comparison method to verify whether two HDR images are from the same scene by using their log histogram similarity. Even though the images carrying the same information, their similarity cannot be detected by pixel-wise comparisons. We name such a pair of HDR images as near-identical images. According to experiments, our detection method is able to identify near-identical HDR images effectively, and our synthesis algorithm is able to recover the correct exposure ratios and compose near-identical HDR images.
Susanto Rahardja, Zhengguo Li, Pasi Fränti
ICIP2
2010 Beamforming performance of circular microphone array on spherical platform near bottom boundary
abstract
Spherical microphone arrays have been extensively studied for multimedia applications by both academic and industrial communities. General assumption of such study is based upon a planar wave traveling in free space model. However, in practical applications, acoustic wave reflections from boundary of a confined environment such as rooms or nearby furniture's carrying the array may degrade array performance significantly. In this paper, we present an approach for beamforming and direction of arrival (DOA) estimation using circular microphone array mounted on the equator of a sphere near a bottom boundary where the bottom reflection is dominant over other reflection, which corresponds to the applications such as spherical microphone array set on table or floor. We examine impact of reverberation to the beamforming performance and propose an algorithm to improve the microphone array performance. We first introduce an acoustic spherical scattering model to provide a theoretical background. The boundary reflection model is subsequently developed based on two-path ray theory. We further propose an approach using corrected steering vector to improve the beamforming performance to compensate mismatch of plane wave assumption in free space that the normal beam-forming is based upon. The effectiveness of the proposed approach is evaluated by numerical simulations.
Susanto Rahardja, Rongshan Yu
ICME2
2010 Audio onset detection using energy-based and pitch-based processing
abstract
Leveraging on the strength of energy-based processing for transient detection and pitch-based processing for softer onsets detection, we present a system that combines both energy and pitch cues for detecting onsets from different instrument categories. Given an audio input from an arbitrary instrument category, the system performs preliminary onset categorization based on the general note characteristics and then performs onsets integration based on the categorization. In addition, the proposed pitch processing technique explores musically relevant features extracted from the chromagram, which are robust for detecting pitch changes. The proposed system showed good detection performance on the MIREX audio onset detection dataset.
Hui Li Tan, Yongwei Zhu, Lekha Chaisorn, Susanto Rahardja
ISCAS4
2010 Complexity Scalable H.264/AVC Encoding
abstract
The H.264/AVC video coding standard encapsulates the most advanced video coding tools. Since the various techniques that lead to better coding efficiency of the coding standard also inevitably increase the complexity of the video encoder, real-time H.264 encoding of video streams is a challenging task. If available computational resource does not allow the entire encoding process to be carried out in time, a complexity scalable technique that ensures a graceful degradation of coding performance will be a valuable tool. We designed a video encoding scheme that allows the rate distortion (R-D) process to be carried out in a complexity scalable fashion. Our proposed singularly parameterized complexity scalable scheme allows the control of complexity-coding performance tradeoff when available resources are limited and the optimal R-D performance is unattainable.
Yih Han Tan, Wei Siong Lee, Jo Yew Tham, Susanto Rahardja, Kin Mun Lye
IEEE Trans. Circuits Syst. Video Technol.4
2009 MPEG-4 scalable lossless audio transparent bitrate and its application
abstract
In this paper, the relevance between the bit-plane levels and the perceptual quality of the MPEG-4 scalable lossless audio is explored. It is observed that only the top 3 bit-planes in SLS with AAC core of 64kbps are closely related to the transparent quality. The result is used in the application of the data hiding process. The proposed hiding method has low complexity as no side information is required at the data extraction process. A data hiding capacity of 98kbps is achieved for SLS lossless bitstream.
Susanto Rahardja
ICASSP2
2009 Speaker diarization in meeting audio
abstract
This paper describes speaker diarization system on a NIST Rich Transcription 2007 (RT-07) meeting recognition evaluation data set for the task of multiple distant microphone (MDM). Our implementation includes three components: initial clustering, non-speech removal and cluster purification. Initial clusters are generated using directional of arrival (DOA) information and bootstrap clustering. Multiple GMM modeling for speech/non-speech classification is employed for non-speech removal component. In addition, a novel system fusion strategy using information from receiver operating curve (ROC) is proposed for non-speech removal component. Finally, consensus clustering approach together with iterative GMM clustering method is employed for speaker cluster purification. The system achieves the overall DER of 10.81%.
Tin Lay Nwe, Hanwu Sun, Haizhou Li 0001, Susanto Rahardja
ICASSP4
2009 Video quality monitoring of streamed videos
abstract
This paper describes a video quality analysis system for inservice monitoring of streamed videos, particularly over mobile/wireless networks. The algorithm adopts the no-reference method, and enables real-time measurement of video quality at any point in the content production and delivery chain using any given video. The technologies developed include no-reference methods for measuring picture freeze, picture loss, and blockiness. The developed system (where the software has not been optimized for speed) is able to process video of CIF size (352×288 pixels) at more than 30 fps on a Pentium-IV 3GHz computer. The experimental results show that the proposed video quality analysis system gives good accuracy for picture freeze, picture loss, and blocking detections.
Ee Ping Ong, Shiqian Wu, Mei Hwan Loke, Susanto Rahardja, Jason Tay, Cheng Kok Tan
ICASSP4
2009 Complexity-Rate-Distortion Optimization for Real-Time H.264/AVC Encoding
abstract
In this paper, we introduce a singularly-parameterized complexity scalable H.264-compliant encoder. Through modeling its complexity-rate-distortion relationships, we derive optimized operating mode of the encoder (rate and complexity) and show through experiment that such optimization can help a video encoder operate within rate and time constraints. The design of the complexity scalable encoding scheme enables the encoder to perform optimization while taking into consideration the availability of computational resource. This extension of traditional rate-distortion optimization is necessary when time or power constraints do not allow a video encoder to achieve rate-distortion optimized coding performance. Our optimization scheme outputs parameters that allow the encoder to be as close to being rate-distortion optimized as possible, within rate and complexity constraints.
Yih Han Tan, Wei Siong Lee, Jo Yew Tham, Susanto Rahardja
ICCCN4
2009 Tennis Space: An Interactive and Immersive Environment for Tennis Simulation
abstract
This paper reports the design and implementation of an interactive and immersive environment (IIE) for tennis simulation. The presented design layout, named Tennis Space, provides the necessary immersive experience without overly restricting the player. To address the instability problem of real-time tracking of fast moving objects, a hybrid tracking solution integrating optical tracking and ultrasound-inertial tracking technologies is developed. An L-shaped IIE has been implemented for tennis simulation and has received positive feedback from users.
Shuhong Xu, Peng Song 0009, Ching-Ling Chin, Gim Guan Chua, Zhiyong Huang 0001, Susanto Rahardja
ICIG6
2009 High dynamic range compression by half quadratic regularization
abstract
This paper presents a new adaptive tone mapping for high dynamic range (HDR) images. In the proposed scheme, the luminance of an HDR image is decomposed into a base layer with large gradients and a detail layer with small gradients by using an adaptive half quadratic regularization method. The base layer is compressed by a novel global mapping to reduce the dynamic range while the detail layer can be amplified to enhance the local contrasts. With the proposed scheme, the generated low dynamic range (LDR) images look more natural and local details are also preserved very well.
Zhengguo Li, Susanto Rahardja, Susu Yao, Jinghong Zheng 0001, Wei Yao 0001
ICIP2
2009 Realistic HDR tone-mapping based on contrast perception matching
abstract
In this paper, a new method to measure human's reaction time on local contrast is introduced. It is believed the reaction time is highly correlated with the response strength on local contrasts in human visual system. Using the proposed method, a large set of subjective viewing data were then collected, and a full-range (from near-threshold to suprathreshold) contrast perception model on low dynamic range display is built. Based on the proposed model and Plainis' reaction time model [17], a new realistic tone-mapping operator, which minimizes the perception difference of local contrasts between HDR and LDR images, is proposed. Experimental results confirmed the performance of the proposed operator.
Zhongkang Lu, Susanto Rahardja
ICIP2
2009 Complexity scalable rate-distortion optimization for H.264/AVC
abstract
The H.264/AVC video coding standard encapsulates the most advanced video coding tools. The various techniques that lead to better coding efficiency of the coding standard also inevitably increase the complexity of the video encoder. Thus, real-time encoding of video streams with H.264 coding standard is a challenging task. If available computational resource does not allow the entire encoding process to be carried out in time, a complexity scalable technique that ensures a graceful degradation of coding performance will be a useful tool. This work proposes a singularly-parameterized complexity scalable rate-distortion framework for H.264/AVC encoders.
Yih Han Tan, Wei Siong Lee, Jo Yew Tham, Susanto Rahardja
ICIP4
2009 Rhythm analysis for personal and social music applications using drum loop patterns
abstract
The development of effective music retrieval and recommendation applications requires meaningful features for characterizing music audio. The rhythm of music audio in particular, besides timbre and melody, is essential in describing the music piece. This paper illustrates how drum loop patterns characterize the salient rhythm structure of music and an approach for drum loop pattern extraction is described. The extracted drum loop patterns are formed based on the temporal regularity (accent on meter) and drum types (bass and snare) of the beats and are good representations of the rhythmic dimension of music. The extracted drum loop patterns are useful for music segmentation and further work would include music classification and similarity matching.
Hui Li Tan, Yongwei Zhu, Susanto Rahardja, Lekha Chaisorn
ICME3
2009 Real-time H.264 encoder implementation on a low-power digital signal processor
abstract
This paper presents a real-time H.264/AVC baseline profile video encoder. The encoder hardware is implemented using a cost-effective, low-power ADI Blackin-561 DSP and related peripherals for real-time video capturing, coding and streaming. The encoder software is developed using a two-stage pipelining framework for efficient parallel video data encoding. A synchronization mechanism with shared memory semaphores is used to schedule the hardware processes and software procedures to acquire the real-time encoding performance. Techniques for reducing the execution time in both stages are also described in this paper. Performance evaluation results verified that the encoder is capable of performing real-time encoding of CIF-resolution and medium-motion VGA-resolution videos, while maintaining good video quality.
Ming-Jiang Yang, Jo Yew Tham, Susanto Rahardja, Dajun Wu
ICME3
2009 Drum loop pattern extraction from polyphonic music audio
abstract
Although drum loops are widely present in many audio recordings of modern style music, there is little research that deals with automatic extraction of drum loops in polyphonic music audio. This paper presents an approach for drum loop pattern extraction, based on a technique of fusing the meter estimation information and onset clustering information. The extracted drum loop patterns are formed based on the temporal regularity (accent on meter) and drum types (bass and snare) of the beats. Despite not achieving very high-precision drum loop transcription, the extracted drum loop pattern characterizes the salient rhythm structure of the music, and can be very useful for categorization or similarity matching of music. Experimental results have shown the effectiveness of the proposed drum loop pattern extraction method.
Yongwei Zhu, Hui Li Tan, Susanto Rahardja
ICME3
2009 The Misadjustment of the Cascaded LMS Prediction Filter
abstract
In this paper, we use a stochastic fixed-point theorem to study the stochastic convergence properties (in mean-squares sense) of the cascaded LMS predictor including conditions on the stepsize for the adaptive algorithm convergence and the misadjustment. An analytic expression for the misadjustment is derived for Gaussian statistical signals and shown to be exponentially dependent on the number of stages in the cascade structure, which is higher than the misadjustment of the conventional LMS filter. It can be observed that a higher misadjusment can be expected if the input signal is extremely uncorrelated.
Dong-Yan Huang, Susanto Rahardja
ISCAS2
2009 Bit-Plane Coding for Source with Generalized Gaussian Distribution
abstract
Bit-plane probability of data is a useful characteristic for many applications, such as entropy coding. Its value may vary according to the distribution of the data. In this paper, the bit-plane probability of data with generalized Gaussian distribution is analyzed. By studying the statistical characteristics of the sources and the bit-plane probability for Laplacian distribution, an approximated analytic description of bit-plane probability for generalized Gaussian distribution is presented. This approximation calculates the bit-plane probability with low computational cost. Simulation results show that, the proposed algorithm can well approximate the actual bit-plane probability. In addition, when the proposed algorithm is applied to audio bit-plane arithmetic coding, the coding efficiency is higher than that based on the Laplacian distribution assumption.
Haiyan Shu, Susanto Rahardja
ISM4
2009 Biologically inspired algorithm for enhancement of speech intelligibility over telephone channel
abstract
This paper describes a method to increase speech intelligibility when the speech signal is being transmitted over telephone lines. In order to detect all factors which affect speech intelligibility, we use telephone simulation tool in ITUT Software Tools Library release 2005 (STL2005) to identify the most problematic telephone-channel deteriorations. Of the various effects considered, additive noise and bandwidth of telephone channel seem to have the greatest influence on human's recognition accuracy. To reduce this degradation, a Lombard production and perception effect model, which represents the spectral changes of speech signal uttered and perceived in noisy environment, is proposed. This model is realized by a threestage modification procedure involving time-scale, pitch-scale, the formants and the consonant-vowel energy ratio modifications of the normal speech signals. From listening test conducted, it is shown that the proposed compensation algorithm increases the intelligibility of speech in telephone channel environments at 0 dB Signal-to-Noise ratio (SNR) and -10 dB SNR by nearly 5.5% and 13% over the unprocessed signals respectively.
Dong-Yan Huang, Susanto Rahardja, Ee Ping Ong
MMSP2
2009 HVS based histogram adjustment for tone mapping
abstract
Global tone mapping operators (TMOs) are better in preserving the naturalness and relative illumination of the HDR scene compared to the local TMOs. However, they cannot preserve the details like the local TMOs. We propose a human visual system (HVS) based histogram method for global tone mapping, which can better preserve the details than the existing global TMOs. Our method can be incorporated to all histogram based global methods. We compare our results with traditional histogram adjustment by Ward et al. [Larson97] and photographic TMO [Reinhard2002], and in most of the cases our method is better in preserving the details of the input HDR images.
Ishtiaq Rasool Khan, Zhiyong Huang 0001, Farzam Farbiz, Corey Manders, Susanto Rahardja
SIGGRAPH ASIA Sketches5
2009 Natural and Seamless Image Composition With Color Control
abstract
While the state-of-the-art image composition algorithms subtly handle the object boundary to achieve seamless image copy-and-paste, it is observed that they are unable to preserve the color fidelity of the source object, often require quite an amount of user interactions, and often fail to achieve realism when there exists salient discrepancy between the background textures in the source and destination images. These observations motivate our research towards color controlled natural and seamless image composition with least user interactions. In particular, based on the Poisson image editing framework, we first propose a variational model that considers both the gradient constraint and the color fidelity. The proposed model allows users to control the coloring effect caused by gradient domain fusion. Second, to have less user interactions, we propose a distance-enhanced random walks algorithm, through which we avoid the necessity of accurate image segmentation while still able to highlight the foreground object. Third, we propose a multiresolution framework to perform image compositions at different subbands so as to separate the texture and color components to simultaneously achieve smooth texture transition and desired color control. The experimental results demonstrate that our proposed framework achieves better and more realistic results for images with salient background color or texture differences, while providing comparable results as the state-of-the-art algorithms for images without the need of preserving the object color fidelity and without significant background texture discrepancy.
Jianmin Zheng, Jianfei Cai 0001, Susanto Rahardja, Chang Wen Chen
IEEE Trans. Image Process.4
2009 Structural Descriptors for Category Level Object Detection
abstract
We propose a new class of descriptors which exhibits the ability to yield meaningful structural descriptions of objects. These descriptors are constructed from two types of image primitives: quadrangles and ellipses. The primitives are extracted from an image based on human cognitive psychology and model local parts of objects. Experiments reveal that these primitives densely cover objects in images. In this regard, structural information of an object can be comprehensively described by these primitives. It is found that a combination of simple spatial relationships between primitives plus a small set of geometrical attributes provide rich and accurate local structural descriptions of objects. Category level object detection of four-legged animals, bicycles, and cars images is demonstrated under scaling, moderate viewpoint variations, and background clutter. Promising results are achieved.
Alex Yong Sang Chia, Susanto Rahardja, Deepu Rajan, Maylor K. H. Leung
IEEE Trans. Multim.2
2009 Fixed Quality Layered Audio Based on Scalable Lossless Coding
abstract
The paper addresses a bitstream scalable coder based on the MPEG-4 scalable lossless (SLS) coding system where, in contrast to SLS, the bitrate of the enhancement layer is not fixed but instead an attempt is made to create a quality-fixed enhancement layer. With a PCM audio input, the proposed structure is able to produce an audio version with near-transparent quality on top of the existing low-quality version. In particular, the proposed fixed quality enhancing process with checking procedures is able to provide the minimum amount of enhancement for the low-quality version to obtain a near-transparent quality that is almost indistinguishable from the CD quality. In addition, a bitrate estimation model is proposed. The model enables the direct estimation of the enhancing bitrate from two parameters extracted from the encoding process of the low-quality version. Evaluation results indicate that a better defined quality level is guaranteed compared to a fixed bitrate setting and that in the mean a lower (approximately 20%) bitrate is attained. It is also shown that the estimation model proposed is able to accurately predict the necessary enhancing bitrate and at the same time, reduce the complexity by around 17%.
Susanto Rahardja, Soo Ngee Koh
IEEE Trans. Multim.2
2008 An improved genetic algorithm for aperiodic array synthesis
abstract
In this paper, a novel algorithm on beam pattern synthesis for linear aperiodic arrays with arbitrary geometrical configuration is proposed. The algorithm is based on an improved genetic algorithm (IGA) that simultaneously adjusts the weight coefficients and inter-sensor spacings of a linear aperiodic array. A novel section-based crossover and a self-supervised mutation process are developed to improve the convergence performance. The results from simulation illustrate that with the IGA, the peak sidelobe level (PSL) of the synthesized beam pattern has been successfully lowered. In addition, the computational cost of the proposed algorithm can be as low as being about 10% of that of a recently reported genetic algorithm based synthesis method. The robustness of the proposed IGA has been illustrated clearly from the statistic of multiple independent runs too. The excellent performance of the IGA makes it a promising optimization algorithm where expensive cost functions are involved.
Ling Cen, Wee Ser, Zhu Liang Yu, Susanto Rahardja
ICASSP4
2008 A fully scalable audio coding structure with embedded psychoacoustic model
abstract
A fully scalable audio coding structure based on a novel combination of the non-core MPEG-4 scalable lossless audio coding (SLS), the state-of-the-art psychoacoustic model, joint stereo coding and the perceptually prioritized bit-plane coding is presented in this paper. The psychoacoustic information is implicitly embedded in the scalable bitstream with negligible amount of side information and trivial modification to the standardized SLS decoder. Results of extensive evaluation show that the subjective quality of scalable audio is improved significantly.
Susanto Rahardja, Soo Ngee Koh
ICASSP2
2008 A New Integrated Spatio-Temporal Framework for Video Error Concealment
abstract
Temporal error concealment (TEC) algorithms do not produce good results in the presence of cut scenes, camera panning or object occlusion in videos due to the lack of a suitable reference frame. Under such circumstances, spatial EC (SEC) may perform better. Thus, EC algorithms should decide adaptively between TEC or SEC. In this paper, we propose an integrated spatio-temporal EC framework which extends our previous work in TEC and SEC. The proposed algorithm decides adaptively on the best EC mode and produces average PSNR improvements of up to 4.59 dB over a coding-mode based EC algorithm. Most importantly, our algorithm produces results of much better perceptual quality.
Ee Sin Ng, Jo Yew Tham, Susanto Rahardja
ICCCN3
2008 A split and merge based ellipse detector
abstract
We present an ellipse detector that continually pools lower level information of the edge pixels together to achieve robust detection of the ellipses present in the image. In addition, the parameters of the detected ellipses are continually refined using a close loop system driven by Gestalt psychology. We highlight that we do not rely on the geometrical properties of the ellipses to detect the ellipses. In this aspect, our algorithm is well suited to detect partially occluded ellipses in the image. Experiments on real and synthetic images demonstrate the robustness of our algorithm in which both complete and incomplete ellipses can be detected. In particular, experimental results show that the mean detection accuracy of our algorithm surpasses 92% even with around 90% outliers in the images. This detection performance is superior to that achieved by the robust regression, least squares and the hough transform based ellipse detectors.
Alex Yong Sang Chia, Deepu Rajan, Maylor K. H. Leung, Susanto Rahardja
ICIP4
2008 Speech enhancement for telephony name speech recognition
abstract
In this paper, we investigate the contribution of the speech enhancement to telephony name speech recognition system that has been of the feature enhancement in noisy environment. Since the masking based subband Kalman filtering (MSKF) method has been shown to have good performance over many existing speech enhancement methods for human listening, we especially examine the MSKF speech enhancement method for the speech recognition system. We evaluate the performance of the speech enhancement methods based on the recorded telephony name speech database. Experimental results show that the recognition error rate of the speech recognition system can be reduced much by using MSKF.
Chang Huai You, Susanto Rahardja, Haizhou Li 0001
ICME2
2008 Re-examination of applying wavelet based progressive image coder for 3D semi-regular mesh compression
abstract
The latest wavelet based 3D mesh coding schemes convert an irregular mesh into a semi-regular mesh and directly apply the zerotree-like image coders to compress the wavelet vectors generated in the remeshing process. The major problem of such type of approaches is that the particular properties of semi-regular meshes are not being considered in the zerotree-like image coders. In this paper, we propose an improved wavelet based 3D mesh coder. The basic idea is to introduce a preprocessing step to scale up the vector wavelets generated in remeshing so that the inherent dependency of wavelets can be truly understood by the zerotree-like image compression algorithms. The weights used in the scaling process are carefully designed through thoroughly analyzing the distortions of wavelets at different refinement levels. Experimental results show that our proposed mesh coder significantly outperforms the state-of-the-art wavelet based 3D mesh compression scheme.
Juyong Zhang, Jianfei Cai 0001, Jianmin Zheng, Susanto Rahardja
ICME5
2008 Early detection of all-zero block in H.264 with new rate-quantization models
abstract
This paper presents a new algorithm for detecting all-zero DCT coefficient blocks (AZB) prior to DCT and quantization for H.264. The early detection criterion is derived from a new rate-quantization model which is established by considering the unique features of quantization process in H.264. The proposed algorithm aims to eliminate the redundant computations in AZB and consequently speed up the encoding process of H.264 video codec. Simulation results show that the proposed algorithm achieves a high detection ratio up to 95.88% while maintain a very low false detection ratio. The results confirm that a more effective AZB detection for H.264 can be achieved by taking the unique feature of quantization into consideration.
Wei Yao 0001, Zhengguo Li, Susanto Rahardja
ISCAS3
2008 Efficient stereo bitrate allocation for fully scalable audio codec
abstract
The bit allocation algorithm for stereo channels in MPEG-4 scalable lossless coding (SLS) is not optimized. A perceptually enhanced stereo bit allocation algorithm for fully scalable audio coding is presented in this paper. According to the energy distribution in different channels, the bitrate is allocated in a much more efficient manner. Experiment results show that the proposed method significantly improves the perceptual quality of the fully scalable audio at various bitrates without introducing any new side information.
Susanto Rahardja, Soo Ngee Koh
MMSP2
2008 Convergence Performance of the Cascaded RLS-LMS Prediction
abstract
In this paper, we use a stochastic fixed-point theorem to analyze the stochastic convergence properties of the cascaded RLS-LMS prediction filter in terms of conditions of convergence and the misadjustment. It is shown that the cascaded RLS-LMS prediction filter converges to almost the same optimal solution of the conventional RLS filter. The misadjustment is shown to be exponentially dependent on the number of stages in the cascade structure and is higher than the misadjustment of the conventional RLS filter. However, the cascaded RLS-LMS prediction filter allows us to build up a low complexity RLS-like predictor with time-varying learning rate, which may be useful in uncertain and non-stationary environments.
Dong-Yan Huang, Susanto Rahardja
VTC Spring2
2008 Cascaded RLS-LMS Prediction in MPEG-4 Lossless Audio Coding
abstract
This paper describes the cascaded recursive least square-least mean square (RLS-LMS) prediction, which is part of the recently published MPEG-4 Audio Lossless Coding international standard. The predictor consists of cascaded stages of simple linear predictors, with the prediction error at the output of one stage passed to the next stage as the input signal. A linear combiner adds up the intermediate estimates at the output of each prediction stage to give a final estimate of the RLS-LMS predictor. In the RLS-LMS predictor, the first prediction stage is a simple first-order predictor with a fixed coefficient value 1. The second prediction stage uses the recursive least square algorithm to adaptively update the predictor coefficients. The subsequent prediction stages use the normalized least mean square algorithm to update the predictor coefficients. The coefficients of the linear combiner are then updated using the sign-sign least mean square algorithm. For stereo audio signals, the RLS-LMS predictor uses both intrachannel prediction and interchannel prediction, which results in a 3% improvement in compression ratio over using only the intrachannel prediction. Through extensive tests, the MPEG-4 Audio Lossless coder using the RLS-LMS predictor has demonstrated a compression ratio that is on par with the best lossless audio coders in the field. In this paper, the structure of the RLS-LMS predictor is described in detail, and the optimal predictor configuration is studied through various experiments.
Pasi Fränti, Dong-Yan Huang, Susanto Rahardja
IEEE Trans. Speech Audio Process.4
2008 Frequency Region-Based Prioritized Bit-Plane Coding for Scalable Audio
abstract
A perceptually enhanced prioritized bit-plane audio coding algorithm is presented in this paper. According to the energy distribution in different frequency regions, the bit-planes are prioritized with optimized parameters. Based on the statistical modeling of the frequency spectrum, a much more simplified implementation of prioritized bit-plane coding is integrated with the recent release of MPEG-4 scalable lossless (SLS) audio coding structure by replacing the sequential bit-plane coding in the enhancement layer. With zero extra side information, trivial added complexity, and modification to the original SLS structure, extensive experimental results show that the perceptual quality of SLS with noncore and very low core bit-rate is improved significantly in a wide range of bit-rate combinations. Fully scalable audio coding up to lossless with much enhanced perceptual quality is thus achieved.
Susanto Rahardja, Soo Ngee Koh
IEEE Trans. Speech Audio Process.2
2008 Simplified Motion-Refined Scheme for Fine-Granularity Scalability
abstract
In this paper, we introduce a low-complexity fine-granularity scalable (FGS) video encoder that refines both residue and motion information in the quality layers. The current scalable video coding (SVC) draft shows that significant gains can be achieved when each enhancement layer undergoes the motion estimation/motion compensation (ME/MC) process with its own motion vector field (MVF). However, given the high computational cost of ME/MC, a motion-refined FGS scheme can be expensive to implement. The proposed scheme controls the macroblock (MB) mode allowed in the base layer and channels computational resources to refine motion in enhancement layers. Through a proper selection of Lagrangian factor for the generation of the first MVF, it is possible to design a low-complexity FGS encoder that has good overall coding performance. A simplified motion-refinement scheme is also adopted for selected MBs in enhancement layers by exploiting the correlation of MB-type information between successive layers to further reduce the complexity of the FGS encoder. Meanwhile, a framework of rate-distortion-complexity optimization is proposed for the FGS by considering the interpolation complexity during the ME/MC in each layer. The FGS decoder can be simplified through the reduction in the number of interpolations.
Yih Han Tan, Zhengguo Li, Keng-Pang Lim, Susanto Rahardja
IEEE Trans. Circuits Syst. Video Technol.4
2008 Orthogonal Data Embedding for Binary Images in Morphological Transform Domain- A High-Capacity Approach
abstract
This paper proposes a data-hiding technique for binary images in morphological transform domain for authentication purpose. To achieve blind watermark extraction, it is difficult to use the detail coefficients directly as a location map to determine the data-hiding locations. Hence, we view flipping an edge pixel in binary images as shifting the edge location one pixel horizontally and vertically. Based on this observation, we propose an interlaced morphological binary wavelet transform to track the shifted edges, which thus facilitates blind watermark extraction and incorporation of cryptographic signature. Unlike existing block-based approach, in which the block size is constrained by 3times3 pixels or larger, we process an image in 2times2 pixel blocks. This allows flexibility in tracking the edges and also achieves low computational complexity. The two processing cases that flipping the candidates of one does not affect theflippabilityconditions of another are employed for orthogonal embedding, which renders more suitable candidates can be identified such that a larger capacity can be achieved. A novel effective Backward-Forward Minimization method is proposed, which considers both backwardly those neighboring processedembeddablecandidates and forwardly those unprocessedflippablecandidates that may be affected by flipping the current pixel. In this way, the total visual distortion can be minimized. Experimental results demonstrate the validity of our arguments.
Huijuan Yang, Alex Chichung Kot, Susanto Rahardja
IEEE Trans. Multim.3
2007 Steganalysis of Binary Cartoon Image using Distortion Measure
abstract
We present a steganalysis technique for data hiding in binary cartoon images. Due to the perturbation from the embedding, the contours of a stego binary cartoon image are distorted. When calculating the distortion of the image based on a distortion measure, the distortion score between a stego image and its de-noised version should be different from that between an original image and its de-noised version. Different binary image distortion measures are used to calculate the distortion scores, which are used as the features for classification. The sequential floating forward search (SFFS) method is used to search for the combination of the features that yields the best classification results.
Jun Cheng 0003, Alex Chichung Kot, Susanto Rahardja
ICASSP (2)3
2007 Autoregressive Parameter Estimation for Kalman Filtering Speech Enhancement
abstract
In this paper, autoregressive parameter estimation for Kalman filtering speech enhancement is studied. In conventional Kalman filtering speech enhancement, spectral subtraction is usually used for speech autoregressive (AR) parameter estimation. We propose log spectral amplitude (LSA) minimum mean-square error (MMSE) instead of spectral subtraction for the estimation of speech AR parameters. Based on an observation that full-band Kalman filtering speech enhancement often causes an unbalanced noise reduction between speech and non-speech segments, a spectral solution is proposed to overcome the unbalanced reduction of noise. This is done by shaping the spectral envelopes of the noise through likelihood ratio. Our simulation results show the effectiveness of the proposed method.
Chang Huai You, Susanto Rahardja, Soo Ngee Koh
ICASSP (4)2
2007 Ellipse Detection with Hough Transform in One Dimensional Parametric Space
abstract
The main advantage of using the Hough Transform to detect ellipses is its robustness against missing data points. However, the storage and computational requirements of the Hough Transform preclude practical applications. Although there are many modifications to the Hough Transform, these modifications still demand significant storage requirement. In this paper, we present a novel ellipse detection algorithm which retains the original advantages of the Hough Transform while minimizing the storage and computation complexity. More specifically, we use an accumulator that is only one dimensional. As such, our algorithm is more effective in terms of storage requirement. In addition, our algorithm can be easily parallelized to achieve good execution time. Experimental results on both synthetic and real images demonstrate the robustness and effectiveness of our algorithm in which both complete and incomplete ellipses can be extracted.
Alex Yong Sang Chia, Maylor K. H. Leung, How-Lung Eng, Susanto Rahardja
ICIP (5)4
2007 Edge Weighted Spatio-Temporal Search for Error Concealment
abstract
In temporal error concealment (EC), the sum of absolute difference (SAD) is commonly used to identify the best replacement macroblock. Even though the use of SAD ensures spatial continuity and produces visually good results, it is insufficient to ensure edge alignment. Other distortion criteria based solely on structural alignment may also perform poorly in the absence of strong edges. In this paper, we propose a spatio-temporal EC search algorithm using an edge weighted SAD distortion criterion. This distortion criterion ensures both edge alignment and spatial continuity. We assume the loss of motion information and use zero motion vector as the starting search point. We show that the proposed algorithm outperforms the use of unweighted SAD in general. Most importantly, the perceptual quality of EC is improved due to edge alignment while ensuring spatial continuity.
Ee Sin Ng, Jo Yew Tham, Susanto Rahardja
ICIP (4)3
2007 Adaptive Bit-Plane Scanning for Scalable Audio
abstract
MPEG-4 scalable lossless (SLS) coding is a unified solution for demands in high compression perceptual audio and high quality lossless audio. It provides a fine-grain scalable extension to the MPEG-4 advanced audio coding (AAC) perceptual audio coder up to fully lossless reconstruction. It is observed that the quality of SLS coded audio is still far from optimal when the core bitrate is low. With this observation, an adaptive bit-plane scanning (ABPS) method is proposed in this paper. In ABPS, the bit-plane scanning order is adaptive to the energy distribution of the signal to be coded. The results show that with ABPS, the perceptual quality of the audio under low core bitrate scenario is significantly improved without any extra payload.
Susanto Rahardja, Soo Ngee Koh
ICME2
2007 An Adaptive Deblocking Filter for ROI-Based Scalable Video Coding
abstract
The Region-of-interest (ROI) based video coding within Scalable Video Coding (SVC) can be implemented by making use of Type 2 Flexible Macroblock Ordering (FMO), which marks independent rectangle regions/slices inside a frame by their top-left and bottom-right coordinates. By employing the proposed scheme, a displaying frame can be separated into several independent regions that are assigned with different Signal-Noise-Ratio (SNR), Spatial and Temporal quality. The scheme can be used to ensure the achievement of high sub-jecitve quality or to fulfill some special functionalities. Owning to the fact that the frame is separated into independent regions, and the regions are assigned with big quality differences, false edge (blockiness) may appear around the ROI boundaries, which cannot be automatically removed by the in-loop filters. The annoyance of such artifact depends on the local visual context, thus a new adaptive deblocking filter is proposed in this paper. The filter includes two steps: first, the complexity of the blocks around ROI boundaries is measured; and different filtering modes and smoothing abilities are then selected accordingly to reduce the annoyance of the blockiness. Experimental results showed that coding quality is improved by the proposed filter in low bitrate Coarse Granular Scalability (CESB) video.
Zhongkang Lu, Jinghong Zheng 0001, Shiqian Wu, Weisi Lin, Susanto Rahardja
ICME5
2007 Wyner-Ziv Image Coding from Random Projections
abstract
In this paper, we present a Wyner-Ziv coding based on random projections for image compression with side information at the decoder. The proposed coder consists of random projections (RPs), nested scalar quantization (NSQ), and Slepian-Wolf coding (SWC). Most of natural images are compressible or sparse in the sense that they are well-approximated by a linear combination of a few coefficients taken from a known basis, e.g., FFT or Wavelet basis. Recent results show that it is surprisingly possible to reconstruct compressible signal to within very high accuracy from limited random projections by solving a simple convex optimization program. Nested quantization provides a practical scheme for lossy source coding with side information at the decoder to achieve further compression. SWC is lossless source coding with side information at the decoder. In this paper, ideal SWC is assumed, thus rates are conditional entropies of NSQ quantization indices. Recently theoretical analysis shows that for the quadratic Gaussian case and at high rate, NSQ with ideal SWC performs the same as conventional entropy-coded quantization with side information available at both the encoder and decoder. We note that the measurements of random projects for a natural large-size image can behave like Gaussian random variables because most of random measurement matrices behave like Gaussian ones if their sizes are large. Hence, by combining random projections with NSQ and SWC, the tradeoff between compression rate and distortion will be improved. Simulation results support the proposed joint codec design and demonstrate considerable performance of the proposed compression systems.
Shoulie Xie, Susanto Rahardja, Zhengguo Li
ICME2
2007 High Capacity Data Hiding for Binary Images in Morphological Wavelet Transform Domain
abstract
This paper investigates the problem of data hiding for binary images authentication in morphological binary wavelet transform domain. Directly using the detail coefficients as a location map to determine the data hiding locations is difficult to achieve blind watermark extraction. Hence, we view flipping an edge pixel as shifting the edge location one pixel horizontally and vertically. Based on this observation, we propose an interlaced transform to track the shifted edges. Different from existing block-based approach, in which the block size is constrained by not less than 3 x 3, we process the image in 2 x 2 blocks. The two processing cases that the flippantly condition of one is not affected by flipping the candidates of another are combined such that an extremely large capacity can be achieved, which slightly sacrifices the visual quality.
Huijuan Yang, Alex Chichung Kot, Susanto Rahardja, Xudong Jiang 0001
ICME3
2007 Using direction of arrival estimate and acoustic feature information in speaker diarization
Chin-Wei Eugene Koh, Hanwu Sun, Tin Lay Nwe, Trung Hieu Nguyen 0001, Bin Ma 0001, Chng Eng Siong, Haizhou Li 0001, Susanto Rahardja
INTERSPEECH8
2007 Balanced Inter-Layer Prediction for Combined Coarse Granular Scalability and Spatial Scalability
abstract
An inter-layer prediction scheme with two base layers and a new concept of auxiliary layer are introduced in this paper. The scheme achieves a good balance among all layers for the combined coarse granular scalability and spatial scalability. The objective is to improve the coding efficiency of layers with higher resolution. Meanwhile, a price based scheme is proposed to determine the auxiliary layer. With the proposed scheme, a new element of price can be integrated into scalable video coding which in turn justifies the necessity of scalable coding
Wei Yao 0001, Zhengguo Li, Susanto Rahardja
ISCAS3
2007 Switchable Bit-Plane Coding for High-Definition Advanced Audio Coding
Susanto Rahardja, Soo Ngee Koh
MMM (1)2
2007 Perceptual Enhancement for Fully Scalable Audio
abstract
MPEG-4 Scalable Lossless (SLS) coding is the latest released ISO international standard for scalable audio coding. Besides its function as an extension of MPEG-4 Advanced Audio Coding (AAC) perceptual audio coder, SLS has a "non-core mode" that is able to offer full scalability. The perceptual audio coder is absent in this mode and scalability is achieved through pure bit-plane coding. In this paper, a perceptually enhanced bit-plane coding method, namely Quad-level Bit-Plane Coding (QBPC) is proposed to enhance the perceptual quality of fully scalable audio at intermediate bitrates. With QBPC structure, the perceptual quality of fully scalable audio coded by SLS is significantly improved in a wide range of intermediate bitrates. Meanwhile this is achieved with trivial added overhead and complexity.
Susanto Rahardja, Soo Ngee Koh
MMSP2
2007 Motion refined medium granular scalability
abstract
In this paper, we propose an interesting scheme to obtain a good tradeoff between motion information and residual information for medium granular scalability (MGS). In this scheme, both motion information and residual information are refined at enhancement layers when the scalable bit rate range is wide, whereas only residual information is refined when the range is narrow. In other words, for the case of wide bit rate range, there can be more than one motion vector fields (MVFs) where one is generated at base layer and others are generated at enhancement layers. When it is narrow, only one MVF is necessary. The layers can either share one MVF or have its own, depending on the bit rate range cross layers. Unlike Coarse Granular Scalability (CGS), the correlation between two adjacent MVFs in MGS is very strong. Hence MGS can be provided in the most important bit rate range to achieve a better tradeoff between motion and residual information and a finer granularity in that range. CGS can be applied in less important bit rate ranges to give a coarse granularity. Experimental results show that the coding efficiency can be improved by up to 1dB compared with existing SNR scalability scheme at high bit rate.
Zhengguo Li, Wei Yao 0001, Susanto Rahardja
VCIP3
2007 A more aggressive prefetching scheme for streaming media delivery over the Internet
abstract
Efficient delivery of streaming media content over the Internet becomes an important area of research as such content is rapidly gaining its popularity. Many research works studied this problem based on the client-proxy-server structure and proposed various mechanisms to address this problem such as proxy caching and prefetching. While the existing techniques can improve the performance of accesses to reused media objects, they are not so effective in reducing the startup delay for first-time accessed objects. In this paper, we try to address this issue by proposing a more aggressive prefetching scheme to reduce the startup delay of first-time accesses. In our proposed scheme, proxy servers aggressively prefetch media objects before they are requested. We make use of servers' knowledge about access patterns to ensure the accuracy of prefetching, and we try to minimize the prefetched data size by prefetching only the initial segments of media objects. Results of trace-driven simulations show that our proposed prefetching scheme can effectively reduce the ratio of delayed requests by up to 38% with very marginal increase in traffic.
Junli Yuan, Qibin Sun, Susanto Rahardja
VCIP3
2007 Subband Kalman filtering incorporating masking properties for noisy speech signal
Chang Huai You, Soo Ngee Koh, Susanto Rahardja
Speech Commun.3
2007 On Integer MDCT for Perceptual Audio Coding
abstract
In MPEG-4 scalable lossless coding (SLS) which was recently published as an ISO standard in June 2006, the integer modified discrete cosine transform (IntMDCT) was adopted to enable efficient lossless reconstruction. In addition, there is an MDCT filterbank which is inherent to the advanced audio coding (AAC) core that is present in the SLS codec. The presence of two filterbanks have undoubtedly increased the complexity of the implementation, and it is for this reason that the MDCT is disabled and the IntMDCT is then the only type of filterbank that is employed in SLS for both lossy and lossless operations. Because of the rounding operations in the IntMDCT, there is a concern if the use of IntMDCT for perceptual audio coding will eventually degrade the fidelity of the audio codec. This paper addresses this concern by analyzing the performance of the IntMDCT in a lossy coding scenario. It is found that noise introduced by the IntMDCT does not affect the perceptual quality of the coded audio under standard playback circumstances. As such, it concludes that the MDCT and IntMDCT filterbanks are interchangeable at lossy bitrate, and the way of using only the IntMDCT filterbank in scalable audio coding is also justified.
Susanto Rahardja, Rongshan Yu, Soo Ngee Koh
IEEE Trans. Speech Audio Process.2
2007 Audible Noise Reduction in Eigendomain for Speech Enhancement
abstract
A signal subspace scheme based on masking properties is proposed for enhancement of speech degraded by additive noise. Since the masking properties are related to the critical frequency band that is derived from the characteristics of human cochlea, the incorporation of masking threshold into a subspace technique requires the transformation between the frequency and eigen domains. We present and apply an invertible transformation between the frequency and eigen domains. In this paper, we use masking properties of the human auditory system to define the audible noise quantity in the eigendomain. We derive the eigen-decomposition of the estimated speech autocorrelation matrix with the assumption of white noise. Subsequently, an audible noise reduction scheme is developed based on a signal subspace technique, and the implementation of our proposed scheme is outlined. We further extend the scheme to the colored noise case. Simulation results show the superiority of our proposed scheme over other existing subspace methods in terms of segmental signal-to-noise ratio (SNR), perceptual evaluation of speech quality (PESQ), modified Bark spectral distortion (MBSD), spectrogram and informal listening tests.
Chang Huai You, Susanto Rahardja, Soo Ngee Koh
IEEE Trans. Speech Audio Process.2
2007 Performance of DS-CDMA Downlink Systems With Orthogonal UCHT Complex Sequences
abstract
This letter investigates a transmitted signaling technique using orthogonal unified complex Hadamard transform (UCHT) spreading sequences and the coherent RAKE receiver in direct-sequence code-division multiple-access (DS-CDMA) downlinks to maintain the orthogonality between users and reduce the effect of multipath fading and interference from other users. A general multipath-fading channel model is assumed. System performance is evaluated by means of signal-to-interference-plus-noise ratio (SINR) at the RAKE receiver. It is shown that the SINR of the system employing UCHT complex sequences is independent of the phase offsets between different paths, while the SINR of the system using Walsh-Hadamard (WH) sequences is related to the squared cosine of path phase offsets. As a result, the bit-error ratio performance of the DS-CDMA downlink system employing UCHT complex sequences is better than that of the system with WH sequences at high SINRs
Shoulie Xie, Susanto Rahardja, Zhenghui Gu
IEEE Trans. Commun.2
2006 Cascaded RLS-LMS Prediction in MPEG-4 Lossless Audio Coding
abstract
A new MPEG-4 standard for lossless audio coding is going to be published in 2006. This coming international standard consists of two parts: the transform-domain scalable to lossless coding (SLS), and the time-domain audio lossless coding (ALS). In ALS, linear prediction is used to compress the dynamic ranges of the input audio signal. The prediction residual is coded by an entropy coder with either Rice code or arithmetic code. There are two prediction modes in ALS: linear predictive coding (LPC) and cascaded RLS-LMS. As the developer of the RLS-LMS prediction, we present this technology in this paper. In RLS-LMS prediction, the input audio samples go through the cascaded DPCM, RLS, and LMS predictors, whose output predictions are linearly combined to generate a prediction for the current input sample. Through MPLG testings, it has been found that ALS with RLS-LMS prediction provides the best lossless compression ratio compared with SLS, ALS with LPC, and several non-MPEG codecs
Susanto Rahardja, Xiao Lin 0001, Rongshan Yu, Pasi Fränti
ICASSP (5)2
2006 Perceptual Kalman Filtering Speech Enhancement
abstract
To enhance the noisy speech signal, a perceptual weighting based Kalman filtering is investigated. Our study is to seek a high perceptual quality of speech enhancement system which optimizes the trade-off between the speech distortion and noise reduction. Using perceptual weighting to replace the masking threshold avoids the frequency domain complexity, it is suitable for time domain Kalman filtering to estimate the state-space vector in only time domain. Through many simulations, it is demonstrated that the proposed perceptual Kalman filtering outperforms the conventional Kalman filtering.
Chang Huai You, Susanto Rahardja, Soo Ngee Koh
ICASSP (1)2
2006 QR-RLS Based Minimum Variance Distortionless Responses Beamformer
abstract
In this paper, a QR-RLS based Minimum Variance Distortionless Responses (MVDR) method, and its systolic array processor, are proposed. The QR-RLS based MVDR has many advantages, such as numerical stability, computational efficiency and pipelined structure in implementation. We also point out that the conventional method, MVDR using QR-RLS method by directly forcing the desired signal to zero is not correct. Numerical experiments are carried out to illustrate the effectiveness of the proposed method.
Zhu Liang Yu, Wee Ser, Susanto Rahardja
ICASSP (3)3
2006 Perceptually Enhanced Bit-Plane Coding for Scalable Audio
abstract
The MPEG-4 scalable to lossless (SLS) audio coding is recently being developed to provide a unified solution for high-compression perceptual audio coding and high-quality lossless audio coding. SLS provides efficient fine granular scalable (FGS) coding from AAC core layer to lossless, and achieves reasonable perceptual quality at its scalable coding range using a sequential bit-plane scanning method, which minimizes the audio distortion according to the spectral shape of the core layer quantization errors. In this paper, it is shown that the perceptual quality performance of SLS at intermediate rates can be further improved by incorporating psycho acoustic model into the bit-plane coding process. In addition, it is also found that such an improvement can be achieved by slightly tweaking the original bit-plane coding process of SLS and hence preserving its nice features such as compatibility to lossless coding and low complexity
Rongshan Yu, Susanto Rahardja
ICME3
2006 Detecting Musical Sounds in Broadcast Audio Based on Pitch Tuning Analysis
abstract
Detecting the presence of musical sounds in broadcast audio is important for content-based indexing and retrieval of auditory and visual information in radio and TV programs. In this paper, we propose a novel approach for musical sounds detection in broadcast audio based on the analysis of the characteristic feature of musical tones, pitch tuning. A spectral analysis method is presented for detecting the evidence of pitch tuning in the audio signal. Unlike the existing methods for discriminating speech and music, the proposed technique is not limited by inadequate training data, and it can deal with the case of music mixed with speech. In addition, the technique can be efficiently implemented for real-time application. Experiments based on TRECVID data set have shown good performance of the proposed technique
Yongwei Zhu, Qibin Sun, Susanto Rahardja
ICME3
2006 Efficient computation of fixed polarity arithmetic expansions for ternary functions
abstract
An efficient algorithm for generating fixed polarity arithmetic expansions for ternary functions is presented. It calculates the required spectral coefficients in a recursive manner based on a developed definition of the polarity matrix. The application of the algorithm for generating both complete polarity matrix and selected fixed polarity arithmetic expansion is given. Computational cost of the algorithm in terms of required number of additions and multiplications is also derived and it is shown to be more efficient than the calculation by matrix multiplication. Fast flow diagrams for implementation of the algorithm on hardware are also shown.
Bogdan J. Falkowski, Cicilia C. Lozano, Susanto Rahardja
ISCAS3
2006 Algorithms for generation of quaternary fixed polarity arithmetic spectra
abstract
Two different algorithms for generating the complete fixed polarity arithmetic transform polarity matrix of a quaternary function are presented. The first approach utilizes relations between coefficient vectors whereas the second one is using relations between the column coefficient vectors to reduce the computational cost. Both algorithms are described and their computational costs are derived and compared.
Cicilia C. Lozano, Bogdan J. Falkowski, Susanto Rahardja
ISCAS3
2006 Generalized Fastest LIA Transform Spectra Calculation by Systolic Processor
abstract
Hardware calculation of generalized fastest linearly independent arithmetic (LIA) expansions using systolic processor is presented in this paper. The relation between the forward flow graph of a particular LIA transform and the systolic processor structure for its spectra calculation is given. In general, a particular systolic hardware structure can be used for more than one fastest LIA transforms with appropriate reordering of inputs and/or outputs
Bogdan J. Falkowski, Cicilia C. Lozano, Susanto Rahardja
ISIT3
2006 Issues, Challenges, and Future Directions in Multimedia Research
abstract
Summary form only given. This article discusses issues, challenges, and future directions in multimedia research, along the following three scale-oriented issues plus the one which is common to them
Masahito Hirakawa, Max Mühlhäuser, Susanto Rahardja, Phillip C.-Y. Sheu, Larry Smarr, Jeffrey J. P. Tsai
ISM3
2006 Perceptually Prioritized Bit-Plane Coding for High-Definition Advanced Audio Coding
abstract
Wide bitrate range scalability is now the latest trend in audio coding. A lot of efforts has been devoted to the development of algorithms for more efficient scalable audio coder that scales from very low bitrate. Scalable audio coding technique such as MPEG-4 scalable lossless coding (SLS) offers a unified solution for high-compression perceptual audio and high-quality lossless audio. SLS provides a fine-grain scalable extension of the well-known MPEG-4 advanced audio coding (AAC) perceptual audio coder up to fully lossless reconstruction. Recently, the combination of SLS and AAC coder is renamed as "high definition advanced audio coding" (HD-AAC). It is observed that HD-AAC can be further improved at intermediate enhancement bitrate when the core bitrate is low. In this paper, a perceptually prioritized bit-plane coding (PPBPC) is proposed. With this novel coding scheme, the bit-plane coding is performed with priorities according to the perceptual information of the signal to be coded. By using this low-complexity structure with trivial extra side information, the bit-plane coding for scalable audio can be implemented in a perceptually more efficient manner and the quality of the audio under aforementioned scenario is greatly improved
Susanto Rahardja, Soo Ngee Koh
ISM2
2006 Asynchronous Multi-Carrier DS-CDMA with UCHT-Based Complex Spreading Sequences
abstract
Quadriphase complex sequences based on unified complex Hadamard transform (UCHT) are orthogonal and easy to generate [6]. There are sixty-four sets in UCHT-based sequences. Some sets of UCHT sequences provide better auto-correlation properties than orthogonal walsh-Hadamard (WH) sequences. In this paper, UCHT-based sequences are applied as spreading sequences in an asynchronous MC-DS-CDMA system. The bit-error rate (BER) performance of the underlying system with complex spreading sequences is investigated and simulation results show that the asynchronous MC-DS-CDMA system spread by UCHT sequences outperforms that spread by WH sequences in frequency selective fading channel.
Zhenghui Gu, Shoulie Xie, Susanto Rahardja
VTC Spring3
2006 Adaptive rate control for H.264
Zhengguo Li, Wen Gao 0001, Feng Pan 0002, S. W. Ma, Keng-Pang Lim, G. N. Feng, Xiao Lin 0001, Susanto Rahardja, H. Q. Lu, Yan Lu 0001
J. Vis. Commun. Image Represent.8
2006 Perceptual quality and objective quality measurements of compressed videos
Ee Ping Ong, Xiaokang Yang 0001, Weisi Lin, Zhongkang Lu, Susu Yao, Xiao Lin 0001, Susanto Rahardja, Choong Seng Boon
J. Vis. Commun. Image Represent.7
2006 Adaptive frame skipping based on spatio-temporal complexity for low bit-rate video coding
Feng Pan 0002, Zhiping Lin 0001, Xiao Lin 0001, Susanto Rahardja, W. Juwono, F. Slamet
J. Vis. Commun. Image Represent.4
2006 Masking-based beta-order MMSE speech enhancement
Chang Huai You, Soo Ngee Koh, Susanto Rahardja
Speech Commun.3
2006 A fine granular scalable to lossless audio coder
abstract
This paper presents Advanced Audio Zip (AAZ), a fine grained scalable to lossless (SLS) audio coder that has recently been adopted as the reference model for MPEG-4 audio SLS work. AAZ integrates the functionalities of high-compression perceptual audio coding, fine granular scalable audio coding, and lossless audio coding in a single framework, and simultaneously provides backward compatibility to MPEG-4 Advanced Audio Coding (AAC). AAZ provides the fine granular bit-rate scalability from lossy to lossless coding, and such a scalability is achieved in a perceptually meaningful way, i.e., better perceptual quality at higher bit-rates. Despite its abundant functionalities, AAZ only introduces negligible overhead in terms of lossless compression performance compared with a nonscalable, lossless only audio coder. As a result, AAZ provides a universal yet efficient solution for digital audio applications such as audio archiving, network audio streaming, portable audio playing, and music downloading which were previously catered for by several different audio coding technologies, and eliminates the need for any transcoding system to facilitate sharing of digital audio contents across these application domains.
Rongshan Yu, Susanto Rahardja, Xiao Lin 0001, Chi Chung Ko
IEEE Trans. Speech Audio Process.2
2006 Implicit Bit Allocation for Combined Coarse Granular Scalability and Spatial Scalability
abstract
A new type of implicit bit allocation (IBA) is studied for the combined coarse granular scalability (CGS) and spatial scalability. Given a region which is a quality level at a spatial resolution and an input of weighting factor in the region that is determined by customers' interest, the IBA is formulated as a multiple objective optimization problem. The IBA exhibits a distinguished feature that allows bits allocation to each region being fixed and only tradeoff between motion information and residual information in each region can be properly set such that coding efficiency of each region is guaranteed in order according to the weighting factor. Due to the nonlinearity of the combined CGS and spatial scalability, this optimization problem is very complex. In this paper, a simple solution to the optimization problem is proposed by using the conventional approach of "Divide and Conquer" or "Hierarchy." Two cross-layer motion estimation/motion compensation (ME/MC) schemes that were introduced for the CGS and spatial scalability are further adopted to support the solution. It is shown that a combined CGS and spatial scalability scheme together with adaptability to customer composition allows the solution to achieve a customer oriented scalable tradeoff (COST)
Zhengguo Li, Susanto Rahardja, Hanwu Sun
IEEE Trans. Circuits Syst. Video Technol.2
2005 Signal Subspace Speech Enhancement for Audible Noise Reduction
abstract
A novel subspace-based speech enhancement scheme, based on a criterion of audible noise reduction, is considered. The masking properties of the human auditory system are used to define the audible noise quantity in the eigen-domain. Subsequently, an audible noise reduction scheme is developed based on a signal subspace technique. We derive the eigendecomposition of the estimated speech autocorrelation matrix with the assumption of white noise and outline the implementation of our proposed scheme. We further extend the scheme to the colored noise case. Simulation results show that our proposed scheme outperforms many existing subspace methods in terms of segmental signal-to-noise ratio (SNR), perceptual evaluation of speech quality (PESQ) and informal listening tests.
Chang Huai You, Soo Ngee Koh, Susanto Rahardja
ICASSP (1)3
2005 Improving coding efficiency for MPEG-4 Audio Scalable Lossless coding
abstract
The recently introduced MPEG standard for lossless audio coding, MPEG-4 Audio Scalable to Lossless (SLS) coding technology, provides a universal audio format that integrates the functionalities of lossy audio coding, lossless audio coding and fine granular scalable audio coding in a single framework. We propose two coding methods that improve the coding efficiency of SLS, namely, a context-based arithmetic code (CBAC) method and a low energy mode code method. These two coding methods work harmonically with the current SLS framework and preserve all its desirable features, such as fine granular scalability, while successfully improving its lossless compression ratio performance.
Rongshan Yu, Xiao Lin 0001, Susanto Rahardja, Chi Chung Ko
ICASSP (3)3
2005 Customer Adaptive Combined SNR and Spatial Scalability
abstract
A new type of region of interest (ROI) is proposed in this paper for scalable video coding with the whole region be the combined SNR and spatial scalability, and a sub-region be a specific choice of spatial resolution and bit rate range. The ROI is applied to design a combined SNR and spatial scalability scheme that is adaptive to customer composition such that an optimal customer oriented scalable tradeoff (COST) can be achieved. The profit can thus be maximized
Zhengguo Li, Susanto Rahardja, Xiao Lin 0001, Wei Yao 0001
MMSP2
2005 MPEG-4 Scalable to Lossless Audio Coding - Emerging International Standard for Digital Audio Compression
abstract
Recently, with the advance of network and storage technologies, it is becoming realistic that people will enjoy high sampling rate, high resolution audio contents with lossless quality. Envisioning of such a need, the international standardization body MPEG has recently introduced a scalable tool for lossless audio coding, namely, MPEG-4 Audio Scalable to Lossless (SLS) coding. MPEG-4 SLS integrates the functionalities of lossless audio coding, perceptual audio coding, and fine granular scalable audio coding in a single framework; meanwhile it provides backward compatibility to MPEG-4 Advanced Audio Coding (AAC) at the bit-stream level. This new tool, in combination with the existing MPEG audio toolset, provides a universal digital audio format that can be used in a variety of application domains such as professional audio, Internet music, consumer electronics, broadcasting
Rongshan Yu, Xiao Lin 0001, Susanto Rahardja
MMSP3
2005 Geometrically determining the leaky bucket parameters for video streaming over constant bit-rate channels
Ping Li 0002, Weisi Lin, Susanto Rahardja, Xiao Lin 0001, Xiaokang Yang 0001, Zhengguo Li
Signal Process. Image Commun.3
2005 An invertible frequency eigendomain transformation for masking-based subspace speech enhancement
abstract
Masking properties have been widely exploited in speech enhancement techniques, especially those implemented in the spectral domain. The incorporation of auditory masking in a subspace technique invariably requires a transformation linking the frequency and eigendomains. In this letter, an invertible transformation between the frequency and eigendomains is derived. The proposed transformation is verified through a conventional masking-based subspace speech-enhancement method. Simulation results show that our proposed transformation for speech enhancement outperforms the conventional transformation in terms of segmental signal-to-noise ratio (SNR), perceptual evaluation of speech quality (PESQ), and listening tests.
Chang Huai You, Soo Ngee Koh, Susanto Rahardja
IEEE Signal Process. Lett.3
2005 beta-order MMSE spectral amplitude estimation for speech enhancement
abstract
This paper proposes /spl beta/-order minimum mean-square error (MMSE) speech enhancement approach for estimating the short time spectral amplitude (STSA) of a speech signal. We analyze the characteristics of the /spl beta/-order STSA MMSE estimator and the relation between the value of /spl beta/ and the spectral amplitude gain function of the MMSE method. We further investigate the effectiveness of a range of fixed-/spl beta/ values in estimating STSA based on the MMSE criterion, and discuss how the /spl beta/ value could be adapted using the frame signal-to-noise ratio (SNR). The performance of the proposed speech enhancement approach is then evaluated through spectrogram inspection, objective speech distortion measures and subjective listening tests using several types of noise sources from the NOISEX-92 database. Evaluation results show that our approach can achieve a more significant noise reduction and a better spectral estimation of weak speech spectral components from a noisy signal as compared to many existing speech enhancement algorithms.
Chang Huai You, Soo Ngee Koh, Susanto Rahardja
IEEE Trans. Speech Audio Process.3
2005 Fast mode decision algorithm for intraprediction in H.264/AVC video coding
abstract
The H.264/AVC video coding standard aims to enable significantly improved compression performance compared to all existing video coding standards. In order to achieve this, a robust rate-distortion optimization (RDO) technique is employed to select the best coding mode and reference frame for each macroblock. As a result, the complexity and computation load increase drastically. This paper presents a fast mode decision algorithm for H.264/AVC intraprediction based on local edge information. Prior to intraprediction, an edge map is created and a local edge direction histogram is then established for each subblock. Based on the distribution of the edge direction histogram, only a small part of intraprediction modes are chosen for RDO calculation. Experimental results show that the fast intraprediction mode decision scheme increases the speed of intracoding significantly with negligible loss of peak signal-to-noise ratio.
Feng Pan 0002, Xiao Lin 0001, Susanto Rahardja, Keng-Pang Lim, Zhengguo Li, Dajun Wu, Si Wu 0004
IEEE Trans. Circuits Syst. Video Technol.3
2005 Fast intermode decision in H.264/AVC video coding
abstract
The new video coding standard, H.264/MPEG-4 AVC, uses variable block sizes ranging from 4/spl times/4 to 16/spl times/16 in interframe coding. This new feature has achieved significant coding gain compared to coding a macroblock (MB) using fixed block size. However, this feature results in extremely high computational complexity when brute force rate distortion optimization (RDO) algorithm is used. This paper proposes a fast intermode decision algorithm to decide the best mode in intercoding. It makes use of the spatial homogeneity and the temporal stationarity characteristics of video objects. Specifically, spatial homogeneity of a MB is decided based on the MB's edge intensity, and temporal stationarity is decided by the difference of the current MB and it colocated counterpart in the reference frame. Based on the homogeneity and stationarity of the video objects, only a small number of intermodes are selected in the RDO process. The experimental results show that the fast intermode decision algorithm is able to reduce on the average 30% encoding time, with a negligible peak signal-to-noise ratio loss of 0.03 dB or, equivalently, a bit rate increment of 0.6%.
Dajun Wu, Feng Pan 0002, Keng-Pang Lim, Si Wu 0004, Zhengguo Li, Xiao Lin 0001, Susanto Rahardja, Chi Chung Ko
IEEE Trans. Circuits Syst. Video Technol.7
2005 Rate Control for Videophone Using Local Perceptual Cues
abstract
We present a method for extracting local visual perceptual cues and its application for rate control of videophone, in order to ensure the scarce bits to be assigned for maximum perceptual coding quality. The optimum quantization step is determined with the rate-distortion model considering the local perceptual cues in the visual signal. For extraction of the perceptual cues, luminance adaptation and texture masking are used as the stimulus-driven factors, while skin color serves as the cognition-driven factor in the current implementation. Both objective and subjective quality evaluations are given by evaluating the proposed perceptual rate control (PRC) scheme in the H.263 platform, and the evaluations show that the proposed PRC scheme achieves significant quality improvement in block-based coding for bandwidth-hungry applications.
Xiaokang Yang 0001, Weisi Lin, Zhongkang Lu, Xiao Lin 0001, Susanto Rahardja, Ee Ping Ong, Susu Yao
IEEE Trans. Circuits Syst. Video Technol.5
2005 Performance evaluation for quaternary DS-SSMA communications with complex signature sequences over Rayleigh-fading channels
abstract
Performance of quaternary direct-sequence spread-spectrum multiple-access (DS-SSMA) systems with complex signature sequences and complex modulators and receivers in flat Rayleigh fading is investigated in this paper. Due to the availability of potentially large sets of complex sequences with good correlation characteristics, the interest of using complex spreading sequences in DS-SSMA has increased dramatically. The complex spreading sequences investigated in this paper include the recently introduced orthogonal unified complex Hadamard transform (UCHT) sequences. In this paper, complex processing in modulators and receivers is also employed in order to take advantage of the correlation properties of complex signature sequences. The average bit error rate (BER) for quaternary synchronous systems is obtained first, and then the BER for quaternary asynchronous systems is evaluated using characteristic function approach. Result based on Gaussian approximation method is also presented for asynchronous systems. The numerical examples illustrate that the systems based on UCHT spreading sequences perform generally better than the Gold sequences and the 4-phase family A-sequences.
Shoulie Xie, Susanto Rahardja
IEEE Trans. Wirel. Commun.2
2004 A fast algorithm of integer MDCT for lossless audio coding
abstract
A new fast algorithm to implement integer modified discrete cosine transform (IntMDCT) is proposed. It is shown that the total rounding operations required for this algorithm are only 2.5N, where N is the block size. As a result, its approximation error is far less than that of directly converted integer transforms. At the same time, the complexity is greatly reduced, which results in improved performance of a lossless audio coding system employing the IntMDCT in terms of compression ratio and complexity costs.
Susanto Rahardja, Rongshan Yu, Xiao Lin 0001
ICASSP (4)2
2004 Geometrically determining leaky bucket parameters for video streaming over constant bit-rate channels
abstract
For streaming of pre-encoded bitstreams over constant bit-rate (CBR) channels, the channel bandwidth, the receiver buffer capacity as well as the latency requirement vary greatly from application to application. We propose an algorithm to determine the minimum buffer size and the minimum start-up delay required for streaming a pre-encoded bitstream over CBR channels at any specific bit rate. The proposed method employs geometric operations to derive the optimal determination for low or high bit rates and sub-optimal determination for medium bit rates. The results have been compared with the H.264/AVC hypothetical reference decoder. The proposed approach provides a theoretical insight and a simple but effective algorithm for determining the leaky bucket parameters for video streaming over CBR channels.
Ping Li 0002, Weisi Lin, Susanto Rahardja, Xiao Lin 0001, Xiaokang Yang 0001, Zhengguo Li
ICASSP (3)3
2004 An MMSE speech enhancement approach incorporating masking properties
abstract
The paper describes a new speech enhancement approach which employs adaptive /spl beta/-order minimum mean square error (MMSE) spectral estimation of the short time spectral amplitude (STSA) of a speech signal. In the proposed approach, the human perceptual auditory masking effect is incorporated into the speech enhancement algorithm. The relationship between the value of /spl beta/ and the noise masking threshold is considered. The algorithm is based on a criterion by which the audible noise may be masked rather than being attenuated, thereby reducing the chance of speech distortion. Performance assessment is given to show that our proposal can achieve a more significant noise reduction and a better spectral estimation of weak speech spectral components from a noisy signal as compared to many existing speech enhancement algorithms.
Chang Huai You, Soo Ngee Koh, Susanto Rahardja
ICASSP (1)3
2004 A scalable lossy to lossless audio coder for MPEG-4 lossless audio coding
abstract
In this paper, we present Advanced Audio Zip (AAZ), a scalable lossless audio coding technology that was recently selected as the reference model for MPEG audio scalable lossless coding (SLS) work. AAZ provides excellent compression performance while delivering fine grain bit-rate scalability from lossy to lossless coding. Moreover, AAZ provides backward compatibility to the MPEG advanced audio coding (AAC) system by embedding an AAC compliant bit-stream into the lossless bit-stream. As a result, AAZ serves as a universal coding solution with functionalities that were previously offered by several distinct audio coding technologies such as lossless audio coding, perceptual audio coding, or scalable audio coding; and maximizes the interchangeability for digital audio contents migrating among these application domains.
Rongshan Yu, Xiao Lin 0001, Susanto Rahardja, Chi Chong Ko
ICASSP (3)3
2004 Adaptive rate control for H.264
abstract
This paper presents a rate control scheme for H.264 by introducing the concept of basic unit and a linear prediction model. The basic unit can be a macroblock (MB), a slice, or a frame. It can be used to obtain a trade-off between the overall coding efficiency and the bits fluctuation. The linear model is used to solve the chicken and egg dilemma existing in the rate control of H.264. Both constant bit rate (CBR) and variable bit rate (VBR) cases are studied. Our scheme has been adopted by H.264.
Zhengguo Li, Feng Pan 0002, Keng-Pang Lim, Xiao Lin 0001, Susanto Rahardja
ICIP5
2004 Fast intra mode decision algorithm for H.264/AVC video coding
abstract
The emerging H.264-AVC video coding standard aims to significantly improve compression performance compared to all existing video coding standards. In order to achieve this, a robust rate-distortion optimization (RDO) technique is employed to select the best coding mode and reference frame for each macroblock. As a result, the complexity and computation load increase drastically. This paper presents a fast mode decision algorithm for H.264 intra prediction based on local edge information. Prior to intra prediction, an edge map is created and a local edge direction histogram is then established for each sub-block. Based on the distribution of the edge direction histogram, only a small part of intra prediction modes are chosen for RDO calculation. Experimental results show that the last intra mode decision scheme increases the speed of intra coding significantly with negligible loss of PSNR.
Feng Pan 0002, Xiao Lin 0001, Susanto Rahardja, Keng-Pang Lim, Zhengguo Li, Dajun Wu, Si Wu 0004, C. All, W. Ye, Z. Liang
ICIP3
2004 An iterative method for hypothetical reference decoder
abstract
We propose a simple iterative method to verify whether a coded bitstream conforms to a hypothetical reference decoder (HRD). A concept of maximum tolerated delay (MTD) is introduced to study the possible low-delay operation, such that we can bound the degree of incorrect motion rendition caused by the variation in end to end delay in the neighborhood of big pictures. Our method can be used in the design of a rate control algorithm to improve the possibility of a coded bitstream that conforms to the HRD.
Zhengguo Li, Nam Ling, Susanto Rahardja, Xiao Lin 0001, Ping Li 0002
ICME3
2004 Complexity adaptive quantization for intra-frames in very low bit rate video coding
abstract
Conventional rate control schemes focus on the problem of finding an optimal quantization value for P- and B-frames. No rate control is available for the encoding of I-frames. This could pose big problems due to the large number of bits an I-frame could generate, and due to the fact that the number of bits varies drastically from sequence to sequence, depending on their image complexity. This problem becomes severe especially at very low bit rate. Therefore, a mechanism to allocate data bits to an I-frame according to its complexity is indispensable in order to have constant coding quality. This work presents a mechanism to establish for I-frames a generic relationship between the quantization value, data bits, and their image complexity. Experimental results show that this generic relationship provides a fairly accurate estimation of quantization value for an I-frame at given data bits and image complexity, and is very useful in controlling the data bits generated by an I-frame.
Feng Pan 0002, Zhengguo Li, Keng-Pang Lim, Xiao Lin 0001, Susanto Rahardja, Dajun Wu, Si Wu 0004
ICME5
2004 A directional field based fast intra mode decision algorithm for H.264 video coding
abstract
The H.264 video coding standard can achieve considerably higher coding efficiency than previous standards. In order to achieve this, a robust rate-distortion optimization (RDO) technique is employed to select the best coding mode for each macroblock. As a result. the encoder complexity is increased considerably. This paper presents a directional field based fast intra mode decision algorithm to improve the encoder's efficiency. Prior to intra prediction, the directional field is calculated for all the block size to decide the dominant edge direction in the blocks. Based on the edge direction information, only a small number of prediction modes are chosen for RDO calculation. Experimental results show that the fast intra mode decision algorithm increases the speed of intra coding significantly with negligible loss of PSNR
Feng Pan 0002, Xiao Lin 0001, Susanto Rahardja, Keng-Pang Lim, Zhengguo Li
ICME3
2004 Proactive frame-skipping decision scheme for variable frame rate video coding
abstract
Many rate control algorithms focus on the adjustment of quantisation values to retain a certain buffer level, and arbitrary frame-skipping is often needed to keep the buffer from overflow at very low bit-rates. A content adaptive rate control is proposed to optimise the balance between spatial and temporal quality via active frame-skipping. The occurrence of frame-skipping is jointly dependent on the temporal and spatial quality of the video, and on the fullness of the buffer. This helps to achieve a consistent spatial and temporal quality and to enhance the overall perceptual quality. Experimental results show that the new scheme is simple but very effective, with large average PSNR gains and consistently improved visual quality, and the improvement in perceptual quality is much more significant than that in average PSNRs.
Feng Pan 0002, Xiao Lin 0001, Susanto Rahardja, Keng-Pang Lim, Zhengguo Li, Dajun Wu, Si Wu 0004
ICME3
2004 Measuring blocking artifacts using edge direction information
abstract
Block-based transform coding is the most popular approach for image and video compression. The objective measurement of blocking artifacts plays an important role in the design, optimization. and assessment of image and video coding systems. This paper presents a new algorithm for measuring blocking artifacts in images and videos. Instead of using the traditional pixel discontinuity along the block boundary, we use the edge directional information of the images. The new algorithm does not need the exact location of the block boundary thus is invariant to the displacement, rotation and scaling of the images. Experiments on various still images and videos show that the new blockiness measure is very efficient in terms of computational complexity and memory usage, and can produce blocking artifact measurement consistent with subjective rating
Feng Pan 0002, Xiao Lin 0001, Susanto Rahardja, Ee Ping Ong, Weisi Lin
ICME3
2004 A statistics study of the MDCT coefficient distribution for audio
abstract
The modified discrete cosine transform (MDCT) has been widely used in many transform audio coding algorithms such as MPEG-1/2 layer III (mp3), MPEG-2/4 AAC, Dolby AC2/AC3, and numerous experimental audio coding algorithms. In this paper, we study the probabilistic distribution properties of the MDCT coefficient for audio signals. It is shown that the generalized Gaussian function with distribution parameter r=0.5 or r=1 (Laplacian) provides a good approximation to the distributions of MDCT coefficients for a variety of audio signals. Results from our study also show that although the distribution of these coefficients is not strictly Laplacian, the divergence between them is in fact very small. Therefore, it leads to only marginal redundancy if these coefficients are simply coded with some low-complexity codes designed for Laplacian sources.
Rongshan Yu, Xiao Lin 0001, Susanto Rahardja, Chi Chung Ko
ICME3
2004 Kalman filtering speech enhancement incorporating masking properties for mobile communication in a car environment
abstract
A single channel speech enhancement system based on Kalman filtering and masking properties of the human auditory system is studied. The main objective of our study is to develop a high quality speech enhancement system for car mobile communication, which optimizes the tradeoff between speech distortion and noise reduction. To develop a robust Kalman filtering speech enhancement system, we propose that the noise variance estimate be modified as a function of masking threshold in each subband. By using such a Kalman filtering with masking method, we can achieve better results as compared to full band Kalman filtering, especially in the case of weak speech spectral components in noise. Through a great deal of simulations, we show the advantages of our proposed system in terms of objective and subjective measurements as compared to many existing speech enhancement methods.
Chang Huai You, Soo Ngee Koh, Susanto Rahardja
ICME3
2004 A locally adaptive algorithm for measuring blocking artifacts in images and videos
Feng Pan 0002, Xiao Lin 0001, Susanto Rahardja, Weisi Lin, Ee Ping Ong, Susu Yao, Zhongkang Lu, Xiaokang Yang 0001
Signal Process. Image Commun.3
2003 An LMI-based decentralized H∞ filtering for interconnected linear systems
abstract
This paper focuses on decentralized H/sub /spl infin// filtering problem for interconnected linear systems. The problem we address is to find a decentralized filter where each local filter is based only on local available information on its own subsystem and the overall filtering error is totally asymptotically stable and the L/sub 2/-gain from the exogenous noise input to the filtering error less than a prespecified level. This paper shows that the decentralized H/sub /spl infin// filtering problem can be solved by using linear matrix inequality (LMI) techniques, which are numerically efficient due to recent advances in convex optimization.
Shoulie Xie, Lihua Xie 0001, Susanto Rahardja
ICASSP (6)3
2003 Adaptive β-order MMSE estimation for speech enhancement
abstract
This paper introduces an adaptive /spl beta/-order minimum mean square error (MMSE) spectral estimator for the short time spectral amplitude (STSA) of speech. The characteristic of /spl beta/-order MMSE attenuation function is introduced and analyzed. The performance of the proposed adaptive /spl beta/-order MMSE has been thoroughly examined by a large number of computer simulations. The new proposed scheme has been found to outperform the conventional a priori SNR Wiener filtering, the Ephraim and Malah (1984) STSA-MMSE and log spectral amplitude (LSA) schemes. It can achieve a more significant noise reduction and a better spectral estimation for weak speech components from a noisy speech signal as compared to the conventional schemes.
Chang Huai You, Soo Ngee Koh, Susanto Rahardja
ICASSP (1)3
2003 Bit-plane Golomb coding for sources with Laplacian distributions
abstract
This paper presents a bit-plane coding algorithm for Laplacian distributed sources that are commonly encountered in signal compression applications. By exploiting the statistical characteristics of the sources, the proposed algorithm achieves a rate-distortion performance that is essentially comparable to an optimal nonscalable scalar quantizer, while at the same time operates at a complexity level suitable for most practical implementations.
Rongshan Yu, Chi Chong Ko, Susanto Rahardja, Xiao Lin 0001
ICASSP (4)3
2003 3D shape modeling by color phase stepping light projection
abstract
Color encoded phase-stepping light projection method is a new and promising technique for 3D shape modeling However, the 3D model acquired is often smeared by large error. The main cause of the error is color coupling amongst the three primary colors RGB. In this paper, we first analyzed the color-coupling problem. It is found that there is a strong coupling between G and R element. The coupling between R and G are proportional to the phase interval between them and the overall intensity of image. Second, we proposed an adaptive phase stepping method to alleviate the color coupling errors efficiently and improve accuracy effectively. An algorithm corresponding to a specific paradigm with R-G-B phase step set to 0-45-180 is given and is applied to measure different objects. Experimental results demonstrate the effectiveness of the method.
Lijun Jiang, Shiqian Wu, Dajun Wu, Ee Ping Ong, Susanto Rahardja
ICME5
2003 Adaptive frame layer rate control for H.264
abstract
This paper proposes an adaptive frame layer rate control scheme for H.264 by introducing a linear model to predict the mean absolute difference (MAD) of current frame by that of previous one. The target bit rate for each frame is computed by adopting a fluid flow traffic model and linear tracking theory. The corresponding quantization parameter is computed by using a quadratic rate-distortion model. The rate distortion optimization (RDO) is then performed for all macroblocks (MBs) in the current frame by the quantization parameter. Both constant bit rate (CBR) and variable bit rate (VBR) cases are studied. The average PSNR is improved up to 0.75 dB compared to an encoder with fixed quantization parameter.
Zhengguo Li, Feng Pan 0002, Keng-Pang Lim, Genan Feng, Xiao Lin 0001, Susanto Rahardja, Dajun Wu
ICME6
2003 Video streaming on embedded devices through GPRS network
abstract
We introduce a PDA-based live video streaming system on GPRS network based on MPEG-4 video compression standard. Due to the limited computational resources of PDA, all the key modules of MPEG-4 codec are efficiently implemented and optimized such as multithreading, buffer design, wireless communication, encoder and decoder. Several novel techniques are developed in the coding, streaming as well as the post- processing stages of the system.
Keng-Pang Lim, Dajun Wu, Si Wu 0004, Susanto Rahardja, Xiao Lin 0001, Lijun Jiang, Rongshan Yu, Feng Pan 0002, Zhengguo Li, Susu Yao, Genan Feng, Chi Chung Ko
ICME4
2003 Binary image watermarking through biased binarization
abstract
This paper presents a watermarking algorithm for binary images. The original binary image is blurred to a gray-level image and we embed the watermark by biasing the threshold in binarization. A loop is used to control the quality of watermarked images and robustness, and a key is generated for extraction. We employ error correction codes to reduce extraction error. This algorithm can be applied to general binary images except dithered images. Experiments show that the distortion in the watermarked image is not obtrusive and the algorithm provides some degree of robustness.
Haiping Lu, Alex Chichung Kot, Susanto Rahardja
ICME3
2003 A fine granular scalable perceptually lossy and lossless audio coder
abstract
This paper presents advanced audio zip (AAZ), an audio codec that provides the fine granular bit-rate scalability from lossy to lossless coding. Perceptually embedded coding principle is employed in AAZ to provide lossy reconstruction with optimal perceptual quality at intermediate bit-rates. AAZ also provides the backward compatibility where the lossless bit-stream embeds a compliant MPEG-4 AAC bit-stream.
Rongshan Yu, Xiao Lin 0001, Susanto Rahardja, Chi Chung Ko
ICME3
2003 Post-processing for JPEG 2000 image coding using recursive line filtering based on a fuzzy control model
Susu Yao, Susanto Rahardja, Xiao Lin 0001, Keng-Pang Lim, Zhongkang Lu
VCIP2
2003 UCHT-based complex sequences for asynchronous CDMA system
abstract
The use of orthogonal spreading codes has attracted much attention due to their ability to suppress interference from other users, compared with the nonorthogonal sequences in the synchronous case. In this paper, new sets of orthogonal sequences derived from the unified complex Hadamard transforms (UCHTs) are investigated. Various correlation properties of the sequences are mathematically derived and analyzed. It is shown that some UCHT sequences provide better autocorrelation properties than orthogonal Walsh-Hadamard sequences. Performance comparisons between UCHT sequences, Gold, small set of Kasami, and m-sequences show that some UCHT sequences outperform these well-known spreading sequences under simulation of systems in the presence of multiple access interference and additive white Gaussian noise.
Susanto Rahardja, Wee Ser, Zinan Lin 0003
IEEE Trans. Commun.1
2000 A Standalone Video Communication System for Wireless Applications
abstract
This paper describes a low bit rate standalone portable video communication system developed for wireless applications. The video codec is based on H.263, and the audio codec is based on G.723.1. The transmission bandwidth is limited to 25 kHz. In order to have reasonable good video and audio quality, the source bit rate is set to 38 kbps. This includes video, audio and synchronization. The bit stream is embedded with 26 kbps for the error protection, resulted in a total of 64 kbps. A 16-QAM is used for the modulation in order to meet the bandwidth requirement. The whole system deployed three general purpose DSP for the real time implementation of video coding, audio coding, channel coding and system control. Besides, an RF module is developed for the transmission on UHF band.
HonCheong Ng, HockLye Toh, Susanto Rahardja, Hui Lan, Xiao Lin 0001
ISCC3
1999 Fast Linearly Independent Arithmetic Expansions
abstract
The concept of Linearly Independent arithmetic (LIA) transforms and expansions is introduced in this paper. The recursive ways of generating forward and inverse fast transforms for LIA are presented. The paper describes basic properties and lists those LIA transforms which have convenient fast forward algorithms and easily defined inverse transforms. In addition, those transforms which require horizontal or vertical permutations to have fast transform are also discussed. The computational advantages and usefulness of new expansions based on LIA logic in comparison to known arithmetic expansions are discussed.
Susanto Rahardja, Bogdan J. Falkowski
IEEE Trans. Computers1
1996 A GA paradigm for learning fuzzy rules
Susanto Rahardja, Bah-Hwee Gwee
Fuzzy Sets Syst.2
1995 Fast Transforms for Orthogonal Logic
abstract
The ways of generation of forward and inverse fast transforms for recently introduced orthogonal logic have been presented. The list of all fast transforms is thoroughly discussed. The paper summarizes those orthogonal transforms which have fast algorithms and easily defined recursive equations. In addition, those transforms which require one or more permutations to have fast transform are also discussed.
Bogdan J. Falkowski, Susanto Rahardja
ISCAS2
1994 Sign Haar Transform
abstract
A non-linear transform, called "Sign Haar Transform" has been introduced. The transform is unique and converts binary/ternary vectors into ternary spectral domain. Recursive definitions and Fast Transforms for the calculation of Sign Haar Transform have been developed. The new transform is extremely computationally effective both in terms of memory requirements and processing time.>
Bogdan J. Falkowski, Susanto Rahardja
ISCAS2