VLDB 2026 Research / reviewers in the wild / expert
Sridhar Krishnan 0001
dblp:184/9465 · also Sridhar Sri Krishnan
· DBLP profile ↗
67ranked-venue papers
0as first author
6since 2021 · last 2025
0000-0002-4659-564XORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 36 · 2 since 2021Artificial intelligence and machine learning · 15 · 4 since 2021Applied, interdisciplinary, general and emerging computing · 10Human-computer interaction and ubiquitous computing · 7Systems, architecture and hardware · 3Computer networks · 2Security and privacy · 2
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Computer graphics and multimedia
9 papers |
Audio and music processing · 90% Image and video processing · 7% Multimedia systems and quality of experience · 2% | |
| Network and information security
2 papers |
Digital forensics and information hiding · 100% |
Topics — the 18 heaviest of 20, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Audio and music processing
audio representation |
0.8 | 1 | 2024 | Time-Frequency Scattergrams for Biomedical Audio Signal Representation and Classification · IEEE ACM Trans. Audio Speech Lang. Process. 2024 |
Audio and music processing
computational auditory scene analysis |
0.6 | 2 | 2022 | Multimodal System for Audio Scene Source Counting and Analysis · IEEE ACM Trans. Audio Speech Lang. Process. 2022 Audio Signal Feature Extraction and Classification Using Local Discriminant Bases · IEEE Trans. Speech Audio Process. 2007 |
Audio and music processing › speaker diarization
speaker counting |
0.6 | 1 | 2022 | Multimodal System for Audio Scene Source Counting and Analysis · IEEE ACM Trans. Audio Speech Lang. Process. 2022 |
Audio and music processing
audio feature extraction |
0.4 | 5 | 2017 | Time-Frequency Matrix Feature Extraction and Classification of Environmental Audio Signals · IEEE Trans. Speech Audio Process. 2011 Combining Temporal Features by Local Binary Pattern for Acoustic Scene Classification · IEEE ACM Trans. Audio Speech Lang. Process. 2017 Audio Signal Feature Extraction and Classification Using Local Discriminant Bases · IEEE Trans. Speech Audio Process. 2007 |
Audio and music processing › audio classification
acoustic scene classification |
0.3 | 1 | 2017 | Combining Temporal Features by Local Binary Pattern for Acoustic Scene Classification · IEEE ACM Trans. Audio Speech Lang. Process. 2017 |
Image and video processing › texture analysis
local binary pattern |
0.3 | 1 | 2017 | Combining Temporal Features by Local Binary Pattern for Acoustic Scene Classification · IEEE ACM Trans. Audio Speech Lang. Process. 2017 |
Audio and music processing
audio classification |
0.3 | 3 | 2011 | Time-Frequency Matrix Feature Extraction and Classification of Environmental Audio Signals · IEEE Trans. Speech Audio Process. 2011 Audio Signal Feature Extraction and Classification Using Local Discriminant Bases · IEEE Trans. Speech Audio Process. 2007 Multigroup classification of audio signals using time-frequency parameters · IEEE Trans. Multim. 2005 |
Audio and music processing › audio classification
environmental sound classification |
0.1 | 1 | 2011 | Time-Frequency Matrix Feature Extraction and Classification of Environmental Audio Signals · IEEE Trans. Speech Audio Process. 2011 |
Digital forensics and information hiding
fingerprinting |
0.1 | 1 | 2010 | A wavelet-PCA-based fingerprinting scheme for peer-to-peer video file sharing · IEEE Trans. Inf. Forensics Secur. 2010 |
Digital forensics and information hiding
watermarking |
0.1 | 1 | 2010 | A wavelet-PCA-based fingerprinting scheme for peer-to-peer video file sharing · IEEE Trans. Inf. Forensics Secur. 2010 |
Audio and music processing › audio feature extraction
mel-frequency cepstral coefficients |
0.1 | 1 | 2017 | Combining Temporal Features by Local Binary Pattern for Acoustic Scene Classification · IEEE ACM Trans. Audio Speech Lang. Process. 2017 |
Audio and music processing › audio analysis › audio content analysis
audio fingerprinting |
0.1 | 1 | 2006 | Gaussian Mixture Modeling of Short-Time Fourier Transform Features for Audio Fingerprinting · IEEE Trans. Inf. Forensics Secur. 2006 |
Audio and music processing
time-frequency analysis |
0.1 | 1 | 2006 | A Robust Audio Watermark Representation Based on Linear Chirps · IEEE Trans. Multim. 2006 |
Digital forensics and information hiding › watermarking
audio watermarking |
0.1 | 1 | 2006 | A Robust Audio Watermark Representation Based on Linear Chirps · IEEE Trans. Multim. 2006 |
Audio and music processing › music information retrieval
music genre classification |
0.1 | 1 | 2005 | Multigroup classification of audio signals using time-frequency parameters · IEEE Trans. Multim. 2005 |
Audio and music processing › audio feature extraction
time-frequency features |
0.1 | 1 | 2005 | Multigroup classification of audio signals using time-frequency parameters · IEEE Trans. Multim. 2005 |
Image and video processing
wavelet transform |
0.0 | 1 | 2010 | A wavelet-PCA-based fingerprinting scheme for peer-to-peer video file sharing · IEEE Trans. Inf. Forensics Secur. 2010 |
Multimedia analysis and retrieval
audio retrieval |
0.0 | 1 | 2006 | Gaussian Mixture Modeling of Short-Time Fourier Transform Features for Audio Fingerprinting · IEEE Trans. Inf. Forensics Secur. 2006 |
Methods — techniques the papers use, named apart from their topics
joint time-frequency scattering transform · 0.8cross-validation · 0.8deep neural network · 0.6local binary pattern · 0.3ensemble classifier · 0.3mel-frequency cepstral coefficients · 0.1linear discriminant analysis · 0.1nonnegative matrix factorization · 0.1matching pursuit · 0.1MFCC · 0.1wavelet transform · 0.1principal component analysis · 0.1fingerprint embedding · 0.1linear chirps · 0.1line detection · 0.1hough-radon transform · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Multimodal Framework for Therapeutic ConsultationsabstractTherapeutic engagement between client and clinician is a key indicator in determining treatment outcomes for clients with mental health disorders. Quantifying this type of engagement provides an opportunity for the development of an engagement quantification framework for therapeutic efficacy, based on a number of data streams including, body movement and synchronicity, speech, and gestures to determine an individual's level of engagement. In this paper, we present a subset of such a framework through the quantification of engagement based on Facial Affect Recognition, Head Motion, and Natural Language Processing. We propose the use of semantic analysis, emotion dynamics and transitions, and head motion to describe a participant's attention over the consultation. For emotion dynamics and transitions we employ seven standard categorical emotions; for head motion we use acute and chronic head movement; and for semantic analysis we employ Robustly Optimized BERT Pretraining Approach. These features derive two engagement levels: low and high. We performed experiments on the AnnoMI dataset, which contains 133 therapeutic consultation videos for low and high quality motivational interviews, and compared the resulting engagement to the level of motivational interviewing. We achieved an 89.1% average accuracy for the Clinician model and an 81.1% average accuracy for the Client model using Gradient Boost as a classifier. Martin Ivanov, Alice Rueda, Venkat Bhat, Sridhar Krishnan 0001 |
IEEE Trans. Affect. Comput. | 4 |
| 2024 | Time-Frequency Scattergrams for Biomedical Audio Signal Representation and ClassificationabstractSpeech, music, and environmental sounds are the main forms of audio signals that are widely studied. There is a certain amount of texture present in every sound, and our human auditory system is not efficient in recognizing and classifying these audio textures present within the sounds. E.g. we are not able to distinguish between two sounds of a fire crackling or two sounds of water-falling. Hence, there is a need for a representation that could model these audio textures. These textures are also present in the audio signals that changes if pathological and pathomorphological conditions are present. To capture and analyze these audio textures, the audio signal is generally transformed into an intermediate time-frequency (t-f) representation such as spectrograms, Mel-spectrograms, and more. But recent studies have shown that joint time-frequency scattering transform is more suitable for classification problems than the standard time-frequency representations because of its inherent property of invariance and invertibility. In this paper, we have investigated the capacity of joint time-frequency scattergrams to capture the audio textures by analyzing the audio structures of biomedical sounds such as COVID-19 cough and breath sounds, pathological speech in children and adults, and infant cry sounds. Accuracy rates up to 96.40% for COVID-19 sounds, 95.50% for pathological speech in children, 94.10% for pathological speech in adults, and 97.30% for infant cry sounds have been achieved with 10-fold cross-validation. The proposed model of using scattergram provides an alternate and an efficient way of representing biomedical audio signals for machine learning applications. Karthikeyan Umapathy, Sridhar Krishnan 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2023 | Digital Phenotype Representation by Statistical, Information Theory, Data-Driven Approach with Digital Health DataabstractDigital phenotyping (DP) is a multidisciplinary field of science that quantifies the individual level phenotype through active and passive data. Although DP is a multidisciplinary field, there lacks a technical and a systematic approach to representing DP. This work proposes the development of digital phenotype profile (DPP) to represent a user’s physical and behavioural health baseline through systematic investigations with an emphasis on robustness and explainability. To achieve this, a Statistical, Information Theory, and Data-driven (SID) pipeline will develop the foundation of the DPP. SID evaluates the non-linearity of the signal to offer inference for domain-specific feature extraction, evaluates the information theory to rank the DPP parameters, and imputes missing data for robust analysis, respectively. SID was applied to a 24-hr Multi-Level dataset and was able to represent individual DPPs. The respective DPPs were visualized and clusters of awake and asleep were used for individual specific modelling. Binh P. Nguyen, Michael Nigro, Alice Rueda, Venkat Bhat, Sridhar Krishnan 0001 |
ICASSP | 5 |
| 2023 | SARdBScene: Dataset and Resnet Baseline for Audio Scene Source Counting and AnalysisabstractThis paper introduces a first of its kind dataset for audio scene analysis (ASA) and presents a baseline approach for audio source counting. SARdBScene is developed to promote research for audio source counting, as a relatively new ASA task, and present a comprehensive dataset that covers a variety of scenarios and audio-based tasks. It contains 80 hours of audio scene mixtures depicting four distinct environments with detailed annotations that make it a unique collection of curated data in the audio analysis landscape. Our baseline approach using ResNet establishes state-of-the-art results of 77.3% and 85.7% accuracy for audio source counting up to 12 sources and speaker counting up to 4 speakers, respectively. Michael Nigro, Sridhar Krishnan 0001 |
ICASSP | 2 |
| 2022 | Empirical Mode Decomposition articulation feature extraction on Parkinson's Diadochokinesia
Alice Rueda, Juan Camilo Vásquez-Correa, Juan Rafael Orozco-Arroyave, Elmar Nöth, Sridhar Krishnan 0001 |
Comput. Speech Lang. | 5 |
| 2022 | Multimodal System for Audio Scene Source Counting and AnalysisabstractAudio scene analysis (ASA) is a challenging and multifaceted task in audio signal processing that uncovers information about the nature of an audio recording. Regardless of the analysis goal, a number of audio sources are observed in any audio scene. However, this consideration is usually not explored or given considerable thought in research. This work aims to demonstrate the utility of audio source counting with a novel solution consisting of a multimodal system for ASA. Both speaker counting and sound event counting techniques use deep neural networks (DNN) to predict the number of sources. We are able to present competitive results for audio source counting by achieving prediction accuracy of 46.03% and 89.57% with a margin of error of$\pm 1$for speaker counting, which outperforms state-of-the-art systems for similar tasks. For sound event counting we achieve 50.55% and 86.59% prediction accuracy and accuracy with a margin of error of$\pm 1$, respectively, that establishes a clear baseline. Our system also demonstrates real-time aspects with an overall processing time of$\sim 0.4614$s per audio recording. Michael Nigro, Sridhar Krishnan 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2020 | Towards an Effective Motor Imagery Based-BCI with Calibration Through Activation of Central and Peripheral Mechanisms of Lower-LimbsabstractStroke is a neurological syndrome that may affect upper and lower limbs functions of post-stroke survivors. Brain-Computer Interfaces (BCIs) are becoming as a promising alter-native to help post-stroke patients rehabilitation, although there are very few associated studies and systems being applied in clinical environment. As a novelty, developing a motor imagery (MI) BCI based on pedal end-effector for motor rehabilitation, we propose to combine pedaling MI and passive pedaling into a Calibration phase. As a result, users would activate continuously their central and peripheral mechanisms linked to lower-limbs throughout BCI intervention. We hypothesize that this strategy enables to obtain a better classification model for our BCI by selecting those feature vectors corresponding to pedaling MI closer to real movements. Therefore, it is expected to have a more effective BCI intervention. Preliminary results show that the proposed method may increase the BCI performance. For almost all participants was noted, during MI tasks, a power decreasing over the foot area (Cz location), corresponding mainly to beta frequency bands, specifically for both low (13 to 22 Hz) and high (23 to 30 Hz) beta bands. Leticia Silva, Denis Delisle Rodríguez, Vivianne Cardoso, Dharmendra Gurve, Sridhar Krishnan 0001, Teodiano Freire Bastos-Filho |
SMC | 5 |
| 2020 | Separation of Fetal-ECG From Single-Channel Abdominal ECG Using Activation Scaled Non-Negative Matrix FactorizationabstractPerforming a fetal electrocardiogram (ECG) analysis, which contains important information about the status of a fetal, can help to detect fetus health even before birth. Since the fetal ECG extracted from the ECG signal recorded from the mother's abdomen, this extraction problem can be seen as a source separation problem, of recovering source signals from signal mixtures. In this paper, a method for separation of fetal ECG from abdominal ECG using activation scaled non-negative matrix factorization (NMF) is proposed. The performance of the proposed method is also compared with independent component analysis. The proposed method is tested under three different scenarios. First, the original abdominal ECG signal is used for fetal separation. Second, the recovered abdominal ECG after compression is used for separation. Third, the fetal ECG is extracted from the compressed domain of the abdominal ECG. We applied scaling on the activation matrix obtained using NMF for emphasizing the fetal ECG present in abdominal ECG. The improved-regularized least-squares [Formula: see text] algorithm is used for signal reconstruction, which provides better reconstruction quality and less processing time in comparison with other existing methods. The proposed algorithm is evaluated and tested on real abdominal recordings obtained from two different datasets from Physionet. The first dataset used for this paper is Silesia dataset for abdominal and direct f-ECG, and the second dataset we considered is Set-A of the Physionet challenge. The obtained outcomes reveal that it is possible to separate fetal ECG from single-channel abdominal ECG signal, which can help us to achieve energy-efficient transmission, and cost-effective fetal ECG remote monitoring for Internet-of-Things applications, where device battery and computational capacity are limited. Dharmendra Gurve, Sridhar Krishnan 0001 |
IEEE J. Biomed. Health Informatics | 2 |
| 2019 | Augmenting Dysphonia Voice Using Fourier-based Synchrosqueezing Transform for a CNN ClassifierabstractThe challenge of dysphonia voice studies is always the small dataset. It is difficult to apply more sophisticated deep learning techniques without overfitting or underfitting. Convolutional neural network (CNN) is a powerful classifier that requires a large amount of training data. Data augmentation techniques for voice are limited. Fourier-based synchrosqueezing transform (FSST) can be used as a data augmentation technique to increase the data size. The results indicated that not only can FSST increase the data size, the CNN can also learn better with FSST than with Short-Time Fourier Transform (STFT) power spectrum. The loss function for FSST converges, but not for STFT. FSST is also more stable and provides more accurate results. Alice Rueda, Sridhar Krishnan 0001 |
ICASSP | 2 |
| 2019 | Feature Representation of Pathophysiology of Parkinsonian Dysarthria
Alice Rueda, Juan Camilo Vásquez-Correa, Cristian D. Ríos-Urrego, Juan Rafael Orozco-Arroyave, Sridhar Krishnan 0001, Elmar Nöth |
INTERSPEECH | 5 |
| 2018 | First-Order Difference Energy Regularization for Enhancing Reconstruction Performance in Compressive Sensing of Foot-Gait SignalsabstractA new method for the regularization of the objective function for the reconstruction of the foot-gait signal from compressively sensed measurements is proposed. The method is based on using the l2 norm of the first-order difference to regularize the objective function. The state-of-the-art first-order difference sparsity promoting algorithms can introduce transient artefacts in the signal. The proposed regularization helps to reduce such artefacts. Involved optimization can be solved by using a sequential optimization procedure. The resulting algorithm is useful for enhancing the quality of reconstructed signal, especially in the situations when the CS system is applied with extremely high compression ratio. Simulation results indicate that the proposed method can offer upto 2.81dB improvement in signal-to-noise ratio, 0.02 units improvement in structural similarity measure, and a marginal increase in the computational effort. Jeevan K. Pant, Sridhar Krishnan 0001 |
ICASSP | 2 |
| 2018 | Sparse Signal Reconstruction Using Multi-Sequential Lp OptimizationabstractAn improved nonconvex optimization based algorithm for the reconstruction of sparse signals for compressive sensing is proposed. The algorithm is based on minimizing nonconvex Lp pseudonorm by using two multiple sequences of sub-optimizations. If the two iterates in between any two sub-optimizations happen to be very similar, one of the iterates is merged with the other. The merged iterate is then set to another randomly chosen vector that is orthogonal to the iterate obtained after merging. The multiple sequences of sub-optimizations can be implemented to run simultaneously in two cores of a multicore computer. Simulation results are presented which indicate that the proposed algorithm can offer improvement in the percentage of perfect signal reconstructions by upto 7% and reduction in the CPU time required to run the algorithm by upto 11%, relative to the competing algorithms. Jeevan K. Pant, Sridhar Krishnan 0001 |
ISCAS | 2 |
| 2017 | Two-pass ℓp-regularized least-squares algorithm for compressive sensingabstractA two-pass algorithm for signal reconstruction in compressive sensing (CS) is proposed. It is based on using a new regularization in the objective function which elevates functional value at a previously obtained optimal or near-optimal point. Elevation in the objective function causes the optimization to converge to a new solution which would be optimal or near-optimal. Either previously obtained solution or the newly obtained solution is selected as the final solution based on which one yields lower value of the objective function. This algorithm is suitable for nonconvex optimization-based sparse signal reconstruction in CS. Simulation results are presented which indicate that the proposed algorithm is effective for not only improving the percentage of perfect reconstructions from noiseless measurements by upto 3.4% but also offering similar performance improvement for the reconstruction from noisy measurements. Jeevan K. Pant, Sridhar Krishnan 0001 |
ISCAS | 2 |
| 2017 | Combining Temporal Features by Local Binary Pattern for Acoustic Scene ClassificationabstractThe popular frequency-domain features Mel-frequency cepstral coefficients (MFCCs) have been widely used for the task of acoustic scene classification (ASC). The MFCC feature vector describes only the power spectral envelope of a single frame, but it seems like environmental audio signal would benefit from information in the temporal dynamics. However, the classic approach of integrating them would lose this important information. Here, we adopt local binary pattern (LBP) as a tool to characterize the latent information on the temporal dynamics. The frame-level MFCC features are viewed as a 2-D image, where we use LBP to encode the evolution process. Besides, some complementary spectral features such as spectral centroid (SC), spectral bandwidth (SBW) is utilized to further improve the ASC performance. The proposed features are then fed into an ensemble classifier called D3C for recognizing environmental sounds. The results show that the proposed method was able to achieve a classification improvement of 8% compared to the baseline system. Our work presented a new method for combing the temporal features, demonstrating the significance of the temporal evolution features for characterizing the environmental sound. Sridhar Krishnan 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 2 |
| 2016 | Compact and robust video fingerprinting using sparse represented featuresabstractIn this paper, we propose a compact and robust video fingerprinting scheme by using sparse represented features (SRF). The SRF are extracted by a two dimensional matching pursuit decomposition (2D-MPD) method. The motivation of using sparse features is that the sparse coding method can significantly reduce the data dimensionality and effectively retain the structure of the images. To further reduce the length of the fingerprint, a two-stage cascade SVD method is applied. The SVD-based feature extraction method can improve the robustness to certain attacks, such as geometric attack. Then, a locally adaptive quantization (AQ) method which considers the local probability distribution of the sample is applied. This method can quantize the real-valued fingerprints into binary bits without degrading the detection performance too much. At last, a secret key based interleaving method is applied to the binary fingerprints. The interleaving can enlarge the Hamming distance of the fingerprints, so the detection performance is improved. According to the experimental results, the proposed method offers a favourable robustness versus discriminability tradeoff over the state-of-the-art video fingerprint methods. Bo Wu 0016, Sridhar Krishnan 0001, Nan Zhang 0015, Li Su 0003 |
ICME | 2 |
| 2015 | Learning Dictionary Via Wavelet Sparse Principal Component Analysis
Shengkun Xie, Anna T. Lawniczak, Sridhar Krishnan 0001 |
ICPRAM (1) | 3 |
| 2014 | Compressive sensing of ECG signals based on mixed pseudonorm of the first- and second-order differencesabstractAn improved algorithm for the reconstruction of electrocardiogram signals in compressive sensing is proposed. The algorithm is based on the minimization of a mixed pseudonorm of first- and second-order differences of the signal. Locations of QRS segments are estimated using a technique based on signal derivatives and the Hilbert transform, and they are used to implement the mixed pseudonorm. Simulation results demonstrate that the proposed algorithm offers approximately 23.5%, 11.4%, 4.4%, and 2.1% improvement in signal-to-noise ratio for a compression ratio of 90%, 80%, 70%, and 60%, respectively, relative to several competitive state-of-the-art algorithms. Jeevan K. Pant, Sridhar Krishnan 0001 |
ICASSP | 2 |
| 2014 | LibD3C: Ensemble classifiers with a clustering and dynamic selection strategy
Chen Lin 0001, Sridhar Krishnan 0001, Quan Zou 0001 |
Neurocomputing | 5 |
| 2013 | A variation of empirical mode decomposition with intelligent peak selection in short time windowsabstractThis paper describes analysis of the behaviour, and establishment of different decomposition properties, of a previously presented modification of the empirical mode decomposition algorithm using fractional Gaussian noise. Importantly, the modified algorithm, called empirical mode decomposition-modified peak selection (EMDMPS), is used to explain certain aspects of the decomposition behaviour of EMD, providing novel insight into the domain. Finally, the utility of EMD-MPS is demonstrated by using it for a novel time-scale based de-trending of signals, using real-world financial time-series as an example. Muhammad Kaleem, Aziz Guergachi, Sridhar Krishnan 0001 |
ICASSP | 3 |
| 2013 | Reconstruction of ECG signals for compressive sensing by promoting sparsity on the gradientabstractA new algorithm for the reconstruction of signals in compressive sensing framework is proposed. The algorithm is based on a least-squares method which incorporates a regularization to promote sparsity on the gradient of the signal. It uses a sequential basic conjugate-gradient method, and it is especially suited for the reconstruction of signals which exhibit temporal correlation, e.g., electrocardiogram (ECG) signals. Simulation results are presented which demonstrate that the proposed algorithm yields upto 80.28% reduction in mean square error and from 49.95% to 65.64% reduction in the required amount of computation, relative to the state-of-the-art block sparse Bayesian learning bound-optimization algorithm. Jeevan K. Pant, Sridhar Krishnan 0001 |
ICASSP | 2 |
| 2013 | Sparse Signal Analysis Using Ramanujan Sums
Guangyi Chen 0001, Sridhar Krishnan 0001, Wen-Fang Xie |
ICIC (2) | 2 |
| 2013 | Illumination Invariant Face Recognition
Guangyi Chen 0001, Sridhar Krishnan 0001, Yongjia Zhao, Wen-Fang Xie |
ICIC (1) | 2 |
| 2013 | Circular Projection for Pattern Recognition
Guangyi Chen 0001, Tien D. Bui, Sridhar Krishnan 0001, Shuling Dai |
ISNN (1) | 3 |
| 2013 | Noise Effects on Spatial Pattern Data Classification Using Wavelet Kernel PCA - A Monte Carlo Simulation Study
Shengkun Xie, Anna T. Lawniczak, Sridhar Krishnan 0001 |
ISNN (1) | 3 |
| 2013 | Visual saliency's modulatory effect on just noticeable distortion profile and its application in image watermarking
Yaqing Niu, Matthew J. Kyan, Azeddine Beghdadi, Sridhar Krishnan 0001 |
Signal Process. Image Commun. | 5 |
| 2013 | Matrix-Based Ramanujan-Sums TransformsabstractIn this letter, we study the Ramanujan Sums (RS) transform by means of matrix multiplication. The RS are orthogonal in nature and therefore offer excellent energy conservation capability. The 1-D and 2-D forward RS transforms are easy to calculate, but their inverse transforms are not defined in the literature for non-even function$ ({\rm mod}~ {\rm M}) $. We solved this problem by using matrix multiplication in this letter. Guangyi Chen 0001, Sridhar Krishnan 0001, Tien D. Bui |
IEEE Signal Process. Lett. | 2 |
| 2012 | Sparse principal component extraction and classification of long-term biomedical signalsabstractThis article focuses on finding a solution of sparse representation for signal classification in long-term observational studies. An approach that involves sparse principal component analysis (SPCA) is proposed. This method first uses a non-overlapping moving window for signal segmentation and makes use of SPCA to select a limited number of signal segments for constructing sparse principal components. A set of supervised predictive models based on sparse principal components of training signal segments is then constructed for signal approximation. Within this approach, their model residuals are estimated and used for signal classification. A nearly perfect classification accuracy is obtained for both the synthetic data and EEG signals that we considered. This highly positive result suggests that the proposed method may be useful for automatic event detection in long-term observational signals. Shengkun Xie, Sridhar Krishnan 0001, Anna T. Lawniczak |
CBMS | 2 |
| 2012 | Analysis of communication network surveillance using functional ANOVA model with unequal variancesabstractAnalysis of simulation result plays an important role in helping decision-making when simulation modeling is used as an approach to investigate complex systems. Analysis of variance (ANOVA) is a popular statistical technique for analyzing simulation results when various designs of simulation experiments are implemented and investigated. However, conventional ANOVA technique is based on the assumption of homogeneous variances of output data, which may not be realistic for data coming from the real-world complex systems. The existence of heterogeneous structures of data coming from complex systems requires a more reliable analysis method, in order to analyze such data. In this paper, a functional ANOVA model with unequal variances is proposed for meeting this goal. The proposed model aims to better capture the heterogeneous data structure that is caused by various experimental simulation setups. The applicability of the proposed method is illustrated by using simulated communication network traffic data. Our work contributes to the development of new approaches for analysis of simulation and real-world complex systems data with heterogeneous structures. The proposed method can be useful for defense problems, e.g. for analysis of communication network surveillance or for the purpose of improving the operational efficiency. Shengkun Xie, Sridhar Krishnan 0001, Anna T. Lawniczak |
CISDA | 2 |
| 2012 | Log-frequency spectrogram for respiratory sound monitoringabstractComputerized patient monitoring provides valuable information on clinical disorders in medical practice, and it triggers the need to simplify the extent of resources required to describe large set of complex biomedical signals. In this paper, we present a new signal quantification method based on block-wise similarity measurement between the neighboring regions in the optimized log-frequency spectrogram of audio signals. Low dimensional cepstral feature set for signal quantification is then formed from the reconstructed similarity matrix using 2D principal component analysis. The effectiveness of the method is verified with real respiratory sound (RS) signals for the purpose of abnormal RS detection towards RS monitoring. Unlike conventional pathological RS detection methods which extract features from well-segmented inspiratory/expiratory phase segments, the proposed scheme is able to perform fast detection of various types of abnormality for unsegmented signals. Farook Sattar, Sridhar Krishnan 0001 |
ICASSP | 3 |
| 2012 | Time-Frequency Analysis via Ramanujan SumsabstractResearch in signal processing shows that a variety of transforms have been introduced to map the data from the original space into the feature space, in order to efficiently analyze a signal. These techniques differ in their basis functions, that is used for projecting the signal into a higher dimensional space. One of the widely used schemes for quasi-stationary and non-stationary signals is the time-frequency (TF) transforms, characterized by specific kernel functions. This work introduces a novel class of Ramanujan Fourier Transform (RFT) based TF transform functions, constituted by Ramanujan sums (RS) basis. The proposed special class of transforms offer high immunity to noise interference, since the computation is carried out only on co-resonant components, during analysis of signals. Further, we also provide a 2-D formulation of the RFT function. Experimental validation using synthetic examples, indicates that this technique shows potential for obtaining relatively sparse TF-equivalent representation and can be optimized for characterization of certain real-life signals. Lakshmi Sugavaneswaran, Shengkun Xie, Karthikeyan Umapathy, Sridhar Krishnan 0001 |
IEEE Signal Process. Lett. | 4 |
| 2011 | Automatic respiratory sound classification using temporal-spectral dominanceabstractRespiratory sound (RS) signals carry significant information about the underlying functioning of the pulmonary system. Auscultation based diagnosis of pulmonary disorders relies on the presence of adventitious sounds. This paper proposes a new method for automatic RS classification based on instantaneous frequency (IF) analysis with the aim to identify various types of pathological RS. The presented method produces a high definition representation of RS signals in the time-frequency (TF) plane. The discarded phase information in spectrogram has been adopted here for the computation of IF and the subsequent temporal-spectral dominance. A new set of features have been extracted to quantify the shapes of the obtained individual TF contour and therefore strongly enhances the identification of multi-components signals such as polyphonic wheezes. An overall accuracy of 92.7 ± 2.9% on real RS recordings shows the promising performance by the presented method. Farook Sattar, Sridhar Krishnan 0001 |
ICME | 3 |
| 2011 | Combining least-squares support vector machines for classification of biomedical signals: a case study with knee-joint vibroarthrographic signalsabstractThe knee-joint vibroarthrographic (VAG) signal could be used as an indicator with regard to the degenerative articular cartilage surfaces of the knee. Computer-aided analysis of VAG signals could provide quantitative indices for the noninvasive diagnosis of knee-joint pathologies at different stages. In this article, we propose a novel multiple classifier system (MCS) based on a recurrent neural network (RNN), to classify a dataset of 89 knee-joint VAG signals. The MCS consists of a group of component classifiers in the form of the least-squares support vector machine. The knowledge generated by the component classifiers is combined with the linear and normalised fusion model, the weights of which are optimised during the energy convergence process of the RNN. The experimental results showed that the proposed MCS was able to provide the classification accuracy of 80.9% and the area of 0.9484 under the receiver operating characteristics curve. The diagnostic performance of the MCS was superior to that obtained with the prevailing fusion approaches, such as the majority vote, the simple average and the median average. Sridhar Krishnan 0001 |
J. Exp. Theor. Artif. Intell. | 2 |
| 2011 | Time-Frequency Matrix Feature Extraction and Classification of Environmental Audio SignalsabstractAudio feature extraction and classification are important tools for audio signal analysis in many applications, such as multimedia indexing and retrieval, and auditory scene analysis. However, due to the nonstationarities and discontinuities exist in these signals, their quantification and classification remains a formidable challenge. In this paper, we develop a new approach for audio feature extraction to effectively quantify these nonstationarities in an attempt to achieve high classification accuracy for environmental audio signals. Our approach consists of three stages: first we propose to construct the time-frequency matrix (TFM) of audio signals using matching-pursuit time-frequency distribution (MP-TFD) technique, and then apply the non-negative matrix decomposition (NMF) technique to decompose the TFM into its significant components. Finally, we propose seven novel features from the spectral and temporal structures of the decomposed vectors in a way that they successfully represent joint TF structure of the audio signal, and combine them with the Mel-frequency cepstral coefficients (MFCCs) features. These features are examined using a database of 192 environmental audio signals which includes 20 aircraft, 17 helicopter, 20 drum, 15 flute, 20 piano, 20 animal, 20 bird, and 20 insect sounds, and the speech of 20 males and 20 females. The results of the numerical simulation support the effectiveness of the proposed approach for environmental audio classification with over 10% accuracy-rate improvement compared to the MFCC features. Behnaz Ghoraani, Sridhar Krishnan 0001 |
IEEE Trans. Speech Audio Process. | 2 |
| 2010 | Discriminative base decomposition for time-frequency matrix decompositionabstractTime-frequency matrix (TFM) decomposition using non-negative matrix factorization (NMF) has been recently considered as a successful tool for time-frequency (TF) quantification. In this paper, we modify the constraints of traditional cost function of NMF to make the method a better fit for TF quantification, and denote the new method with NMF discriminant base (NMFDB) decomposition. We evaluate the proposed method, and show that it successfully identifies the discriminant bases. Additionally, we measure the discrimination ability of NMFDB over the signals with very low discriminations, and compare it with the discrimination of the decomposed bases derived using traditional NMF. It is concluded that the proposed method is able to locate the region of difference with 20% better performance compared to the conventional NMF. Behnaz Ghoraani, Sridhar Krishnan 0001 |
ICASSP | 2 |
| 2010 | Combined just noticeable difference model guided image watermarkingabstractPerceptual Watermarking should take full advantage of the results from human visual system (HVS) studies. Just noticeable difference (JND), which refers to the maximum difference that the HVS does not perceive, gives us a way to model the HVS accurately. In this paper, we exploit a combined JND model which represents additional accurate perceptual visibility threshold profile to guide watermarking for digital images. The proposed combined JND model guided watermarking scheme, where visual models are fully used to determine image dependent upper bounds on watermark insertion, allows us to provide the maximum strength transparent watermark. Experimental results confirm the improved performance of our combined JND model. Our combined JND model is capable of yielding higher injected-watermark energy without introducing noticeable difference to the original image and outperforms the relevant existing visual models. Robustness results show the proposed JND model guided watermarking scheme performs much better than other algorithms based on Watson's perceptual model. Yaqing Niu, Sridhar Krishnan 0001, Qin Zhang 0009 |
ICME | 3 |
| 2010 | Towards robust speech-based emotion recognitionabstractMaintaining the robustness of a speech processing system in the presence of noise is a challenge. This paper shows frequency subband architecture can improve robustness of an emotion recognition system when signals are corrupted by selective noise. This paper also demonstrates that feature selection based on some mutual-information criterion (maximum-relevant minimum-redundancy) can give us the most effective subset of features to get a better result. Talieh Seyed Tabatabaei, Sridhar Krishnan 0001 |
SMC | 2 |
| 2010 | SVM-based classification of digital modulation signalsabstractModulation recognition systems have to be able to correctly classify the incoming signal's modulation scheme in the presence of noise. This paper addresses the problem of automatic modulation recognition of digital communication signals using support vector machines (SVM). Three digital modulation schemes have been considered and four features have been used as inputs to the SVM. A fuzzy multi-class classification method has been proposed and the overall accuracy of 77.0% at signal-to-noise ratio (SNR) of 10dB has been achieved. Talieh Seyed Tabatabaei, Sridhar Krishnan 0001, Alagan Anpalagan |
SMC | 2 |
| 2010 | A wavelet-PCA-based fingerprinting scheme for peer-to-peer video file sharingabstractIn order to utilize peer-to-peer (P2P) networks in legal content distribution to benefit the legal content providers, copyright protection needs to be enhanced. In this paper, a fingerprint generation and embedding method is proposed for complex P2P file sharing networks. In this method, wavelet and principal component analysis (PCA) techniques are used for fingerprint generation. First, the wavelet technique obtains a low-frequency representation of the test image (or source file, which is assumed to be one I frame of a video with a DVD quality) and PCA finds the features of the representation. Then, a set of fingerprint matrices can be created based on a proposed algorithm. Finally, each matrix combines with the low-frequency representative to become a unique fingerprinted matrix. The fingerprinted matrix is not only much smaller than the original image in size but also contains the most important information. Without this information, the quality of the reconstructed image will be very poor. Thus, the fingerprinted file is more suitable for distribution in P2P networks, because, in the distribution stage, the uniquely fingerprinted matrix will only be dispensed by the source host and leave the rest for P2P networks to handle. On the other hand, among other frames of the same video which are not decomposed, some will be embedded with sharable fingerprints. The relationship between unique fingerprint and sharable fingerprint and the purpose of using it will be discussed in the paper. Our result indicates that the proposed fingerprint has shown strong robustness against common attacks such as Gaussian noise, median filter, and lossy compression. Xiaoli Li 0004, Sridhar Krishnan 0001, Ngok-Wah Ma |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2009 | Identifying the potential for Failure of Businesses in the Technology, Pharmaceutical and Banking Sectors using Kernel-based Machine Learning MethodsabstractThe objective of this paper is to analyze the performance of a kernel-based method in identifying the potential for collapse (or survival) of a firm operating in three different sectors of the economy-Technology, Pharmaceutical and Banking. The analysis uses the actual stock market data, collected on a weekly basis in a common time-series interval for the active and dead companies in each of the three sectors. The basic idea is to apply the concept of Fisher kernels and visualization to reduce the data from a time-series format to two-dimensional plots that can be visually inspected and potentially segregate the 'collapse' class from the 'survival' one. From our experiments we observe that our method fits well for the Technology and Banking sectors, but is not able to provide a visually clear classification for the Pharmaceuticals sector. Depending on the range of data we use as input, and its distribution, the classification pattern varies from an ideally separable case to a non separable one, in a two dimensional feature space. Yashodhan Rajiv Athavale, Aziz Guergachi, Pouyan Hosseinizadeh, Sridhar Krishnan 0001 |
SMC | 4 |
| 2008 | An audio watermarking method based on molecular matching pursuitabstractIn this paper we introduce a new watermarking model combining a joint time frequency (TF) representation using the molecular matching pursuit (MMP) algorithm and a psychoacoustic model. We take advantage of the notion of structure of the signal introduced by the MMP to get a precise representation of audio signals, and then by using a psychoacoustic model we can embed a watermark efficiently on the signal. By selecting atoms of TF components that are not perceptible by the human ear we ensure the security and imperceptibility of the watermark. Then by judicious selection of the watermark host spots we ensure the robustness of the watermark to main kind of signal attacks, including lossy compression. The robustness of the proposed method proves the potential of joint TF representation techniques as viable watermarking schemes. Mathieu Parvaix, Sridhar Krishnan 0001, Cornel Ioana |
ICASSP | 2 |
| 2008 | Gaussian Mixture Modeling of Keystroke Patterns for Biometric ApplicationsabstractThe keystroke patterns produced during typing have been shown to be unique biometric signatures. Therefore, these patterns can be used asdigitalsignaturesto verify the identity of computer users remotely over the Internet or locally at a specific workstation. In particular, keystroke recognition can enhance the username and password security model by monitoring the way that these strings are typed. To this end, this paper proposes a novel up--up keystroke latency (UUKL) feature and compares its performance with existing features using a Gaussian mixture model (GMM)-based verification system that utilizes an adaptive and user-specific threshold based on the leave-one-out method (LOOM). The results show that the UUKL feature significantly outperforms the commonly used key hold-down time (KD) and down--down keystroke latency (DDKL) features. Overall, the inclusion of the UUKL feature led to an equal error rate (EER) of 4.4% based on a database of 41 users, which is a 2.1% improvement as compared to the existing features. Comprehensive results are also presented for a two-stage authentication system that has shown significant benefits. Lastly, due to many inconsistencies in previous works, a formal keystroke protocol is recommended that consolidates a number of parameters concerning how to improve performance, reliability, and accuracy of keystroke-recognition systems. Danoush Hosseinzadeh, Sridhar Krishnan 0001 |
IEEE Trans. Syst. Man Cybern. Part C | 2 |
| 2007 | A Watermarking Method for Speech Signals Based on the Time-Warping Signal Processing ConceptabstractThis paper deals with the watermarking of audio speech signals which consists in introducing an imperceptible mark in a signal. To this end, we suggest 10 use an amplitude modulated signal that mimics a formantic structure present in the signal. This allows to exploit the time-masking effect occurring when two signals are close in the time-frequency plane. From this embedding scheme, a watermark extraction method based on nonstationary linear filtering and matched filler detection is proposed in order to recover information carried by the watermark. Numerical results conducted on a real speech signal show that the watermark is likely not hearable and informations carried by the watermark are easily retrievable. Cornel Ioana, Arnaud Jarrot, André Quinquis, Sridhar Krishnan 0001 |
ICASSP (2) | 4 |
| 2007 | Interference Detection in Spread Spectrum Communication Using Polynomial Phase TransformabstractWe propose an interference detection technique for detecting time varying jamming signals in spread spectrum communication systems. The technique is based on discrete polynomial phase transform (DPPT), where the jamming signal is synthesized from the modulated spread spectrum signal using the DPPT. The technique has shown good performance under low interference conditions with 2dB SJR, when correlation coefficient between the synthesized chirp signal and the reference chirp is 0.9. The computational complexity of the proposed technique is low compared to other techniques such as Hough-Radon transform. This interference detection technique can be applied for different interference excision methods in military and wireless communication applications. Randa Zarifeh, Nandini Alinier, Sridhar Krishnan 0001, Alagan Anpalagan |
ICC | 3 |
| 2007 | Emotion Recognition Using Novel Speech Signal FeaturesabstractAutomatic emotion recognition (AER) is a very recent research topic in the human-computer interaction (HCI) field which still has much room to grow. In this contribution a set of novel acoustic features and least square-support vector machines (LS-SVMs) are proposed to set up a speaker-independent automatic human emotion recognition system. Six discrete emotional states are classified throughout this work: happiness, sadness, anger, surprise, fear, and disgust. Different multi-class SVM methods are implemented in order to get the best result. The result achieved by LS-SVM is then compared by that of a linear classifier. We achieved an overall accuracy of 81.3%. Talieh Seyed Tabatabaei, Sridhar Krishnan 0001, Aziz Guergachi |
ISCAS | 2 |
| 2007 | Combining Vocal Source and MFCC Features for Enhanced Speaker Recognition Performance Using GMMsabstractThis work presents seven novel spectral features for speaker recognition. These features are the spectral centroid (SC), spectral bandwidth (SBW), spectral band energy (SBE), spectral crest factor (SCF), spectral flatness measure (SFM), Shannon entropy (SE) and Renyi entropy (RE). The proposed spectral features can quantify some of the characteristics of the vocal source or the excitation component of speech. This is useful for speaker recognition since vocal source information is known to be complementary to the vocal tract transfer function, which is usually obtained using the Mel frequency cepstral coefficients (MFCC) or linear predication cepstral coefficients (LPCC). To evaluate the performance of the spectral features, experiments were performed using a text-independent cohort Gaussian mixture model (GMM) speaker identification system. Based on 623 users from the TIMIT database, the spectral features achieved an identification accuracy of 99.33% when combined with the MFCC based features and when using undistorted speech. This represents a 4.03% improvement over the baseline system trained with only MFCC and ΔMFCC features. Danoush Hosseinzadeh, Sridhar Krishnan 0001 |
MMSP | 2 |
| 2007 | Chaotic time series prediction using knowledge based Green's Kernel and least-squares support vector machinesabstractThis paper proposes a novel prior knowledge based Green's kernel for long term chaotic time series prediction. A mathematical framework is presented to obtain the domain knowledge about the magnitude of the Fourier transform of the function to be predicted and design a prior knowledge based Green's kernel that exhibits optimal regularization properties by using the concept of matched filters. The matched filter behavior of the proposed kernel function provides the optimal regularization. Simulation results on a chaotic benchmark time series indicate that the knowledge based Green's kernel shows good prediction performance compared to the other existing support vector kernels for the time series prediction task considered in this paper. Tahir Farooq, Aziz Guergachi, Sridhar Krishnan 0001 |
SMC | 3 |
| 2007 | Audio Signal Feature Extraction and Classification Using Local Discriminant BasesabstractAudio feature extraction plays an important role in analyzing and characterizing audio content. Auditory scene analysis, content-based retrieval, indexing, and fingerprinting of audio are few of the applications that require efficient feature extraction. The key to extract strong features that characterize the complex nature of audio signals is to identify their discriminatory subspaces. In this paper, we propose an audio feature extraction and a multigroup classification scheme that focuses on identifying discriminatory time-frequency subspaces using the local discriminant bases (LDB) technique. Two dissimilarity measures were used in the process of selecting the LDB nodes and extracting features from them. The extracted features were then fed to a linear discriminant analysis-based classifier for a three-level hierarchical classification of audio signals into ten classes. In the first level, the audio signals were grouped into artificial and natural sounds. Each of the first level groups were subdivided to form the second level groups viz. instrumental, automobile, human, and nonhuman sounds. The third level was formed by subdividing the four groups of the second level into the final ten groups (drums, flute, piano, aircraft, helicopter, male, female, animals, birds and insects). A database of 213 audio signals were used in this study and an average classification accuracy of 83% for the first level (113 artificial and 100 natural sounds), 92% for the second level (73 instrumental and 40 automobile sounds; 40 human and 60 nonhuman sounds), and 89% for the third level (27 drums, 15 flute, and 31 piano sounds; 23 aircraft and 17 helicopter sounds; 20 male and 20 female speech; 20 animals, 20 birds and 20 insects sounds) were achieved. In addition to the above, a separate classification was also performed combining the LDB features with the mel-frequency cepstral coefficients. The average classification accuracies achieved using the combined features were 91% for the first level, 99% for the second level, and 95% for the third level Karthikeyan Umapathy, Sridhar Krishnan 0001, R. K. Rao |
IEEE Trans. Speech Audio Process. | 2 |
| 2006 | Keystroke Identification Based on Gaussian Mixture ModelsabstractMany computer systems rely on the username and password model to authenticate users. This method is widely used, yet it can be highly insecure if a user's login information has been compromised. To increase security, some authors have proposed keystroke patterns as a biometric tool for user authentication; they can be used to recognize users based on how they type. This paper introduces a novel method that applies GMMs to keystroke identification. The major benefit of this method is the ability to update the user's model each time he or she is authenticated. Therefore, as time goes on, each user model accurately reflects the changes in that user's keystroke pattern. Using this method, a FAR and a FRR rate of approximately 2% was achieved. However, it should be noted that 50% of the test subjects were the traditional "two finger" typists and therefore, this had a disproportionately negative impact on the results Danoush Hosseinzadeh, Sridhar Krishnan 0001, April Khademi |
ICASSP (3) | 2 |
| 2006 | Soccer Video Retrival Using Adaptive Time-Frequency MethodsabstractThe retrieval of soccer highlights is a suitable technique for video indexing, required by the multimedia database management or for the development of television on demand. For these purposes, it should be interesting to have an automatic annotation of events happened in soccer games. One solution consists in analyzing the audio soundtrack associated to the soccer video and to detect the interesting frames. In this paper we use the adaptive time-frequency decomposition of the soundtrack as a feature extraction procedure. This decomposition is based on the Matching Pursuit concept and a dictionary composed of Gabor functions. The parameters provided by these transformations constitute the input of the classification stage. The results provided for real soccer video will prove the efficiency of the adaptive time-frequency representation as a feature extraction stage. Jonathan Marchal, Cornel Ioana, Emanuel Radoi, André Quinquis, Sridhar Krishnan 0001 |
ICASSP (5) | 5 |
| 2006 | Computational Intellegence Techniques and their Applications in Content-Based Image RetrievalabstractThe main focus of this paper is to present a methodology for optimizing relevance identification in content-based image retrieval (CBIR) systems through the principle of feature weight detection. The purpose of relevance identification is to find a collection of images that are statistically similar to, or match with, an original query image within a large visual database. The novelty of this scheme is two-fold: using a base-10 genetic algorithm method to accurately determine the contribution of individual feature vectors for a successful retrieval in the so-called feature weight detection process, and defining a new unsupervised learning algorithm, the directed self-organizing tree map (DSOTM), for the purpose of classification in the automatic relevance identification module of the search engine. Comprehensive experiments demonstrate feasibility of the proposed methodology Kambiz Jarrah, Matthew J. Kyan, Sridhar Krishnan 0001, Ling Guan |
ICME | 3 |
| 2006 | Discrete Polynomial Transform for Digital Imagewatermarking ApplicationabstractIn this study, we propose a new way to detect the image watermark messages modulated as linear chirp signals. The spread spectrum image watermarking algorithm embeds linear chirps as watermark messages. The phase of the chirp represents watermark message such that each phase corresponds to a different message. We extract the watermark message using a phase detection algorithm based on discrete polynomial phase transform (DPT). The DPT models the signal as polynomial and uses ambiguity function to estimate the signal parameters. The proposed method not only detects the presence of watermark, but also extracts the embedded watermark bits and ensures the message is received correctly. The robustness of the proposed detection scheme has been evaluated using checkmark benchmark attacks, and we found a guaranteed maximum bit error rate of 15%, which watermark message is correctly detected using DPT Lam Le, Sridhar Krishnan 0001, Behnaz Ghoraani |
ICME | 2 |
| 2006 | Automatic Content-Based Image Retrieval Using Hierarchical Clustering AlgorithmsabstractThe overall objective of this paper is to present a methodology for guiding adaptations of an RBF based relevance feedback network, embedded in automatic content-based image retrieval (CBIR) systems, through the principle of unsupervised hierarchical clustering. The self organizing tree map (SOTM) is essentially attractive for our approach since it not only extracts global intuition from an input pattern space but also injects some degree of localization into the discriminative process such that maximal discrimination becomes a priority at any given resolution. The main focus of this paper is two-fold: introducing a new member of SOTM family, the Directed SOTM (DSOTM) that not only provides a partial supervision on duster generation by forcing divisions away from the query class, but also presents a flexible verdict on resemblance of the input pattern as its tree structure grows; and modifying the current structure of the normalised graph cuts (Ncut) process by enabling the algorithm to determine appropriate number of clusters within an unknown dataset prior to its recursive clustering scheme through the principle of self-organizing normalized graph cuts (SONcut). Comprehensive comparisons with the Self-Organizing feature Map (SOFM), SOTM, and Ncut algorithms demonstrate feasibility of the proposed methods. Kambiz Jarrah, Sridhar Krishnan 0001, Ling Guan |
IJCNN | 2 |
| 2006 | Gaussian Mixture Modeling of Short-Time Fourier Transform Features for Audio FingerprintingabstractIn audio fingerprinting, an audio clip must be recognized by matching an extracted fingerprint to a database of previously computed fingerprints. The fingerprints should reduce the dimensionality of the input significantly, provide discrimination among different audio clips, and, at the same time, be invariant to distorted versions of the same audio clip. In this paper, we design fingerprints addressing the above issues by modeling an audio clip by Gaussian mixture models (GMM). We evaluate the performance of many easy-to-compute short-time Fourier transform features, such as Shannon entropy, Renyi entropy, spectral centroid, spectral bandwidth, spectral flatness measure, spectral crest factor, and Mel-frequency cepstral coefficients in modeling audio clips using GMM for fingerprinting. We test the robustness of the fingerprints under a large number of distortions. To make the system robust, we use some of the distorted versions of the audio for training. However, we show that the audio fingerprints modeled using GMM are not only robust to the distortions used in training but also to distortions not used in training. Among the features tested, spectral centroid performs best with an identification rate of 99.2% at a false positive rate of 10-4. All of the features give an identification rate of more than 90% at a false positive rate of 10-3 Arunan Ramalingam, Sridhar Krishnan 0001 |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2006 | A Robust Audio Watermark Representation Based on Linear ChirpsabstractIn this paper, we introduce a novel watermark representation for audio watermarking, where we embed linear chirps as watermark signals. Different chirp rates, i.e., slopes on the time–frequency (TF) plane, represent watermark messages such that each slope corresponds to a unique message. These watermark signals, i.e., linear chirps, are embedded and extracted using an existing watermarking algorithm. The extracted chirps are then postprocessed at the receiver using a line detection algorithm based on the Hough–Radon transform (HRT). The HRT is an optimal line-detection algorithm, which detects directional components that satisfy a parametric constraint equation in the image of a TF plane, even at discontinuities corresponding to bit errors. Simulation results show that HRT correctly detects the embedded watermark message after common signal processing operations for bit error rates up to 20%. The new watermark representation and the postprocessing stage based on HRT significantly improve the performance of the watermark detection process and can be combined with existing watermark embedding/extraction algorithms for increased robustness. Serhat Erküçük, Sridhar Krishnan 0001, Mehmet Zeytinoglu |
IEEE Trans. Multim. | 2 |
| 2005 | Data embedding in μ-law speech with spread spectrum techniquesabstractThis paper explores data embedding in G.711 mu-law speech signals with the spread spectrum techniques. Based on an optimized spread spectrum scheme, a simple but effective solution is presented for high-capacity embedding. Simulations show that the proposed scheme, when incorporated with the measure of the frequency masking effects, can achieve an embedding rate of about 100 bits per second with a 7% bit error rate (BER), or 1000 bps with a 10% BER Heping Ding, Sridhar Krishnan 0001 |
GLOBECOM | 3 |
| 2005 | Indexing of NFL Video using MPEG-7 Descriptors and MFCC featuresabstractIn this paper, we propose an application system to classify American football (NFL) video shots into 4 categories, namely: pass plays, run plays, field goal/extra point plays (FG/XP) and kickoff/punt plays (K/P). The proposed system consists of two stages. The first stage is responsible for play event localization and the latter stage is responsible for feature mapping and classification. For play event localization we have proposed an algorithm that uses MPEG-7 motion activity descriptor and mean of the magnitudes of motion vectors, in a collaborative manner to detect the starting point of a play event within a video shot with 83% accuracy. The indexing and classification stage uses MPEG-7 motion and audio descriptors along with Mel Frequency Cepstrum Coefficients (MFCC) features to classify the events into 4 categories using Fisher's LDA. We obtain indexing accuracy of 92.5% by using a leave-one-out classification technique on a database of 200 video shots taken from 4 different games obtained from 4 different networks. Syed G. Quadri, Sridhar Krishnan 0001, Ling Guan |
ICASSP (2) | 2 |
| 2005 | A signal classification approach using time-width vs frequency band sub-energy distributionsabstractTime-frequency (TF) signal decompositions provide us with ample information and extreme flexibility for signal analysis. By applying suitable processing on the TF decomposition parameters, even subtle signal characteristics can be revealed. In many real world applications, identification of these subtle differences make a significant impact in signal analysis. Particularly in classification applications using TF approaches, there may be situations where a localized high discriminative signal structure is diluted due to the presence of other overlapping signal structures. To address this problem we propose a novel approach to construct multiple time-width vs frequency band mappings based on the energy decomposition pattern of the signal. These mapping are then analyzed to locate the highly discriminative features for classification. Initial results with two real world biomedical signal databases: (1) vibroarthrographic (VAG) signals; and (2) pathological speech signals, indicate high potential for the proposed technique. Karthikeyan Umapathy, Sridhar Krishnan 0001 |
ICASSP (5) | 2 |
| 2005 | Gaussian Mixture Modeling Using Short Time Fourier Transform Features for Audio FingerprintingabstractIn audio fingerprinting, an audio clip must be recognized by matching an extracted fingerprint to a database of previously computed fingerprints. The fingerprints should reduce the dimensionality of the input significantly, provide discrimination among different audio clips, and at the same time, invariant to the distorted versions of the same audio clip. In this paper, we design fingerprints addressing the above issues by modeling an audio clip by Gaussian mixture models (GMM) using a wide range of easy-to-compute short time Fourier transform features such as Shannon entropy, Renyi entropy, spectral centroid, spectral bandwidth, spectral flatness measure, spectral crest factor, and Mel-frequency cepstral coefficients. We test the robustness of the fingerprints under a large number of distortions. To make the system robust, we use some of the distorted versions of the audio for training. However, we show that the audio fingerprints modeled using GMM are not only robust to the distortions used in training but also to distortions not used in training. Using spectral centroid as feature, we obtain the highest identification rate of 99.1% with a false positive rate of 10-4 Arunan Ramalingam, Sridhar Krishnan 0001 |
ICME | 2 |
| 2005 | Multigroup classification of audio signals using time-frequency parametersabstractThe ongoing advancements in the multimedia technologies drive the need for efficient classification of the audio signals to make the content-based retrieval process more accurate and much easier from huge databases. The challenge of this task lies in an accurate extraction of signal characteristics so as to derive a strong discriminatory feature suitable for classification. In this paper, a time-frequency (TF) approach for audio classification is proposed. Audio signals are nonstationary in nature and TF approach is the best way to analyze them. The audio signals were decomposed using an adaptive TF decomposition algorithm, and the signal decomposition parameter based on octave (scaling) was used to generate a set of 42 features over three frequency bands within the auditory range. These features were analyzed using linear discriminant functions and classified into six music groups (rock, classical, country, jazz, folk and pop). Overall classification accuracies as high as 97.6 % was achieved by linear discriminant analysis of 170 audio signals. Karthikeyan Umapathy, Sridhar Krishnan 0001, Shihab A. Jimaa |
IEEE Trans. Multim. | 2 |
| 2004 | Content based audio classification and retrieval using joint time-frequency analysisabstractWe present an audio classification and retrieval technique that exploits the non-stationary behavior of music signals and extracts features that characterize their spectral change over time. Audio classification provides a solution to incorrect and inefficient manual labelling of audio files on computers by allowing users to extract music files based on content similarity rather than labels. In our technique, classification is performed using time-frequency analysis and sounds are classified into 6 music groups consisting of rock, classical, folk, jazz and pop. For each 5 second music segment, the features that are extracted include entropy, centroid, centroid ratio, bandwidth, silence ratio, energy ratio, and location of minimum and maximum energy. Using a database of 143 signals, a set of 10 time-frequency features are extracted and an accuracy of classification of around 93% using regular linear discriminant analysis or 92.3% using the leave-one-out method is achieved. Shahrzad Esmaili, Sridhar Krishnan 0001, Kaamran Raahemifar |
ICASSP (5) | 2 |
| 2004 | Modified local discriminant bases and its applications in signal classification [biomedical signal examples]abstractOne of the major challenges in classification problems, based on the signal decomposition approach, is to identify the right basis function and its derivatives that can provide optimal features to distinguish the classes. With the vast amount of available libraries of orthonormal bases, it is hard to select an optimal set of basis functions for a specific dataset. To address this problem, pruning algorithms based on certain selection criteria, are needed. The local discriminant bases (LDB) algorithm is one such algorithm, which efficiently selects a set of significant basis functions from the library of orthonormal bases based on a certain defined dissimilarity measure. The selection of this dissimilarity measure is critical as they indirectly contribute to the performance accuracy of the LDB algorithm. In this paper, we study the impact of the dissimilarity measures on the performance of the LDB algorithm with two classification examples. Two biomedical signal databases used are: 1) vibroarthographic signals (VAG) - 89 signals with 51 normal and 38 abnormal; and 2) pathological speech signals - 100 signals with 50 normal and 50 pathological. Classification accuracies of 76.4% with the VAG database and 96% with the pathological speech database were obtained. This modified method of signal analysis using LDB has shown its powerfulness in analyzing non-stationary signals. Karthikeyan Umapathy, Sridhar Krishnan 0001 |
ICASSP (2) | 2 |
| 2003 | Robust audio watermarking using a chirp based techniqueabstractIn this study, we propose a new spread spectrum audio watermarking algorithm that embeds linear chirps as watermark messages. Different chirp rates, i.e., slopes on the time-frequency (TF) plane, represent watermark messages such that each slope corresponds to a different message. We extract the watermark message using a line detection algorithm based on the Hough-Radon transform (HRT). The HRT detects the directional elements that satisfy a parametric constraint in the image of a TF plane. The proposed method not only detects the presence of watermark, but also extracts the embedded watermark bits and ensures the message is received correctly. The results show that the HRT detects the embedded watermark message even after common signal processing operations such as MPEG audio coding, resampling, lowpass filtering and amplitude re-scaling. Serhat Erküçük, Sridhar Krishnan 0001, Mehmet Zeytinoglu |
ICME | 2 |
| 2002 | Interference excision in spread spectrum communications using adaptive positive time-frequency distributionsabstractThere have been several techniques proposed to excise the interference in spread spectrum communications using time-frequency distributions (TFDs). TFDs localize any interference both in time and frequency domain, and are idealy suited for interference excision. Unfortunately, the commonly used TFDs suffer from a trade-off between time-frequency (TF) resolution and cross-terms suppression. This paper focuses on a new excision technique based on constructing a positive TFD of the received spread spectrum signal using an adaptive signal decomposition technique. By decomposing a signal into components, the interaction between components can be avoided, and the TFD constructed by combining the TFDs of the individual components would be free of cross-terms. Also, by using Gaussian functions as bases for decomposition, a high TF resolution of interference signals can be achieved. Construction of positive TFDs by signal decomposition techniques facilitates automatic denoising, and extraction of marginal and local properties of a signal such as instantaneous energy, power spectral density, instantaneous frequency and group delay. Interference excision is then achieved by suitably thresholding energy values in the TF plane. Initial results with synthetic models have shown successful performance with linear and quadratic chirp interferences. The interference excisions are highly localized in the TF plane with no cross-terms Serhat Erküçük, Sridhar Krishnan 0001 |
ICASSP | 2 |
| 2002 | Detection of linear chirp and non-linear chirp interferences in a spread spectrum signal by using Hough-Radon transformabstractThe time-frequency distribution (TFD) of a spread spectrum signal looks more like a noise, and the energy distribution occupies the full two-dimensional time-frequency (TF) plane. Any jammer or interference will be well localized in the TF plane. By treating the TF plane as an image, the interference patterns can be detected by using the image analysis technique of Hough-Radon transform (HRT). Curves with mathematical equations can be easily detected by transforming the shapes into Hough domain, and searching for dominant peaks (maximum values). The co-ordinates of the dominant peaks provide the parameters of the shape. For example, in case of a straight line, the Hough domain would be the “rho, theta” space, where “rho and theta” are the parameters of a straight line. The maximum value in the rho, theta plane would correspond to the exact parmeters of the straight line. If a high resolution TFD for a spread spectrum signal is achieved, then any linear chirp or non-linear chirp interference will show up as straight lines and curves in the TF plane. By applying the HRT on the TF plane, chirp interferences can be identified. Evaluation of the proposed techniques show successful detection of both linear and hyperbolic (nonlinear) chirp interferences in spread spectrum signals even under very low SNR conditions of 0 dB. The method detects any localized interference as along as the interference pattern in the TF plane can be represented by a Shynimol Thayilchira, Sridhar Krishnan 0001 |
ICASSP | 2 |
| 2002 | Discrimination of pathological voices using an adaptive time-frequency approachabstractAcoustic measures of vocal function are routinely used for the assessment of disordered voice, and for monitoring patient's progress over the course of therapy. In current clinical practice, acoustic measures extracted from sustained vowels are used for vocal function characterization. However, the measures derived from continuous speech samples are required for accurate assessment of voice quality. In this paper, a time-frequency approach for pathological voice discrimination has been proposed. The speech signals were decomposed using an adaptive time-frequency transform algorithm, and the signal decomposition parameters such as the octave (scale) maximum, octave mean, energy rate, and length ratio were analyzed using the maximum likelihood method and Jack-knife algorithm for classification. A classification accuracy of 90% was obtained with a database of 40 speech signals (20 normal and 20 pathological cases). Karthikeyan Umapathy, Sridhar Krishnan 0001, Vijay Parsa, Donald G. Jamieson |
ICASSP | 2 |
| 2002 | Audio signal classification using time-frequency parametersabstractThe ongoing advancements in the multimedia technologies drive the need for efficient classification of the audio signals to make the content-based retrieval process more accurate and much easier from huge databases. The challenge of this task lies in an accurate extraction of signal characteristics so as to derive a strong discriminatory feature suitable for retrieval process. A time-frequency approach for audio classification is proposed. The audio signals were decomposed using an adaptive time-frequency decomposition algorithm, and the signal decomposition parameter octave (scale) was used to create patterns based on a similarity measure of the audio signals. These patterns were used to generate templates to classify the audio signals into different categories. Initial studies have yielded a overall correct classification accuracy of 90% with a database of 64 audio segments. Karthikeyan Umapathy, Sridhar Krishnan 0001, Shihab A. Jimaa |
ICME (2) | 2 |
| 2001 | Feature identification in the time-frequency plane by using the Hough-Radon transform
Rangaraj M. Rangayyan, Sridhar Krishnan 0001 |
Pattern Recognit. | 2 |