Anindya Sarkar

dblp:87/2477 · DBLP profile ↗
← Back
26ranked-venue papers
19as first author
12since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 15 · 11 first-author · 5 since 2021Artificial intelligence and machine learning · 13 · 10 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 2Security and privacy · 1 · 1 first-author
YearPublicationVenuePosition
2025 Active Geospatial Search for Efficient Tenant Eviction Outreach
abstract
Tenant evictions threaten housing stability and are a major concern for many cities. An open question concerns whether data-driven methods enhance outreach programs that target at-risk tenants to mitigate their risk of eviction. We propose a novel active geospatial search (AGS) modeling framework for this problem. AGS integrates property-level information in a search policy that identifies a sequence of rental units to canvas to both determine their eviction risk and provide support if needed. We propose a hierarchical reinforcement learning approach to learn a search policy for AGS that scales to large urban areas containing thousands of parcels, balancing exploration and exploitation and accounting for travel costs and a budget constraint. Crucially, the search policy adapts online to newly discovered information about evictions. Evaluation using eviction data for a large urban area demonstrates that the proposed framework and algorithmic approach are considerably more effective at sequentially identifying eviction cases than baseline methods.
Anindya Sarkar, Alex DiChristofano, Sanmay Das, Patrick J. Fowler, Nathan Jacobs, Yevgeniy Vorobeychik
AAAI1
2025 Learning Policy Committees for Effective Personalization in MDPs with Diverse Tasks
abstract
Many dynamic decision problems, such as robotic control, involve a series of tasks, many of which are unknown at training time. Typical approaches for these problems, such as multi-task and meta reinforcement learning, do not generalize well when the tasks are diverse. On the other hand, approaches that aim to tackle task diversity, such as using task embedding as policy context and task clustering, typically lack performance guarantees and require a large number of training tasks. To address these challenges, we propose a novel approach for learning a policy committee that includes at least one near-optimal policy with high probability for tasks encountered during execution. While we show that this problem is in general inapproximable, we present two practical algorithmic solutions. The first yields provable approximation and task sample complexity guarantees when tasks are low-dimensional (the best we can do due to inapproximability), whereas the second is a general and practical gradient-based approach. In addition, we provide a provable sample complexity bound for few-shot learning. Our experiments on MuJoCo and Meta-World show that the proposed approach outperforms state-of-the-art multi-task, meta-, and task clustering baselines in training, generalization, and few-shot learning, often by a large margin. Our code is available at https://github.com/CERL-WUSTL/PACMAN.
Luise Ge, Michael Lanier, Anindya Sarkar, Bengisu Guresti, Chongjie Zhang, Yevgeniy Vorobeychik
ICML3
2025 Online Feedback Efficient Active Target Discovery in Partially Observable Environments
abstract
In various scientific and engineering domains, where data acquisition is costly—such as in medical imaging, environmental monitoring, or remote sensing—strategic sampling from unobserved regions, guided by prior observations, is essential to maximize target discovery within a limited sampling budget. In this work, we introduce Diffusion-guided Active Target Discovery (DiffATD), a novel method that leverages diffusion dynamics for active target discovery. DiffATD maintains a belief distribution over each unobserved state in the environment, using this distribution to dynamically balance exploration-exploitation. Exploration reduces uncertainty by sampling regions with the highest expected entropy, while exploitation targets areas with the highest likelihood of discovering the target, indicated by the belief distribution and an incrementally trained reward model designed to learn the characteristics of the target. DiffATD enables efficient target discovery in a partially observable environment within a fixed sampling budget, all without relying on any prior supervised training. Furthermore, DiffATD offers interpretability, unlike existing black-box policies that require extensive supervised training. Through extensive experiments and ablation studies across diverse domains, including medical imaging, species discovery and remote sensing, we show that DiffATD performs significantly better than baselines and competitively with supervised methods that operate under full environmental observability.
Anindya Sarkar, Binglin Ji, Yevgeniy Vorobeychik
NeurIPS1
2025 Active Target Discovery under Uninformative Priors: The Power of Permanent and Transient Memory
abstract
In many scientific and engineering fields, where acquiring high-quality data is expensive—such as medical imaging, environmental monitoring, and remote sensing—strategic sampling of unobserved regions based on prior observations is crucial for maximizing discovery rates within a constrained budget. The rise of powerful generative models, such as diffusion models, has enabled active target discovery in partially observable environments by leveraging learned priors—probabilistic representations that capture underlying structure from data. With guidance from sequentially gathered task-specific observations, these models can progressively refine exploration and efficiently direct queries toward promising regions. However, in domains where learning a strong prior is infeasible due to extremely limited data or high sampling cost (such as rare species discovery, diagnostics for emerging diseases, etc.), these methods struggle to generalize. To overcome this limitation, we propose a novel approach that enables effective active target discovery even in settings with uninformative priors, ensuring robust exploration and adaptability in complex real-world scenarios. Our framework is theoretically principled and draws inspiration from neuroscience to guide its design. Unlike black-box policies, our approach is inherently interpretable, providing clear insights into decision-making. Furthermore, it guarantees a strong, monotonic improvement in prior estimates with each new observation, leading to increasingly accurate sampling and reinforcing both reliability and adaptability in dynamic settings. Through comprehensive experiments and ablation studies across various domains, including species distribution modeling and remote sensing, we demonstrate that our method substantially outperforms baseline approaches.
Anindya Sarkar, Binglin Ji, Yevgeniy Vorobeychik
NeurIPS1
2024 GOMAA-Geo: GOal Modality Agnostic Active Geo-localization
abstract
We consider the task of active geo-localization (AGL) in which an agent uses a sequence of visual cues observed during aerial navigation to find a target specified through multiple possible modalities. This could emulate a UAV involved in a search-and-rescue operation navigating through an area, observing a stream of aerial images as it goes. The AGL task is associated with two important challenges. Firstly, an agent must deal with a goal specification in one of multiple modalities (e.g., through a natural language description) while the search cues are provided in other modalities (aerial imagery). The second challenge is limited localization time (e.g., limited battery life, urgency) so that the goal must be localized as efficiently as possible, i.e. the agent must effectively leverage its sequentially observed aerial views when searching for the goal. To address these challenges, we propose GOMAA-Geo -- a goal modality agnostic active geo-localization agent -- for zero-shot generalization between different goal modalities. Our approach combines cross-modality contrastive learning to align representations across modalities with supervised foundation model pretraining and reinforcement learning to obtain highly effective navigation and localization policies. Through extensive evaluations, we show that GOMAA-Geo outperforms alternative learnable approaches and that it generalizes across datasets -- e.g., to disaster-hit areas without seeing a single disaster scenario during training -- and goal modalities -- e.g., to ground-level imagery or textual descriptions, despite only being trained with goals specified as aerial views. Our code is available at: https://github.com/mvrl/GOMAA-Geo.
Anindya Sarkar, Srikumar Sastry, Aleksis Pirinen, Chongjie Zhang, Nathan Jacobs, Yevgeniy Vorobeychik
NeurIPS1
2024 A Visual Active Search Framework for Geospatial Exploration
abstract
Many problems can be viewed as forms of geospatial search aided by aerial imagery, with examples ranging from detecting poaching activity to human trafficking. We model this class of problems in a visual active search (VAS) framework, which has three key inputs: (1) an image of the entire search area, which is subdivided into regions, (2) a local search function, which determines whether a previously unseen object class is present in a given region, and (3) a fixed search budget, which limits the number of times the local search function can be evaluated. The goal is to maximize the number of objects found within the search budget. We propose a reinforcement learning approach for VAS that learns a meta-search policy from a collection of fully annotated search tasks. This meta-search policy is then used to dynamically search for a novel target-object class, leveraging the outcome of any previous queries to determine where to query next. Through extensive experiments on several large-scale satellite imagery datasets, we show that the proposed approach significantly outperforms several strong baselines. We also propose novel domain adaptation techniques that improve the policy at decision time when there is a significant domain gap with the training data. Code is publicly available at this link.
Anindya Sarkar, Michael Lanier, Scott Alfeld, Jiarui Feng, Roman Garnett, Nathan Jacobs, Yevgeniy Vorobeychik
WACV1
2023 A Partially-Supervised Reinforcement Learning Framework for Visual Active Search
abstract
Visual active search (VAS) has been proposed as a modeling framework in which visual cues are used to guide exploration, with the goal of identifying regions of interest in a large geospatial area. Its potential applications include identifying hot spots of rare wildlife poaching activity, search-and-rescue scenarios, identifying illegal trafficking of weapons, drugs, or people, and many others. State of the art approaches to VAS include applications of deep reinforcement learning (DRL), which yield end-to-end search policies, and traditional active search, which combines predictions with custom algorithmic approaches. While the DRL framework has been shown to greatly outperform traditional active search in such domains, its end-to-end nature does not make full use of supervised information attained either during training, or during actual search, a significant limitation if search tasks differ significantly from those in the training distribution. We propose an approach that combines the strength of both DRL and conventional active search approaches by decomposing the search policy into a prediction module, which produces a geospatial distribution of regions of interest based on task embedding and search history, and a search module, which takes the predictions and search history as input and outputs the search distribution. In addition, we develop a novel meta-learning approach for jointly learning the resulting combined policy that can make effective use of supervised information obtained both at training and decision time. Our extensive experiments demonstrate that the proposed representation and meta-learning frameworks significantly outperform state of the art in visual active search on several problem domains.
Anindya Sarkar, Nathan Jacobs, Yevgeniy Vorobeychik
NeurIPS1
2022 A Framework for Learning Ante-hoc Explainable Models via Concepts
abstract
Self-explaining deep models are designed to learn the latent concept-based explanations implicitly during training, which eliminates the requirement of any post-hoc explanation generation technique. In this work, we propose one such model that appends an explanation generation module on top of any basic network and jointly trains the whole module that shows high predictive performance and generates meaningful explanations in terms of concepts. Our training strategy is suitable for unsupervised concept learning with much lesser parameter space requirements compared to baseline methods. Our proposed model also has provision for leveraging self-supervision on concepts to extract better explanations. However, with full concept supervision, we achieve the best predictive performance compared to recently proposed concept-based explainable models. We report both qualitative and quantitative results with our method, which shows better performance than recently proposed concept-based explainability methods. We reported exhaustive results with two datasets without ground truth concepts, i.e., CIFAR10, ImageNet, and two datasets with ground truth concepts, i.e., AwA2, CUB-200, to show the effectiveness of our method for both cases. To the best of our knowledge, we are the first ante-hoc explanation generation method to show results with a large-scale dataset such as ImageNet.
Anirban Sarkar 0001, Deepak Vijaykeerthy, Anindya Sarkar, Vineeth N. Balasubramanian
CVPR3
2022 How Powerful are K-hop Message Passing Graph Neural Networks
abstract
The most popular design paradigm for Graph Neural Networks (GNNs) is 1-hop message passing---aggregating information from 1-hop neighbors repeatedly. However, the expressive power of 1-hop message passing is bounded by the Weisfeiler-Lehman (1-WL) test. Recently, researchers extended 1-hop message passing to $K$-hop message passing by aggregating information from $K$-hop neighbors of nodes simultaneously. However, there is no work on analyzing the expressive power of $K$-hop message passing. In this work, we theoretically characterize the expressive power of $K$-hop message passing. Specifically, we first formally differentiate two different kernels of $K$-hop message passing which are often misused in previous works. We then characterize the expressive power of $K$-hop message passing by showing that it is more powerful than 1-WL and can distinguish almost all regular graphs. Despite the higher expressive power, we show that $K$-hop message passing still cannot distinguish some simple regular graphs and its expressive power is bounded by 3-WL. To further enhance its expressive power, we introduce a KP-GNN framework, which improves $K$-hop message passing by leveraging the peripheral subgraph information in each hop. We show that KP-GNN can distinguish many distance regular graphs which could not be distinguished by previous distance encoding or 3-WL methods. Experimental results verify the expressive power and effectiveness of KP-GNN. KP-GNN achieves competitive results across all benchmark datasets.
Jiarui Feng, Yixin Chen 0001, Fuhai Li 0001, Anindya Sarkar, Muhan Zhang
NeurIPS4
2022 Leveraging Test-Time Consensus Prediction for Robustness against Unseen Noise
abstract
We propose a method to improve DNN robustness against unseen noisy corruptions, such as Gaussian noise, Shot Noise, Impulse Noise, Speckle noise with different levels of severity by leveraging ensemble technique through a consensus based prediction method using self-supervised learning at inference time. We also propose to enhance the model training by considering other aspects of the issue i.e. noise in data and better representation learning which shows even better generalization performance with the consensus based prediction strategy. We report results of each noisy corruption on the standard CIFAR10-C and ImageNet-C benchmark which shows significant boost in performance over previous methods. We also introduce results for MNIST-C and TinyImagenet-C to show usefulness of our method across datasets of different complexities to provide robustness against unseen noise. We show results with different architectures to validate our method against other baseline methods, and also conduct experiments to show the usefulness of each part of our method.
Anindya Sarkar, Anirban Sarkar 0001, Vineeth N. Balasubramanian
WACV1
2021 Enhanced Regularizers for Attributional Robustness
abstract
Deep neural networks are the default choice of learning models for computer vision tasks. Extensive work has been carried out in recent years on explaining deep models for vision tasks such as classification. However, recent work has shown that it is possible for these models to produce substantially different attribution maps even when two very similar images are given to the network, raising serious questions about trustworthiness. To address this issue, we propose a robust attribution training strategy to improve attributional robustness of deep neural networks. Our method carefully analyzes the requirements for attributional robustness and introduces two new regularizers that preserve a model's attribution map during attacks. Our method surpasses state-of-the-art attributional robustness methods by a margin of approximately 3% to 9% in terms of attribution robustness measures on several datasets including MNIST, FMNIST, Flower and GTSRB.
Anindya Sarkar, Anirban Sarkar 0001, Vineeth N. Balasubramanian
AAAI1
2021 Adversarial Robustness without Adversarial Training: A Teacher-Guided Curriculum Learning Approach
abstract
Current SOTA adversarially robust models are mostly based on adversarial training (AT) and differ only by some regularizers either at inner maximization or outer minimization steps. Being repetitive in nature during the inner maximization step, they take a huge time to train. We propose a non-iterative method that enforces the following ideas during training. Attribution maps are more aligned to the actual object in the image for adversarially robust models compared to naturally trained models. Also, the allowed set of pixels to perturb an image (that changes model decision) should be restricted to the object pixels only, which reduces the attack strength by limiting the attack space. Our method achieves significant performance gains with a little extra effort (10-20%) over existing AT models and outperforms all other methods in terms of adversarial as well as natural accuracy. We have performed extensive experimentation with CIFAR-10, CIFAR-100, and TinyImageNet datasets and reported results against many popular strong adversarial attacks to prove the effectiveness of our method.
Anindya Sarkar, Anirban Sarkar 0001, Sowrya Gali, Vineeth N. Balasubramanian
NeurIPS1
2020 Enforcing Linearity in DNN Succours Robustness and Adversarial Image Generation
Anindya Sarkar, Raghu Sesha Iyengar
ICANN (1)1
2020 Neural Data Augmentation Techniques for Time Series Data and its Benefits
abstract
Exploring adversarial attacks and studying their effects on machine learning algorithms has been of interest to researchers. Deep neural networks working with time series data have received lesser interest compared to their image counterparts in this context. In a recent finding, it has been revealed that current state-of-the-art deep learning time series classifiers are vulnerable to adversarial attacks. In this paper, we introduce neural data augmentation techniques and show that classifier trained with such augmented data obtains state-of-the-art classification accuracy as well as adversarial accuracy against Fast Gradient Sign Method (FGSM) and Basic Iterative Method (BIM) on various time series benchmarks.
Anindya Sarkar, Anirudh Sunder Raj, Raghu Sesha Iyengar
ICMLA1
2014 Prostate Cancer Grading: Use of Graph Cut and Spatial Arrangement of Nuclei
abstract
Tissue image grading is one of the most important steps in prostate cancer diagnosis, where the pathologist relies on the gland structure to assign a Gleason grade to the tissue image. In this grading scheme, the discrimination between grade 3 and grade 4 is the most difficult, and receives the most attention from researchers. In this study, we propose a novel method (called nuclei-based method) that 1) utilizes graph theory techniques to segment glands and 2) computes a gland-score (based on the spatial arrangement of nuclei) to estimate how similar a segmented region is to a gland. Next, we create a fusion method by combining this nuclei-based method with the lumen-based method presented in our previous work to improve the performance of grade 3 versus grade 4 classification problem (the accuracy is now improved to 87.3% compared to 81.1% of the lumen-based method alone). To segment glands, we build a graph of nuclei and lumina in the image, and use the normalized cut method to partition the graph into different components, each corresponding to a gland. Unlike most state-of-the-art lumen-based gland segmentation method, the nuclei-based method is able to segment glands without lumen or glands with multiple lumina. Moreover, another important contribution in this research is the development of a set of measures to exploit the difference in nuclei spatial arrangement between grade 3 images (where nuclei form closed chain structure on the gland boundary) and grade 4 image (where nuclei distribute more randomly in the gland). These measures are combined to generate a single gland-score value, which estimates how similar a segmented region (which is a set of nuclei and lumina) is to a gland.
Kien Nguyen 0005, Anindya Sarkar, Anil K. Jain 0001
IEEE Trans. Medical Imaging2
2012 Structure and Context in Prostatic Gland Segmentation and Classification
Kien Nguyen 0005, Anindya Sarkar, Anil K. Jain 0001
MICCAI (1)2
2010 Precise localization of key-points to identify local regions for robust data hiding
abstract
We propose a novel data hiding system where data is embedded in local non-overlapping regions in an image. To survive cropping, the encoder embeds the same data in multiple regions of fixed dimensions, while the decoder's challenge is to independently retrieve these regions. Salient feature points are computed on an image and the local regions are centered around them. To obtain non-overlapping regions, the points are pruned based on their corner strength and the size of the region. The decoder can retrieve the data only if it can precisely identify one or more of the same key-points. We present suitable key-point pruning methods such that even after considering a reduced number of key-points, the receiver is successful in identifying the same key-point locations as the encoder. We perform experimental comparison of various corner detectors and also study the performance of segmentation methods to obtain robust key-points.
Lakshmanan Nataraj, Anindya Sarkar, B. S. Manjunath
ICIP2
2010 Efficient and Robust Detection of Duplicate Videos in a Large Database
abstract
We present an efficient and accurate method for duplicate video detection in a large database using video fingerprints. We have empirically chosen the color layout descriptor, a compact and robust frame-based descriptor, to create fingerprints which are further encoded by vector quantization (VQ). We propose a new nonmetric distance measure to find the similarity between the query and a database video fingerprint and experimentally show its superior performance over other distance measures for accurate duplicate detection. Efficient search cannot be performed for high-dimensional data using a nonmetric distance measure with existing indexing techniques. Therefore, we develop novel search algorithms based on precomputed distances and new dataset pruning techniques yielding practical retrieval times. We perform experiments with a database of 38 000 videos, worth 1600 h of content. For individual queries with an average duration of 60 s (about 50% of the average database video length), the duplicate video is retrieved in 0.032 s, on Intel Xeon with CPU 2.33 GHz, with a very high accuracy of 97.5%.
Anindya Sarkar, Vishwakarma Singh, Pratim Ghosh, B. S. Manjunath, Ambuj K. Singh
IEEE Trans. Circuits Syst. Video Technol.1
2010 Matrix embedding with pseudorandom coefficient selection and error correction for robust and secure steganography
abstract
In matrix embedding (ME)-based steganography, the host coefficients are minimally perturbed such that the transmitted bits fall in a coset of a linear code, with the syndrome conveying the hidden bits. The corresponding embedding distortion and vulnerability to steganalysis are significantly less than that of conventional quantization index modulation (QIM)-based hiding. However, ME is less robust to attacks, with a single host bit error leading to multiple decoding errors for the hidden bits. In this paper, we employ the ME-RA scheme, a combination of ME-based hiding with powerful repeat accumulate (RA) codes for error correction, to address this problem. A key contribution of this paper is to compute log likelihood ratios for RA decoding, taking into account the many-to-one mapping between the host coefficients and an encoded bit, for ME. To reduce detectability, we hide in randomized blocks, as in the recently proposed Yet Another Steganographic Scheme (YASS), replacing the QIM-based embedding in YASS by the proposed ME-RA scheme. We also show that the embedding performance can be improved by employing punctured RA codes. Through experiments based on a couple of thousand images, we show that for the same embedded data rate and a moderate attack level, the proposed ME-based method results in a lower detection rate than that obtained for QIM-based YASS.
Anindya Sarkar, Upamanyu Madhow, B. S. Manjunath
IEEE Trans. Inf. Forensics Secur.1
2009 Adding Gaussian noise to "denoise" JPEG for detecting image resizing
abstract
A common problem affecting most image resizing detection algorithms is that they are susceptible to JPEG compression. This is because JPEG introduces periodic artifacts, as it works on 8×8 blocks. We propose a novel yet counter intuitive technique to "denoise" JPEG images by adding Gaussian noise. We add a suitable amount of Gaussian noise to a resized and JPEG compressed image so that the periodicity due to JPEG compression is suppressed while that due to resizing is retained. The controlled Gaussian noise addition works better than median filtering and weighted averaging based filtering for suppressing the JPEG induced periodicity.
Lakshmanan Nataraj, Anindya Sarkar, B. S. Manjunath
ICIP2
2009 Double embedding in the quantization index modulation framework
abstract
Quantization index modulation (QIM) is a commonly used data hiding technique where a single bit is embedded per coefficient. Here, we propose the use of double embedding in the QIM framework where a single coefficient is modified twice, using two quantizers, to embed two bits. The motivation behind substituting single embedding with double embedding in the QIM framework for a certain steganographic scheme is to increase its hiding rate without significantly increasing the embedding distortion and the stego scheme's detectability against steganalysis. We empirically determine the best way to couple the double embedding framework with a repeat accumulate code based error correction scheme. For moderate noise levels, the use of double embedding is seen to be significantly advantageous over single embedding.
Anindya Sarkar, B. S. Manjunath
ICIP1
2008 Estimation of optimum coding redundancy and frequency domain analysis of attacks for YASS - a randomized block based hiding scheme
abstract
Our recently introduced JPEG steganographic method called yet another steganographic scheme (YASS) can resist blind steganalysis by embedding data in the discrete cosine transform (DCT) domain in randomly chosen image blocks. To maximize the embedding rate for a given image and a specified attack channel, the redundancy factor used by the repeat- accumulate (RA) code based error correction framework in YASS is optimally chosen by the encoder. An efficient method is suggested for the decoder to accurately compute this redundancy factor. We also show experimentally which DCT coefficients are better suited for hiding and detection under various attacks. The effectiveness of YASS for robust steganography is demonstrated for certain attacks.
Anindya Sarkar, Lakshmanan Nataraj, B. S. Manjunath, Upamanyu Madhow
ICIP1
2007 Secure Steganography: Statistical Restoration of the Second Order Dependencies for Improved Security
abstract
We present practical approaches for steganography that can provide improved security by closely matching the second-order statistics of the host rather than just the marginal distribution. The methods are based on the framework of statistical restoration, wherein a fraction of the host symbols available for hiding is actually used to restore the statistics; thus reducing the rate, but providing security against steganalysis. We establish correspondence between steganography and the earth-mover's distance (EMD), a popular distance metric used in computer vision applications. The EMD framework can be used to define the optimum flow (modifications) of the host symbols for compensation. This formulation is used for image steganography by restoring the second-order statistics of the blockwise discrete cosine transform (DCT) coefficients. Some practical limitations of this approach (such as computational complexity and difficulty in dealing with overlapping coefficient pairs) are noted, and a new method is proposed that alleviates these deficiencies by identifying the coefficients to modify based on a local compensation criterion. Experimental results on several thousand natural images demonstrate the utility of the presented methods.
Anindya Sarkar, Kaushal Solanki, Upamanyu Madhow, Shivkumar Chandrasekaran, B. S. Manjunath
ICASSP (2)1
2007 Estimating Steganographic Capacity for Odd-Even Based Embedding and its Use in Individual Compensation
abstract
We present a method to compute the steganographic capacity for images, with odd-even based hiding in the quantized discrete cosine transform domain. The method has been generalized for varying orders of co-occurrence statistics for statistical restoration based steganography. We further utilize this capacity estimate to hide the maximum possible data per individual frequency stream, while ensuring that the first order histograms of individual frequency coefficients remain matched. We also show that certain frequency components are more useful for steganalysis after first order statistical restoration is performed for a certain band of select frequencies.
Anindya Sarkar, B. S. Manjunath
ICIP (1)1
2005 Automatic Speech Segmentation Using Average Level Crossing Rate Information
abstract
We explore new methods of determining automatically derived units for classification of speech into segments. For detecting signal changes, temporal features are more reliable than the standard feature vector domain methods, since both magnitude and phase information are retained. Motivated by auditory models, we have presented a method based on average level crossing rate (ALCR) of the signal, to detect significant temporal changes in the signal. An adaptive level allocation scheme has been used in this technique that allocates levels, depending on the signal pdf and SNR. We compare the segmentation performance to manual phonemic segmentation and also that provided by maximum likelihood (ML) segmentation for 100 TIMIT sentences. The ALCR method matches the best segmentation performance without a priori knowledge of number of segments, as in ML segmentation.
Anindya Sarkar, Thippur V. Sreenivas
ICASSP (1)1
2005 Dynamic programming based segmentation approach to LSF matrix reconstruction
Anindya Sarkar, Thippur V. Sreenivas
INTERSPEECH1