Arindam Jati

dblp:117/4725 · DBLP profile ↗
← Back
22ranked-venue papers
11as first author
10since 2021 · last 2025
0000-0002-9498-8536ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 16 · 8 first-author · 8 since 2021Graphics, computer vision, multimedia, augmented reality and games · 11 · 6 first-author · 3 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021
YearPublicationVenuePosition
2025 Towards Unbiased Evaluation of Time-series Anomaly Detector
abstract
Time series anomaly detection (TSAD) is an evolving area of research motivated by its critical applications, such as detecting seismic activity, sensor failures in industrial plants, predicting crashes in the stock market, and so on. Across domains, anomalies occur significantly less frequently than normal data, making the F1-score the most commonly adopted metric for anomaly detection. However, in the case of time series, it is not straightforward to use standard F1-score because of the dissociation between ‘time points’ and ‘time events’. To accommodate this, anomaly predictions are adjusted, called as point adjustment (PA), before the F1-score evaluation. However, these adjustments are heuristics-based, and biased towards true positive detection, resulting in over-estimated detector performance. In this work, we propose an alternative adjustment protocol called "Balanced point adjustment" (BA). It addresses the limitations of existing point adjustment methods and provides guarantees of fairness backed by axiomatic definitions of TSAD evaluation. Code and implementation details: https://github.com/summukhe/balanced_f1score.
Debarpan Bhattacharya, Sumanta Mukherjee, Chandramouli K, Vijay Ekambaram, Arindam Jati, Pankaj Dayama 0001
ICASSP5
2024 AutoMixer for Improved Multivariate Time-Series Forecasting on Business and IT Observability Data
abstract
The efficiency of business processes relies on business key performance indicators (Biz-KPIs), that can be negatively impacted by IT failures. Business and IT Observability (BizITObs) data fuses both Biz-KPIs and IT event channels together as multivariate time series data. Forecasting Biz-KPIs in advance can enhance efficiency and revenue through proactive corrective measures. However, BizITObs data generally exhibit both useful and noisy inter-channel interactions between Biz-KPIs and IT events that need to be effectively decoupled. This leads to suboptimal forecasting performance when existing multivariate forecasting models are employed. To address this, we introduce AutoMixer, a time-series Foundation Model (FM) approach, grounded on the novel technique of channel-compressed pretrain and finetune workflows. AutoMixer leverages an AutoEncoder for channel-compressed pretraining and integrates it with the advanced TSMixer model for multivariate time series forecasting. This fusion greatly enhances the potency of TSMixer for accurate forecasts and also generalizes well across several downstream tasks. Through detailed experiments and dashboard analytics, we show AutoMixer's capability to consistently improve the Biz-KPI's forecasting accuracy (by 11-15%) which directly translates to actionable business insights.
Santosh Palaskar, Vijay Ekambaram, Arindam Jati, Neelamadhav Gantayat, Avirup Saha, Seema Nagar, Nam H. Nguyen, Pankaj Dayama 0001, Renuka Sindhgatta, Prateeti Mohapatra, Jayant Kalagnanam, Nandyala Hemachandra, Narayan Rangaraj
AAAI3
2024 Tiny Time Mixers (TTMs): Fast Pre-trained Models for Enhanced Zero/Few-Shot Forecasting of Multivariate Time Series
abstract
Large pre-trained models excel in zero/few-shot learning for language and vision tasks but face challenges in multivariate time series (TS) forecasting due to diverse data characteristics. Consequently, recent research efforts have focused on developing pre-trained TS forecasting models. These models, whether built from scratch or adapted from large language models (LLMs), excel in zero/few-shot forecasting tasks. However, they are limited by slow performance, high computational demands, and neglect of cross-channel and exogenous correlations. To address this, we introduce Tiny Time Mixers (TTM), a compact model (starting from 1M parameters) with effective transfer learning capabilities, trained exclusively on public TS datasets. TTM, based on the light-weight TSMixer architecture, incorporates innovations like adaptive patching, diverse resolution sampling, and resolution prefix tuning to handle pre-training on varied dataset resolutions with minimal model capacity. Additionally, it employs multi-level modeling to capture channel correlations and infuse exogenous signals during fine-tuning. TTM outperforms existing popular benchmarks in zero/few-shot forecasting by (4-40\%), while reducing computational requirements significantly. Moreover, TTMs are lightweight and can be executed even on CPU-only machines, enhancing usability and fostering wider adoption in resource-constrained environments. The model weights for reproducibility and research use are available at https://huggingface.co/ibm/ttm-research-r2/, while enterprise-use weights under the Apache license can be accessed as follows: the initial TTM-Q variant at https://huggingface.co/ibm-granite/granite-timeseries-ttm-r1, and the latest variants (TTM-B, TTM-E, TTM-A) weights are available at https://huggingface.co/ibm-granite/granite-timeseries-ttm-r2. The source code for the TTM model along with the usage scripts are available at https://github.com/ibm-granite/granite-tsfm/tree/main/tsfm_public/models/tinytimemixer
Vijay Ekambaram, Arindam Jati, Pankaj Dayama 0001, Sumanta Mukherjee, Wesley M. Gifford, Chandra Reddy, Jayant Kalagnanam
NeurIPS2
2023 TSMixer: Lightweight MLP-Mixer Model for Multivariate Time Series Forecasting
abstract
Transformers have gained popularity in time series forecasting for their ability to capture long-sequence interactions. However, their memory and compute-intensive requirements pose a critical bottleneck for long-term forecasting, despite numerous advancements in compute-aware self-attention modules. To address this, we propose TSMixer, a lightweight neural architecture exclusively composed of multi-layer perceptron (MLP) modules. TSMixer is designed for multivariate forecasting and representation learning on patched time series, providing an efficient alternative to Transformers. Our model draws inspiration from the success of MLP-Mixer models in computer vision. We demonstrate the challenges involved in adapting Vision MLP-Mixer for time series and introduce empirically validated components to enhance accuracy. This includes a novel design paradigm of attaching online reconciliation heads to the MLP-Mixer backbone, for explicitly modeling the time-series properties such as hierarchy and channel-correlations. We also propose a Hybrid channel modeling approach to effectively handle noisy channel interactions and generalization across diverse datasets, a common challenge in existing patch channel-mixing methods. Additionally, a simple gated attention mechanism is introduced in the backbone to prioritize important features. By incorporating these lightweight components, we significantly enhance the learning capability of simple MLP structures, outperforming complex Transformer models with minimal computing usage. Moreover, TSMixer's modular design enables compatibility with both supervised and masked self-supervised learning methods, making it a promising building block for time-series Foundation Models. TSMixer outperforms state-of-the-art MLP and Transformer models in forecasting by a considerable margin of 8-60%. It also outperforms the latest strong benchmarks of Patch-Transformer models (by 1-2%) with a significant reduction in memory and runtime (2-3X).
Vijay Ekambaram, Arindam Jati, Phanwadee Sinthong, Jayant Kalagnanam
KDD2
2023 Hierarchical Proxy Modeling for Improved HPO in Time Series Forecasting
abstract
Selecting the right set of hyperparameters is crucial in time series forecasting. The classical temporal cross-validation framework for hyperparameter optimization (HPO) often leads to poor test performance because of a possible mismatch between validation and test periods. To address this test-validation mismatch, we propose a novel technique, H-Pro to drive HPO via test proxies by exploiting data hierarchies often associated with time series datasets. Since higher-level aggregated time series often show less irregularity and better predictability as compared to the lowest-level time series which can be sparse and intermittent, we optimize the hyperparameters of the lowest-level base-forecaster by leveraging the proxy forecasts for the test period generated from the forecasters at higher levels. H-Pro can be applied on any off-the-shelf machine learning model to perform HPO. We validate the efficacy of our technique with extensive empirical evaluation on five publicly available hierarchical forecasting datasets. Our approach outperforms existing state-of-the-art methods in Tourism, Wiki, and Traffic datasets, and achieves competitive result in Tourism-L dataset, without any model-specific enhancements. Moreover, our method outperforms the winning method of the M5 forecast accuracy competition.
Arindam Jati, Vijay Ekambaram, Shaonli Pal, Brian Quanz, Wesley M. Gifford, Pavithra Harsha, Stuart Siegel, Sumanta Mukherjee, Chandrasekhar Narayanaswami 0001
KDD1
2022 Physics-based multiple time-series univariate forecasting
abstract
The Koopman operator theory provides a recipe for data-driven analysis of dynamical systems. The forecasting problem addresses the future value estimation of an observation from its recent past. One common modeling approach is producing forecasts via modeling the data-generating process, where observations are a function of the state variables of the data-generating process. One major challenge to this is the lack of knowledge of the data-generating process. Delay embedding is a common approach in dynamical system analysis to approximate the state-space manifold from a few state value measurements. In this work, we propose a deep learning framework that employs delay embedding, and Koopman operator theory in the context of univariate forecasting to produce a long-term stable forecast. In this work, we empirically show the correctness of the proposed framework. Our study shows the proposed model can generalize across data-set and is capable of producing a stable long-range forecast.
Sumanta Mukherjee, Arindam Jati, Kanthi K. Sarpatwar, Roman Vaculín
IEEE Big Data2
2022 Distributed Incremental Machine Learning for Big Time Series Data
abstract
Today’s highly instrumented systems generate large amounts of time series data from many different domains. In order to create meaningful insights from these data, techniques are needed to handle the collection, processing, and analysis at scale. The high frequency and volume of data that is generated introduces several challenges including data transformation, managing concept drift, the operational cost of model re-training and tracking, and scaling hyperparameter optimization.Incremental machine learning can provide a viable solution to handle these kinds of data. Further, distributed machine learning can be an efficient technique to improve performance, increase accuracy, and scale to larger input sizes.In this paper, we introduce a framework that combines the computational capabilities of Apache Spark and the workflow parallelization of Ray for distributed incremental learning. We conduct an empirical analysis of our framework for time series forecasting using the Walmart M5 dataset. The system can perform a parameter search on streaming data with concept drift producing a robust pipeline that fits high-volume data effectively. The results are encouraging and substantiate system proficiency over traditional big data analysis approaches that exclusively use either offline or online training.
Dhaval Salwala, Seshu Tirupathi, Brian Quanz, Wesley M. Gifford, Stuart Siegel, Vijay Ekambaram, Arindam Jati
IEEE Big Data7
2021 Adversarial Defense for Deep Speaker Recognition Using Hybrid Adversarial Training
abstract
Deep neural network based speaker recognition systems can easily be deceived by an adversary using minuscule imperceptible perturbations to the input speech samples. These adversarial attacks pose serious security threats to the speaker recognition systems that use speech biometric. To address this concern, in this work, we propose a new defense mechanism based on a hybrid adversarial training (HAT) setup. In contrast to existing works on countermeasures against adversarial attacks in deep speaker recognition that only use class-boundary information by supervised cross-entropy (CE) loss, we propose to exploit additional information from supervised and unsupervised cues to craft diverse and stronger perturbations for adversarial training. Specifically, we employ multi-task objectives using CE, feature-scattering (FS), and margin losses to create adversarial perturbations and include them for adversarial training to enhance the robustness of the model. We conduct speaker recognition experiments on the Librispeech dataset, and compare the performance with state-of-the-art projected gradient descent (PGD)-based adversarial training which employs only CE objective. The proposed HAT improves adversarial accuracy by absolute 3.29% and 3.18% for PGD and Carlini-Wagner (CW) attacks respectively, while retaining high accuracy on benign examples.
Monisankha Pal, Arindam Jati, Raghuveer Peri, Chin-Cheng Hsu, Wael Abd-Almageed, Shri Narayanan
ICASSP2
2021 Adversarial attack and defense strategies for deep speaker recognition systems
Arindam Jati, Chin-Cheng Hsu, Monisankha Pal, Raghuveer Peri, Wael Abd-Almageed, Shri Narayanan
Comput. Speech Lang.1
2021 Temporal Dynamics of Workplace Acoustic Scenes: Egocentric Analysis and Prediction
abstract
Identification of the acoustic environment from an audio recording, also known as acoustic scene classification, is an active area of research. In this paper, we study dynamically-changing background acoustic scenes from the egocentric perspective of an individual in a workplace. In a novel data collection setup, wearable sensors were deployed on individuals to collect audio signals within a built environment, while Bluetooth-based hubs continuously tracked the individual's location which represents the acoustic scene at a certain time. The data of this paper come from 170 hospital workers gathered continuously during work shifts for a 10 week period. In the first part of our study, we investigate temporal patterns in the egocentric sequence of acoustic scenes encountered by an employee, and the association of those patterns with factors such as job-role and daily routine of the individual. Motivated by evidence of multifaceted effects of ambient sounds on human psychology, we also analyze the association of the temporal dynamics of the perceived acoustic scenes with particular behavioral traits of the individual. Experiments reveal rich temporal patterns in the acoustic scenes experienced by the individuals during their work shifts, and a strong association of those patterns with various constructs related to job-roles and behavior of the employees. In the second part of our study, we employ deep learning models to predict the temporal sequence of acoustic scenes from the egocentric audio signal. We propose a two-stage framework where a recurrent neural network is trained on top of the latent acoustic representations learned by a segment-level neural network. The experimental results show the efficacy of the proposed system in predicting sequence of acoustic scenes, highlighting the existence of underlying temporal patterns in the acoustic scenes experienced in workplace.
Arindam Jati, Amrutha Nadarajan, Raghuveer Peri, Karel Mundnich, Tiantian Feng, Benjamin Girault, Shri Narayanan
IEEE ACM Trans. Audio Speech Lang. Process.1
2020 Supervised Deep Hashing for Efficient Audio Event Retrieval
abstract
Efficient retrieval of audio events can facilitate real-time implementation of numerous query and search-based systems. This work investigates the potency of different hashing techniques for efficient audio event retrieval. Multiple state-of-the-art weak audio embeddings are employed for this purpose. The performance of four classical unsupervised hashing algorithms is explored as part of off-the-shelf analysis. Then, we propose a partially supervised deep hashing framework that transforms the weak embeddings into a low-dimensional space while optimizing for efficient hash codes. The model uses only a fraction of the available labels and is shown here to significantly improve the retrieval accuracy on two widely employed audio event datasets. The extensive analysis and comparison between supervised and unsupervised hashing methods presented here, give insights on the quantizability of audio embeddings. This work provides a first look in efficient audio event retrieval systems and hopes to set baselines for future research.
Arindam Jati, Dimitra Emmanouilidou
ICASSP1
2020 Robust Speaker Recognition Using Unsupervised Adversarial Invariance
abstract
In this paper, we address the problem of speaker recognition in challenging acoustic conditions using a novel method to extract robust speaker-discriminative speech representations. We adopt a recently proposed unsupervised adversarial invariance architecture to train a network that maps speaker embeddings extracted using a pretrained model onto two lower dimensional embedding spaces. The embedding spaces are learnt to disentangle speaker-discriminative information from all other information present in the audio recordings, without supervision about the acoustic conditions. We analyze the robustness of the proposed embeddings to various sources of variability present in the signal for speaker verification and unsupervised clustering tasks on a large-scale speaker recognition corpus. Our analyses show that the proposed system substantially outperforms the baseline in a variety of challenging acoustic scenarios. Furthermore, for the task of speaker diarization on a real-world meeting corpus, our system shows a relative improvement of 36% in the diarization error rate compared to the state-of-the-art baseline.
Raghuveer Peri, Monisankha Pal, Arindam Jati, Krishna Somandepalli, Shri Narayanan
ICASSP3
2019 Hierarchy-aware Loss Function on a Tree Structured Label Space for Audio Event Detection
abstract
The paper introduces a hierarchy-aware loss function in a Deep Neural Network for an audio event detection task that has a bi-level tree structured label space. The goal is not only to improve audio event detection performance at all levels in the label hierarchy, but also to produce better audio embeddings. We exploit the label tree structure to preserve that information in the hierarchy-aware loss function. Two different loss functions are separately employed. First, a triplet loss with probabilistic multi-level batch mining is introduced. Second, a quadruplet learning method is applied, which is a special case of generalized triplet learning for bi-level label taxonomy. The training is performed in a multi-task learning framework by jointly optimizing cross entropy based loss and hierarchy-aware loss function. The proposed method is found to outperform the baseline cross entropy based models at both levels of the hierarchy. The multi-task model is also able to learn better audio representations as observed in our clustering experiments. Moreover, the model is shown to transfer well when an out-of-domain dataset is used for evaluation.
Arindam Jati, Naveen Kumar 0004, Ruxin Chen, Panayiotis G. Georgiou
ICASSP1
2019 Multi-Task Discriminative Training of Hybrid DNN-TVM Model for Speaker Verification with Noisy and Far-Field Speech
Arindam Jati, Raghuveer Peri, Monisankha Pal, Tae Jin Park, Naveen Kumar 0004, Ruchir Travadi, Panayiotis G. Georgiou, Shri Narayanan
INTERSPEECH1
2019 Multiview Shared Subspace Learning Across Speakers and Speech Commands
Krishna Somandepalli, Naveen Kumar 0004, Arindam Jati, Panayiotis G. Georgiou, Shri Narayanan
INTERSPEECH3
2019 Neural Predictive Coding Using Convolutional Neural Networks Toward Unsupervised Learning of Speaker Characteristics
abstract
Learning speaker-specific features is vital in many applications like speaker recognition, diarization, and speech recognition. This paper provides a novel approach, we term neural predictive coding (NPC), to learn speaker-specific characteristics in a completely unsupervised manner from large amounts of unlabeled training data that even contain many non-speech events and multi-speaker audio streams. The NPC framework exploits the proposed short-term active-speaker stationarity hypothesis which assumes two temporally close short speech segments belong to the same speaker, and thus a common representation that can encode the commonalities of both the segments, should capture the vocal characteristics of that speaker. We train a convolutional deep siamese network to produce “speaker embeddings” by learning to separate “same” versus “different” speaker pairs which are generated from an unlabeled data of audio streams. Two sets of experiments are done in different scenarios to evaluate the strength of NPC embeddings and compare with state-of-the-art in-domain supervised methods. First, two speaker identification experiments with different context lengths are performed in a scenario with comparatively limited within-speaker channel variability. NPC embeddings are found to perform the best at short duration experiment, and they provide complementary information to i-vectors for full utterance experiments. Second, a large-scale speaker verification task having a wide range of within-speaker channel variability is adopted as an upper-bound experiment where comparisons are drawn with in-domain supervised methods.
Arindam Jati, Panayiotis G. Georgiou
IEEE ACM Trans. Audio Speech Lang. Process.1
2018 Towards Predicting Physiology from Speech During Stressful Conversations: Heart Rate and Respiratory Sinus Arrhythmia
abstract
Being affected by mental stress during conversations might have a direct or indirect effect on our speech acoustics as well as on our physiological responses. This paper presents a study on finding the relationship between these two modalities, speech acoustics and physiology, during stressful conversations between humans. Heart rate and respiratory sinus arrhythmia have been considered as physiological variables in the present study. Two datasets, one from stress induction sessions and the other one from in-lab discussions of relationship conflicts between couples, have been analyzed. A series of experiments have been performed separately on the two datasets, as well as on the combined dataset. The research finds acoustic features that are significantly correlated with the physiological variables during stressful conversations. It also predicts the physiological signals from speech features through a nonlinear regression analysis. The results take us one step forward towards building an extremely non-intrusive and relatively inexpensive method of predicting physiological responses from speech, and thus detecting the presence and quantifying the intensity of stress during stressful conversations.
Arindam Jati, Paula G. Williams, Brian R. Baucom, Panayiotis G. Georgiou
ICASSP1
2018 An Unsupervised Neural Prediction Framework for Learning Speaker Embeddings Using Recurrent Neural Networks
Arindam Jati, Panayiotis G. Georgiou
INTERSPEECH1
2017 Speaker2Vec: Unsupervised Learning and Adaptation of a Speaker Manifold Using Deep Neural Networks with an Evaluation on Speaker Segmentation
Arindam Jati, Panayiotis G. Georgiou
INTERSPEECH1
2014 Object-Shape Recognition by tactile Image Analysis using Support Vector Machine
abstract
The sense of touch is important to human to understand shape, texture, and hardness of the objects. An object under grip, i.e. object exploration by enclosure, provides a unique pressure distribution on the different regions of palm depending on its shape. This paper utilizes the above experience for recognition of object shapes by tactile image analysis. The high pressure regions (HPRs) are segmented and analyzed for object shape recognition rather than analyzing the entire image. Tactile images are acquired by capacitive tactile sensor while grasping a particular object. Geometrical features are extracted from the chain codes obtained by polygon approximation of the contours of the segmented HPRs. Two-level classification scheme using linear support vector machine (LSVM) is employed to classify the input feature vector in respective object shape classes with an average classification accuracy of 93.46% and computational time of 1.19 s for 12 different object shape classes. Our proposed two-level LSVM reduces the misclassification rates, thus efficiently recognizes various object shapes from the tactile images.
Anwesha Khasnobish, Arindam Jati, Amit Konar, D. N. Tibarewala
Int. J. Pattern Recognit. Artif. Intell.2
2012 A hybridisation of Improved Harmony Search and Bacterial Foraging for multi-robot motion planning
abstract
This paper provides a new approach to include the chemotactic behavior of Bacterial Foraging Algorithm (BFOA) in the existing Improved Harmony Search (IHS) algorithm. Extensive computer simulations with CEC-2005 benchmark functions reveal that the proposed algorithm outperforms the existing one with respect to accuracy in determining the optima. The proposed algorithm has successfully been implemented for multi-robot motion planning application. Performance has been studied using the proposed IHS-BFO algorithm and compared with existing IHS and Particle Swarm Optimization (PSO) algorithm.
Arindam Jati, Pratyusha Rakshit, Amit Konar, Atulya K. Nagar
IEEE Congress on Evolutionary Computation1
2012 Object-shape recognition from tactile images using a feed-forward neural network
abstract
The sense of touch is an extremely important sensory system in the human body which helps to understand object shape, texture, hardness in the world around us. Incorporating artificial haptic sensory systems in rehabilitative aids and in various other human computer interfaces is a thrust area of research presently. This paper presents a novel approach of shape recognition and classification from the tactile pressure images by touching the surface of various real life objects. Here four objects (viz. a planar surface, object with one edge, a cuboid i.e. object with two edges and a cylindrical object) are used for shape recognition. The obtained tactile pressure images of the object surfaces are subjected to segmentation, edge detection and a mapping procedure to finally reconstruct the particular object shapes. The reconstructed images are used as features. The processed tactile pressure images are classified with feed- forward neural network (FFNN) using extracted features. The classifier performance is tested with different signal-to-noise (SNR) ratios. Is is observed that classifier accuracy decreases with decrease in SNR, but at SNR value 6 i.e. when the noise power is one sixth of the signal power, the mean classification accuracy of the classifier is 88%. This shows the robustness of feed-forward neural network in the classification purpose. The performance of FFNN is compared with four classifiers (Linear Discriminant Analysis, Linear Support vector machine, Radial Basis Function SVM, k-Nearest Neighbor). FFNN performed best acquiring first rank with a average classification accuracy of 94.0%.
Anwesha Khasnobish, Arindam Jati, Saugat Bhattacharyya, Amit Konar, D. N. Tibarewala, Atulya K. Nagar
IJCNN2