Samarjit Das

dblp:96/6697 · DBLP profile ↗
← Back
20ranked-venue papers
5as first author
7since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 18 · 4 first-author · 7 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 3 since 2021
YearPublicationVenuePosition
2024 CLAP4Emo: ChatGPT-Assisted Speech Emotion Retrieval with Natural Language Supervision
abstract
Speech emotion retrieval is an important technique for large-scale and high-quality data collection. Conventional approach using ensemble of classification models might limit the retrieved emotion diversity and/or underperform in out-of-domain acoustic conditions. Natural language is diverse and agnostic to specific acoustic concepts, embedding a huge potential for developing language-based speech emotion retrieval system. In this paper we introduce CLAP4Emo, a novel framework to retrieve emotional speech via natural language prompts based on contrastive language-audio pretraining. To compensate for the absence of training captions in existing public datasets, we propose a systematic framework that applies ChatGPT to generate emotion captions. The experimental results demonstrate that our method can effectively improve the retrieved sample diversity while maintaining high precision across five benchmark datasets. By leveraging large language models, we establish a connection between audio and language for emotion description, culminating in an intuitive and interactive retrieval system. We release the generated emotion captions at: https://github.com/boschresearch/soundsee-emo-caps
Shabnam Ghaffarzadegan, Luca Bondi, Abinaya Kumar, Samarjit Das, Ho-Hsiang Wu
ICASSP5
2024 Sound of Traffic: A Dataset for Acoustic Traffic Identification and Counting
Shabnam Ghaffarzadegan, Luca Bondi, Abinaya Kumar, Ho-Hsiang Wu, Hans-Georg Horst, Samarjit Das
INTERSPEECH7
2023 Active Learning for Abnormal Lung Sound Data Curation and Detection in Asthma
Shabnam Ghaffarzadegan, Luca Bondi, Ho-Hsiang Wu, Sirajum Munir, Kelly J. Shields, Samarjit Das, Joseph Aracri
INTERSPEECH6
2022 Acoustic Imaging Aboard The International Space Station (ISS): Challenges and Preliminary Results
abstract
Design and execution of high fidelity acoustic sensing in complex environments poses a number of practical challenges, from accurately measuring the geometry of the setup and estimating the channel response, to time synchronization amongst the sources and receivers. When acoustic experiments are performed on-board the International Space Stations (ISS), the number of constraints and obstacles vastly increases, due to the combination of a highly unpredictable acoustic environment, and restricted availability of crew time. In this paper, we present our preliminary results with a first-of-a-kind acoustic imaging experiment performed aboard the ISS, highlighting the difference between simulations, laboratory measurements, and in-space experiments. We hope that these experiments and results will help the research community in realizing high performance acoustic imaging capabilities in complex environments.
Luca Bondi, Gabriel Chuang, Christopher Ick, Adarsh Dave, Charles Shelton, Brian Coltin, Trey Smith, Samarjit Das
ICASSP8
2022 Urban Sound & Sight: Dataset And Benchmark For Audio-Visual Urban Scene Understanding
abstract
Automatic audio-visual urban traffic understanding is a growing area of research with many potential applications of value to industry, academia, and the public sector. Yet, the lack of well-curated resources for training and evaluating models to research in this area hinders their development. To address this we present a curated audio-visual dataset, Urban Sound & Sight (Urbansas), developed for investigating the detection and localization of sounding vehicles in the wild. Urbansas consists of 12 hours of unlabeled data along with 3 hours of manually annotated data, including bounding boxes with classes and unique id of vehicles, and strong audio labels featuring vehicle types and indicating off-screen sounds. We discuss the challenges presented by the dataset and how to use its annotations for the localization of vehicles in the wild through audio models.
Magdalena Fuentes, Bea Steers, Pablo Zinemanas, Martín Rocamora, Luca Bondi, Julia Wilkins, Qianyi Shi, Yao Hou, Samarjit Das, Xavier Serra, Juan Pablo Bello
ICASSP9
2022 Learning to Adapt to Domain Shifts with Few-shot Samples in Anomalous Sound Detection
abstract
Anomaly detection has many important applications, such as monitoring industrial equipment. Despite recent advances in anomaly detection with deep-learning methods, it is unclear how existing solutions would perform under out-of-distribution scenarios, e.g., due to shifts in machine load or environmental noise. Grounded in the application of machine health monitoring, we propose a framework that adapts to new conditions with few-shot samples. Building upon prior work, we adopt a classification-based approach for anomaly detection and show its equivalence to mixture density estimation of the normal samples. We incorporate an episodic training procedure to match the few-shot setting during inference. We define multiple auxiliary classification tasks based on meta-information and leverage gradient-based meta-learning to improve generalization to different shifts. We evaluate our proposed method on a recently-released dataset of audio measurements from different machine types. It improved upon two baselines by around 10% and is on par with best-performing model reported on the dataset.
Bingqing Chen, Luca Bondi, Samarjit Das
ICPR3
2021 Synthetic Aperture Acoustic Imaging with Deep Generative Model Based Source Distribution Prior
abstract
Acoustic imaging has a wide range of real-world applications such as machine health monitoring. Conventionally, large microphone arrays are utilized to achieve useful spatial resolution in the imaging process. The advent of location-aware autonomous mobile robotic platforms opens up unique opportunity to apply synthetic aperture techniques to the acoustic imaging problem. By leveraging motion and location cues as well as some available prior information on the source distribution, a small moving microphone array has the potential to achieve imaging resolution far beyond the physical aperture limits. In this work, we propose to image large acoustic sources with a combination of synthetic aperture and their geometric structures modeled by a conditional generative adversarial network (cGAN). The acoustic imaging problem is formulated as a linear inverse problem and solved with the gradient-based method. Numerical simulations show that our synthetic aperture imaging framework can reconstruct the acoustic source distribution from microphone recordings and outperform static microphone arrays.
Boqiang Fan, Samarjit Das
ICASSP2
2018 A Light-Weight Multimodal Framework for Improved Environmental Audio Tagging
abstract
The lack of strong labels has severely limited the state-of-the-art fully supervised audio tagging systems to be scaled to larger dataset. Meanwhile, audio-visual learning models based on unlabeled videos have been successfully applied to audio tagging, but they are inevitably resource hungry and require a long time to train. In this work, we propose a light-weight, multimodal framework for environmental audio tagging. The audio branch of the framework is a convolutional and recurrent neural network (CRNN) based on multiple instance learning (MIL). It is trained with the audio tracks of a large collection of weakly labeled YouTube video excerpts; the video branch uses pretrained state-of-the-art image recognition networks and word embeddings to extract information from the video track and to map visual objects to sound events. Experiments on the audio tagging task of the DCASE 2017 challenge show that the incorporation of video information improves a strong baseline audio tagging system by 5.3% in terms of F1score. The entire system can be trained within 6 hours on a single GPU, and can be easily carried over to other audio tasks such as speech sentimental analysis.
Juncheng Li 0001, Yun Wang 0005, Joseph Szurley, Florian Metze, Samarjit Das
ICASSP5
2018 Eventness: Object Detection on Spectrograms for Temporal Localization of Audio Events
abstract
In this paper, we introduce the concept of Eventness for audio event detection, which can, in part, be thought of as an analogue to Objectness from computer vision. The key observation behind the eventness concept is that audio events reveal themselves as 2-dimensional time-frequency patterns with specific textures and geometric structures in spectrograms. These time-frequency patterns can then be viewed analogously to objects occurring in natural images (with the exception that scaling and rotation invariance properties do not apply). With this key observation in mind, we pose the problem of detecting monophonic or polyphonic audio events as an equivalent visual object(s) detection problem under partial occlusion and clutter in spectrograms. We adapt a state-of-the-art visual object detection model to evaluate the audio event detection task on publicly available datasets. The proposed network has comparable results with a state-of-the-art baseline and is more robust on minority events. Provided large-scale datasets, we hope that our proposed conceptual model of eventness will be beneficial to the audio signal processing community towards improving performance of audio event detection.
Phuong Pham, Juncheng Li 0001, Joseph Szurley, Samarjit Das
ICASSP4
2018 Multiple Instance Deep Learning for Weakly Supervised Small-Footprint Audio Event Detection
abstract
State-of-the-art audio event detection (AED) systems rely on supervised learning using strongly labeled data. However, this dependence severely limits scalability to large-scale datasets where fine resolution annotations are too expensive to obtain. In this paper, we propose a small-footprint multiple instance learning (MIL) framework for multi-class AED using weakly annotated labels. The proposed MIL framework uses audio embeddings extracted from a pre-trained convolutional neural network as input features. We show that by using audio embeddings the MIL framework can be implemented using a simple DNN with performance comparable to recurrent neural networks. We evaluate our approach by training an audio tagging system using a subset of AudioSet, which is a large collection of weakly labeled YouTube video excerpts. Combined with a late-fusion approach, we improve the F1 score of a baseline audio tagging system by 17%. We show that audio embeddings extracted by the convolutional neural networks significantly boost the performance of all MIL models. This framework reduces the model complexity of the AED system and is suitable for applications where computational resources are limited.
Shao-Yen Tseng, Juncheng Li 0001, Yun Wang 0005, Florian Metze, Joseph Szurley, Samarjit Das
INTERSPEECH6
2017 Very deep convolutional neural networks for raw waveforms
abstract
Learning acoustic models directly from the raw waveform data with minimal processing is challenging. Current waveform-based models have generally used very few (~2) convolutional layers, which might be insufficient for building high-level discriminative features. In this work, we propose very deep convolutional neural networks (CNNs) that directly use time-domain waveforms as inputs. Our CNNs, with up to 34 weight layers, are efficient to optimize over very long sequences (e.g., vector of size 32000), necessary for processing acoustic waveforms. This is achieved through batch normalization, residual learning, and a careful design of down-sampling in the initial layers. Our networks are fully convolutional, without the use of fully connected layers and dropout, to maximize representation learning. We use a large receptive field in the first convolutional layer to mimic bandpass filters, but very small receptive fields subsequently to control the model capacity. We demonstrate the performance gains with the deeper models. Our evaluation shows that the CNN with 18 weight layers outperforms the CNN with 3 weight layers by over 15% in absolute accuracy for an environmental sound recognition task and is competitive with the performance of models using log-mel features.
Chia Dai, Shuhui Qu, Juncheng Li 0001, Samarjit Das
ICASSP5
2017 A comparison of Deep Learning methods for environmental sound detection
abstract
Environmental sound detection is a challenging application of machine learning because of the noisy nature of the signal, and the small amount of (labeled) data that is typically available. This work thus presents a comparison of several state-of-the-art Deep Learning models on the IEEE challenge on Detection and Classification of Acoustic Scenes and Events (DCASE) 2016 challenge task and data, classifying sounds into one of fifteen common indoor and outdoor acoustic scenes, such as bus, cafe, car, city center, forest path, library, train, etc. In total, 13 hours of stereo audio recordings are available, making this one of the largest datasets available.
Juncheng Li 0001, Florian Metze, Shuhui Qu, Samarjit Das
ICASSP5
2017 Improved scene identification and object detection on egocentric vision of daily activities
Gonzalo Vaca-Castano, Samarjit Das, Joao P. Sousa, Niels da Vitoria Lobo, Mubarak Shah
Comput. Vis. Image Underst.2
2015 Improving egocentric vision of daily activities
abstract
In this paper, we investigates the interplay between scene and objects on daily activities under egocentric vision constraints. The nature of egocentric vision implies that the identity of the current scene remains consistent for several frames. We showed that this constraint can be used to improve several scene identification baselines including the current state of the art scene identification method. We also show that the scene identity can be used to improve the object detection. In generic object detection, models for objects typically only considers local context, ignoring the global scene context; however in daily activities, objects are typically associated to particular types of scenes. We exploited this context clue to re-score the object detectors. Re-scoring function is learned from scene classifiers and object detectors in a validation set. In testing time, models of objects are weighted according to the scene identity score (context) of the tested frame, improving the object detection as measured by mAP, respect to object detectors without the scene identity clue. Our experiments were performed in the Activities of Daily Living (ADL) public dataset [1] which is a standard benchmark for egocentric vision.
Gonzalo Vaca-Castano, Samarjit Das, Joao P. Sousa
ICIP2
2013 Tracking sparse signal sequences from nonlinear/non-Gaussian measurements and applications in illumination-motion tracking
abstract
In this work, we develop algorithms for tracking time sequences of sparse spatial signals with slowly changing sparsity patterns, and other unknown states, from a sequence of nonlinear observations corrupted by (possibly) non-Gaussian noise. A key example of the above problem occurs in tracking moving objects across spatially varying illumination changes, where motion is the small dimensional state while the illumination image is the sparse spatial signal satisfying the slow-sparsity-pattern-change property.
Rituparna Sarkar, Samarjit Das, Namrata Vaswani
ICASSP2
2012 Multimodal feature analysis for quantitative performance evaluation of endotracheal intubation (ETI)
abstract
Endotracheal intubation (ETI) is a crucial medical procedure performed on critically ill patients. It involves insertion of a breathing tube into the trachea i.e. the windpipe connecting the larynx and the lungs. Often, this procedure is performed by the paramedics (aka providers) under challenging prehospital settings e.g. roadside, ambulances or helicopters. Successful intubations could be lifesaving, whereas, failed intubation could potentially be fatal. Under prehospital environments, ETI success rates among the paramedics are surprisingly low and this necessitates better training and performance evaluation of ETI skills. Currently, few objective metrics exist to quantify the differences in ETI techniques between providers. In this pilot study, we develop a quantitative framework for discriminating the kinematic characteristics of providers with different experience levels. The system utilizes statistical analysis on spatio-temporal multimodal features extracted from optical motion capture, accelerometers and electromyography (EMG) sensors. Our experiments involved three individuals performing intubations on a dummy, each with different levels of training. Quantitative performance analysis on multimodal features revealed distinctive differences among different skill levels. In future work, the feedback from these analysis could potentially be harnessed for enhanced ETI training.
Samarjit Das, Jestin N. Carlson, Fernando De la Torre, Paul E. Phrampus, Jessica K. Hodgins
ICASSP1
2012 Particle Filter With a Mode Tracker for Visual Tracking Across Illumination Changes
abstract
In this correspondence, our goal is to develop a visual tracking algorithm that is able to track moving objects in the presence of illumination variations in the scene and that is robust to occlusions. We treat the illumination and motion ( x-y translation and scale) parameters as the unknown "state" sequence. The observation is the entire image, and the observation model allows for occasional occlusions (modeled as outliers). The nonlinearity and multimodality of the observation model necessitate the use of a particle filter (PF). Due to the inclusion of illumination parameters, the state dimension increases, thus making regular PFs impractically expensive. We show that the recently proposed approach using a PF with a mode tracker can be used here since, even in most occlusion cases, the posterior of illumination conditioned on motion and the previous state is unimodal and quite narrow. The key idea is to importance sample on the motion states while approximating importance sampling by posterior mode tracking for estimating illumination. Experiments demonstrate the advantage of the proposed algorithm over existing PF-based approaches for various face and vehicle tracking. We are also able to detect illumination model changes, e.g., those due to transition from shadow to sunlight or vice versa by using the generalized expected log-likelihood statistics and successfully compensate for it without ever loosing track.
Samarjit Das, Amit A. Kale, Namrata Vaswani
IEEE Trans. Image Process.1
2010 Hiding information inside structured shapes
abstract
This paper describes a new technique for embedding a message within structured shapes. It is desired that any changes in the shape owing to the embedded message are invisible to a casual observer but detectable by a specialized decoder. The message embedding algorithm represents shape outlines as a set of cubic Bezier curves and straight line segments. By slightly perturbing the Bezier curves, a single shape can spawn a library of similar-looking shapes each corresponding to a unique message. This library is efficiently stored using Adaptively Sampled Distance Fields (ADFs) which also facilitate rendering of the modified shapes at the desired resolution and fidelity. Given any modified shape, a forensic detector applies Procrustes analysis to determine the embedded message. Results of an extensive subjective test confirm that the shape modifications are indeed unobtrusive. Further, to test the recovery of the message bits in noisy physical environments, a text document is put through a print-photocopy-scan process. Message recovery is found to be stable even after multiple rounds of photocopying.
Samarjit Das, Shantanu Rane, Anthony Vetro
ICASSP1
2010 Nonstationary Shape Activities: Dynamic Models for Landmark Shape Change and Applications
abstract
Our goal is to develop statistical models for the shape change of a configuration of "landmark" points (key points of interest) over time and to use these models for filtering and tracking to automatically extract landmarks, synthesis, and change detection. The term "shape activity" was introduced in recent work to denote a particular stochastic model for the dynamics of landmark shapes (dynamics after global translation, scale, and rotation effects are normalized for). In that work, only models for stationary shape sequences were proposed. But most "activities" of a set of landmarks, e.g., running, jumping, or crawling, have large shape changes with respect to initial shape and hence are nonstationary. The key contribution of this work is a novel approach to define a generative model for both 2D and 3D nonstationary landmark shape sequences. Greatly improved performance using the proposed models is demonstrated for sequentially filtering noise-corrupted landmark configurations to compute Minimum Mean Procrustes Square Error (MMPSE) estimates of the true shape and for tracking human activity videos, i.e., for using the filtering to predict the locations of the landmarks (body parts) and using this prediction for faster and more accurate landmarks extraction from the current image.
Samarjit Das, Namrata Vaswani
IEEE Trans. Pattern Anal. Mach. Intell.1
2008 Model-based compression of nonstationary landmark shape sequences
abstract
We have proposed a novel model-based compression technique for nonstationary landmark shape data extracted from video sequences. The main goal is to develop a technique for the compact storage of landmark shape data. We use nonstationary shape activity (NSSA) to model the shape sequences. The shape data is encoded by applying differential pulse code modulation (DPCM) on the shape velocity coefficients under the NSSA model. We have studied the system performance in terms of compressibility-distortion trade off. NSSA based compression technique has been compared with two other methods based on existing shape modeling techniques namely, stationary shape activity (SSA) and active shape model (ASM). We tested our system with landmark shape data extracted from multiple video sequences of the CMU mocap database. It was found that NSSA outperforms both SSA and ASM in terms of compressibility for a given distortion tolerance. Thus NSSA based compression technique could be very useful in the applications like storage of large volumes of biomedical landmarks' data.
Samarjit Das, Namrata Vaswani
ICIP1