Chandrashekhar Lavania

dblp:119/0800 · DBLP profile ↗
← Back
12ranked-venue papers
6as first author
8since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 7 · 3 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 2 · 2 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
YearPublicationVenuePosition
2025 Improving Lip-synchrony in Direct Audio-Visual Speech-to-Speech Translation
abstract
Audio-Visual Speech-to-Speech Translation (AVS2S) typically prioritizes improving translation quality and naturalness. However, an equally critical aspect in audio-visual content is lip-synchrony—ensuring that the movements of the lips match the spoken content—essential for maintaining realism in dubbed videos. Despite its importance, the inclusion of lip-synchrony constraints in AVS2S models has been largely overlooked. This study addresses this gap by integrating a lip-synchrony loss into the training process of AVS2S models. Our proposed method significantly enhances lip-synchrony in direct audio-visual speechto-speech translation, achieving an average LSE-D score of 10.67, representing a 9.2% reduction in LSE-D over a strong baseline across four language pairs. Additionally, it maintains the naturalness and high quality of the translated speech when overlaid onto the original video, without any degradation in translation quality.
Lucas Goncalves, Prashant Mathur, Xing Niu 0001, Chandrashekhar Lavania, Brady Houston, Srikanth Vishnubhotla, Lijia Sun, Anthony Ferritto
ICASSP4
2024 Perceptual Evaluation of Audio-Visual Synchrony Grounded in Viewers' Opinion Scores
Lucas Goncalves, Prashant Mathur, Chandrashekhar Lavania, Metehan Cekic, Marcello Federico, Kyu J. Han
ECCV (79)3
2024 Tackling Missing Modalities in Audio-Visual Representation Learning Using Masked Autoencoders
Georgios Chochlakis, Chandrashekhar Lavania, Prashant Mathur, Kyu J. Han
INTERSPEECH2
2023 Multi-Scale Compositional Constraints for Representation Learning on Videos
abstract
Combining simple concepts to form structured thoughts and decomposing complex concepts into their constituents is one key characteristic of human cognition. In this work we extract video representations by combining multi-scale processing with compositional constraints, i.e., we constrain the latent space created by the network so that coarse grained video features are composed from a set of fine-grained video features using simple functions. We integrate the proposed constraints in a state-of-the-art contrastive learning frame-work. In our ablations, we evaluate different formulations of the compositional constraints and composition functions. We evaluate the proposed approach for the downstream tasks of action detection in UCF-101, and video summarization in the SumMe dataset. We achieve significant improvements over the baseline, i.e., 3.9% and 6.3% relative improvements for UCF-101 and SumMe respectively, showcasing the importance of compositional video representations.
Georgios Paraskevopoulos, Chandrashekhar Lavania, Lovish Chum, Shiva Sundaram
ICASSP2
2023 Utility-Preserving Privacy-Enabled Speech Embeddings for Emotion Detection
Chandrashekhar Lavania, Sanjiv Das, Kyu J. Han
INTERSPEECH1
2022 Enhancing Contrastive Learning with Temporal Cognizance for Audio-Visual Representation Generation
abstract
Audio-visual data allows us to leverage different modalities for downstream tasks. The idea being individual streams can complement each other in the given task, thereby resulting in a model with improved performance. In this work, we present our experimental results on action recognition and video summarization tasks. The proposed modeling approach builds upon the recent advances in contrastive loss based audio-visual representation learning. Temporally cognizant audio-visual discrimination is achieved in a Transformer model by learning with a masked feature reconstruction loss over a fixed time window in addition to learning via contrastive loss. Overall, our results indicate that the addition of temporal information significantly improved the performance of the contrastive loss based framework. We achieve an action classification accuracy of 66.2% versus the next best baseline at 64.7% on the HMDB dataset. For video summarization, we attain an F1 score of 43.5 verses 42.2 on the SumMe dataset.
Chandrashekhar Lavania, Shiva Sundaram, Sundararajan Srinivasan, Katrin Kirchhoff
ICASSP1
2021 Constrained Robust Submodular Partitioning
abstract
In the robust submodular partitioning problem, we aim to allocate a set of items into $m$ blocks, so that the evaluation of the minimum block according to a submodular function is maximized. Robust submodular partitioning promotes the diversity of every block in the partition. It has many applications in machine learning, e.g., partitioning data for distributed training so that the gradients computed on every block are consistent. We study an extension of the robust submodular partition problem with additional constraints (e.g., cardinality, multiple matroids, and/or knapsack) on every block. For example, when partitioning data for distributed training, we can add a constraint that the number of samples of each class is the same in each partition block, ensuring data balance. We present two classes of algorithms, i.e., Min-Block Greedy based algorithms (with an $\Omega(1/m)$ bound), and Round-Robin Greedy based algorithms (with a constant bound) and show that under various constraints, they still have good approximation guarantees. Interestingly, while normally the latter runs in only weakly polynomial time, we show that using the two together yields strongly polynomial running time while preserving the approximation guarantee. Lastly, we apply the algorithms on a real-world machine learning data partitioning problem showing good results.
Shengjie Wang 0001, Tianyi Zhou 0001, Chandrashekhar Lavania, Jeff A. Bilmes
NeurIPS3
2021 A Practical Online Framework for Extracting Running Video Summaries under a Fixed Memory Budget
abstract
We study the problem of summarizing a video stream (potentially of unbounded duration) on the fly, where the summarization system must operate under a fixed memory budget and should produce an appropriate summary of its past at each time step. This problem is motivated by applications that have access to only limited memory and compute resources (e.g., sensor networks, smart devices and phones). We approach this problem as constrained streaming maximization of a natural submodular objective function. In particular, we propose a novel feature-based submodular function for the summarization objective — this function is instantiated by deep-model-learnt features. We solve the constrained streaming maximization problem with an algorithm that abides by the memory budget. Our streaming algorithm, is unique in that it uses both an “adding gain” (to determine something new) and a “swapping gain” (to determine something better) relative to a current summary. Based on a time dependent F-measure method to gauge the performance of streaming summarization techniques, we demonstrate that our approach provides significant improvement over a number of state-of-the-art baseline methods that utilize comparable resources on both the TVSum50 and SumMe data sets.
Chandrashekhar Lavania, Rishabh Iyer 0001, Jeff A. Bilmes
SDM1
2019 Fixing Mini-batch Sequences with Hierarchical Robust Partitioning
abstract
We propose a general and efficient hierarchical robust partitioning framework to generate a deterministic sequence of mini-batches, one that offers assurances of being high quality, unlike a randomly drawn sequence. We compare our deterministically generated mini-batch sequences to randomly generated sequences; we show that, on a variety of deep learning tasks, the deterministic sequences significantly beat the mean and worst case performance of the random sequences, and often outperforms the best of the random sequences. Our theoretical contributions include a new algorithm for the robust submodular partition problem subject to cardinality constraints (which is used to construct mini-batch sequences), and show in general that the algorithm is fast and has good theoretical guarantees; we also show a more efficient hierarchical variant of the algorithm with similar guarantees under mild assumptions.
Shengjie Wang 0001, Wenruo Bai, Chandrashekhar Lavania, Jeff A. Bilmes
AISTATS3
2019 Auto-Summarization: A Step Towards Unsupervised Learning of a Submodular Mixture
abstract
We introduce an approach that requires the specification of only a handful of hyperparameters to determine a mixture of submodular functions for use in data science applications. Two techniques, applied in succession, are used to achieve this. The first involves training an autoencoder neural network constrainedly so that the bottleneck features have the following characteristic: the larger a feature's value, the more an input sample should have an automatically learnt property. This is analogous to bag of-words features, but where the “words” are learnt automatically. The second technique instantiates a mixture of submodular functions, each of which consists of a concave composed with a modular function comprised of the learnt neural network features. We introduce a mixture weight learning approach that does not (as is common) directly utilize supervised summary information. Instead, it optimizes a set of meta-objectives each of which corresponds to a likely necessary condition on what constitutes a good summarization objective. While hyperparameter optimization is often the bane of unsupervised methods, our approach reduces the learning of a summarization function (which most generally involves learning 2n parameters) down to the problem of selecting only a handful of hyperparameters. Empirical results on three very different modalities of data (i.e., image, text, and machine learning training data) show that our method produces functions that perform significantly better than a variety of unsupervised baseline methods.
Chandrashekhar Lavania, Jeff A. Bilmes
SDM1
2017 Reducing total latency in online real-time inference and decoding via combined context window and model smoothing latencies
abstract
Real-time low-latency online inference and decoding in sequential probabilistic models are important in many interactive systems, including automatic speech recognition (ASR) and streaming environments. We study total inference latency (TL) in such systems, the additively combined latency of the inherent look-ahead of a deep neural network's (DNN) contextual window (CWL) in a DNN-HMM hybrid system and the latency incurred during Kalman-style smoothing in a dynamic probabilistic model (MSL) (hence, TL = CWL + MSL). For a fixed TL, the best accuracy can occur with a strictly positive MSL, often by quite a bit, a surprising result given the DNN's power. Furthermore, we find that accuracy is often improved with smaller TL and larger MSL. These results suggest that for optimal low-latency real-time decoding, the size of a DNN context window along with model smoothing should be jointly considered.
Chandrashekhar Lavania, Jeff A. Bilmes
ICASSP1
2016 A weakly supervised activity recognition framework for real-time synthetic biology laboratory assistance
abstract
We describe the design of a hybrid system -- a combination of a Dynamic Graphical Model (DGM) with a Deep Neural Network (DNN) -- to identify activities performed during synthetic biology experiments. The purpose is to provide real-time feedback to experimenters, thus helping to reduce human errors and improve experimental reproducibility. The data consists of unlabeled videos of recorded experiments and "weakly supervised" information (i.e., "theoretical" and asynchronous knowledge of sets of high level activity sequences in the experiment) used to train the system. Multiple activity sequences are modeled using a trellis, and deep features are extracted from video images. Model performance is accessed using real-time online statistical inference. The trellis incorporates variations during experiment execution, making our model very general and capable of high performance.
Chandrashekhar Lavania, Sunil Thulasidasan, Anthony LaMarca, Jeffrey Scofield, Jeff A. Bilmes
UbiComp1