VLDB 2026 Research / reviewers in the wild / expert
Vikram Gupta
dblp:65/6215
· DBLP profile ↗
17ranked-venue papers
11as first author
6since 2021 · last 2022
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 8 · 5 first-author · 5 since 2021Artificial intelligence and machine learning · 7 · 4 first-author · 5 since 2021Computer networks · 4 · 3 first-authorSystems, architecture and hardware · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2022 | Deep Clustering of Text Representations for Supervision-Free Probing of SyntaxabstractWe explore deep clustering of multilingual text representations for unsupervised model interpretation and induction of syntax. As these representations are high-dimensional, out-of-the-box methods like K-means do not work well. Thus, our approach jointly transforms the representations into a lower-dimensional cluster-friendly space and clusters them. We consider two notions of syntax: Part of Speech Induction (POSI) and Constituency Labelling (CoLab) in this work. Interestingly, we find that Multilingual BERT (mBERT) contains surprising amount of syntactic knowledge of English; possibly even as much as English BERT (E-BERT). Our model can be used as a supervision-free probe which is arguably a less-biased way of probing. We find that unsupervised probes show benefits from higher layers as compared to supervised probes. We further note that our unsupervised probe utilizes E-BERT and mBERT representations differently, especially for POSI. We validate the efficacy of our probe by demonstrating its capabilities as a unsupervised syntax induction technique. Our probe works well for both syntactic formalisms by simply adapting the input representations. We report competitive performance of our probe on 45-tag English POSI, state-of-the-art performance on 12-tag POSI across 10 languages, and competitive results on CoLab. We also perform zero-shot syntax induction on resource impoverished languages and report strong results. Vikram Gupta, Freda Shi, Kevin Gimpel, Mrinmaya Sachan |
AAAI | 1 |
| 2022 | 3MASSIV: Multilingual, Multimodal and Multi-Aspect dataset of Social Media Short VideosabstractWe present 3MASSIV, a multilingual, multimodal and multi-aspect, expertly-annotated dataset of diverse short videos extracted from short-video social media platform - Moj. 3MASSIV comprises of 50k short videos (20 seconds average duration) and 100K unlabeled videos in 11 different languages and captures popular short video trends like pranks, fails, romance, comedy expressed via unique audio-visual formats like self-shot videos, reaction videos, lip-synching, self-sung songs, etc. 3MASSIV presents an opportunity for multimodal and multilingual semantic understanding on these unique videos by annotating them for concepts, affective states, media types, and audio language. We present a thorough analysis of 3MASSIV and highlight the variety and unique aspects of our dataset compared to other contemporary popular datasets with strong baselines. We also show how the social media content in 3MASSIV is dynamic and temporal in nature, which can be used for semantic understanding tasks and cross-lingual analysis. Vikram Gupta, Trisha Mittal, Puneet Mathur, Mayank Maheshwari, Aniket Bera, Debdoot Mukherjee, Dinesh Manocha |
CVPR | 1 |
| 2022 | ADIMA: Abuse Detection In Multilingual AudioabstractAbusive content detection in spoken text can be addressed by performing Automatic Speech Recognition (ASR) and leveraging advancements in natural language processing. However, ASR models introduce latency and often perform sub-optimally for abusive words as they are underrepresented in training corpora and not spoken clearly or completely. Exploration of this problem entirely in the audio domain has largely been limited by the lack of audio datasets. Building on these challenges, we propose ADIMA, a novel, linguistically diverse, ethically sourced, expert annotated and well- balanced multilingual abuse detection audio dataset comprising of 11,775 audio samples in 10 Indic languages spanning 65 hours and spoken by 6,446 unique users. Through quantitative experiments across monolingual and cross-lingual zeroshot settings, we take the first step in democratizing audio based content moderation in Indic languages and set forth our dataset to pave future work. Dataset and code are available at: https://github.com/ShareChatAI/Adima Vikram Gupta, Rini A. Sharon, Ramit Sawhney, Debdoot Mukherjee |
ICASSP | 1 |
| 2022 | Multilingual and Multimodal Abuse DetectionabstractThe presence of abusive content on social media platforms is undesirable as it severely impedes healthy and safe social media interactions.While automatic abuse detection has been widely explored in textual domain, audio abuse detection still remains unexplored.In this paper, we attempt abuse detection in conversational audio from a multimodal perspective in a multilingual social media setting.Our key hypothesis is that along with the modelling of audio, incorporating discriminative information from other modalities can be highly beneficial for this task.Our proposed method, MADA, explicitly focuses on two modalities other than the audio itself, namely, the underlying emotions expressed in the abusive audio and the semantic information encapsulated in the corresponding textual form.Observations prove that MADA demonstrates gains over audio-only approaches on the ADIMA dataset.We test the proposed approach on 10 different languages and observe consistent gains in the range 0.6%-5.2%by leveraging multiple modalities.We also perform extensive ablation experiments for studying the contributions of every modality and observe the best results while leveraging all the modalities together.Additionally, we perform experiments to empirically confirm that there is a strong correlation between underlying emotions and abusive behaviour. Rini A. Sharon, Heet Shah, Debdoot Mukherjee, Vikram Gupta |
INTERSPEECH | 4 |
| 2022 | Multilingual Abusive Comment Detection at Scale for Indic LanguagesabstractSocial media platforms were conceived to act as online town squares' where people could get together, share information and communicate with each other peacefully. However, harmful content borne out of bad actors are constantly plaguing these platforms slowly converting them intomosh pits' where the bad actors take the liberty to extensively abuse various marginalised groups. Accurate and timely detection of abusive content on social media platforms is therefore very important for facilitating safe interactions between users. However, due to the small scale and sparse linguistic coverage of Indic abusive speech datasets, development of such algorithms for Indic social media users (one-sixth of global population) is severely impeded.To facilitate and encourage research in this important direction, we contribute for the first time MACD - a large-scale (150K), human-annotated, multilingual (5 languages), balanced (49\% abusive content) and diverse (70K users) abuse detection dataset of user comments, sourced from a popular social media platform - ShareChat. We also release AbuseXLMR, an abusive content detection model pretrained on large number of social media comments in 15+ Indic languages which outperforms XLM-R and MuRIL on multiple Indic datasets. Along with the annotations, we also release the mapping between comment, post and user id's to facilitate modelling the relationship between them. We share competitive monolingual, cross-lingual and few-shot baselines so that MACD can be used as a dataset benchmark for future research. Vikram Gupta, Sumegh Roychowdhury, Mithun Das, Somnath Banerjee 0002, Punyajoy Saha, Binny Mathew, Hastagiri Prakash Vanchinathan, Animesh Mukherjee 0001 |
NeurIPS | 1 |
| 2021 | MeGA-CDA: Memory Guided Attention for Category-Aware Unsupervised Domain Adaptive Object DetectionabstractExisting approaches for unsupervised domain adaptive object detection perform feature alignment via adversarial training. While these methods achieve reasonable improvements in performance, they typically perform category-agnostic domain alignment, thereby resulting in negative transfer of features. To overcome this issue, in this work, we attempt to incorporate category information into the domain adaptation process by proposing Memory Guided Attention for Category-Aware Domain Adaptation (MeGA-CDA). The proposed method consists of employing category-wise discriminators to ensure category-aware feature alignment for learning domain-invariant discriminative features. However, since the category information is not available for the target samples, we propose to generate memory-guided category-specific attention maps which are then used to route the features appropriately to the corresponding category discriminator. The proposed method is evaluated on several benchmark datasets and is shown to outperform existing approaches. Vibashan VS, Vikram Gupta, Poojan Oza, Vishwanath A. Sindagi, Vishal M. Patel |
CVPR | 2 |
| 2019 | Progression Modelling for Online and Early Gesture DetectionabstractOnline and Early detection of gestures is crucial for building touchless gesture based interfaces. These interfaces should operate on a stream of video frames instead of the complete video and detect the presence of gestures at an earlier stage than post-completion for providing real time user experience. To achieve this, it is important to recognize the progression of the gesture across different stages so that appropriate responses can be triggered on reaching the desired execution stage. To address this, we propose a simple yet effective multi-task learning framework which models the progression of the gesture along with frame level recognition. The proposed framework recognizes the gestures at an early stage with high precision and also achieves state-of-the-art recognition accuracy of 87.8% which is closer to human accuracy of 88.4% on the NVIDIA gesture dataset in the offline configuration and advances the state-of-the-art by more than 4%. We also introduce tightly segmented annotations for the NVIDIA gesture dataset and setup a strong baseline for gesture localization for this dataset. We also evaluate our framework on the Montalbano dataset and report competitive results. Vikram Gupta, Saikumar Dwivedi, Rishabh Dabral, Arjun Jain |
3DV | 1 |
| 2019 | Out-Of-Distribution Detection for Generalized Zero-Shot Action RecognitionabstractGeneralized zero-shot action recognition is a challenging problem, where the task is to recognize new action categories that are unavailable during the training stage, in addition to the seen action categories. Existing approaches suffer from the inherent bias of the learned classifier towards the seen action categories. As a consequence, unseen category samples are incorrectly classified as belonging to one of the seen action categories. In this paper, we set out to tackle this issue by arguing for a separate treatment of seen and unseen action categories in generalized zero-shot action recognition. We introduce an out-of-distribution detector that determines whether the video features belong to a seen or unseen action category. To train our out-of-distribution detector, video features for unseen action categories are synthesized using generative adversarial networks trained on seen action category features. To the best of our knowledge, we are the first to propose an out-of-distribution detector based GZSL framework for action recognition in videos. Experiments are performed on three action recognition datasets: Olympic Sports, HMDB51 and UCF101. For generalized zero-shot action recognition, our proposed approach outperforms the baseline with absolute gains (in classification accuracy) of 7.0%, 3.4%, and 4.9%, respectively, on these datasets. Devraj Mandal, Sanath Narayan, Saikumar Dwivedi, Vikram Gupta, S. Shuaib Ahmed, Fahad Shahbaz Khan, Ling Shao 0001 |
CVPR | 4 |
| 2019 | A Hybrid Framework for Functional Verification using Reinforcement Learning and Deep LearningabstractIn this paper, we propose a novel hybrid verification framework (HVF) which uses Reinforcement Learning (RL) and Deep Neural Networks (DNNs) to accelerate the verification of complex systems. More precisely, our HVF incorporates RL to generate all possible sequences of vectors needed to approach a target state as well as the corresponding path to the target state which contains a potential design error. Furthermore, HVF utilizes DNNs to accelerate the verification of complex data paths in the target states. We have tested our framework on several circuits including multi-core designs as well as bus-arbiters and confirmed its significant verification speedup when compared to prior work. For example, HVF provides a total speedup of 4.5x for a quad-core MIPS processor verification. Karunveer Singh, Vikram Gupta, Arash Fayyazi, Massoud Pedram, Shahin Nazarian |
ACM Great Lakes Symposium on VLSI | 3 |
| 2014 | Feature Extraction in Densely Sensed EnvironmentsabstractWith the reduction in size and cost of sensor nodes, dense sensor networks are becoming more popular in a wide-range of applications. Many such applications with dense deployments are geared towards finding various patterns or features such as peaks, boundaries and shapes in the spread of sensed physical quantities over an area. However, collecting all the data from individual sensor nodes can be impractical both in terms of timing requirements and the overall resource consumption. Hence, it is imperative to devise distributed information processing techniques that can help in identifying such features with a high accuracy and within certain time constraints. In this paper, we exploit the prioritized channel-access mechanism of dominance-based Medium Access Control (MAC) protocols to efficiently obtain exterma of the sensed quantities. We show how by the use of simple transforms that sensor nodes employ on local data it is also possible to efficiently extract certain features such as local extrema and boundaries of events. Using these transformations, we show through extensive evaluations that our proposed technique is fast and efficient at retrieving only sensor data point with the most constructive information, independent of the number of sensor nodes in the network. Maryam Vahabi, Vikram Gupta, Michele Albano, Eduardo Tovar |
DCOSS | 2 |
| 2014 | Poster abstract: a harmony of sensors: achieving determinism in multi-application sensor networks
Vikram Gupta, Nuno Pereira 0001, Eduardo Tovar, Ragunathan Rajkumar |
IPSN | 1 |
| 2014 | Network-Harmonized Scheduling for multi-application sensor networksabstractSupport for multiple concurrent applications is an important enabler for promoting the use of sensor networks as an infrastructure technology, where multiple users can deploy their applications independently. In such a scenario, different applications on a node may transmit packets at distinct periods, causing the node to change from sleep to active state more often, which negatively impacts the energy consumption of the whole network. In this paper, we propose to batch the transmissions together by defining a harmonizing period to align the transmissions from multiple applications at periodic boundaries. This harmonizing period is then leveraged to design a protocol that coordinates the transmissions across nodes and provides real-time guarantees in a multi-hop network. This protocol, which we call Network- Harmonized Scheduling (NHS), takes advantage of the periodicity introduced to assign offsets to nodes at different hop-levels such that collisions are always avoided, and deterministic behavior is enforced. NHS is a light-weight and distributed protocol that does not require any global state-keeping mechanism. We implemented NHS on the Contiki operating system and show how it can achieve a duty-cycle comparable to an ideal TDMA approach. Vikram Gupta, Nuno Pereira 0001, Shashank Gaur, Eduardo Tovar, Ragunathan Rajkumar |
RTCSA | 1 |
| 2011 | A Language Independent Approach to Audio Search
Vikram Gupta, Jitendra Ajmera, Arun Kumar 0007, Ashish Verma 0001 |
INTERSPEECH | 1 |
| 2011 | A Framework for Programming Sensor Networks with Scheduling and Resource-Sharing OptimizationsabstractSeveral projects in the recent past have aimed at promoting Wireless Sensor Networks as an infrastructure technology, where several independent users can submit applications that execute concurrently across the network. Concurrent multiple applications cause significant energy-usage overhead on sensor nodes, that cannot be eliminated by traditional schemes optimized for single-application scenarios. In this paper, we outline two main optimization techniques for reducing power consumption across applications. First, we describe a compiler based approach that identifies redundant sensing requests across applications and eliminates those. Second, we cluster the radio transmissions together by concatenating packets from independent applications based on Rate-Harmonized Scheduling. Vikram Gupta, Eduardo Tovar, Karthik Lakshmanan, Ragunathan Rajkumar |
RTCSA (2) | 1 |
| 2011 | Nano-CF: A coordination framework for macro-programming in Wireless Sensor NetworksabstractWireless Sensor Networks (WSN) are being used for a number of applications involving infrastructure monitoring, building energy monitoring and industrial sensing. The difficulty of programming individual sensor nodes and the associated overhead have encouraged researchers to design macro-programming systems which can help program the network as a whole or as a combination of subnets. Most of the current macro-programming schemes do not support multiple users seamlessly deploying diverse applications on the same shared sensor network. As WSNs are becoming more common, it is important to provide such support, since it enables higher-level optimizations such as code reuse, energy savings, and traffic reduction. In this paper, we propose a macro-programming framework called Nano-CF, which, in addition to supporting in-network programming, allows multiple applications written by different programmers to be executed simultaneously on a sensor networking infrastructure. This framework enables the use of a common sensing infrastructure for a number of applications without the users being concerned about the applications already deployed on the network. The framework also supports timing constraints and resource reservations using the Nano-RK operating system. Nano-CF is efficient at improving WSN performance by (a) combining multiple user programs, (b) aggregating packets for data delivery, and (c) satisfying timing and energy specifications using Rate-Harmonized Scheduling. Using representative applications, we demonstrate that Nano-CF achieves 90% reduction in Source Lines-of-Code (SLoC) and 50% energy savings from aggregated data delivery. Vikram Gupta, Junsung Kim 0001, Aditi Pandya, Karthik Lakshmanan, Ragunathan Rajkumar, Eduardo Tovar |
SECON | 1 |
| 2009 | Low-power clock synchronization using electromagnetic energy radiating from AC power linesabstractClock synchronization is highly desirable in many sensor networking applications. It enables event ordering, coordinated actuation, energy-efficient communication and duty cycling. This paper presents a novel low-power hardware module for achieving global clock synchronization by tuning to the magnetic field radiating from existing AC power lines. This signal can be used as a global clock source for battery-operated sensor nodes to eliminate drift between nodes over time even when they are not passing messages. With this scheme, each receiver is frequency-locked with each other, but there is typically a phase-offset between them. Since these phase offsets tend to be constant, a higher-level compensation protocol can be used to globally synchronize a sensor network. We present the design of an LC tank receiver circuit tuned to the AC 60Hz signal which we call a Syntonistor. The Syntonistor incorporates a low-power microcontroller that filters the signal induced from AC power lines generating a pulse-per-second output for easy interfacing with sensor nodes. The hardware consumes less than 58μW which is 2--3 times lower than the idle state of most sensor networking MAC protocols. Next, we evaluate a software clock-recovery technique running on the local microcontroller that minimizes timing jitter and provides robustness to noise. Finally, we provide a protocol that sets a global notion of time by accounting for phase-offsets. We evaluate the synchronization accuracy and energy performance as compared to in-band message passing schemes. The use of out-of-band signals for clock synchronization has the useful property of decoupling the synchronization scheme from any particular MAC protocol. Our experiments show that over a 11 day period, eight nodes distributed across the floor of the CIC building on Carnegie Mellon's campus remained synchronized on an average to less than 1ms without exchanging any radio messages beyond the initialization phase. Anthony Rowe 0001, Vikram Gupta, Ragunathan Rajkumar |
SenSys | 2 |
| 2004 | Improving the Performance of TCP in the Presence of Interacting UDP Flows in Ad Hoc Networks
Vikram Gupta, Srikanth V. Krishnamurthy, Michalis Faloutsos |
NETWORKING | 1 |