VLDB 2026 Research / reviewers in the wild / expert
Agrim Gupta
dblp:200/8282
· DBLP profile ↗
27ranked-venue papers
11as first author
21since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 7 first-author · 12 since 2021Computer networks · 11 · 4 first-author · 9 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author · 1 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | PhaseMO: A Universal Massive MIMO Architecture for Sustainable NextG
Adel Heidari, Agrim Gupta, Ish Kumar Jain, Dinesh Bharadia |
INFOCOM | 2 |
| 2025 | Exploring Diffusion Transformer Designs via GraftingabstractDesigning model architectures requires decisions such as selecting operators (e.g., attention, convolution) and configurations (e.g., depth, width). However, evaluating the impact of these decisions on model quality requires costly pretraining, limiting architectural investigation.
Inspired by how new software is built on existing code, we ask: can new architecture designs be studied using pretrained models? To this end, we present *grafting*, a simple approach for editing pretrained diffusion transformers (DiTs) to materialize new architectures under small compute budgets. Informed by our analysis of activation behavior and attention locality, we construct a testbed based on the DiT-XL/2 design to study the impact of grafting on model quality. Using this testbed, we develop a family of hybrid designs via grafting: replacing softmax attention with gated convolution, local attention, and linear attention, and replacing MLPs with variable expansion ratio and convolutional variants. Notably, many hybrid designs achieve good quality (FID: 2.38–2.64 vs. 2.27 for DiT-XL/2)
using $<2$% pretraining compute. We then graft a text-to-image model (PixArt-$\Sigma$), achieving a 1.43$\times$ speedup with less than a 2% drop in GenEval score. Finally, we present a case study that restructures DiT-XL/2 by converting every pair of sequential transformer blocks into parallel blocks via grafting. This reduces model depth by 2$\times$ and yields better quality (FID: 2.77) than other models of comparable depth. Together, we show that new diffusion model designs can be explored by grafting pretrained DiTs, with edits ranging from operator replacement to architecture restructuring. Code and grafted models: https://grafting.stanford.edu. Keshigeyan Chandrasegaran, Michael Poli, Daniel Y. Fu, Lea M. Hadzic, Manling Li, Agrim Gupta, Stefano Massaroli, Azalia Mirhoseini, Juan Carlos Niebles, Stefano Ermon, Li Fei-Fei 0001 |
NeurIPS | 7 |
| 2025 | Demo Abstract - SIGAR: Sensor Integration Gateway using Augmented RealityabstractWe introduce SIGAR, a Sensor Integration Gateway using Augmented Reality, which combines RFID-based passive sensing with AR for real-time visualization. Using batteryless, wireless RFID sensors, SIGAR eliminates the need for power sources, enabling sustainable and cost-effective monitoring. A mobile app automatically detects sensors within the camera's field of view and overlays realtime sensory data onto the physical environment. Demonstrated through applications like force, soil moisture and light sensing, SIGAR provides intuitive, context-aware insights for environmental monitoring, inventory management, and more. This fusion of AR and passive sensing bridges digital and physical worlds, offering scalable, low-power IoT solutions. Ishan Bansal, Nagarjun Bhat, Agrim Gupta, Harine Govindarajan, Dinesh Bharadia |
SenSys | 3 |
| 2024 | Photorealistic Video Generation with Diffusion Models
Agrim Gupta, Lijun Yu, Kihyuk Sohn, Xiuye Gu, Meera Hahn, Li Fei-Fei 0001, Irfan A. Essa, Lu Jiang 0004, José Lezama |
ECCV (79) | 1 |
| 2024 | Language Model Beats Diffusion - Tokenizer is key to visual generationabstractWhile Large Language Models (LLMs) are the dominant models for generative tasks in language, they do not perform as well as diffusion models on image and video generation. To effectively use LLMs for visual generation, one crucial component is the visual tokenizer that maps pixel-space inputs to discrete tokens appropriate for LLM learning. In this paper, we introduce \modelname{}, a video tokenizer designed to generate concise and expressive tokens for both videos and images using a common token vocabulary. Equipped with this new tokenizer, we show that LLMs outperform diffusion models on standard image and video generation benchmarks including ImageNet and Kinetics. In addition, we demonstrate that our tokenizer surpasses the previously top-performing video tokenizer on two more tasks: (1) video compression comparable to the next-generation video codec (VCC) according to human evaluations, and (2) learning effective representations for action recognition tasks. Lijun Yu, José Lezama, Nitesh Bharadwaj Gundavarapu, Luca Versari, Kihyuk Sohn, David Minnen, Yong Cheng 0003, Agrim Gupta, Xiuye Gu, Alex Hauptmann 0001, Boqing Gong, Ming-Hsuan Yang 0001, Irfan A. Essa, David A. Ross, Lu Jiang 0004 |
ICLR | 8 |
| 2024 | VideoPoet: A Large Language Model for Zero-Shot Video GenerationabstractWe present VideoPoet, a language model capable of synthesizing high-quality video from a large variety of conditioning signals. VideoPoet employs a decoder-only transformer architecture that processes multimodal inputs – including images, videos, text, and audio. The training protocol follows that of Large Language Models (LLMs), consisting of two stages: pretraining and task-specific adaptation. During pretraining, VideoPoet incorporates a mixture of multimodal generative objectives within an autoregressive Transformer framework. The pretrained LLM serves as a foundation that can be adapted for a range of video generation tasks. We present empirical results demonstrating the model’s state-of-the-art capabilities in zero-shot video generation, specifically highlighting the ability to generate high-fidelity motions. Project page: http://sites.research.google/videopoet/ Dan Kondratyuk, Lijun Yu, Xiuye Gu, José Lezama, Jonathan Huang, Grant Schindler, Rachel Hornung, Vighnesh Birodkar, Jimmy Yan, Ming-Chang Chiu, Krishna Somandepalli, Hassan Akbari, Yair Alon, Yong Cheng 0003, Joshua V. Dillon, Agrim Gupta, Meera Hahn, Anja Hauth, David Hendon, Alonso Martinez, David Minnen, Mikhail Sirotenko, Kihyuk Sohn, Hartwig Adam, Ming-Hsuan Yang 0001, Irfan A. Essa, Huisheng Wang, David A. Ross, Bryan Seybold, Lu Jiang 0004 |
ICML | 16 |
| 2024 | Open X-Embodiment: Robotic Learning Datasets and RT-X Models : Open X-Embodiment CollaborationabstractLarge, high-capacity models trained on diverse datasets have shown remarkable successes on efficiently tackling downstream applications. In domains from NLP to Computer Vision, this has led to a consolidation of pretrained models, with general pretrained backbones serving as a starting point for many applications. Can such a consolidation happen in robotics? Conventionally, robotic learning methods train a separate model for every application, every robot, and even every environment. Can we instead train "generalist" X-robot policy that can be adapted efficiently to new robots, tasks, and environments? In this paper, we provide datasets in standardized data formats and models to make it possible to explore this possibility in the context of robotic manipulation, alongside experimental results that provide an example of effective X-robot policies. We assemble a dataset from 22 different robots collected through a collaboration between 21 institutions, demonstrating 527 skills (160266 tasks). We show that a high-capacity model trained on this data, which we call RT-X, exhibits positive transfer and improves the capabilities of multiple robots by leveraging experience from other platforms. The project website is robotics-transformer-x.github.io. Abigail O'Neill, Abhiram Maddukuri, Abhishek Gupta 0004, Abhishek Padalkar, Abraham Lee, Acorn Pooley, Agrim Gupta, Ajay Mandlekar, Ajinkya Jain, Albert Tung, Alex Bewley, Alex Irpan, Alexander Khazatsky, Anant Rai, Anchit Gupta, Andrew E. Wang, Anikait Singh, Animesh Garg, Aniruddha Kembhavi, Annie Xie, Anthony Brohan, Antonin Raffin, Archit Sharma, Arefeh Yavary, Arhan Jain, Ashwin Balakrishna, Ayzaan Wahid, Ben Burgess-Limerick, Bernhard Schölkopf, Blake Wulfe, Brian Ichter, Cewu Lu, Charles Xu 0003, Charlotte Le, Chelsea Finn, Chen Wang 0053, Chenfeng Xu, Cheng Chi 0001, Chenguang Huang, Christine Chan, Christopher Agia, Chuer Pan, Chuyuan Fu, Coline Devin, Danfei Xu, Daniel Morton, Danny Drieß, Daphne Chen, Deepak Pathak, Dhruv Shah, Dieter Büchler, Dinesh Jayaraman, Dmitry Kalashnikov, Dorsa Sadigh, Edward Johns, Ethan Paul Foster, Fangchen Liu, Federico Ceola, Fei Xia 0002, Feiyu Zhao, Freek Stulp, Gaoyue Zhou, Gaurav S. Sukhatme, Gautam Salhotra, Gilbert Feng, Giulio Schiavi, Glen Berseth, Gregory Kahn, Guanzhi Wang, Hao Su 0001, Haoshu Fang, Henghui Bao, Heni Ben Amor, Henrik I. Christensen, Hiroki Furuta, Homer Walke, Hongjie Fang, Huy Ha, Igor Mordatch, Ilija Radosavovic, Isabel Leal, Jacky Liang, Jad Abou-Chakra, Jaehyung Kim 0001, Jaimyn Drake, Jan Peters 0001, Jan Schneider 0007, Jasmine Hsu, Jeannette Bohg, Jeffrey T. Bingham, Jensen Gao, Jiaheng Hu, Jiajun Wu 0001, Jiankai Sun, Jianlan Luo, Jiayuan Gu, Jie Tan 0001, Jihoon Oh, Jimmy Wu, Jingpei Lu, Jitendra Malik, João Silvério, Joey Hejna, Jonathan Booher, Jonathan Tompson, Jonathan Yang, Jordi Salvador, Joseph J. Lim, Junhyek Han, Kanishka Rao, Karl Pertsch, Karol Hausman, Keegan Go, Keerthana Gopalakrishnan, Kenneth Y. Goldberg, Kendra Byrne, Kenneth Oslund, Kento Kawaharazuka, Kevin Black, Kevin Zhang 0002, Kiana Ehsani, Kiran Lekkala, Kirsty Ellis, Krishan Rana, Krishnan Srinivasan, Kuan Fang, Kunal Pratap Singh, Kuo-Hao Zeng, Kyle Hatch, Kyle Hsu, Laurent Itti, Yunliang Chen 0001, Lerrel Pinto, Li Fei-Fei 0001, Liam Tan, Linxi Fan, Lionel Ott, Lisa Lee, Luca Weihs, Magnum Chen, Marion Lepert, Marius Memmel, Masayoshi Tomizuka, Masha Itkina, Mateo Guaman Castro, Max Spero, Maximilian Du, Michael Ahn, Michael C. Yip, Mingtong Zhang 0003, Mingyu Ding, Minho Heo, Mohan Kumar Srirama, Mohit Sharma 0001, Moo Jin Kim, Naoaki Kanazawa, Nicklas Hansen 0001, Nicolas Heess, Nikhil J. Joshi, Niko Sünderhauf, Norman Di Palo, Nur Muhammad Shafiullah, Oier Mees, Oliver Kroemer, Osbert Bastani, Pannag R. Sanketi, Patrick Tree Miller, Patrick Yin, Paul Wohlhart, Peng Xu 0010, Peter David Fagan, Peter Mitrano, Pierre Sermanet, Pieter Abbeel, Priya Sundaresan, Qiuyu Chen, Rafael Rafailov, Ria Doshi, Roberto Martin Martin, Rohan Baijal, Rosario Scalise, Rose Hendrix, Roy Lin, Runjia Qian, Russell Mendonca, Rutav Shah, Ryan Hoque, Ryan Julian, Samuel Bustamante-Gomez, Sean Kirmani, Sergey Levine, Sherry Moore, Shikhar Bahl, Shivin Dass, Shubham D. Sonawani, Shuran Song, Sichun Xu, Siddhant Haldar, Siddharth Karamcheti, Simeon Adebola, Simon Guist, Soroush Nasiriany, Stefan Schaal, Stefan Welker, Stephen Tian, Subramanian Ramamoorthy, Sudeep Dasari, Suneel Belkhale, Sungjae Park, Suraj Nair 0003, Suvir Mirchandani, Takayuki Osa, Tanmay Gupta, Tatsuya Harada, Tatsuya Matsushima, Ted Xiao, Thomas Kollar, Tianhe Yu, Tianli Ding, Todor Davchev, Tony Z. Zhao, Travis Armstrong, Trevor Darrell, Trinity Chung, Vidhi Jain, Vincent Vanhoucke, Wolfram Burgard, Xiaolong Wang 0004, Xinghao Zhu, Xinyang Geng, Liangwei Xu, Yecheng Jason Ma 0001, Yejin Kim 0003, Yevgen Chebotar, Yilin Wu 0003, Yonatan Bisk, Yoonyoung Cho, Youngwoon Lee, Yuchen Cui, Yueh-Hua Wu, Yujin Tang, Yuke Zhu, Yunchu Zhang, Yunfan Jiang 0001, Yunshuang Li, Yunzhu Li, Yusuke Iwasawa, Yutaka Matsuo, Zehan Ma, Zichen Jeff Cui, Zichen Zhang 0016, Zipeng Lin |
ICRA | 8 |
| 2024 | ZenseTag: An RFID assisted Twin-Tag Single Antenna COTS Sensor InterfaceabstractSensing allows us to interact with and quantify the natural world. Despite the advancements in sensor versatility, sensing systems still suffer from limited adoption due to their dependence on batteries, complex interfaces, energy-harvesting modules, and readout latency. To address these challenges, we present ZenseTag --- a miniaturized, sticker-like platform that can interface commercial sensors directly with COTS RFID tags. ZenseTag exploits the impedance response of COTS sensors to the measured stimulus at Radio Frequencies, tuned to the UHF RFID band. It combines reliable hardware realization of differential analog sensing with robust software for accurate, low-latency sensor readouts, even in the presence of multipath effects. Ishan Bansal, Nagarjun Bhat, Agrim Gupta, Harine Govindarajan, Dinesh Bharadia |
MobiCom | 3 |
| 2024 | 3 W's of smartphone power consumption: Who, Where and How much is draining my battery?abstractWith 6.5 billion smartphones in use worldwide, each relying on a battery for key subsystems like display, compute, and cellular connectivity, previous studies on power consumption often used invalidated indirect estimates that failed to isolate specific hardware usage. We address this by utilizing Google's On Device Power Rails Monitor (ODPM) tool for precise power measurements of individual components. Our findings indicate that connectivity (Wi-Fi, 4G/5G) and screen display are the primary power consumers, as shown with the Google Pixel 7A. We also confirmed similar power consumption trends using an energy estimation method on the Samsung S23+. Given the prevalence of smartphones, we discuss the challenges and opportunities for optimizing power usage. Agrim Gupta, Adel Heidari, Avyakta Kalipattapu, Ish Kumar Jain, Dinesh Bharadia |
MobiCom | 1 |
| 2024 | HourVideo: 1-Hour Video-Language UnderstandingabstractWe present HourVideo, a benchmark dataset for hour-long video-language understanding. Our dataset consists of a novel task suite comprising summarization, perception (recall, tracking), visual reasoning (spatial, temporal, predictive, causal, counterfactual), and navigation (room-to-room, object retrieval) tasks. HourVideo includes 500 manually curated egocentric videos from the Ego4D dataset, spanning durations of 20 to 120 minutes, and features 12,976 high-quality, five-way multiple-choice questions. Benchmarking results reveal that multimodal models, including GPT-4 and LLaVA-NeXT, achieve marginal improvements over random chance. In stark contrast, human experts significantly outperform the state-of-the-art long-context multimodal model, Gemini Pro 1.5 (85.0\% vs. 37.3\%), highlighting a substantial gap in multimodal capabilities. Our benchmark, evaluation toolkit, prompts, and documentation are available at https://hourvideo.stanford.edu. Keshigeyan Chandrasegaran, Agrim Gupta, Lea M. Hadzic, Taran Kota, Jimming He, Cristóbal Eyzaguirre, Zane Durante, Manling Li, Jiajun Wu 0001, Li Fei-Fei 0001 |
NeurIPS | 2 |
| 2024 | A Versatile Diffusion Transformer with Mixture of Noise Levels for Audiovisual GenerationabstractTraining diffusion models for audiovisual sequences allows for a range of generation tasks by learning conditional distributions of various input-output combinations of the two modalities. Nevertheless, this strategy often requires training a separate model for each task which is expensive. Here, we propose a novel training approach to effectively learn arbitrary conditional distributions in the audiovisual space. Our key contribution lies in how we parameterize the diffusion timestep in the forward diffusion process. Instead of the standard fixed diffusion timestep, we propose applying variable diffusion timesteps across the temporal dimension and across modalities of the inputs. This formulation offers flexibility to introduce variable noise levels for various portions of the input, hence the term mixture of noise levels. We propose a transformer-based audiovisual latent diffusion model and show that it can be trained in a task-agnostic fashion using our approach to enable a variety of audiovisual generation tasks at inference time. Experiments demonstrate the versatility of our method in tackling cross-modal and multimodal interpolation tasks in the audiovisual space. Notably, our proposed approach surpasses baselines in generating temporally and perceptually consistent samples conditioned on the input. Project page: neurips13025.github.io Gwanghyun Kim, Alonso Martinez, Yu-Chuan Su, Brendan Jou, José Lezama, Agrim Gupta, Lijun Yu, Lu Jiang 0004, Aren Jansen, Jacob Walker, Krishna Somandepalli |
NeurIPS | 6 |
| 2024 | ZenseTag: An RFID assisted Twin-Tag Single Antenna COTS Sensor InterfaceabstractSensors enable us to digitally capture stimuli like moisture, light, and force. Despite their low cost, reliability, and scalability, the lack of widespread adoption of IoT has hindered the realization of true ubiquitous sensing. A likely reason is that the current sensor platforms are bulky due to the batteries and complex electronics needed to interface sensors communication systems. In this work, we present a fully-passive, miniaturized, flexible form factor sensor interface titled ZenseTag that uses minimal electronics to read and communicate analog sensor data, directly at radio frequencies (RF). We exploit the fundamental principle of resonance, where a sensor's terminal impedance becomes most sensitive to the measured stimulus at its resonant frequency. This enables ZenseTag to read out the sensor variation using only energy harvested from wireless signals. We demonstrate its implementation with a 15x10mm flexible PCB that connects sensors to a printed antenna and passive RFID ICs, enabling near real-time readout through a performant GUI-enabled software. Nagarjun Bhat, Agrim Gupta, Ishan Bansal, Harine Govindarajan, Dinesh Bharadia |
SenSys | 2 |
| 2023 | MaskViT: Masked Visual Pre-Training for Video Prediction
Agrim Gupta, Stephen Tian, Jiajun Wu 0001, Roberto Martin Martin, Li Fei-Fei 0001 |
ICLR | 1 |
| 2023 | VIMA: Robot Manipulation with Multimodal PromptsabstractPrompt-based learning has emerged as a successful paradigm in natural language processing, where a single general-purpose language model can be instructed to perform any task specified by input prompts. Yet task specification in robotics comes in various forms, such as imitating one-shot demonstrations, following language instructions, and reaching visual goals. They are often considered different tasks and tackled by specialized models. We show that a wide spectrum of robot manipulation tasks can be expressed with multimodal prompts, interleaving textual and visual tokens. Accordingly, we develop a new simulation benchmark that consists of thousands of procedurally-generated tabletop tasks with multimodal prompts, 600K+ expert trajectories for imitation learning, and a four-level evaluation protocol for systematic generalization. We design a transformer-based robot agent, VIMA, that processes these prompts and outputs motor actions autoregressively. VIMA features a recipe that achieves strong model scalability and data efficiency. It outperforms alternative designs in the hardest zero-shot generalization setting by up to $2.9\times$ task success rate given the same training data. With $10\times$ less training data, VIMA still performs $2.7\times$ better than the best competing variant. Code and video demos are available at https://vimalabs.github.io Yunfan Jiang 0001, Agrim Gupta, Zichen Zhang 0011, Guanzhi Wang, Yongqiang Dou, Li Fei-Fei 0001, Anima Anandkumar, Yuke Zhu, Linxi Fan |
ICML | 2 |
| 2023 | GreenMO: Enabling Virtualized, Sustainable Massive MIMO with a Single RF ChainabstractWith the turn of new decade, wireless communications face a major challenge on connecting many more new users and devices, at the same time being energy efficient and minimizing its carbon footprint. However, the current approaches to address the growing number of users and spectrum demands, like Massive MIMO, demand exorbitant energy consumption. The reason is that traditionally Massive MIMO requires a digital beamforming architecture that needs a separate RF chain per antenna, so the power consumption scales with number of antennas. Instead, GreenMO creates a new Massive MIMO architecture with just a single physically laid RF chain, shared by all the antennas and introduces for the first time, the concept of virtualizing the RF chain hardware. That is, GreenMO creates an optimal number of virtual RF chains to serve a given number of spatial streams, depending on channel conditions and network load. Due to efficient, softwarized control over the number of virtual RF chains, GreenMO paves the way for green and flexible massive MIMO. We prototype GreenMO on a PCB with eight antennas and evaluate it with a WARPv3 SDR platform in an office environment. The results demonstrate that GreenMO is 3× more power-efficient than traditional Massive MIMO and 4× more spectrum-efficient than traditional OFDMA systems, while multiplexing 4 spatial streams, and can save upto 50% power in modern 5G NR base stations. Agrim Gupta, Sajjad Nassirpour, Manideep Dunna, Eamon Patamasing, Alireza Vahid, Dinesh Bharadia |
MobiCom | 1 |
| 2023 | Siamese Masked AutoencodersabstractEstablishing correspondence between images or scenes is a significant challenge in computer vision, especially given occlusions, viewpoint changes, and varying object appearances. In this paper, we present Siamese Masked Autoencoders (SiamMAE), a simple extension of Masked Autoencoders (MAE) for learning visual correspondence from videos. SiamMAE operates on pairs of randomly sampled video frames and asymmetrically masks them. These frames are processed independently by an encoder network, and a decoder composed of a sequence of cross-attention layers is tasked with predicting the missing patches in the future frame. By masking a large fraction (95%) of patches in the future frame while leaving the past frame unchanged, SiamMAE encourages the network to focus on object motion and learn object-centric representations. Despite its conceptual simplicity, features learned via SiamMAE outperform state-of-the-art self-supervised methods on video object segmentation, pose keypoint propagation, and semantic part propagation tasks. SiamMAE achieves competitive results without relying on data augmentation, handcrafted tracking-based pretext tasks, or other techniques to prevent representational collapse. Agrim Gupta, Jiajun Wu 0001, Jia Deng 0001, Li Fei-Fei 0001 |
NeurIPS | 1 |
| 2023 | Holistic Evaluation of Text-to-Image ModelsabstractThe stunning qualitative improvement of text-to-image models has led to their widespread attention and adoption. However, we lack a comprehensive quantitative understanding of their capabilities and risks. To fill this gap, we introduce a new benchmark, Holistic Evaluation of Text-to-Image Models (HEIM). Whereas previous evaluations focus mostly on image-text alignment and image quality, we identify 12 aspects, including text-image alignment, image quality, aesthetics, originality, reasoning, knowledge, bias, toxicity, fairness, robustness, multilinguality, and efficiency. We curate 62 scenarios encompassing these aspects and evaluate 26 state-of-the-art text-to-image models on this benchmark. Our results reveal that no single model excels in all aspects, with different models demonstrating different strengths. We release the generated images and human evaluation results for full transparency at https://crfm.stanford.edu/heim/latest and the code at https://github.com/stanford-crfm/helm, which is integrated with the HELM codebase Michihiro Yasunaga, Chenlin Meng, Yifan Mai 0001, Joon Sung Park 0001, Agrim Gupta, Deepak Narayanan, Hannah Teufel, Marco Bellagente, Minguk Kang, Taesung Park, Jure Leskovec, Jun-Yan Zhu, Li Fei-Fei 0001, Jiajun Wu 0001, Stefano Ermon, Percy Liang |
NeurIPS | 6 |
| 2023 | Power-Efficient Analog Front-End Interference Suppression With Binary AntennasabstractDigital and analog beamforming are well-known methods to suppress interference using multiple-antenna structures, but they have practical limitations: (i) Digital beamforming requires multiple analog-to-digital converters (ADCs) to enable digital conversion, which increases the cost and complexity; (ii) Although analog beamforming does not require expensive ADCs, it uses phase shifters, which cause quantization errors, insertion losses, and reduced power efficiency. In this paper, we consider a$K$-user uplink interference channel and propose a low-complexity algorithmic interference-suppression solution relying on simple switch-based reconfigurable antennas at the receivers. We utilize switches to enable/disable antennas to maximize each user’s signal-to-interference-plus-noise ratio (SINR). We present an optimization approach to approximate the optimal solution. To evaluate the results, we compare our method with relevant benchmarks. Moreover, we derive a lower bound on the minimum number of antenna elements per receiver to attain the desired SINR and verify the findings via simulations. Sajjad Nassirpour, Agrim Gupta, Alireza Vahid, Dinesh Bharadia |
IEEE Trans. Wirel. Commun. | 2 |
| 2022 | MetaMorph: Learning Universal Controllers with Transformers
Agrim Gupta, Linxi Fan, Surya Ganguli, Li Fei-Fei 0001 |
ICLR | 1 |
| 2021 | WiForce: Wireless Sensing and Localization of Contact Forces on a Space Continuum
Agrim Gupta, Cédric Girerd, Manideep Dunna, Qiming Zhang 0003, Raghav Subbaraman, Tania K. Morimoto, Dinesh Bharadia |
NSDI | 1 |
| 2021 | Flag Manifold-Based Precoder Interpolation Techniques for MIMO-OFDM SystemsabstractThe use of channel state information (CSI) at the transmitter significantly enhances the performance of wireless communication systems. However, the requirement of CSI feedback places an undue burden on the reverse link, especially in links that employ multiple-input multiple-output (MIMO) and orthogonal frequency division multiplexing (OFDM), where CSI takes the form of a precoding matrix (precoder) for each subcarrier. Typical deployments use quantization and feedback of CSI at certain subcarriers, with interpolation to fill in missing CSI at the transmitter. Past work has used the orthogonal structure of precoders with Flag manifolds for quantization and interpolation of CSI, although interpolation is complicated due to the absence of analytic expressions for geodesics on Flag manifolds. Other approaches have involved the parameterization of the precoder into scalar parameters that are amenable to quantization and interpolation. In this paper, we present efficient methods to quantize and interpolate on Flag manifolds, using both optimal algorithms as well as simplified suboptimal algorithms. Further, we unify these with the parameterization based approaches and show that these translate directly to low-complexity quantization and interpolation on Flag manifolds. Simulations reveal that the proposed precoder quantization and interpolation effectively enhance achievable rates with limited complexity. Sarthak Nijhawan, Agrim Gupta, Kumar Appaiah, Rahul Vaze, Nikhil Karamchandani |
IEEE Trans. Commun. | 2 |
| 2020 | BluBLE, space-time social distancing to monitor the spread of COVID-19: poster abstractabstractSocial distancing has been the key factor which has helped control the COVID-19 pandemic spread. We present BluBLE, which utilizes Bluetooth Low Energy (BLE) based mobile sensing to help monitor these social distancing protocols. Specifically, we formulate the problem in two parts - spatial and temporal social distancing. The spatial distancing formulation aims to enforce the 6 feet distance recommended by various public health organization around the world. The temporal distancing formulation aims to inform and prevent users from entering high-occupancy regions (hotspots) in buildings. BluBLE achieved more than 80 % classification accuracy in both the tasks, that is, predicting if a user is within '6' feet of another user as well as characterizing the user's location within a particular hotspot. Aditya Arun 0002, Agrim Gupta, Shivani Bhatka, Saikiran Komatineni, Dinesh Bharadia |
SenSys | 2 |
| 2019 | LVIS: A Dataset for Large Vocabulary Instance SegmentationabstractProgress on object detection is enabled by datasets that focus the research community’s attention on open challenges. This process led us from simple images to complex scenes and from bounding boxes to segmentation masks. In this work, we introduce LVIS (pronounced ‘el-vis’): a new dataset for Large Vocabulary Instance Segmentation. We plan to collect 2.2 million high-quality instance segmentation masks for over 1000 entry-level object categories in 164k images. Due to the Zipfian distribution of categories in natural images, LVIS naturally has a long tail of categories with few training samples. Given that state-of-the-art deep learning methods for object detection perform poorly in the low-sample regime, we believe that our dataset poses an important and exciting new scientific challenge. LVIS is available at http://www.lvisdataset.org. Agrim Gupta, Piotr Dollár, Ross B. Girshick |
CVPR | 1 |
| 2019 | Predictive Quantization and Joint Time-Frequency Interpolation Technique for MIMO-OFDM PrecodingabstractPrecoding transmissions in wireless MIMO systems is essential to enable optimal utilization of the spatial degrees of freedom. However, communicating the precoding matrices from the receiver is challenging, owing to large feedback requirements. Past work has shown that predictive quantization in time, as well as interpolation over frequency can be used to reconstruct the precoders over a wide band, although these techniques have not been used jointly. We propose both a predictive quantization as well as a joint time-frequency interpolation strategy for precoding matrices over the Stiefel manifold. The key insight that we use is that local tangent spaces in the manifold permit effective combination of both temporal and frequency domain information for more accurate precoder reconstruction. Simulations reveal that we obtain a significant improvement in achievable rate as well as BER reduction when compared to existing strategies. Agrim Gupta, Kumar Appaiah, Rahul Vaze |
ICC | 1 |
| 2018 | Social GAN: Socially Acceptable Trajectories With Generative Adversarial NetworksabstractUnderstanding human motion behavior is critical for autonomous moving platforms (like self-driving cars and social robots) if they are to navigate human-centric environments. This is challenging because human motion is inherently multimodal: given a history of human motion paths, there are many socially plausible ways that people could move in the future. We tackle this problem by combining tools from sequence prediction and generative adversarial networks: a recurrent sequence-to-sequence model observes motion histories and predicts future behavior, using a novel pooling mechanism to aggregate information across people. We predict socially plausible futures by training adversarially against a recurrent discriminator, and encourage diverse predictions with a novel variety loss. Through experiments on several datasets we demonstrate that our approach outperforms prior work in terms of accuracy, variety, collision avoidance, and computational complexity. Agrim Gupta, Justin Johnson 0001, Li Fei-Fei 0001, Silvio Savarese, Alexandre Alahi |
CVPR | 1 |
| 2018 | Image Generation From Scene GraphsabstractTo truly understand the visual world our models should be able not only to recognize images but also generate them. To this end, there has been exciting recent progress on generating images from natural language descriptions. These methods give stunning results on limited domains such as descriptions of birds or flowers, but struggle to faithfully reproduce complex sentences with many objects and relationships. To overcome this limitation we propose a method for generating images from scene graphs, enabling explicitly reasoning about objects and their relationships. Our model uses graph convolution to process input graphs, computes a scene layout by predicting bounding boxes and segmentation masks for objects, and converts the layout to an image with a cascaded refinement network. The network is trained adversarially against a pair of discriminators to ensure realistic outputs. We validate our approach on Visual Genome and COCO-Stuff, where qualitative results, ablations, and user studies demonstrate our method's ability to generate complex images with multiple objects. Justin Johnson 0001, Agrim Gupta, Li Fei-Fei 0001 |
CVPR | 2 |
| 2017 | Characterizing and Improving Stability in Neural Style Transfer
Agrim Gupta, Justin Johnson 0001, Alexandre Alahi, Li Fei-Fei 0001 |
ICCV | 1 |