Sunav Choudhary

dblp:01/7607 · DBLP profile ↗
← Back
19ranked-venue papers
7as first author
8since 2021 · last 2026
0000-0002-7711-487XORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 4 first-author · 1 since 2021Databases, data management, data science and information retrieval · 4 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-authorSystems, architecture and hardware · 1 · 1 first-authorComputer networks · 1Security and privacy · 1 · 1 since 2021Theory of computation · 1 · 1 since 2021
YearPublicationVenuePosition
2026 TeamFusion: Supporting Open-ended Teamwork with Multi-Agent Systems
abstract
Jiale Liu, Victor Bursztyn, Lin Ai, Haoliang Wang, Sunav Choudhary, Saayan Mitra, Qingyun Wu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Victor S. Bursztyn, Lin Ai, Sunav Choudhary, Saayan Mitra, Qingyun Wu
ACL (1)5
2024 Fake or Compromised? Making Sense of Malicious Clients in Federated Learning
Hamid Mozaffari, Sunav Choudhary, Amir Houmansadr
ESORICS (1)2
2024 Thinking Forward: Memory-Efficient Federated Finetuning of Language Models
abstract
Finetuning large language models (LLMs) in federated learning (FL) settings has become increasingly important as it allows resource-constrained devices to finetune a model using private data. However, finetuning LLMs using backpropagation requires excessive memory (especially from intermediate activations) for resource-constrained devices. While Forward-mode Auto-Differentiation (AD) can significantly reduce memory footprint from activations, we observe that directly applying it to LLM finetuning results in slow convergence and poor accuracy. In this paper, we introduce Spry, an FL algorithm that splits trainable weights of an LLM among participating clients, such that each client computes gradients using forward-mode AD that are closer estimations of the true gradients. Spry achieves a low memory footprint, high accuracy, and fast convergence. We formally prove that the global gradients in Spry are unbiased estimators of true global gradients for homogeneous data distributions across clients, while heterogeneity increases bias of the estimates. We also derive Spry's convergence rate, showing that the gradients decrease inversely proportional to the number of FL rounds, indicating the convergence up to the limits of heterogeneity. Empirically, Spry reduces the memory footprint during training by 1.4-7.1$\times$ in contrast to backpropagation, while reaching comparable accuracy, across a wide range of language tasks, models, and FL settings. Spry reduces the convergence time by 1.2-20.3$\times$ and achieves 5.2-13.5\% higher accuracy against state-of-the-art zero-order methods. When finetuning Llama2-7B with LoRA, compared to the peak memory consumption of 33.9GB of backpropagation, Spry only consumes 6.2GB of peak memory. For OPT13B, the reduction is from 76.5GB to 10.8GB. Spry makes feasible previously impossible FL deployments on commodity mobile and edge devices. Our source code is available for replication at https://github.com/Astuary/Spry.
Kunjal Panchal, Nisarg Parikh, Sunav Choudhary, Lijun Zhang 0005, Yuriy Brun, Hui Guan 0001
NeurIPS3
2023 Delivery Optimized Discovery in Behavioral User Segmentation under Budget Constraint
abstract
Users' behavioral footprints online enable firms to discover behavior-based user segments (or, segments) and deliver segment specific messages to users. Following the discovery of segments, delivery of messages to users through preferred media channels like Facebook and Google can be challenging, as only a portion of users in a behavior segment find match in a medium, and only a fraction of those matched actually see the message (exposure). Even high quality discovery becomes futile when delivery fails. Many sophisticated algorithms exist for discovering behavioral segments; however, these ignore the delivery component. The problem is compounded because (i) the discovery is performed on the behavior data space in firms' data (e.g., user clicks), while the delivery is predicated on the static data space (e.g., geo, age) as defined by media; and (ii) firms work under budget constraint. We introduce a stochastic optimization based algorithm for delivery optimized discovery of behavioral user segments and offer new metrics to address the joint optimization. We leverage optimization under a budget constraint for delivery combined with a learning-based component for discovery. Extensive experiments on a public dataset from Google and a proprietary dataset show the effectiveness of our approach by simultaneously improving delivery metrics, reducing budget spend and achieving strong predictive performance in discovery.
Harshita Chopra, Atanu R. Sinha, Sunav Choudhary, Ryan Rossi, Paavan Kumar Indela, Veda Pranav Parwatala, Srinjayee Paul, Aurghya Maiti
CIKM3
2023 Flash: Concept Drift Adaptation in Federated Learning
abstract
In Federated Learning (FL), adaptive optimization is an effective approach to addressing the statistical heterogeneity issue but cannot adapt quickly to concept drifts. In this work, we propose a novel adaptive optimizer called Flash that simultaneously addresses both statistical heterogeneity and the concept drift issues. The fundamental insight is that a concept drift can be detected based on the magnitude of parameter updates that are required to fit the global model to each participating client's local data distribution. Flash uses a two-pronged approach that synergizes client-side early-stopping training to facilitate detection of concept drifts and the server-side drift-aware adaptive optimization to effectively adjust effective learning rate. We theoretically prove that Flash matches the convergence rate of state-of-the-art adaptive optimizers and further empirically evaluate the efficacy of Flash on a variety of FL benchmarks using different concept drift settings.
Kunjal Panchal, Sunav Choudhary, Subrata Mitra, Koyel Mukherjee 0001, Somdeb Sarkhel, Saayan Mitra, Hui Guan 0001
ICML2
2023 Flow: Per-instance Personalized Federated Learning
abstract
Federated learning (FL) suffers from data heterogeneity, where the diverse data distributions across clients make it challenging to train a single global model effectively. Existing personalization approaches aim to address the data heterogeneity issue by creating a personalized model for each client from the global model that fits their local data distribution. However, these personalized models may achieve lower accuracy than the global model in some clients, resulting in limited performance improvement compared to that without personalization. To overcome this limitation, we propose a per-instance personalization FL algorithm Flow. Flow creates dynamic personalized models that are adaptive not only to each client’s data distributions but also to each client’s data instances. The personalized model allows each instance to dynamically determine whether it prefers the local parameters or its global counterpart to make correct predictions, thereby improving clients’ accuracy. We provide theoretical analysis on the convergence of Flow and empirically demonstrate the superiority of Flow in improving clients’ accuracy compared to state-of-the-art personalization approaches on both vision and language-based tasks.
Kunjal Panchal, Sunav Choudhary, Nisarg Parikh, Lijun Zhang 0005, Hui Guan 0001
NeurIPS2
2022 EI-CLIP: Entity-aware Interventional Contrastive Learning for E-commerce Cross-modal Retrieval
abstract
Cross language-image modality retrieval in E-commerce is a fundamental problem for product search, recommendation, and marketing services. Extensive efforts have been made to conquer the cross-modal retrieval problem in the general domain. When it comes to E-commerce, a com-mon practice is to adopt the pretrained model and finetune on E-commerce data. Despite its simplicity, the performance is sub-optimal due to overlooking the uniqueness of E-commerce multimodal data. A few recent efforts [10], [72] have shown significant improvements over generic methods with customized designs for handling product images. Unfortunately, to the best of our knowledge, no existing method has addressed the unique challenges in the e-commerce language. This work studies the outstanding one, where it has a large collection of special meaning entities, e.g., “Di s s e l (brand)”, “Top (category)”, “relaxed (fit)” in the fashion clothing business. By formulating such out-of-distribution finetuning process in the Causal Inference paradigm, we view the erroneous semantics of these special entities as confounders to cause the retrieval failure. To rectify these semantics for aligning with e-commerce do-main knowledge, we propose an intervention-based entity-aware contrastive learning framework with two modules, i.e., the Confounding Entity Selection Module and Entity-Aware Learning Module. Our method achieves competitive performance on the E-commerce benchmark Fashion-Gen. Particularly, in top-1 accuracy (R@l), we observe 10.3% and 10.5% relative improvements over the closest baseline in image-to-text and text-to-image retrievals, respectively.
Handong Zhao, Zhe Lin 0001, Ajinkya Kale, Zhangyang Wang, Tong Yu 0001, Jiuxiang Gu, Sunav Choudhary, Xiaohui Xie
CVPR8
2022 Correlated Stochastic Knapsack with a Submodular Objective
abstract
We study the correlated stochastic knapsack problem of a submodular target function, with optional additional constraints. We utilize the multilinear extension of submodular function, and bundle it with an adaptation of the relaxed linear constraints from Ma [Mathematics of Operations Research, Volume 43(3), 2018] on correlated stochastic knapsack problem. The relaxation is then solved by the stochastic continuous greedy algorithm, and rounded by a novel method to fit the contention resolution scheme (Feldman et al. [FOCS 2011]). We obtain a pseudo-polynomial time $(1 - 1/\sqrt{e})/2 \simeq 0.1967$ approximation algorithm with or without those additional constraints, eliminating the need of a key assumption and improving on the $(1 - 1/\sqrt[4]{e})/2 \simeq 0.1106$ approximation by Fukunaga et al. [AAAI 2019].
Sheng Yang 0005, Samir Khuller, Sunav Choudhary, Subrata Mitra, Kanak Mahadik
ESA3
2019 Mentor Pattern Identification from Product Usage Logs
Ankur Garg, Aman Kharb, Yash H. Malviya, J. P. Sagar, Atanu R. Sinha, Iftikhar Ahamath Burhanuddin, Sunav Choudhary
PAKDD (3)7
2018 Sparse Decomposition for Time Series Forecasting and Anomaly Detection
abstract
Anomaly detection and forecasting are two fundamental problems in time series analysis that are relevant to a wide range of academic and industrial disciplines. Although these problems have been investigated in the literature previously, the assumptions therein are too restrictive for autonomous analysis. Common examples of limiting assumptions include perfect knowledge about the time series seasonality and/or presence of anomaly (spikes and level changes) free time windows. Current practice is to manually input this knowledge into anomaly detection and forecasting systems which negate any possibility of autonomous analysis. This paper relaxes these assumptions by jointly estimating the latent components (viz. seasonality, level changes, and spikes) in the observed time series without assuming the availability of anomaly-free time windows. The novel and flexible two stage approach proposed herein is based on (a) sparse modeling of the different latent components of the time series and (b) ARMA modeling for fitting the error. The approach leads to a solution for anomaly detection with control over type-I errors. Further, by design, the method is robust against anomalies in the observation window when it is used to solve the forecasting problem by extrapolation. Experiments are conducted with both synthetic and real datasets to demonstrate the efficacy of the proposed method. We compare our approach to various popular baselines. The presented approach outperforms baseline algorithms for anomaly detection in all our experiments and performs favorably for the forecasting task.
Sunav Choudhary, Gaurush Hiranandani, Shiv Kumar Saini
SDM1
2017 Smart Geo-fencing with Location Sensitive Product Affinity
abstract
Geo-fencing is a location based service that allows sending of messages to users who enter/exit a specified geographical area, known as a geo-fence. Today, it has become one of the popular location based mobile marketing strategies. However, the process of designing geo-fences is presently manual, i.e. a retailer must specify the location and the radius of area around it to setup the geo-fences. Moreover, this process does not consider the user's preference towards the targeted product/service and thus, can compromise his/her experience of the app that sends these communications. We attempt to solve this problem by presenting a novel end-to-end system for automated design of affinity based smart geo-fences. Affinity towards a product/service refers to the user's interest in a product/service. Our unique formulation to estimate affinity, using historical app usage data, is sensitive to a user's location and thus, the affinity is termed as location sensitive product affinity (LSPA). The geo-fence logic tries to capture contiguous groups of locations where the affinity high. Experiments on real world e-commerce dataset reveals that geo-fences designed by our approach performs significantly better at accurately targeting the users who are interested in a product. We thus show that, using historical app usage data, geo-fences can be designed in an automated manner and can help enterprises target interested users with better accuracy as compared to the present industry practices.
Ankur Garg, Sunav Choudhary, Payal Bajaj, Sweta Agrawal, Abhishek Kedia
SIGSPATIAL/GIS2
2016 Delay-Doppler estimation via structured low-rank matrix recovery
abstract
The estimation of a narrowband time-varying channel under finite block length and transmission bandwidth is investigated. A novel method is proposed for estimation in the delay-Doppler domain by exploiting structural constraints on low-rank matrix recovery. The proposed algorithm uses Gauss-Seidel iterations on the low-rank parameterization under noisy training signal measurements. Theoretical global identifiability results for the channel leakage (due to finite block length and transmission bandwidth) are stated and the necessity of considering Doppler shift induced structure is demonstrated. Justification is provided for the choice of simulation parameters and initialization strategies to achieve good convergence rates and some ill-posed scenarios are also described. It is further shown that simple sparsity-based algorithms like basis pursuit/nuclear norm minimization do not perform well on the said constraint set for measurement operators arising out of training sequences.
Sunav Choudhary, Sajjad Beygi, Urbashi Mitra
ICASSP1
2016 On target localization with communication costs via tensor completion: A multi-modal approach
abstract
The problem of active target detection using low rank methods is explored. In prior work, a strategy was proposed based on matrix completion for randomly sampling a field combined with binary search to localize a target. Herein, two innovations are explored: the consideration of tensor-completion in order to exploit multi-modal data and the examination of the costs associated with communication. In particular, the random samples are collected in neighborhoods wherein the quality of the observation is a function of the distance of the sampling point to the centroid of the neighborhood. Due to the tradeoff between communication quality and sampling quality, there is an optimal neighborhood size.
Sagar Honnungar, Sunav Choudhary, Urbashi Mitra
ICASSP2
2015 Analysis of target detection via matrix completion
abstract
A problem of broad interest is the detection and localization of a target or object from its generated field. In this paper, a detection and localization strategy which exploits the structure of target fields is designed and analyzed. In taking advantage of this structure, one is able to reduce sample complexity requirements while maintaining good performance. In particular, an exploration-exploitation approach to target detection is proposed utilizing the theory of low-rank matrix completion for a decaying separable target field. The assumptions on the field are fairly generic and are applicable to many decay profiles. Our approach does not require specific knowledge of the field, only that it admits a rank-one representation. A performance analysis for localization is presented that characterizes a trade-off with sample complexity in the presence of noise.
Sunav Choudhary, Urbashi Mitra
ICASSP1
2014 Active target detection with mobile agents
abstract
A strategy for active target detection suitable for the use of mobile agents in a field is presented. In particular, there is an interest in autonomous underwater vehicles. By exploiting notions from group testing, the proposed algorithm decides when to collect new samples depending on whether the mobile agent perceives the sensor measurements correspond to noise or a target pattern. Under suitable assumptions about the field emanated by the target, i.e. the target signature is locally low rank in the field, one can efficiently sample the field to locate the target using O(m log m log n) samples on an n × n grid where m ≪ n is a parameter specifying the group size.
Sunav Choudhary, Naveen Kumar 0004, Srikanth Narayanan, Urbashi Mitra
ICASSP1
2014 Sparse blind deconvolution: What cannot be done
abstract
Identifiability is a key concern in ill-posed blind deconvolution problems arising in wireless communications and image processing. The single channel version of the problem is the most challenging and there have been efforts to use sparse models for regularizing the problem. Identifiability of the sparse blind deconvolution problem is analyzed and it is established that a simple sparsity assumption in the canonical basis is insufficient for unique recovery; a surprising negative result. The proof technique involves lifting the deconvolution problem into a rank one matrix recovery problem and analyzing the rank two null-space of the resultant linear operator. A DoF (degrees of freedom) wise tight parametrized subset of this rank two null-space is constructed to establish the results.
Sunav Choudhary, Urbashi Mitra
ISIT1
2013 On identifiability in bilinear inverse problems
abstract
This paper considers identifiability and recoverability in bilinear inverse problems which is relevant to blind deconvolution and matrix factorization. It is shown that bilinear inverse problems can be posed as rank-1 matrix recovery problems subject to linear constraints. Sufficient conditions for identifiability are developed for the cases when rank-2 matrices are present in the null space of the linear operator. Signal recovery using the nuclear norm heuristic for rank-1 matrix recovery is considered and simple conditions for success are provided.
Sunav Choudhary, Urbashi Mitra
ICASSP1
2012 Underwater Data Collection Using Robotic Sensor Networks
abstract
We examine the problem of utilizing an autonomous underwater vehicle (AUV) to collect data from an underwater sensor network. The sensors in the network are equipped with acoustic modems that provide noisy, range-limited communication. The AUV must plan a path that maximizes the information collected while minimizing travel time or fuel expenditure. We propose AUV path planning methods that extend algorithms for variants of the Traveling Salesperson Problem (TSP). While executing a path, the AUV can improve performance by communicating with multiple nodes in the network at once. Such multi-node communication requires a scheduling protocol that is robust to channel variations and interference. To this end, we examine two multiple access protocols for the underwater data collection scenario, one based on deterministic access and another based on random access. We compare the proposed algorithms to baseline strategies through simulated experiments that utilize models derived from experimental test data. Our results demonstrate that properly designed communication models and scheduling protocols are essential for choosing the appropriate path planning algorithms for data collection.
Geoffrey A. Hollinger, Sunav Choudhary, Parastoo Qarabaqi, Chris Murphy, Urbashi Mitra, Gaurav S. Sukhatme, Milica Stojanovic, Hanumant Singh, Franz S. Hover
IEEE J. Sel. Areas Commun.2
2010 A SPT treatment to the bit serial realization of the sign-LMS based adaptive filter
abstract
This paper presents a bit serial realization of the sign-LMS based adaptive filter which enjoys multiplier free weight update loop. To reduce the complexity of the multipliers that arise in the filtering process, the filter weights are represented in the so-called canonic SPT form which guarantees presence of at least one zero between every two non-zero power-of-two terms. As the filter weights are not fixed but updated in time, it is essential to ensure that the canonic SPT format is retained in the updated filter coefficients. For this, a bit serial adder is proposed that takes as input two numbers in canonic SPT and produces an output also in canonic SPT. It is further shown how the canonic SPT property of the input can be used to reduce the complexity of the adder. For the filtering part, a bit serial multiplier is developed that takes one input (i.e., data bits) in 2's complement form and the other input (i.e., weight bits) in canonic SPT, producing the result in 2's complement. The multiplication can not, however, be realized using a few fixed shift and add operations, since the position of the non-zero SPT terms in the canonic SPT expression of each coefficient changes with time. The proposed multiplier instead multiplies the 2's complement number with pairs of consecutive SPT bits of the other number. The resulting partial products can be realized using simple AND-OR logic.
Sunav Choudhary, Pritam Mukherjee, Mrityunjoy Chakraborty
ISCAS1