Supriyo Chakraborty

dblp:14/7588 · DBLP profile ↗
← Back
42ranked-venue papers
8as first author
12since 2021 · last 2026
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 15 · 1 first-author · 9 since 2021Databases, data management, data science and information retrieval · 12 · 3 first-author · 3 since 2021Computer networks · 6 · 1 first-authorSecurity and privacy · 6 · 1 first-author · 2 since 2021Human-computer interaction and ubiquitous computing · 5 · 1 first-authorSystems, architecture and hardware · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Theory of computation · 2 · 1 first-authorGraphics, computer vision, multimedia, augmented reality and games · 1
YearPublicationVenuePosition
2026 Routing with Generated Data: Annotation-Free LLM Skill Estimation and Expert Selection
abstract
Tianyi Niu, Justin Chen, Genta Indra Winata, Shi-Xiong Zhang, Supriyo Chakraborty, Sambit Sahu, Yue Zhang, Elias Stengel-Eskin, Mohit Bansal. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Tianyi Niu, Justin Chih-Yao Chen, Genta Indra Winata, Supriyo Chakraborty, Sambit Sahu, Yue Zhang 0004, Elias Stengel-Eskin, Mohit Bansal
ACL (1)5
2025 Training Dynamics Underlying Language Model Scaling Laws: Loss Deceleration and Zero-Sum Learning
abstract
Andrei Mircea, Supriyo Chakraborty, Nima Chitsazan, Irina Rish, Ekaterina Lobacheva. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Andrei Mircea, Supriyo Chakraborty, Nima Chitsazan, Irina Rish, Ekaterina Lobacheva
ACL (1)2
2025 SPEAR-MM: Selective Parameter Evaluation and Restoration via Model Merging for Efficient Financial LLM Adaptation
Berkcan Kapusuzoglu, Supriyo Chakraborty, Renkun Ni, Stephen Rawls, Sambit Sahu
IEEE Big Data2
2025 SaViD: Spectravista Aesthetic Vision Integration for Robust and Discerning 3D Object Detection in Challenging Environments
abstract
The fusion of LiDAR and camera sensors has demonstrated significant effectiveness in achieving accurate detection for short-range tasks in autonomous driving. However, this fusion approach could face challenges when dealing with long-range detection scenarios due to disparity between sparsity of LiDAR and high-resolution camera data. Moreover, sensor corruption introduces complexities that affect the ability to maintain robustness, despite the growing adoption of sensor fusion in this domain. We present SaViD, a novel framework comprised of a three-stage fusion alignment mechanism designed to address long-range detection challenges in the presence of natural corruption. The SaViD framework consists of three key elements: the Global Memory Attention Network (GMAN), which enhances the extraction of image features through offering a deeper understanding of global patterns; the Attentional Sparse Memory Network (ASMN), which enhances the inte-gration of LiDAR and image features; and the KNNnectivity Graph Fusion (KGF), which enables the entire fusion of spatial information. SaViD achieves superior performance on the long-range detection Argoverse-2 (AV2) dataset with a performance improvement of 9.87% in AP value and an improvement of 2.39% in mAPH for L2 difficulties on the Waymo Open dataset (WOD). Comprehensive experiments are carried out to showcase its robustness against 14 natural sensor corruptions. SaViD exhibits a robust performance improvement of 31.43% for AV2 and 16.13% for WOD in RCE value compared to other existing fusion-based methods while considering all the corruptions for both datasets. Our code is available at SaVil).
Tanmoy Dam, Sanjay Bhargav Dharavath, Sameer Alam, Nimrod Lilith, Aniruddha Maiti, Supriyo Chakraborty, Mir Feroskhan
ICRA6
2025 Dense Backpropagation Improves Training for Sparse Mixture-of-Experts
abstract
Mixture of Experts (MoE) pretraining is more scalable than dense Transformer pretraining, because MoEs learn to route inputs to a sparse set of their feedforward parameters. However, this means that MoEs only receive a sparse backward update, leading to training instability and suboptimal performance. We present a lightweight approximation method that gives the MoE router a dense gradient update while continuing to sparsely activate its parameters. Our method, which we refer to as Default MoE, substitutes missing expert activations with default outputs consisting of an exponential moving average of expert outputs previously seen over the course of training. This allows the router to receive signals from every expert for each token, leading to significant improvements in training performance. Our Default MoE outperforms standard TopK routing in a variety of settings without requiring significant computational overhead.
Ashwinee Panda, Vatsal Baherwani, Zain Sarwar, Benjamin Thérien, Sambit Sahu, Tom Goldstein, Supriyo Chakraborty
NeurIPS7
2024 Quantum Inverse Contextual Vision Transformers (Q-ICVT): A New Frontier in 3D Object Detection for AVs
abstract
The field of autonomous vehicles (AVs) predominantly leverages multi-modal integration of LiDAR and camera data to achieve better performance compared to using a single modality. However, the fusion process encounters challenges in detecting distant objects due to the disparity between the high resolution of cameras and the sparse data from LiDAR. Insufficient integration of global perspectives with local-level details results in sub-optimal fusion performance.To address this issue, we have developed an innovative two-stage fusion process called Quantum Inverse Contextual Vision Transformers (Q-ICVT). This approach leverages adiabatic computing in quantum concepts to create a novel reversible vision transformer known as the Global Adiabatic Transformer (GAT). GAT aggregates sparse LiDAR features with semantic features in dense images for cross-modal integration in a global form. Additionally, the Sparse Expert of Local Fusion (SELF) module maps the sparse LiDAR 3D proposals and encodes position information of the raw point cloud onto the dense camera feature space using a gating point fusion approach. Our experiments show that Q-ICVT achieves an mAPH of 82.54 for L2 difficulties on the Waymo dataset, improving by 1.88% over current state-of-the-art fusion methods. We also analyze GAT and SELF in ablation studies to highlight the impact of Q-ICVT. Our code is available at https://github.com/sanjay-810/Qicvt
Sanjay Bhargav Dharavath, Tanmoy Dam, Supriyo Chakraborty, Prithwiraj Roy, Aniruddha Maiti
CIKM3
2024 DNS Exfiltration Guided by Generative Adversarial Networks
abstract
Today, DNS exfiltration attacks are detected by checking for anomalies present in the traffic, such as unusu-ally high transmission rates to a single domain and/or DNS query patterns that are very different from those in benign queries. While such approaches are seemingly robust, we show in this paper that our carefully designed and novel DNS exfiltration attack, Dolos, that uses a generative adversarial network (GAN), can guide the encoding of sensitive data in a manner that both evades these detectors and significantly speeds up the exfiltration rate compared to prior methods. At its core, Dolos divides the exfiltration data into smaller chunks, and projects each chunk into a representation that is very similar to benign queries. In addition, Dolosadaptively tunes its exfiltration rate to conform with benign DNS traffic from the compromised host, and introduces proper levels of spurious traffic to reduce entropy. Importantly, Dolos evades machine learning (ML) based detectors with no prior knowledge of their architectures or training sets (i.e., it is a blackbox exfiltration). We perform extensive evaluations using multiple datasets and also have a real im-plementation of DOLOS. Our evaluations show that DOLOS has a 12% detection probability even if 6 out of the 9 state-of-the-art defenses that we consider, are jointly used to detect exfiltration; if any of today's baseline exfiltration techniques try to achieve the same rate as Dolos in this setting, they are almost surely detected. If we reduce the rates of the baselines to achieve even a low albeit slightly higher detection probability than Dolos (0.15), we see that they take 25 x longer to achieve the exfiltration. With the other three defenses, we find that baselines are almost surely detected while Dolos remains relatively unaffected regardless of the rate of exfiltration.
Abdulrahman Fahim, Shitong Zhu, Zhiyun Qian, Chengyu Song, Evangelos E. Papalexakis, Supriyo Chakraborty, Kevin S. Chan, Paul L. Yu, Trent Jaeger, Srikanth V. Krishnamurthy
EuroS&P6
2024 AYDIV: Adaptable Yielding 3D Object Detection via Integrated Contextual Vision Transformer
abstract
Combining LiDAR and camera data has shown potential in enhancing short-distance object detection in autonomous driving systems. Yet, the fusion encounters difficulties with extended distance detection due to the contrast between LiDAR’s sparse data and the dense resolution of cameras. Besides, discrepancies in the two data representations further complicate fusion methods. We introduce AYDIV, a novel framework integrating a tri-phase alignment process specifically designed to enhance long-distance detection even amidst data discrepancies. AYDIV consists of the Global Contextual Fusion Alignment Transformer (GCFAT), which improves the extraction of camera features and provides a deeper understanding of large-scale patterns; the Sparse Fused Feature Attention (SFFA), which fine-tunes the fusion of LiDAR and camera details; and the Volumetric Grid Attention (VGA) for a comprehensive spatial data fusion. AYDIV’s performance on the Waymo Open Dataset (WOD) with an improvement of 1.24% in mAPH value(L2 difficulty) and the Argoverse2 Dataset with a performance improvement of 7.40% in AP value demonstrates its efficacy in comparison to other existing fusion-based methods. Our code is publicly available at https://github.com/sanjay-810/AYDIV2
Tanmoy Dam, Sanjay Bhargav Dharavath, Sameer Alam, Nimrod Lilith, Supriyo Chakraborty, Mir Feroskhan
ICRA5
2022 SparseFed: Mitigating Model Poisoning Attacks in Federated Learning with Sparsification
abstract
Federated learning is inherently vulnerable to model poisoning attacks because its decentralized nature allows attackers to participate with compromised devices. In model poisoning attacks, the attacker reduces the model’s performance on targeted sub-tasks (e.g. classifying planes as birds) by uploading "poisoned" updates. In this paper we introduce SparseFed, a novel defense that uses global top-k update sparsification and device-level gradient clipping to mitigate model poisoning attacks. We propose a theoretical framework for analyzing the robustness of defenses against poisoning attacks, and provide robustness and convergence analysis of our algorithm. To validate its empirical efficacy we conduct an open-source evaluation at scale across multiple benchmark datasets for computer vision and federated learning.
Ashwinee Panda, Saeed Mahloujifar, Arjun Nitin Bhagoji, Supriyo Chakraborty, Prateek Mittal
AISTATS4
2022 Identification of Thin Layer Via Source Wave-Field Dictionary Learning
abstract
Seismic data is an aggregation of traces that shows a spatial and temporal sampling of reflected wavefields—the identification of reservoir localization based on thin layer selection from the earth layers. The challenge to finding thin layers is that the information acquired gets distorted limiting reservoir localization due to subsurface complexity and random environmental noise. Dictionary elements from structure functions represent reflectivity patterns used in inversion seismic data. The seismic traces is a superposition of constituting reflectivity patterns. Traditionally noise separation is done by utilizing different properties in the various transform domains. The results lack adaptability and consideration of physical layer properties. In this paper, the dictionary is carefully chosen based on spatial source wavefield considers horizontal and vertical wavefield components having translational and rotational components. The horizontal source wavefield components coupled with wedges formed from reflectivity pair then convolved with seismic wavelet. The generated output is reflectivity inversion is based on the source wavefield dictionary. The proposed method is robust to noise as the dictionary used in inversion is built with the same spatial source wavefield; hence it could significantly improve the resolution of seismic data.
Supriyo Chakraborty, Aurobinda Routray, Ritesh Chandra Tewari
IECON1
2021 On Exploring Attention-based Explanation for Transformer Models in Text Classification
abstract
The Transformer models have achieved unprecedented breakthroughs in text classification, and have become the foundation of most state-of-the-art NLP systems. The core function that drives the success is the attention mechanism, which provides the ability to dynamically focus on different parts of the input sequence when producing the predictions. Several previous works have investigated the usage of attention weights to explain the model predictions, because intuitively, attention weights reflect the importance of the input positions in the output. Specifically, the objective for explanation is to compute a relevance score for each input token, such that the key input words that are most important to the prediction can be identified. However, previous efforts produced mixed results. We find that the key reason why attention weights cannot be directly used as effective relevance indications is because they do not contain the directional information for relevance (i.e., whether the input tokens contribute towards or against the prediction). We then propose two novel explanation techniques, namely AGrad and RePAGrad, that produce directional relevance scores based on attention weights. To evaluate the explanation performance, we propose three properties that an effective explanation method should satisfy (i.e., faithfulness, resilience, and consistency), and design the corresponding test to quantify each property. Through extensive evaluations with Transformer models and pre-trained BERT models on multiple public text classification datasets, we show that AGrad and RePAGrad significantly outperform existing state-of-the-art explanation methods in faithfulness and consistency, at the cost of nominal degradation on resilience compared to attention weights. In addition, we reveal that elements of a model architecture can play an important role towards explainability.
Shengzhong Liu, Franck Le, Supriyo Chakraborty, Tarek F. Abdelzaher
IEEE BigData3
2021 WristPrint: Characterizing User Re-identification Risks from Wrist-worn Accelerometry Data
abstract
Public release of wrist-worn motion sensor data is growing. They enable and accelerate research in developing new algorithms to passively track daily activities, resulting in improved health and wellness utilities of smartwatches and activity trackers. But, when combined with sensitive attribute inference attack and linkage attack via re-identification of the same user in multiple datasets, undisclosed sensitive attributes can be revealed to unintended organizations with potentially adverse consequences for unsuspecting data contributing users. To guide both users and data collecting researchers, we characterize the re-identification risks inherent in motion sensor data collected from wrist-worn devices in users' natural environment. For this purpose, we use an open-set formulation, train a deep learning architecture with a new loss function, and apply our model to a new data set consisting of 10 weeks of daily sensor wearing by 353 users. We find that re-identification risk increases with an increase in the activity intensity. On average, such risk is 96% for a user when sharing a full day of sensor data.
Nazir Saleheen, Md. Azim Ullah, Supriyo Chakraborty, Deniz S. Ones, Mani Srivastava 0001, Santosh Kumar 0001
CCS3
2020 Sanity Checks for Saliency Metrics
abstract
Saliency maps are a popular approach to creating post-hoc explanations of image classifier outputs. These methods produce estimates of the relevance of each pixel to the classification output score, which can be displayed as a saliency map that highlights important pixels. Despite a proliferation of such methods, little effort has been made to quantify how good these saliency maps are at capturing the true relevance of the pixels to the classifier output (i.e. their “fidelity”). We therefore investigate existing metrics for evaluating the fidelity of saliency methods (i.e. saliency metrics). We find that there is little consistency in the literature in how such metrics are calculated, and show that such inconsistencies can have a significant effect on the measured fidelity. Further, we apply measures of reliability developed in the psychometric testing literature to assess the consistency of saliency metrics when applied to individual saliency maps. Our results show that saliency metrics can be statistically unreliable and inconsistent, indicating that comparative rankings between saliency methods generated using such metrics can be untrustworthy.
Richard Tomsett, Dan Harborne, Supriyo Chakraborty, Prudhvi Gurram, Alun D. Preece
AAAI3
2019 Fair Transfer Learning with Missing Protected Attributes
abstract
Risk assessment is a growing use for machine learning models. When used in high-stakes applications, especially ones regulated by anti-discrimination laws or governed by societal norms for fairness, it is important to ensure that learned models do not propagate and scale any biases that may exist in training data. In this paper, we add on an additional challenge beyond fairness: unsupervised domain adaptation to covariate shift between a source and target distribution. Motivated by the real-world problem of risk assessment in new markets for health insurance in the United States and mobile money-based loans in East Africa, we provide a precise formulation of the machine learning with covariate shift and score parity problem. Our formulation focuses on situations in which protected attributes are not available in either the source or target domain. We propose two new weighting methods: prevalence-constrained covariate shift (PCCS) which does not require protected attributes in the target domain and target-fair covariate shift (TFCS) which does not require protected attributes in the source domain. We empirically demonstrate their efficacy in two applications.
Amanda Coston, Karthikeyan Natesan Ramamurthy, Dennis Wei, Kush R. Varshney, Skyler Speakman, Zairah Mustahsan, Supriyo Chakraborty
AIES7
2019 GenAttack: practical black-box attacks with gradient-free optimization
abstract
Deep neural networks are vulnerable to adversarial examples, even in the black-box setting, where the attacker is restricted solely to query access. Existing black-box approaches to generating adversarial examples typically require a significant number of queries, either for training a substitute network or performing gradient estimation. We introduce GenAttack, a gradient-free optimization technique that uses genetic algorithms for synthesizing adversarial examples in the black-box setting. Our experiments on different datasets (MNIST, CIFAR-10, and ImageNet) show that GenAttack can successfully generate visually imperceptible adversarial examples against state-of-the-art image recognition models with orders of magnitude fewer queries than previous approaches. Against MNIST and CIFAR-10 models, GenAttack required roughly 2,126 and 2,568 times fewer queries respectively, than ZOO, the prior state-of-the-art black-box attack. In order to scale up the attack to large-scale high-dimensional ImageNet models, we perform a series of optimizations that further improve the query efficiency of our attack leading to 237 times fewer queries against the Inception-v3 model than ZOO. Furthermore, we show that GenAttack can successfully attack some state-of-the-art ImageNet defenses, including ensemble adversarial training and non-differentiable or randomized input transformations. Our results suggest that evolutionary algorithms open up a promising area of research into effective black-box attacks.
Moustafa Farid Alzantot, Yash Sharma 0001, Supriyo Chakraborty, Huan Zhang 0001, Cho-Jui Hsieh, Mani Srivastava 0001
GECCO3
2019 Analyzing Federated Learning through an Adversarial Lens
abstract
Federated learning distributes model training among a multitude of agents, who, guided by privacy concerns, perform training using their local data but share only model parameter updates, for iterative aggregation at the server to train an overall global model. In this work, we explore how the federated learning setting gives rise to a new threat, namely model poisoning, which differs from traditional data poisoning. Model poisoning is carried out by an adversary controlling a small number of malicious agents (usually 1) with the aim of causing the global model to misclassify a set of chosen inputs with high confidence. We explore a number of strategies to carry out this attack on deep neural networks, starting with targeted model poisoning using a simple boosting of the malicious agent’s update to overcome the effects of other agents. We also propose two critical notions of stealth to detect malicious updates. We bypass these by including them in the adversarial objective to carry out stealthy model poisoning. We improve its stealth with the use of an alternating minimization strategy which alternately optimizes for stealth and the adversarial objective. We also empirically demonstrate that Byzantine-resilient aggregation strategies are not robust to our attacks. Our results indicate that highly constrained adversaries can carry out model poisoning attacks while maintaining stealth, thus highlighting the vulnerability of the federated learning setting and the need to develop effective defense strategies.
Arjun Nitin Bhagoji, Supriyo Chakraborty, Prateek Mittal, Seraphin B. Calo
ICML2
2018 Learning Light-Weight Edge-Deployable Privacy Models
abstract
Privacy becomes one of the important issues in data-driven applications. The advent of non-PC devices such as Internet-of-Things (IoT) devices for data-driven applications leads to needs for light-weight data anonymization. In this paper, we develop an anonymization framework that expedites model learning in parallel and generates deployable models for devices with low computing capability. We evaluate our framework with various settings such as different data schema and characteristics. Our results exhibit that our framework learns anonymization models up to 16 times faster than a sequential anonymization approach and that it preserves enough information in anonymized data for data-driven applications.
Yeon-Sup Lim, Mudhakar Srivatsa, Supriyo Chakraborty, Ian J. Taylor
IEEE BigData3
2018 Learning and Reasoning in Complex Coalition Information Environments: A Critical Analysis
abstract
In this paper we provide a critical analysis with metrics that will inform guidelines for designing distributed systems for Collective Situational Understanding (CSU). CSU requires both collective insight-i.e., accurate and deep understanding of a situation derived from uncertain and often sparse data and collective foresight-i.e., the ability to predict what will happen in the future. When it comes to complex scenarios, the need for a distributed CSU naturally emerges, as a single monolithic approach not only is unfeasible: it is also undesirable. We therefore propose a principled, critical analysis of AI techniques that can support specific tasks for CSU to derive guidelines for designing distributed systems for CSU.
Federico Cerutti 0001, Moustafa Farid Alzantot, Tianwei Xing, Dan Harborne, Jonathan Z. Bakdash, Dave Braines, Supriyo Chakraborty, Lance M. Kaplan, Angelika Kimmig, Alun D. Preece, Ramya Raghavendra, Murat Sensoy, Mani Srivastava 0001
FUSION7
2018 Why the Failure? How Adversarial Examples Can Provide Insights for Interpretable Machine Learning
abstract
Recent advances in Machine Learning (ML) have profoundly changed many detection, classification, recognition and inference tasks. Given the complexity of the battlespace, ML has the potential to revolutionise how Coalition Situation Understanding is synthesised and revised. However, many issues must be overcome before its widespread adoption. In this paper we consider two - interpretability and adversarial attacks. Interpretability is needed because military decision-makers must be able to justify their decisions. Adversarial attacks arise because many ML algorithms are very sensitive to certain kinds of input perturbations. In this paper, we argue that these two issues are conceptually linked, and insights in one can provide insights in the other. We illustrate these ideas with relevant examples from the literature and our own experiments.
Richard Tomsett, Amy Widdicombe, Tianwei Xing, Supriyo Chakraborty, Simon J. Julier, Prudhvi Gurram, Raghuveer M. Rao, Mani Srivastava 0001
FUSION4
2018 Self-Generation of Access Control Policies
abstract
Access control for information has primarily focused on access statically granted to subjects by administrators usually in the context of a specific system. Even if mechanisms are available for access revocation, revocations must still be executed manually by an administrator. However, as physical devices become increasingly embedded and interconnected, access control needs to become an integral part of the resource being protected and be generated dynamically by resources depending on the context in which the resource is being used. In this paper, we discuss a set of scenarios for access control needed in current and future systems and use that to argue that an approach for resources to generate and manage their access control policies dynamically on their own is needed. We discuss some approaches for generating such access control policies that may address the requirements of the scenarios.
Seraphin B. Calo, Dinesh C. Verma, Supriyo Chakraborty, Elisa Bertino, Emil C. Lupu, Gregory H. Cirincione
SACMAT3
2017 LightSpy: Optical eavesdropping on displays using light sensors on mobile devices
abstract
Light emanations from flat-panel displays are a side channel hinting towards the displayed content. Optical eavesdropping requires sensors in the proximity of such displays, necessitating physical access to the the target's environment. This requirement may be eliminated by exploiting the light sensor on the target's mobile device, though there are significant challenges. Such sensors measure one-dimensional light intensity, provide no chromatic information, and have very low sampling rate (normally up to 10Hz). In this paper, we demonstrate that in spite of these challenges, it is possible - based on intensity measurements from a mobile device's light sensor - to make quality inferences regarding the displayed content. We do so by selecting features of measured light that capture information related to transitions between samples. Such features are resilient to ambient noise. In our experiments, involving over 60 hours of collected data and 140 movie clips, we were able to (i) classify content into categories (game, movie, etc) with approximately 90% and 70% accuracy for two-class and four-class classification, respectively; and (ii) identify specific movies or TV programs being played with > 85% accuracy. These findings suggest that access to raw light-sensor readings, which can currently be done without special access controls, may carry nontrivial security ramifications.
Supriyo Chakraborty, Wentao Robin Ouyang, Mani Srivastava 0001
IEEE BigData1
2017 Deep learning for situational understanding
abstract
Situational understanding (SU) requires a combination of insight - the ability to accurately perceive an existing situation - and foresight - the ability to anticipate how an existing situation may develop in the future. SU involves information fusion as well as model representation and inference. Commonly, heterogenous data sources must be exploited in the fusion process: often including both hard and soft data products. In a coalition context, data and processing resources will also be distributed and subjected to restrictions on information sharing. It will often be necessary for a human to be in the loop in SU processes, to provide key input and guidance, and to interpret outputs in a way that necessitates a degree of transparency in the processing: systems cannot be “black boxes”. In this paper, we characterize the Coalition Situational Understanding (CSU) problem in terms of fusion, temporal, distributed, and human requirements. There is currently significant interest in deep learning (DL) approaches for processing both hard and soft data. We analyze the state-of-the-art in DL in relation to these requirements for CSU, and identify areas where there is currently considerable promise, and key gaps.
Supriyo Chakraborty, Alun D. Preece, Moustafa Farid Alzantot, Tianwei Xing, Dave Braines, Mani Srivastava 0001
FUSION1
2017 PrOLoc: resilient localization with private observers using partial homomorphic encryption: demo abstract
abstract
This demo abstract presents PrOLoc, a localization system that combines partially homomorphic encryption with a new way of structuring the localization problem to enable efficient and accurate computation of a target's location while preserving the privacy of the observers.
Amr Al-Anwar 0001, Yasser Shoukry, Supriyo Chakraborty, Bharathan Balaji, Paul Martin 0008, Paulo Tabuada, Mani Srivastava 0001
IPSN3
2017 PrOLoc: resilient localization with private observers using partial homomorphic encryption
abstract
Aided by advances in sensors and algorithms, systems for localizing and tracking target objects or events have become ubiquitous in recent years. Most of these systems operate on the principle of fusing measurements of distance and/or direction to the target made by a set of spatially distributed observers using sensors that measure signals such as RF, acoustic, or optical. The computation of the target's location is done using multilateration and multiangulation algorithms, typically running at an aggregation node that, in addition to the distance/direction measurements, also needs to know the observers' locations. This presents a privacy risk for an observer that does not trust the aggregation node or other observers and could in turn lead to lack of participation. For example, consider a crowd-sourced sensing system where citizens are required to report security threats, or a smart car, stranded with a malfunctioning GPS, sending out localization requests to neighboring cars - in both cases, observer (i.e., citizens and cars respectively) participation can be increased by keeping their location private. This paper presents PrOLoc, a localization system that combines partially homomorphic encryption with a new way of structuring the localization problem to enable efficient and accurate computation of a target's location without requiring observers to make public their locations or measurements. Moreover, and unlike previously proposed perturbation based techniques, PrOLoc is also resilient to malicious active false data injection attacks. We present two realizations of our approach, provide rigorous theoretical guarantees, and also compare the performance of each against traditional methods. Our experiments on real hardware demonstrate that PrOLoc yields location estimates that are accurate while being at least 500x faster than state-of-art secure function evaluation techniques.
Amr Al-Anwar 0001, Yasser Shoukry, Supriyo Chakraborty, Paul Martin 0008, Paulo Tabuada, Mani Srivastava 0001
IPSN3
2017 Computing messages that reveal selected inferences while protecting others
abstract
We study a lossy coding scenario, posed as an algorithmic optimization problem, where we trade off between the conflicting goals of accuracy and privacy, motivated by scenarios such as the public release of estimates of data that accurately reflect some aspects of the raw data without revealing other sensitive confidential aspects of it, or permitting it to be inferred. More precisely, given a discrete probability distribution p(D, X, Y), where X represents the whitelisted inferences from D and Y represents the blacklisted inferences, we seek to construct a conditional distribution p(M|D) with the dual goals of making I(M; X) large and I(M; Y) small. Chakraborty et al. 2013 [1] provided optimal solutions within this model to the two extreme points on the two objectives' Pareto frontier: maximizing I(M; X) subject to the constraint that I(M; Y) be as small as possible (“perfect privacy”, using linear programming (LP)) and vice versa (“perfect utility”, which is trivial). In this paper we provide a faster combinatorial optimal algorithm for the perfect privacy problem, which does not require the use of an LP solver. Moreover, this algorithm any be used to compute Pareto-optimal solutions at any point on the Pareto frontier. (En route to this algorithm, we also provide a mathematical programming-based solution.) This solves the primary open problem posed by [1].
Matthew P. Johnson 0001, Supriyo Chakraborty
ITW3
2016 mSieve: differential behavioral privacy in time series of mobile sensor data
abstract
Differential privacy concepts have been successfully used to protect anonymity of individuals in population-scale analysis. Sharing of mobile sensor data, especially physiological data, raise different privacy challenges, that of protecting private behaviors that can be revealed from time series of sensor data. Existing privacy mechanisms rely on noise addition and data perturbation. But the accuracy requirement on inferences drawn from physiological data, together with well-established limits within which these data values occur, render traditional privacy mechanisms inapplicable. In this work, we define a new behavioral privacy metric based on differential privacy and propose a novel data substitution mechanism to protect behavioral privacy. We evaluate the efficacy of our scheme using 660 hours of ECG, respiration, and activity data collected from 43 participants and demonstrate that it is possible to retain meaningful utility, in terms of inference accuracy (90%), while simultaneously preserving the privacy of sensitive behaviors.
Nazir Saleheen, Supriyo Chakraborty, Nasir Ali, Syed Monowar Hossain, Rummana Bari, Eugene H. Buder, Mani Srivastava 0001, Santosh Kumar 0001
UbiComp2
2016 Dependence Makes You Vulnberable: Differential Privacy Under Dependent Tuples
Changchang Liu, Supriyo Chakraborty, Prateek Mittal
NDSS2
2015 AnonyCast: privacy-preserving location distribution for anonymous crowd tracking systems
abstract
Fusion of infrastructure-based pedestrian tracking systems and embedded sensors on mobile devices holds promise for providing accurate positioning in large public buildings. However, privacy concerns regarding handling of sensitive user location data potentially disrupt the adoption of such systems. This paper presents AnonyCast, a novel privacy-aware mechanism for delivering precise location information measured by crowd-tracking systems to individual pedestrians' smartphones. AnonyCast uses sparsely placed Bluetooth Low Energy transmitters to advertise location-dependent, time-varying keys. Using location measurements, AnonyCast estimates a subset of keys that each pedestrian's phone receives along its path. By combining a cryptography scheme called CP-ABE with a novel greedy algorithm for key selection, it encrypts each path before publishing, allowing users to decrypt only their own trajectories. The results from field experiments show that AnonyCast delivers accurate locations over 84% of time, bounding probability of unauthorized access to one's location below 1%.
Takamasa Higuchi, Paul Martin 0008, Supriyo Chakraborty, Mani Srivastava 0001
UbiComp3
2014 Towards a rich sensing stack for IoT devices
abstract
The broad spectrum of interconnected sensors and actuators, available on various mobile devices and smartphones and collectively defined as the Internet of Things (IoT), have evolved into platforms with ability to both collect personal sensory data and also change the users' immediate environment. The continuous streams of richly annotated sensory data on these IoT devices have also enabled the emergence of a new class of context-aware apps that use the data to infer user context and accordingly customize their responses in real-time. However, this growth in the number of apps has not been complemented with adequate system support on the IoT devices resulting in monolithic apps that each implement and execute their own customized sensing pipelines. In this paper, we outline our vision of a sensing stack, akin to a networking stack, that can facilitate the development and execution of context-aware apps on IoT devices. There are several advantages to building a rich sensing stack. First, it allows apps to reuse stages of the sensing pipeline easing their development. Second, the layers of the stack allow for both in- and cross-layer resource optimization. Finally, it allows better control over the shared data as instead of raw-sensor data, higher-level semantic abstractions, such as inferences can now be shared with apps. We describe our initial efforts towards creating the different building blocks of such a sensing stack.
Chenguang Shen, Haksoo Choi, Supriyo Chakraborty, Mani Srivastava 0001
ICCAD3
2014 ipShield: A Framework For Enforcing Context-Aware Privacy
Supriyo Chakraborty, Chenguang Shen, Kasturi Rangan Raghavan, Yasser Shoukry, Matt Millar, Mani Srivastava 0001
NSDI1
2014 Inference management, trust and obfuscation principles for quality of information in emerging pervasive environments
Chatschik Bisdikian, Christopher Gibson, Supriyo Chakraborty, Mani Srivastava 0001, Murat Sensoy, Timothy J. Norman
Pervasive Mob. Comput.3
2013 Reasoning under uncertainty: Variations of subjective logic deduction
Lance M. Kaplan, Murat Sensoy, Supriyo Chakraborty, Chatschik Bisdikian, Geeth de Mel
FUSION4
2013 Protecting data against unwanted inferences
abstract
We study the competing goals of utility and privacy as they arise when a provider delegates the processing of its personal information to a recipient who is better able to handle this data. We formulate our goals in terms of the inferences which can be drawn using the shared data. A whitelist describes the inferences that are desirable, i.e., providing utility. A blacklist describes the unwanted inferences which the provider wants to keep private. We formally define utility and privacy parameters using elementary information-theoretic notions and derive a bound on the region spanned by these parameters. We provide constructive schemes for achieving certain boundary points of this region. Finally, we improve the region by sharing data over aggregated time slots.
Supriyo Chakraborty, Nicolas Bitouze, Mani Srivastava 0001, Lara Dolecek
ITW1
2012 Model-based context privacy for personal data streams
abstract
Smart phones with increased computation and sensing capabilities have enabled the growth of a new generation of applications which are organic and designed to react depending on the user contexts. These contexts typically define the personal, social, work and urban spaces of an individual and are derived from the underlying sensor measurements. The shared context streams therefore embed in them information, which when stitched together can reveal behavioral patterns and possible sensitive inferences, raising serious privacy concerns. In this paper, we propose a model based technique to capture the relationship between these contexts, and better understand the privacy implications of sharing them. We further demonstrate that by using a generative model of the context streams we can simultaneously meet the utility objectives of the context-aware applications while maintaining individual privacy. We present our current implementation which uses offline model learning with online inferencing performed on the smart phone. Preliminary results are presented to provide proof-of-concept of our proposed technique.
Supriyo Chakraborty, Kasturi Rangan Raghavan, Mani Srivastava 0001, Harris Teague
CCS1
2012 Balancing value and risk in information sharing through obfuscation
Supriyo Chakraborty, Kasturi Rangan Raghavan, Mani Srivastava 0001, Chatschik Bisdikian, Lance M. Kaplan
FUSION1
2012 Subjective logic with uncertain partial observations
Lance M. Kaplan, Supriyo Chakraborty, Chatschik Bisdikian
FUSION2
2012 Design and Evaluation of SensorSafe: A Framework for Achieving Behavioral Privacy in Sharing Personal Sensory Information
abstract
Continuous collection of sensory information using smartphones and body-worn sensors is now feasible with recent advancement of technologies. Sharing such personal information enables many useful applications such as medical behavioral studies, personal health-care, and participatory sensing. However, sharing such information along with inferences that can be drawn from the data increases user's various privacy concerns. This paper proposes SensorSafe, an application framework that enables users to share adequate amounts of their private data and supports obfuscation of sensitive information to protect user privacy. Our framework provides rule-based sharing with context-awareness and conflicting rule detection. In addition, our framework includes several optimization techniques for database processing of rule-based sharing and data obfuscation. We evaluate the optimization techniques with a large amount of accelerometer data from the fine-grained posture recognition application, which is about 6.25 GB.
Haksoo Choi, Supriyo Chakraborty, Mani Srivastava 0001
TrustCom2
2012 Balancing behavioral privacy and information utility in sensory data flows
Supriyo Chakraborty, Zainul Charbiwala, Haksoo Choi, Kasturi Rangan Raghavan, Mani Srivastava 0001
Pervasive Mob. Comput.1
2011 Distributed coordination for fast iterative optimization in wireless sensor/actuator networks
abstract
Large-scale coordination and control problems in sensor/actuator networks are often expressed within the networked optimization model. While significant advances have taken place in both first- and higher-order optimization techniques, their widespread adoption in practical implementations has been hindered by a lack of adequate programming and evaluation support. This motivates the two major contributions of this paper. First, we extend the distributed programming framework proposed in with a synchronization primitive to implement different versions of the subgradient technique and perform extensive evaluation with varying deployment and algorithmic parameters. Second, the insights - obtained by observing the variability in practical metrics such as response time and incurred message cost - lead us to exploit the spatial locality inherent in these large-scale actuator control applications, and propose a novel consensus algorithm applied to the subgradient method. We show using simulations that there is at least 99% improvement in response time and the message cost is reduced by more than 90% over prior consensus based algorithms.
Rahul Balani, Mohamed Nabil Hajj Chehade, Supriyo Chakraborty, Mani Srivastava 0001
SECON3
2011 Neighborhood based fast graph search in large networks
abstract
Complex social and information network search becomes important with a variety of applications. In the core of these applications, lies a common and critical problem: Given a labeled network and a query graph, how to efficiently search the query graph in the target network. The presence of noise and the incomplete knowledge about the structure and content of the target network make it unrealistic to find an exact match. Rather, it is more appealing to find the top-k approximate matches.
Arijit Khan 0001, Xifeng Yan, Ziyu Guan, Supriyo Chakraborty, Shu Tao
SIGMOD Conference5
2011 Estimating network link characteristics using packet-pair dispersion: A discrete-time queueing theoretic analysis
Bikash Kumar Dey, D. Manjunath, Supriyo Chakraborty
Comput. Networks3
2010 Compressive Oversampling for Robust Data Transmission in Sensor Networks
abstract
Data loss in wireless sensing applications is inevitable and while there have been many attempts at coping with this issue, recent developments in the area of Compressive Sensing (CS) provide a new and attractive perspective. Since many physical signals of interest are known to be sparse or compressible, employing CS, not only compresses the data and reduces effective transmission rate, but also improves the robustness of the system to channel erasures. This is possible because reconstruction algorithms for compressively sampled signals are not hampered by the stochastic nature of wireless link disturbances, which has traditionally plagued attempts at proactively handling the effects of these errors. In this paper, we propose that if CS is employed for source compression, then CS can further be exploited as an application layer erasure coding strategy for recovering missing data. We show that CS erasure encoding (CSEC) with random sampling is efficient for handling missing data in erasure channels, paralleling the performance of BCH codes, with the added benefit of graceful degradation of the reconstruction error even when the amount of missing data far exceeds the designed redundancy. Further, since CSEC is equivalent to nominal oversampling in the incoherent measurement basis, it is computationally cheaper than conventional erasure coding. We support our proposal through extensive performance studies.
Zainul Charbiwala, Supriyo Chakraborty, Sadaf Zahedi, Younghun Kim, Ting He 0001, Chatschik Bisdikian, Mani Srivastava 0001
INFOCOM2