EDBT 2026 Demo / reviewers in the wild / expert
Soumik Sarkar
dblp:33/7053
· DBLP profile ↗
40ranked-venue papers
4as first author
25since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 24 · 16 since 2021Graphics, computer vision, multimedia, augmented reality and games · 18 · 3 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DeCAF: Decentralized consensus-and-factorization for low-rank adaptation of foundation models
Nastaran Saadati, Zhanhong Jiang, Joshua R. Waite, Shreyan Ganguly, Aditya Balu, Chinmay Hegde, Soumik Sarkar |
Neural Networks | 7 |
| 2026 | Enhancing PPO With Trajectory-Aware Hybrid PoliciesabstractProximal policy optimization (PPO) is one of the most popular state-of-the-art on-policy algorithms that has become a standard baseline in modern reinforcement learning with applications in numerous fields. Though it delivers stable performance with theoretical policy improvement guarantees, high variance and high sample complexity still remain critical challenges in on-policy algorithms. To alleviate these issues, we propose a hybrid-policy PPO (HP3O), which utilizes a trajectory replay buffer to make efficient use of trajectories generated by recent policies. Particularly, the buffer applies the "first in, first out" (FIFO) strategy so as to keep only the recent trajectories to attenuate the data distribution drift. A batch consisting of the trajectory with the best return and other randomly sampled ones from the buffer is used for updating the policy networks. The strategy helps the agent to improve its capability on top of the most recent best performance and, in turn, reduce variance empirically. We theoretically construct the policy improvement guarantees for the proposed algorithm. HP3O is validated and compared against several baseline algorithms using multiple continuous control environments. Our code is available at https://anonymous.4open.science/r/HP30-EB61/HP3O_train.py. Qisai Liu, Zhanhong Jiang, Hsin-Jung Yang, Mahsa Khosravi, Joshua R. Waite, Soumik Sarkar |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2025 | Foundation Model Efficient Fine-Tuning in Centralized and Federated SettingsabstractParameter-efficient fine-tuning (PEFT) is a critical approach for adapting large models to downstream tasks with reduced computational cost. LoRA, a popular PEFT method, fine-tunes low-rank matrices and has shown strong empirical results. Recently, ReFT emerged as a more effective alternative by updating hidden representations instead of parameters. However, the convergence behavior of both methods remains poorly understood. In this work, we bridge LoRA and ReFT under a unified meta-framework called Model Efficient Fine-Tuning (MeFT). MeFT offers provable convergence guarantees with stochastic gradient descent and reveals how low-rank structures impact convergence error. We extend MeFT to federated learning as FedMeFT, with theoretical analysis and validation across foundation models and benchmarks in centralized and federated settings. Nastaran Saadati, Zhanhong Jiang, Aditya Balu, Chao Liu 0028, Chinmay Hegde, Soumik Sarkar |
IEEE Big Data | 6 |
| 2025 | Latent Safety-Constrained Policy Approach for Safe Offline Reinforcement LearningabstractIn safe offline reinforcement learning, the objective is to develop a policy that maximizes cumulative rewards while strictly adhering to safety constraints, utilizing only offline data. Traditional methods often face difficulties in balancing these constraints, leading to either diminished performance or increased safety risks. We address these issues with a novel approach that begins by learning a conservatively safe policy through the use of Conditional Variational Autoencoders, which model the latent safety constraints. Subsequently, we frame this as a Constrained Reward-Return Maximization problem, wherein the policy aims to optimize rewards while complying with the inferred latent safety constraints. This is achieved by training an encoder with a reward-Advantage Weighted Regression objective within the latent constraint space. Our methodology is supported by theoretical analysis, including bounds on policy performance and sample complexity. Extensive empirical evaluation on benchmark datasets, including challenging autonomous driving scenarios, demonstrates that our approach not only maintains safety compliance but also excels in cumulative reward optimization, surpassing existing methods. Additional visualizations provide further insights into the effectiveness and underlying mechanisms of our approach. Prajwal Koirala, Zhanhong Jiang, Soumik Sarkar, Cody H. Fleming |
ICLR | 3 |
| 2025 | Leveraging Vision Language Models for Specialized Agricultural TasksabstractAs Vision Language Models (VLMs) become increasingly accessible to farmers and agricultural experts, there is a growing need to evaluate their potential in specialized tasks. We present AgEval, a comprehensive benchmark for assessing VLMs' capabilities in plant stress phenotyping, offering a solution to the challenge of limited annotated data in agriculture. Our study explores how general-purpose VLMs can be leveraged for domain-specific tasks with only a few annotated examples, providing insights into their behavior and adaptability. AgEval encompasses 12 diverse plant stress phenotyping tasks, evaluating zero-shot and few-shot in-context learning performance of state-of-the-art models including Claude, GPT, Gemini, and LLaVA. Our results demonstrate VLMs' rapid adaptability to specialized tasks, with the best-performing model showing an increase in F1 scores from 46.24% to 73.37% in 8-shot identification. To quantify performance disparities across classes, we introduce metrics such as the coefficient of variation (CV), revealing that VLMs' training impacts classes differently, with CV ranging from 26.02% to 58.03%. We also find that strategic example selection enhances model reliability, with exact category examples improving F1 scores by 15.38% on average. AgEval establishes a framework for assessing VLMs in agricultural applications, offering valuable benchmarks for future evaluations. Our findings suggest that VLMs, with minimal few-shot examples, show promise as a viable alternative to traditional specialized models in plant stress phenotyping, while also highlighting areas for further refinement. Results and benchmark details are available at: https://github.com/arbab-ml/AgEval Muhammad Arbab Arshad, Talukder Z. Jubery, Tirtho Roy, Rim Nassiri, Asheesh K. Singh, Arti Singh, Chinmay Hegde, Baskar Ganapathysubramanian, Aditya Balu, Adarsh Krishnamurthy, Soumik Sarkar |
WACV | 11 |
| 2024 | DIMAT: Decentralized Iterative Merging-And-Training for Deep Learning ModelsabstractRecent advances in decentralized deep learning algorithms have demonstrated cutting-edge performance on various tasks with large pretrained models. However, a pivotal prerequisite for achieving this level of competitiveness is the significant communication and computation overheads when updating these models, which prohibits the applications of them to real-world scenarios. To address this issue, drawing inspiration from advanced model merging techniques without requiring additional training, we introduce the Decentralized Iterative Merging-And-Training (DIMAT) paradigm-a novel decentralized deep learning framework. Within DIMAT, each agent is trained on their local data and periodically merged with their neighboring agents using advanced model merging techniques like activation matching until convergence is achieved. DIMAT provably converges with the best available rate for non-convex functions with various first-order methods, while yielding tighter error bounds compared to the popular existing approaches. We conduct a comprehensive empirical analysis to validate DIMAT's superiority over baselines across diverse computer vision tasks sourced from multiple datasets. Empirical results validate our theoretical claims by showing that DIMAT attains faster and higher initial gain in accuracy with independent and identically distributed (IID) and non-IID data, incurring lower communication overhead. This DIMAT paradigm presents a new op-portunity for the future decentralized learning, enhancing its adaptability to real-world with sparse and lightweight communication and computation. Nastaran Saadati, Minh Pham 0005, Nasla Saleem, Joshua R. Waite, Aditya Balu, Zhanhong Jiang, Chinmay Hegde, Soumik Sarkar |
CVPR | 8 |
| 2024 | BioTrove: A Large Curated Image Dataset Enabling AI for BiodiversityabstractWe introduce BioTrove, the largest publicly accessible dataset designed to advance AI applications in biodiversity. Curated from the iNaturalist platform and vetted to include only research-grade data, BioTrove contains 161.9 million images, offering unprecedented scale and diversity from three primary kingdoms: Animalia ("animals"), Fungi ("fungi"), and Plantae ("plants"), spanning approximately 366.6K species. Each image is annotated with scientific names, taxonomic hierarchies, and common names, providing rich metadata to support accurate AI model development across diverse species and ecosystems.We demonstrate the value of BioTrove by releasing a suite of CLIP models trained using a subset of 40 million captioned images, known as BioTrove-Train. This subset focuses on seven categories within the dataset that are underrepresented in standard image recognition models, selected for their critical role in biodiversity and agriculture: Aves ("birds"), Arachnida} ("spiders/ticks/mites"), Insecta ("insects"), Plantae ("plants"), Fungi ("fungi"), Mollusca ("snails"), and Reptilia ("snakes/lizards"). To support rigorous assessment, we introduce several new benchmarks and report model accuracy for zero-shot learning across life stages, rare species, confounding species, and multiple taxonomic levels.We anticipate that BioTrove will spur the development of AI models capable of supporting digital tools for pest control, crop monitoring, biodiversity assessment, and environmental conservation. These advancements are crucial for ensuring food security, preserving ecosystems, and mitigating the impacts of climate change. BioTrove is publicly available, easily accessible, and ready for immediate use. Chih-Hsuan Yang, Benjamin Feuer, Talukder Z. Jubery, Zi K. Deng, Andre Nakkab, Md. Zahid Hasan, Shivani Chiranjeevi, Kelly O. Marshall, Nirmal Baishnab, Asheesh Kumar Singh, Arti Singh, Soumik Sarkar, Nirav C. Merchant, Chinmay Hegde, Baskar Ganapathysubramanian |
NeurIPS | 12 |
| 2024 | Latent Diffusion Models for Structural Component Design
Ethan Herron, Jaydeep Rade, Anushrut Jignasu, Baskar Ganapathysubramanian, Aditya Balu, Soumik Sarkar, Adarsh Krishnamurthy |
Comput. Aided Des. | 6 |
| 2024 | Neural PDE Solvers for Irregular Domains
Biswajit Khara, Ethan Herron, Aditya Balu, Dhruv Gamdha, Chih-Hsuan Yang, Anushrut Jignasu, Zhanhong Jiang, Soumik Sarkar, Chinmay Hegde, Baskar Ganapathysubramanian, Adarsh Krishnamurthy |
Comput. Aided Des. | 9 |
| 2024 | Dominating Set Model Aggregation for communication-efficient decentralized deep learning
Fateme Fotouhi, Aditya Balu, Zhanhong Jiang, Yasaman Esfandiari, Salman Jahani, Soumik Sarkar |
Neural Networks | 6 |
| 2024 | Vision-Language Models Can Identify Distracted Driver Behavior From Naturalistic VideosabstractRecognizing the activities causing distraction in real-world driving scenarios is critical for ensuring the safety and reliability of both drivers and pedestrians on the roadways. Conventional computer vision techniques are typically data-intensive and require a large volume of annotated training data to detect and classify various distracted driving behaviors, thereby limiting their generalization ability, efficiency and scalability. We aim to develop a generalized framework that showcases robust performance with access to limited or no annotated training data. Recently, vision-language models have offered large-scale visual-textual pretraining that can be adapted to task-specific learning like distracted driving activity recognition. Vision-language pretraining models like CLIP have shown significant promise in learning natural language-guided visual representations. This paper proposes a CLIP-based driver activity recognition approach that identifies driver distraction from naturalistic driving images and videos. CLIP’s vision embedding offers zero-shot transfer and task-based finetuning, which can classify distracted activities from naturalistic driving video. Our results show that this framework offers state-of-the-art performance on zero-shot transfer, finetuning and video-based models for predicting the driver’s state on four public datasets. We propose frame-based and video-based frameworks developed on top of the CLIP’s visual representation for distracted driving detection and classification tasks and report the results. Our code is available at https://github.com/zahid-isu/DriveCLIP Md. Zahid Hasan, Jiajing Chen, Jiyang Wang, Mohammed Shaiqur Rahman, Ameya Joshi, Senem Velipasalar, Chinmay Hegde, Anuj Sharma 0001, Soumik Sarkar |
IEEE Trans. Intell. Transp. Syst. | 9 |
| 2023 | Deep learning-based 3D multigrid topology optimization of manufacturable designs
Jaydeep Rade, Anushrut Jignasu, Ethan Herron, Ashton M. Corpuz, Baskar Ganapathysubramanian, Soumik Sarkar, Aditya Balu, Adarsh Krishnamurthy |
Eng. Appl. Artif. Intell. | 6 |
| 2022 | MDPGT: Momentum-Based Decentralized Policy Gradient TrackingabstractWe propose a novel policy gradient method for multi-agent reinforcement learning, which leverages two different variance-reduction techniques and does not require large batches over iterations. Specifically, we propose a momentum-based decentralized policy gradient tracking (MDPGT) where a new momentum-based variance reduction technique is used to approximate the local policy gradient surrogate with importance sampling, and an intermediate parameter is adopted to track two consecutive policy gradient surrogates. MDPGT provably achieves the best available sample complexity of O(N -1 e -3) for converging to an e-stationary point of the global average of N local performance functions (possibly nonconcave). This outperforms the state-of-the-art sample complexity in decentralized model-free reinforcement learning and when initialized with a single trajectory, the sample complexity matches those obtained by the existing decentralized policy gradient methods. We further validate the theoretical claim for the Gaussian policy function. When the required error tolerance e is small enough, MDPGT leads to a linear speed up, which has been previously established in decentralized stochastic optimization, but not for reinforcement learning. Lastly, we provide empirical results on a multi-agent reinforcement learning benchmark environment to support our theoretical findings. Zhanhong Jiang, Xian Yeow Lee, Sin Yong Tan, Kai Liang Tan, Aditya Balu, Young M. Lee, Chinmay Hegde, Soumik Sarkar |
AAAI | 8 |
| 2022 | NURBS-Diff: A Differentiable Programming Module for NURBS
Anjana Deva Prasad, Aditya Balu, Harshil Shah, Soumik Sarkar, Chinmay Hegde, Adarsh Krishnamurthy |
Comput. Aided Des. | 4 |
| 2022 | The Stochastic Augmented Lagrangian method for domain adaptation
Zhanhong Jiang, Chao Liu 0028, Young M. Lee, Chinmay Hegde, Soumik Sarkar, Dongxiang Jiang |
Knowl. Based Syst. | 5 |
| 2021 | Decentralized Deep Learning Using Momentum-Accelerated ConsensusabstractWe consider the problem of decentralized deep learning where multiple agents collaborate to learn from a distributed dataset. While several decentralized deep learning approaches exist, the majority consider a central parameter-server topology for aggregating the model parameters from the agents. However, such a topology may be inapplicable in networked systems such as ad-hoc mobile networks, field robotics, and power network systems where direct communication with the central parameter server may be inefficient. In this context, we propose and analyze a novel decentralized deep learning algorithm where the agents interact over a fixed communication topology (without a central server). Our algorithm is based on the heavy-ball acceleration method used in gradient-based optimization. We propose a novel consensus protocol where each agent shares with its neighbors its model parameters and gradient-momentum values during the optimization process. We consider nonconvex objective functions and theoretically analyze our algorithm’s performance. We present several empirical comparisons with competing decentralized learning methods to demonstrate the efficacy of our approach under different communication topologies. Aditya Balu, Zhanhong Jiang, Sin Yong Tan, Chinmay Hegde, Young M. Lee, Soumik Sarkar |
ICASSP | 6 |
| 2021 | Spatiotemporal Attention for Multivariate Time Series Prediction and InterpretationabstractMultivariate time series modeling and prediction problems are abundant in many machine learning application domains. Accurate interpretation of the prediction outcomes from the model can significantly benefit the domain experts. In addition to isolating the important time-steps, spatial interpretation is also critical to understand the contributions of different variables on the model output. We propose a novel deep learning architecture, called spatiotemporal attention mechanism (STAM) for simultaneous learning of the most important time steps and variables. STAM is a causal (i.e., only depends on past inputs and does not use future inputs) and scalable (i.e., scales well with an increase in the number of variables) approach that is comparable to the state-of-the-art models in terms of computational tractability. We demonstrate our models’ performance on a popular public dataset and a domain-specific dataset, where the learned attention weights are validated from a domain knowledge perspective. When compared with the baseline models, the results show that STAM maintains state-of-the-art prediction accuracy while offering the benefit of accurate spatiotemporal interpretability. Tryambak Gangopadhyay, Sin Yong Tan, Zhanhong Jiang, Soumik Sarkar |
ICASSP | 5 |
| 2021 | Cross-Gradient Aggregation for Decentralized Learning from Non-IID DataabstractDecentralized learning enables a group of collaborative agents to learn models using a distributed dataset without the need for a central parameter server. Recently, decentralized learning algorithms have demonstrated state-of-the-art results on benchmark data sets, comparable with centralized algorithms. However, the key assumption to achieve competitive performance is that the data is independently and identically distributed (IID) among the agents which, in real-life applications, is often not applicable. Inspired by ideas from continual learning, we propose Cross-Gradient Aggregation (CGA), a novel decentralized learning algorithm where (i) each agent aggregates cross-gradient information, i.e., derivatives of its model with respect to its neighbors’ datasets, and (ii) updates its model using a projected gradient based on quadratic programming (QP). We theoretically analyze the convergence characteristics of CGA and demonstrate its efficiency on non-IID data distributions sampled from the MNIST and CIFAR-10 datasets. Our empirical comparisons show superior learning performance of CGA over existing state-of-the-art decentralized learning algorithms, as well as maintaining the improved performance under information compression to reduce peer-to-peer communication overhead. The code is available here on GitHub. Yasaman Esfandiari, Sin Yong Tan, Zhanhong Jiang, Aditya Balu, Ethan Herron, Chinmay Hegde, Soumik Sarkar |
ICML | 7 |
| 2021 | Differentiable Spline ApproximationsabstractThe paradigm of differentiable programming has significantly enhanced the scope of machine learning via the judicious use of gradient-based optimization. However, standard differentiable programming methods (such as autodiff) typically require that the machine learning models be differentiable, limiting their applicability. Our goal in this paper is to use a new, principled approach to extend gradient-based optimization to functions well modeled by splines, which encompass a large family of piecewise polynomial models. We derive the form of the (weak) Jacobian of such functions and show that it exhibits a block-sparse structure that can be computed implicitly and efficiently. Overall, we show that leveraging this redesigned Jacobian in the form of a differentiable "layer'' in predictive models leads to improved performance in diverse applications such as image segmentation, 3D point cloud reconstruction, and finite element analysis. We also open-source the code at \url{https://github.com/idealab-isu/DSA}. Minsu Cho, Aditya Balu, Ameya Joshi, Anjana Deva Prasad, Biswajit Khara, Soumik Sarkar, Baskar Ganapathysubramanian, Adarsh Krishnamurthy, Chinmay Hegde |
NeurIPS | 6 |
| 2021 | Distributed multigrid neural solvers on megavoxel domainsabstractWe consider the distributed training of large scale neural networks that serve as PDE (partial differential equation) solvers producing full field outputs. We specifically consider neural solvers for the generalized 3D Poisson equation over megavoxel domains. A scalable framework is presented that integrates two distinct advances. First, we accelerate training a large model via a method analogous to the multigrid technique used in numerical linear algebra. Here, the network is trained using a hierarchy of increasing resolution inputs in sequence, analogous to the `V', `W', `F' and `Half-V' cycles used in multigrid approaches. In conjunction with the multi-grid approach, we implement a distributed deep learning framework which significantly reduces the time to solve. We show scalability of this approach on both GPU (Azure VMs on Cloud) and CPU clusters (PSC Bridges2). This approach is deployed to train a generalized 3D Poisson solver that scales well to predict output full field solutions up to the resolution of 512 X 512 X 512 for a high dimensional family of inputs. This strategy opens up the possibility of fast and scalable training of neural PDE solvers on heterogeneous clusters. Aditya Balu, Sergio Botelho, Biswajit Khara, Vinay Rao, Soumik Sarkar, Chinmay Hegde, Adarsh Krishnamurthy, Santi Adavani, Baskar Ganapathysubramanian |
SC | 5 |
| 2021 | Multi-resolution 3D CNN for learning multi-scale spatial features in CAD models
Sambit Ghadai, Xian Yeow Lee, Aditya Balu, Soumik Sarkar, Adarsh Krishnamurthy |
Comput. Aided Geom. Des. | 4 |
| 2021 | Algorithmically-consistent deep learning frameworks for structural topology optimization
Jaydeep Rade, Aditya Balu, Ethan Herron, Jay Pathak, Rishikesh Ranade, Soumik Sarkar, Adarsh Krishnamurthy |
Eng. Appl. Artif. Intell. | 6 |
| 2021 | Segmentation of natural images based on super pixel and graph mergingabstractAbstract The task of natural image segmentation is one of the most researched topics of computer vision. There are mainly two principal approaches for the task, the statistical approach and the supervised approach. The proposed methodology segments natural images combining a set of statistical algorithms. First, the image is preprocessed to enhance the edges. Weighted average of the denoised image and its derivatives is the preprocessed output. Thereafter, an energy based super pixelation is applied to over segment the image. Finally, a connectivity graph is built where nodes correspond to super pixels and edges connect the adjacent super pixels. The adjacent super pixels are merged based on the confidence value defined in terms of their textural and colour similarity. Proposed methodology has been applied on the images of BSDS500 dataset. Performance of the proposed work has been compared with that of other works based on detected edge maps. Few works generate ultrametric contour maps (UCM). To compare the performance with those works, UCM is also generated by the proposed methodology. To do so images at multiple scales are considered. It is observed that the output of segmentation is better in case of the proposed methodology. Proposed methodology is much faster than others. Thus, makes it suitable for real time application in robot vision. Aritra Mukherjee, Soumik Sarkar, Sanjoy Kumar Saha 0001 |
IET Comput. Vis. | 2 |
| 2021 | Root-cause analysis for time-series anomalies via spatiotemporal graphical modeling in distributed complex systems
Chao Liu 0028, Kin Gwn Lore, Zhanhong Jiang, Soumik Sarkar |
Knowl. Based Syst. | 4 |
| 2021 | A fast saddle-point dynamical system approach to robust deep learning
Yasaman Esfandiari, Aditya Balu, Keivan Ebrahimi, Umesh Vaidya, Nicola Elia, Soumik Sarkar |
Neural Networks | 6 |
| 2020 | InvNet: Encoding Geometric and Statistical Invariances in Deep Generative ModelsabstractGenerative Adversarial Networks (GANs), while widely successful in modeling complex data distributions, have not yet been sufficiently leveraged in scientific computing and design. Reasons for this include the lack of flexibility of GANs to represent discrete-valued image data, as well as the lack of control over physical properties of generated samples. We propose a new conditional generative modeling approach (InvNet) that efficiently enables modeling discrete-valued images, while allowing control over their parameterized geometric and statistical properties. We evaluate our approach on several synthetic and real world problems: navigating manifolds of geometric shapes with desired sizes; generation of binary two-phase materials; and the (challenging) problem of generating multi-orientation polycrystalline microstructures. Ameya Joshi, Minsu Cho, Viraj Shah, Balaji Sesha Sarath Pokuri, Soumik Sarkar, Baskar Ganapathysubramanian, Chinmay Hegde |
AAAI | 5 |
| 2020 | Spatiotemporally Constrained Action Space Attacks on Deep Reinforcement Learning AgentsabstractRobustness of Deep Reinforcement Learning (DRL) algorithms towards adversarial attacks in real world applications such as those deployed in cyber-physical systems (CPS) are of increasing concern. Numerous studies have investigated the mechanisms of attacks on the RL agent's state space. Nonetheless, attacks on the RL agent's action space (corresponding to actuators in engineering systems) are equally perverse, but such attacks are relatively less studied in the ML literature. In this work, we first frame the problem as an optimization problem of minimizing the cumulative reward of an RL agent with decoupled constraints as the budget of attack. We propose the white-box Myopic Action Space (MAS) attack algorithm that distributes the attacks across the action space dimensions. Next, we reformulate the optimization problem above with the same objective function, but with a temporally coupled constraint on the attack budget to take into account the approximated dynamics of the agent. This leads to the white-box Look-ahead Action Space (LAS) attack algorithm that distributes the attacks across the action and temporal dimensions. Our results showed that using the same amount of resources, the LAS attack deteriorates the agent's performance significantly more than the MAS attack. This reveals the possibility that with limited resource, an adversary can utilize the agent's dynamics to malevolently craft attacks that causes the agent to fail. Additionally, we leverage these attack strategies as a possible tool to gain insights on the potential vulnerabilities of DRL agents. Xian Yeow Lee, Sambit Ghadai, Kai Liang Tan, Chinmay Hegde, Soumik Sarkar |
AAAI | 5 |
| 2020 | Detecting and Tracking Unsafe Lane Departure Events for Predicting Driver Safety in Challenging Naturalistic Driving DataabstractOur goal is to improve driver safety predictions in at-risk medical or aging populations from naturalistic driving video data. To meet this goal, we developed a novel model capable of detecting and tracking unsafe lane departure events (e.g., changes and incursions), which may occur more frequently in at-risk driver populations. The model detects and tracks roadway lane markings in challenging, low-resolution driving videos using a semantic lane detection pre-processor (Mask R-CNN) utilizing the driver's forward lane region, demarking the convex hull that represents the driver's lane. The hull centroid is tracked over time, improving lane tracking over approaches which detect lane markers from single video frames. The lane time series was denoised using a Fix-lag Kalman filter. Preliminary results show promise for robust lane departure event detection. Overall recall for detecting lane departure events was 81.82%. The F1 score was 75% (precision 69.23%) and 70.59% (precision 62.07%) for left and right lane departures, respectively. Future investigations include exploring (1) horizontal offset as a means to detect lead vehicle proximity, even when image perspectives are known to have a chirp effect and (2) Long Short Term Memory (LSTM) models to detect peaks instead of a peak detection algorithm. Luis G. Riera, Koray Ozcan, Jennifer Merickel, Matthew Rizzo, Soumik Sarkar, Anuj Sharma 0001 |
IV | 5 |
| 2019 | Semantic Adversarial Attacks: Parametric Transformations That Fool Deep ClassifiersabstractDeep neural networks have been shown to exhibit an intriguing vulnerability to adversarial input images corrupted with imperceptible perturbations. However, the majority of adversarial attacks assume global, fine-grained control over the image pixel space. In this paper, we consider a different setting: what happens if the adversary could only alter specific attributes of the input image? These would generate inputs that might be perceptibly different, but still natural-looking and enough to fool a classifier. We propose a novel approach to generate such ``semantic'' adversarial examples by optimizing a particular adversarial loss over the range-space of a parametric conditional generative model. We demonstrate implementations of our attacks on binary classifiers trained on face images, and show that such natural-looking semantic adversarial examples exist. We evaluate the effectiveness of our attack on synthetic and real data, and present detailed comparisons with existing attack methods. We supplement our empirical results with theoretical bounds that demonstrate the existence of such parametric adversarial examples. Ameya Joshi, Amitangshu Mukherjee, Soumik Sarkar, Chinmay Hegde |
ICCV | 3 |
| 2018 | Online Robust Policy Learning in the Presence of Unknown AdversariesabstractThe growing prospect of deep reinforcement learning (DRL) being used in cyber-physical systems has raised concerns around safety and robustness of autonomous agents. Recent work on generating adversarial attacks have shown that it is computationally feasible for a bad actor to fool a DRL policy into behaving sub optimally. Although certain adversarial attacks with specific attack models have been addressed, most studies are only interested in off-line optimization in the data space (e.g., example fitting, distillation). This paper introduces a Meta-Learned Advantage Hierarchy (MLAH) framework that is attack model-agnostic and more suited to reinforcement learning, via handling the attacks in the decision space (as opposed to data space) and directly mitigating learned bias introduced by the adversary. In MLAH, we learn separate sub-policies (nominal and adversarial) in an online manner, as guided by a supervisory master agent that detects the presence of the adversary by leveraging the advantage function for the sub-policies. We demonstrate that the proposed algorithm enables policy learning with significantly lower bias as compared to the state-of-the-art policy learning approaches even in the presence of heavy state information attacks. We present algorithm analysis and simulation results using popular OpenAI Gym environments. Aaron J. Havens, Zhanhong Jiang, Soumik Sarkar |
NeurIPS | 3 |
| 2018 | Learning localized features in 3D CAD models for manufacturability analysis of drilled holes
Sambit Ghadai, Aditya Balu, Soumik Sarkar, Adarsh Krishnamurthy |
Comput. Aided Geom. Des. | 3 |
| 2018 | A deep learning framework for causal shape transformation
Kin Gwn Lore, Daniel Stoecklein, Baskar Ganapathysubramanian, Soumik Sarkar |
Neural Networks | 5 |
| 2018 | Hierarchical symbolic dynamic filtering of streaming non-stationary time series data
Adedotun Akintayo, Soumik Sarkar |
Signal Process. | 2 |
| 2017 | Collaborative Deep Learning in Fixed Topology NetworksabstractThere is significant recent interest to parallelize deep learning algorithms in order to handle the enormous growth in data and model sizes. While most advances focus on model parallelization and engaging multiple computing agents via using a central parameter server, aspect of data parallelization along with decentralized computation has not been explored sufficiently. In this context, this paper presents a new consensus-based distributed SGD (CDSGD) (and its momentum variant, CDMSGD) algorithm for collaborative deep learning over fixed topology networks that enables data parallelization as well as decentralized computation. Such a framework can be extremely useful for learning agents with access to only local/private data in a communication constrained environment. We analyze the convergence properties of the proposed algorithm with strongly convex and nonconvex objective functions with fixed and diminishing step sizes using concepts of Lyapunov function construction. We demonstrate the efficacy of our algorithms in comparison with the baseline centralized SGD and the recently proposed federated averaging algorithm (that also enables data parallelism) based on benchmark datasets such as MNIST, CIFAR-10 and CIFAR-100. Zhanhong Jiang, Aditya Balu, Chinmay Hegde, Soumik Sarkar |
NIPS | 4 |
| 2017 | LLNet: A deep autoencoder approach to natural low-light image enhancement
Kin Gwn Lore, Adedotun Akintayo, Soumik Sarkar |
Pattern Recognit. | 3 |
| 2016 | A composite discretization scheme for symbolic identification of complex systems
Soumik Sarkar, Abhishek Srivastav |
Signal Process. | 1 |
| 2012 | Optimization of symbolic feature extraction for pattern classification
Soumik Sarkar, Kushal Mukherjee, Xin Jin 0016, Dheeraj S. Singh, Asok Ray |
Signal Process. | 1 |
| 2012 | Statistical Mechanics-Inspired Modeling of Heterogeneous Packet Transmission in Communication NetworksabstractThis paper presents the qualitative nature of communication network operations as abstraction of typical thermodynamic parameters (e.g., order parameter, temperature, and pressure). Specifically, statistical mechanics-inspired models of critical phenomena (e.g., phase transitions and size scaling) for heterogeneous packet transmission are developed in terms of multiple intensive parameters, namely, the external packet load on the network system and the packet transmission probabilities of heterogeneous packet types. Network phase diagrams are constructed based on these traffic parameters, and decision and control strategies are formulated for heterogeneous packet transmission in the network system. In this context, decision functions and control objectives are derived in closed forms, and the pertinent results of test and validation on a simulated network system are presented. Soumik Sarkar, Kushal Mukherjee, Asok Ray, Abhishek Srivastav, Thomas A. Wettergren |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2009 | Statistical estimation of multiple parameters via symbolic dynamic filtering
Chinmay Rao, Kushal Mukherjee, Soumik Sarkar, Asok Ray |
Signal Process. | 3 |
| 2009 | Generalization of Hilbert transform for symbolic analysis of noisy signals
Soumik Sarkar, Kushal Mukherjee, Asok Ray |
Signal Process. | 1 |