Deep Patel

dblp:226/5175 · DBLP profile ↗
← Back
12ranked-venue papers
4as first author
10since 2021 · last 2025
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 8 · 1 first-author · 6 since 2021Artificial intelligence and machine learning · 5 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2025 Solving Neural Min-Max Games: The Role of Architecture, Initialization & Dynamics
abstract
Many emerging applications—such as adversarial training, AI alignment, and robust optimization—can be framed as zero-sum games between neural nets, with von Neumann–Nash equilibria (NE) capturing the desirable system behavior. While such games often involve non-convex non-concave objectives, empirical evidence shows that simple gradient methods frequently converge, suggesting a hidden geometric structure. In this paper, we provide a theoretical framework that explains this phenomenon through the lens of \emph{hidden convexity} and \emph{overparameterization}. We identify sufficient conditions spanning initialization, training dynamics, and network width—that guarantee global convergence to a NE in a broad class of non-convex min-max games. To our knowledge, this is the first such result for games that involve two-layer neural networks. Technically, our approach is twofold: (a) we derive a novel path-length bound for alternating gradient-descent-ascent scheme in min-max games; and (b) we show that games with hidden convex–concave geometry reduce to settings satisfying two-sided Polyak–Łojasiewicz (PL) and smoothness conditions, which hold with high probability under overparameterization, using tools from random matrix theory.
Deep Patel, Emmanouil V. Vlatakis-Gkaragkounis
NeurIPS1
2025 Data Driven Approach for Rapid Collision Checking Using Cross-Attention on Robot and Obstacle Sets
abstract
Collision checking is the process of determining whether a candidate robot configuration or trajectory intersects any obstacles in the workspace, and it underpins all safe motionplanning algorithms. Reliable real-time collision checking is critical for safe, autonomous robot manipulation in dynamic industrial settings. This paper presents a learning-based collision checking framework that leverages Set Transformers with cross-attention mechanisms to rapidly and accurately predict collisions between robotic manipulators and complex, variable obstacle environments. Unlike traditional geometric approaches, the proposed method efficiently processes a robot's configuration and a dynamic obstacle set to estimate collision likelihood in real time. A synthetic dataset generated via CoppeliaSim for three industrial robot models (UR5, UR10, ABB IRB140) is used for training and evaluation. Experimental results demonstrate an average classification accuracy of 89.6%, with the model significantly outperforming established learning-based methods such as Fastron and CollisionGP, while achieving up to 7.6× faster inference than PyBullet-based checks. The proposed approach offers a scalable, generalizable, and computationally efficient solution for safe robot operation in dynamic and unstructured environments.
Chayan Maiti, Deep Patel, Sreekumar Muthuswamy
TENCON2
2025 Exploiting VLM Localizability and Semantics for Open Vocabulary Action Detection
abstract
Action detection aims to detect (recognize and localize) human actions spatially and temporally in videos. Existing approaches focus on the closed-set setting where an action detector is trained and tested on videos from a fixed set of action categories. However, this constrained setting is not viable in an open world where test videos inevitably come beyond the trained action categories. In this paper, we address the practical yet challenging Open-Vocabulary Action Detection (OVAD) problem. It aims to detect any action in test videos while training a model on a fixed set of action categories. To achieve such an open-vocabulary capability, we propose a novel method OpenMixer that exploits the inherent semantics and localizability of large vision-language models (VLM) within the family of query-based detection transformers (DETR). Specifically, the OpenMixer is developed by spatial and temporal OpertMixer blocks (S-OMB and T-OMB), and a dynamically fused alignment (DFA) module. The three components collectively enjoy the merits of strong generalization from pretrained VLMs and end-to-end learning from DETR design. Moreover, we established OVAD benchmarks under various settings, and the experimental results show that the OpenMixer performs the best over baselines for detecting seen and unseen actions. We release the codes, models, and dataset splits at https://github.com/Cogito2012/0penMixer.
Wentao Bao, Kai Li 0012, Yuxiao Chen 0002, Deep Patel, Martin Renqiang Min, Yu Kong 0001
WACV4
2024 Learning from Synthetic Human Group Activities
abstract
The study of complex human interactions and group activities has become a focal point in human-centric computer vision. However, progress in related tasks is often hindered by the challenges of obtaining large-scale labeled datasets from real-world scenarios. To address the limitation, we introduce M3 Act, a synthetic data generator for multi-view multi-group multi-person human atomic actions and group activities. Powered by Unity Engine, M3 Act features mul-tiple semantic groups, highly diverse and photorealistic images, and a comprehensive set of annotations, which facilitates the learning of human-centered tasks across single-person, multi-person, and multi-group conditions. We demonstrate the advantages of M3 Act across three core experiments. The results suggest our synthetic dataset can significantly improve the performance of several downstream methods and replace real-world datasets to reduce cost. Notably, M3 Act improves the state-of-the-art MOTRv2 on DanceTrack dataset, leading to a hop on the leaderboard from 10thto 2ndplace. Moreover, M3 Act opens new research for controllable 3D group activity generation. We define multiple metrics and propose a competitive baseline for the novel task. Our code and data are available at our project page: http://cjerry1243.github.io/M3Act.
Che-Jui Chang, Danrui Li, Deep Patel, Parth Goel, Honglu Zhou, Seonghyeon Moon, Samuel S. Sohn, Sejong Yoon, Vladimir Pavlovic 0001, Mubbasir Kapadia
CVPR3
2024 Learning to Localize Actions in Instructional Videos with LLM-Based Multi-pathway Text-Video Alignment
Yuxiao Chen 0002, Kai Li 0012, Wentao Bao, Deep Patel, Yu Kong 0001, Martin Renqiang Min, Dimitris N. Metaxas
ECCV (82)4
2024 An Open Framework for Smart Retrofitting of Legacy Industrial Machines Towards Industry 4.0
abstract
This paper presents a six-layer framework for smart retrofitting of legacy industrial machines to align with Industry 4.0 standards, using open-source technologies for cost-effectiveness and scalability. The layers include edge data acquisition, data preprocessing, Digital Twin technology, long-term data storage, AI-based data analysis, and real-time visual-ization. Implemented and evaluated through a case study on a legacy CNC lathe, the framework demonstrated reliable communication and effective performance across various layers. The results validate the framework's feasibility and highlight the need for further fine-tuning of open large language models with manufacturing domain knowledge.
Deep Patel, Sreekumar Muthuswamy
TENCON1
2024 Differentiable JPEG: The Devil is in the Details
abstract
JPEG remains one of the most widespread lossy image coding methods. However, the non-differentiable nature of JPEG restricts the application in deep learning pipelines. Several differentiable approximations of JPEG have recently been proposed to address this issue. This paper conducts a comprehensive review of existing diff. JPEG approaches and identifies critical details that have been missed by previous methods. To this end, we propose a novel diff. JPEG approach, overcoming previous limitations. Our approach is differentiable w.r.t. the input image, the JPEG quality, the quantization tables, and the color conversion parameters. We evaluate the forward and backward performance of our diff. JPEG approach against existing methods. Additionally, extensive ablations are performed to evaluate crucial design choices. Our proposed diff. JPEG resembles the (non-diff.) reference implementation best, significantly surpassing the recent-best diff. approach by 3.47dB (PSNR) on average. For strong compression rates, we can even improve PSNR by 9.51dB. Strong adversarial attack results are yielded by our diff. JPEG, demonstrating the effective gradient approximation. Our code is available at https://github.com/necla-ml/Diff-JPEG.
Christoph Reich, Biplob Debnath, Deep Patel, Srimat T. Chakradhar
WACV3
2023 Source-Free Video Domain Adaptation with Spatial-Temporal-Historical Consistency Learning
abstract
Source-free domain adaptation (SFDA) is an emerging research topic that studies how to adapt a pretrained source model using unlabeled target data. It is derived from unsupervised domain adaptation but has the advantage of not requiring labeled source data to learn adaptive models. This makes it particularly useful in real-world applications where access to source data is restricted. While there has been some SFDA work for images, little attention has been paid to videos. Naively extending image-based methods to videos without considering the unique properties of videos often leads to unsatisfactory results. In this paper, we propose a simple and highly flexible method for Source-Free Video Domain Adaptation (SFVDA), which extensively exploits consistency learning for videos from spatial, temporal, and historical perspectives. Our method is based on the assumption that videos of the same action category are drawn from the same low-dimensional space, regardless of the spatio-temporal variations in the high-dimensional space that cause domain shifts. To overcome domain shifts, we simulate spatio-temporal variations by applying spatial and temporal augmentations on target videos and encourage the model to make consistent predictions from a video and its augmented versions. Due to the simple design, our method can be applied to various SFVDA settings, and experiments show that our method achieves state-of-the-art performance for all the settings.
Kai Li 0012, Deep Patel, Erik Kruus, Martin Renqiang Min
CVPR2
2023 Adaptive Sample Selection for Robust Learning under Label Noise
abstract
Deep Neural Networks (DNNs) have been shown to be susceptible to memorization or overfitting in the presence of noisily-labelled data. For the problem of robust learning under such noisy data, several algorithms have been proposed. A prominent class of algorithms rely on sample selection strategies wherein, essentially, a fraction of samples with loss values below a certain threshold are selected for training. These algorithms are sensitive to such thresholds, and it is difficult to fix or learn these thresholds. Often, these algorithms also require information such as label noise rates which are typically unavailable in practice. In this paper, we propose an adaptive sample selection strategy that relies only on batch statistics of a given mini-batch to provide robustness against label noise. The algorithm does not have any additional hyperparameters for sample selection, does not need any information on noise rates and does not need access to separate data with clean labels. We empirically demonstrate the effectiveness of our algorithm on benchmark datasets.1
Deep Patel, P. S. Sastry 0001
WACV1
2021 Memorization in Deep Neural Networks: Does the Loss Function Matter?
Deep Patel, P. S. Sastry 0001
PAKDD (2)1
2019 Comparison of Speech Tasks and Recording Devices for Voice Based Automatic Classification of Healthy Subjects and Patients with Amyotrophic Lateral Sclerosis
Suhas B. N., Deep Patel, Nithin Rao Koluguri, Yamini Belur, Pradeep Reddy, Atchayaram Nalini, Dipanjan Gope, Prasanta Kumar Ghosh
INTERSPEECH2
2018 Comparison of Speech Tasks for Automatic Classification of Patients with Amyotrophic Lateral Sclerosis and Healthy Subjects
abstract
In this work, we consider the task of acoustic and articulatory feature based automatic classification of Amyotrophic Lateral Sclerosis (ALS) patients and healthy subjects using speech tasks. In particular, we compare the roles of different types of speech tasks, namely rehearsed speech, spontaneous speech and repeated words for this purpose. Simultaneous articulatory and speech data were recorded from 8 healthy controls and 8 ALS patients using AG501 for the classification experiments. In addition to typical acoustic and articulatory features, new articulatory features are proposed for classification. As classifiers, both Deep Neural Networks (DNN) and Support Vector Machines (SVM) are examined. Classification experiments reveal that the proposed articulatory features outperform other acoustic and articulatory features using both DNN and SVM classifier. However, SVM performs better than DNN classifier using the proposed feature. Among three different speech tasks considered, the rehearsed speech was found to provide the highest F-score of 1, followed by an F-score of 0.92 when both repeated words and spontaneous speech are used for classification.
Aravind Illa, Deep Patel, Yamini Belur, Meera SS, N. Shivashankar, Preethish-Kumar Veeramani, Seena Vengalil, Kiran Polavarapu, Saraswati Nashi, Atchayaram Nalini, Prasanta Kumar Ghosh
ICASSP2