EDBT 2026 Demo / reviewers in the wild / expert
Shayok Chakraborty
dblp:70/908
· DBLP profile ↗
46ranked-venue papers
16as first author
16since 2021 · last 2026
0000-0001-6378-8286ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 29 · 10 first-author · 12 since 2021Graphics, computer vision, multimedia, augmented reality and games · 27 · 10 first-author · 10 since 2021Databases, data management, data science and information retrieval · 4 · 2 first-author · 2 since 2021Software engineering, systems software and programming languages · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | LLIE-CvT: A Convolutional Vision Transformer for Low Light Image Enhancement
Debanjan Goswami, Bishal Bashyal, Shayok Chakraborty |
ICPR (15) | 3 |
| 2025 | Adapting Vehicle Detectors for Aerial Imagery to Unseen Domains with Weak SupervisionabstractDetecting vehicles in aerial imagery is a critical task with applications in traffic monitoring, urban planning, and defense intelligence. Deep learning methods have provided state-of-the-art (SOTA) results for this application. However, a significant challenge arises when models trained on data from one geographic region fail to generalize effectively to other areas. Variability in factors such as environmental conditions, urban layouts, road networks, vehicle types, and image acquisition parameters (e.g., resolution, lighting, and angle) leads to domain shifts that degrade model performance. This paper proposes a novel method that uses generative AI to synthesize high-quality aerial images and their labels, improving detector training through data augmentation. Our key contribution is the development of a multi-stage, multi-modal knowledge transfer framework utilizing fine-tuned latent diffusion models (LDMs) to mitigate the distribution gap between the source and target environments. Extensive experiments across diverse aerial imagery domains show consistent performance improvements in AP50 over supervised learning on source domain data, weakly supervised adaptation methods, unsupervised domain adaptation methods, and open-set object detectors by 4-23%, 6-10%, 7-40%, and more than 50%, respectively. Furthermore, we introduce two newly annotated aerial datasets from New Zealand and Utah to support further research in this field. Project page is available at: https://humansensinglab.github.io/AGenDA Minhyek Jeon, Shuowen Hu, Zheyang Qin, Shayok Chakraborty, Stanislav Panev, Celso de Melo, Fernando De la Torre |
ICCV | 5 |
| 2025 | Active Learning for Image Segmentation with Binary User FeedbackabstractDeep learning algorithms have depicted commendable performance in a variety of computer vision applications. However, training a robust deep neural network necessitates a large amount of labeled training data, which is timeconsuming and labor-intensive to acquire. This problem is even more serious for an application like image segmentation, as the human oracle has to hand-annotate each and every pixel in a given training image, which is extremely laborious. Active learning algorithms automatically identify the salient and exemplar samples from large amounts of unlabeled data, and tremendously reduce human annotation effort in inducing a machine learning model. In this paper, we propose a novel active learning algorithm for image segmentation, with the goal of further reducing the labeling burden on the human oracles. Our framework identifies a batch of informative images, together with a list of semantic classes for each, and the human annotator merely needs to answer whether a given semantic class is present or absent in a given image. To the best of our knowledge, this is the first research effort to develop an active learning framework for image segmentation, which poses only binary (yes/no) queries to the users. We pose the image and class selection as a constrained optimization problem and derive a linear programming relaxation to select a batch of (image-class) pairs, which are maximally informative to the underlying deep neural network. Our extensive empirical studies on three challenging datasets corroborate the potential of our method in substantially reducing human annotation effort for real-world image segmentation applications. Debanjan Goswami, Shayok Chakraborty |
WACV | 2 |
| 2024 | ACIL: Active Class Incremental Learning for Image Classification
Aditya R. Bhattacharya, Debanjan Goswami, Shayok Chakraborty |
BMVC | 3 |
| 2024 | Multi-source Deep Domain Adaptation for Deepfake Detection
Md Shamim Seraj, Shayok Chakraborty |
ICPR (21) | 2 |
| 2024 | Knowledge Distillation in Deep Networks Under a Constrained Query Budget
Ankita Singh, Shayok Chakraborty |
ICPR (1) | 2 |
| 2024 | Empowering Active Learning for 3D Molecular Graphs with Geometric Graph IsomorphismabstractMolecular learning is pivotal in many real-world applications, such as drug discovery. Supervised learning requires heavy human annotation, which is particularly challenging for molecular data, e.g., the commonly used density functional theory (DFT) is highly computationally expensive. Active learning (AL) automatically queries labels for most informative samples, thereby remarkably alleviating the annotation hurdle. In this paper, we present a principled AL paradigm for molecular learning, where we treat molecules as 3D molecular graphs. Specifically, we propose a new diversity sampling method to eliminate mutual redundancy built on distributions of 3D geometries. We first propose a set of new 3D graph isometries for 3D graph isomorphism analysis. Our method is provably at least as expressive as the Geometric Weisfeiler-Lehman (GWL) test. The moments of the distributions of the associated geometries are then extracted for efficient diversity computing. To ensure our AL paradigm selects samples with maximal uncertainties, we carefully design a Bayesian geometric graph neural network to compute uncertainties specifically for 3D molecular graphs. We pose active sampling as a quadratic programming (QP) problem using the proposed components. Experimental results demonstrate the effectiveness of our AL paradigm, as well as the proposed diversity and uncertainty methods. Ronast Subedi, Wenhan Gao 0002, Shayok Chakraborty, Yi Liu 0059 |
NeurIPS | 4 |
| 2024 | FedAR: Addressing Client Unavailability in Federated Learning with Local Update Approximation and Rectification
Chutian Jiang, Hansong Zhou, Xiaonan Zhang 0001, Shayok Chakraborty |
ECML/PKDD (3) | 4 |
| 2024 | Active Batch Sampling for Multi-label Classification with Binary User FeedbackabstractMulti-label classification is a generalization of multiclass classification, where a single data sample can have multiple labels. While deep neural networks have depicted commendable performance for multi-label learning, they require a large amount of manually annotated training data to attain good generalization capability. However, annotating a multi-label data sample requires a human oracle to consider the presence/absence of every single class individually, which is extremely laborious. Active learning algorithms automatically identify the salient and exemplar instances from large amounts of unlabeled data and are effective in reducing human annotation effort in inducing a machine learning model. In this paper, we propose a novel active learning framework for multi-label learning, which queries a batch of (image-label) pairs and for each pair, poses the question whether the queried label is present in the corresponding image; the human annotators merely need to provide a binary feedback ("yes/no") in response to each query, which involves much less manual work. We pose the image and label selection as a constrained optimization problem and derive a linear programming relaxation to select a batch of (image-label) pairs, which are maximally informative to the underlying deep neural network. Our extensive empirical studies on three challenging datasets corroborate the potential of our method for real-world multi-label classification applications. Debanjan Goswami, Shayok Chakraborty |
WACV | 2 |
| 2024 | D3GU: Multi-target Active Domain Adaptation via Enhancing Domain AlignmentabstractUnsupervised domain adaptation (UDA) for image classification has made remarkable progress in transferring classification knowledge from a labeled source domain to an unlabeled target domain, thanks to effective domain alignment techniques. Recently, in order to further improve performance on a target domain, many Single-Target Active Domain Adaptation (ST-ADA) methods have been proposed to identify and annotate the salient and exemplar target samples. However, it requires one model to be trained and deployed for each target domain and the domain label associated with each test sample. This largely restricts its application in the ubiquitous scenarios with multiple target domains. Therefore, we propose a Multi-Target Active Domain Adaptation (MT-ADA) framework for image classification, named D3GU, to simultaneously align different domains and actively select samples from them for annotation. This is the first research effort in this field to our best knowledge. D3GU applies Decomposed Domain Discrimination (D3) during training to achieve both source-target and target-target domain alignments. Then during active sampling, a Gradient Utility (GU) score is designed to weight every unlabeled target image by its contribution towards classification and domain alignment tasks, and is further combined with KMeans clustering to form GU-KMeans for diverse image sampling. Extensive experiments on three benchmark datasets, Office31, OfficeHome, and DomainNet, have been conducted to validate consistently superior performance of D3GU for MT-ADA1. Lin Zhang 0040, Linghan Xu, Saman Motamed, Shayok Chakraborty, Fernando De la Torre |
WACV | 4 |
| 2023 | Active Learning for Video Classification with Frame Level QueriesabstractDeep learning algorithms have pushed the boundaries of computer vision research and have depicted commendable performance in a variety of applications. However, training a robust deep neural network necessitates a large amount of labeled training data, acquiring which involves significant time and human effort. This problem is even more serious for an application like video classification, where a human annotator has to watch an entire video end-to-end to furnish a label. Active learning algorithms automatically identify the most informative samples from large amounts of unlabeled data; this tremendously reduces the human annotation effort in inducing a machine learning model, as only the few samples that are identified by the algorithm, need to be labeled manually. In this paper, we propose a novel active learning framework for video classification, with the goal of further reducing the labeling onus on the human annotators. Our framework identifies a batch of exemplar videos, together with a set of informative frames for each video; the human annotator needs to merely review the frames and provide a label for each video. This involves much less manual work than watching the complete video to come up with a label. We formulate a criterion based on uncertainty and diversity to identify the informative videos and exploit representative sampling techniques to extract a set of exemplar frames from each video. To the best of our knowledge, this is the first research effort to develop an active learning framework for video classification, where the annotators need to inspect only a few frames to produce a label, rather than watching the end-to-end video. Our extensive empirical analyses corroborate the potential of our method to substantially reduce human annotation effort in applications like video classification, where annotating a single data instance can be extremely tedious. Debanjan Goswami, Shayok Chakraborty |
IJCNN | 2 |
| 2022 | Active Sampling for Text Classification with Subinstance Level QueriesabstractActive learning algorithms are effective in identifying the salient and exemplar samples from large amounts of unlabeled data. This tremendously reduces the human annotation effort in inducing a machine learning model as only a few samples, which are identified by the algorithm, need to be labeled manually. In problem domains like text mining and video classification, human oracles peruse the data instances incrementally to derive an opinion about their class labels (such as reading a movie review progressively to assess its sentiment). In such applications, it is not necessary for the human oracles to review an unlabeled sample end-to-end in order to provide a label; it may be more efficient to identify an optimal subinstance size (percentage of the sample from the start) for each unlabeled sample, and request the human annotator to label the sample by analyzing only the subinstance, instead of the whole data sample. In this paper, we propose a novel framework to address this challenging problem, in an effort to further reduce the labeling burden on the human oracles and utilize the available labeling budget more efficiently. We pose the sample and subinstance size selection as a constrained optimization problem and derive a linear programming relaxation to select a batch of exemplar samples, together with the optimal subinstance size of each, which can potentially augment maximal information to the underlying classification model. Our extensive empirical studies on six challenging datasets from the text mining domain corroborate the practical usefulness of our framework over competing baselines. Shayok Chakraborty, Ankita Singh |
AAAI | 1 |
| 2022 | Deep Active Learning with Range Feedback for Facial Age EstimationabstractDeep learning has achieved unprecedented break-throughs in machine learning research. However, the paucity of labeled training data, together with the human effort and time associated with obtaining a large amount of labeled data, poses significant challenges in training a reliable deep learning model. Active Learning (AL) algorithms automatically select the exemplar instances from large amounts of unlabeled data, and are instrumental in reducing the human annotation effort in inducing a machine learning model. However, in certain applications providing the exact label to queried unlabeled instances may be challenging even for human annotators. Vision based facial age estimation is one such application where it is difficult to estimate the exact age of a person merely from a face image; it maybe easier, and more practical, to provide other forms of annotation such as the best estimated lower and upper bounds on the age of the person within a given span. In this paper, we propose DALRange, a novel deep active learning framework, where annotators merely need to provide an estimated range on the label of an unlabeled sample, rather than the exact label. We formulate a loss function relevant to the research task and exploit the gradient descent algorithm to optimize the loss and train the network. To the best of our knowledge, this is the first research effort to develop an active learning algorithm to train a deep neural network, which poses only range label queries to the oracles. Our extensive empirical studies on human-annotated data corroborate the practical usefulness of our framework in applications where providing the exact labels to queried samples can be challenging. Aditya R. Bhattacharya, Shayok Chakraborty |
IJCNN | 2 |
| 2022 | A Machine-Learning Based Approach for Predicting Older Adults' Adherence to Technology-Based Cognitive Training
Zhe He 0001, Shubo Tian, Ankita Singh, Shayok Chakraborty, Mia Liza A. Lustria, Neil Charness, Nelson A. Roque, Erin Harrell, Walter R. Boot |
Inf. Process. Manag. | 4 |
| 2021 | Deterministic Mini-batch Sequencing for Training Deep Neural NetworksabstractRecent advancements in the field of deep learning have dramatically improved the performance of machine learning models in a variety of applications, including computer vision, text mining, speech processing and fraud detection among others. Mini-batch gradient descent is the standard algorithm to train deep models, where mini-batches of a fixed size are sampled randomly from the training data and passed through the network sequentially. In this paper, we present a novel algorithm to generate a deterministic sequence of mini-batches to train a deep neural network (rather than a random sequence). Our rationale is to select a mini-batch by minimizing the Maximum Mean Discrepancy (MMD) between the already selected mini-batches and the unselected training samples. We pose the mini-batch selection as a constrained optimization problem and derive a linear programming relaxation to determine the sequence of mini-batches. To the best of our knowledge, this is the first research effort that uses the MMD criterion to determine a sequence of mini-batches to train a deep neural network. The proposed mini-batch sequencing strategy is deterministic and independent of the underlying network architecture and prediction task. Our extensive empirical analyses on three challenging datasets corroborate the merit of our framework over competing baselines. We further study the performance of our framework on two other applications besides classification (regression and semantic segmentation) to validate its generalizability. Subhankar Banerjee, Shayok Chakraborty |
AAAI | 2 |
| 2021 | Deep Active Learning with Relative Label Feedback: An Application to Facial Age EstimationabstractDeep learning has emerged as an effective machine learning algorithm to automatically learn a representative set of features and has revolutionized multimedia computing research. However, training a reliable deep model necessitates a large amount of labeled training data, which is time-consuming and labor-intensive to acquire. Active Learning (AL) algorithms address this challenge by automatically identifying the salient and exemplar samples from large amounts of unlabeled data; this drastically reduces human annotation effort, as only a handful of samples, that are identified by the algorithm, need to be labeled manually. However, in applications like vision-based facial age estimation, providing the exact labels (age of a person) may be challenging even for human annotators, as it maybe difficult to accurately estimate the age of a person merely from a facial image; it maybe much easier to provide relative label feedback, such as whether a particular subject is older than another subject. In this paper, we propose a novel deep active learning algorithm (DALRel) which requires only relative label feedback in response to the queried samples. We formulate a loss function relevant to the research task and exploit the gradient descent algorithm to optimize the loss and train the deep network. To the best of our knowledge, this is the first research effort to develop an active learning framework to train a deep neural network, which poses only relative label queries to the labeling oracles. Our extensive empirical studies demonstrate the promise and potential of this method for real-world active learning applications, where providing the exact labels to queried instances can be challenging. Ankita Singh, Shayok Chakraborty |
IJCNN | 2 |
| 2020 | Asking the Right Questions to the Right Users: Active Learning with Imperfect OraclesabstractActive learning algorithms automatically identify the salient and exemplar samples from large amounts of unlabeled data and tremendously reduce human annotation effort in inducing a machine learning model. In a traditional active learning setup, the labeling oracles are assumed to be infallible, that is, they always provide correct answers (in terms of class labels) to the queried unlabeled instances. However, in real-world applications, oracles are often imperfect and provide incorrect label annotations. Oracles also have diverse expertise and while they may be noisy, certain oracles may provide accurate annotations to certain specific instances. In this paper, we propose a novel framework to address the challenging problem of active learning in the presence of multiple imperfect oracles. We pose the optimal sample and oracle selection as a constrained optimization problem and derive a linear programming relaxation to select a batch of (sample-oracle) pairs, which can potentially augment maximal information to the underlying classification model. Our extensive empirical studies on 9 challenging datasets (from a variety of application domains) corroborate the usefulness of our framework over competing baselines. Shayok Chakraborty |
AAAI | 1 |
| 2020 | A Study of Deep Learning for Predicting Freeze of Gait in Patients with Parkinson's DiseaseabstractFreezing of gait (FOG) is a gait impairment, common in patients with advanced Parkinson's disease. Predicting FOG before its onset enables preemptive cueing that can prevent FOG or reduce its intensity and duration. Deep learning models have recently been proposed to predict FOG. Such models have feature learning capabilities and do not require the use of hand-crafted features. However, some intricacies that are specific to this approach have not been carefully studied. In particular, the implication of the lack of accurately labelled pre-FOG data, which can have a significant impact on model development and evaluation, has not been fully understood.In this work, we discuss the challenges in deep learning for predicting FOG, illustrate the impact of the lack of accurate pre-FOG data on model development and evaluation, and present a more reliable evaluation method that is independent of the labelling of pre-FOG data. Using this new evaluation method, we study the deep learning schemes for FOG prediction by performing extensive experiments on a public domain dataset. The main conclusions of the study include the following: 1) even without accurate pre-FOG data, deep learning techniques can achieve very high FOG prediction performance while not introducing significant false alarms; 2) traditional deep learning performance metrics such as accuracy, sensitivity, and specificity may not be indicative of the FOG prediction performance; 3) human gait data have high subject-dependent variability, and it requires different deep learning models to achieve the best performance for different individuals; and finally 4) transfer learning is an effective technique for predicting FOG. To the best of our knowledge, this is the first research effort to derive these conclusions via extensive empirical analysis. Alexander M. Yuan, Shayok Chakraborty |
ICMLA | 2 |
| 2020 | Budgeted Subset Selection for Fine-tuning Deep Learning Architectures in Resource-Constrained ApplicationsabstractThe growing success and popularity of deep learning in computer vision have resulted in the availability of several pretrained deep learning architectures (such as AlexNet, ResNet, VGGNet among others). A common practice in deep learning research is to use one of the pre-trained models and fine-tune it to a given target task, using training data from the target. However, training a deep learning model efficiently necessitates expensive, high-quality GPUs and distributed computing infrastructures. Some applications (such as those running on mobile platforms) are severely limited in terms of memory and computational resources; in these applications, it is a significant challenge to fine-tune a pre-trained deep learning model to a target task, using large amounts of target training data. Cloud services can be leveraged for training, but involve issues with data privacy and cost. In such applications, it is important to select an informative subset of the training data and fine-tune the deep model using only the selected subset. In this paper, we propose a novel framework to address this problem. We pose subset selection as a constrained NP-hard integer quadratic programming problem and derive an efficient linear relaxation to select a subset of exemplar instances. Our extensive empirical studies on three challenging vision datasets (from different application domains) using three commonly used pre-trained deep learning models corroborate the potential of our framework for real-world, resource-constrained applications. Subhankar Banerjee, Shayok Chakraborty |
IJCNN | 2 |
| 2020 | Deep Active Transfer Learning for Image RecognitionabstractIn recent years, deep learning has revolutionized the field of computer vision and has achieved state-of-the-art performance in a variety of applications. However, training a robust deep neural network necessitates a large amount of hand-labeled training data, which is time-consuming and labor-intensive to acquire. Active learning and transfer learning are two popular methodologies to address the problem of learning with limited labeled data. Active learning attempts to select the salient and exemplar instances from large amounts of unlabeled data; transfer learning leverages knowledge from a labeled source domain to develop a model for a (related) target domain, where labeled data is scarce. In this paper, we propose a novel active transfer learning algorithm with the objective of learning informative feature representations from a given dataset using a deep convolutional neural network, under the constraint of weak supervision. We formulate a loss function relevant to the research task and exploit the gradient descent algorithm to optimize the loss and train the deep network. To the best of our knowledge, this is the first research effort to propose a task-specific loss function integrating active and transfer learning, with the goal of learning informative feature representations using a deep neural network, under weak human supervision. Our extensive empirical studies on a variety of challenging, real-world applications depict the merit of our framework over competing baselines. Ankita Singh, Shayok Chakraborty |
IJCNN | 2 |
| 2020 | Active Learning for Multimedia Computing: Survey, Recent Trends and ApplicationsabstractThe widespread emergence and deployment of inexpensive sensors has resulted in the generation of enormous amounts of digital data in today's world. While this has expanded the possibilities of solving real world problems using computational learning frameworks, selecting the salient data samples from such huge collections of data has proved to be a significant and practical challenge. Further, to train a reliable classification model, it is important to have a large quantity of labeled training data. Manual annotation of large amounts of data is an expensive process in terms of time, labor and human expertise. This has set the stage for research in the field of active learning. Active learning algorithms automatically select the salient and exemplar instances from large quantities of unlabeled data and thereby tremendously reduce human annotation effort in training an effective classifier. It can be applied across all existing classification / regression methods and with any kind of data, thus making it a very generalizable approach. The success of active learning in several applications (such as image retrieval, image recognition) has resulted in the extension of the framework to problem settings beyond regular classification / regression. Active learning concepts have been extended to newer problem settings (such as feature selection, video summarization, matrix completion) and have also been combined with other learning paradigms such as deep learning and transfer learning. This tutorial will seek to present a comprehensive overview of active learning with a focus on multimedia computing applications, including historical perspectives, theoretical analysis and novel paradigms. The novelty of this tutorial lies in its focus on the emerging trends, algorithms and applications of active learning. It will aim at introducing concepts and open perspectives that motivate further work in this domain, ranging from fundamentals to applications and systems. Shayok Chakraborty |
ACM Multimedia | 1 |
| 2019 | A Generic Active Learning Framework for Class Imbalance Applications
Aditya R. Bhattacharya, Ji Liu 0002, Shayok Chakraborty |
BMVC | 3 |
| 2019 | Deepsub: A Novel Subset Selection Framework for Training Deep Learning ArchitecturesabstractDeep learning algorithms automatically learn a set of informative features from a given dataset and have depicted commendable performance on a variety of computer vision applications. However, efficient training of deep learning architectures with a large number of hidden layers is mostly dependent on high-end GPUs and distributed computing infrastructures. Some applications (such as applications running on mobile platforms), have limited access to computational resources and face a fundamental challenge in handling large scale training data. Cloud services can be leveraged for training, but involve challenges with data privacy and cost. In such applications, it is crucial to select a subset of informative training samples and use only the extracted subset to induce the deep model. In this paper, we propose a novel subset selection algorithm, DeepSub, to address this practical challenge. Our framework is computationally efficient, easy to implement and also enjoys nice theoretical properties. Our extensive empirical studies on three challenging computer vision applications (face, handwritten digits and object recognition), using three popular deep learning architectures (AlexNet, GoogleNet and ResNet) corroborate the potential of DeepSub over competing baselines. Subhankar Banerjee, Shayok Chakraborty |
ICIP | 2 |
| 2019 | When to Pull Starting Pitchers in Major League Baseball? A Data Mining ApproachabstractOne of the most important decisions made by managers in a baseball game is when to pull the starting pitcher. It has a direct consequence on the outcome of the game and also on the physical fitness of the pitcher. Traditionally, managers rely on various heuristics for this decision. In this paper, we propose a machine learning based approach to determine when to replace the starting pitcher. We curate a large dataset of more than one million samples, spanning more than 10 years of baseball games (2007 - 2017), and study the performance of various classification algorithms on this dataset. We further perform feature analysis to gain insights on the most important features influencing the replacement of starting pitchers. To the best of our knowledge, this is the first research effort to leverage machine learning and data analytics to model managers' decisions of pulling a starting pitcher from historic data. Such a system can be immensely useful in assisting managers make more informed decisions during an ongoing game and has the potential to reduce the risk of baseball related injuries. We hope that our curated dataset and initial research findings will promote further work toward this important problem of deciding when to pull a starting pitcher in an ongoing baseball game. Michael Woodham, Jason Hawkins, Ankita Singh, Shayok Chakraborty |
ICMLA | 4 |
| 2019 | Tracing with Less Data: Active Learning for Classification-Based Traceability Link RecoveryabstractPrevious work has established both the importance and difficulty of establishing and maintaining adequate software traceability. While it has been shown to support essential maintenance and evolution tasks, recovering traceability links between related software artifacts is a time consuming and error prone task. As such, substantial research has been done to reduce this barrier to adoption by at least partially automating traceability link recovery. In particular, recent work has shown that supervised machine learning can be effectively used for automating traceability link recovery, as long as there is sufficient data (i.e., labeled traceability links) to train a classification model. Unfortunately, the amount of data required by these techniques is a serious limitation, given that most software systems rarely have traceability information to begin with. In this paper we address this limitation of previous work and propose an approach based on active learning, which substantially reduces the amount of training data needed by supervised classification approaches for traceability link recovery while maintaining similar performance. Chris Mills, Javier Escobar-Avila, Aditya R. Bhattacharya, Grigoriy Kondyukov, Shayok Chakraborty, Sonia Haiduc |
ICSME | 5 |
| 2019 | Active Learning with n-ary Queries for Image RecognitionabstractActive learning algorithms automatically identify the salient and informative samples from large amounts of unlabeled data and tremendously reduce human annotation effort in inducing a machine learning model. In a multi-class classification problem, however, the human oracle has to provide the precise category label of each unlabeled sample to be annotated. In an application with a significantly large (and possibly unknown) number of classes (such as object recognition), providing the exact class label may be time consuming and error prone. In this paper, we propose a novel active learning framework where the annotator merely needs to identify which of the selected n categories a given unlabeled sample belongs to (where n is much smaller than the actual number of classes). We pose the active sample selection as an NP-hard integer quadratic programming problem and exploit the Iterative Truncated Power algorithm to derive an efficient solution. To the best of our knowledge, this is the first research effort to propose a generic n-ary query framework for active sample selection. Our extensive empirical results on six challenging vision datasets (from four different application domains and varied number of classes ranging from 10 to 369) corroborate the potential of the framework in further reducing human annotation effort in real-world active learning applications. Aditya R. Bhattacharya, Shayok Chakraborty |
WACV | 2 |
| 2018 | Multi-Label Deep Active Learning with Label CorrelationabstractAnnotating a data sample in a multi-label learning problem requires a human oracle to consider the presence/absence of every possible label separately, which is extremely labor intensive. Active learning algorithms automatically identify the informative samples from large amounts of unlabeled data and significantly reduce human annotation efforts in inducing a classification model. Further, deep models have gained popularity to automatically learn representative features from a given dataset and have depicted promising empirical performance in a variety of applications. In this paper, we exploit the feature learning capabilities of deep neural networks and propose a novel framework to address the problem of multi-label active learning with label correlation. We integrate an active selection criterion to the objective function and train deep networks to optimize the function. Our extensive empirical studies on five benchmark multi-label datasets show that our methods outperform the state-of-the-art algorithms, corroborating their potential for real-world image classification applications. Hiranmayi Ranganathan, Hemanth Venkateswara, Shayok Chakraborty, Sethuraman Panchanathan |
ICIP | 3 |
| 2018 | Deep Domain Adaptation to Predict Freezing of Gait in Patients with Parkinson's DiseaseabstractFreezing of gait (FoG) is a common gait impairment in patients with advanced Parkinson's disease (PD), which manifests as sudden difficulties in starting or continuing locomotion. FoG often results in falls and negatively affect a patient's quality of life. Real-time detection algorithms have been developed, which detect FoG events using signals derived from wearable sensors. However, predicting FoG before it actually occurs opens the possibility of preemptive cueing, which can potentially avoid (or reduce the intensity and duration of) the episodes. Moreover, human gait involves significant subject-based variability and a machine learning model trained on a particular patient's data may not generalize well to other patients. In this paper, we study the performance of advanced deep learning algorithms to predict FoG events in short time durations before their occurrence. We further study the performance of domain adaptation (or transfer learning) algorithms to address the domain disparity between data from different subjects, in order to develop a better prediction model for a particular subject. To the best of our knowledge, this is the first research effort to study domain adaptation algorithms to predict FoG episodes in patients with PD. Our extensive empirical studies on a publicly available dataset (collected from 10 PD patients) demonstrate the potential of our algorithms to accurately identify FoG events before their onset. We believe this research will serve as a stepping stone toward the development of more advanced FoG prediction algorithms for patients with PD. Vishwas G. Torvi, Aditya R. Bhattacharya, Shayok Chakraborty |
ICMLA | 3 |
| 2018 | Distributed Active Learning for Image RecognitionabstractDeveloping intelligent learning algorithms under the constraint of limited manual labor is a fundamental research challenge and a problem of immense practical importance. Active learning algorithms alleviate this problem by automatically identifying the salient and exemplar instances from large amounts of unlabeled data; this tremendously reduces the human annotation effort as only a small subset of the samples, identi ed by the algorithm, needs to be labeled manually. Further, the unprecedented increase in the volume of digital data has necessitated the usage of multiple, independent computers for its storage and processing, in a given application. The need of the hour is therefore an active learning framework which can operate in a distributed setup, where the unlabeled data is partitioned across multiple computers. In this paper, we propose a novel algorithm to address this important challenge. Our algorithm requires minimal communication among the computers (over which the data is stored) and also enjoys nice theoretical properties. Our extensive empirical studies on a variety of challenging, real-world vision datasets, from different application domains, corroborate the potential of the proposed framework. Shayok Chakraborty |
WACV | 1 |
| 2017 | Deep Hashing Network for Unsupervised Domain AdaptationabstractIn recent years, deep neural networks have emerged as a dominant machine learning tool for a wide variety of application domains. However, training a deep neural network requires a large amount of labeled data, which is an expensive process in terms of time, labor and human expertise. Domain adaptation or transfer learning algorithms address this challenge by leveraging labeled data in a different, but related source domain, to develop a model for the target domain. Further, the explosive growth of digital data has posed a fundamental challenge concerning its storage and retrieval. Due to its storage and retrieval efficiency, recent years have witnessed a wide application of hashing in a variety of computer vision applications. In this paper, we first introduce a new dataset, Office-Home, to evaluate domain adaptation algorithms. The dataset contains images of a variety of everyday objects from multiple domains. We then propose a novel deep learning framework that can exploit labeled source data and unlabeled target data to learn informative hash codes, to accurately classify unseen target data. To the best of our knowledge, this is the first research effort to exploit the feature learning capabilities of deep neural networks to learn representative hash codes to address the domain adaptation problem. Our extensive empirical studies on multiple transfer tasks corroborate the usefulness of the framework in learning efficient hash codes which outperform existing competitive baselines for unsupervised domain adaptation. Hemanth Venkateswara, Jose Eusebio, Shayok Chakraborty, Sethuraman Panchanathan |
CVPR | 3 |
| 2017 | Deep active learning for image classificationabstractIn the recent years, deep learning algorithms have achieved state-of-the-art performance in a variety of computer vision applications. In this paper, we propose a novel active learning framework to select the most informative unlabeled samples to train a deep belief network model. We introduce a loss function specific to the active learning task and train the model to minimize the loss function. To the best of our knowledge, this is the first research effort to integrate an active learning based criterion in the loss function used to train a deep belief network. Our extensive empirical studies on a wide variety of uni-modal and multi-modal vision datasets corroborate the potential of the method for real-world image recognition applications. Hiranmayi Ranganathan, Hemanth Venkateswara, Shayok Chakraborty, Sethuraman Panchanathan |
ICIP | 3 |
| 2016 | Smart Stadium for Smarter Living: Enriching the Fan ExperienceabstractRapid urbanization has led to more people residing in cities than ever before, and projections estimate that 64% of the global population will be urban by 2050. Cities are beginning to explore Smart City initiatives to reduce expenses and complexities while increasing efficiency and quality of life for its citizens. To achieve this goal, advances in technology and policies are needed together with rethinking traditional solutions to transportation, safety, sustainability, among other priority areas. We propose the use of a Smart Stadium as a 'living laboratory' to identify, deploy and test Internet of Things technologies and Smart City solutions in an environment small enough to practically trial but large enough to evaluate effectiveness and scalability. The Smart Stadium for Smarter Living initiative brings together Arizona State University, Dublin City University, Intel Corporation, Gaelic Athletic Association, Sun Devil Stadium and Croke Park to explore smart environment solutions. Sethuraman Panchanathan, Shayok Chakraborty, Troy McDaniel, Matt Bunch, Noel E. O'Connor, Suzanne Little, Kevin McGuinness, Mark Marsden |
ISM | 2 |
| 2016 | Multimodal emotion recognition using deep learning architecturesabstractEmotion analysis and recognition has become an interesting topic of research among the computer vision research community. In this paper, we first present the emoF-BVP database of multimodal (face, body gesture, voice and physiological signals) recordings of actors enacting various expressions of emotions. The database consists of audio and video sequences of actors displaying three different intensities of expressions of 23 different emotions along with facial feature tracking, skeletal tracking and the corresponding physiological data. Next, we describe four deep belief network (DBN) models and show that these models generate robust multimodal features for emotion classification in an unsupervised manner. Our experimental results show that the DBN models perform better than the state of the art methods for emotion recognition. Finally, we propose convolutional deep belief network (CDBN) models that learn salient multimodal features of expressions of emotions. Our CDBN models give better recognition accuracies when recognizing low intensity or subtle expressions of emotions when compared to state of the art methods. Hiranmayi Ranganathan, Shayok Chakraborty, Sethuraman Panchanathan |
WACV | 2 |
| 2015 | BatchRank: A Novel Batch Mode Active Learning Framework for Hierarchical ClassificationabstractActive learning algorithms automatically identify the salient and exemplar instances from large amounts of unlabeled data and thus reduce human annotation effort in inducing a classification model. More recently, Batch Mode Active Learning (BMAL) techniques have been proposed, where a batch of data samples is selected simultaneously from an unlabeled set. Most active learning algorithms assume a flat label space, that is, they consider the class labels to be independent. However, in many applications, the set of class labels are organized in a hierarchical tree structure, with the leaf nodes as outputs and the internal nodes as clusters of outputs at multiple levels of granularity. In this paper, we propose a novel BMAL algorithm (BatchRank) for hierarchical classification. The sample selection is posed as an NP-hard integer quadratic programming problem and a convex relaxation (based on linear programming) is derived, whose solution is further improved by an iterative truncated power method. Finally, a deterministic bound is established on the quality of the solution. Our empirical results on several challenging, real-world datasets from multiple domains, corroborate the potential of the proposed framework for real-world hierarchical classification applications. Shayok Chakraborty, Vineeth N. Balasubramanian, Adepu Ravi Sankar, Sethuraman Panchanathan, Jieping Ye |
KDD | 1 |
| 2015 | Towards Distributed Video SummarizationabstractVideo summarization is a fertile topic in multimedia research. While the advent of modern video cameras and several social networking and video sharing websites (like YouTube, Flickr, Facebook) has led to the generation of humongous amounts of redundant video data, video summarization has emerged as an effective methodology to automatically extract a succinct and condensed representation of a given video. The unprecedented increase in the volume of video data necessitates the usage of multiple, independent computers for its storage and processing. In order to understand the overall essence of a video, it is therefore necessary to develop an algorithm which can summarize a video distributed across multiple computers. In this paper, we propose a novel algorithm for distributed video summarization. Our algorithm requires minimal communication among the computers (over which the video is stored) and also enjoys nice theoretical properties. Our empirical results on several challenging, unconstrained videos corroborate the potential of the proposed framework for real-world distributed video summarization applications. Shayok Chakraborty, Omesh Tickoo, Ravi R. Iyer 0001 |
ACM Multimedia | 1 |
| 2015 | Adaptive Keyframe Selection for Video SummarizationabstractThe explosive growth of video data in the modern era has set the stage for research in the field of video summarization, which attempts to abstract the salient frames in a video in order to provide an easily interpreted synopsis. Existing work on video summarization has primarily been static - that is, the algorithms require the summary length to be specified as an input parameter. However, video streams are inherently dynamic in nature, while some of them are relatively simple in terms of visual content, others are much more complex due to camera/object motion, changing illumination, cluttered scenes and low quality. This necessitates the development of adaptive summarization techniques, which adapt to the complexity of a video and generate a summary accordingly. In this paper, we propose a novel algorithm to address this problem. We pose the summary selection as an optimization problem and derive an efficient technique to solve the summary length and the specific frames to be selected, through a single formulation. Our extensive empirical studies on a wide range of challenging, unconstrained videos demonstrate tremendous promise in using this method for real-world video summarization applications. Shayok Chakraborty, Omesh Tickoo, Ravi R. Iyer 0001 |
WACV | 1 |
| 2015 | Active Batch Selection via Convex Relaxations with Guaranteed Solution BoundsabstractActive learning techniques have gained popularity to reduce human effort in labeling data instances for inducing a classifier. When faced with large amounts of unlabeled data, such algorithms automatically identify the exemplar instances for manual annotation. More recently, there have been attempts towards a batch mode form of active learning, where a batch of data points is simultaneously selected from an unlabeled set. In this paper, we propose two novel batch mode active learning (BMAL) algorithms: BatchRank and BatchRand. We first formulate the batch selection task as an NP-hard optimization problem; we then propose two convex relaxations, one based on linear programming and the other based on semi-definite programming to solve the batch selection problem. Finally, a deterministic bound is derived on the solution quality for the first relaxation and a probabilistic bound for the second. To the best of our knowledge, this is the first research effort to derive mathematical guarantees on the solution quality of the BMAL problem. Our extensive empirical studies on 15 binary, multi-class and multi-label challenging datasets corroborate that the proposed algorithms perform at par with the state-of-the-art techniques, deliver high quality solutions and are robust to real-world issues like label noise and class imbalance. Shayok Chakraborty, Vineeth N. Balasubramanian, Qian Sun 0002, Sethuraman Panchanathan, Jieping Ye |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2015 | Adaptive Batch Mode Active LearningabstractActive learning techniques have gained popularity to reduce human effort in labeling data instances for inducing a classifier. When faced with large amounts of unlabeled data, such algorithms automatically identify the exemplar and representative instances to be selected for manual annotation. More recently, there have been attempts toward a batch mode form of active learning, where a batch of data points is simultaneously selected from an unlabeled set. Real-world applications require adaptive approaches for batch selection in active learning, depending on the complexity of the data stream in question. However, the existing work in this field has primarily focused on static or heuristic batch size selection. In this paper, we propose two novel optimization-based frameworks for adaptive batch mode active learning (BMAL), where the batch size as well as the selection criteria are combined in a single formulation. We exploit gradient-descent-based optimization strategies as well as properties of submodular functions to derive the adaptive BMAL algorithms. The solution procedures have the same computational complexity as existing state-of-the-art static BMAL techniques. Our empirical results on the widely used VidTIMIT and the mobile biometric (MOBIO) data sets portray the efficacy of the proposed frameworks and also certify the potential of these approaches in being used for real-world biometric recognition applications. Shayok Chakraborty, Vineeth N. Balasubramanian, Sethuraman Panchanathan |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2013 | Active Matrix CompletionabstractRecovering a matrix from a sampling of its entries is a problem of rapidly growing interest and has been studied under the name of matrix completion. It occurs in many areas of engineering and applied science. In most machine learning and data mining applications, it is possible to leverage the expertise of human oracles to improve the performance of the system. It is therefore natural to extend this idea of "human-in-the-loop" to the matrix completion problem. However, considering the enormity of data in the modern era, manually completing all the entries in a matrix will be an expensive process in terms of time, labor and human expertise, human oracles can only provide selective supervision to guide the solution process. Thus, appropriately identifying a subset of missing entries (for manual annotation) in an incomplete matrix is of paramount practical importance, this can potentially lead to better reconstructions of the incomplete matrix with minimal human effort. In this paper, we propose novel algorithms to address this issue. Since the query locations are actively selected by the algorithms, we refer to these methods as active matrix completion algorithms. The proposed techniques are generic and the same frameworks can be used in a wide variety of applications including recommendation systems, transductive / multi-label active learning, active learning in regression and active feature acquisition among others. Our extensive empirical analysis on several challenging real-world datasets certify the merit and versatility of the proposed frameworks in efficiently exploiting human intelligence in data mining / machine learning applications. Shayok Chakraborty, Vineeth N. Balasubramanian, Sethuraman Panchanathan, Ian Davidson, Jieping Ye |
ICDM | 1 |
| 2013 | Generalized batch mode active learning for face-based biometric recognition
Shayok Chakraborty, Vineeth N. Balasubramanian, Sethuraman Panchanathan |
Pattern Recognit. | 1 |
| 2012 | Batch Mode Active Learning for Multimedia Pattern RecognitionabstractMultimedia applications like face recognition and facial expression recognition inherently rely on the availability of a large amount of labeled data to train a robust recognition system. In order to induce a reliable classification model for a multimedia pattern recognition application, the data is typically labeled by human experts based on some domain knowledge. However, manual annotation of a large number of images is an expensive process in terms of time, labor and human expertise. This has led to the development of active learning algorithms, which automatically identify the salient instances from a given set of unlabeled data and are effective in reducing the human annotation effort to train a classification model. Further, to address the possible presence of multiple labeling oracles, there have been efforts towards a batch form of active learning, where a set of unlabeled images are selected simultaneously for labeling instead of a single image at a time. Existing algorithms on batch mode active learning concentrate only on the development of a batch selection criterion and assume that the batch size (number of samples to be queried from an unlabeled set) to be specified in advance. However, in multimedia applications like face/facial expression recognition, it is difficult to decide on a batch size in advance because of the dynamic nature of video streams. Further, multimedia applications like facial expression recognition involve a fuzzy label space because of the imprecision and the vagueness in the class label boundaries. This necessitates a BMAL framework, for fuzzy label problems. To address these fundamental challenges, we propose two novel BMAL techniques in this work: (i) a framework for dynamic batch mode active learning, which adaptively selects the batch size and the specific instances to be queried based on the complexity of the data stream being analyzed and (ii) a BMAL algorithm for fuzzy label classification problems. To the best of our knowledge, this is the first attempt to develop such techniques in the active learning literature. Shayok Chakraborty, Vineeth N. Balasubramanian, Sethuraman Panchanathan |
ISM | 1 |
| 2011 | Dynamic Batch Mode Active Learning via L1 RegularizationabstractWe propose a method for dynamic batch mode active learning where the batch size and selection criteria are integrated into a single formulation. Shayok Chakraborty, Vineeth N. Balasubramanian, Sethuraman Panchanathan |
AAAI | 1 |
| 2011 | Dynamic batch mode active learningabstractActive learning techniques have gained popularity to reduce human effort in labeling data instances for inducing a classifier. When faced with large amounts of unlabeled data, such algorithms automatically identify the exemplar and representative instances to be selected for manual annotation. More recently, there have been attempts towards a batch mode form of active learning, where a batch of data points is simultaneously selected from an unlabeled set. Real-world applications require adaptive approaches for batch selection in active learning. However, existing work in this field has primarily been heuristic and static. In this work, we propose a novel optimization-based framework for dynamic batch mode active learning, where the batch size as well as the selection criteria are combined in a single formulation. The solution procedure has the same computational complexity as existing state-of-the-art static batch mode active learning techniques. Our results on four challenging biometric datasets portray the efficacy of the proposed framework and also certify the potential of this approach in being used for real world biometric recognition applications. Shayok Chakraborty, Vineeth N. Balasubramanian, Sethuraman Panchanathan |
CVPR | 1 |
| 2011 | Optimal batch selection for active learning in multi-label classificationabstractMulti-label classification is a generalization of conventional classification, where it is possible for a single data point to have multiple labels. Manual annotation of a multi-label data point requires a human oracle to consider the presence/absence of every possible class separately, which involves significant labor. Active learning techniques are effective in reducing human labeling effort to induce a classification model. When exposed to large quantities of unlabeled data, such algorithms automatically select the salient and representative instances for manual annotation. Further, to address the high redundancy in data such as image or video sequences as well as the availability of multiple labeling agents, there have been recent attempts towards a batch mode form of active learning, where a batch of data points is selected simultaneously from an unlabeled set. In this work, we propose a novel optimization based batch mode active learning strategy to minimize human labeling effort in multi-label classification problems. To the best of our knowledge, this is the first attempt to develop such a scheme primarily intended for the multi-label context. The proposed framework is computationally simple, easy to implement and can be suitably modified to perform batch mode active learning in other formulations, such as single-label classification or problems involving hierarchical label spaces. Our results corroborate the efficacy of the proposed algorithm and certify the potential of the framework in being used for real world applications. Shayok Chakraborty, Vineeth N. Balasubramanian, Sethuraman Panchanathan |
ACM Multimedia | 1 |
| 2010 | Kernel Learning for Efficiency Maximization in the Conformal Predictions FrameworkabstractThe Conformal Predictions framework is a recent development in machine learning to associate reliable measures of confidence with results in classification and regression. This framework is founded on the principles of algorithmic randomness (Kolmogorov complexity), transductive inference and hypothesis testing. While the formulation of the framework guarantees validity, the efficiency of the framework depends greatly on the choice of the classifier and appropriate kernel functions or parameters. While this framework has extensive potential to be useful in several applications, the lack of efficiency can limit its usability. In this paper, we propose a novel kernel learning methodology to maximize efficiency in the CP framework. This method is validated using the k-Nearest Neighbors classifier on three different datasets, and our results show immense promise in applying this method to obtain efficient conformal predictors that can be practically useful. Vineeth N. Balasubramanian, Shayok Chakraborty, Sethuraman Panchanathan, Jieping Ye |
ICMLA | 2 |
| 2010 | Dynamic Batch Size Selection for Batch Mode Active Learning in BiometricsabstractRobust biometric recognition is of paramount importance in security and surveillance applications. In face based biometric systems, data is usually collected using a video camera with high frame rate and thus the captured data has high redundancy. Selecting the appropriate instances from this data to update a classification model, is a significant, yet valuable challenge. Active learning methods have gained popularity in identifying the salient and exemplar data instances from superfluous sets. Batch mode active learning schemes attempt to select a batch of samples simultaneously rather than updating the model after selecting every single data point. Existing work on batch mode active learning assume a fixed batch size, which is not a practical assumption in biometric recognition applications. In this paper, we propose a novel framework to dynamically select the batch size using clustering based unsupervised learning techniques. We also present a batch mode active learning strategy specially suited to handle the high redundancy in biometric datasets. The results obtained on the challenging VidTIMIT and MOBIO datasets corroborate the superiority of dynamic batch size selection over static batch size and also certify the potential of the proposed active learning scheme in being used for real world biometric recognition applications. Shayok Chakraborty, Vineeth N. Balasubramanian, Sethuraman Panchanathan |
ICMLA | 1 |