EDBT 2026 Demo / reviewers in the wild / expert
Arghya Pal
dblp:210/9831
· DBLP profile ↗
21ranked-venue papers
6as first author
17since 2021 · last 2026
0000-0003-4918-2565ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 13 · 5 first-author · 10 since 2021Artificial intelligence and machine learning · 8 · 4 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Systems, architecture and hardware · 1Computer networks · 1 · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FLAG-4D: Flow-Guided Local-Global Dual-Deformation Model for 4D ReconstructionabstractWe introduce FLAG-4D, a novel framework for generating novel views of dynamic scenes by reconstructing how 3D Gaussian primitives evolve through space and time. Existing methods typically rely on a single Multilayer Perceptron(MLP) to model temporal deformations, and they often struggle to capture complex point motions and fine-grained dynamic details consistently over time, especially from sparse input views. Our approach, FLAG-4D overcomes this by employing a dual-deformation network that dynamically warps a canonical set of 3D Gaussians over time into new positions and anisotropic shapes. This dual-deformation network consists of an Instantaneous Deformation Network (IDN) for modeling fine-grained, local deformations, and Global Motion Network (GMN) for capturing long-range dynamics, refined via mutual learning. To ensure these deformations are both accurate and temporally smooth, FLAG-4D incorporates dense motion features from a pretrained optical flow backbone. We fuse these motion cues from adjacent timeframes and use a deformation-guided attention mechanism to align this flow information with the current state of each evolving 3D Gaussian. Extensive experiments demonstrate that FLAG-4D achieves higher-fidelity and more temporally coherent reconstructions with finer detail preservation than state-of-the-art methods. Guan Yuan Tan, Ngoc Tuan Vu, Arghya Pal, Sailaja Rajanala, Raphael C.-W. Phan, Mettu Srinivas, Chee-Ming Ting |
AAAI | 3 |
| 2026 | Causal-Ex: Causal graph-based micro and macro expression spottingabstractDetecting concealed emotions within apparently normal expressions is crucial for identifying potential mental health issues and facilitating timely support and intervention. The task of spotting macro- and micro-expressions involves predicting the emotional timeline within a video by identifying the onset (i.e., the beginning), apex (the peak of emotion), and offset (the end of emotion) frames of the displayed emotions. More particularly, closely monitoring the key emotion-conveying regions of the face; namely, the foundational muscle-movement cues known as facial action units (AUs)–greatly aids in the clear identification of micro-expressions. One major roadblock is the inadvertent introduction of biases into the training process, which degrades performance regardless of feature quality. Biases are spurious factors that falsely inflate or deflate performance metrics. For instance, the neural networks tend to falsely attribute certain AUs in specific facial regions to particular emotion classes, a phenomenon also termed as Inductive biases. To remove these false attributions, we must identify and mitigate biases that arise from mere correlation between some features and the output class labels. We hence introduce action-unit causal graphs. Unlike the traditional action-unit graph, which connects AUs based solely on spatial adjacency, the causal AU graph is derived from statistical tests and retains edges between AUs only when there is significant evidence that one AU causally influences another. Our model, named Causal-Ex ( Causal -based Ex pression spotting), employs a fast causal inference algorithm to construct a causal graph of facial region of interests (ROIs). This enables us to select causally relevant facial action units in the ROIs. Our work demonstrates improvement in overall F1-scores compared to state-of-the-art approaches with 0.388 on CAS(ME) 2 and 0.3701 on SAMM-Long Video datasets. Our code can be found at: https://github.com/noobasuna/causal_ex.git . Pei-Sze Tan, Sailaja Rajanala, Arghya Pal, Raphael C.-W. Phan, Huey Fang Ong |
Pattern Recognit. Lett. | 3 |
| 2026 | A Unified Framework for Sparse Reconstruction via Preconditioning and Nonconvex RegularizationabstractCompressed Sensing (CS) is an effective technique to recover sparse signals with fewer samples than what is required by the classical Shannon Nyquist sampling theorem. The sensing matrix, sparsifying transform, and sparse recovery algorithm are three key factors for accurate reconstruction in CS. Traditional CS uses a convex $l_{1}$-norm sparse regularizer which may lead to biased estimates and is suboptimal in promoting sparsity. Another challenge is the design of incoherent sensing matrices which is crucial for accurate sparse recovery. In this paper, we propose a novel CS framework combining a preconditioned sensing matrix and nonconvex regularization for improved sparse signal recovery. First, we formulate an optimization problem to find an incoherent sensing matrix via a preconditioner. It allows for a direct computation of the optimal preconditioner and preconditioned sensing matrix, simultaneously. Secondly, we consider a generalized CS model for signal recovery based on the incoherent sensing matrix and a nonconvex $\ell _{1/2}$-norm regularizer. We then derive an Alternating Direction Method of Multipliers (ADMM) algorithm to solve this nonconvex optimization problem. The proposed model is applied to sparse-view Computed Tomography (CT) reconstruction with highly-undersampled and noisy data. Qualitative and quantitative results show significantly better image reconstruction using the preconditioned sensing matrix and $\ell _{1/2}$ regularizer, compared to methods without preconditioning and using the $\ell _{1}$ regularizer. Prasad Theeda, Fuad Noman, Arghya Pal, Raphael C.-W. Phan, Hernando C. Ombao, Chee-Ming Ting |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | GENIE: Socially Unbiased Generative Text-to-Image EditingabstractGenerative diffusion models often exhibit societal biases in sensitive personal attributes such as age, gender, and race. In this work, we describe GENIE – a method to reduce such biases in a variety of classifier-free diffusion models used for image editing. Our method implicitly incorporates debiasing terms together with the user’s explicit edit instruction to reduce bias. This automatic method relieves the user from needing to modify edit instructions in order to avoid bias. Further, no additional training is needed. Experimental results are provided based on modifications to four diffusion models, namely InstructPix2Pix, Stable Diffusion 1.5, Stable Diffusion 2.1, and Stable Diffusion XL. We show that, on average, bias is reduced by 31% in gender, 15% in age, 39% in race. Julia Kaiwen Lau, Raphael C.-W. Phan, Sailaja Rajanala, Ingemar J. Cox, Arghya Pal |
ICASSP | 5 |
| 2025 | Post-Hoc Adversarial Stickers Against Micro-Expression LeakageabstractSecuring micro-expressions against leakage is crucial for privacy, as these subtle facial movements convey genuine emotions and are inherently personal. This study aims to protect micro-expression data from potential adversarial attacks, ensuring the preservation of individuals’ privacy and preventing unauthorized access or misuse of sensitive emotional information. Unlike traditional methods, which often require training and extensive access to models, this research introduces a novel post-hoc method that does not require additional training. We focus on physical adversarial attacks in micro-expression recognition, involving intentional manipulation of visual cues to deceive recognition systems and protect individual emotional privacy. Our approach leverages a causal discovery algorithm to identify causal relationships between facial parts, enabling rapid identification of the optimal locations for adversarial patches in frames with triggered micro-expressions. This method exhibits a more consistent attack success rate than randomly placed adversarial stickers, demonstrating effective generalization across different emotions, stickers, and models. Particularly relevant in scenarios with restricted access to the model, our technique requires only a single interaction during the attack process, highlighting its efficiency and minimal need for querying the target model. The proposed method effectively balances privacy protection with high generalization capability, setting a new standard for defending against adversarial threats in micro-expression recognition. The code is available at https://github.com/noobasuna/au-sticker. Pei-Sze Tan, Sailaja Rajanala, Yee-Fan Tan, Arghya Pal, Chun-Ling Tan, Raphael C.-W. Phan, Huey Fang Ong |
ICASSP | 4 |
| 2025 | Polyfit generative model: can a group of lower-order polynomials generate high resolution diverse images?abstractImplicit neural representations (INRs) have recently gained popularity as a means to model images as continuous functions of spatial coordinates, synthesizing each pixel independently and yielding impressive results in tasks such as scene reconstruction and image generation. A notable advancement, PolyINR, utilizes element-wise multiplications between features and affine-transformed coordinates to achieve higher-order polynomial functions, eliminating the need for positional encodings. However, the finite encoding capacity of INRs, coupled with PolyINR’s recursive polynomial estimation, necessitates substantial training parameters, resulting in high computational costs and limiting applicability across diverse computer vision domains. In this work, we address these challenges by representing images as grids of smaller patches, within which we fit low-degree polynomials to capture local intensity variations. Our approach substantially reduces parameter requirements and computational demands. We evaluate our model qualitatively and quantitatively on large-scale datasets, ImageNet, CelebA, LSUN Bedroom, and Flower102; thus demonstrating competitive performance with state-of-the-art generative models, despite the absence of convolutional, normalization, or self-attention layers. Arghya Pal, Ai-Fang Chai, Sailaja Rajanala, Raphael C.-W. Phan, Koksheik Wong, Chee-Ming Ting |
IJCNN | 1 |
| 2025 | Res-SH: Unbiased Residual Learning for Self-Healing Interface Toughness Prediction with Limited DataabstractThe development of self-healing materials is often hindered by the high costs and material waste associated with traditional characterization methods. Current approaches to toughness prediction, primarily based on convolutional neural networks (CNNs), are limited by their tendency to capture only surface-level features, which can lead to biased predictions. Moreover, working with small datasets, which is common in materials science, further increases the risk of biased training due to overfitting, posing a critical challenge to the reliability and generalizability of predictive models. This study introduces an unbiased residual learning framework designed explicitly for predicting self-healing interface toughness under limiteddata conditions. Our approach, ResNet-inspired approach for predicting self-healing material toughness, named Res-SH, used the power of residual networks to capture deeper, more complex patterns in the data, thereby addressing critical challenges in materials research. Res-SH minimises resource consumption and experimental overhead by focusing on unbiased learning, achieving accurate predictions with fewer training epochs and lower R2score and root mean square prediction errors compared to conventional CNN and lightweight model MobileNetv2. This novel framework provides a cost-effective and resource-efficient alternative to traditional material characterization methods, reducing material waste and accelerating the discovery and optimization of self-healing material systems. Pei-Sze Tan, Karen Jia-Jun Koh, Sailaja Rajanala, Arghya Pal, Raphael C.-W. Phan, Nan Ze, Fuad Noman, Chee-Ming Ting, Norfadilah Dolmat, Nik Nur Wahidah Nik Hashim, Afidalina Tumian |
TENCON | 4 |
| 2024 | Causally Uncovering Bias in Video Micro-Expression RecognitionabstractDetecting microexpressions presents formidable challenges, primarily due to their fleeting nature and the limited diversity in existing datasets. Our studies find that these datasets exhibit a pronounced bias towards specific ethnicities and suffer from significant imbalances in terms of both class and gender representation among the samples. These disparities create fertile ground for various biases to permeate deep learning models, leading to skewed results and inadequate portrayal of specific demographic groups. Our research is driven by a compelling need to identify and rectify these biases within model architectures. To achieve this, we commence by constructing a causal graph that elucidates the intricate relationships between the model, input features, and training outcomes. This graphical representation forms the foundation for our analytical framework. Leveraging this causal framework, we conduct comprehensive case studies, employing counterfactuals as a diagnostic tool to unveil biases arising from dataset-induced class imbalances, gender inequalities, and variations in facial action units. Our final step involves a highly efficient counterfactual debiasing process, eliminating the necessity for additional data collection or model retraining. Our results showcase superior performance compared to state-of-the-art methods across the CASME II, SAMM, and SMIC datasets. Pei-Sze Tan, Sailaja Rajanala, Arghya Pal, Shu-Min Leong, Raphael C.-W. Phan, Huey Fang Ong |
ICASSP | 3 |
| 2024 | A Preconditioning Approach To Optimizing Sensing Matrix For Improved Compressed Sensing CT ReconstructionabstractCompressed sensing (CS) exploiting inherent sparsity prior of signals has been proven effective for sparse-view computed tomography (CT) image reconstruction from undersampled projection data. However, most CS-based CT studies focused on formulating different sparsity regularizers, e.g., total variation (TV) minimization, and neglect design of an incoherent sensing matrix - a key factor of CS performance. The sensing matrix formed by an incomplete set of Radon projections in CT typically exhibits large coherence. In this paper, we propose a novel method for optimizing the sensing matrix via preconditioning to improve CS-CT reconstruction. A well-conditioned preconditioner is designed to optimally reduce the coherence of the sensing matrix and thus improving the CS systems. The desired preconditioner is obtained by solving a nonconvex optimization problem via gradient descent method. The preconditioned systems solved by TV-based sparse recovery algorithms can provide better reconstruction accuracy with fewer measurements even in noisy settings. Evaluated on brain and COVID-19 chest CT datasets, the proposed method when used for preconditioning of Radon sensing matrix reconstructed images with substantially higher quality with faster speed than baselines without preconditioning. Prasad Theeda, Chee-Ming Ting, Arghya Pal, Hernando C. Ombao |
ICIP | 3 |
| 2024 | Distingusic: Distinguishing Synthesized Music from HumanabstractIn this paper we focus on a problem that is increasingly plaguing the music industry; to a large extent due to the proliferation of generative AI models that enable the generation of new realistic and indistinguishable content for diverse modalities: text, image, audio, video. We address this problem from the perspective of audio watermarking; to our best knowledge, this is the first-known watermarking based approach to solve the problem of distinguishing realistic songs synthesized from generative AI models from real songs sung by humans. In more detail, our approach specifically utilizes the SHA-256 hash function, Singular Value Decomposition (SVD) and Discrete Wavelet Transform (DWT) for robust audio watermarking of synthesized songs. Before embedding, the audio is subjected to an attack phase to pinpoint less vulnerable regions for QR watermark placement. During the embedding process, the audio chunks first undergo a 1-level Discrete Wavelet Transform (DWT), and then the resulting approximate coefficients go through Singular Value Decompo-sition (SVD). Additionally, the watermarked array is subjected to SHA-256 hashing for collision-resistant conciseness, which is subsequently embedded into the singular values of the audio. Experimental findings demonstrate the superiority of our method over existing audio watermarking approaches under various signal attack scenarios. Zi Qian Yong, Shu-Min Leong, Sailaja Rajanala, Arghya Pal, Raphael C.-W. Phan |
SMC | 4 |
| 2023 | $\mathrm{C}\eta\iota \text{DAE}$: Cryptographically Distinguishing Autoencoder for Cipher CryptanalysisabstractWe propose a new autoencoder (AE) construction$\mathrm{C}\eta\iota \text{DAE}$(Cryptographically Distinguishing AE) based on a novel loss formulation to solve the cipher cryptanalysis distinguishing problem in the domain of cryptology. Vanilla AE and variational AE are unable to address this problem as they are designed to draw new samples which are either similar to the input sample or are from the same distribution. Such generated samples do not facilitate the cryptanalysis task. We show that our AE construction enables the discovery of cipher distinguishers, which are the fundamental building blocks that make or break new cipher design proposals. This also answers an open question on the applicability of autoencoders for cipher cryptanalysis; as to date, only discriminative models have been applied for cryptanalysis problems. To the best of our knowledge,$\mathrm{C}\eta\iota \text{DAE}$is the first-known generative model designed to solve crypt-analysis problems. We apply our$\mathrm{C}\eta\iota \text{DAE}$model to discover distinguishing properties for up to 10 rounds of the NSA-designed Speck32/64 cipher that allows to distinguish it from a random permutation. This contrasts with the best-known machine learning-discovered neural distinguisher in the literature that covers up to 8 rounds of Speck32/64. Unlike these recent related work which leverage on white box analysis and human-guided differential or linear analysis in order for machine learning models to be applicable, our$\mathrm{C}\eta\iota \text{DAE}$distinguisher does not require prior human cryptanalytic knowledge. This motivates the new direction of human-unsupervised machine learning-based cryptanalysis techniques. Raphael C.-W. Phan, Arghya Pal, Koksheik Wong, Sailaja Rajanala |
GLOBECOM | 2 |
| 2023 | Self Supervised Bert for Legal Text ClassificationabstractCritical BERT-based text classification tasks, such as legal text classification, require huge amounts of accurately labeled data. Legal text classification faces two trivial problems: labeling legal data is a sensitive process and can only be carried out by skilled professionals, and legal text is prone to privacy issues hence not all the data can be made available in the public domain. This means that we have limited diversity in the textual data, and to account for this data paucity, we propose a self-supervision approach to train Legal-BERT classifiers. We use the BERT text classifier’s knowledge of the class boundaries and perform gradient ascent w.r.t. class logits. Synthetic latent texts are generated through activation maximization. The main advantages over existing SOTAs are that our model: is easy to train, does not require much data but instead uses the synthesized data as fake samples; has less variance that helps to generate texts with good sample quality and diversity. We show the efficacy of the proposed method on the ECHR Violation (Multi-Label) Dataset and the Over-ruling Task Dataset. Arghya Pal, Sailaja Rajanala, Raphael C.-W. Phan, Koksheik Wong |
ICASSP | 1 |
| 2023 | USURP: Universal Single-Source Adversarial Perturbations on Multimodal Emotion RecognitionabstractThe field of affective computing has progressed from traditional unimodal analysis to more complex multimodal analysis due to the proliferation of videos posted online. Multimodal learning has shown remarkable performance in emotion recognition tasks, but its robustness in an adversarial setting remains unknown. This paper investigates the robustness of multimodal emotion recognition models against worst-case adversarial perturbations on a single modality. We found that standard multimodal models are susceptible to single-source adversaries and can be easily fooled by perturbations on any single modality. We draw some key observations that serve as guidelines for designing universal adversarial attacks on multimodal emotion recognition models. Motivated by these findings, we propose a novel universal single-source adversarial perturbations framework on multimodal emotion recognition models: USURP. Through our analysis of adversarial robustness, we demonstrate the necessity of studying adversarial attacks on multimodal models. Our experimental results show that the proposed USURP method achieves high attack success rates and significantly improves adversarial transferability in multimodal settings. The observations and novel attack methods presented in this paper provide a new understanding of the adversarial robustness of multimodal models, contributing to their safe and reliable deployment in more real- world scenarios. Yin Yin Low, Raphael C.-W. Phan, Arghya Pal, Xiaojun Chang |
ICIP | 3 |
| 2022 | Distilling the Undistillable: Learning from a Nasty Teacher
Surgan Jandial, Yash Khasbage, Arghya Pal, Vineeth N. Balasubramanian, Balaji Krishnamurthy |
ECCV (13) | 3 |
| 2022 | Guess-It-Generator: Generating in a Lewis Signaling Framework through Logical ReasoningabstractHuman minds spontaneously integrate two inherited cognitive capabilities: perception and reasoning to accomplish cognitive tasks such as problem solving, imagination, and causation. It is observed in the primate brains that perception offers the assistance required for problem comprehension, whilst the reasoning elucidates upon the facts recovered during perception in order to make a decision. The field of artificial intelligence (AI) thus considers perception and reasoning as two complementary areas that are realized by machine learning and logic programming, respectively. In this work, we propose a generative model using a collaborative guessing game of the kind first introduced by David Lewis in his famous work called the Lewis signaling game that is synonymous with the "20 Questions'' game. Our proposed model, Guess-It-Generator (GIG) is a collaborative framework that engages two recurrent neural networks in a guessing game. GIG unifies perception and reasoning with a view to generating labeled images by capturing, (X, y), the underlying density of a data distribution, i.e. (X, y) - p(X, y). An encoder attends to a region of the input image and encodes that onto a latent variable that acts as a perception signal to a decoder. In contrast, the decoder leverages on the perception signals to guess the image and verifies the guess by reasoning with logical facts derived from the domain knowledge. Our experiments and comprehensive studies on seven datasets: PCAM, Chest-Xray-14, FIRE, HAM10000 from the medical domain, and CIFAR 10, LSUN, ImageNet, among standard benchmark datasets, show significant promise for the proposed method. Arghya Pal, Sailaja Rajanala, Raphael C.-W. Phan, Koksheik Wong |
ACM Multimedia | 1 |
| 2022 | DeSCoVeR: Debiased Semantic Context Prior for Venue RecommendationabstractWe present a novel semantic context prior-based venue recommendation system that uses only the title and the abstract of a paper. Based on the intuition that the text in the title and abstract have both semantic and syntactic components, we demonstrate that a joint training of a semantic feature extractor and syntactic feature extractor collaboratively leverages meaningful information that helps to provide venues for papers. The proposed methodology that we call DeSCoVeR at first elicits these semantic and syntactic features using a Neural Topic Model and text classifier respectively. The model then executes a transfer learning optimization procedure to perform a contextual transfer between the feature distributions of the Neural Topic Model and the text classifier during the training phase. DeSCoVeR also mitigates the document-level label bias using a Causal back-door path criterion and a sentence-level keyword bias removal technique. Experiments on the DBLP dataset show that DeSCoVeR outperforms the state-of-the-art methods. Sailaja Rajanala, Arghya Pal, Manish Singh 0002, Raphael C.-W. Phan, Koksheik Wong |
SIGIR | 2 |
| 2021 | Synthesize-It-Classifier: Learning a Generative Classifier Through Recurrent Self-AnalysisabstractWe show the generative capability of an image classifier network by synthesizing high-resolution, photo-realistic, and diverse images at scale. The overall methodology, called Synthesize-It-Classifier (STIC), does not require an explicit generator network to estimate the density of the data distribution and sample images from that, but instead uses the classifier’s knowledge of the boundary to perform gradient ascent w.r.t. class logits and then synthesizes images using the Gram Matrix Metropolis Adjusted Langevin Algorithm (GRMALA) by drawing on a blank canvas. During training, the classifier iteratively uses these synthesized images as fake samples and re-estimates the class boundary in a recurrent fashion to improve both the classification accuracy and quality of synthetic images. The STIC shows that mixing of the hard fake samples (i.e. those synthesized by the one-hot class conditioning), and the soft fake samples (which are synthesized as a convex combination of classes, i.e. a mixup of classes [36]) improves class interpolation. We demonstrate an Attentive-STIC network that shows iterative drawing of synthesized images on the ImageNet dataset that has thousands of classes. In addition, we introduce the synthesis using a class conditional score classifier (Score-STIC) instead of a normal image classifier and show improved results on several real world datasets, i.e. ImageNet, LSUN and CIFAR 10. Arghya Pal, Raphael C.-W. Phan, Koksheik Wong |
CVPR | 1 |
| 2019 | Zero-Shot Task TransferabstractIn this work, we present a novel meta-learning algorithm that regresses model parameters for novel tasks for which no ground truth is available (zero-shot tasks). In order to adapt to novel zero-shot tasks, our meta-learner learns from the model parameters of known tasks (with ground truth) and the correlation of known tasks to zero-shot tasks. Such intuition finds its foothold in cognitive science, where a subject (human baby) can adapt to a novel concept (depth understanding) by correlating it with old concepts (hand movement or self-motion), without receiving an explicit supervision. We evaluated our model on the Taskonomy dataset, with four tasks as zero-shot: surface normal, room layout, depth and camera pose estimation. These tasks were chosen based on the data acquisition complexity and the complexity associated with the learning process using a deep network. Our proposed methodolgy outperforms state-of-the-art models (which use ground truth) on each of our zero-shot tasks, showing promise on zero-shot task transfer. We also conducted extensive experiments to study the various choices of our methodology, as well as showed how the proposed method can also be used in transfer learning. To the best of our knowledge, this is the first such effort on zero-shot learning in the task space. Arghya Pal, Vineeth N. Balasubramanian |
CVPR | 1 |
| 2019 | C4Synth: Cross-Caption Cycle-Consistent Text-to-Image SynthesisabstractGenerating an image from its description is a challenging task worth solving because of its numerous practical applications ranging from image editing to virtual reality. All existing methods use one single caption to generate a plausible image. A single caption by itself, can be limited and may not be able to capture the variety of concepts and behavior that would be present in the image. We propose two deep generative models that generate an image by making use of multiple captions describing it. This is achieved by ensuring 'Cross-Caption Cycle Consistency' between the multiple captions and the generated image(s). We report quantitative and qualitative results on the standard Caltech-UCSD Birds (CUB) and Oxford-102 Flowers datasets to validate the efficacy of the proposed approach. K. J. Joseph, Arghya Pal, Sailaja Rajanala, Vineeth N. Balasubramanian |
WACV | 2 |
| 2018 | Adversarial Data Programming: Using GANs to Relax the Bottleneck of Curated Labeled DataabstractPaucity of large curated hand-labeled training data forms a major bottleneck in the deployment of machine learning models in computer vision and other fields. Recent work (Data Programming) has shown how distant supervision signals in the form of labeling functions can be used to obtain labels for given data in near-constant time. In this work, we present Adversarial Data Programming (ADP), which presents an adversarial methodology to generate data as well as a curated aggregated label, given a set of weak labeling functions. We validated our method on the MNIST, Fashion MNIST, CIFAR 10 and SVHN datasets, and it outperformed many state-of-the-art models. We conducted extensive experiments to study its usefulness, as well as showed how the proposed ADP framework can be used for transfer learning as well as multi-task learning, where data from two domains are generated simultaneously using the framework along with the label information. Our future work will involve understanding the theoretical implications of this new framework from a game-theoretic perspective, as well as explore the performance of the method on more complex datasets. Arghya Pal, Vineeth N. Balasubramanian |
CVPR | 1 |
| 2017 | Have i reached the intersection: A deep learning-based approach for intersection detection from monocular camerasabstractLong-short term memory networks(LSTM) models have shown considerable performance on variety of problems dealing with sequential data. In this paper, we propose a variant of Long-Term Recurrent Convolutional Network(LRCN) to detect road intersection. We call this network as IntersectNet. We pose road intersection detection as binary classification task over sequence of frames. The model combines deep hierarchical visual feature extractor with recurrent sequence model. The model is end to end trainable with capability of capturing the temporal dynamics of the system. We exploit this capability to identify road intersection in a sequence of temporally consistent images. The model has been rigorously trained and tested on various different datasets. We think that our findings could be useful to model behavior of autonomous agent in the real-world. Dhaivat Bhatt, Danish Sodhi, Arghya Pal, Vineeth N. Balasubramanian, K. Madhava Krishna |
IROS | 3 |