VLDB 2026 Research / reviewers in the wild / expert
Sailaja Rajanala
dblp:214/7106
· DBLP profile ↗
16ranked-venue papers
2as first author
13since 2021 · last 2026
0000-0003-0529-4389ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 7 · 6 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Computer networks · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | FLAG-4D: Flow-Guided Local-Global Dual-Deformation Model for 4D ReconstructionabstractWe introduce FLAG-4D, a novel framework for generating novel views of dynamic scenes by reconstructing how 3D Gaussian primitives evolve through space and time. Existing methods typically rely on a single Multilayer Perceptron(MLP) to model temporal deformations, and they often struggle to capture complex point motions and fine-grained dynamic details consistently over time, especially from sparse input views. Our approach, FLAG-4D overcomes this by employing a dual-deformation network that dynamically warps a canonical set of 3D Gaussians over time into new positions and anisotropic shapes. This dual-deformation network consists of an Instantaneous Deformation Network (IDN) for modeling fine-grained, local deformations, and Global Motion Network (GMN) for capturing long-range dynamics, refined via mutual learning. To ensure these deformations are both accurate and temporally smooth, FLAG-4D incorporates dense motion features from a pretrained optical flow backbone. We fuse these motion cues from adjacent timeframes and use a deformation-guided attention mechanism to align this flow information with the current state of each evolving 3D Gaussian. Extensive experiments demonstrate that FLAG-4D achieves higher-fidelity and more temporally coherent reconstructions with finer detail preservation than state-of-the-art methods. Guan Yuan Tan, Ngoc Tuan Vu, Arghya Pal, Sailaja Rajanala, Raphael C.-W. Phan, Mettu Srinivas, Chee-Ming Ting |
AAAI | 4 |
| 2026 | Causal-Ex: Causal graph-based micro and macro expression spottingabstractDetecting concealed emotions within apparently normal expressions is crucial for identifying potential mental health issues and facilitating timely support and intervention. The task of spotting macro- and micro-expressions involves predicting the emotional timeline within a video by identifying the onset (i.e., the beginning), apex (the peak of emotion), and offset (the end of emotion) frames of the displayed emotions. More particularly, closely monitoring the key emotion-conveying regions of the face; namely, the foundational muscle-movement cues known as facial action units (AUs)–greatly aids in the clear identification of micro-expressions. One major roadblock is the inadvertent introduction of biases into the training process, which degrades performance regardless of feature quality. Biases are spurious factors that falsely inflate or deflate performance metrics. For instance, the neural networks tend to falsely attribute certain AUs in specific facial regions to particular emotion classes, a phenomenon also termed as Inductive biases. To remove these false attributions, we must identify and mitigate biases that arise from mere correlation between some features and the output class labels. We hence introduce action-unit causal graphs. Unlike the traditional action-unit graph, which connects AUs based solely on spatial adjacency, the causal AU graph is derived from statistical tests and retains edges between AUs only when there is significant evidence that one AU causally influences another. Our model, named Causal-Ex ( Causal -based Ex pression spotting), employs a fast causal inference algorithm to construct a causal graph of facial region of interests (ROIs). This enables us to select causally relevant facial action units in the ROIs. Our work demonstrates improvement in overall F1-scores compared to state-of-the-art approaches with 0.388 on CAS(ME) 2 and 0.3701 on SAMM-Long Video datasets. Our code can be found at: https://github.com/noobasuna/causal_ex.git . Pei-Sze Tan, Sailaja Rajanala, Arghya Pal, Raphael C.-W. Phan, Huey Fang Ong |
Pattern Recognit. Lett. | 2 |
| 2025 | GENIE: Socially Unbiased Generative Text-to-Image EditingabstractGenerative diffusion models often exhibit societal biases in sensitive personal attributes such as age, gender, and race. In this work, we describe GENIE – a method to reduce such biases in a variety of classifier-free diffusion models used for image editing. Our method implicitly incorporates debiasing terms together with the user’s explicit edit instruction to reduce bias. This automatic method relieves the user from needing to modify edit instructions in order to avoid bias. Further, no additional training is needed. Experimental results are provided based on modifications to four diffusion models, namely InstructPix2Pix, Stable Diffusion 1.5, Stable Diffusion 2.1, and Stable Diffusion XL. We show that, on average, bias is reduced by 31% in gender, 15% in age, 39% in race. Julia Kaiwen Lau, Raphael C.-W. Phan, Sailaja Rajanala, Ingemar J. Cox, Arghya Pal |
ICASSP | 3 |
| 2025 | Post-Hoc Adversarial Stickers Against Micro-Expression LeakageabstractSecuring micro-expressions against leakage is crucial for privacy, as these subtle facial movements convey genuine emotions and are inherently personal. This study aims to protect micro-expression data from potential adversarial attacks, ensuring the preservation of individuals’ privacy and preventing unauthorized access or misuse of sensitive emotional information. Unlike traditional methods, which often require training and extensive access to models, this research introduces a novel post-hoc method that does not require additional training. We focus on physical adversarial attacks in micro-expression recognition, involving intentional manipulation of visual cues to deceive recognition systems and protect individual emotional privacy. Our approach leverages a causal discovery algorithm to identify causal relationships between facial parts, enabling rapid identification of the optimal locations for adversarial patches in frames with triggered micro-expressions. This method exhibits a more consistent attack success rate than randomly placed adversarial stickers, demonstrating effective generalization across different emotions, stickers, and models. Particularly relevant in scenarios with restricted access to the model, our technique requires only a single interaction during the attack process, highlighting its efficiency and minimal need for querying the target model. The proposed method effectively balances privacy protection with high generalization capability, setting a new standard for defending against adversarial threats in micro-expression recognition. The code is available at https://github.com/noobasuna/au-sticker. Pei-Sze Tan, Sailaja Rajanala, Yee-Fan Tan, Arghya Pal, Chun-Ling Tan, Raphael C.-W. Phan, Huey Fang Ong |
ICASSP | 2 |
| 2025 | Polyfit generative model: can a group of lower-order polynomials generate high resolution diverse images?abstractImplicit neural representations (INRs) have recently gained popularity as a means to model images as continuous functions of spatial coordinates, synthesizing each pixel independently and yielding impressive results in tasks such as scene reconstruction and image generation. A notable advancement, PolyINR, utilizes element-wise multiplications between features and affine-transformed coordinates to achieve higher-order polynomial functions, eliminating the need for positional encodings. However, the finite encoding capacity of INRs, coupled with PolyINR’s recursive polynomial estimation, necessitates substantial training parameters, resulting in high computational costs and limiting applicability across diverse computer vision domains. In this work, we address these challenges by representing images as grids of smaller patches, within which we fit low-degree polynomials to capture local intensity variations. Our approach substantially reduces parameter requirements and computational demands. We evaluate our model qualitatively and quantitatively on large-scale datasets, ImageNet, CelebA, LSUN Bedroom, and Flower102; thus demonstrating competitive performance with state-of-the-art generative models, despite the absence of convolutional, normalization, or self-attention layers. Arghya Pal, Ai-Fang Chai, Sailaja Rajanala, Raphael C.-W. Phan, Koksheik Wong, Chee-Ming Ting |
IJCNN | 3 |
| 2025 | Res-SH: Unbiased Residual Learning for Self-Healing Interface Toughness Prediction with Limited DataabstractThe development of self-healing materials is often hindered by the high costs and material waste associated with traditional characterization methods. Current approaches to toughness prediction, primarily based on convolutional neural networks (CNNs), are limited by their tendency to capture only surface-level features, which can lead to biased predictions. Moreover, working with small datasets, which is common in materials science, further increases the risk of biased training due to overfitting, posing a critical challenge to the reliability and generalizability of predictive models. This study introduces an unbiased residual learning framework designed explicitly for predicting self-healing interface toughness under limiteddata conditions. Our approach, ResNet-inspired approach for predicting self-healing material toughness, named Res-SH, used the power of residual networks to capture deeper, more complex patterns in the data, thereby addressing critical challenges in materials research. Res-SH minimises resource consumption and experimental overhead by focusing on unbiased learning, achieving accurate predictions with fewer training epochs and lower R2score and root mean square prediction errors compared to conventional CNN and lightweight model MobileNetv2. This novel framework provides a cost-effective and resource-efficient alternative to traditional material characterization methods, reducing material waste and accelerating the discovery and optimization of self-healing material systems. Pei-Sze Tan, Karen Jia-Jun Koh, Sailaja Rajanala, Arghya Pal, Raphael C.-W. Phan, Nan Ze, Fuad Noman, Chee-Ming Ting, Norfadilah Dolmat, Nik Nur Wahidah Nik Hashim, Afidalina Tumian |
TENCON | 3 |
| 2025 | Disentangle Class Imbalance in Micro-Expression Recognition with Causal Structure LearningabstractMicro-expression recognition is a critical task in affective computing, yet it is often hindered by biases inherent in datasets, leading to skewed and unreliable outcomes. This work introduces a novel approach to disentangling bias in microexpression recognition using causal structure learning. By modeling causal relationships within the data, we identify and mitigate sources of bias that traditional machine learning models often overlook. Our framework integrates causal discovery techniques to uncover biased patterns in widely used micro-expression datasets. We employ debiasing strategies to enhance the fairness and accuracy of recognition models and generate counterfactual examples to address sensitive attributes such as gender and age, allowing us to observe the effects of imbalanced classes on classification results. Experiments conducted on baseline microexpression recognition models demonstrate comparable results after undersampling to create emotion class balance, revealing label bias in current training datasets including CASME2, SAMM, and SMIC. Further evaluation on balanced gender and age classes using generated counterfactual data as additional training instances showed performance improvements for the 4DME dataset. Pei-Sze Tan, Sailaja Rajanala, Raphael C.-W. Phan |
TENCON | 2 |
| 2024 | Causally Uncovering Bias in Video Micro-Expression RecognitionabstractDetecting microexpressions presents formidable challenges, primarily due to their fleeting nature and the limited diversity in existing datasets. Our studies find that these datasets exhibit a pronounced bias towards specific ethnicities and suffer from significant imbalances in terms of both class and gender representation among the samples. These disparities create fertile ground for various biases to permeate deep learning models, leading to skewed results and inadequate portrayal of specific demographic groups. Our research is driven by a compelling need to identify and rectify these biases within model architectures. To achieve this, we commence by constructing a causal graph that elucidates the intricate relationships between the model, input features, and training outcomes. This graphical representation forms the foundation for our analytical framework. Leveraging this causal framework, we conduct comprehensive case studies, employing counterfactuals as a diagnostic tool to unveil biases arising from dataset-induced class imbalances, gender inequalities, and variations in facial action units. Our final step involves a highly efficient counterfactual debiasing process, eliminating the necessity for additional data collection or model retraining. Our results showcase superior performance compared to state-of-the-art methods across the CASME II, SAMM, and SMIC datasets. Pei-Sze Tan, Sailaja Rajanala, Arghya Pal, Shu-Min Leong, Raphael C.-W. Phan, Huey Fang Ong |
ICASSP | 2 |
| 2024 | Distingusic: Distinguishing Synthesized Music from HumanabstractIn this paper we focus on a problem that is increasingly plaguing the music industry; to a large extent due to the proliferation of generative AI models that enable the generation of new realistic and indistinguishable content for diverse modalities: text, image, audio, video. We address this problem from the perspective of audio watermarking; to our best knowledge, this is the first-known watermarking based approach to solve the problem of distinguishing realistic songs synthesized from generative AI models from real songs sung by humans. In more detail, our approach specifically utilizes the SHA-256 hash function, Singular Value Decomposition (SVD) and Discrete Wavelet Transform (DWT) for robust audio watermarking of synthesized songs. Before embedding, the audio is subjected to an attack phase to pinpoint less vulnerable regions for QR watermark placement. During the embedding process, the audio chunks first undergo a 1-level Discrete Wavelet Transform (DWT), and then the resulting approximate coefficients go through Singular Value Decompo-sition (SVD). Additionally, the watermarked array is subjected to SHA-256 hashing for collision-resistant conciseness, which is subsequently embedded into the singular values of the audio. Experimental findings demonstrate the superiority of our method over existing audio watermarking approaches under various signal attack scenarios. Zi Qian Yong, Shu-Min Leong, Sailaja Rajanala, Arghya Pal, Raphael C.-W. Phan |
SMC | 3 |
| 2023 | $\mathrm{C}\eta\iota \text{DAE}$: Cryptographically Distinguishing Autoencoder for Cipher CryptanalysisabstractWe propose a new autoencoder (AE) construction$\mathrm{C}\eta\iota \text{DAE}$(Cryptographically Distinguishing AE) based on a novel loss formulation to solve the cipher cryptanalysis distinguishing problem in the domain of cryptology. Vanilla AE and variational AE are unable to address this problem as they are designed to draw new samples which are either similar to the input sample or are from the same distribution. Such generated samples do not facilitate the cryptanalysis task. We show that our AE construction enables the discovery of cipher distinguishers, which are the fundamental building blocks that make or break new cipher design proposals. This also answers an open question on the applicability of autoencoders for cipher cryptanalysis; as to date, only discriminative models have been applied for cryptanalysis problems. To the best of our knowledge,$\mathrm{C}\eta\iota \text{DAE}$is the first-known generative model designed to solve crypt-analysis problems. We apply our$\mathrm{C}\eta\iota \text{DAE}$model to discover distinguishing properties for up to 10 rounds of the NSA-designed Speck32/64 cipher that allows to distinguish it from a random permutation. This contrasts with the best-known machine learning-discovered neural distinguisher in the literature that covers up to 8 rounds of Speck32/64. Unlike these recent related work which leverage on white box analysis and human-guided differential or linear analysis in order for machine learning models to be applicable, our$\mathrm{C}\eta\iota \text{DAE}$distinguisher does not require prior human cryptanalytic knowledge. This motivates the new direction of human-unsupervised machine learning-based cryptanalysis techniques. Raphael C.-W. Phan, Arghya Pal, Koksheik Wong, Sailaja Rajanala |
GLOBECOM | 4 |
| 2023 | Self Supervised Bert for Legal Text ClassificationabstractCritical BERT-based text classification tasks, such as legal text classification, require huge amounts of accurately labeled data. Legal text classification faces two trivial problems: labeling legal data is a sensitive process and can only be carried out by skilled professionals, and legal text is prone to privacy issues hence not all the data can be made available in the public domain. This means that we have limited diversity in the textual data, and to account for this data paucity, we propose a self-supervision approach to train Legal-BERT classifiers. We use the BERT text classifier’s knowledge of the class boundaries and perform gradient ascent w.r.t. class logits. Synthetic latent texts are generated through activation maximization. The main advantages over existing SOTAs are that our model: is easy to train, does not require much data but instead uses the synthesized data as fake samples; has less variance that helps to generate texts with good sample quality and diversity. We show the efficacy of the proposed method on the ECHR Violation (Multi-Label) Dataset and the Over-ruling Task Dataset. Arghya Pal, Sailaja Rajanala, Raphael C.-W. Phan, Koksheik Wong |
ICASSP | 2 |
| 2022 | Guess-It-Generator: Generating in a Lewis Signaling Framework through Logical ReasoningabstractHuman minds spontaneously integrate two inherited cognitive capabilities: perception and reasoning to accomplish cognitive tasks such as problem solving, imagination, and causation. It is observed in the primate brains that perception offers the assistance required for problem comprehension, whilst the reasoning elucidates upon the facts recovered during perception in order to make a decision. The field of artificial intelligence (AI) thus considers perception and reasoning as two complementary areas that are realized by machine learning and logic programming, respectively. In this work, we propose a generative model using a collaborative guessing game of the kind first introduced by David Lewis in his famous work called the Lewis signaling game that is synonymous with the "20 Questions'' game. Our proposed model, Guess-It-Generator (GIG) is a collaborative framework that engages two recurrent neural networks in a guessing game. GIG unifies perception and reasoning with a view to generating labeled images by capturing, (X, y), the underlying density of a data distribution, i.e. (X, y) - p(X, y). An encoder attends to a region of the input image and encodes that onto a latent variable that acts as a perception signal to a decoder. In contrast, the decoder leverages on the perception signals to guess the image and verifies the guess by reasoning with logical facts derived from the domain knowledge. Our experiments and comprehensive studies on seven datasets: PCAM, Chest-Xray-14, FIRE, HAM10000 from the medical domain, and CIFAR 10, LSUN, ImageNet, among standard benchmark datasets, show significant promise for the proposed method. Arghya Pal, Sailaja Rajanala, Raphael C.-W. Phan, Koksheik Wong |
ACM Multimedia | 2 |
| 2022 | DeSCoVeR: Debiased Semantic Context Prior for Venue RecommendationabstractWe present a novel semantic context prior-based venue recommendation system that uses only the title and the abstract of a paper. Based on the intuition that the text in the title and abstract have both semantic and syntactic components, we demonstrate that a joint training of a semantic feature extractor and syntactic feature extractor collaboratively leverages meaningful information that helps to provide venues for papers. The proposed methodology that we call DeSCoVeR at first elicits these semantic and syntactic features using a Neural Topic Model and text classifier respectively. The model then executes a transfer learning optimization procedure to perform a contextual transfer between the feature distributions of the Neural Topic Model and the text classifier during the training phase. DeSCoVeR also mitigates the document-level label bias using a Causal back-door path criterion and a sentence-level keyword bias removal technique. Experiments on the DBLP dataset show that DeSCoVeR outperforms the state-of-the-art methods. Sailaja Rajanala, Arghya Pal, Manish Singh 0002, Raphael C.-W. Phan, Koksheik Wong |
SIGIR | 1 |
| 2020 | FLY: Venue Recommendation using Limited ContextabstractRecommendation of publication venues is very much the need of the hour. There is a sea of options for research conferences and journals. As a result, it is often found that researchers are not even aware of many publication venues in their research area. Existing work uses information, such as all the references in a given paper, past publication venues of the author(s), co-author relationships, and the paper content to do venue recommendation. However, in this paper, we propose a system that uses very limited context, namely the paper title, abstract, and a venue network, which is constructed using only a small subset of authors' publication history, to do venue recommendation. Our venue recommendation system, FLY, gives 30% higher accuracy compared to the current state-of-the-art limited context venue recommendation system. Sailaja Rajanala, Manish Singh 0002 |
ICTAI | 1 |
| 2019 | Evaluating the Choice of Tags in CQA Sites
Rohan Banerjee, Sailaja Rajanala, Manish Singh 0002 |
DASFAA (1) | 2 |
| 2019 | C4Synth: Cross-Caption Cycle-Consistent Text-to-Image SynthesisabstractGenerating an image from its description is a challenging task worth solving because of its numerous practical applications ranging from image editing to virtual reality. All existing methods use one single caption to generate a plausible image. A single caption by itself, can be limited and may not be able to capture the variety of concepts and behavior that would be present in the image. We propose two deep generative models that generate an image by making use of multiple captions describing it. This is achieved by ensuring 'Cross-Caption Cycle Consistency' between the multiple captions and the generated image(s). We report quantitative and qualitative results on the standard Caltech-UCSD Birds (CUB) and Oxford-102 Flowers datasets to validate the efficacy of the proposed approach. K. J. Joseph, Arghya Pal, Sailaja Rajanala, Vineeth N. Balasubramanian |
WACV | 3 |