VLDB 2026 Research / reviewers in the wild / expert
Balaji Veeramani
dblp:08/840
· DBLP profile ↗
9ranked-venue papers
2as first author
7since 2021 · last 2025
0009-0002-5263-1210ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 6 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Citations and Trust in LLM Generated ResponsesabstractQuestion answering systems are rapidly advancing, but their opaque nature may impact user trust. We explored trust through an anti-monitoring framework, where trust is predicted to be correlated with presence of citations and inversely related to checking citations. We tested this hypothesis with a live question-answering experiment that presented text responses generated using a commercial Chatbot along with varying citations (zero, one, or five), both relevant and random, and recorded if participants checked the citations and their self-reported trust in the generated responses. We found a significant increase in trust when citations were present, a result that held true even when the citations were random; we also found a significant decrease in trust when participants checked the citations. These results highlight the importance of citations in enhancing trust in AI-generated content. Yifan Ding 0001, Matthew Facciani, Ellen Joyce, Amrit Poudel, Sanmitra Bhattacharya, Balaji Veeramani, Salvador Aguiñaga, Tim Weninger |
AAAI | 6 |
| 2025 | MentorPDM: Learning Data-Driven Curriculum for Multi-Modal Predictive MaintenanceabstractPredictive Maintenance (PDM) systems are essential for preemptive monitoring of sensor signals to detect potential machine component failures in industrial assets such as bearings in rotating machinery. Existing PDM systems face two primary challenges: 1) Irregular Signal Acquisition, where data collection from the sensors is intermittent, and 2) Signal Heterogeneity, where the full spectrum of sensor modalities is not effectively integrated. To address these challenges, we propose a Curriculum Learning Framework for Multi-Modal Predictive Maintenance - MentorPDM. MentorPDM consists of 1) a graph-augmented pretraining module that captures intrinsic and structured temporal correlations across time segments via a temporal contrastive learning objective and 2) a bi-level curriculum learning module that captures task complexities for weighing the importance of signal modalities and samples via modality and sample curricula. Empirical results from MentorPDM show promising performance with better generalizability in PDM tasks compared to existing benchmarks. The efficacy of the MentorPDM model will be further demonstrated in real industry testbeds and platforms. Shuaicheng Zhang, Sanmitra Bhattacharya, Sunil Reddy Tiyyagura, Edward Bowen, Balaji Veeramani, Dawei Zhou 0003 |
KDD (1) | 7 |
| 2024 | Data Composition for Continual Learning in Application of Cyberattack Detection
Jiayi Lian, Kevin Choi, Balaji Veeramani, Sathvik Murli, Alison Hu, Laura J. Freeman, Edward Bowen, Xinwei Deng |
ASONAM (4) | 4 |
| 2024 | EvoluNet: Advancing Dynamic Non-IID Transfer Learning on GraphsabstractNon-IID transfer learning on graphs is crucial in many high-stakes domains. The majority of existing works assume stationary distribution for both source and target domains. However, real-world graphs are intrinsically dynamic, presenting challenges in terms of domain evolution and dynamic discrepancy between source and target domains. To bridge the gap, we shift the problem to the dynamic setting and pose the question: given the *label-rich* source graphs and the *label-scarce* target graphs both observed in previous $T$ timestamps, how can we effectively characterize the evolving domain discrepancy and optimize the generalization performance of the target domain at the incoming $T+1$ timestamp? To answer it, we propose a generalization bound for *dynamic non-IID transfer learning on graphs*, which implies the generalization performance is dominated by domain evolution and domain discrepancy between source and target graphs. Inspired by the theoretical results, we introduce a novel generic framework named EvoluNet. It leverages a transformer-based temporal encoding module to model temporal information of the evolving domains and then uses a dynamic domain unification module to efficiently learn domain-invariant representations across the source and target domains. Finally, EvoluNet outperforms the state-of-the-art models by up to 12.1%, demonstrating its effectiveness in transferring knowledge from dynamic source graphs to dynamic target graphs. Haohui Wang, Yuzhen Mao, Yujun Yan, Yaoqing Yang 0002, Jianhui Sun, Kevin Choi, Balaji Veeramani, Alison Hu, Edward Bowen, Tyler Cody, Dawei Zhou 0003 |
ICML | 7 |
| 2024 | Toward Robust Generative AI Text Detection: Generalizable Neural ModelabstractLarge Language Models (LLMs) have demonstrated remarkable capabilities in generating text that closely resembles human writing across wide range of styles and genres. However, such capabilities are prone to potential misuse, such as fake news generation, spam email creation, and misuse in academic assignments. Hence, it is essential to build automated approaches capable of distinguishing between Artificial Intelligence-generated text and human-authored text. In this paper, we proposed a fine-tuning based neural model which is a combination of transformer models, linguistic features and state-of-the-art embedding models. We have also curated a training dataset encompassing diverse samples from different LLMs and domains to fine-tune pretrained language models. We evaluated our model's performance against state-of-the-art methods, and the comparative analysis demonstrates that our approach consistently outperforms other methods across various datasets using established evaluation metrics. Harika Abburi, Nirmala Pudota, Balaji Veeramani, Edward Bowen, Sanmitra Bhattacharya |
ICMLA | 3 |
| 2024 | Pre-train. Mixup and Fine-tune: A Simple Strategy to Handle Domain ShiftabstractTransfer learning leverages models trained on large source datasets to target domains with limited datasets by fine- tuning pre-trained models. These approaches work well with minimal distribution shifts. However, the nature of distribution shift is unknown in real-world applications which may lead to worse performance of models in the field. Domain adaptation approaches handles this issue explicitly by adapting source trained models using discrepancy, adversarial or reconstruction based approaches, but doesn't tackle the limited target data issue. A data augmentation approach such as mixup helps train resilient models with limited datasets, and is recently being considered for its ability to handle distribution shifts. In this work, we investigate how mixup can be used along with transfer learning to improve model performance on target domains with distribution shifts. Our experimental results shows mixup is complementary to transfer learning which we demonstrate by varying the percentage of available target training data with publicly available source and target datasets. Our proposed approach of using mixup along with fine tuning shows improved performance than just fine tuning or mixup across varying percentage of target dataset sizes. Haider Ilyas, Harika Abburi, Edward Bowen, Balaji Veeramani |
ICMLA | 4 |
| 2023 | An Ensemble-Based Approach for Generative Language Model Attribution
Harika Abburi, Michael Suesserman, Nirmala Pudota, Balaji Veeramani, Edward Bowen, Sanmitra Bhattacharya |
WISE | 4 |
| 2018 | DeepSort: deep convolutional networks for sorting haploid maize seedsabstractBACKGROUND: Maize is a leading crop in the modern agricultural industry that accounts for more than 40% grain production worldwide. THe double haploid technique that uses fewer breeding generations for generating a maize line has accelerated the pace of development of superior commercial seed varieties and has been transforming the agricultural industry. In this technique the chromosomes of the haploid seeds are doubled and taken forward in the process while the diploids marked for elimination. Traditionally, selective visual expression of a molecular marker within the embryo region of a maize seed has been used to manually discriminate diploids from haploids. Large scale production of inbred maize lines within the agricultural industry would benefit from the development of computer vision methods for this discriminatory task. However the variability in the phenotypic expression of the molecular marker system and the heterogeneity arising out of the maize genotypes and image acquisition have been an enduring challenge towards such efforts. RESULTS: In this work, we propose a novel application of a deep convolutional network (DeepSort) for the sorting of haploid seeds in these realistic settings. Our proposed approach outperforms existing state-of-the-art machine learning classifiers that uses features based on color, texture and morphology. We demonstrate the network derives features that can discriminate the embryo regions using the activations of the neurons in the convolutional layers. Our experiments with different architectures show that the performance decreases with the decrease in the depth of the layers. CONCLUSION: Our proposed method DeepSort based on the convolutional network is robust to the variation in the phenotypic expression, shape of the corn seeds, and the embryo pose with respect to the camera. In the era of modern digital agriculture, deep learning and convolutional networks will continue to play an important role in advancing research and product development within the agricultural industry. Balaji Veeramani, John W. Raymond, Pritam Chanda |
BMC Bioinform. | 1 |
| 2004 | Measuring the direction and the strength of coupling in nonlinear Systems-a modeling approach in the State spaceabstractWe present a novel signal processing methodology to determine the direction and the strength of coupling between coupled nonlinear systems. The methodology is based on multivariate local linear prediction in the reconstructed state spaces of the observed variables from each multivariable nonlinear system. Application of the method is illustrated with systems of coupled Rossler and Lorenz oscillators in various coupling configurations. The obtained results are compared with ones produced by the use of the directed transfer function, a model-based method in the time domain. Through a surrogate analysis, it is shown that the proposed method is more reliable than the directed transfer function in identifying the direction and strength of the involved interactions. Balaji Veeramani, K. Narayanan, Awadhesh Prasad, Leonidas D. Iasemidis, Andreas Spanias, Kostas Tsakalis |
IEEE Signal Process. Lett. | 1 |