VLDB 2026 Research / reviewers in the wild / expert
Adam D. Cobb
dblp:206/6601
· DBLP profile ↗
15ranked-venue papers
4as first author
12since 2021 · last 2026
0000-0003-2868-6983ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 13 · 4 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Privacy Preserving In-Context-Learning Framework for Large Language ModelsabstractLarge language models (LLMs) have significantly transformed natural language understanding and generation, but they raise privacy concerns due to potential exposure of sensitive information. Studies have highlighted the risk of information leakage, where adversaries can extract sensitive information embedded in the prompts. In this work, we introduce a novel private prediction framework for generating high-quality synthetic text with strong privacy guarantees. Our approach leverages the Differential Privacy (DP) framework to ensure worst-case theoretical bounds on information leakage without requiring any fine-tuning of the underlying models. The proposed method performs inference on private records and aggregates the resulting per-token output distributions. This enables the generation of longer and coherent synthetic text while maintaining privacy guarantees. Additionally, we propose a simple blending operation that combines private and public inference to further enhance utility. Empirical evaluations demonstrate that our approach outperforms previous state-of-the-art methods on in-context-learning (ICL) tasks, making it a promising direction for privacy-preserving text generation while maintaining high utility. Bishnu Bhusal, Manoj Acharya, Ramneet Kaur, Colin Samplawski, Adam D. Cobb, Rohit Chadha, Susmit Jha |
AAAI | 6 |
| 2025 | Polysemantic Dropout: Conformal OOD Detection for Specialized LLMsabstractWe propose a novel inference-time out-ofdomain (OOD) detection algorithm for specialized large language models (LLMs).Despite achieving state-of-the-art performance on in-domain tasks through fine-tuning, specialized LLMs remain vulnerable to incorrect or unreliable outputs when presented with OOD inputs, posing risks in critical applications.Our method leverages the Inductive Conformal Anomaly Detection (ICAD) framework, using a new non-conformity measure based on the model's dropout tolerance.Motivated by recent findings on polysemanticity and redundancy in LLMs, we hypothesize that in-domain inputs exhibit higher dropout tolerance than OOD inputs.We aggregate dropout tolerance across multiple layers via a valid ensemble approach, improving detection while maintaining theoretical false alarm bounds from ICAD.Experiments with medical-specialized LLMs show that our approach detects OOD inputs better than baseline methods, with AUROC improvements of 2% to 37% when treating OOD datapoints as positives and in-domain test datapoints as negatives. Ayush Gupta 0001, Ramneet Kaur, Adam D. Cobb, Rama Chellappa, Susmit Jha |
EMNLP | 4 |
| 2025 | SpikingVTG: A Spiking Detection Transformer for Video Temporal GroundingabstractVideo Temporal Grounding (VTG) aims to retrieve precise temporal segments in a video conditioned on natural language queries. Unlike conventional neural frameworks that rely heavily on computationally expensive dense matrix multiplications, Spiking Neural Networks (SNNs)—previously underexplored in this domain—offer a unique opportunity to tackle VTG tasks through bio-plausible spike-based communication and an event-driven accumulation-based computational paradigm. We introduce SpikingVTG, a multi-modal spiking detection transformer, designed to harness the computational simplicity and sparsity of SNNs for VTG tasks. Leveraging the temporal dynamics of SNNs, our model introduces a Saliency Feedback Gating (SFG) mechanism that assigns dynamic saliency scores to video clips and applies multiplicative gating to highlight relevant clips while suppressing less informative ones. SFG enhances performance and reduces computational overhead by minimizing neural activity. We analyze the layer-wise convergence dynamics of SFG-enabled model and apply implicit differentiation at equilibrium to enable efficient, BPTT-free training. To improve generalization and maximize performance, we enable knowledge transfer by optimizing a Cos-L2 representation matching loss that aligns the layer-wise representation and attention maps of a non-spiking teacher with those of our student SpikingVTG. Additionally, we present Normalization-Free (NF)-SpikingVTG, which eliminates non-local operations like softmax and layer normalization, and an extremely quantized 1-bit (NF)-SpikingVTG variant for potential deployment on edge devices. Our models achieve competitive results on QVHighlights, Charades-STA, TACoS, and YouTube Highlights, establishing a strong baseline for multi-modal spiking VTG solutions. Malyaban Bal, Brian Matejek, Susmit Jha, Adam D. Cobb |
NeurIPS | 4 |
| 2025 | Scalable Bayesian Low-Rank Adaptation of Large Language Models via Stochastic Variational Subspace InferenceabstractDespite their widespread use, large language models (LLMs) are known to hallucinate incorrect information and be poorly calibrated. This makes the uncertainty quantification of these models of critical importance, especially in high-stakes domains, such as autonomy and healthcare. Prior work has made Bayesian deep learning-based approaches to this problem more tractable by performing inference over the low-rank adaptation (LoRA) parameters of a fine-tuned model. While effective, these approaches struggle to scale to larger LLMs due to requiring further additional parameters compared to LoRA. In this work we present $\textbf{Scala}$ble $\textbf{B}$ayesian $\textbf{L}$ow-Rank Adaptation via Stochastic Variational Subspace Inference (ScalaBL). We perform Bayesian inference in an $r$-dimensional subspace, for LoRA rank $r$. By repurposing the LoRA parameters as projection matrices, we are able to map samples from this subspace into the full weight space of the LLM. This allows us to learn all the parameters of our approach using stochastic variational inference. Despite the low dimensionality of our subspace, we are able to achieve competitive performance with state-of-the-art approaches while only requiring ${\sim}1000$ additional parameters. Furthermore, it allows us to scale up to the largest Bayesian LLM to date, with four times as a many base parameters as prior work. Colin Samplawski, Adam D. Cobb, Manoj Acharya, Ramneet Kaur, Susmit Jha |
UAI | 2 |
| 2025 | Zero-Shot Detection of Out-of-Context Objects Using Foundation ModelsabstractWe address the problem of detecting out-of-context (OOC) objects in a scene. Given an image, we aim to detect whether the image has objects that are not present in their usual context and localize such OOC objects. Existing approaches for OOC detection rely on defining the common context in terms of the manually constructed features, such as the co-occurrence of objects, spatial relations between objects, and shape and size of the objects, and then learning such context for a given dataset. But context is often nu-anced ranging from very common to very surprising. Further, learned context from specific datasets may not be generalized as datasets may not truly represent the human notion of what is in context. Motivated by the success of large language models and more generally, foundation models (FMs) in common sense reasoning, we investigate the FM's ability to capture a more generalized notion of context. We find that a pre-trained FM, such as GPT-4, provides a more nuanced notion of OOC and enables zero-shot OOC detection when coupled with other pre-trained FMs for caption generation such as BLIP-2, and image in-painting with Sta-ble Diffusion 2.0. Our approach does not need any dataset-specific training. We demonstrate the efficacy of our approach on two OOC object detection datasets, achieving 90.8% zero-shot accuracy on the MIT-OOC dataset and 87.26% on the IJCAI22-COCO-OOC dataset. Adam D. Cobb, Ramneet Kaur, Sumit Kumar Jha 0001, Nathaniel D. Bastian, Alexander M. Berenbeim, Robert Thomson 0001, Iain Cruickshank, Alvaro Velasquez, Susmit Jha |
WACV | 2 |
| 2024 | Direct Amortized Likelihood Ratio EstimationabstractWe introduce a new amortized likelihood ratio estimator for likelihood-free simulation-based inference (SBI). Our estimator is simple to train and estimates the likelihood ratio using a single forward pass of the neural estimator. Our approach directly computes the likelihood ratio between two competing parameter sets which is different from the previous approach of comparing two neural network output values. We refer to our model as the direct neural ratio estimator (DNRE). As part of introducing the DNRE, we derive a corresponding Monte Carlo estimate of the posterior. We benchmark our new ratio estimator and compare to previous ratio estimators in the literature. We show that our new ratio estimator often outperforms these previous approaches. As a further contribution, we introduce a new derivative estimator for likelihood ratio estimators that enables us to compare likelihood-free Hamiltonian Monte Carlo (HMC) with random-walk Metropolis-Hastings (MH). We show that HMC is equally competitive, which has not been previously shown. Finally, we include a novel real-world application of SBI by using our neural ratio estimator to design a quadcopter. Code is available at https://github.com/SRI-CSL/dnre. Adam D. Cobb, Brian Matejek, Daniel Elenius, Susmit Jha |
AAAI | 1 |
| 2023 | AircraftVerse: A Large-Scale Multimodal Dataset of Aerial Vehicle DesignsabstractWe present AircraftVerse, a publicly available aerial vehicle design dataset. Aircraft design encompasses different physics domains and, hence, multiple modalities of representation. The evaluation of these designs requires the use of scientific analytical and simulation models ranging from computer-aided design tools for structural and manufacturing analysis, computational fluid dynamics tools for drag and lift computation, battery models for energy estimation, and simulation models for flight control and dynamics. AircraftVerse contains $27{,}714$ diverse air vehicle designs - the largest corpus of designs with this level of complexity. Each design comprises the following artifacts: a symbolic design tree describing topology, propulsion subsystem, battery subsystem, and other design details; a STandard for the Exchange of Product (STEP) model data; a 3D CAD design using a stereolithography (STL) file format; a 3D point cloud for the shape of the design; and evaluation results from high fidelity state-of-the-art physics models that characterize performance metrics such as maximum flight distance and hover-time. We also present baseline surrogate models that use different modalities of design representation to predict design performance metrics, which we provide as part of our dataset release. Finally, we discuss the potential impact of this dataset on the use of learning in aircraft design, and more generally, in the emerging field of deep learning for scientific design. AircraftVerse is accompanied by a datasheet as suggested in the recent literature, and it is released under Creative Commons Attribution-ShareAlike (CC BY-SA) license. The dataset with baseline models are hosted at http://doi.org/10.5281/zenodo.6525446, code at https://github.com/SRI-CSL/AircraftVerse, and the dataset description at https://uavdesignverse.onrender.com/. Adam D. Cobb, Daniel Elenius, F. Michael Heim, Brian Swenson, Sydney Whittington, James D. Walker, Ted Bapty, Joseph Hite, Karthik Ramani, Christopher McComb, Susmit Jha |
NeurIPS | 1 |
| 2023 | Decentralized Bayesian learning with Metropolis-adjusted Hamiltonian Monte Carlo
Vyacheslav Kungurtsev, Adam D. Cobb, Tara Javidi, Brian Jalaian |
Mach. Learn. | 2 |
| 2022 | Principal Component FlowsabstractNormalizing flows map an independent set of latent variables to their samples using a bijective transformation. Despite the exact correspondence between samples and latent variables, their high level relationship is not well understood. In this paper we characterize the geometric structure of flows using principal manifolds and understand the relationship between latent variables and samples using contours. We introduce a novel class of normalizing flows, called principal component flows (PCF), whose contours are its principal manifolds, and a variant for injective flows (iPCF) that is more efficient to train than regular injective flows. PCFs can be constructed using any flow architecture, are trained with a regularized maximum likelihood objective and can perform density estimation on all of their principal manifolds. In our experiments we show that PCFs and iPCFs are able to learn the principal manifolds over a variety of datasets. Additionally, we show that PCFs can perform density estimation on data that lie on a manifold with variable dimensionality, which is not possible with existing normalizing flows. Edmond Cunningham, Adam D. Cobb, Susmit Jha |
ICML | 2 |
| 2021 | Improving Differential Evolution through Bayesian Hyperparameter OptimizationabstractWe propose a novel Evolutionary Algorithm (EA) based on the Differential Evolution algorithm for solving global numerical optimization problem in real-valued continuous parameter space. The proposed MadDE algorithm leverages the power of the multiple adaptation strategy with respect to the control parameters and search mechanisms, and is tested on the benchmark functions taken from the CEC 2021 special session & competition on single-objective bound-constrained optimization. Experimental results indicate that MadDE is able to achieve superior performance on global numerical optimization problems when compared against state-of-the-art real-parameter optimizers. We also provide a hyperparameter optimization algorithm SUBHO for improving the search performance of any EA by finding an optimal set of control parameters, and demonstrate its efficacy in enhancing MadDE's performance on the same benchmark. The source code of our implementation is publicly available at https://github.com/subhodipbiswas/MadDE. Subhodip Biswas, Debanjan Saha, Shuvodeep De, Adam D. Cobb, Swagatam Das, Brian Jalaian |
CEC | 4 |
| 2021 | Automatic Acoustic Mosquito Tagging with Bayesian Neural Networks
Ivan Kiskin, Adam D. Cobb, Marianne Sinka, Kathy Willis, Stephen J. Roberts |
ECML/PKDD (4) | 2 |
| 2021 | Scaling Hamiltonian Monte Carlo inference for Bayesian neural networks with symmetric splittingabstractHamiltonian Monte Carlo (HMC) is a Markov chain Monte Carlo (MCMC) approach that exhibits favourable exploration properties in high-dimensional models such as neural networks. Unfortunately, HMC has limited use in large-data regimes and little work has explored suitable approaches that aim to preserve the entire Hamiltonian. In our work, we introduce a new symmetric integration scheme for split HMC that does not rely on stochastic gradients. We show that our new formulation is more efficient than previous approaches and is easy to implement with a single GPU. As a result, we are able to perform full HMC over common deep learning architectures using entire data sets. In addition, when we compare with stochastic gradient MCMC, we show that our method achieves better performance in both accuracy and uncertainty quantification. Our approach demonstrates HMC as a feasible option when considering inference schemes for large-scale machine learning problems. Adam D. Cobb, Brian Jalaian |
UAI | 1 |
| 2020 | Humbug Zooniverse: A Crowd-Sourced Acoustic Mosquito DatasetabstractMosquitoes are the only known vector of malaria, which leads to hundreds of thousands of deaths each year. Understanding the number and location of potential mosquito vectors is of paramount importance to aid the reduction of malaria transmission cases. In recent years, deep learning has become widely used for bioacoustic classification tasks. In order to enable further research applications in this field, we release a new dataset of mosquito audio recordings. With over a thousand contributors, we obtained 195,434 labels of two second duration, of which approximately 10 percent signify mosquito events. We present an example use of the dataset, in which we train a convolutional neural network on log-Mel features, showcasing the information content of the labels. We hope this will become a vital resource for those researching all aspects of malaria, and add to the existing audio datasets for bioacoustic detection and signal processing. Ivan Kiskin, Adam D. Cobb, Lawrence Wang, Stephen J. Roberts |
ICASSP | 2 |
| 2020 | BayesOpt Adversarial Attack
Binxin Ru, Adam D. Cobb, Arno Blaas, Yarin Gal |
ICLR | 2 |
| 2018 | Identifying Sources and Sinks in the Presence of Multiple Agents with Gaussian Process Vector CalculusabstractIn systems of multiple agents, identifying the cause of observed agent dynamics is challenging. Often, these agents operate in diverse, non-stationary environments, where models rely on hand-crafted environment-specific features to infer influential regions in the system's surroundings. To overcome the limitations of these inflexible models, we present GP-LAPLACE, a technique for locating sources and sinks from trajectories in time-varying fields. Using Gaussian processes, we jointly infer a spatio-temporal vector field, as well as canonical vector calculus operations on that field. Notably, we do this from only agent trajectories without requiring knowledge of the environment, and also obtain a metric for denoting the significance of inferred causal features in the environment by exploiting our probabilistic method. To evaluate our approach, we apply it to both synthetic and real-world GPS data, demonstrating the applicability of our technique in the presence of multiple agents, as well as its superiority over existing methods. Adam D. Cobb, Richard Everett 0001, Andrew Markham, Stephen J. Roberts |
KDD | 1 |