VLDB 2026 Research / reviewers in the wild / expert
George D. Montañez
dblp:115/6536
· DBLP profile ↗
23ranked-venue papers
8as first author
12since 2021 · last 2025
0000-0002-1333-4611ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 19 · 6 first-author · 10 since 2021Human-computer interaction and ubiquitous computing · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 2 first-author · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021Theory of computation · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Model Characterization with Inductive Orientation Vectors
Kerria Pang-Naylor, George D. Montañez |
ICAART (2) | 3 |
| 2024 | From Targets to Rewards: Continuous Target Sets in the Algorithmic Search Framework
Milo Knell, Sahil Rane, Forrest Bicker, Tiger Che, George D. Montañez |
ICAART (3) | 6 |
| 2023 | Finite-Sample Bounds for Two-Distribution Hypothesis TestsabstractWith the rapid growth of large language models, big data, and malicious online attacks, it has become increasingly important to have tools for anomaly detection that can distinguish machine from human, fair from unfair, and dangerous from safe. Prior work has shown that two-distribution (specified complexity) hypothesis tests are useful tools for such tasks, aiding in detecting bias in datasets and providing artificial agents with the ability to recognize artifacts that are likely to have been designed by humans and pose a threat. However, existing work on two-distribution hypothesis tests requires exact values for the specification function, which can often be costly or impossible to compute. In this work, we prove novel finite-sample bounds that allow for two-distribution hypothesis tests with only estimates of required quantities, such as specification function values. Significantly, the resulting bounds do not require knowledge of the true distribution, distinguishing them from traditional p-values. We apply our bounds to detect student cheating on multiple-choice tests, as an example where the exact specification function is unknown. We additionally apply our results to detect representational bias in machine-learning datasets and provide artificial agents with intention perception, showing that our results are consistent with prior work despite only requiring a finite sample of the space. Finally, we discuss additional applications and provide guidance for those applying these bounds to their own work. Cynthia Hom, William Yik, George D. Montañez |
DSAA | 3 |
| 2022 | Identifying Bias in Data Using Two-Distribution Hypothesis TestsabstractAs machine learning models become more widely used in important decision-making processes, the need for identifying and mitigating potential sources of bias has increased substantially. Using two-distribution (specified complexity) hypothesis tests, we identify biases in training data with respect to proposed distributions and without the need to train a model, distinguishing our methods from common output-based fairness tests. Furthermore, our methods allow us to return a "closest plausible explanation" for a given dataset, potentially revealing underlying biases in the processes that generated them. We also show that a binomial variation of this hypothesis test could be used to identify bias in certain directions, or towards certain outcomes, and again return a closest plausible explanation. The benefits of this binomial variation are compared with other hypothesis tests, including the exact binomial. Lastly, potential industrial applications of our methods are shown using two real-world datasets. William Yik, Limnanthes Serafini, Timothy Lindsey, George D. Montañez |
AIES | 4 |
| 2022 | Vectorization of Bias in Machine Learning Algorithms
Sophie Bekerman, Lily Lin, George D. Montañez |
ICAART (2) | 4 |
| 2022 | The Gopher Grounds: Testing the Link between Structure and Function in Simple Machines
Anshul Kamath, Nick Grisanti, George D. Montañez |
ICAART (2) | 4 |
| 2022 | Bounding Generalization Error Through Bias and CapacityabstractWe derive generalization bounds on learning algorithms through algorithm capacity and a vector representation of inductive bias. Leveraging the algorithmic search framework, a formalism for casting machine learning as a type of search, we present a unified interpretation of the upper bounds of generalization error in terms of a vector representation of bias and the mutual information between the hypothesis and the dataset. Ramya Ramalingam, Nicolas A. Espinosa Dice, Megan L. Kaye, George D. Montañez |
IJCNN | 4 |
| 2021 | The Predator's Purpose: Intention Perception in Simulated Agent EnvironmentsabstractWe evaluate the benefits of intention perception, the ability of an agent to perceive the intentions and plans of others, in improving a software agent's survival likelihood in a simulated virtual environment. To model intention perception, we set up a multi-agent predator and prey model, where the prey agents search for food and the predator agents seek to eat the prey. We then analyze the difference in average survival rates between prey with intention perception-knowledge of which predators are targeting them-and those without. We find that intention perception provides significant survival advantages in almost all cases tested, agreeing with other recent studies investigating intention perception in adversarial situations and environmental danger assessment. Amani R. Maina-Kilaas, Cynthia Hom, Kevin Ginta, George D. Montañez |
CEC | 4 |
| 2021 | The Hero's Dilemma: Survival Advantages of Intention Perception in Virtual Agent GamesabstractConjecturing that an agent's ability to perceive the intentions of others can increase its chances of survival, we introduce a simple game, the Hero's Dilemma, which simulates interactions between two virtual agents to investigate whether an agent's ability to detect the intentional stance of a second agent provides a measurable survival advantage. We test whether agents able to make decisions based on the perceived intention of an adversarial agent have advantages over agents without such perception, but who instead rely on a variety of different game-playing strategies. In the game, an agent must decide whether to remain hidden or attack an often more powerful agent based on the perceived intention of the other agent. We compare the survival rates of agents with and without intention perception, and find that intention perception provides significant survival advantages and is the most successful strategy in the majority of situations tested. Amani R. Maina-Kilaas, George D. Montañez, Cynthia Hom, Kevin Ginta, Cindy Lay |
CoG | 2 |
| 2021 | A Probabilistic Theory of Abductive Reasoning
Nicolas A. Espinosa Dice, Megan L. Kaye, Hana Ahmed, George D. Montañez |
ICAART (2) | 4 |
| 2021 | The Gopher's Gambit: Survival Advantages of Artifact-based Intention Perception
Cynthia Hom, Amani R. Maina-Kilaas, Kevin Ginta, Cindy Lay, George D. Montañez |
ICAART (1) | 5 |
| 2021 | Hyperparameter Choice as Search Bias in AlphaZeroabstractThe AlphaZero algorithm has achieved remarkable success in a variety of sequential, perfect information games including Go, Shogi and chess. To better understand how AlphaZero works and leverage that understanding when deploying the system, we study the properties of the α hyperparameter that governs exploration noise in AlphaZero’s search, the only hyperparameter the system’s creators modified when moving among the three aforementioned games. First, we build a formal intuition for its behavior on a simple example meant to isolate the influence of the hyperparameter. Then, by comparing performance of AlphaZero agents with different α values on the game Connect 4, we show that the performance of AlphaZero improves considerably with a good choice of α. This all highlights the importance of α as an interpretable hyperparameter which allows for cross-game tuning that more opaque hyperparameters like model architecture may not. Eric M. Weiner, George D. Montañez, Aaron Trujillo, Abtin Molavi |
SMC | 2 |
| 2020 | The Bias-Expressivity Trade-offabstractLearning algorithms need bias to generalize and perform better than random guessing. We examine the flexibility (expressivity) of biased algorithms. An expressive algorithm can adapt to changing training data, altering its outcome based on changes in its input. We measure expressivity by using an information-theoretic notion of entropy on algorithm outcome distributions, demonstrating a trade-off between bias and expressivity. To the degree an algorithm is biased is the degree to which it can outperform uniform random sampling, but is also the degree to which is becomes inflexible. We derive bounds relating bias to expressivity, proving the necessary trade-offs inherent in trying to create strongly performing yet flexible algorithms. Julius Lauw, Dominique Macias, Akshay Trikha, Julia Vendemiatti, George D. Montañez |
ICAART (2) | 5 |
| 2020 | Decomposable Probability-of-Success Metrics in Algorithmic SearchabstractPrevious studies have used a specific success metric within an algorithmic search framework to prove machine learning impossibility results. However, this specific success metric prevents us from applying these results on other forms of machine learning, e.g. transfer learning. We define decomposable metrics as a category of success metrics for search problems which can be expressed as a linear operation on a probability distribution to solve this issue. Using an arbitrary decomposable metric to measure the success of a search, we demonstrate theorems which bound success in various ways, generalizing several existing results in the literature. Tyler Sam, Jake Williams, Abel Tadesse, Huey Sun, George D. Montañez |
ICAART (2) | 5 |
| 2020 | The Labeling Distribution Matrix (LDM): A Tool for Estimating Machine Learning Algorithm CapacityabstractAlgorithm performance in supervised learning is a combination of memorization, generalization, and luck. By estimating how much information an algorithm can memorize from a dataset, we can set a lower bound on the amount of performance due to other factors such as generalization and luck. With this goal in mind, we introduce the Labeling Distribution Matrix (LDM) as a tool for estimating the capacity of learning algorithms. The method attempts to characterize the diversity of possible outputs by an algorithm for different training datasets, using this to measure algorithm flexibility and responsiveness to data. We test the method on several supervised learning algorithms, and find that while the results are not conclusive, the LDM does allow us to gain potentially valuable insight into the prediction behavior of algorithms. We also introduce the Label Recorder as an additional tool for estimating algorithm capacity, with more promising initial results. Pedro Sandoval Segura, Julius Lauw, Daniel Bashir, Kinjal Shah, Sonia Sehra, Dominique Macias, George D. Montañez |
ICAART (2) | 7 |
| 2017 | The LICORS cabinet: Nonparametric light cone methods for spatio-temporal modelingabstractSpatio-temporal data is intrinsically high dimensional, so unsupervised modeling is only feasible if we can exploit structure in the process. When the dynamics are local in both space and time, this structure can be exploited by splitting the global field into many lower-dimensional “light cones”. We review light cone decompositions for predictive state reconstruction, introducing three simple light cone algorithms. These methods allow for tractable inference of spatio-temporal data, such as full-frame video. The algorithms make few assumptions on the underlying process yet have good predictive performance and can provide distributions over spatio-temporal data, enabling sophisticated probabilistic inference. George D. Montañez, Cosma Rohilla Shalizi |
IJCNN | 1 |
| 2017 | The famine of forte: Few search problems greatly favor your algorithmabstractCasting machine learning as a type of search, we demonstrate that the proportion of problems that are favorable for a fixed algorithm is strictly bounded, such that no single algorithm can perform well over a large fraction of them. If an algorithm greatly excels on a class of problems (e.g., convex problems), that class must necessarily be small. We give an upper bound on the expected performance for a search algorithm as a function of the mutual information between the target and the information resource (e.g., training dataset), proving the importance of certain types of dependence for machine learning. Lastly, given that the expected per-query probability of success for an algorithm is mathematically equivalent to a single-query probability of success under a distribution (called a search strategy), we prove that the proportion of favorable strategies is also strictly bounded. Thus, whether one holds fixed the search algorithm and considers all possible problems or one fixes the search problem and looks at all possible search strategies, favorable matches are exceedingly rare. The forte (strength) of any algorithm is quantifiably restricted. George D. Montañez |
SMC | 1 |
| 2016 | Detecting Intelligence - The Turing Test and Other Design Detection Methodologiesabstract“Can machines think?” When faced with this “meaningless” question, Alan Turing suggested we ask a different, more precise question: can a machine reliably fool a human interviewer into believing the machine is human? To answer this question, Turing outlined what came to be known as the Turing Test for artificial intelligence, namely, an imitation game where machines and humans interacted from remote locations and human judges had to distinguish between the human and machine participants. According to the test, machines that consistently fool human judges are to be viewed as intelligent. While popular culture champions the Turing Test as a scientific procedure for detecting artificial intelligence, doing so raises significant issues. First, a simple argument establishes the equivalence of the Turing Test to intelligent design methodology in several fundamental respects. Constructed with similar goals, shared assumptions and identical observational models, both projects attempt to detect intelligent agents through the examination of generated artifacts of uncertain origin. Second, if the Turing Test rests on scientifically defensible assumptions then design inferences become possible and cannot, in general, be wholly unscientific. Third, if passing the Turing Test reliably indicates intelligence, this implies the likely existence of a designing intelligence in nature. George D. Montañez |
ICAART (2) | 1 |
| 2015 | Inertial Hidden Markov Models: Modeling Change in Multivariate Time SeriesabstractFaced with the problem of characterizing systematic changes in multivariate time series in an unsupervised manner, we derive and test two methods of regularizing hidden Markov models for this task. Regularization on state transitions provides smooth transitioning among states, such that the sequences are split into broad, contiguous segments. Our methods are compared with a recent hierarchical Dirichlet process hidden Markov model (HDP-HMM) and a baseline standard hidden Markov model, of which the former suffers from poor performance on moderate-dimensional data and sensitivity to parameter settings, while the latter suffers from rapid state transitioning, over-segmentation and poor performance on a segmentation task involving human activity accelerometer data from the UCI Repository. The regularized methods developed here are able to perfectly characterize change of behavior in the human activity data for roughly half of the real-data test cases, with accuracy of 94% and low variation of information. In contrast to the HDP-HMM, our methods provide simple, drop-in replacements for standard hidden Markov model update rules, allowing standard expectation maximization (EM) algorithms to be used for learning. George D. Montañez, Saeed Amizadeh, Nikolay Laptev |
AAAI | 1 |
| 2014 | Cross-Device SearchabstractOwnership and use of multiple devices such as desktop computers, smartphones, and tablets is increasing rapidly. Search is popular and people often perform search tasks that span device boundaries. Understanding how these devices are used and how people transition between them during information seeking is essential in developing search support for a multi-device world. In this paper, we study search across devices and propose models to predict aspects of cross-device search transitions. We characterize multi-device search across four device types, including aspects of search behavior on each device (e.g., topics of interest) and characteristics of device transitions. Building on the characterization, we learn models to predict various aspects of cross-device search, including the next device used for search. This enables many applications. For example, accurately forecasting the device used for the next query lets search engines proactively retrieve device-appropriate content (e.g., short documents for smartphones), while knowledge of the current device combined with device-specific topical interest models may assist in better query-sense disambiguation. %to help the searcher once they transition to the target device. George D. Montañez, Ryen W. White |
CIKM | 1 |
| 2013 | Information transmission through genetic algorithm fitness mapsabstractTo bound the amount of information transmitted from a fitness map to a genetic algorithm population, we use a method suggested by Abu-Mostafa et al. [1] for measuring the information storage capacity of general forms of memory and represent the genetic algorithm as a communication channel. Our results show that a number of bits linear in the size of the search space can be stored in a fitness map, but on average only a logarithmic number of bits can be stored within a genetic algorithm population of bounded size and finite precision representation. Our results place an upper bound on the rate at which information can be transmitted through, or generated by and later extracted from, a genetic algorithm under fairly general conditions. George D. Montañez |
IEEE Congress on Evolutionary Computation | 1 |
| 2013 | Bounding the number of favorable functions in stochastic searchabstractAccording to the No Free Lunch theorems for search, when uniformly averaged over all possible search functions, every search algorithm has identical search performance for a wide variety of common performance metrics [1], [2], [3], [4]. Differences in performance can arise, however, between two algorithms when performance is measured over non-closed under permutation sets of functions, such as sets consisting of a single function. Using uniform random sampling with replacement as a baseline, we ask how many functions exist such that a search algorithm has better expected performance than random sampling. We define favorable functions as those that allow an algorithm to locate a search target with higher probability than uniform random sampling with replacement, and we bound the proportion of favorable functions for stochastic search methods, including genetic algorithms. Using active information [5] as our divergence measure, we demonstrate that no more than 2-bof all functions are favorable by b or more bits, for b ≥ 2 and reasonably sized search spaces (n ≥ 19). Thus, the proportion of functions for which an algorithm performs relatively well by a moderate degree is strictly bounded. Our results can be viewed as statement of information conservation [6], [7], [1], [8], [5], since identifying a favorable function of b or more bits requires at least b bits of information, under the conditions given. George D. Montañez |
IEEE Congress on Evolutionary Computation | 1 |
| 2012 | Assessing reliability of protein-protein interactions by gene ontology integrationabstractRecent advances in genome-wide identification of protein-protein interactions (PPIs) have produced an abundance of interaction data which give an insight into functional associations among proteins. However, it is known that the PPI datasets determined by high-throughput experiments or inferred by computational methods include an extremely large number of false positives. Using Gene Ontology (GO) and its annotations, we assess reliability of the PPIs by considering the semantic similarity of interacting proteins. Protein pairs with high semantic similarity are considered highly likely to share common functions, and therefore, are more likely to interact. We analyze the performance of existing semantic similarity measures in terms of functional consistency and propose a combined method that achieves improved performance over existing methods. The semantic similarity measures are applied to identify false positive PPIs. The classification results show that the combined hybrid method has higher accuracy than the other existing measures. Furthermore, the combined hybrid classifier predicts that 59.6% of the S. cerevisiae PPIs from the BioGRID database are false positives. George D. Montañez, Young-Rae Cho |
CIBCB | 1 |