EDBT 2026 Demo / reviewers in the wild / expert
Bert Huang
dblp:93/10793
· DBLP profile ↗
37ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0002-8548-7246ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 32 · 2 first-author · 4 since 2021Databases, data management, data science and information retrieval · 8 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6Applied, interdisciplinary, general and emerging computing · 4Human-computer interaction and ubiquitous computing · 2Systems, architecture and hardware · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
16 papers |
Probabilistic and Bayesian machine learning · 38% Generative modeling · 17% Learning theory · 11% | |
| Databases, data mining, and information retrieval
4 papers |
Data mining · 43% Recommender systems · 32% Web and social media mining · 14% | |
| Interdisciplinary, comprehensive, and emerging computing
4 papers |
Computational social science and digital humanities · 55% Computing education · 26% Energy systems and smart grids · 19% |
Topics — the 30 heaviest of 45, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Probabilistic and Bayesian machine learning › structured models
graphical models |
0.9 | 4 | 2019 | Block Belief Propagation for Parameter Learning in Markov Random Fields · AAAI 2019 Hinge-Loss Markov Random Fields and Probabilistic Soft Logic · J. Mach. Learn. Res. 2017 Paired-Dual Learning for Fast Training of Latent Variable Hinge-Loss MRFs · ICML 2015 |
Machine learning › Learning paradigms
weakly supervised learning |
0.9 | 2 | 2021 | A General Framework for Adversarial Label Learning · J. Mach. Learn. Res. 2021 Adversarial Label Learning · AAAI 2019 |
Machine learning › Generative modeling
normalizing flow |
0.9 | 2 | 2020 | Woodbury Transformations for Deep Generative Flows · NeurIPS 2020 Structured Output Learning with Conditional Generative Flows · AAAI 2020 |
Machine learning › Probabilistic and Bayesian machine learning
structured prediction |
0.9 | 3 | 2020 | Structured Output Learning with Conditional Generative Flows · AAAI 2020 Stability and Generalization in Structured Prediction · J. Mach. Learn. Res. 2016 Collective Stability in Structured Prediction: Generalization from One Example · ICML (3) 2013 |
Machine learning › Learning theory
generalization bounds |
0.8 | 3 | 2019 | Adversarial Label Learning · AAAI 2019 Stability and Generalization in Structured Prediction · J. Mach. Learn. Res. 2016 Collective Stability in Structured Prediction: Generalization from One Example · ICML (3) 2013 |
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models
markov random field |
0.7 | 3 | 2019 | Block Belief Propagation for Parameter Learning in Markov Random Fields · AAAI 2019 Paired-Dual Learning for Fast Training of Latent Variable Hinge-Loss MRFs · ICML 2015 Stability and Generalization in Structured Prediction · J. Mach. Learn. Res. 2016 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › probabilistic reasoning
hinge-loss markov random fields |
0.5 | 2 | 2017 | Hinge-Loss Markov Random Fields and Probabilistic Soft Logic · J. Mach. Learn. Res. 2017 Paired-Dual Learning for Fast Training of Latent Variable Hinge-Loss MRFs · ICML 2015 |
Machine learning › Optimization for machine learning
game-theoretic optimization |
0.5 | 1 | 2021 | A General Framework for Adversarial Label Learning · J. Mach. Learn. Res. 2021 |
Knowledge, reasoning and agents › Multi-agent systems › game theory
zero-sum games |
0.5 | 1 | 2021 | A General Framework for Adversarial Label Learning · J. Mach. Learn. Res. 2021 |
Machine learning › Generative modeling › normalizing flow
conditional normalizing flow |
0.4 | 1 | 2020 | Structured Output Learning with Conditional Generative Flows · AAAI 2020 |
Machine learning › Generative modeling › normalizing flow
invertible transformation |
0.4 | 1 | 2020 | Woodbury Transformations for Deep Generative Flows · NeurIPS 2020 |
Machine learning › Probabilistic and Bayesian machine learning › structured prediction
structured output learning |
0.4 | 1 | 2020 | Structured Output Learning with Conditional Generative Flows · AAAI 2020 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
approximate inference |
0.4 | 1 | 2019 | Block Belief Propagation for Parameter Learning in Markov Random Fields · AAAI 2019 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
belief propagation |
0.4 | 1 | 2019 | Block Belief Propagation for Parameter Learning in Markov Random Fields · AAAI 2019 |
Machine learning › Trustworthy machine learning
fairness |
0.3 | 1 | 2017 | Beyond Parity: Fairness Objectives for Collaborative Filtering · NIPS 2017 |
Machine learning › Probabilistic and Bayesian machine learning
probabilistic programming |
0.3 | 1 | 2017 | Hinge-Loss Markov Random Fields and Probabilistic Soft Logic · J. Mach. Learn. Res. 2017 |
Knowledge, reasoning and agents › Knowledge representation and reasoning › probabilistic reasoning › probabilistic logic
probabilistic soft logic |
0.3 | 1 | 2017 | Hinge-Loss Markov Random Fields and Probabilistic Soft Logic · J. Mach. Learn. Res. 2017 |
Knowledge, reasoning and agents › Knowledge representation and reasoning
statistical relational learning |
0.3 | 1 | 2017 | Hinge-Loss Markov Random Fields and Probabilistic Soft Logic · J. Mach. Learn. Res. 2017 |
Machine learning › Probabilistic and Bayesian machine learning
structured models |
0.3 | 1 | 2017 | Hinge-Loss Markov Random Fields and Probabilistic Soft Logic · J. Mach. Learn. Res. 2017 |
Recommender systems
collaborative filtering |
0.3 | 1 | 2017 | Beyond Parity: Fairness Objectives for Collaborative Filtering · NIPS 2017 |
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model |
0.3 | 2 | 2015 | Paired-Dual Learning for Fast Training of Latent Variable Hinge-Loss MRFs · ICML 2015 Learning Latent Engagement Patterns of Students in Online Courses · AAAI 2014 |
Machine learning › Learning theory › generalization bounds
PAC-Bayes bounds |
0.2 | 1 | 2016 | Stability and Generalization in Structured Prediction · J. Mach. Learn. Res. 2016 |
Machine learning › Learning theory › generalization
stability and generalization |
0.2 | 1 | 2016 | Stability and Generalization in Structured Prediction · J. Mach. Learn. Res. 2016 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference
marginal inference |
0.2 | 1 | 2015 | The Benefits of Learning with Strongly Convex Approximate Inference · ICML 2015 |
Natural language and speech › Information extraction and text analysis
stance detection |
0.2 | 1 | 2015 | Joint Models of Disagreement and Stance in Online Debate · ACL (1) 2015 |
Machine learning › Probabilistic and Bayesian machine learning › probabilistic inference › approximate inference
variational inference |
0.2 | 1 | 2015 | The Benefits of Learning with Strongly Convex Approximate Inference · ICML 2015 |
Data mining › predictive modeling
event prediction |
0.2 | 1 | 2014 | 'Beating the news' with EMBERS: forecasting civil unrest using open source indicators · KDD 2014 |
Data mining
predictive modeling |
0.2 | 1 | 2014 | 'Beating the news' with EMBERS: forecasting civil unrest using open source indicators · KDD 2014 |
Machine learning › Learning theory › generalization bounds
uniform convergence |
0.2 | 1 | 2013 | Collective Stability in Structured Prediction: Generalization from One Example · ICML (3) 2013 |
Machine learning › Trustworthy machine learning › model monitoring
failure prediction |
0.1 | 1 | 2012 | Machine Learning for theNew York City Power Grid · IEEE Trans. Pattern Anal. Mach. Intell. 2012 |
Methods — techniques the papers use, named apart from their topics
parameter learning · 0.7zero-sum game · 0.5primal-dual subgradient descent · 0.5non-zero sum game optimization · 0.5time normalization · 0.4probabilistic soft logic · 0.4key phrase learning · 0.4woodbury matrix identity · 0.4variational inference · 0.4sylvester's determinant identity · 0.4conditional glow · 0.4weak supervision · 0.4projected primal-dual subgradient descent · 0.4suppression engine · 0.4data fusion · 0.4feature fusion · 0.3SVM ensemble · 0.3HOG · 0.3
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Weakly supervised label learning flows
You Lu 0003, Wenzhuo Song, Chidubem Arachie, Bert Huang |
Neural Networks | 4 |
| 2022 | Firebolt: Weak Supervision Under Weaker AssumptionsabstractModern machine learning demands a large amount of training data. Weak supervision is a promising approach to meet this demand. It aggregates multiple labeling functions (LFs)–noisy, user-provided labeling heuristics—to rapidly and cheaply curate probabilistic labels for large-scale unlabeled data. However, standard assumptions in weak supervision—such as user-specified class balance, similar accuracy of an LF in classifying different classes, and full knowledge of LF dependency at inference time—might be undesirable in practice. In response, we present Firebolt, a new weak supervision framework that seeks to operate under weaker assumptions. In particular, Firebolt learns the class balance and class-specific accuracy of LFs jointly from unlabeled data. It carries out inference in an efficient and interpretable manner. We analyze the parameter estimation error of Firebolt and characterize its impact on downstream model performance. Furthermore, we show that on five publicly available datasets, Firebolt outperforms a state-of-the-art weak supervision method by up to 5.8 points in AUC. We also provide a case study in the production setting of a tech company, where a Firebolt-supervised model outperforms the existing weakly-supervised production model by 1.3 points in AUC and speedup label model training and inference from one hour to three minutes. Zhaobin Kuang, Chidubem Arachie, Bangyong Liang, Pradyumna Narayana, Giulia DeSalvo, Michael S. Quinn, Bert Huang, Geoffrey Downs |
AISTATS | 7 |
| 2021 | Personalized Regularization Learning for Fairer Matrix Factorization
Sirui Yao, Bert Huang |
PAKDD (2) | 2 |
| 2021 | Constrained labeling for weakly supervised learningabstractCuration of large fully supervised datasets has become one of the major roadblocks for machine learning. Weak supervision provides an alternative to supervised learning by training with cheap, noisy, and possibly correlated labeling functions from varying sources. The key challenge in weakly supervised learning is combining the different weak supervision signals while navigating misleading correlations in their errors. In this paper, we propose a simple data-free approach for combining weak supervision signals by defining a constrained space for the possible labels of the weak signals and training with a random labeling within this constrained space. Our method is efficient and stable, converging after a few iterations of gradient descent. We prove theoretical conditions under which the worst-case error of the randomized label decreases with the rank of the linear constraints. We show experimentally that our method outperforms other weak supervision methods on various text- and image-classification tasks. Chidubem Arachie, Bert Huang |
UAI | 2 |
| 2021 | A General Framework for Adversarial Label LearningabstractWe consider the task of training classifiers without fully labeled data. We propose a weakly supervised method---adversarial label learning---that trains classifiers to perform well when noisy and possibly correlated labels are provided. Our framework allows users to provide different weak labels and multiple constraints on these labels. Our model then attempts to learn parameters for the data by solving a zero-sum game for the binary problems and a non-zero sum game optimization for multi-class problems. The game is between an adversary that chooses labels for the data and a model that minimizes the error made by the adversarial labels. The weak supervision constrains what labels the adversary can choose. The method therefore minimizes an upper bound of the classifier's error rate using projected primal-dual subgradient descent. Minimizing this bound protects against bias and dependencies in the weak supervision. We first show the performance of our framework on binary classification tasks then we extend our algorithm to show its performance on multiclass datasets. Our experiments show that our method can train without labels and outperforms other approaches for weakly supervised learning. Chidubem Arachie, Bert Huang |
J. Mach. Learn. Res. | 2 |
| 2020 | Structured Output Learning with Conditional Generative FlowsabstractTraditional structured prediction models try to learn the conditional likelihood, i.e., p(y|x), to capture the relationship between the structured output y and the input features x. For many models, computing the likelihood is intractable. These models are therefore hard to train, requiring the use of surrogate objectives or variational inference to approximate likelihood. In this paper, we propose conditional Glow (c-Glow), a conditional generative flow for structured output learning. C-Glow benefits from the ability of flow-based models to compute p(y|x exactly and efficiently. Learning with c-Glow does not require a surrogate objective or performing inference during training. Once trained, we can directly and efficiently generate conditional samples. We develop a sample-based prediction method, which can use this advantage to do efficient and effective inference. In our experiments, we test c-Glow on five different tasks. C-Glow outperforms the state-of-the-art baselines in some tasks and predicts comparable outputs in the other tasks. The results show that c-Glow is versatile and is applicable to many different structured prediction problems. You Lu 0003, Bert Huang |
AAAI | 2 |
| 2020 | Woodbury Transformations for Deep Generative FlowsabstractNormalizing flows are deep generative models that allow efficient likelihood calculation and sampling. The core requirement for this advantage is that they are constructed using functions that can be efficiently inverted and for which the determinant of the function's Jacobian can be efficiently computed. Researchers have introduced various such flow operations, but few of these allow rich interactions among variables without incurring significant computational costs. In this paper, we introduce Woodbury transformations, which achieve efficient invertibility via the Woodbury matrix identity and efficient determinant calculation via Sylvester's determinant identity. In contrast with other operations used in state-of-the-art normalizing flows, Woodbury transformations enable (1) high-dimensional interactions, (2) efficient sampling, and (3) efficient likelihood evaluation. Other similar operations, such as 1x1 convolutions, emerging convolutions, or periodic convolutions allow at most two of these three advantages. In our experiments on multiple image datasets, we find that Woodbury transformations allow learning of higher-likelihood models than other flow architectures while still enjoying their efficiency advantages. You Lu 0003, Bert Huang |
NeurIPS | 2 |
| 2020 | Attention-Based Graph Evolution
Shuangfei Fan, Bert Huang |
PAKDD (1) | 2 |
| 2019 | Adversarial Label LearningabstractWe consider the task of training classifiers without labels. We propose a weakly supervised method—adversarial label learning—that trains classifiers to perform well against an adversary that chooses labels for training data. The weak supervision constrains what labels the adversary can choose. The method therefore minimizes an upper bound of the classifier’s error rate using projected primal-dual subgradient descent. Minimizing this bound protects against bias and dependencies in the weak supervision. Experiments on real datasets show that our method can train without labels and outperforms other approaches for weakly supervised learning. Chidubem Arachie, Bert Huang |
AAAI | 2 |
| 2019 | Block Belief Propagation for Parameter Learning in Markov Random Fields
You Lu 0003, Zhiyuan Liu 0003, Bert Huang |
AAAI | 3 |
| 2019 | Recurrent collective classification
Shuangfei Fan, Bert Huang |
Knowl. Inf. Syst. | 2 |
| 2018 | Weakly Supervised Cyberbullying Detection Using Co-Trained Ensembles of Embedding ModelsabstractSocial media has become an inevitable part of individuals personal and business lives. Its benefits come with various negative consequences. One major concern is the prevalence of detrimental online behavior on social media, such as online harassment and cyberbullying. In this study, we aim to address the computational challenges associated with harassment detection in social media by developing a machine learning framework with three distinguishing characteristics. (1) It uses minimal supervision in the form of expert-provided key phrases that are indicative of bullying or non-bullying. (2) It detects harassment with an ensemble of two learners that co-train one another; one learner examines the language content in the message, and the other learner considers the social structure. (3) It incorporates distributed word and graph-node representations by training nonlinear deep models. The model is trained by optimizing an objective function that balances a co-training loss with a weak-supervision loss. We evaluate the effectiveness of our approach using post-hoc, crowdsourced annotation of Twitter, Ask.fm, and Instagram data, finding that our deep ensembles outperform previous non-deep methods for weakly supervised harassment detection. Elaheh Raisi, Bert Huang |
ASONAM | 2 |
| 2018 | Best-Choice Edge Grafting for Efficient Structure Learning of Markov Random FieldsabstractIncremental methods for structure learning of pairwise Markov random fields (MRFs), such as grafting, improve scalability by avoiding inference over the entire feature space in each optimization step. Instead, inference is performed over an incrementally grown active set of features. In this paper, we address key computational bottlenecks that current incremental techniques still suffer by introducing best-choice edge grafting, an incremental, structured method that activates edges as groups of features in a streaming setting. The method uses a reservoir of edges that satisfy an activation condition, approximating the search for the optimal edge to activate. It also reorganizes the search space using search-history and structure heuristics. Experiments show a significant speedup for structure learning and a controllable trade-off between the speed and quality of learning. Walid Chaabene, Bert Huang |
IEEE BigData | 2 |
| 2018 | Sparse-Matrix Belief Propagation
Reid Bixler, Bert Huang |
UAI | 2 |
| 2017 | Cyberbullying Detection with Weakly Supervised Machine LearningabstractDetrimental online behavior such as harassment and cyberbullying is becoming a serious, large-scale problem damaging people's lives. This phenomenon is creating a need for automated, data-driven techniques for analyzing and detecting such behaviors. We propose a machine learning method for simultaneously inferring user roles in harassment-based bullying and new vocabulary indicators of bullying. The learning algorithm considers social structure and infers which users tend to bully and which tend to be victimized. To address the elusive nature of cyberbullying, the learning algorithm only requires weak supervision. Experts provide a small seed vocabulary of bullying indicators, and the algorithm uses a large, unlabeled corpus of social media interactions to extract bullying roles of users and additional vocabulary indicators of bullying. The model estimates whether each social interaction is bullying based on who participates and based on what language is used, and it tries to maximize the agreement between these estimates, i.e., participant-vocabulary consistency (PVC). We evaluate PVC on three social media data sets, demonstrating quantitatively and qualitatively its effectiveness in cyberbullying detection. Elaheh Raisi, Bert Huang |
ASONAM | 2 |
| 2017 | Beyond Parity: Fairness Objectives for Collaborative FilteringabstractWe study fairness in collaborative-filtering recommender systems, which are sensitive to discrimination that exists in historical data. Biased data can lead collaborative-filtering methods to make unfair predictions for users from minority groups. We identify the insufficiency of existing fairness metrics and propose four new metrics that address different forms of unfairness. These fairness metrics can be optimized by adding fairness terms to the learning objective. Experiments on synthetic and real data show that our new metrics can better measure fairness than the baseline, and that the fairness objectives effectively help reduce unfairness. Sirui Yao, Bert Huang |
NIPS | 2 |
| 2017 | Hinge-Loss Markov Random Fields and Probabilistic Soft LogicabstractA fundamental challenge in developing high-impact machine learning technologies is balancing the need to model rich, structured domains with the ability to scale to big data. Many important problem areas are both richly structured and large scale, from social and biological networks, to knowledge graphs and the Web, to images, video, and natural language. In this paper, we introduce two new formalisms for modeling structured data, and show that they can both capture rich structure and scale to big data. The first, hinge-loss Markov random fields (HL-MRFs), is a new kind of probabilistic graphical model that generalizes different approaches to convex inference. We unite three approaches from the randomized algorithms, probabilistic graphical models, and fuzzy logic communities, showing that all three lead to the same inference objective. We then define HL- MRFs by generalizing this unified objective. The second new formalism, probabilistic soft logic (PSL), is a probabilistic programming language that makes HL-MRFs easy to define using a syntax based on first-order logic. We introduce an algorithm for inferring most-probable variable assignments (MAP inference) that is much more scalable than general-purpose convex optimization methods, because it uses message passing to take advantage of sparse dependency structures. We then show how to learn the parameters of HL-MRFs. The learned HL-MRFs are as accurate as analogous discrete models, but much more scalable. Together, these algorithms enable HL-MRFs and PSL to model rich, structured data at scales not previously possible. Stephen H. Bach, Matthias Broecheler, Bert Huang, Lise Getoor |
J. Mach. Learn. Res. | 3 |
| 2016 | Stability and Generalization in Structured PredictionabstractStructured prediction models have been found to learn effectively from a few large examples--- sometimes even just one. Despite empirical evidence, canonical learning theory cannot guarantee generalization in this setting because the error bounds decrease as a function of the number of examples. We therefore propose new PAC-Bayesian generalization bounds for structured prediction that decrease as a function of both the number of examples and the size of each example. Our analysis hinges on the stability of joint inference and the smoothness of the data distribution. We apply our bounds to several common learning scenarios, including max-margin and soft-max training of Markov random fields. Under certain conditions, the resulting error bounds can be far more optimistic than previous results and can even guarantee generalization from a single large example. Ben London 0001, Bert Huang, Lise Getoor |
J. Mach. Learn. Res. | 2 |
| 2015 | Planned Protest Modeling in News and Social MediaabstractCivil unrest (protests, strikes, and “occupy” events) is a common occurrence in both democracies and authoritarian regimes. The study of civil unrest is a key topic for political scientists as it helps capture an important mechanism by which citizenry express themselves. In countries where civil unrest is lawful, qualitative analysis has revealed that more than 75% of the protests are planned, organized, and/or announced in advance; therefore detecting future time mentions in relevant news and social media is a direct way to develop a protest forecasting system. We develop such a system in this paper, using a combination of key phrase learning to identify what to look for, probabilistic soft logic to reason about location occurrences in extracted results, and time normalization to resolve future tense mentions. We illustrate the application of our system to 10 countries in Latin America, viz. Argentina, Brazil, Chile, Colombia, Ecuador, El Salvador, Mexico, Paraguay, Uruguay, and Venezuela. Results demonstrate our successes in capturing significant societal unrest in these countries with an average lead time of 4.08 days. We also study the selective superiorities of news media versus social media (Twitter, Facebook) to identify relevant tradeoffs. Sathappan Muthiah, Bert Huang, Jaime Arredondo, David Mares, Lise Getoor, Graham Katz, Naren Ramakrishnan |
AAAI | 2 |
| 2015 | Joint Models of Disagreement and Stance in Online DebateabstractDhanya Sridhar, James Foulds, Bert Huang, Lise Getoor, Marilyn Walker. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Dhanya Sridhar, James R. Foulds, Bert Huang, Lise Getoor, Marilyn A. Walker |
ACL (1) | 3 |
| 2015 | Unifying Local Consistency and MAX SAT Relaxations for Scalable Inference with Rounding GuaranteesabstractWe prove the equivalence of first-order local consistency relaxations and the MAX SAT relaxation of Goemans and Williamson (1994) for a class of MRFs we refer to as logical MRFs. This allows us to combine the advantages of each into a single MAP inference technique: solving the local consistency relaxation with any of a number of highly scalable message-passing algorithms, and then obtaining a high-quality discrete solution via a guaranteed rounding procedure when the relaxation is not tight. Logical MRFs are a general class of models that can incorporate many common dependencies, such as logical implications and mixtures of supermodular and submodular potentials. They can be used for many structured prediction tasks, including natural language processing, computer vision, and computational social science. We show that our new inference technique can improve solution quality by as much as 20% without sacrificing speed on problems with over one million dependencies. Stephen H. Bach, Bert Huang, Lise Getoor |
AISTATS | 2 |
| 2015 | Paired-Dual Learning for Fast Training of Latent Variable Hinge-Loss MRFsabstractLatent variables allow probabilistic graphical models to capture nuance and structure in important domains such as network science, natural language processing, and computer vision. Naive approaches to learning such complex models can be prohibitively expensive—because they require repeated inferences to update beliefs about latent variables—so lifting this restriction for useful classes of models is an important problem. Hinge-loss Markov random fields (HL-MRFs) are graphical models that allow highly scalable inference and learning in structured domains, in part by representing structured problems with continuous variables. However, this representation leads to challenges when learning with latent variables. We introduce paired-dual learning, a framework that greatly speeds up training by using tractable entropy surrogates and avoiding repeated inferences. Paired-dual learning optimizes an objective with a pair of dual inference problems. This allows fast, joint optimization of parameters and dual variables. We evaluate on social-group detection, trust prediction in social networks, and image reconstruction, finding that paired-dual learning trains models as accurate as those trained by traditional methods in much less time, often before traditional methods make even a single parameter update. Stephen H. Bach, Bert Huang, Jordan L. Boyd-Graber, Lise Getoor |
ICML | 2 |
| 2015 | The Benefits of Learning with Strongly Convex Approximate InferenceabstractWe explore the benefits of strongly convex free energies in variational inference, providing both theoretical motivation and a new meta-algorithm. Using the duality between strong convexity and stability, we prove a high-probability bound on the error of learned marginals that is inversely proportional to the modulus of convexity of the free energy, thereby motivating free energies whose moduli are constant with respect to the size of the graph. We identify sufficient conditions for Ω(1)-strong convexity in two popular variational techniques: tree-reweighted and counting number entropies. Our insights for the latter suggest a novel counting number optimization framework, which guarantees strong convexity for any given modulus. Our experiments demonstrate that learning with a strongly convex free energy, using our optimization framework to guarantee a given modulus, results in substantially more accurate marginal probabilities, thereby validating our theoretical claims and the effectiveness of our framework. Ben London 0001, Bert Huang, Lise Getoor |
ICML | 2 |
| 2014 | Learning Latent Engagement Patterns of Students in Online CoursesabstractMaintaining and cultivating student engagement is critical for learning. Understanding factors affecting student engagement will help in designing better courses and improving student retention. The large number of participants in massive open online courses (MOOCs) and data collected from their interaction with the MOOC open up avenues for studying student engagement at scale. In this work, we develop a framework for modeling and understanding student engagement in online courses based on student behavioral cues. Our first contribution is the abstraction of student engagement types using latent representations and using that in a probabilistic model to connect student behavior with course completion. We demonstrate that the latent formulation for engagement helps in predicting student survival across three MOOCs. Next, in order to initiate better instructor interventions, we need to be able to predict student survival early in the course. We demonstrate that we can predict student survival early in the course reliably using the latent model. Finally, we perform a closer quantitative analysis of user interaction with the MOOC and identify student activities that are good indicators for survival at different points in the course. Arti Ramesh, Dan Goldwasser, Bert Huang, Hal Daumé III, Lise Getoor |
AAAI | 3 |
| 2014 | PAC-Bayesian Collective StabilityabstractRecent results have shown that the generalization error of structured predictors decreases with both the number of examples and the size of each example, provided the data distribution has weak dependence and the predictor exhibits a smoothness property called collective stability. These results use an especially strong definition of collective stability that must hold uniformly over all inputs and all hypotheses in the class. We investigate whether weaker definitions of collective stability suffice. Using the PAC-Bayes framework, which is particularly amenable to our new definitions, we prove that generalization is indeed possible when uniform collective stability happens with high probability over draws of predictors (and inputs). We then derive a generalization bound for a class of structured predictors with variably convex inference, which suggests a novel learning objective that optimizes collective stability. Ben London 0001, Bert Huang, Ben Taskar, Lise Getoor |
AISTATS | 2 |
| 2014 | 'Beating the news' with EMBERS: forecasting civil unrest using open source indicatorsabstractWe describe the design, implementation, and evaluation of EMBERS, an automated, 24x7 continuous system for forecasting civil unrest across 10 countries of Latin America using open source indicators such as tweets, news sources, blogs, economic indicators, and other data sources. Unlike retrospective studies, EMBERS has been making forecasts into the future since Nov 2012 which have been (and continue to be) evaluated by an independent T&E team (MITRE). Of note, EMBERS has successfully forecast the June 2013 protests in Brazil and Feb 2014 violent protests in Venezuela. We outline the system architecture of EMBERS, individual models that leverage specific data sources, and a fusion and suppression engine that supports trading off specific evaluation criteria. EMBERS also provides an audit trail interface that enables the investigation of why specific predictions were made along with the data utilized for forecasting. Through numerous evaluations, we demonstrate the superiority of EMBERS over baserate methods and its capability to forecast significant societal happenings. Naren Ramakrishnan, Patrick Butler, Sathappan Muthiah, Nathan Self, Rupinder Paul Khandpur, Parang Saraf, Wei Wang 0064, Jose Cadena, Anil Vullikanti, Gizem Korkmaz, Chris J. Kuhlman, Achla Marathe, Liang Zhao 0002, Ting Hua, Feng Chen 0001, Chang-Tien Lu, Bert Huang, Aravind Srinivasan, Khoa Trinh, Lise Getoor, Graham Katz, Andy Doyle, Chris Ackermann, Ilya Zavorin, Jim Ford, Kristen Maria Summers, Youssef Fayed, Jaime Arredondo, Dipak Gupta, David Mares |
KDD | 17 |
| 2014 | Uncovering hidden engagement patterns for predicting learner performance in MOOCsabstractMaintaining and cultivating student engagement is a prerequisite for MOOCs to have broad educational impact. Understanding student engagement as a course progresses helps characterize student learning patterns and can aid in minimizing dropout rates, initiating instructor intervention. In this paper, we construct a probabilistic model connecting student behavior and class performance, formulating student engagement types as latent variables. We show that our model identifies course success indicators that can be used by instructors to initiate interventions and assist students. Arti Ramesh, Dan Goldwasser, Bert Huang, Hal Daumé III, Lise Getoor |
L@S | 3 |
| 2014 | Network-Based Drug-Target Interaction Prediction with Probabilistic Soft LogicabstractDrug-target interaction studies are important because they can predict drugs' unexpected therapeutic or adverse side effects. In silico predictions of potential interactions are valuable and can focus effort on in vitro experiments. We propose a prediction framework that represents the problem using a bipartite graph of drug-target interactions augmented with drug-drug and target-target similarity measures and makes predictions using probabilistic soft logic (PSL). Using probabilistic rules in PSL, we predict interactions with models based on triad and tetrad structures. We apply (blocking) techniques that make link prediction in PSL more efficient for drug-target interaction prediction. We then perform extensive experimental studies to highlight different aspects of the model and the domain, first comparing the models with different structures and then measuring the effect of the proposed blocking on the prediction performance and efficiency. We demonstrate the importance of rule weight learning in the proposed PSL model and then show that PSL can effectively make use of a variety of similarity measures. We perform an experiment to validate the importance of collective inference and using multiple similarity measures for accurate predictions in contrast to non-collective and single similarity assumptions. Finally, we illustrate that our PSL model achieves state-of-the-art performance with simple, interpretable rules and evaluate our novel predictions using online data sets. Shobeir Fakhraei, Bert Huang, Louiqa Raschid, Lise Getoor |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |
| 2013 | A hypergraph-partitioned vertex programming approach for large-scale consensus optimizationabstractIn modern data science problems, techniques for extracting value from big data require performing large-scale optimization over heterogenous, irregularly structured data. Much of this data is best represented as multi-relational graphs, making vertex-programming abstractions such as those of Pregel and GraphLab ideal fits for modern large-scale data analysis. In this paper, we describe a vertex-programming implementation of a popular consensus optimization technique known as the alternating direction method of multipliers (ADMM) [1]. ADMM consensus optimization allows the elegant solution of complex objectives such as inference in rich probabilistic models. We also introduce a novel hypergraph partitioning technique that improves over the state-of-the-art vertex programming framework and significantly reduces the communication cost by reducing the number of replicated nodes by an order of magnitude. We implement our algorithm in GraphLab and measure scaling performance on a variety of realistic bipartite graphs and a large synthetic voter-opinion analysis application. We show a 50% improvement in running time over the current GraphLab partitioning scheme. Hui Miao 0001, Bert Huang, Lise Getoor |
IEEE BigData | 3 |
| 2013 | Collective Stability in Structured Prediction: Generalization from One ExampleabstractStructured predictors enable joint inference over multiple interdependent output variables. These models are often trained on a small number of examples with large internal structure. Existing distribution-free generalization bounds do not guarantee generalization in this setting, though this contradicts a large body of empirical evidence from computer vision, natural language processing, social networks and other fields. In this paper, we identify a set of natural conditions – weak dependence, hypothesis complexity and a new measure, collective stability – that are sufficient for generalization from even a single example, without imposing an explicit generative model of the data. We then demonstrate that the complexity and stability conditions are satisfied by a broad class of models, including marginal inference in templated graphical models. We thus obtain uniform convergence rates that can decrease significantly faster than previous bounds, particularly when each structured example is sufficiently large and the number of training examples is constant, even one. Ben London 0001, Bert Huang, Ben Taskar, Lise Getoor |
ICML (3) | 2 |
| 2013 | Hinge-loss Markov Random Fields: Convex Inference for Structured Prediction
Stephen H. Bach, Bert Huang, Ben London 0001, Lise Getoor |
UAI | 2 |
| 2012 | Machine Learning for theNew York City Power GridabstractPower companies can benefit from the use of knowledge discovery methods and statistical machine learning for preventive maintenance. We introduce a general process for transforming historical electrical grid data into models that aim to predict the risk of failures for components and systems. These models can be used directly by power companies to assist with prioritization of maintenance and repair work. Specialized versions of this process are used to produce 1) feeder failure rankings, 2) cable, joint, terminator, and transformer rankings, 3) feeder Mean Time Between Failure (MTBF) estimates, and 4) manhole events vulnerability rankings. The process in its most general form can handle diverse, noisy, sources that are historical (static), semi-real-time, or realtime, incorporates state-of-the-art machine learning algorithms for prioritization (supervised ranking or MTBF), and includes an evaluation of results via cross-validation and blind test. Above and beyond the ranked lists and MTBF estimates are business management interfaces that allow the prediction capability to be integrated directly into corporate planning and decision support; such interfaces rely on several important properties of our general modeling approach: that machine learning features are meaningful to domain experts, that the processing of data is transparent, and that prediction results are accurate enough to support sound decision making. We discuss the challenges in working with historical electrical grid data that were not designed for predictive purposes. The “rawness” of these data contrasts with the accuracy of the statistical models that can be obtained from the process; these models are sufficiently accurate to assist in maintaining New York City’s electrical grid. Cynthia Rudin, David L. Waltz, Roger Anderson, Albert Boulanger, Ansaf Salleb-Aouissi, Maggie Chow, Haimonti Dutta, Philip Gross, Bert Huang, Steve Ierome |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2012 | Semantic Model Vectors for Complex Video Event RecognitionabstractWe propose semantic model vectors, an intermediate level semantic representation, as a basis for modeling and detecting complex events in unconstrained real-world videos, such as those from YouTube. The semantic model vectors are extracted using a set of discriminative semantic classifiers, each being an ensemble of SVM models trained from thousands of labeled web images, for a total of 280 generic concepts. Our study reveals that the proposed semantic model vectors representation outperforms-and is complementary to-other low-level visual descriptors for video event modeling. We hence present an end-to-end video event detection system, which combines semantic model vectors with other static or dynamic visual descriptors, extracted at the frame, segment, or full clip level. We perform a comprehensive empirical study on the 2010 TRECVID Multimedia Event Detection task (http://www.nist.gov/itl/iad/mig/med10.cfm), which validates the semantic model vectors representation not only as the best individual descriptor, outperforming state-of-the-art global and local static features as well as spatio-temporal HOG and HOF descriptors, but also as the most compact. We also study early and late feature fusion across the various approaches, leading to a 15% performance boost and an overall system performance of 0.46 mean average precision. In order to promote further research in this direction, we made our semantic model vectors for the TRECVID MED 2010 set publicly available for the community to use (http://www1.cs.columbia.edu/~mmerler/SMV.html). Michele Merler, Bert Huang, Lexing Xie, Gang Hua 0001, Apostol Natsev |
IEEE Trans. Multim. | 2 |
| 2011 | Learning a Distance Metric from a NetworkabstractMany real-world networks are described by both connectivity information and features for every node. To better model and understand these networks, we present structure preserving metric learning (SPML), an algorithm for learning a Mahalanobis distance metric from a network such that the learned distances are tied to the inherent connectivity structure of the network. Like the graph embedding algorithm structure preserving embedding, SPML learns a metric which is structure preserving, meaning a connectivity algorithm such as k-nearest neighbors will yield the correct connectivity when applied using the distances from the learned metric. We show a variety of synthetic and real-world experiments where SPML predicts link patterns from node features more accurately than standard techniques. We further demonstrate a method for optimizing SPML based on stochastic gradient descent which removes the running-time dependency on the size of the network and allows the method to easily scale to networks of thousands of nodes and millions of edges. Blake Shaw, Bert Huang, Tony Jebara |
NIPS | 2 |
| 2009 | Exact Graph Structure Estimation with Degree PriorsabstractWe describe a generative model for graph edges under specific degree distributions which admits an exact and efficient inference method for recovering the most likely structure. This binary graph structure is obtained by reformulating the inference problem as a generalization of the polynomial time combinatorial optimization known as b-matching. Standard b-matching recovers a constant-degree constrained maximum weight subgraph from an original graph instead of a distribution over degrees. After this mapping, the most likely graph structure can be found in cubic time with respect to the number of nodes using max flow methods. Furthermore, in some instances, the combinatorial optimization problem can be solved exactly in near quadratic time by loopy belief propagation and max product updates even if the original input graph is dense. We show an example application to post-processing of recommender system predictions. Bert Huang, Tony Jebara |
ICMLA | 1 |
| 2009 | Alive on Back-feed Culprit Identification via Machine LearningabstractWe describe an application of machine learning techniques toward the problem of predicting which network protector switch is the cause of an alive on back-feed (ABF) event in the New York City power distribution system. When an electrical feeder is shut down, all network protector switches connected to the feeder should open to isolate the feeder. When a switch malfunctions and does not open, electrical current flows into the feeder, which remains energized. This causes the feeder to be "alive" on back-feed current, and maintenance cannot proceed. Our goal is to provide a ranking of network protector switches according to their susceptibility to such malfunction. Such a ranking can assist prioritization of which switches to repair when an ABF event occurs. We compare three methods for computing a ranking: an SVM classification approach, a maximum entropy density estimation approach and an SVM-ranking approach. Bert Huang, Ansaf Salleb-Aouissi, Philip Gross |
ICMLA | 1 |
| 2009 | Discovering Characterization Rules from RankingsabstractFor many ranking applications we would like to understand not only which items are top-ranked, but also why they are top-ranked. However, many of the best ranking algorithms (e. g., SVMs) are black boxes that give little information about the factors for their rankings. We describe and demonstrate a new approach that can work in conjunction with any ranking algorithm to discover explanations for the items at the top of the rankings. These explanations are in the form of rules expressed as boolean combinations of attribute-value expressions. These rules are discovered by contrasting attributes of items drawn from both the top and bottom of a ranking list, looking for items that have high leverage, corresponding to rules with broad coverage and sharp differentiations. We include empirical results to demonstrate the utility of our method. Ansaf Salleb-Aouissi, Bert Huang, David L. Waltz |
ICMLA | 2 |