VLDB 2026 Research / reviewers in the wild / expert
Wray L. Buntine
dblp:72/3885 · also Wray Lindsay Buntine
· DBLP profile ↗
122ranked-venue papers
30as first author
32since 2021 · last 2026
0000-0001-9292-1015ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 92 · 23 first-author · 24 since 2021Databases, data management, data science and information retrieval · 37 · 7 first-author · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 13 · 4 first-author · 8 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 2 first-author · 4 since 2021Systems, architecture and hardware · 5 · 2 first-authorHuman-computer interaction and ubiquitous computing · 3 · 1 first-author · 2 since 2021Theory of computation · 2 · 1 first-authorSoftware engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Uncertainty-Based Methods for Automated Process Reward Data Construction and Output Aggregation in Mathematical ReasoningabstractLarge language models have demonstrated remarkable capabilities in complex mathematical reasoning tasks, but they inevitably generate errors throughout multi-step solutions. Process-level Reward Models (PRMs) have shown great promise by providing supervision and evaluation at each intermediate step, thereby effectively improving the models’ reasoning abilities. However, training effective PRMs requires high-quality process reward data, yet existing methods for constructing such data are often labour-intensive or inefficient. In this paper, we propose an uncertainty-driven framework for automated process reward data construction, encompassing both data generation and annotation processes for PRMs. Additionally, we identify the limitations of both majority vote and PRMs, and introduce two generic uncertainty-aware output aggregation methods: Hybrid Majority Reward Vote and Weighted Reward Frequency Vote, which combine the strengths of majority vote with PRMs. Extensive experiments on ProcessBench, MATH, and GSMPlus show the effectiveness and efficiency of the proposed PRM data construction framework, and demonstrate that the two output aggregation methods further improve the mathematical reasoning abilities across diverse PRMs. Jiuzhou Han, Wray L. Buntine, Ehsan Shareghi |
AAAI | 2 |
| 2026 | ViMedCSS: A Vietnamese Medical Code-Switching Speech Dataset & BenchmarkabstractCode-switching (CS), which is when Vietnamese speech uses English words like drug names or procedures, is a common phenomenon in Vietnamese medical communication. This creates challenges for Automatic Speech Recognition (ASR) systems, especially in low-resource languages like Vietnamese. Current most ASR systems struggle to recognize correctly English medical terms within Vietnamese sentences, and no benchmark addresses this challenge. In this paper, we construct a 34-hour Vietnamese Medical Code-Switching Speech dataset (ViMedCSS) containing 16,576 utterances. Each utterance includes at least one English medical term drawn from a curated bilingual lexicon covering five medical topics. Using this dataset, we evaluate several state-of-the-art ASR models and examine different specific fine-tuning strategies for improving medical term recognition to investigate the best approach to solve in the dataset. Experimental results show that Vietnamese-optimized models perform better on general segments, while multilingual pretraining helps capture English insertions. The combination of both approaches yields the best balance between overall and code-switched accuracy. This work provides the first benchmark for Vietnamese medical code-switching and offers insights into effective domain adaptation for low-resource, multilingual ASR systems. Tung X. Nguyen, Nhu Vo, Giang-Son Nguyen, Duy Mai Hoang, Chien Dinh Huynh, Inigo Jauregi Unanue, Massimo Piccardi, Wray L. Buntine, Dung D. Le |
LREC | 8 |
| 2026 | Ensembled Bayesian tabular data generator
Yishuo Zhang, Nayyar Abbas Zaidi, Jiahui Zhou, Gang Li 0009, Wray L. Buntine |
Knowl. Inf. Syst. | 5 |
| 2025 | Leveraging Deep AUC Maximisation for Enhanced Active Learning in Named Entity Recognition
Dan Nguyen, Wray L. Buntine, Haifeng Zhao 0002, Lan Du 0002 |
ADMA (2) | 4 |
| 2025 | Logical Reasoning with Outcome Reward Models for Test-Time ScalingabstractLogical reasoning is a critical benchmark for evaluating the capabilities of large language models (LLMs), as it reflects their ability to derive valid conclusions from given premises.While the combination of test-time scaling with dedicated outcome or process reward models has opened up new avenues to enhance LLMs performance in complex reasoning tasks, this space is under-explored in deductive logical reasoning.We present a set of Outcome Reward Models (ORMs) for deductive reasoning.To train the ORMs we mainly generate data using Chain-of-Thought (CoT) with single and multiple samples.Additionally, we propose a novel tactic to further expand the type of errors covered in the training dataset of the ORM.In particular, we propose an echo generation technique that leverages LLMs' tendency to reflect incorrect assumptions made in prompts to extract additional training data, covering previously unexplored error types.While a standard CoT chain may contain errors likely to be made by the reasoner, the echo strategy deliberately steers the model toward incorrect reasoning.We show that ORMs trained on CoT and echoaugmented data demonstrate improved performance on the FOLIO, JustLogic, and ProverQA datasets across four different LLMs. Ramya Keerthy Thatikonda, Wray L. Buntine, Ehsan Shareghi |
EMNLP | 2 |
| 2025 | Navigating Conflicting Views: Harnessing Trust for LearningabstractResolving conflicts is critical for improving the reliability of multi-view classification. While prior work focuses on learning consistent and informative representations across views, it often assumes perfect alignment and equal importance of all views, an assumption rarely met in real-world scenarios, as some views may express distinct information. To address this, we develop a computational trust-based discounting method that enhances the Evidential Multi-view framework by accounting for the instance-wise reliability of each view through a probability-sensitive trust mechanism. We evaluate our method on six real-world datasets using Top-1 Accuracy, Fleiss’ Kappa, and a new metric, Multi-View Agreement with Ground Truth, to assess prediction reliability. We also assess the effectiveness of uncertainty in indicating prediction correctness via AUROC. Additionally, we test the scalability of our method through end-to-end training on a large-scale dataset. The experimental results show that computational trust can effectively resolve conflicts, paving the way for more reliable multi-view classification models in real-world applications. Codes available at: https://github.com/OverfitFlow/Trust4Conflict Jueqing Lu, Wray L. Buntine, Joanna Dipnall, Belinda Gabbe, Lan Du 0002 |
ICML | 2 |
| 2025 | HOPE: A Memory-Based and Composition-Aware Framework for Zero-Shot Learning with Hopfield Network and Soft Mixture of ExpertsabstractCompositional Zero-Shot Learning (CZSL) has emerged as an essential paradigm in machine learning, aiming to overcome the constraints of traditional zero-shot learning by incorporating compositional thinking into its method-ology. Conventional zero-shot learning has difficulty managing unfamiliar combinations of seen and unseen classes because it depends on pre-defined class embeddings. In contrast, Compositional Zero-Shot Learning leverages the inherent hierarchies and structural connections among classes, creating new class representations by combining at-tributes, components, or other semantic elements. In our paper, we propose a novel framework that for the first time combines the Modern Hopfield Network with a Mixture of Experts (HOPE) to classify the compositions of previously unseen objects. Specifically, the Modern Hopfield Network creates a memory that stores label prototypes and identifies relevant labels for a given input image. Subsequently, the Mixture of Expert models integrates the image with the appropriate prototype to produce the final composition classi-fication. Our approach achieves SOTA performance on sev-eral benchmarks, including MIT-States and UT-Zappos. We also examine how each component contributes to improved generalization. Do Huu Dat, Po Yuan Mao, Tien Hoang Nguyen, Wray L. Buntine, Mohammed Bennamoun |
WACV | 4 |
| 2025 | LLM Reading Tea Leaves: Automatically Evaluating Topic Models with Large Language ModelsabstractAbstract Topic modeling has been a widely used tool for unsupervised text analysis. However, comprehensive evaluations of a topic model remain challenging. Existing evaluation methods are either less comparable across different models (e.g., perplexity) or focus on only one specific aspect of a model (e.g., topic quality or document representation quality) at a time, which is insufficient to reflect the overall model performance. In this paper, we propose WALM (Word Agreement with Language Model), a new evaluation method for topic modeling that considers the semantic quality of document representations and topics in a joint manner, leveraging the power of Large Language Models (LLMs). With extensive experiments involving different types of topic models, WALM is shown to align with human judgment and can serve as a complementary evaluation method to the existing ones, bringing a new perspective to topic modeling. Our software package is available at https://github.com/Xiaohao-Yang/Topic_Model_Evaluation. Xiaohao Yang, He Zhao 0001, Dinh Q. Phung, Wray L. Buntine, Lan Du 0002 |
Trans. Assoc. Comput. Linguistics | 4 |
| 2025 | Context-driven cold-start Web traffic forecastingabstractAbstract Cold-start forecasting is critical in dynamic scenarios where early-stage forecasting drives key decisions, such as content prioritization, resource allocation, and demand estimation before observable trends emerge. In this work, we explore the potential of multimodal forecasting techniques for cold-start forecasting and offer insights into designing more scalable and adaptive models. In particular, we address context-driven cold-start web traffic forecasting that includes textual content and historical web traffic of relevant web pages to generate forecasts when no historical data is available for the target new web page. To advance research in this area, we collect, clean, and align a high-dimensional, multimodal web traffic dataset. We adopt a Retrieval-Augmented Generation framework, and propose the use of large language models (LLMs) for this task. Our experiments demonstrate that the LLM-based strategy consistently outperforms the statistical baseline across multiple forecasting horizons. The best-performing LLM-based model reduces WRMSPE by 0.81% and WAPE by 4.5%, compared with other methods. Furthermore, LLM-based feature extraction enhances contextual understanding, leading to greater stability in long-horizon forecasts. Xin Zhou 0023, Weiqing Wang 0001, Wray L. Buntine, Christoph Bergmeir |
World Wide Web (WWW) | 3 |
| 2024 | Harnessing the Power of Beta Scoring in Deep Active Learning for Multi-Label Text ClassificationabstractWithin the scope of natural language processing, the domain of multi-label text classification is uniquely challenging due to its expansive and uneven label distribution. The complexity deepens due to the demand for an extensive set of annotated data for training an advanced deep learning model, especially in specialized fields where the labeling task can be labor-intensive and often requires domain-specific knowledge. Addressing these challenges, our study introduces a novel deep active learning strategy, capitalizing on the Beta family of proper scoring rules within the Expected Loss Reduction framework. It computes the expected increase in scores using the Beta Scoring Rules, which are then transformed into sample vector representations. These vector representations guide the diverse selection of informative sample, directly linking this process to the model's expected proper score. Comprehensive evaluations across both synthetic and real datasets reveal our method's capability to often outperform established acquisition techniques in multi-label text classification, presenting encouraging outcomes across various architectural and dataset scenarios. Ngoc Dang Nguyen, Lan Du 0002, Wray L. Buntine |
AAAI | 4 |
| 2024 | Scalable Transformer for High Dimensional Multivariate Time Series ForecastingabstractDeep models for Multivariate Time Series (MTS) forecasting have recently demonstrated significant success. Channel-dependent models capture complex dependencies that channel-independent models cannot capture. However, the number of channels in real-world applications outpaces the capabilities of existing channel-dependent models, and contrary to common expectations, some models underperform the channel-independent models in handling high-dimensional data, which raises questions about the performance of channel-dependent models. To address this, our study first investigates the reasons behind the suboptimal performance of these channel-dependent models on high-dimensional MTS data. Our analysis reveals that two primary issues lie in the introduced noise from unrelated series that increases the difficulty of capturing the crucial inter-channel dependencies, and challenges in training strategies due to high-dimensional data. To address these issues, we propose STHD, the Scalable Transformer for High-Dimensional Multivariate Time Series Forecasting. STHD has three components: a) Relation Matrix Sparsity that limits the noise introduced and alleviates the memory issue; b) ReIndex applied as a training strategy to enable a more flexible batch size setting and increase the diversity of training data; and c) Transformer that handles 2-D inputs and captures channel dependencies. These components jointly enable STHD to manage the high-dimensional MTS while maintaining computational feasibility. Furthermore, experimental results show STHD's considerable improvement on three high-dimensional datasets: Crime-Chicago, Wiki-People, and Traffic. The source code and dataset are publicly available https://github.com/xinzzzhou/ScalableTransformer4HighDimensionMTSF.git. Xin Zhou 0023, Weiqing Wang 0001, Wray L. Buntine, Shilin Qu, Abishek Sriramulu, Weicong Tan, Christoph Bergmeir |
CIKM | 3 |
| 2024 | Improving Vietnamese-English Medical Machine TranslationabstractMachine translation for Vietnamese-English in the medical domain is still an under-explored research area. In this paper, we introduce MedEV—a high-quality Vietnamese-English parallel dataset constructed specifically for the medical domain, comprising approximately 360K sentence pairs. We conduct extensive experiments comparing Google Translate, ChatGPT (gpt-3.5-turbo), state-of-the-art Vietnamese-English neural machine translation models and pre-trained bilingual/multilingual sequence-to-sequence models on our new MedEV dataset. Experimental results show that the best performance is achieved by fine-tuning “vinai-translate” for each translation direction. We publicly release our dataset to promote further research. Nhu Vo, Dat Quoc Nguyen, Dung D. Le, Massimo Piccardi, Wray L. Buntine |
LREC/COLING | 5 |
| 2024 | Bayesian Estimate of Mean Proper Scores for Diversity-Enhanced Active LearningabstractThe effectiveness of active learning largely depends on the sampling efficiency of the acquisition function. Expected Loss Reduction (ELR) focuses on a Bayesian estimate of the reduction in classification error, and more general costs fit in the same framework. We propose Bayesian Estimate of Mean Proper Scores (BEMPS) to estimate the increase in strictly proper scores such as log probability or negative mean square error within this framework. We also prove convergence results for this general class of costs. To facilitate better experimentation with the new acquisition functions, we develop a complementary batch AL algorithm that encourages diversity in the vector of expected changes in scores for unlabeled data. To allow high-performance classifiers, we combine deep ensembles, and dynamic validation set construction on pretrained models, and further speed up the ensemble process with the idea of Monte Carlo Dropout. Extensive experiments on both texts and images show that the use of mean square error and log probability with BEMPS yields robust acquisition functions and well-calibrated classifiers, and consistently outperforms the others tested. The advantages of BEMPS over the others are further supported by a set of qualitative analyses, where we visualise their sampling behaviour using data maps and t-SNE plots. Lan Du 0002, Wray L. Buntine |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | OntoMedRec: Logically-pretrained model-agnostic ontology encoders for medication recommendationabstractAbstract Recommending medications with electronic health records (EHRs) is a challenging task for data-driven clinical decision support systems. Most existing models learnt representations for medical concepts based on EHRs and make recommendations with the learnt representations. However, most medications appear in EHR datasets for limited times (the frequency distribution of medications follows power law distribution), resulting in insufficient learning of their representations of the medications. Medical ontologies are the hierarchical classification systems for medical terms where similar terms will be in the same class on a certain level. In this paper, we propose OntoMedRec, the logically-pretrained and model-agnostic medical Ontology Encoders for Medication Recommendation that addresses data sparsity problem with medical ontologies. We conduct comprehensive experiments on real-world EHR datasets to evaluate the effectiveness of OntoMedRec by integrating it into various existing downstream medication recommendation models. The result shows the integration of OntoMedRec improves the performance of various models in both the entire EHR datasets and the admissions with few-shot medications. We provide the GitHub repository for the source code. ( https://github.com/WaicongTam/OntoMedRec ) Weicong Tan, Weiqing Wang 0001, Xin Zhou 0023, Wray L. Buntine, Gordon Bingham, Hongzhi Yin |
World Wide Web (WWW) | 4 |
| 2023 | AUC Maximization for Low-Resource Named Entity RecognitionabstractCurrent work in named entity recognition (NER) uses either cross entropy (CE) or conditional random fields (CRF) as the objective/loss functions to optimize the underlying NER model. Both of these traditional objective functions for the NER problem generally produce adequate performance when the data distribution is balanced and there are sufficient annotated training examples. But since NER is inherently an imbalanced tagging problem, the model performance under the low-resource settings could suffer using these standard objective functions. Based on recent advances in area under the ROC curve (AUC) maximization, we propose to optimize the NER model by maximizing the AUC score. We give evidence that by simply combining two binary-classifiers that maximize the AUC score, significant performance improvement over traditional loss functions is achieved under low-resource NER settings. We also conduct extensive experiments to demonstrate the advantages of our method under the low-resource and highly-imbalanced data distribution settings. To the best of our knowledge, this is the first work that brings AUC maximization to the NER setting. Furthermore, we show that our method is agnostic to different types of NER embeddings, models and domains. The code of this work is available at https://github.com/dngu0061/NER-AUC-2T. Ngoc Dang Nguyen, Lan Du 0002, Wray L. Buntine, Richard Beare, Changyou Chen |
AAAI | 4 |
| 2023 | Cross-Domain Graph Anomaly Detection via Anomaly-Aware Contrastive AlignmentabstractCross-domain graph anomaly detection (CD-GAD) describes the problem of detecting anomalous nodes in an unlabelled target graph using auxiliary, related source graphs with labelled anomalous and normal nodes. Although it presents a promising approach to address the notoriously high false positive issue in anomaly detection, little work has been done in this line of research. There are numerous domain adaptation methods in the literature, but it is difficult to adapt them for GAD due to the unknown distributions of the anomalies and the complex node relations embedded in graph data. To this end, we introduce a novel domain adaptation approach, namely Anomaly-aware Contrastive alignmenT (ACT), for GAD. ACT is designed to jointly optimise: (i) unsupervised contrastive learning of normal representations of nodes in the target graph, and (ii) anomaly-aware one-class alignment that aligns these contrastive node representations and the representations of labelled normal nodes in the source graph, while enforcing significant deviation of the representations of the normal nodes from the labelled anomalous nodes in the source graph. In doing so, ACT effectively transfers anomaly-informed knowledge from the source graph to learn the complex node relations of the normal class for GAD on the target graph without any specification of the anomaly distributions. Extensive experiments on eight CD-GAD settings demonstrate that our approach ACT achieves substantially improved detection performance over 10 state-of-the-art GAD methods. Code is available at https://github.com/QZ-WANG/ACT. Qizhou Wang 0001, Guansong Pang, Mahsa Salehi, Wray L. Buntine, Christopher Leckie |
AAAI | 4 |
| 2023 | Robust Educational Dialogue Act Classifiers with Low-Resource and Imbalanced Datasets
Jionghao Lin, Ngoc Dang Nguyen, David Lang, Lan Du 0002, Wray L. Buntine, Richard Beare, Guanliang Chen, Dragan Gasevic |
AIED | 6 |
| 2023 | Does Informativeness Matter? Active Learning for Educational Dialogue Act Classification
Jionghao Lin, David Lang, Guanliang Chen, Dragan Gasevic, Lan Du 0002, Wray L. Buntine |
AIED | 7 |
| 2023 | Low-Resource Named Entity Recognition: Can One-vs-All AUC Maximization Help?abstractNamed entity recognition (NER), a task that identifies and categorizes named entities such as persons or organizations from text, is traditionally framed as a multi-class classification problem. However, this approach often overlooks the issues of imbalanced label distributions, particularly in low-resource settings, which is common in certain NER contexts, like biomedical NER (bioNER). To address these issues, we propose an innovative reformulation of the multi-class problem as a one-vs-all (OVA) learning problem and introduce a loss function based on the area under the receiver operating characteristic curve (AUC). To enhance the efficiency of our OVA-based approach, we propose two training strategies: one groups labels with similar linguistic characteristics, and another employs meta-learning. The superiority of our approach is confirmed by its performance, which surpasses traditional NER learning in varying NER settings. Ngoc Dang Nguyen, Lan Du 0002, Wray L. Buntine, Richard Beare, Changyou Chen |
ICDM | 4 |
| 2023 | MEG: Masked Ensemble Tabular Data GeneratorabstractTabular data generation has seen renewed interest with the advent of Generative Adversarial Networks (GAN). Recently, it has been shown that one can use a Bayesian network as either a generator or a discriminator in the GAN framework, resulting in an algorithm known as GANBLR. It has been shown that GANBLR gives state of the art results for tabular data generation. However, the model has one limitation. It uses class attributes during model training. For example, a supervised Bayesian network is needed as a generator at training time. This makes GANBLR inapplicable for cases where we do not have access to class information. Addressing this shortcoming of GANBLR has been the main motivation of this work. In this work, we have proposed a new model of tabular data generation – Masked Ensemble Tabular Generator (MEG), which does not require class labels to generate tabular data. The proposed models rely on a novel strategy of using a collection of Bayesian networks as part of the generator, and relies on masking operations to train the generator efficiently. It also uses a group-based similarity measure to adjust the number of samples generated from each Bayesian network in the collection. We perform extensive experiments on a variety of datasets and demonstrate that MEG not only outperforms baselines that do not have class information during training, such as CTGAN and TVAE, but also outperforms baselines that provide access to class information during training, such as TableGAN and CtabGANmethods. It has almost similar performance in terms of machine learning utility to GANBLR, and of course is greatly advantaged by being truly unsupervised in nature. We highlight this by demonstrating its applicability to a clustering task. We also investigate the privacy preserving capabilities of MEG and demonstrate its superior performance compared to other baselines. Yishuo Zhang, Nayyar Abbas Zaidi, Gang Li 0009, Wray L. Buntine |
ICDM | 4 |
| 2023 | A Survey on Out-of-Distribution Evaluation of Neural NLP ModelsabstractAdversarial robustness, domain generalization and dataset biases are three active lines of research contributing to out-of-distribution (OOD) evaluation on neural NLP models. However, a comprehensive, integrated discussion of the three research lines is still lacking in the literature. This survey will 1) compare the three lines of research under a unifying definition; 2) summarize their data-generating processes and evaluation protocols for each line of research; and 3) emphasize the challenges and opportunities for future work. Ming Liu 0028, Shang Gao 0003, Wray L. Buntine |
IJCAI | 4 |
| 2022 | Hardness-guided domain adaptation to recognise biomedical named entities under low-resource scenariosabstractDomain adaptation is an effective solution to data scarcity in low-resource scenarios.However, when applied to token-level tasks such as bioNER, domain adaptation methods often suffer from the challenging linguistic characteristics that clinical narratives possess, which leads to unsatsifactory performance.In this paper, we present a simple yet effective hardnessguided domain adaptation (HGDA) framework for bioNER tasks that can effectively leverage the domain hardness information to improve the adaptability of the learnt model in the low-resource scenarios.Experimental results on biomedical datasets show that our model can achieve significant performance improvement over the recently published state-of-theart (SOTA) MetaNER model. Ngoc Dang Nguyen, Lan Du 0002, Wray L. Buntine, Changyou Chen, Richard Beare |
EMNLP | 3 |
| 2022 | ENDASh: Embedding Neighbourhood Dissimilarity with Attribute Shuffling for Graph Anomaly Detection
Qizhou Wang 0001, Mahsa Salehi, Jia Shun Low, Wray L. Buntine, Christopher Leckie |
PAKDD (2) | 4 |
| 2022 | SQAPlanner: Generating Data-Informed Software Quality Improvement PlansabstractSoftware Quality Assurance (SQA) planning aims to define proactive plans, such as defining maximum file size, to prevent the occurrence of software defects in future releases. To aid this,defect prediction modelshave been proposed to generate insights as the most important factors that are associated with software quality. Such insights that are derived from traditional defect models are far from actionable—i.e., practitioners still do not know what they should do or avoid to decrease the risk of having defects, and what is the risk threshold for each metric. A lack of actionable guidance and risk threshold can lead to inefficient and ineffective SQA planning processes. In this paper, we investigate the practitioners’ perceptions of current SQA planning activities, current challenges of such SQA planning activities, and propose four types of guidance to support SQA planning. We then propose and evaluate our AI-Driven SQAPlanner approach, a novel approach for generating four types of guidance and their associated risk thresholds in the form of rule-based explanations for the predictions of defect prediction models. Finally, we develop and evaluate a visualization for our SQAPlanner approach. Through the use of qualitative survey and empirical evaluation, our results lead us to conclude that SQAPlanner is needed, effective, stable, and practically applicable. We also find that 80 percent of our survey respondents perceived that our visualization is more actionable. Thus, our SQAPlanner paves a way for novel research in actionable software analytics—i.e., generating actionable guidance on what should practitioners do and not do to decrease the risk of having defects to support SQA planning. Dilini Rajapaksha, Chakkrit Tantithamthavorn, Jirayus Jiarpakdee, Christoph Bergmeir, John C. Grundy, Wray L. Buntine |
IEEE Trans. Software Eng. | 6 |
| 2021 | All Labels Are Not Created Equal: Enhancing Semi-Supervision via Label Grouping and Co-TrainingabstractPseudo-labeling is a key component in semi-supervised learning (SSL). It relies on iteratively using the model to generate artificial labels for the unlabeled data to train against. A common property among its various methods is that they only rely on the model’s prediction to make labeling decisions without considering any prior knowledge about the visual similarity among the classes. In this paper, we demonstrate that this degrades the quality of pseudo-labeling as it poorly represents visually similar classes in the pool of pseudo-labeled data. We propose SemCo, a method which leverages label semantics and co-training to address this problem. We train two classifiers with two different views of the class labels: one classifier uses the one-hot view of the labels and disregards any potential similarity among the classes, while the other uses a distributed view of the labels and groups potentially similar classes together. We then co-train the two classifiers to learn based on their disagreements. We show that our method achieves state-of-the-art performance across various SSL tasks including 5.6% accuracy improvement on Mini-ImageNet dataset with 1000 labeled examples. We also show that our method requires smaller batch size and fewer training iterations to reach its best performance. We make our code available at https://github.com/islam-nassar/semco. Islam Nassar, Samitha Herath, Ehsan Abbasnejad, Wray L. Buntine, Gholamreza Haffari |
CVPR | 4 |
| 2021 | Neural Attention-Aware Hierarchical Topic ModelabstractNeural topic models (NTMs) apply deep neural networks to topic modelling.Despite their success, NTMs generally ignore two important aspects: (1) only document-level word count information is utilized for the training, while more fine-grained sentence-level information is ignored, and (2) external semantic knowledge regarding documents, sentences and words are not exploited for the training.To address these issues, we propose a variational autoencoder (VAE) NTM model that jointly reconstructs the sentence and document word counts using combinations of bag-of-words (BoW) topical embeddings and pre-trained semantic embeddings.The pre-trained embeddings are first transformed into a common latent topical space to align their semantics with the BoW embeddings.Our model also features hierarchical KL divergence to leverage embeddings of each document to regularize those of their sentences, thereby paying more attention to semantically relevant sentences.Both quantitative and qualitative experiments have shown the efficacy of our model in 1) lowering the reconstruction errors at both the sentence and document levels, and 2) discovering more coherent topics from real-world datasets. He Zhao 0001, Ming Liu 0028, Lan Du 0002, Wray L. Buntine |
EMNLP (1) | 5 |
| 2021 | Neural Topic Model via Optimal Transport
He Zhao 0001, Dinh Q. Phung, Viet Huynh, Trung Le 0001, Wray L. Buntine |
ICLR | 5 |
| 2021 | Topic Modelling Meets Deep Neural Networks: A SurveyabstractTopic modelling has been a successful technique for text analysis for almost twenty years. When topic modelling met deep neural networks, there emerged a new and increasingly popular research area, neural topic models, with nearly a hundred models developed and a wide range of applications in neural language understanding such as text generation, summarisation and language models. There is a need to summarise research developments and discuss open problems and future directions. In this paper, we provide a focused yet comprehensive overview of neural topic models for interested researchers in the AI community, so as to facilitate them to navigate and innovate in this fast-growing research area. To the best of our knowledge, ours is the first review on this specific topic. He Zhao 0001, Dinh Q. Phung, Viet Huynh, Lan Du 0002, Wray L. Buntine |
IJCAI | 6 |
| 2021 | Topic Model or Topic Twaddle? Re-evaluating Semantic Interpretability MeasuresabstractWhen developing topic models, a critical question that should be asked is: How well will this model work in an applied setting?Because standard performance evaluation of topic interpretability uses automated measures modeled on human evaluation tests that are dissimilar to applied usage, these models' generalizability remains in question.In this paper, we probe the issue of validity in topic model evaluation and assess how informative coherence measures are for specialized collections used in an applied setting.Informed by the literature, we propose four understandings of interpretability.We evaluate these using a novel experimental framework reflective of varied applied settings, including human evaluations using open labeling, typical of applied research.These evaluations show that for some specialized collections, standard coherence measures may not inform the most appropriate topic model or the optimal number of topics, and current interpretability performance validation methods are challenged as a means to confirm model quality in the absence of ground truth data. Caitlin Doogan, Wray L. Buntine |
NAACL-HLT | 2 |
| 2021 | Diversity Enhanced Active Learning with Strictly Proper Scoring RulesabstractWe study acquisition functions for active learning (AL) for text classification. The Expected Loss Reduction (ELR) method focuses on a Bayesian estimate of the reduction in classification error, recently updated with Mean Objective Cost of Uncertainty (MOCU). We convert the ELR framework to estimate the increase in (strictly proper) scores like log probability or negative mean square error, which we call Bayesian Estimate of Mean Proper Scores (BEMPS). We also prove convergence results borrowing techniques used with MOCU. In order to allow better experimentation with the new acquisition functions, we develop a complementary batch AL algorithm, which encourages diversity in the vector of expected changes in scores for unlabelled data. To allow high performance text classifiers, we combine ensembling and dynamic validation set construction on pretrained language models. Extensive experimental evaluation then explores how these different acquisition functions perform. The results show that the use of mean square error and log probability with BEMPS yields robust acquisition functions, which consistently outperform the others tested. Lan Du 0002, Wray L. Buntine |
NeurIPS | 3 |
| 2021 | Recommending content using side information
Rabeh Ravanifard, Wray L. Buntine, Abdolreza Mirzaei |
Appl. Intell. | 2 |
| 2021 | Content-Aware Listwise Collaborative Filtering
Rabeh Ravanifard, Abdolreza Mirzaei, Wray L. Buntine, Mehran Safayani |
Neurocomputing | 3 |
| 2020 | Variational Autoencoders for Sparse and Overdispersed Discrete DataabstractMany applications, such as text modelling, high-throughput sequencing, and recommender systems, require analysing sparse, high-dimensional, and overdispersed discrete (count or binary) data. Recent deep probabilistic models based on variational autoencoders (VAE) have shown promising results on discrete data but may have inferior modelling performance due to the insufficient capability in modelling overdispersion and model misspecification. To address these issues, we develop a VAE-based framework using the negative binomial distribution as the data distribution. We also provide an analysis of its properties vis-à-vis other models. We conduct extensive experiments on three problems from discrete data analysis: text analysis/topic modelling, collaborative filtering, and multi-label learning. Our models outperform state-of-the-art approaches on these problems, while also capturing the phenomenon of overdispersion more effectively. He Zhao 0001, Piyush Rai, Lan Du 0002, Wray L. Buntine, Dinh Q. Phung, Mingyuan Zhou |
AISTATS | 4 |
| 2020 | Collective Wisdom: Improving Low-resource Neural Machine Translation using Adaptive Knowledge DistillationabstractScarcity of parallel sentence-pairs poses a significant hurdle for training high-quality Neural Machine Translation (NMT) models in bilingually low-resource scenarios.A standard approach is transfer learning, which involves taking a model trained on a high-resource language-pair and fine-tuning it on the data of the low-resource MT condition of interest.However, it is not clear generally which high-resource language-pair offers the best transfer learning for the target MT setting.Furthermore, different transferred models may have complementary semantic and/or syntactic strengths, hence using only one model may be sub-optimal.In this paper, we tackle this problem using knowledge distillation, where we propose to distill the knowledge of ensemble of teacher models to a single student model.As the quality of these teacher models varies, we propose an effective adaptive knowledge distillation approach to dynamically adjust the contribution of the teacher models during the distillation process.Experiments on transferring from a collection of six language pairs from IWSLT to five low-resource language-pairs from TED Talks demonstrate the effectiveness of our approach, achieving up to +0.9 BLEU score improvement compared to strong baselines. Fahimeh Saleh, Wray L. Buntine, Gholamreza Haffari |
COLING | 2 |
| 2020 | MedGraph: Structural and Temporal Representation Learning of Electronic Medical RecordsabstractElectronic medical record (EMR) data contains historical sequences of visits of patients, and each visit contains rich information, such as patient demographics, hospital utilisation and medical codes, including diagnosis, procedure and medication codes. Most existing EMR embedding methods capture visit-code associations by constructing input visit representations as binary vectors with a static vocabulary of medical codes. With this limited representation, they fail in encapsulating rich attribute information of visits (demographics and utilisation information) and/or codes (e.g., medical code descriptions). Furthermore, current work considers visits of the same patient as discrete-time events and ignores time gaps between them. However, the time gaps between visits depict dynamics of the patient's medical history inducing varying influences on future visits. To address these limitations, we present MedGraph, a supervised EMR embedding method that captures two types of information: (1) the visit-code associations in an attributed bipartite graph, and (2) the temporal sequencing of visits through a point process. MedGraph produces Gaussian embeddings for visits and codes to model the uncertainty. We evaluate the performance of MedGraph through an extensive experimental study and show that MedGraph outperforms state-of-the-art EMR embedding methods in several medical risk prediction tasks. Bhagya Hettige, Weiqing Wang 0001, Yuan-Fang Li, Suong Le, Wray L. Buntine |
ECAI | 5 |
| 2020 | Robust Attribute and Structure Preserving Graph Embedding
Bhagya Hettige, Weiqing Wang 0001, Yuan-Fang Li, Wray L. Buntine |
PAKDD (2) | 4 |
| 2020 | Hierarchical Gradient Smoothing for Probability Estimation Trees
He Zhang 0010, François Petitjean, Wray L. Buntine |
PAKDD (1) | 3 |
| 2020 | Machine learning after the deep learning revolution
Wray L. Buntine |
Frontiers Comput. Sci. | 1 |
| 2020 | LoRMIkA: Local rule-based model interpretability with k-optimal associations
Dilini Rajapaksha, Christoph Bergmeir, Wray L. Buntine |
Inf. Sci. | 3 |
| 2020 | Bayesian network classifiers using ensembles and smoothing
He Zhang 0010, François Petitjean, Wray L. Buntine |
Knowl. Inf. Syst. | 3 |
| 2019 | Leveraging Meta Information in Short Text AggregationabstractAnalysing topics in short texts (e.g., tweets and new headings) is a challenging task because short texts often contain insufficient word co-occurrence information, which is important to learn good topics in conventional topic topics.To deal with the insufficiency, we propose a generative model that aggregates short texts into clusters by leveraging the associated meta information.Our model can generate more interpretable topics as well as document clusters.We develop an effective Gibbs sampling algorithm favoured by the fully local conjugacy in the model.Extensive experiments demonstrate that our model achieves better performance in terms of document clustering and topic coherence. He Zhao 0001, Lan Du 0002, Guanfeng Liu 0001, Wray L. Buntine |
ACL (1) | 4 |
| 2019 | Leveraging external information in topic modelling
He Zhao 0001, Lan Du 0002, Wray L. Buntine, Gang Liu 0021 |
Knowl. Inf. Syst. | 3 |
| 2018 | Learning How to Actively Learn: A Deep Imitation Learning ApproachabstractHeuristic-based active learning (AL) methods are limited when the data distribution of the underlying learning problems vary.We introduce a method that learns an AL policy using imitation learning (IL).Our IL-based approach makes use of an efficient and effective algorithmic expert, which provides the policy learner with good actions in the encountered AL situations.The AL strategy is then learned with a feedforward network, mapping situations to most informative query datapoints.We evaluate our method on two different tasks: text classification and named entity recognition.Experimental results show that our IL-based AL strategy is more effective than strong previous methods using heuristics and reinforcement learning. Ming Liu 0028, Wray L. Buntine, Gholamreza Haffari |
ACL (1) | 2 |
| 2018 | Distinguishing Question Subjectivity from Difficulty for Improved CrowdsourcingabstractThe questions in a crowdsourcing task typically exhibit varying degrees of difficulty and subjectivity. Their joint effects give rise to the variation in responses to the same question by different crowd-workers. This variation is low when the question is easy to answer and objective, and high when it is difficult and subjective. Unfortunately, current quality control methods for crowdsourcing consider only the question difficulty to account for the variation. As a result, these methods cannot distinguish workers' ,personal preferences for different correct answers of a partially subjective question from their ability to avoid objectively incorrect answers for that question. To address this issue, we present a probabilistic model which (i) explicitly encodes question difficulty as a model parameter and (ii) implicitly encodes question subjectivity via latent preference factors for crowd-workers. We show that question subjectivity induces grouping of crowd-workers, revealed through clustering of their latent preferences. Moreover, we develop a quantitative measure for the question subjectivity. Experiments show that our model (1) improves both the question true answer prediction and the unseen worker response prediction, and (2) can potentially provide rankings of questions coherent with human assessment in terms of difficulty and subjectivity. Mark J. Carman, Ye Zhu 0002, Wray L. Buntine |
ACML | 4 |
| 2018 | Bayesian Multi-label Learning with Sparse Features and Labels, and Label Co-occurrencesabstractWe present a probabilistic, fully Bayesian framework for multi-label learning. Our framework is based on the idea of learning a joint low-rank embedding of the label matrix and the label co-occurrence matrix. The proposed framework has the following appealing aspects: (1) It leverages the sparsity in the label matrix and the feature matrix, which results in very efficient inference, especially for sparse datasets, commonly encountered in multi-label learning problems, and (2) By effectively utilizing the label co-occurrence information, the model yields improved prediction accuracies, especially in the case where the amount of training data is low and/or the label matrix has a significant fraction of missing labels. Our framework enjoys full local conjugacy and admits a simple inference procedure via a scalable Gibbs sampler. We report experimental results on a number of benchmark datasets, on which it outperforms several state-of-the-art multi-label learning models. He Zhao 0001, Piyush Rai, Lan Du 0002, Wray L. Buntine |
AISTATS | 4 |
| 2018 | Learning to Actively Learn Neural Machine TranslationabstractTraditional active learning (AL) methods for machine translation (MT) rely on heuristics.However, these heuristics are limited when the characteristics of the MT problem change due to e.g. the language pair or the amount of the initial bitext.In this paper, we present a framework to learn sentence selection strategies for neural MT.We train the AL query strategy using a high-resource language-pair based on AL simulations, and then transfer it to the lowresource language-pair of interest.The learned query strategy capitalizes on the shared characteristics between the language pairs to make an effective use of the AL budget.Our experiments on three language-pairs confirms that our method is more effective than strong heuristic-based methods in various conditions, including cold-start and warm-start as well as small and extremely small data conditions. Ming Liu 0028, Wray L. Buntine, Gholamreza Haffari |
CoNLL | 2 |
| 2018 | Inter and Intra Topic Structure Learning with Word EmbeddingsabstractOne important task of topic modeling for text analysis is interpretability. By discovering structured topics one is able to yield improved interpretability as well as modeling accuracy. In this paper, we propose a novel topic model with a deep structure that explores both inter-topic and intra-topic structures informed by word embeddings. Specifically, our model discovers inter topic structures in the form of topic hierarchies and discovers intra topic structures in the form of sub-topics, each of which is informed by word embeddings and captures a fine-grained thematic aspect of a normal topic. Extensive experiments demonstrate that our model achieves the state-of-the-art performance in terms of perplexity, document classification, and topic quality. Moreover, with topic hierarchies and sub-topics, the topics discovered in our model are more interpretable, providing an illuminating means to understand text data. He Zhao 0001, Lan Du 0002, Wray L. Buntine, Mingyuan Zhou |
ICML | 3 |
| 2018 | Dirichlet belief networks for topic structure learningabstractRecently, considerable research effort has been devoted to developing deep architectures for topic models to learn topic structures. Although several deep models have been proposed to learn better topic proportions of documents, how to leverage the benefits of deep structures for learning word distributions of topics has not yet been rigorously studied. Here we propose a new multi-layer generative process on word distributions of topics, where each layer consists of a set of topics and each topic is drawn from a mixture of the topics of the layer above. As the topics in all layers can be directly interpreted by words, the proposed model is able to discover interpretable topic hierarchies. As a self-contained module, our model can be flexibly adapted to different kinds of topic models to improve their modelling accuracy and interpretability. Extensive experiments on text corpora demonstrate the advantages of the proposed model. He Zhao 0001, Lan Du 0002, Wray L. Buntine, Mingyuan Zhou |
NeurIPS | 3 |
| 2018 | A Left-to-Right Algorithm for Likelihood Estimation in Gamma-Poisson Factor Analysis
Joan Capdevila, Jesús Cerquides, Jordi Torres, François Petitjean, Wray L. Buntine |
ECML/PKDD (2) | 5 |
| 2018 | Accurate parameter estimation for Bayesian network classifiers using hierarchical Dirichlet processes
François Petitjean, Wray L. Buntine, Geoffrey I. Webb, Nayyar Abbas Zaidi |
Mach. Learn. | 2 |
| 2017 | A Word Embeddings Informed Focused Topic ModelabstractIn natural language processing and related fields, it has been shown that the word embeddings can successfully capture both the semantic and syntactic features of words. They can serve as complementary information to topics models, especially for the cases where word co-occurrence data is insufficient, such as with short texts. In this paper, we propose a focused topic model where how a topic focuses on words is informed by word embeddings. Our models is able to discover more informed and focused topics with more representative words, leading to better modelling accuracy and topic quality. With the data argumentation technique, we can derive an efficient Gibbs sampling algorithm that benefits from the fully local conjugacy of the model. We conduct extensive experiments on several real world datasets, which demonstrate that our model achieves comparable or improved performance in terms of both perplexity and topic coherence, particularly in handling short text data. He Zhao 0001, Lan Du 0002, Wray L. Buntine |
ACML | 3 |
| 2017 | MetaLDA: A Topic Model that Efficiently Incorporates Meta InformationabstractBesides the text content, documents and their associated words usually come with rich sets of meta information, such as categories of documents and semantic/syntactic features of words, like those encoded in word embeddings. Incorporating such meta information directly into the generative process of topic models can improve modelling accuracy and topic quality, especially in the case where the word-occurrence information in the training data is insufficient. In this paper, we present a topic model, called MetaLDA, which is able to leverage either document or word meta information, or both of them jointly. With two data argumentation techniques, we can derive an efficient Gibbs sampling algorithm, which benefits from the fully local conjugacy of the model. Moreover, the algorithm is favoured by the sparsity of the meta information. Extensive experiments on several real world datasets demonstrate that our model achieves comparable or improved performance in terms of both perplexity and topic quality, particularly in handling sparse texts. In addition, compared with other models using meta information, our model runs significantly faster. He Zhao 0001, Lan Du 0002, Wray L. Buntine, Gang Liu 0021 |
ICDM | 3 |
| 2017 | Leveraging Node Attributes for Incomplete Relational DataabstractRelational data are usually highly incomplete in practice, which inspires us to leverage side information to improve the performance of community detection and link prediction. This paper presents a Bayesian probabilistic approach that incorporates various kinds of node attributes encoded in binary form in relational models with Poisson likelihood. Our method works flexibly with both directed and undirected relational networks. The inference can be done by efficient Gibbs sampling which leverages sparsity of both networks and node attributes. Extensive experiments show that our models achieve the state-of-the-art link prediction results, especially with highly incomplete relational data. He Zhao 0001, Lan Du 0002, Wray L. Buntine |
ICML | 3 |
| 2017 | Efficient parameter learning of Bayesian network classifiersabstractRecent advances have demonstrated substantial benefits from learning with both generative and discriminative parameters. On the one hand, generative approaches address the estimation of the parameters of the joint distribution— $$\mathrm{P}(y,\mathbf{x})$$ , which for most network types is very computationally efficient (a notable exception to this are Markov networks) and on the other hand, discriminative approaches address the estimation of the parameters of the posterior distribution—and, are more effective for classification, since they fit $$\mathrm{P}(y|\mathbf{x})$$ directly. However, discriminative approaches are less computationally efficient as the normalization factor in the conditional log-likelihood precludes the derivation of closed-form estimation of parameters. This paper introduces a new discriminative parameter learning method for Bayesian network classifiers that combines in an elegant fashion parameters learned using both generative and discriminative methods. The proposed method is discriminative in nature, but uses estimates of generative probabilities to speed-up the optimization process. A second contribution is to propose a simple framework to characterize the parameter learning task for Bayesian network classifiers. We conduct an extensive set of experiments on 72 standard datasets and demonstrate that our proposed discriminative parameterization provides an efficient alternative to other state-of-the-art parameterizations. Nayyar Abbas Zaidi, Geoffrey I. Webb, Mark J. Carman, François Petitjean, Wray L. Buntine, Mike Hynes, Hans De Sterck |
Mach. Learn. | 5 |
| 2016 | PULP: A System for Exploratory Search of Scientific LiteratureabstractDespite the growing importance of exploratory search, information retrieval (IR) systems tend to focus on lookup search. Lookup searches are well served by optimising the precision and recall of search results, however, for exploratory search this may be counterproductive if users are unable to formulate an appropriate search query. We present a system called PULP that supports exploratory search for scientific literature, though the system can be easily adapted to other types of literature. PULP uses reinforcement learning (RL) to avert the user from context traps resulting from poorly chosen search queries, trading off between exploration (presenting the user with diverse topics) and exploitation (moving towards more specific topics). Where other RL-based systems suffer from the "cold start" problem, requiring sufficient time to adjust to a user's information needs, PULP initially presents the user with an overview of the dataset using temporal topic models. Topic models are displayed in an interactive alluvial diagram, where topics are shown as ribbons that change thickness with a given topics relative prevalence over time. Interactive, exploratory search sessions can be initiated by selecting topics as a starting point. Alan Medlar, Kalle Ilves, Wray L. Buntine, Dorota Glowacka |
SIGIR | 4 |
| 2016 | Nonparametric Bayesian topic modelling with the hierarchical Pitman-Yor processes
Kar Wai Lim, Wray L. Buntine, Changyou Chen, Lan Du 0002 |
Int. J. Approx. Reason. | 2 |
| 2016 | Bibliographic analysis on research publications using authors, categorical labels and the citation network
Kar Wai Lim, Wray L. Buntine |
Mach. Learn. | 2 |
| 2015 | Special session on trends & controversies in data science (TCDS)abstractAs an emerging area, data science is facing great opportunities as well as challenges. Often arguments exist: What is data science? Why data science? We have information science already, why do we need data science? Do we need analytics science? Is analytics new? What is the difference between statistics and data analytics? What makes a data scientist? We believe that a special session on Trends and Controversy about data science and advanced analytics could bring insights from different mindsets for the healthy development of the science and society. Accordingly, this T&C special session will host talks by invitation to outline different views about today and future of data science. Invited speakers can contribute a paper (in the same format as the main conference submissions but could be less than 10 pages) to the special session, which will be handled by program co-chairs and accepted into the main conference proceeding probably by addressing comments from the program cochairs. Florence Forbes, Wray L. Buntine |
DSAA | 2 |
| 2015 | Introduction: special issue of selected papers of ACML 2013
Cheng Soon Ong, Wray L. Buntine, Masashi Sugiyama, Geoffrey I. Webb |
Mach. Learn. | 2 |
| 2015 | Differential Topic ModelsabstractIn applications we may want to compare different document collections: they could have shared content but also different and unique aspects in particular collections. This task has been called comparative text mining or cross-collection modeling. We present a differential topic model for this application that models both topic differences and similarities. For this we use hierarchical Bayesian nonparametric models. Moreover, we found it was important to properly model power-law phenomena in topic-word distributions and thus we used the full Pitman-Yor process rather than just a Dirichlet process. Furthermore, we propose the transformed Pitman-Yor process (TPYP) to incorporate prior knowledge such as vocabulary variations in different collections into the model. To deal with the non-conjugate issue between model prior and likelihood in the TPYP, we thus propose an efficient sampling algorithm using a data augmentation technique based on the multinomial theorem. Experimental results show the model discovers interesting aspects of different collections. We also show the proposed MCMC based algorithm achieves a dramatically reduced test perplexity compared to some existing topic models. Finally, we show our model outperforms the state-of-the-art for document classification/ideology prediction on a number of text collections. Changyou Chen, Wray L. Buntine, Nan Ding 0002, Lexing Xie, Lan Du 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2014 | Bibliographic Analysis with the Citation Network Topic Model
Kar Wai Lim, Wray L. Buntine |
ACML | 2 |
| 2014 | Twitter Opinion Topic Model: Extracting Product Opinions from Tweets by Leveraging Hashtags and Sentiment LexiconabstractAspect-based opinion mining is widely applied to review data to aggregate or summarize opinions of a product, and the current state-of-the-art is achieved with Latent Dirichlet Allocation (LDA)-based model. Although social media data like tweets are laden with opinions, their "dirty" nature (as natural language) has discouraged researchers from applying LDA-based opinion model for product review mining. Tweets are often informal, unstructured and lacking labeled data such as categories and ratings, making it challenging for product opinion mining. In this paper, we propose an LDA-based opinion model named Twitter Opinion Topic Model (TOTM) for opinion mining and sentiment analysis. TOTM leverages hashtags, mentions, emoticons and strong sentiment words that are present in tweets in its discovery process. It improves opinion prediction by modeling the target-opinion interaction directly, thus discovering target specific opinion words, neglected in existing approaches. Moreover, we propose a new formulation of incorporating sentiment prior information into a topic model, by utilizing an existing public sentiment lexicon. This is novel in that it learns and updates with the data. We conduct experiments on 9 million tweets on electronic products, and demonstrate the improved performance of TOTM in both quantitative evaluations and qualitative analysis. We show that aspect-based opinion analysis on massive volume of tweets provides useful opinions on products. Kar Wai Lim, Wray L. Buntine |
CIKM | 2 |
| 2014 | Experiments with non-parametric topic modelsabstractIn topic modelling, various alternative priors have been developed, for instance asymmetric and symmetric priors for the document-topic and topic-word matrices respectively, the hierarchical Dirichlet process prior for the document-topic matrix and the hierarchical Pitman-Yor process prior for the topic-word matrix. For information retrieval, language models exhibiting word burstiness are important. Indeed, this burstiness effect has been show to help topic models as well, and this requires additional word probability vectors for each document. Here we show how to combine these ideas to develop high-performing non-parametric topic models exhibiting burstiness based on standard Gibbs sampling. Experiments are done to explore the behavior of the models under different conditions and to compare the algorithms with previously published. The full non-parametric topic models with burstiness are only a small factor slower than standard Gibbs sampling for LDA and require double the memory, making them very competitive. We look at the comparative behaviour of different models and present some experimental insights. Wray L. Buntine, Swapnil Mishra |
KDD | 1 |
| 2013 | Dependent Normalized Random MeasuresabstractIn this paper we propose two constructions of dependent normalized random measures, a class of nonparametric priors over dependent probability measures. Our constructions, which we call mixed normalized random measures (MNRM) and thinned normalized random measures (TNRM), involve (respectively) weighting and thinning parts of a shared underlying Poisson process before combining them together. We show that both MNRM and TNRM are marginally normalized random measures, resulting in well understood theoretical properties. We develop marginal and slice samplers for both models, the latter necessary for inference in TNRM. In time-varying topic modelling experiments, both models exhibit superior performance over related dependent models such as the hierarchical Dirichlet process and the spatial normalized Gamma process. Changyou Chen, Vinayak A. Rao, Wray L. Buntine, Yee Whye Teh |
ICML (3) | 3 |
| 2013 | Topic Segmentation with a Structured Topic Model
Lan Du 0002, Wray L. Buntine, Mark Johnson 0001 |
HLT-NAACL | 2 |
| 2013 | Improving LDA topic models for microblogs via tweet pooling and automatic labelingabstractTwitter, or the world of 140 characters poses serious challenges to the efficacy of topic models on short, messy text. While topic models such as Latent Dirichlet Allocation (LDA) have a long history of successful application to news articles and academic abstracts, they are often less coherent when applied to microblog content like Twitter. In this paper, we investigate methods to improve topics learned from Twitter content without modifying the basic machinery of LDA; we achieve this through various pooling schemes that aggregate tweets in a data preprocessing step for LDA. We empirically establish that a novel method of tweet pooling by hashtags leads to a vast improvement in a variety of measures for topic coherence across three diverse Twitter datasets in comparison to an unmodified LDA baseline and a variety of pooling schemes. An additional contribution of automatic hashtag labeling further improves on the hashtag pooling results for a subset of metrics. Overall, these two novel schemes lead to significantly improved LDA topic models on Twitter content. Rishabh Mehrotra, Scott Sanner, Wray L. Buntine, Lexing Xie |
SIGIR | 3 |
| 2013 | Introduction: special issue of selected papers of ACML 2012
Zhi-Hua Zhou, Wee Sun Lee, Steven C. H. Hoi, Wray L. Buntine, Hiroshi Motoda |
Mach. Learn. | 4 |
| 2013 | Introduction to the special issue on social web miningabstractNo abstract available. Francesco Bonchi, Wray L. Buntine, Ricard Gavaldà, Shengbo Guo |
ACM Trans. Intell. Syst. Technol. | 2 |
| 2012 | Modelling Sequential Text with an Adaptive Topic Model
Lan Du 0002, Wray L. Buntine, Huidong Jin 0001 |
EMNLP-CoNLL | 2 |
| 2012 | Dependent Hierarchical Normalized Random Measures for Dynamic Topic Modeling
Changyou Chen, Nan Ding 0002, Wray L. Buntine |
ICML | 3 |
| 2012 | Score-Based Bayesian Skill Learning
Shengbo Guo, Scott Sanner, Thore Graepel, Wray L. Buntine |
ECML/PKDD (1) | 4 |
| 2012 | Sequential latent Dirichlet allocation
Lan Du 0002, Wray L. Buntine, Huidong Jin 0001, Changyou Chen |
Knowl. Inf. Syst. | 2 |
| 2011 | Improving Topic Coherence with Regularized Topic ModelsabstractTopic models have the potential to improve search and browsing by extracting useful semantic themes from web pages and other text documents. When learned topics are coherent and interpretable, they can be valuable for faceted browsing, results set diversity analysis, and document retrieval. However, when dealing with small collections or noisy text (e.g. web search result snippets or blog posts), learned topics can be less coherent, less interpretable, and less useful. To overcome this, we propose two methods to regularize the learning of topic models. Our regularizers work by creating a structured prior over words that reflect broad patterns in the external data. Using thirteen datasets we show that both regularizers improve topic coherence and interpretability while learning a faithful representation of the collection of interest. Overall, this work makes topic models more useful across a broader range of text data. Edwin V. Bonilla, Wray L. Buntine |
NIPS | 3 |
| 2011 | Sampling Table Configurations for the Hierarchical Poisson-Dirichlet Process
Changyou Chen, Lan Du 0002, Wray L. Buntine |
ECML/PKDD (1) | 3 |
| 2010 | Sequential Latent Dirichlet Allocation: Discover Underlying Topic Structures within a DocumentabstractUnderstanding how topics within a document evolve over its structure is an interesting and important problem. In this paper, we address this problem by presenting a novel variant of Latent Dirichlet Allocation (LDA): Sequential LDA (SeqLDA). This variant directly considers the underlying sequential structure, i.e., a document consists of multiple segments (e.g., chapters, paragraphs), each of which is correlated to its previous and subsequent segments. In our model, a document and its segments are modelled as random mixtures of the same set of latent topics, each of which is a distribution over words; and the topic distribution of each segment depends on that of its previous segment, the one for first segment will depend on the document topic distribution. The progressive dependency is captured by using the nested two-parameter Poisson Dirichlet process (PDP). We develop an efficient collapsed Gibbs sampling algorithm to sample from the posterior of the PDP. Our experimental results on patent documents show that by taking into account the sequential structure within a document, our SeqLDA model has a higher fidelity over LDA in terms of perplexity (a standard measure of dictionary-based compressibility). The SeqLDA model also yields a nicer sequential topic structure than LDA, as we show in experiments on books such as Melville's "The Whale". Lan Du 0002, Wray L. Buntine, Huidong Jin 0001 |
ICDM | 2 |
| 2010 | Word Features for Latent Dirichlet AllocationabstractWe extend Latent Dirichlet Allocation (LDA) by explicitly allowing for the encoding of side information in the distribution over words. This results in a variety of new capabilities, such as improved estimates for infrequently occurring words, as well as the ability to leverage thesauri and dictionaries in order to boost topic cohesion within and across languages. We present experiments on multi-language topic synchronisation where dictionary information is used to bias corresponding words towards similar topics. Results indicate that our model substantially improves topic cohesion when compared to the standard LDA model. James Petterson, Alexander J. Smola, Tibério S. Caetano, Wray L. Buntine, Shravan M. Narayanamurthy |
NIPS | 4 |
| 2010 | Unsupervised Object Discovery: A ComparisonabstractThe goal of this paper is to evaluate and compare models and methods for learning to recognize basic entities in images in an unsupervised setting. In other words, we want to discover the objects present in the images by analyzing unlabeled data and searching for re-occurring patterns. We experiment with various baseline methods, methods based on latent variable models, as well as spectral clustering methods. The results are presented and compared both on subsets of Caltech256 and MSRC2, data sets that are larger and more challenging and that include more object classes than what has previously been reported in the literature. A rigorous framework for evaluating unsupervised object discovery methods is proposed. Tinne Tuytelaars, Christoph H. Lampert, Matthew B. Blaschko, Wray L. Buntine |
Int. J. Comput. Vis. | 4 |
| 2010 | A segmented topic model based on the two-parameter Poisson-Dirichlet process
Lan Du 0002, Wray L. Buntine, Huidong Jin 0001 |
Mach. Learn. | 2 |
| 2009 | Estimating Likelihoods for Topic Models
Wray L. Buntine |
ACML | 1 |
| 2009 | Kernel Conditional Quantile Estimation via Reduction RevisitedabstractQuantile regression refers to the process of estimating the quantiles of a conditional distribution and has many important applications within econometrics and data mining, among other domains. In this paper, we show how to estimate these conditional quantile functions within a Bayes risk minimization framework using a Gaussian process prior. The resulting non-parametric probabilistic model is easy to implement and allows non-crossing quantile functions to be enforced. Moreover, it can directly be used in combination with tools and extensions of standard Gaussian processes such as principled hyperparameter estimation, sparsification, and quantile regression with input-dependent noise rates. No existing approach enjoys all of these desirable properties. Experiments on benchmark datasets show that our method is competitive with state-of-the-art approaches. Novi Quadrianto, Kristian Kersting, Mark D. Reid, Tibério S. Caetano, Wray L. Buntine |
ICDM | 5 |
| 2009 | Exploring Scale-Induced Feature Hierarchies in Natural ImagesabstractRecently there has been considerable interest in topic models based on the bag-of-features representation of images. The strong independence assumption inherent in the bag-of-features representation is not realistic however: patches often overlap and share underlying image structures. Moreover, important information with respect to relative scales of the features is completely ignored, for the sake of scale invariance. Considering both spatial and scale-based constraints one can derive spatially constrained natural feature hierarchies within images. We explore the use of topic models that build such spatially constrained scale-induced hierarchies of the features in an unsupervised fashion. Our model uses standard topic models as a starting point. We then incorporate information about the hierarchical and spatial relations of the features into the model. We illustrate the hierarchical nature of the resulting models using datasets of natural images, including the MSRC2 dataset as well as a challenging set of images of trees collected from the Internet. Jukka Perkiö, Tinne Tuytelaars, Wray L. Buntine |
ICMLA | 3 |
| 2009 | Guest editors' introduction: special issue of selected papers from ECML PKDD 2009
Alek Kolcz, Dunja Mladenic, Wray L. Buntine, Marko Grobelnik, John Shawe-Taylor |
Data Min. Knowl. Discov. | 3 |
| 2009 | Guest editors' introduction: Special Issue from ECML PKDD 2009
Alek Kolcz, Dunja Mladenic, Wray L. Buntine, Marko Grobelnik, John Shawe-Taylor |
Mach. Learn. | 3 |
| 2008 | Natural language retrieval of grocery productsabstractIn this paper we describe modifications to a natural language grocery retrieval system, introduced in our earlier work. We also compare our system against an off-the-shelf retrieval tool, and show that our system is significantly better for top-ranked retrieval results. Petteri Nurmi, Eemil Lagerspetz, Wray L. Buntine, Patrik Floréen, Joonas Kukkonen, Peter Peltonen |
CIKM | 3 |
| 2008 | Product retrieval for grocery storesabstractWe introduce a grocery retrieval system that maps shopping lists written in natural language into actual products in a grocery store. We have developed the system using nine months of shopping basket data from a large Finnish supermarket. To evaluate the system, we used 70 real shopping lists gathered from customers of the supermarket. Our system achieves over 80% precision for products at rank one, and the precision is around 70% for products at rank 5. Petteri Nurmi, Eemil Lagerspetz, Wray L. Buntine, Patrik Floréen, Joonas Kukkonen |
SIGIR | 3 |
| 2005 | A temporally adaptive content-based relevance ranking algorithmabstractIn information retrieval relevance ranking of the results is one of the most important single tasks there are. There are many diffierent ranking algorithms based on the content of the documents or on some external properties e.g. link structure of html documents.We present a temporally adaptive content-based relevance ranking algorithm that explicitly takes into account the temporal behavior of the underlying statistical properties of the documents in the form of a statistical topic model. more we state that our algorithm can be used on top of any ranking algorithm. Jukka Perkiö, Wray L. Buntine, Henry Tirri |
SIGIR | 2 |
| 2005 | Opportunities from Open Source SearchabstractInternet search has a strong business model that permits a free service to users, so it is difficult to see why, if at all, there should be open source offerings as well. This paper first discusses open source search and a rationale for the computer science community at large to get involved. Because there is no shortage of core open source components for at least some of the tasks involved, the Alvis Consortium is building infrastructure for open source search engines using peer-to-peer and subject specific technology as its core, based on this rationale. We view open source search as a rich future playground in which information extraction and retrieval components can be used and intelligent agents can operate. Wray L. Buntine, Karl Aberer, Ivana Podnar Zarko, Martin Rajman |
Web Intelligence | 1 |
| 2005 | Multi-Faceted Information Retrieval System for Large Scale Email ArchivesabstractWe profile a system for search and analysis of large-scale email archives. The system builds around four facets: content-based search engine, statistical topic model, automatically inferred social networks, and time-series analysis. The facets correspond to the types of information available in email data. The presented system allows chaining or combining the facets flexibly. Results of one facet may be used as input to another yielding remarkable combinatorial power. In information retrieval point of view, the system provides support for exploration, approximate textual searches and data visualization. We present some experimental results based on a large real-world email corpus. Jukka Perkiö, Ville H. Tuulos, Wray L. Buntine, Henry Tirri |
Web Intelligence | 3 |
| 2004 | Automated Synthesis of Data Analysis Programs: Learning in Logic
Wray L. Buntine |
ILP | 1 |
| 2004 | Applying Discrete PCA in Data Analysis
Wray L. Buntine, Aleks Jakulin |
UAI | 1 |
| 2004 | A Scalable Topic-Based Open Source Search EngineabstractSite-based or topic-specific search engines work with mixed success because of the general difficulty of the information retrieval task, and the lack of good link information to allow authorities to be identified. We are advocating an open source approach to the problem due to its scope and need for software components. We have adopted a topic-based search engine because it represents the next generation of capability. This paper outlines our scalable system for site-based or topic-specific search, and demonstrates the developing system on a small 250,000 document collection of EU and UN web pages. Wray L. Buntine, Jaakko Löfström, Jukka Perkiö, Sami Perttu, Vladimir Poroshin, Tomi Silander, Henry Tirri, Antti J. Tuominen, Ville H. Tuulos |
Web Intelligence | 1 |
| 2004 | Exploring Independent Trends in a Topic-Based Search EngineabstractTopic-based search engines are an alternative to simple keyword search engines that are common in today's intranets. The temporal behaviour of the topics in a topic model based search engine can be used for trend analysis, which is an important research goal on its own. We apply topic modelling to an online financial newspaper data and show that some of the trends in the topics are consistent with common understanding. Jukka Perkiö, Wray L. Buntine, Sami Perttu |
Web Intelligence | 2 |
| 2002 | Variational Extensions to EM and Multinomial PCA
Wray L. Buntine |
ECML | 1 |
| 2002 | Automatic Derivation of Statistical Algorithms: The EM Family and BeyondabstractMachine learning has reached a point where many probabilistic meth- ods can be understood as variations, extensions and combinations of a much smaller set of abstract themes, e.g., as different instances of the EM algorithm. This enables the systematic derivation of algorithms cus- tomized for different models. Here, we describe the AUTO BAYES sys- tem which takes a high-level statistical model specification, uses power- ful symbolic techniques based on schema-based program synthesis and computer algebra to derive an efficient specialized algorithm for learning that model, and generates executable code implementing that algorithm. This capability is far beyond that of code collections such as Matlab tool- boxes or even tools for model-independent optimization such as BUGS for Gibbs sampling: complex new algorithms can be generated with- out new programming, algorithms can be highly specialized and tightly crafted for the exact structure of the model and data, and efficient and commented code can be generated for different languages or systems. We present automatically-derived algorithms ranging from closed-form solutions of Bayesian textbook problems to recently-proposed EM algo- rithms for clustering, regression, and a multinomial form of PCA. 1 Automatic Derivation of Statistical Algorithms Overview. We describe a symbolic program synthesis system which works as a “statistical algorithm compiler:” it compiles a statistical model specification into a custom algorithm design and from that further down into a working program implementing the algorithm design. This system, AUTOBAYES, can be loosely thought of as “part theorem prover, part Mathematica, part learning textbook, and part Numerical Recipes.” It provides much more flexibility than a fixed code repository such as a Matlab toolbox, and allows the creation of efficient algorithms which have never before been implemented, or even written down. AUTOBAYES is intended to automate the more routine application of complex methods in novel contexts. For example, recent multinomial extensions to PCA [2, 4] can be derived in this way. The algorithm design problem. Given a dataset and a task, creating a learning method can be characterized by two main questions: 1. What is the model? 2. What algorithm will optimize the model parameters? The statistical algorithm (i.e., a parameter optimization algorithm for the statistical model) can then be implemented manually. The system in this paper answers the algorithm question given that the user has chosen a model for the data,and continues through to implementation. Performing this task at the state-of-the-art level requires an intertwined meld of probability theory, computational mathematics, and software engineering. However, a number of factors unite to allow us to solve the algorithm design problem computationally: 1. The existence of fundamental building blocks (e.g., standardized probability distributions, standard optimization procedures, and generic data structures). 2. The existence of common representations (i.e., graphical models [3, 13] and program schemas). 3. The formalization of schema applicability constraints as guards.1 The challenges of algorithm design. The design problem has an inherently combinatorial nature, since subparts of a function may be optimized recursively and in different ways. It also involves the use of new data structures or approximations to gain performance. As the research in statistical algorithms advances, its creative focus should move beyond the ultimately mechanical aspects and towards extending the abstract applicability of already existing schemas (algorithmic principles like EM), improving schemas in ways that gener- alize across anything they can be applied to, and inventing radically new schemas. 2 Combining Schema-based Synthesis and Bayesian Networks with 0 < n_points; 1 model mog as ’Mixture of Gaussians’; with 0 < nclasses with nclasses << n_points; with 1 = sum(I := 1..n_classes, phi(I)); 7 double phi(1..nclasses) as ’weights’ 8 9 double mu(1..nclasses); 9 double sigma(1..n_classes); 2 const int npoints as ’nr. of data points’ 3 4 const int nclasses := 3 as ’nr. classes’ 5 6 Statistical Models. Externally, AUTOBAYES has the look and feel of a compiler. Users specify their model of interest in a high-level specification language (as opposed to a program- ming language). The figure shows the specification of the mixture of Gaus- sians example used throughout this paper.2 Note the constraint that the sum of the class probabilities must equal one (line 8) along with others (lines 3 and 5) that make optimization of the model well-defined. Also note the ability to specify assumptions of the kind in line 6, which may be used by some algorithms. The last line specifies the goal 10 int c(1..npoints) as ’class labels’; 11 c ˜ disc(vec(I := 1..nclasses, phi(I))); 12 data double x(1..n_points) as ’data’; 13 x(I) ˜ gauss(mu(c(I)), sigma(c(I))); 14 max pr(x| phi,mu,sigma ) wrt phi,mu,sigma ; inference task: maximize the conditional probability pr rameters Alexander G. Gray, Bernd Fischer 0002, Johann Schumann, Wray L. Buntine |
NIPS | 4 |
| 2001 | Learning as applied to stochastic optimization for standard-cellplacementabstractStochastic combinatorial optimization techniques, such as simulated annealing and genetic algorithms, have become increasingly important in design automation as the size of design problems have grown and the design objectives have become increasingly complex. However, stochastic algorithms are, often slow since a large number of random design perturbations are required to achieve an acceptable result-they have no built-in "intelligence". In this paper, it is shown that statistical learning techniques can improve the quality of results and reduce the number of expensive cost-function evaluations for stochastic optimization for a particular solution quality. In particular, simulated annealing was selected as a representative stochastic optimization approach and a two-dimensional cell-based layout placement problem was used to evaluate the utility of such a learning-based approach. In this paper, we used regression to learn the properties of the solution space and tested the trained algorithm on a number of examples to demonstrate the improvement gained. A general response model is constructed by learning from the annealing of benchmark circuits. This model is then used in the trained simulated annealing, which returns significantly better annealing quality than the untrained algorithm for the same number of moves in the solution space. The annealing quality improvement was 15%-43% for the set of examples used in training and 7%-21% when the trained algorithm was applied to new examples. Lixin Su, Wray L. Buntine, A. Richard Newton, Bradley S. Peters |
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst. | 2 |
| 1999 | Towards Automated Synthesis of Data Mining ProgramsabstractCode synthesis is routinely used in industry to generate GUIs, form filling applications, and database support code and is even used with COBOL. In this paper we consider the question of whether code synthesis could also be applied to the data mining phase of knowledge discovery. We view this as a rapid prototyping method. Rapid prototyping of statistical data analysis algorithms would allow experienced analysts to experiment with different statistical models before choosing one, but without requiring prohibitively expensive programming efforts. It would also smooth the steep learning curve often faced by novice users of data mining tools and libraries. Finally, it would accelerate dissemination of essential research results and the development of applications. In this paper, we present a framework and the basic software for the automated synthesis of data analysis programs. We use a specification language that generalizes Bayesian networks, a popular notation used in many communities... Wray L. Buntine, Bernd Fischer 0002, Thomas Pressburger |
KDD | 1 |
| 1998 | Learning as applied to stochastic optimization for standard cell placementabstractAlthough becoming increasingly important, stochastic algorithms are often slow since a large number of random design perturbations are required to achieve an acceptable result-they have no built-in "intelligence". In this work, we used regression to learn the swap evaluation function while simulated annealing is applied to 2D standard-cell placement problem. The learned evaluation function is then applied to the trained simulated annealing algorithm (TSA). The annealing quality improvement of TSA was 15%/spl sim/43% for the set of examples used in learning and 7%/spl sim/21% for new examples. With the same amount of CPU time, TSA could improve the annealing quality by up to 28% for some benchmark circuits we tested. In addition the use of the evaluation function successfully predicted the effect of the windowed sampling technique and derived the informally accepted advantages of windowing from the test set automatically. Lixin Su, Wray L. Buntine, A. Richard Newton, Bradley S. Peters |
ICCD | 2 |
| 1998 | Analysing Rock Samples for the Mars Lander
Jonathan Oliver, Ted Roush, Paul Gazis, Wray L. Buntine, Rohan A. Baxter, Steven R. Waterhouse |
KDD | 4 |
| 1997 | Adaptive methods for netlist partitioningabstractAn algorithm that remains in use at the core of many partitioning systems is the Kemighan-Lin algorithm and a variant the Fidducia-Matheysses (FM) algorithm. To understand the FM algorithm we applied principles of data engineering where visualization and statistical analysis are used to analyze the run-time behavior. We identified two improvements to the algorithm which, without clustering or an improved heuristic function, bring the performance of the algorithm near that of more sophisticated algorithms. One improvement is based on the observation, explored empirically, that the full passes in the FM algorithm appear comparable to a stochastic local restart in the search. We motivate this observation with a discussion of recent improvements in Monte Carlo Markov Chain methods in statistics. The other improvement is based on the observation that when an FM-like algorithm is run 20 times and the best run chosen, the performance trace of the algorithm on earlier runs is useful data for learning when to abort a later run. These improvements, implemented with a simple adaptive scheme, are orthogonal to techniques used in state-of-the-art implementations, and therefore should be applicable to other VLSI optimization algorithms. Wray L. Buntine, Lixin Su, A. Richard Newton, Andrew Mayer |
ICCAD | 1 |
| 1996 | A Guide to the Literature on Learning Probabilistic Networks from DataabstractThe literature review presented discusses different methods under the general rubric of learning Bayesian networks from data, and includes some overlapping work on more general probabilistic networks. Connections are drawn between the statistical, neural network, and uncertainty communities, and between the different methodological communities, such as Bayesian, description length, and classical statistics. Basic concepts for learning and Bayesian networks are introduced and methods are then reviewed. Methods are discussed for learning parameters of a probabilistic network, for learning the structure, and for learning hidden variables. The article avoids formal definitions and theorems, as these are plentiful in the literature, and instead illustrates key concepts with simplified examples. Wray L. Buntine |
IEEE Trans. Knowl. Data Eng. | 1 |
| 1995 | Intelligent Instruments: Discovering How to Turn Spectral Data into Information
Wray L. Buntine, Tarang Patel |
KDD | 1 |
| 1995 | Chain graphs for learning
Wray L. Buntine |
UAI | 1 |
| 1994 | On Solving Equations and DisequationsabstractWe are interested in the problem of solving a system 〈s l = t l : 1 ≤ i ≤ n, p j ≠ q j : 1 ≤ j ≤ m 〉 of equations and disequations, also known as disunification . Solutions to disunification problems are substitutions for the variables of the problem that make the two terms of each equation equal, but leave those of the disequations different. We investigate this in both algebraic and logical contexts where equality is defined by an equational theory and more generally by a definitive clause equality theory E. We show how E-disunification can be reduced to E-unification, that is, solving equations only, and give a disunification algorithm for theories given a unification algorithm. In fact, this result shows that for theories in which the solutions of all unification problems can also be represented finitely. We sketch how disunification can be applied to handle negation in logic programming with equality in a similar style to Colmerauer's logic programming with rational trees, and to represent many solutions to AC-unification problems by a few solutions to ACI-disunification problems. Wray L. Buntine, Hans-Jürgen Bürckert |
J. ACM | 1 |
| 1994 | Operations for Learning with Graphical ModelsabstractThis paper is a multidisciplinary review of empirical, statistical learning from a graphical model perspective. Well-known examples of graphical models include Bayesian networks, directed graphs representing a Markov chain, and undirected networks representing a Markov field. These graphical models are extended to model data analysis and empirical learning using the notation of plates. Graphical operations for simplifying and manipulating a problem are provided including decomposition, differentiation, andthe manipulation of probability models from the exponential family. Two standard algorithm schemas for learning are reviewed in a graphical framework: Gibbs sampling and the expectation maximizationalgorithm. Using these operations and schemas, some popular algorithms can be synthesized from their graphical specification. This includes versions of linear regression, techniques for feed-forward networks, and learning Gaussian and discrete Bayesian networks from data. The paper concludes by sketching some implications for data analysis and summarizing how some popular algorithms fall within the framework presented. The main original contributions here are the decompositiontechniques and the demonstration that graphical models provide a framework for understanding and developing complex learning algorithms. Wray L. Buntine |
J. Artif. Intell. Res. | 1 |
| 1994 | Guest Editorial
Katharina Morik, Francesco Bergadano, Wray L. Buntine |
Mach. Learn. | 3 |
| 1994 | Computing second derivatives in feed-forward networks: a reviewabstractThe calculation of second derivatives is required by recent training and analysis techniques of connectionist networks, such as the elimination of superfluous weights, and the estimation of confidence intervals both for weights and network outputs. We review and develop exact and approximate algorithms for calculating second derivatives. For networks with |w| weights, simply writing the full matrix of second derivatives requires O(|w|(2)) operations. For networks of radial basis units or sigmoid units, exact calculation of the necessary intermediate terms requires of the order of 2h+2 backward/forward-propagation passes where h is the number of hidden units in the network. We also review and compare three approximations (ignoring some components of the second derivative, numerical differentiation, and scoring). The algorithms apply to arbitrary activation functions, networks, and error functions. Wray L. Buntine, Andreas S. Weigend |
IEEE Trans. Neural Networks | 1 |
| 1992 | A Further Comparison of Splitting Rules for Decision-Tree InductionabstractOne approach to learning classification rules from examples is to build decision trees. A review and comparison paper by Mingers (Mingers, 1989) looked at the first stage of tree building, which uses a “splitting rule” to grow trees with a greedy recursive partitioning algorithm. That paper considered a number of different measures and experimentally examined their behavior on four domains. The main conclusion was that a random splitting rule does not significantly decrease classificational accuracy. This note suggests an alternative experimental method and presents additional results on further domains. Our results indicate that random splitting leads to increased error. These results are at variance with those presented by Mingers. Wray L. Buntine, Tim Niblett |
Mach. Learn. | 1 |
| 1991 | Classifiers: A Theoretical and Empirical Study
Wray L. Buntine |
IJCAI | 1 |
| 1991 | Some Properties of Plausible Reasoning
Wray L. Buntine |
UAI | 1 |
| 1991 | Theory Refinement on Bayesian Networks
Wray L. Buntine |
UAI | 1 |
| 1990 | Myths and Legends in Learning Classification Rules
Wray L. Buntine |
AAAI | 1 |
| 1989 | Learning Classification Rules Using Bayes
Wray L. Buntine |
ML | 1 |
| 1989 | A Critique of the Valiant Model
Wray L. Buntine |
IJCAI | 1 |
| 1989 | Inductive knowledge acquisition and induction methodologies
Wray L. Buntine |
Knowl. Based Syst. | 1 |
| 1988 | Machine Invention of First Order Predicates by Inverting Resolution
Stephen H. Muggleton, Wray L. Buntine |
ML | 2 |
| 1988 | Generalized Subsumption and Its Applications to Induction and Redundancy
Wray L. Buntine |
Artif. Intell. | 1 |
| 1988 | Decision tree induction systems: A Bayesian analysis
Wray L. Buntine |
Int. J. Approx. Reason. | 1 |
| 1987 | Decision Tree Induction Systems: A Bayesian Analysis
Wray L. Buntine |
UAI | 1 |
| 1987 | Induction of Horn Clauses: Methods and the Plausible Generalization Algorithm
Wray L. Buntine |
Int. J. Man Mach. Stud. | 1 |
| 1986 | Generalised Subsumption and its Applications to Induction and Redundancy
Wray L. Buntine |
ECAI | 1 |
| 1976 | Design rule checking and analysis of IC mask designsabstractAn efficient method of producing logical combinations of integrated circuit (IC) masks in numerical form leads to a generalized design rule checking program. The union (OR), intersection (AND) and the complements, as well as topological classification and simple geometric operations, are provided through a set of LOGical MASk Checking (LOGMASC) commands, allowing the designer to construct, for the given IC technology, a tailored set of design rule checks. These range from simple tolerance checks to complex analysis of mask geometries based on pattern recognition. Wray L. Buntine, Bryan Preas |
DAC | 1 |
| 1976 | Automatic circuit analysis based on mask informationabstractA Circuit MAsk Translator (CMAT) code has been developed which converts integrated circuit mask information into a circuit schematic. Logical operations, pattern recognition, and special functions are used to identify and interconnect diodes, transistors, capacitors, and resistances. The circuit topology provided by the translator is compatible with the input required for a circuit analysis program. Bryan Preas, Wray L. Buntine, Charles W. Gwyn |
DAC | 2 |