Hai Leong Chieu

dblp:38/4132 · DBLP profile ↗
← Back
35ranked-venue papers
9as first author
5since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 31 · 7 first-author · 5 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
21 papers
Information extraction and text analysis · 39% Language models and text generation · 31% Trustworthy machine learning · 10%
Databases, data mining, and information retrieval
6 papers
Web and social media mining · 61% Data mining · 35% Information retrieval · 4%

Topics — the 30 heaviest of 50, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis
stance detection
1.122023
Guiding Computational Stance Detection with Expanded Stance Triangle Framework · ACL (1) 2023
Coupled Hierarchical Transformer for Stance-Aware Rumor Verification in Social Media Conversations · EMNLP (1) 2020
Natural language and speech › Language models and text generation
controllable text generation
0.912025
Colloquial Singaporean English Style Transfer with Fine-Grained Explainable Control · ACL (1) 2025
Natural language and speech › Language models and text generation
retrieval-augmented generation
0.912025
Measuring and Enhancing Trustworthiness of LLMs in RAG through Grounded Attributions and Learning to Refuse · ICLR 2025
Natural language and speech › Language models and text generation › controllable text generation
text style transfer
0.912025
Colloquial Singaporean English Style Transfer with Fine-Grained Explainable Control · ACL (1) 2025
Web and social media mining › misinformation detection
rumor detection
0.412020
Interpretable Rumor Detection in Microblogs by Attending to User Interactions · AAAI 2020
Machine learning › Transfer learning and domain adaptation
domain adaptation
0.422016
A General Regularization Framework for Domain Adaptation · EMNLP 2016
Domain Adaptation for Coreference Resolution: An Adaptive Ensemble Approach · EMNLP-CoNLL 2012
Web and social media mining › social media user profiling
user attribute inference
0.412019
Twitter Homophily: Network Based Prediction of User's Occupation · ACL (1) 2019
Natural language and speech › Information extraction and text analysis › syntactic parsing › dependency parsing
cross-lingual dependency parsing
0.312017
Universal Dependencies Parsing for Colloquial Singaporean English · ACL (1) 2017
Natural language and speech › Information extraction and text analysis › syntactic parsing
dependency parsing
0.312017
Universal Dependencies Parsing for Colloquial Singaporean English · ACL (1) 2017
Natural language and speech › Language models and text generation › text summarization
sentence compression
0.312017
Can Syntax Help? Improving an LSTM-based Sentence Compression Model for New Domains · ACL (1) 2017
Natural language and speech › Information extraction and text analysis
syntactic parsing
0.312017
Universal Dependencies Parsing for Colloquial Singaporean English · ACL (1) 2017
Machine learning › Probabilistic and Bayesian machine learning › structured prediction
conditional random field
0.322014
Conditional random field with high-order dependencies for sequence labeling and segmentation · J. Mach. Learn. Res. 2014
Conditional Random Fields with High-Order Features for Sequence Labeling · NIPS 2009
Natural language and speech › Information extraction and text analysis
sequence labeling
0.322014
Conditional random field with high-order dependencies for sequence labeling and segmentation · J. Mach. Learn. Res. 2014
Conditional Random Fields with High-Order Features for Sequence Labeling · NIPS 2009
Natural language and speech › Information extraction and text analysis
text classification
0.222023
Guiding Computational Stance Detection with Expanded Stance Triangle Framework · ACL (1) 2023
Active Learning for Probabilistic Hypotheses Using the Maximum Gibbs Error Criterion · NIPS 2013
Natural language and speech › Information extraction and text analysis › relation extraction
domain adaptation for relation extraction
0.212014
Robust Domain Adaptation for Relation Extraction via Clustering Consistency · ACL (1) 2014
Natural language and speech › Information extraction and text analysis
relation extraction
0.212014
Robust Domain Adaptation for Relation Extraction via Clustering Consistency · ACL (1) 2014
Natural language and speech › Information extraction and text analysis
named entity recognition
0.232013
Domain adaptive bootstrapping for named entity recognition · EMNLP 2009
Active Learning for Probabilistic Hypotheses Using the Maximum Gibbs Error Criterion · NIPS 2013
Teaching a Weaker Classifier: Named Entity Recognition on Upper Case Text · ACL 2002
Data mining
clustering
0.222012
A Split-Merge Framework for Comparing Clusterings · ICML 2012
Cooled and Relaxed Survey Propagation for MRFs · NIPS 2007
Machine learning › Efficient and distributed learning
active learning
0.212013
Active Learning for Probabilistic Hypotheses Using the Maximum Gibbs Error Criterion · NIPS 2013
Machine learning › Probabilistic and Bayesian machine learning › experimental design › bayesian experimental design
bayesian active learning
0.212013
Active Learning for Probabilistic Hypotheses Using the Maximum Gibbs Error Criterion · NIPS 2013
Machine learning › Efficient and distributed learning › active learning › active data collection
pool-based active learning
0.212013
Active Learning for Probabilistic Hypotheses Using the Maximum Gibbs Error Criterion · NIPS 2013
Natural language and speech › Information extraction and text analysis
coreference resolution
0.112012
Domain Adaptation for Coreference Resolution: An Adaptive Ensemble Approach · EMNLP-CoNLL 2012
Machine learning › Learning theory › classification
f-measure optimization
0.112012
Optimizing F-measure: A Tale of Two Approaches · ICML 2012
Machine learning › Optimization for machine learning › optimization › metric optimization
non-decomposable performance measure optimization
0.112012
Optimizing F-measure: A Tale of Two Approaches · ICML 2012
Data mining › clustering › clustering evaluation
clustering comparison
0.112012
A Split-Merge Framework for Comparing Clusterings · ICML 2012
Machine learning › Trustworthy machine learning › interpretability › attention analysis
attention-based explanation
0.112020
Interpretable Rumor Detection in Microblogs by Attending to User Interactions · AAAI 2020
Machine learning › Deep learning architectures and training › transformer
hierarchical transformer
0.112020
Coupled Hierarchical Transformer for Stance-Aware Rumor Verification in Social Media Conversations · EMNLP (1) 2020
Machine learning › Trustworthy machine learning
interpretability
0.112020
Interpretable Rumor Detection in Microblogs by Attending to User Interactions · AAAI 2020
Knowledge, reasoning and agents › Knowledge representation and reasoning › knowledge acquisition › knowledge extraction
unsupervised information extraction
0.112011
Unsupervised Information Extraction with Distributional Prior Knowledge · EMNLP 2011
Machine learning › Graph learning › graph neural network
graph convolutional network
0.112019
Twitter Homophily: Network Based Prediction of User's Occupation · ACL (1) 2019

Methods — techniques the papers use, named apart from their topics

in-context learning · 0.9fine-grained control · 0.9alignment · 0.9transformer · 0.9multi-head attention · 0.9hierarchical attention · 0.9stance triangle framework · 0.7data augmentation · 0.7hierarchical transformer · 0.4attention · 0.4social homophily · 0.4graph convolutional network · 0.4split-merge framework · 0.1survey propagation · 0.1sum-product algorithm · 0.1tree-reweighted belief propagation · 0.1relaxed survey propagation · 0.1sentence extraction · 0.0
YearPublicationVenuePosition
2025 Colloquial Singaporean English Style Transfer with Fine-Grained Explainable Control
abstract
Jinggui Liang, Dung Vo, Yap Hong Xian, Hai Leong Chieu, Kian Ming A. Chai, Jing Jiang, Lizi Liao. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Jinggui Liang, Dung Vo 0002, Yap Hong Xian, Hai Leong Chieu, Kian Ming A. Chai, Jing Jiang 0001, Lizi Liao
ACL (1)4
2025 Measuring and Enhancing Trustworthiness of LLMs in RAG through Grounded Attributions and Learning to Refuse
abstract
LLMs are an integral component of retrieval-augmented generation (RAG) systems. While many studies focus on evaluating the overall quality of end-to-end RAG systems, there is a gap in understanding the appropriateness of LLMs for the RAG task. To address this, we introduce Trust-Score, a holistic metric that evaluates the trustworthiness of LLMs within the RAG framework. Our results show that various prompting methods, such as in-context learning, fail to effectively adapt LLMs to the RAG task as measured by Trust-Score. Consequently, we propose Trust-Align, a method to align LLMs for improved Trust-Score performance. 26 out of 27 models aligned using Trust-Align substantially outperform competitive baselines on ASQA, QAMPARI, and ELI5. Specifically, in LLaMA-3-8b, Trust-Align outperforms FRONT on ASQA (↑12.56), QAMPARI (↑36.04), and ELI5 (↑17.69). Trust-Align also significantly enhances models’ ability to correctly refuse and provide quality citations. We also demonstrate the effectiveness of Trust-Align across different open-weight models, including the LLaMA series (1b to 8b), Qwen-2.5 series (0.5b to 7b), and Phi3.5 (3.8b). We release our code at https://github.com/declare-lab/trust-align.
Maojia Song, Shang Hong Sim, Rishabh Bhardwaj, Hai Leong Chieu, Navonil Majumder, Soujanya Poria
ICLR4
2023 Guiding Computational Stance Detection with Expanded Stance Triangle Framework
abstract
Stance detection determines whether the author of a piece of text is in favor of, against, or neutral towards a specified target, and can be used to gain valuable insights into social media.The ubiquitous indirect referral of targets makes this task challenging, as it requires computational solutions to model semantic features and infer the corresponding implications from a literal statement.Moreover, the limited amount of available training data leads to subpar performance in out-of-domain and cross-target scenarios, as data-driven approaches are prone to rely on superficial and domain-specific features.In this work, we decompose the stance detection task from a linguistic perspective, and investigate key components and inference paths in this task.The stance triangle is a generic linguistic framework previously proposed to describe the fundamental ways people express their stance.We further expand it by characterizing the relationship between explicit and implicit objects.We then use the framework to extend one single training corpus with additional annotation.Experimental results show that strategically-enriched data can significantly improve the performance on out-of-domain and cross-target evaluation.
Zhengyuan Liu, Yong Keong Yap, Hai Leong Chieu, Nancy F. Chen
ACL (1)3
2023 Interpretable Sock Puppet Attribution
abstract
The intentional spread of misinformation can have serious consequences in our society. This motivates us to address the task of identifying sock puppet accounts (i.e. fabricated online personas) created by individuals or organizations with the intention of deceiving their target audience. By approaching the problem as authorship attribution, we develop a sock puppet detection framework that relies solely on the texts posted by sock puppets without using any meta-information. We employed a large pre-trained language model and used interpretability methods to extract linguistic cues to answer our research question. We curated a high-quality sock puppet dataset to enable a comprehensive study for the research communities where existing datasets may include false positives due to their approximate approaches in identifying sock puppets. This dataset enables us to study the research question of whether authorship attribution works on sock puppets’ written texts. Our experiment shows that our method remarkably outperformed human annotators by 31.8% points. Furthermore, we employed an interpretability method, Integrated Gradient, to extract linguistic cues from our model for explanations of why an account is a sock puppet or not.
Chun-Wei Seah, Chia-Yu Hung, Hai Leong Chieu, Roy Ka-Wei Lee
IEEE Big Data4
2021 Cross-Topic Rumor Detection using Topic-Mixtures
abstract
There has been much interest in rumor detection using deep learning models in recent years.A well-known limitation of deep learning models is that they tend to learn superficial patterns, which restricts their generalization ability.We find that this is also true for cross-topic rumor detection.In this paper, we propose a method inspired by the "mixture of experts" paradigm.We assume that the prediction of the rumor class label given an instance is dependent on the topic distribution of the instance.After deriving a vector representation for each topic, given an instance, we derive a "topic mixture" vector for the instance based on its topic distribution.This topic mixture is combined with the vector representation of the instance itself to make rumor predictions.Our experiments show that our proposed method can outperform two baseline debiasing methods in a cross-topic setting.In a synthetic setting when we removed topic-specific words, our method also works better than the baselines, showing that our method does not rely on superficial features.
Xiaoying Ren, Jing Jiang 0001, Ling Min Serena Khoo, Hai Leong Chieu
EACL4
2020 Interpretable Rumor Detection in Microblogs by Attending to User Interactions
abstract
We address rumor detection by learning to differentiate between the community's response to real and fake claims in microblogs. Existing state-of-the-art models are based on tree models that model conversational trees. However, in social media, a user posting a reply might be replying to the entire thread rather than to a specific user. We propose a post-level attention model (PLAN) to model long distance interactions between tweets with the multi-head attention mechanism in a transformer network. We investigated variants of this model: (1) a structure aware self-attention model (StA-PLAN) that incorporates tree structure information in the transformer network, and (2) a hierarchical token and post-level attention model (StA-HiTPLAN) that learns a sentence representation with token-level self-attention. To the best of our knowledge, we are the first to evaluate our models on two rumor detection data sets: the PHEME data set as well as the Twitter15 and Twitter16 data sets. We show that our best models outperform current state-of-the-art models for both data sets. Moreover, the attention mechanism allows us to explain rumor detection predictions at both token-level and post-level.
Ling Min Serena Khoo, Hai Leong Chieu, Zhong Qian 0001, Jing Jiang 0001
AAAI2
2020 Coupled Hierarchical Transformer for Stance-Aware Rumor Verification in Social Media Conversations
abstract
The prevalent use of social media enables rapid spread of rumors on a massive scale, which leads to the emerging need of automatic rumor verification (RV). A number of previous studies focus on leveraging stance classification to enhance RV with multi-task learning (MTL) methods. However, most of these methods failed to employ pre-trained contextualized embeddings such as BERT, and did not exploit inter-task dependencies by using predicted stance labels to improve the RV task. Therefore, in this paper, to extend BERT to obtain thread representations, we first propose a Hierarchical Transformer, which divides each long thread into shorter subthreads, and employs BERT to separately represent each subthread, followed by a global Transformer layer to encode all the subthreads. We further propose a Coupled Transformer Module to capture the inter-task interactions and a Post-Level Attention layer to use the predicted stance labels for RV, respectively. Experiments on two benchmark datasets show the superiority of our Coupled Hierarchical Transformer model over existing MTL approaches.
Jianfei Yu, Jing Jiang 0001, Ling Min Serena Khoo, Hai Leong Chieu
EMNLP (1)4
2020 Probabilistic Decision Modeling in Social Networks
abstract
Bayesian approaches have been successfully applied in social network analysis to study group behaviors such as online information dissemination and voting pattern. The focus has been on estimating the structure and strength of peer influence and its impact on the decisions of an individual. Less attention has been given to incorporating contextual information and individuals' hidden characteristics (or bias). In this work, we examine the social dynamics where social influence and contextual information play pivotal roles in driving one's decision. We design a probabilistic graphical model called CLAP to understand users' decision behavior in a social network, with an emphasis on both social-level and individual-level factors. To this end, the proposed model introduces hidden bias states associated with each actor and jointly estimates each actor's hidden bias state together with the social influence network. We demonstrate the effectiveness of CLAP on two types of social networks, a real-world US Congress network where senators vote on new bills, and online Twitter networks where users debate on the effectiveness of vaccine and lockdown policy during COVID-19. The experiment results show that CLAP outperforms state-of-the-art game theoretic approaches in predicting user decision. Further, the estimated social influence networks by CLAP has high edge homogeniety ratios.
Tangqing Li, Wynne Hsu, Mong-Li Lee, Hai Leong Chieu
ICTAI4
2019 Twitter Homophily: Network Based Prediction of User's Occupation
abstract
In this paper, we investigate the importance of social network information compared to content information in the prediction of a Twitter user's occupational class.We show that the content information of a user's tweets, the profile descriptions of a user's follower/following community, and the user's social network provide useful information for classifying a user's occupational group.In our study, we extend an existing dataset for this problem, and we achieve significantly better performance by using social network homophily that has not been fully exploited in previous work.In our analysis, we found that by using the graph convolutional network to exploit social homophily, we can achieve competitive performance on this dataset with just a small fraction of the training data.
Rishabh Bhardwaj, Wei Lu 0011, Hai Leong Chieu, Xinghao Pan, Ni Yi Puay
ACL (1)4
2017 Can Syntax Help? Improving an LSTM-based Sentence Compression Model for New Domains
abstract
Liangguo Wang, Jing Jiang, Hai Leong Chieu, Chen Hui Ong, Dandan Song, Lejian Liao. Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2017.
Liangguo Wang, Jing Jiang 0001, Hai Leong Chieu, Chen Hui Ong, Lejian Liao
ACL (1)3
2017 Universal Dependencies Parsing for Colloquial Singaporean English
abstract
Singlish can be interesting to the ACL community both linguistically as a major creole based on English, and computationally for information extraction and sentiment analysis of regional social media.We investigate dependency parsing of Singlish by constructing a dependency treebank under the Universal Dependencies scheme, and then training a neural network model by integrating English syntactic knowledge into a state-ofthe-art parser trained on the Singlish treebank.Results show that English knowledge can lead to 25% relative error reduction, resulting in a parser of 84.47% accuracies.To the best of our knowledge, we are the first to use neural stacking to improve cross-lingual dependency parsing on low-resource languages.We make both our annotation and parser available for further research.
Hongmin Wang, Yue Zhang 0004, GuangYong Leonard Chan, Jie Yang 0039, Hai Leong Chieu
ACL (1)5
2016 A General Regularization Framework for Domain Adaptation
abstract
We propose a domain adaptation framework, and formally prove that it generalizes the feature augmentation technique in (Daumé III, 2007) and the multi-task regularization framework in (Evgeniou and Pontil, 2004).We show that our framework is strictly more general than these approaches and allows practitioners to tune hyper-parameters to encourage transfer between close domains and avoid negative transfer between distant ones.
Wei Lu 0011, Hai Leong Chieu, Jonathan Löfgren
EMNLP2
2016 Learning to Capitalize with Character-Level Recurrent Neural Networks: An Empirical Study
abstract
In this paper, we investigate case restoration for text without case information.Previous such work operates at the word level.We propose an approach using character-level recurrent neural networks (RNN), which performs competitively compared to language modeling and conditional random fields (CRF) approaches.We further provide quantitative and qualitative analysis on how RNN helps improve truecasing.
Raymond Hendy Susanto, Hai Leong Chieu, Wei Lu 0011
EMNLP2
2015 Troll detection by domain-adapting sentiment analysis
Chun-Wei Seah, Hai Leong Chieu, Kian Ming A. Chai, Loo-Nin Teow, Lee Wei Yeong
FUSION2
2014 Robust Domain Adaptation for Relation Extraction via Clustering Consistency
abstract
We propose a two-phase framework to adapt existing relation extraction classifiers to extract relations for new target domains. We address two challenges: negative transfer when knowledge in source domains is used without considering the differences in relation distributions; and lack of adequate labeled samples for rarer relations in the new domain, due to a small labeled data set and imbalance relation distributions. Our framework leverages on both labeled and unlabeled data in the target domain. First, we determine the relevance of each source domain to the target domain for each relation type, using the consistency between the clustering given by the target domain labels and the clustering given by the predictors trained for the source domain. To overcome the lack of labeled samples for rarer relations, these clusterings operate on both the labeled and unlabeled data in the target domain. Second, we trade-off between using relevance-weighted sourcedomain predictors and the labeled target data. Again, to overcome the imbalance distribution, the source-domain predictors operate on the unlabeled target data. Our method outperforms numerous baselines and a weakly-supervised relation extraction method on ACE 2004 and YAGO. © 2014 Association for Computational Linguistics.
Minh Luan Nguyen, Ivor W. Tsang, Kian Ming A. Chai, Hai Leong Chieu
ACL (1)4
2014 Language modeling with sum-product networks
abstract
Sum product networks (SPNs) are a new class of deep probabilistic models. They can contain multiple hidden layers while keeping their inference and training times tractable. An SPN consists of interleaving layers of sum nodes and product nodes. A sum node can be interpreted as a hidden variable, and a product node can be viewed as a feature capturing rich interactions among an SPN’s inputs. We show that the ability of SPN to use hidden layers to model complex dependencies among words, and its tractable inference and learning times, make it a suitable framework for a language model. Even though SPNs have been applied to a variety of vision problems [1, 2], we are the first to use it for language modeling. Our empirical comparisons with six previous language models indicate that our SPN has superior performance.
Wei-Chen Cheng, Stanley Kok, Hoai Vu Pham, Hai Leong Chieu, Kian Ming A. Chai
INTERSPEECH4
2014 Conditional random field with high-order dependencies for sequence labeling and segmentation
Viet Cuong Nguyen, Wee Sun Lee, Hai Leong Chieu
J. Mach. Learn. Res.4
2013 Active Learning for Probabilistic Hypotheses Using the Maximum Gibbs Error Criterion
abstract
We introduce a new objective function for pool-based Bayesian active learning with probabilistic hypotheses. This objective function, called the policy Gibbs error, is the expected error rate of a random classifier drawn from the prior distribution on the examples adaptively selected by the active learning policy. Exact maximization of the policy Gibbs error is hard, so we propose a greedy strategy that maximizes the Gibbs error at each iteration, where the Gibbs error on an instance is the expected error of a random classifier selected from the posterior label distribution on that instance. We apply this maximum Gibbs error criterion to three active learning scenarios: non-adaptive, adaptive, and batch active learning. In each scenario, we prove that the criterion achieves near-maximal policy Gibbs error when constrained to a fixed budget. For practical implementations, we provide approximations to the maximum Gibbs error criterion for Bayesian conditional random fields and transductive Naive Bayes. Our experimental results on a named entity recognition task and a text classification task show that the maximum Gibbs error criterion is an effective active learning criterion for noisy models.
Viet Cuong Nguyen, Wee Sun Lee, Kian Ming A. Chai, Hai Leong Chieu
NIPS5
2012 Domain Adaptation for Coreference Resolution: An Adaptive Ensemble Approach
Jian-Bo Yang, Qi Mao 0001, Qiaoliang Xiang, Ivor W. Tsang, Kian Ming A. Chai, Hai Leong Chieu
EMNLP-CoNLL6
2012 Combining local and non-local information with dual decomposition for named entity recognition from text
Hai Leong Chieu, Loo-Nin Teow
FUSION1
2012 Optimizing F-measure: A Tale of Two Approaches
Kian Ming A. Chai, Wee Sun Lee, Hai Leong Chieu
ICML4
2012 A Split-Merge Framework for Comparing Clusterings
Qiaoliang Xiang, Qi Mao 0001, Kian Ming A. Chai, Hai Leong Chieu, Ivor W. Tsang, Zhendong Zhao
ICML4
2011 Unsupervised Information Extraction with Distributional Prior Knowledge
Cane Wing-ki Leung, Jing Jiang 0001, Kian Ming A. Chai, Hai Leong Chieu, Loo-Nin Teow
EMNLP4
2011 Extracting Relation Descriptors with Conditional Random Fields
Yaliang Li, Jing Jiang 0001, Hai Leong Chieu, Kian Ming A. Chai
IJCNLP3
2009 Domain adaptive bootstrapping for named entity recognition
Wee Sun Lee, Hai Leong Chieu
EMNLP4
2009 Conditional Random Fields with High-Order Features for Sequence Labeling
abstract
Dependencies among neighbouring labels in a sequence is an important source of information for sequence labeling problems. However, only dependencies between adjacent labels are commonly exploited in practice because of the high computational complexity of typical inference algorithms when longer distance dependencies are taken into account. In this paper, we show that it is possible to design efficient inference algorithms for a conditional random field using features that depend on long consecutive label sequences (high-order features), as long as the number of distinct label sequences in the features used is small. This leads to efficient learning algorithms for these conditional random fields. We show experimentally that exploiting dependencies using high-order features can lead to substantial performance improvements for some problems and discuss conditions under which high-order features can be effective.
Wee Sun Lee, Hai Leong Chieu
NIPS3
2009 Relaxed Survey Propagation for The Weighted Maximum Satisfiability Problem
abstract
The survey propagation (SP) algorithm has been shown to work well on large instances of the random 3-SAT problem near its phase transition. It was shown that SP estimates marginals over covers that represent clusters of solutions. The SP-y algorithm generalizes SP to work on the maximum satisfiability (Max-SAT) problem, but the cover interpretation of SP does not generalize to SP-y. In this paper, we formulate the relaxed survey propagation (RSP) algorithm, which extends the SP algorithm to apply to the weighted Max-SAT problem. We show that RSP has an interpretation of estimating marginals over covers violating a set of clauses with minimal weight. This naturally generalizes the cover interpretation of SP. Empirically, we show that RSP outperforms SP-y and other state-of-the-art Max-SAT solvers on random Max-SAT instances. RSP also outperforms state-of-the-art weighted Max-SAT solvers on random weighted Max-SAT instances.
Hai Leong Chieu, Wee Sun Lee
J. Artif. Intell. Res.1
2008 Relaxed Survey Propagation: A Sum-Product Algorithm for Max-SAT
Hai Leong Chieu, Wee Sun Lee
AAAI1
2007 Cooled and Relaxed Survey Propagation for MRFs
abstract
We describe a new algorithm, Relaxed Survey Propagation (RSP), for finding MAP configurations in Markov random fields. We compare its performance with state-of-the-art algorithms including the max-product belief propagation, its se- quential tree-reweighted variant, residual (sum-product) belief propagation, and tree-structured expectation propagation. We show that it outperforms all ap- proaches for Ising models with mixed couplings, as well as on a web person disambiguation task formulated as a supervised clustering problem.
Hai Leong Chieu, Wee Sun Lee, Yee Whye Teh
NIPS1
2004 Query based event extraction along a timeline
abstract
In this paper, we present a framework and a system that extracts events relevant to a query from a collection C of documents, and places such events along a timeline. Each event is represented by a sentence extracted from C, based on the assumption that "important" events are widely cited in many documents for a period of time within which these events are of interest. In our experiments, we used queries that are event types ("earthquake") and person names (e.g. "George Bush"). Evaluation was performed using G8 leader names as queries: comparison made by human evaluators between manually and system generated timelines showed that although manually generated timelines are on average more preferable, system generated timelines are sometimes judged to be better than manually constructed ones.
Hai Leong Chieu, Yoong Keok Lee
SIGIR1
2003 Closing the Gap: Learning-Based Information Extraction Rivaling Knowledge-Engineering Methods
abstract
In this paper, we present a learning approach to the scenario template task of information extraction, where information filling one template could come from multiple sentences. When tested on the MUC-4 task, our learning approach achieves accuracy competitive to the best of the MUC-4 systems, which were all built with manually engineered rules. Our analysis reveals that our use of full parsing and state-of-the-art learning algorithms have contributed to the good performance. To our knowledge, this is the first research to have demonstrated that a learning approach to the full-scale information extraction task could achieve performance rivaling that of the knowledge engineering approach.
Hai Leong Chieu, Hwee Tou Ng, Yoong Keok Lee
ACL1
2003 Named Entity Recognition with a Maximum Entropy Approach
Hai Leong Chieu, Hwee Tou Ng
CoNLL1
2002 Teaching a Weaker Classifier: Named Entity Recognition on Upper Case Text
abstract
This paper describes how a machine-learning named entity recognizer (NER) on upper case text can be improved by using a mixed case NER and some unlabeled text. The mixed case NER can be used to tag some unlabeled mixed case text, which are then used as additional training material for the upper case NER. We show that this approach reduces the performance gap between the mixed case NER and the upper case NER substantially, by 39% for MUC-6 and 22% for MUC-7 named entity test data. Our method is thus useful in improving the accuracy of NERs on upper case text, such as transcribed text from automatic speech recognizers where case information is missing.
Hai Leong Chieu, Hwee Tou Ng
ACL1
2002 Named Entity Recognition: A Maximum Entropy Approach Using Global Information
Hai Leong Chieu, Hwee Tou Ng
COLING1
2002 Bayesian online classifiers for text classification and filtering
abstract
This paper explores the use of Bayesian online classifiers to classify text documents. Empirical results indicate that these classifiers are comparable with the best text classification systems. Furthermore, the online approach offers the advantage of continuous learning in the batch-adaptive text filtering task.
Kian Ming A. Chai, Hai Leong Chieu, Hwee Tou Ng
SIGIR2