Hwee Tou Ng

dblp:97/3037 · DBLP profile ↗
← Back
135ranked-venue papers
14as first author
25since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 126 · 13 first-author · 24 since 2021Graphics, computer vision, multimedia, augmented reality and games · 22 · 3 first-author · 4 since 2021Databases, data management, data science and information retrieval · 7 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 3 · 1 since 2021Theory of computation · 1 · 1 first-author
YearPublicationVenuePosition
2026 SlideTailor: Personalized Presentation Slide Generation for Scientific Papers
abstract
Automatic presentation slide generation can greatly streamline content creation. However, since preferences of each user may vary, existing under-specified formulations often lead to suboptimal results that fail to align with individual user needs. We introduce a novel task that conditions paper-to-slides generation on user-specified preferences. We propose a human behavior-inspired agentic framework, SlideTailor, that progressively generates editable slides in a user-aligned manner. Instead of requiring users to write their preferences in detailed textual form, our system only asks for a paper-slides example pair and a visual template—natural and easy-to-provide artifacts that implicitly encode rich user preferences across content and visual style. Despite the implicit and unlabeled nature of these inputs, our framework effectively distills and generalizes the preferences to guide customized slide generation. We also introduce a novel chain-of-speech mechanism to align slide content with planned oral narration. Such a design significantly enhances the quality of generated slides and enables downstream applications like video presentations. To support this new task, we construct a benchmark dataset that captures diverse user preferences, with carefully designed interpretable metrics for robust evaluation. Extensive experiments demonstrate the effectiveness of our framework.
Wenzheng Zeng, Mingyu Ouyang, Langyuan Cui, Hwee Tou Ng
AAAI4
2025 Just What You Desire: Constrained Timeline Summarization with Self-Reflection for Enhanced Relevance
abstract
Given news articles about an entity, such as a public figure or organization, timeline summarization (TLS) involves generating a timeline that summarizes the key events about the entity. However, the TLS task is too underspecified, since what is of interest to each reader may vary, and hence there is not a single ideal or optimal timeline. In this paper, we introduce a novel task, called Constrained Timeline Summarization (CTLS), where a timeline is generated in which all events in the timeline meet some constraint. An example of a constrained timeline concerns the legal battles of Tiger Woods, where only events related to his legal problems are selected to appear in the timeline. We collected a new human-verified dataset of constrained timelines involving 47 entities and 5 constraints per entity. We propose an approach that employs a large language model (LLM) to summarize news articles according to a specified constraint and cluster them to identify key events to include in a constrained timeline. In addition, we propose a novel self-reflection method during summary generation, demonstrating that this approach successfully leads to improved performance.
Muhammad Reza Qorib, Qisheng Hu, Hwee Tou Ng
AAAI3
2025 Think&Cite: Improving Attributed Text Generation with Self-Guided Tree Search and Progress Reward Modeling
abstract
Despite their outstanding capabilities, large language models (LLMs) are prone to hallucination and producing factually incorrect information. This challenge has spurred efforts in attributed text generation, which prompts LLMs to generate content with supporting evidence. In this paper, we propose a novel framework, called Think&Cite, and formulate attributed text generation as a multi-step reasoning problem integrated with search. Specifically, we propose Self-Guided Monte Carlo Tree Search (SG-MCTS), which capitalizes on the self-reflection capability of LLMs to reason about the intermediate states of MCTS for guiding the tree expansion process. To provide reliable and comprehensive feedback, we introduce Progress Reward Modeling to measure the progress of tree search from the root to the current state from two aspects, i.e., generation and attribution progress. We conduct extensive experiments on three datasets and the results show that our approach significantly outperforms baseline approaches.
Junyi Li 0001, Hwee Tou Ng
ACL (1)2
2025 Just Go Parallel: Improving the Multilingual Capabilities of Large Language Models
abstract
Large language models (LLMs) have demonstrated impressive translation capabilities even without being explicitly trained on parallel data.This remarkable property has led some to believe that parallel data is no longer necessary for building multilingual language models.While some attribute this to the emergent abilities of LLMs due to scale, recent work suggests that it is actually caused by incidental bilingual signals present in the training data.Various methods have been proposed to maximize the utility of parallel data to enhance the multilingual capabilities of multilingual encoder-based and encoder-decoder language models.However, some decoder-based LLMs opt to ignore parallel data instead.In this work, we conduct a systematic study on the impact of adding parallel data on LLMs' multilingual capabilities, focusing specifically on translation and multilingual common-sense reasoning.Through controlled experiments, we demonstrate that parallel data can significantly improve LLMs' multilingual capabilities.1
Muhammad Reza Qorib, Junyi Li 0001, Hwee Tou Ng
ACL (1)3
2025 Finding the Sweet Spot: Preference Data Construction for Scaling Preference Optimization
abstract
Yao Xiao, Hai Ye, Linyao Chen, Hwee Tou Ng, Lidong Bing, Xiaoli Li, Roy Ka-Wei Lee. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Hai Ye, Linyao Chen, Hwee Tou Ng, Lidong Bing, Roy Ka-Wei Lee
ACL (1)4
2025 Factorized Learning for Temporally Grounded Video-Language Models
Wenzheng Zeng, Difei Gao, Zheng Shou 0001, Hwee Tou Ng
ICCV4
2025 Reasoning Models Hallucinate More: Factuality-Aware Reinforcement Learning for Large Reasoning Models
abstract
Large language models (LLMs) have significantly advanced in reasoning tasks through reinforcement learning (RL) optimization, achieving impressive capabilities across various challenging benchmarks. However, our empirical analysis reveals a critical drawback: reasoning-oriented RL fine-tuning significantly increases the prevalence of hallucinations. We theoretically analyze the RL training dynamics, identifying high-variance gradient, entropy-induced randomness, and susceptibility to spurious local optima as key factors leading to hallucinations. To address this drawback, we propose Factuality-aware Step-wise Policy Optimization (FSPO), an innovative RL fine-tuning algorithm incorporating explicit factuality verification at each reasoning step. FSPO leverages automated verification against given evidence to dynamically adjust token-level advantage values, incentivizing factual correctness throughout the reasoning process. Experiments across mathematical reasoning and hallucination benchmarks using Qwen2.5 and Llama models demonstrate that FSPO effectively reduces hallucinations while enhancing reasoning accuracy, substantially improving both reliability and performance.
Junyi Li 0001, Hwee Tou Ng
NeurIPS2
2024 From Moments to Milestones: Incremental Timeline Summarization Leveraging Large Language Models
abstract
Timeline summarization (TLS) is essential for distilling coherent narratives from a vast collection of texts, tracing the progression of events and topics over time.Prior research typically focuses on either event or topic timeline summarization, neglecting the potential synergy of these two forms.In this study, we bridge this gap by introducing a novel approach that leverages large language models (LLMs) for generating both event and topic timelines.Our approach diverges from conventional TLS by prioritizing event detection, leveraging LLMs as pseudo-oracles for incremental event clustering and construction of timelines from a text stream.As a result, it produces a more interpretable pipeline.Empirical evaluation across four TLS benchmarks reveals that our approach outperforms the best prior published approaches, highlighting the potential of LLMs in timeline summarization for real-world applications.1
Qisheng Hu, Geonsik Moon, Hwee Tou Ng
ACL (1)3
2024 Preference-Guided Reflective Sampling for Aligning Language Models
abstract
Iterative data generation and model re-training can effectively align large language models (LLMs) to human preferences.The process of data sampling is crucial, as it significantly influences the success of policy improvement.Repeated random sampling is a widely used method that independently queries the model multiple times to generate outputs.In this work, we propose a more effective sampling method, named Preference-Guided Reflective Sampling (PRS).Unlike random sampling, PRS employs a tree-based generation framework to enable more efficient sampling.It leverages adaptive self-refinement techniques to better explore the sampling space.By specifying user preferences in natural language, PRS can further optimize response generation according to these preferences.As a result, PRS can align models to diverse user preferences.Our experiments demonstrate that PRS generates higher-quality responses with significantly higher rewards.On AlpacaEval and Arena-Hard, PRS substantially outperforms repeated random sampling in bestof-N sampling.Moreover, PRS shows strong performance when applied in iterative offline RL training 1 .
Hai Ye, Hwee Tou Ng
EMNLP2
2023 Towards Benchmarking and Improving the Temporal Reasoning Capability of Large Language Models
abstract
Reasoning about time is of fundamental importance.Many facts are time-dependent.For example, athletes change teams from time to time, and different government officials are elected periodically.Previous time-dependent question answering (QA) datasets tend to be biased in either their coverage of time spans or question types.In this paper, we introduce a comprehensive probing dataset TEMPREASON to evaluate the temporal reasoning capability of large language models.Our dataset includes questions of three temporal reasoning levels.In addition, we also propose a novel learning framework to improve the temporal reasoning capability of large language models, based on temporal span extraction and time-sensitive reinforcement learning.We conducted experiments in closed book QA, open book QA, and reasoning QA settings and demonstrated the effectiveness of our approach 1 .
Hwee Tou Ng, Lidong Bing
ACL (1)2
2023 Multi-Source Test-Time Adaptation as Dueling Bandits for Extractive Question Answering
abstract
In this work, we study multi-source test-time model adaptation from user feedback, where K distinct models are established for adaptation.To allow efficient adaptation, we cast the problem as a stochastic decision-making process, aiming to determine the best adapted model after adaptation.We discuss two frameworks: multi-armed bandit learning and multi-armed dueling bandits.Compared to multi-armed bandit learning, the dueling framework allows pairwise collaboration among K models, which is solved by a novel method named Co-UCB proposed in this work.Experiments on six datasets of extractive question answering (QA) show that the dueling framework using Co-UCB is more effective than other strong baselines for our studied problem 1 .
Hai Ye, Qizhe Xie, Hwee Tou Ng
ACL (1)3
2023 Mitigating Exposure Bias in Grammatical Error Correction with Data Augmentation and Reweighting
abstract
The most popular approach in grammatical error correction (GEC) is based on sequence-tosequence (seq2seq) models.Similar to other autoregressive generation tasks, seq2seq GEC also faces the exposure bias problem, i.e., the context tokens are drawn from different distributions during training and testing, caused by the teacher forcing mechanism.In this paper, we propose a novel data manipulation approach to overcome this problem, which includes a data augmentation method during training to mimic the decoder input at inference time, and a data reweighting method to automatically balance the importance of each kind of augmented samples.Experimental results on benchmark GEC datasets show that our method achieves significant improvements compared to prior approaches. 1
Hannan Cao, Wenmian Yang, Hwee Tou Ng
EACL3
2023 Unsupervised Grammatical Error Correction Rivaling Supervised Methods
abstract
State-of-the-art grammatical error correction (GEC) systems rely on parallel training data (ungrammatical sentences and their manually corrected counterparts), which are expensive to construct.In this paper, we employ the Break-It-Fix-It (BIFI) method to build an unsupervised GEC system.The BIFI framework generates parallel data from unlabeled text using a fixer to transform ungrammatical sentences into grammatical ones, and a critic to predict sentence grammaticality.We present an unsupervised approach to build the fixer and the critic, and an algorithm that allows them to iteratively improve each other.We evaluate our unsupervised GEC system on English and Chinese GEC.Empirical results show that our GEC system outperforms previous unsupervised GEC systems, and achieves performance comparable to supervised GEC systems without ensemble.Furthermore, when combined with labeled training data, our system achieves new state-of-the-art results on the CoNLL-2014 and NLPCC-2018 test sets. 1
Hannan Cao, Liping Yuan, Hwee Tou Ng
EMNLP4
2023 System Combination via Quality Estimation for Grammatical Error Correction
abstract
Quality estimation models have been developed to assess the corrections made by grammatical error correction (GEC) models when the reference or gold-standard corrections are not available.An ideal quality estimator can be utilized to combine the outputs of multiple GEC systems by choosing the best subset of edits from the union of all edits proposed by the GEC base systems.However, we found that existing GEC quality estimation models are not good enough in differentiating good corrections from bad ones, resulting in a low F 0.5 score when used for system combination.In this paper, we propose GRECO 1 , a new state-of-the-art quality estimation model that gives a better estimate of the quality of a corrected sentence, as indicated by having a higher correlation to the F 0.5 score of a corrected sentence.It results in a combined GEC system with a higher F 0.5 score.We also propose three methods for utilizing GEC quality estimation models for system combination with varying generality: modelagnostic, model-agnostic with voting bias, and model-dependent method.The combined GEC system outperforms the state of the art on the CoNLL-2014 test set and the BEA-2019 test set, achieving the highest F 0.5 scores published to date.
Muhammad Reza Qorib, Hwee Tou Ng
EMNLP2
2023 Grammatical Error Correction: A Survey of the State of the Art
abstract
Abstract Grammatical Error Correction (GEC) is the task of automatically detecting and correcting errors in text. The task not only includes the correction of grammatical errors, such as missing prepositions and mismatched subject–verb agreement, but also orthographic and semantic errors, such as misspellings and word choice errors, respectively. The field has seen significant progress in the last decade, motivated in part by a series of five shared tasks, which drove the development of rule-based methods, statistical classifiers, statistical machine translation, and finally neural machine translation systems, which represent the current dominant state of the art. In this survey paper, we condense the field into a single article and first outline some of the linguistic challenges of the task, introduce the most popular datasets that are available to researchers (for both English and other languages), and summarize the various methods and techniques that have been developed with a particular focus on artificial error generation. We next describe the many different approaches to evaluation as well as concerns surrounding metric reliability, especially in relation to subjective human judgments, before concluding with an overview of recent progress and suggestions for future work and remaining challenges.We hope that this survey will serve as a comprehensive resource for researchers who are new to the field or who want to be kept apprised of recent developments.
Christopher Bryant 0001, Zheng Yuan 0003, Muhammad Reza Qorib, Hannan Cao, Hwee Tou Ng, Ted Briscoe
Comput. Linguistics5
2023 My Tenure as the Editor-in-Chief of Computational Linguistics
abstract
Abstract Times flies and it has been close to five and a half years since I became the editor-in-chief of Computational Linguistics on 15 July 2018. In this editorial, I will describe the changes that I have introduced at the journal, and highlight the achievements and challenges of the journal.
Hwee Tou Ng
Comput. Linguistics1
2023 Inferring cancer disease response from radiology reports using large language models with data augmentation and prompting
abstract
OBJECTIVE: To assess large language models on their ability to accurately infer cancer disease response from free-text radiology reports. MATERIALS AND METHODS: We assembled 10 602 computed tomography reports from cancer patients seen at a single institution. All reports were classified into: no evidence of disease, partial response, stable disease, or progressive disease. We applied transformer models, a bidirectional long short-term memory model, a convolutional neural network model, and conventional machine learning methods to this task. Data augmentation using sentence permutation with consistency loss as well as prompt-based fine-tuning were used on the best-performing models. Models were validated on a hold-out test set and an external validation set based on Response Evaluation Criteria in Solid Tumors (RECIST) classifications. RESULTS: The best-performing model was the GatorTron transformer which achieved an accuracy of 0.8916 on the test set and 0.8919 on the RECIST validation set. Data augmentation further improved the accuracy to 0.8976. Prompt-based fine-tuning did not further improve accuracy but was able to reduce the number of training reports to 500 while still achieving good performance. DISCUSSION: These models could be used by researchers to derive progression-free survival in large datasets. It may also serve as a decision support tool by providing clinicians an automated second opinion of disease response. CONCLUSIONS: Large clinical language models demonstrate potential to infer cancer disease response from radiology reports at scale. Data augmentation techniques are useful to further improve performance. Prompt-based fine-tuning can significantly reduce the size of the training dataset.
Ryan Shea Ying Cong Tan, Guat Hwa Low, Ruixi Lin, Tzer Chew Goh, Christopher Chu En Chang, Fung Fung Lee, Wei Yin Chan, Wei Chong Tan, Han Jieh Tey, Fun Loon Leong, Hong Qi Tan, Wen Long Nei, Wen Yee Chay, David Wai Meng Tai, Gillianne Geet Yi Lai, Lionel Tim-Ee Cheng, Fuh Yong Wong, Matthew Chua 0001, Melvin Lee Kiang Chua, Daniel Shao-Weng Tan, Choon Hua Thng, Iain Bee Huat Tan, Hwee Tou Ng
J. Am. Medical Informatics Assoc.24
2022 A Semi-supervised Learning Approach with Two Teachers to Improve Breakdown Identification in Dialogues
abstract
Identifying breakdowns in ongoing dialogues helps to improve communication effectiveness. Most prior work on this topic relies on human annotated data and data augmentation to learn a classification model. While quality labeled dialogue data requires human annotation and is usually expensive to obtain, unlabeled data is easier to collect from various sources. In this paper, we propose a novel semi-supervised teacher-student learning framework to tackle this task. We introduce two teachers which are trained on labeled data and perturbed labeled data respectively. We leverage unlabeled data to improve classification in student training where we employ two teachers to refine the labeling of unlabeled data through teacher-student learning in a bootstrapping manner. Through our proposed training approach, the student can achieve improvements over single-teacher performance. Experimental results on the Dialogue Breakdown Detection Challenge dataset DBDC5 and Learning to Identify Follow-Up Questions dataset LIF show that our approach outperforms all previous published approaches as well as other supervised and semi-supervised baseline methods.
Hwee Tou Ng
AAAI2
2022 On the Robustness of Question Rewriting Systems to Questions of Varying Hardness
abstract
In conversational question answering (CQA), the task of question rewriting (QR) in context aims to rewrite a context-dependent question into an equivalent self-contained question that gives the same answer.In this paper, we are interested in the robustness of a QR system to questions varying in rewriting hardness or difficulty.Since there is a lack of questions classified based on their rewriting hardness, we first propose a heuristic method to automatically classify questions into subsets of varying hardness, by measuring the discrepancy between a question and its rewrite.To find out what makes questions hard or easy for rewriting, we then conduct a human evaluation to annotate the rewriting hardness of questions.Finally, to enhance the robustness of QR systems to questions of varying hardness, we propose a novel learning framework for QR that first trains a QR model independently on each subset of questions of a certain level of hardness, then combines these QR models as one joint model for inference.Experimental results on two datasets show that our framework improves the overall performance compared to the baselines 1 .
Hai Ye, Hwee Tou Ng, Wenjuan Han
ACL (1)2
2022 Grammatical Error Correction: Are We There Yet?
abstract
There has been much recent progress in natural language processing, and grammatical error correction (GEC) is no exception. We found that state-of-the-art GEC systems (T5 and GECToR) outperform humans by a wide margin on the CoNLL-2014 test set, a benchmark GEC test corpus, as measured by the standard F0.5 evaluation metric. However, a careful examination of their outputs reveals that there are still classes of errors that they fail to correct. This suggests that creating new test data that more accurately measure the true performance of GEC systems constitutes important future work.
Muhammad Reza Qorib, Hwee Tou Ng
COLING2
2022 Domain Generalization for Text Classification with Memory-Based Supervised Contrastive Learning
abstract
While there is much research on cross-domain text classification, most existing approaches focus on one-to-one or many-to-one domain adaptation. In this paper, we tackle the more challenging task of domain generalization, in which domain-invariant representations are learned from multiple source domains, without access to any data from the target domains, and classification decisions are then made on test documents in unseen target domains. We propose a novel framework based on supervised contrastive learning with a memory-saving queue. In this way, we explicitly encourage examples of the same class to be closer and examples of different classes to be further apart in the embedding space. We have conducted extensive experiments on two Amazon review sentiment datasets, and one rumour detection dataset. Experimental results show that our domain generalization method consistently outperforms state-of-the-art domain adaptation methods.
Ruidan He, Lidong Bing, Hwee Tou Ng
COLING4
2022 Revisiting DocRED - Addressing the False Negative Problem in Relation Extraction
abstract
The DocRED dataset is one of the most popular and widely used benchmarks for documentlevel relation extraction (RE).It adopts a recommend-revise annotation scheme so as to have a large-scale annotated dataset.However, we find that the annotation of DocRED is incomplete, i.e., false negative samples are prevalent.We analyze the causes and effects of the overwhelming false negative problem in the DocRED dataset.To address the shortcoming, we re-annotate 4,053 documents in the DocRED dataset by adding the missed relation triples back to the original DocRED.We name our revised DocRED dataset Re-DocRED.We conduct extensive experiments with state-ofthe-art neural models on both datasets, and the experimental results show that the models trained and evaluated on our Re-DocRED achieve performance improvements of around 13 F1 points.Moreover, we conduct a comprehensive analysis to identify the potential areas for further improvement.1
Lu Xu 0007, Lidong Bing, Hwee Tou Ng, Sharifah Mahani Aljunied
EMNLP4
2022 Frustratingly Easy System Combination for Grammatical Error Correction
abstract
In this paper, we formulate system combination for grammatical error correction (GEC) as a simple machine learning task: binary classification.We demonstrate that with the right problem formulation, a simple logistic regression algorithm can be highly effective for combining GEC models.Our method successfully increases the F 0.5 score from the highest base GEC system by 4.2 points on the CoNLL-2014 test set and 7.2 points on the BEA-2019 test set.Furthermore, our method outperforms the state of the art by 4.0 points on the BEA-2019 test set, 1.2 points on the CoNLL-2014 test set with original annotation, and 3.4 points on the CoNLL-2014 test set with alternative annotation.We also show that our system combination generates better corrections with higher F 0.5 scores than the conventional ensemble.1
Muhammad Reza Qorib, Seung-Hoon Na, Hwee Tou Ng
NAACL-HLT3
2021 Do Multi-Hop Question Answering Systems Know How to Answer the Single-Hop Sub-Questions?
abstract
Multi-hop question answering (QA) requires a model to retrieve and integrate information from multiple passages to answer a question.Rapid progress has been made on multi-hop QA systems with regard to standard evaluation metrics, including EM and F1.However, by simply evaluating the correctness of the answers, it is unclear to what extent these systems have learned the ability to perform multihop reasoning.In this paper, we propose an additional sub-question evaluation for the multihop QA dataset HotpotQA, in order to shed some light on explaining the reasoning process of QA systems in answering complex questions.We adopt a neural decomposition model to generate sub-questions for a multi-hop question, followed by extracting the corresponding sub-answers.Contrary to our expectation, multiple state-of-the-art multi-hop QA models fail to answer a large portion of sub-questions, although the corresponding multi-hop questions are correctly answered.Our work takes a step forward towards building a more explainable multi-hop QA system.
Hwee Tou Ng, Anthony K. H. Tung
EACL2
2021 Diversity-Driven Combination for Grammatical Error Correction
abstract
Grammatical error correction (GEC) is the task of detecting and correcting errors in a written text. The idea of combining multiple system outputs has been successfully used in GEC. To achieve successful system combination, multiple component systems need to produce corrected sentences that are both diverse and of comparable quality. However, most existing state-of-the-art GEC approaches are based on similar sequence-to-sequence neural networks, so the gains are limited from combining the outputs of component systems similar to one another. In this paper, we present Diversity-Driven Combination (DDC) for GEC, a system combination strategy that encourages diversity among component systems. We evaluate our system combination strategy on the CoNLL-2014 shared task and the BEA-2019 shared task. On both benchmarks, DDC achieves significant performance gain with a small number of training examples and outperforms the component systems by a large margin. Our source code is available at https://github.com/nusnlp/gec-ddc.
Wenjuan Han, Hwee Tou Ng
ICTAI2
2020 Effective Modeling of Encoder-Decoder Architecture for Joint Entity and Relation Extraction
abstract
A relation tuple consists of two entities and the relation between them, and often such tuples are found in unstructured text. There may be multiple relation tuples present in a text and they may share one or both entities among them. Extracting such relation tuples from a sentence is a difficult task and sharing of entities or overlapping entities among the tuples makes it more challenging. Most prior work adopted a pipeline approach where entities were identified first followed by finding the relations among them, thus missing the interaction among the relation tuples in a sentence. In this paper, we propose two approaches to use encoder-decoder architecture for jointly extracting entities and relations. In the first approach, we propose a representation scheme for relation tuples which enables the decoder to generate one word at a time like machine translation models and still finds all the tuples present in a sentence with full entity names of different length and with overlapping entities. Next, we propose a pointer network-based decoding approach where an entire tuple is generated at every time step. Experiments on the publicly available New York Times corpus show that our proposed approaches outperform previous work and achieve significantly higher F1 scores.
Tapas Nayak, Hwee Tou Ng
AAAI2
2020 Learning to Identify Follow-Up Questions in Conversational Question Answering
abstract
Despite recent progress in conversational question answering, most prior work does not focus on follow-up questions.Practical conversational question answering systems often receive follow-up questions in an ongoing conversation, and it is crucial for a system to be able to determine whether a question is a follow-up question of the current conversation, for more effective answer finding subsequently.In this paper, we introduce a new follow-up question identification task.We propose a three-way attentive pooling network that determines the suitability of a follow-up question by capturing pair-wise interactions between the associated passage, the conversation history, and a candidate follow-up question.It enables the model to capture topic continuity and topic shift while scoring a particular candidate follow-up question.Experiments show that our proposed three-way attentive pooling network outperforms all baseline systems by significant margins.
Souvik Kundu 0003, Hwee Tou Ng
ACL3
2020 A Survey of Unsupervised Dependency Parsing
abstract
Syntactic dependency parsing is an important task in natural language processing.Unsupervised dependency parsing aims to learn a dependency parser from sentences that have no annotation of their correct parse trees.Despite its difficulty, unsupervised parsing is an interesting research direction because of its capability of utilizing almost unlimited unannotated text data.It also serves as the basis for other research in low-resource parsing.In this paper, we survey existing approaches to unsupervised dependency parsing, identify two major classes of approaches, and discuss recent trends.We hope that our survey can provide insights for researchers and facilitate future research on this topic.
Wenjuan Han, Yong Jiang 0005, Hwee Tou Ng, Kewei Tu
COLING3
2020 A Co-Attentive Cross-Lingual Neural Model for Dialogue Breakdown Detection
abstract
Ensuring smooth communication is essential in a chat-oriented dialogue system, so that a user can obtain meaningful responses through interactions with the system.Most prior work on dialogue research does not focus on preventing dialogue breakdown.One of the major challenges is that a dialogue system may generate an undesired utterance leading to a dialogue breakdown, which degrades the overall interaction quality.Hence, it is crucial for a machine to detect dialogue breakdowns in an ongoing conversation.In this paper, we propose a novel dialogue breakdown detection model that jointly incorporates a pretrained cross-lingual language model and a co-attention network.Our proposed model leverages effective word embeddings trained on one hundred different languages to generate contextualized representations.Co-attention aims to capture the interaction between the latest utterance and the conversation history, and thereby determines whether the latest utterance causes a dialogue breakdown.Experimental results show that our proposed model outperforms all previous approaches on all evaluation metrics in both the Japanese and English tracks in Dialogue Breakdown Detection Challenge 4 (DBDC4 at IWSDS2019).
Souvik Kundu 0003, Hwee Tou Ng
COLING3
2020 Feature Adaptation of Pre-Trained Language Models across Languages and Domains with Robust Self-Training
abstract
Adapting pre-trained language models (PrLMs) (e.g., BERT) to new domains has gained much attention recently.Instead of fine-tuning PrLMs as done in most previous work, we investigate how to adapt the features of PrLMs to new domains without fine-tuning.We explore unsupervised domain adaptation (UDA) in this paper.With the features from PrLMs, we adapt the models trained with labeled data from the source domain to the unlabeled target domain.Self-training is widely used for UDA, and it predicts pseudo labels on the target domain data for training.However, the predicted pseudo labels inevitably include noise, which will negatively affect training a robust model.To improve the robustness of self-training, in this paper we present class-aware feature self-distillation (CFd) to learn discriminative features from PrLMs, in which PrLM features are self-distilled into a feature adaptation module and the features from the same class are more tightly clustered.We further extend CFd to a cross-language setting, in which language discrepancy is studied.Experiments on two monolingual and multilingual Amazon review datasets show that CFd can consistently improve the performance of self-training in cross-domain and cross-language settings.
Hai Ye, Ruidan He, Juntao Li 0005, Hwee Tou Ng, Lidong Bing
EMNLP (1)5
2020 Unsupervised Domain Adaptation of a Pretrained Cross-Lingual Language Model
abstract
Recent research indicates that pretraining cross-lingual language models on large-scale unlabeled texts yields significant performance improvements over various cross-lingual and low-resource tasks. Through training on one hundred languages and terabytes of texts, cross-lingual language models have proven to be effective in leveraging high-resource languages to enhance low-resource language processing and outperform monolingual models. In this paper, we further investigate the cross-lingual and cross-domain (CLCD) setting when a pretrained cross-lingual language model needs to adapt to new domains. Specifically, we propose a novel unsupervised feature decomposition method that can automatically extract domain-specific features and domain-invariant features from the entangled pretrained cross-lingual representations, given unlabeled raw texts in the source language. Our proposed model leverages mutual information estimation to decompose the representations computed by a cross-lingual model into domain-invariant and domain-specific parts. Experimental results show that our proposed method achieves significant performance improvements over the state-of-the-art pretrained cross-lingual language model in the CLCD setting.
Juntao Li 0005, Ruidan He, Hai Ye, Hwee Tou Ng, Lidong Bing, Rui Yan 0001
IJCAI4
2019 Cross-Sentence Grammatical Error Correction
abstract
Automatic grammatical error correction (GEC) research has made remarkable progress in the past decade.However, all existing approaches to GEC correct errors by considering a single sentence alone and ignoring crucial cross-sentence context.Some errors can only be corrected reliably using cross-sentence context and models can also benefit from the additional contextual information in correcting other errors.In this paper, we address this serious limitation of existing approaches and improve strong neural encoder-decoder models by appropriately modeling wider contexts.We employ an auxiliary encoder that encodes previous sentences and incorporate the encoding in the decoder via attention and gating mechanisms.Our approach results in statistically significant improvements in overall GEC performance over strong baselines across multiple test sets.Analysis of our cross-sentence GEC model on a synthetic dataset shows high performance in verb tense corrections that require cross-sentence context.
Shamil Chollampatt, Hwee Tou Ng
ACL (1)3
2019 Improving the Robustness of Question Answering Systems to Question Paraphrasing
abstract
Despite the advancement of question answering (QA) systems and rapid improvements on held-out test sets, their generalizability is a topic of concern.We explore the robustness of QA models to question paraphrasing by creating two test sets consisting of paraphrased SQuAD questions.Paraphrased questions from the first test set are very similar to the original questions designed to test QA models' over-sensitivity, while questions from the second test set are paraphrased using context words near an incorrect answer candidate in an attempt to confuse QA models.We show that both paraphrased test sets lead to significant decrease in performance on multiple state-of-the-art QA models.Using a neural paraphrasing model trained to generate multiple paraphrased questions for a given source question and a set of paraphrase suggestions, we propose a data augmentation approach that requires no human intervention to re-train the models for improved robustness to question paraphrasing.
Wee Chung Gan, Hwee Tou Ng
ACL (1)2
2019 An Interactive Multi-Task Learning Network for End-to-End Aspect-Based Sentiment Analysis
abstract
Aspect-based sentiment analysis produces a list of aspect terms and their corresponding sentiments for a natural language sentence.This task is usually done in a pipeline manner, with aspect term extraction performed first, followed by sentiment predictions toward the extracted aspect terms.While easier to develop, such an approach does not fully exploit joint information from the two subtasks and does not use all available sources of training information that might be helpful, such as document-level labeled sentiment corpus.In this paper, we propose an interactive multi-task learning network (IMN) which is able to jointly learn multiple related tasks simultaneously at both the token level as well as the document level.Unlike conventional multi-task learning methods that rely on learning common features for the different tasks, IMN introduces a message passing architecture where information is iteratively passed to different tasks through a shared set of latent variables.Experimental results demonstrate superior performance of the proposed method against multiple baselines on three benchmark datasets.
Ruidan He, Wee Sun Lee, Hwee Tou Ng, Daniel Dahlmeier
ACL (1)3
2019 Effective Attention Modeling for Neural Relation Extraction
abstract
Relation extraction is the task of determining the relation between two entities in a sentence.Distantly-supervised models are popular for this task.However, sentences can be long and two entities can be located far from each other in a sentence.The pieces of evidence supporting the presence of a relation between two entities may not be very direct, since the entities may be connected via some indirect links such as a third entity or via coreference.Relation extraction in such scenarios becomes more challenging as we need to capture the long-distance interactions among the entities and other words in the sentence.Also, the words in a sentence do not contribute equally in identifying the relation between the two entities.To address this issue, we propose a novel and effective attention model which incorporates syntactic information of the sentence and a multi-factor attention mechanism.Experiments on the New York Times corpus show that our proposed model outperforms prior state-of-the-art models.
Tapas Nayak, Hwee Tou Ng
CoNLL2
2019 Improved Word Sense Disambiguation Using Pre-Trained Contextualized Word Representations
abstract
Christian Hadiwinoto, Hwee Tou Ng, Wee Chung Gan. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Christian Hadiwinoto, Hwee Tou Ng, Wee Chung Gan
EMNLP/IJCNLP (1)2
2018 A Multilayer Convolutional Encoder-Decoder Neural Network for Grammatical Error Correction
abstract
We improve automatic correction of grammatical, orthographic, and collocation errors in text using a multilayer convolutional encoder-decoder neural network. The network is initialized with embeddings that make use of character N-gram information to better suit this task. When evaluated on common benchmark test data sets (CoNLL-2014 and JFLEG), our model substantially outperforms all prior neural approaches on this task as well as strong statistical machine translation-based systems with neural and task-specific features trained on the same data. Our analysis shows the superiority of convolutional neural networks over recurrent neural networks such as long short-term memory (LSTM) networks in capturing the local context via attention, and thereby improving the coverage in correcting grammatical errors. By ensembling multiple models, and incorporating an N-gram language model and edit features via rescoring, our novel method becomes the first neural approach to outperform the current state-of-the-art statistical machine translation-based approach, both in terms of grammaticality and fluency.
Shamil Chollampatt, Hwee Tou Ng
AAAI2
2018 A Question-Focused Multi-Factor Attention Network for Question Answering
abstract
Neural network models recently proposed for question answering (QA) primarily focus on capturing the passage-question relation. However, they have minimal capability to link relevant facts distributed across multiple sentences which is crucial in achieving deeper understanding, such as performing multi-sentence reasoning, co-reference resolution, etc. They also do not explicitly focus on the question and answer type which often plays a critical role in QA. In this paper, we propose a novel end-to-end question-focused multi-factor attention network for answer extraction. Multi-factor attentive encoding using tensor-based transformation aggregates meaningful facts even when they are located in multiple sentences. To implicitly infer the answer type, we also propose a max-attentional question aggregation mechanism to encode a question vector based on the important words in a question. During prediction, we incorporate sequence-level encoding of the first wh-word and its immediately following word as an additional source of question type information. Our proposed model achieves significant improvements over the best prior state-of-the-art results on three large-scale challenging QA datasets, namely NewsQA, TriviaQA, and SearchQA.
Souvik Kundu 0003, Hwee Tou Ng
AAAI2
2018 A Reassessment of Reference-Based Grammatical Error Correction Metrics
abstract
Several metrics have been proposed for evaluating grammatical error correction (GEC) systems based on grammaticality, fluency, and adequacy of the output sentences. Previous studies of the correlation of these metrics with human quality judgments were inconclusive, due to the lack of appropriate significance tests, discrepancies in the methods, and choice of datasets used. In this paper, we re-evaluate reference-based GEC metrics by measuring the system-level correlations with humans on a large dataset of human judgments of GEC outputs, and by properly conducting statistical significance tests. Our results show no significant advantage of GLEU over MaxMatch (M2), contradicting previous studies that claim GLEU to be superior. For a finer-grained analysis, we additionally evaluate these metrics for their agreement with human judgments at the sentence level. Our sentence-level analysis indicates that comparing GLEU and M2, one metric may be more useful than the other depending on the scenario. We further qualitatively analyze these metrics and our findings show that apart from being less interpretable and non-deterministic, GLEU also produces counter-intuitive scores in commonly occurring test examples.
Shamil Chollampatt, Hwee Tou Ng
COLING2
2018 Effective Attention Modeling for Aspect-Level Sentiment Classification
abstract
Aspect-level sentiment classification aims to determine the sentiment polarity of a review sentence towards an opinion target. A sentence could contain multiple sentiment-target pairs; thus the main challenge of this task is to separate different opinion contexts for different targets. To this end, attention mechanism has played an important role in previous state-of-the-art neural models. The mechanism is able to capture the importance of each context word towards a target by modeling their semantic associations. We build upon this line of research and propose two novel approaches for improving the effectiveness of attention. First, we propose a method for target representation that better captures the semantic meaning of the opinion target. Second, we introduce an attention model that incorporates syntactic information into the attention mechanism. We experiment on attention-based LSTM (Long Short-Term Memory) models using the datasets from SemEval 2014, 2015, and 2016. The experimental results show that the conventional attention-based LSTM can be substantially improved by incorporating the two approaches.
Ruidan He, Wee Sun Lee, Hwee Tou Ng, Daniel Dahlmeier
COLING3
2018 Neural Quality Estimation of Grammatical Error Correction
abstract
Grammatical error correction (GEC) systems deployed in language learning environments are expected to accurately correct errors in learners' writing.However, in practice, they often produce spurious corrections and fail to correct many errors, thereby misleading learners.This necessitates the estimation of the quality of output sentences produced by GEC systems so that instructors can selectively intervene and re-correct the sentences which are poorly corrected by the system and ensure that learners get accurate feedback.We propose the first neural approach to automatic quality estimation of GEC output sentences that does not employ any hand-crafted features.Our system is trained in a supervised manner on learner sentences and corresponding GEC system outputs with quality score labels computed using human-annotated references.Our neural quality estimation models for GEC show significant improvements over a strong feature-based baseline.We also show that a state-of-the-art GEC system can be improved when quality scores are used as features for reranking the N-best candidates.
Shamil Chollampatt, Hwee Tou Ng
EMNLP2
2018 Adaptive Semi-supervised Learning for Cross-domain Sentiment Classification
abstract
We consider the cross-domain sentiment classification problem, where a sentiment classifier is to be learned from a source domain and to be generalized to a target domain.Our approach explicitly minimizes the distance between the source and the target instances in an embedded feature space.With the difference between source and target minimized, we then exploit additional information from the target domain by consolidating the idea of semi-supervised learning, for which, we jointly employ two regularizations -entropy minimization and self-ensemble bootstrapping -to incorporate the unlabeled target data for classifier refinement.Our experimental results demonstrate that the proposed approach can better leverage unlabeled data from the target domain and achieve substantial improvements over baseline methods in various experimental settings.
Ruidan He, Wee Sun Lee, Hwee Tou Ng, Daniel Dahlmeier
EMNLP3
2018 A Nil-Aware Answer Extraction Framework for Question Answering
abstract
Recently, there has been a surge of interest in reading comprehension-based (RC) question answering (QA).However, current approaches suffer from an impractical assumption that every question has a valid answer in the associated passage.A practical QA system must possess the ability to determine whether a valid answer exists in a given text passage.In this paper, we focus on developing QA systems that can extract an answer for a question if and only if the associated passage contains an answer.If the associated passage does not contain any valid answer, the QA system will correctly return Nil.We propose a nil-aware answer span extraction framework that is capable of returning Nil or a text span from the associated passage as an answer in a single step.We show that our proposed framework can be easily integrated with several recently proposed QA models developed for reading comprehension and can be trained in an endto-end fashion.Our proposed nil-aware answer extraction neural network decomposes pieces of evidence into relevant and irrelevant parts and then combines them to infer the existence of any answer.Experiments on the NewsQA dataset show that the integration of our proposed framework significantly outperforms several strong baseline systems that use pipeline or threshold-based approaches.
Souvik Kundu 0003, Hwee Tou Ng
EMNLP2
2018 Upping the Ante: Towards a Better Benchmark for Chinese-to-English Machine Translation
Christian Hadiwinoto, Hwee Tou Ng
LREC2
2017 A Dependency-Based Neural Reordering Model for Statistical Machine Translation
abstract
In machine translation (MT) that involves translating between two languages with significant differences in word order, determining the correct word order of translated words is a major challenge. The dependency parse tree of a source sentence can help to determine the correct word order of the translated words. In this paper, we present a novel reordering approach utilizing a neural network and dependency-based embeddings to predict whether the translations of two source words linked by a dependency relation should remain in the same order or should be swapped in the translated sentence. Experiments on Chinese-to-English translation show that our approach yields a statistically significant improvement of 0.57 BLEU point on benchmark NIST test sets, compared to our prior state-of-the-art statistical MT system that uses sparse dependency-based reordering features.
Christian Hadiwinoto, Hwee Tou Ng
AAAI2
2017 An Unsupervised Neural Attention Model for Aspect Extraction
abstract
Aspect extraction is an important and challenging task in aspect-based sentiment analysis.Existing works tend to apply variants of topic models on this task.While fairly successful, these methods usually do not produce highly coherent aspects.In this paper, we present a novel neural approach with the aim of discovering coherent aspects.The model improves coherence by exploiting the distribution of word co-occurrences through the use of neural word embeddings.Unlike topic models which typically assume independently generated words, word embedding models encourage words that appear in similar contexts to be located close to each other in the embedding space.In addition, we use an attention mechanism to de-emphasize irrelevant words during training, further improving the coherence of aspects.Experimental results on real-life datasets demonstrate that our approach discovers more meaningful and coherent aspects, and substantially outperforms baseline methods on several evaluation tasks.
Ruidan He, Wee Sun Lee, Hwee Tou Ng, Daniel Dahlmeier
ACL (1)3
2016 To Swap or Not to Swap? Exploiting Dependency Word Pairs for Reordering in Statistical Machine Translation
abstract
Reordering poses a major challenge in machine translation (MT) between two languages with significant differences in word order. In this paper, we present a novel reordering approach utilizing sparse features based on dependency word pairs. Each instance of these features captures whether two words, which are related by a dependency link in the source sentence dependency parse tree, follow the same order or are swapped in the translation output. Experiments on Chinese-to-English translation show a statistically significant improvement of 1.21 BLEU point using our approach, compared to a state-of-the-art statistical MT system that incorporates prior reordering approaches.
Christian Hadiwinoto, Yang Liu 0005, Hwee Tou Ng
AAAI3
2016 Adapting Grammatical Error Correction Based on the Native Language of Writers with Neural Network Joint Models
abstract
An important aspect for the task of grammatical error correction (GEC) that has not yet been adequately explored is adaptation based on the native language (L1) of writers, despite the marked influences of L1 on second language (L2) writing.In this paper, we adapt a neural network joint model (NNJM) using L1-specific learner text and integrate it into a statistical machine translation (SMT) based GEC system.Specifically, we train an NNJM on general learner text (not L1-specific) and subsequently train on L1-specific data using a Kullback-Leibler divergence regularized objective function in order to preserve generalization of the model.We incorporate this adapted NNJM as a feature in an SMT-based English GEC system and show that adaptation achieves significant F 0.5 score gains on English texts written by L1 Chinese, Russian, and Spanish writers.
Shamil Chollampatt, Duc Tam Hoang, Hwee Tou Ng
EMNLP3
2016 A Neural Approach to Automated Essay Scoring
Kaveh Taghipour, Hwee Tou Ng
EMNLP2
2016 Neural Network Translation Models for Grammatical Error Correction
Shamil Chollampatt, Kaveh Taghipour, Hwee Tou Ng
IJCAI3
2016 Exploiting N-Best Hypotheses to Improve an SMT Approach to Grammatical Error Correction
Duc Tam Hoang, Shamil Chollampatt, Hwee Tou Ng
IJCAI3
2016 Source Language Adaptation Approaches for Resource-Poor Machine Translation
abstract
Most of the world languages are resource-poor for statistical machine translation; still, many of them are actually related to some resource-rich language. Thus, we propose three novel, language-independent approaches to source language adaptation for resource-poor statistical machine translation. Specifically, we build improved statistical machine translation models from a resource-poor language POOR into a target language TGT by adapting and using a large bitext for a related resource-rich language RICH and the same target language TGT. We assume a small POOR–TGT bitext from which we learn word-level and phrase-level paraphrases and cross-lingual morphological variants between the resource-rich and the resource-poor language. Our work is of importance for resource-poor machine translation because it can provide a useful guideline for people building machine translation systems for resource-poor languages. Our experiments for Indonesian/Malay–English translation show that using the large adapted resource-rich bitext yields 7.26 BLEU points of improvement over the unadapted one and 3.09 BLEU points over the original small bitext. Moreover, combining the small POOR–TGT bitext with the adapted bitext outperforms the corresponding combinations with the unadapted bitext by 1.93–3.25 BLEU points. We also demonstrate the applicability of our approaches to other languages and domains.
Pidong Wang, Preslav Nakov, Hwee Tou Ng
Comput. Linguistics3
2015 How Far are We from Fully Automatic High Quality Grammatical Error Correction?
abstract
Christopher Bryant, Hwee Tou Ng. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Christopher Bryant 0001, Hwee Tou Ng
ACL (1)2
2015 One Million Sense-Tagged Instances for Word Sense Disambiguation and Induction
abstract
Supervised word sense disambiguation (WSD) systems are usually the best performing systems when evaluated on standard benchmarks.However, these systems need annotated training data to function properly.While there are some publicly available open source WSD systems, very few large annotated datasets are available to the research community.The two main goals of this paper are to extract and annotate a large number of samples and release them for public use, and also to evaluate this dataset against some word sense disambiguation and induction tasks.We show that the open source IMS WSD system trained on our dataset achieves stateof-the-art results in standard disambiguation tasks and a recent word sense induction task, outperforming several task submissions and strong baselines.
Kaveh Taghipour, Hwee Tou Ng
CoNLL2
2015 Flexible Domain Adaptation for Automated Essay Scoring Using Correlated Linear Regression
abstract
Most of the current automated essay scoring (AES) systems are trained using manually graded essays from a specific prompt.These systems experience a drop in accuracy when used to grade an essay from a different prompt.Obtaining a large number of manually graded essays each time a new prompt is introduced is costly and not viable.We propose domain adaptation as a solution to adapt an AES system from an initial prompt to a new prompt.We also propose a novel domain adaptation technique that uses Bayesian linear ridge regression.We evaluate our domain adaptation technique on the publicly available Automated Student Assessment Prize (ASAP) dataset and show that our proposed technique is a competitive default domain adaptation algorithm for the AES task.
Peter Phandi, Kian Ming A. Chai, Hwee Tou Ng
EMNLP3
2015 Semi-Supervised Word Sense Disambiguation Using Word Embeddings in General and Specific Domains
abstract
One of the weaknesses of current supervised word sense disambiguation (WSD) systems is that they only treat a word as a discrete entity. However, a continuous-space representation of words (word embeddings) can provide valuable information and thus improve generalization accuracy. Since word embeddings are typically obtained from unlabeled data using unsupervised methods, this method can be seen as a semi-supervised word sense disambiguation approach. This paper investigates two ways of incorporating word embeddings in a word sense disambiguation setting and evaluates these two methods on some SensEval/SemEval lexical sample and all-words tasks and also a domain-specific lexical sample task. The obtained results show that such representations consistently improve the accuracy of the selected supervised WSD system. Moreover, our experiments on a domainspecific dataset show that our supervised baseline system beats the best knowledge-based systems by a large margin.
Kaveh Taghipour, Hwee Tou Ng
HLT-NAACL2
2014 A Beam-Search Decoder for Disfluency Detection
Xuancong Wang, Hwee Tou Ng, Khe Chai Sim
COLING2
2014 A Constituent-Based Approach to Argument Labeling with Joint Inference in Discourse Parsing
abstract
Discourse parsing is a challenging task and plays a critical role in discourse analysis.In this paper, we focus on labeling full argument spans of discourse connectives in the Penn Discourse Treebank (PDTB).Previous studies cast this task as a linear tagging or subtree extraction problem.In this paper, we propose a novel constituent-based approach to argument labeling, which integrates the advantages of both linear tagging and subtree extraction.In particular, the proposed approach unifies intra-and intersentence cases by treating the immediately preceding sentence as a special constituent.Besides, a joint inference mechanism is introduced to incorporate global information across arguments into our constituent-based approach via integer linear programming.Evaluation on PDT-B shows significant performance improvements of our constituent-based approach over the best state-of-the-art system.It also shows the effectiveness of our joint inference mechanism in modeling global information across arguments.
Fang Kong 0001, Hwee Tou Ng, Guodong Zhou 0001
EMNLP2
2014 System Combination for Grammatical Error Correction
abstract
Different approaches to high-quality grammatical error correction have been proposed recently, many of which have their own strengths and weaknesses.Most of these approaches are based on classification or statistical machine translation (SMT).In this paper, we propose to combine the output from a classification-based system and an SMT-based system to improve the correction quality.We adopt the system combination technique of Heafield and Lavie (2010).We achieve an F 0.5 score of 39.39% on the test set of the CoNLL-2014 shared task, outperforming the best system in the shared task.
Raymond Hendy Susanto, Peter Phandi, Hwee Tou Ng
EMNLP3
2014 Combining Punctuation and Disfluency Prediction: An Empirical Study
abstract
Punctuation prediction and disfluency prediction can improve downstream natural language processing tasks such as machine translation and information extraction.Combining the two tasks can potentially improve the efficiency of the overall pipeline system and reduce error propagation.In this work 1 , we compare various methods for combining punctuation prediction (PU) and disfluency prediction (DF) on the Switchboard corpus.We compare an isolated prediction approach with a cascade approach, a rescoring approach, and three joint model approaches.For the cascade approach, we show that the soft cascade method is better than the hard cascade method.We also use the cascade models to generate an n-best list, use the bi-directional cascade models to perform rescoring, and compare that with the results of the cascade models.For the joint model approach, we compare mixedlabel Linear-chain Conditional Random Field (LCRF), cross-product LCRF and 2layer Factorial Conditional Random Field (FCRF) with soft-cascade LCRF.Our results show that the various methods linking the two tasks are not significantly different from one another, although they perform better than the isolated prediction method by 0.5-1.5% in the F1 score.Moreover, the clique order of features also shows a marked difference.
Xuancong Wang, Khe Chai Sim, Hwee Tou Ng
EMNLP3
2014 A PDTB-styled end-to-end discourse parser
abstract
Abstract Since the release of the large discourse-level annotation of the Penn Discourse Treebank (PDTB), research work has been carried out on certain subtasks of this annotation, such as disambiguating discourse connectives and classifying Explicit or Implicit relations. We see a need to construct a full parser on top of these subtasks and propose a way to evaluate the parser. In this work, we have designed and developed an end-to-end discourse parser-to-parse free texts in the PDTB style in a fully data-driven approach. The parser consists of multiple components joined in a sequential pipeline architecture, which includes a connective classifier, argument labeler, explicit classifier, non-explicit classifier, and attribution span labeler. Our trained parser first identifies all discourse and non-discourse relations, locates and labels their arguments, and then classifies the sense of the relation between each pair of arguments. For the identified relations, the parser also determines the attribution spans, if any, associated with them. We introduce novel approaches to locate and label arguments, and to identify attribution spans. We also significantly improve on the current state-of-the-art connective classifier. We propose and present a comprehensive evaluation from both component-wise and error-cascading perspectives, in which we illustrate how each component performs in isolation, as well as how the pipeline performs with errors propagated forward. The parser gives an overall system F1 score of 46.80 percent for partial matching utilizing gold standard parses, and 38.18 percent with full automation.
Ziheng Lin, Hwee Tou Ng, Min-Yen Kan
Nat. Lang. Eng.2
2013 Grammatical Error Correction Using Integer Linear Programming
Yuanbin Wu, Hwee Tou Ng
ACL (1)2
2013 Towards Robust Linguistic Analysis using OntoNotes
Sameer Pradhan, Alessandro Moschitti, Nianwen Xue, Hwee Tou Ng, Anders Björkelund, Olga Uryupina
CoNLL4
2013 Exploiting Zero Pronouns to Improve Chinese Coreference Resolution
abstract
Coreference resolution plays a critical role in discourse analysis.This paper focuses on exploiting zero pronouns to improve Chinese coreference resolution.In particular, a simplified semantic role labeling framework is proposed to identify clauses and to detect zero pronouns effectively, and two effective methods (refining syntactic parser and refining learning example generation) are employed to exploit zero pronouns for Chinese coreference resolution.Evaluation on the CoNLL-2012 shared task data set shows that zero pronouns can significantly improve Chinese coreference resolution.
Fang Kong 0001, Hwee Tou Ng
EMNLP2
2013 A Beam-Search Decoder for Normalization of Social Media Text with Application to Machine Translation
Pidong Wang, Hwee Tou Ng
HLT-NAACL2
2012 Combining Coherence Models and Machine Translation Evaluation Metrics for Summarization Evaluation
Ziheng Lin, Hwee Tou Ng, Min-Yen Kan
ACL (1)3
2012 Character-Level Machine Translation Evaluation for Languages with Ambiguous Word Boundaries
Hwee Tou Ng
ACL (1)2
2012 Word Sense Disambiguation Improves Information Retrieval
Hwee Tou Ng
ACL (1)2
2012 A Beam-Search Decoder for Grammatical Error Correction
Daniel Dahlmeier, Hwee Tou Ng
EMNLP-CoNLL2
2012 Source Language Adaptation for Resource-Poor Machine Translation
Pidong Wang, Preslav Nakov, Hwee Tou Ng
EMNLP-CoNLL3
2012 Dynamic Conditional Random Fields for Joint Sentence Boundary and Punctuation Prediction
abstract
13th Annual Conference of the International Speech Communication Association 2012, INTERSPEECH 2012
Xuancong Wang, Hwee Tou Ng, Khe Chai Sim
INTERSPEECH2
2012 Better Evaluation for Grammatical Error Correction
Daniel Dahlmeier, Hwee Tou Ng
HLT-NAACL2
2012 Improving Statistical Machine Translation for a Resource-Poor Language Using Related Resource-Rich Languages
abstract
We propose a novel language-independent approach for improving machine translation for resource-poor languages by exploiting their similarity to resource-rich ones. More precisely, we improve the translation from a resource-poor source language X_1 into a resource-rich language Y given a bi-text containing a limited number of parallel sentences for X_1-Y and a larger bi-text for X_2-Y for some resource-rich language X_2 that is closely related to X_1. This is achieved by taking advantage of the opportunities that vocabulary overlap and similarities between the languages X_1 and X_2 in spelling, word order, and syntax offer: (1) we improve the word alignments for the resource-poor language, (2) we further augment it with additional translation options, and (3) we take care of potential spelling differences through appropriate transliteration. The evaluation for Indonesian- >English using Malay and for Spanish -> English using Portuguese and pretending Spanish is resource-poor shows an absolute gain of up to 1.35 and 3.37 BLEU points, respectively, which is an improvement over the best rivaling approaches, while using much less additional data. Overall, our method cuts the amount of necessary "real'' training data by a factor of 2--5.
Preslav Nakov, Hwee Tou Ng
J. Artif. Intell. Res.2
2011 Grammatical Error Correction with Alternating Structure Optimization
Daniel Dahlmeier, Hwee Tou Ng
ACL2
2011 Automatically Evaluating Text Coherence Using Discourse Relations
Ziheng Lin, Hwee Tou Ng, Min-Yen Kan
ACL2
2011 Translating from Morphologically Complex Languages: A Paraphrase-Based Approach
Preslav Nakov, Hwee Tou Ng
ACL2
2011 Correcting Semantic Collocation Errors with L1-induced Paraphrases
Daniel Dahlmeier, Hwee Tou Ng
EMNLP2
2011 Better Evaluation Metrics Lead to Better Machine Translation
Daniel Dahlmeier, Hwee Tou Ng
EMNLP3
2011 A Probabilistic Forest-to-String Model for Language Generation from Typed Lambda Calculus Expressions
Wei Lu 0011, Hwee Tou Ng
EMNLP2
2011 Enriching document representation via translation for improved monolingual information retrieval
abstract
Word ambiguity and vocabulary mismatch are critical problems in information retrieval. To deal with these problems, this paper proposes the use of translated words to enrich document representation, going beyond the words in the original source language to represent a document. In our approach, each original document is automatically translated into an auxiliary language, and the resulting translated document serves as a semantically enhanced representation for supplementing the original bag of words. The core of our translation representation is the expected term frequency of a word in a translated document, which is calculated by averaging the term frequencies over all possible translations, rather than focusing on the 1-best translation only. To achieve better efficiency of translation, we do not rely on full-fledged machine translation, but instead use monotonic translation by removing the time-consuming reordering component. Experiments carried out on standard TREC test collections show that our proposed translation representation leads to statistically significant improvements over using only the original language of the document collection.
Seung-Hoon Na, Hwee Tou Ng
SIGIR2
2010 Joint Syntactic and Semantic Parsing of Chinese
Junhui Li 0001, Guodong Zhou 0001, Hwee Tou Ng
ACL3
2010 Maximum Metric Score Training for Coreference Resolution
Shanheng Zhao, Hwee Tou Ng
COLING2
2010 PEM: A Paraphrase Evaluation Metric Exploiting Parallel Texts
Daniel Dahlmeier, Hwee Tou Ng
EMNLP3
2010 Better Punctuation Prediction with Dynamic Conditional Random Fields
Wei Lu 0011, Hwee Tou Ng
EMNLP2
2010 Domain adaptation for semantic role labeling in the biomedical domain
abstract
MOTIVATION: Semantic role labeling (SRL) is a natural language processing (NLP) task that extracts a shallow meaning representation from free text sentences. Several efforts to create SRL systems for the biomedical domain have been made during the last few years. However, state-of-the-art SRL relies on manually annotated training instances, which are rare and expensive to prepare. In this article, we address SRL for the biomedical domain as a domain adaptation problem to leverage existing SRL resources from the newswire domain. RESULTS: We evaluate the performance of three recently proposed domain adaptation algorithms for SRL. Our results show that by using domain adaptation, the cost of developing an SRL system for the biomedical domain can be reduced significantly. Using domain adaptation, our system can achieve 97% of the performance with as little as 60 annotated target domain abstracts. AVAILABILITY: Our BioKIT system that performs SRL in the biomedical domain as described in this article is implemented in Python and C and operates under the Linux operating system. BioKIT can be downloaded at http://nlp.comp.nus.edu.sg/software. The domain adaptation software is available for download at http://www.mysmu.edu/faculty/jingjiang/software/DALR.html. The BioProp corpus is available from the Linguistic Data Consortium http://www.ldc.upenn.edu.
Daniel Dahlmeier, Hwee Tou Ng
Bioinform.2
2010 The State of the Journal
abstract
No abstract available.
Hwee Tou Ng
ACM Trans. Asian Lang. Inf. Process.1
2010 Statistical lattice-based spoken document retrieval
abstract
Recent research efforts on spoken document retrieval have tried to overcome the low quality of 1-best automatic speech recognition transcripts, especially in the case of conversational speech, by using statistics derived from speech lattices containing multiple transcription hypotheses as output by a speech recognizer. We present a method for lattice-based spoken document retrieval based on a statistical n -gram modeling approach to information retrieval. In this statistical lattice-based retrieval (SLBR) method, a smoothed statistical model is estimated for each document from the expected counts of words given the information in a lattice, and the relevance of each document to a query is measured as a probability under such a model. We investigate the efficacy of our method under various parameter settings of the speech recognition and lattice processing engines, using the Fisher English Corpus of conversational telephone speech. Experimental results show that our method consistently achieves better retrieval performance than using only the 1-best transcripts in statistical retrieval, outperforms a recently proposed lattice-based vector space retrieval method, and also compares favorably with a lattice-based retrieval method based on the Okapi BM25 model.
Tee Kiah Chia, Khe Chai Sim, Haizhou Li 0001, Hwee Tou Ng
ACM Trans. Inf. Syst.4
2009 Joint Learning of Preposition Senses and Semantic Roles of Prepositional Phrases
Daniel Dahlmeier, Hwee Tou Ng, Tanja Schultz
EMNLP2
2009 Recognizing Implicit Discourse Relations in the Penn Discourse Treebank
Ziheng Lin, Min-Yen Kan, Hwee Tou Ng
EMNLP3
2009 Natural Language Generation with Tree Conditional Random Fields
Wei Lu 0011, Hwee Tou Ng, Wee Sun Lee
EMNLP2
2009 Improved Statistical Machine Translation for Resource-Poor Languages Using Related Resource-Rich Languages
Preslav Nakov, Hwee Tou Ng
EMNLP2
2009 Word Sense Disambiguation for All Words without Hard Labor
Hwee Tou Ng
IJCAI2
2009 A 2-poisson model for probabilistic coreference of named entities for improved text retrieval
abstract
Text retrieval queries frequently contain named entities. The standard approach of term frequency weighting does not work well when estimating the term frequency of a named entity, since anaphoric expressions (like he, she, the movie, etc) are frequently used to refer to named entities in a document, and the use of anaphoric expressions causes the term frequency of named entities to be underestimated. In this paper, we propose a novel 2-Poisson model to estimate the frequency of anaphoric expressions of a named entity, without explicitly resolving the anaphoric expressions. Our key assumption is that the frequency of anaphoric expressions is distributed over named entities in a document according to the probabilities of whether the document is elite for the named entities. This assumption leads us to formulate our proposed Co-referentially Enhanced Entity Frequency (CEEF). Experimental results on the text collection of TREC Blog Track show that CEEF achieves significant and consistent improvements over state-of-the-art retrieval methods using standard term frequency estimation. In particular, we achieve a 3% increase of MAP over the best performing run of TREC 2008 Blog Track.
Seung-Hoon Na, Hwee Tou Ng
SIGIR2
2009 MaxSim: performance and effects of translation fluency
Yee Seng Chan, Hwee Tou Ng
Mach. Transl.2
2008 MAXSIM: A Maximum Similarity Metric for Machine Translation Evaluation
Yee Seng Chan, Hwee Tou Ng
ACL2
2008 Decomposability of Translation Metrics for Improved Evaluation and Efficient Algorithms
David Chiang 0001, Steve DeNeefe, Yee Seng Chan, Hwee Tou Ng
EMNLP4
2008 A Generative Model for Parsing Natural Language to Meaning Representations
Wei Lu 0011, Hwee Tou Ng, Wee Sun Lee, Luke Zettlemoyer
EMNLP2
2008 Word Sense Disambiguation Using OntoNotes: An Empirical Study
Hwee Tou Ng, Yee Seng Chan
EMNLP2
2008 A lattice-based approach to query-by-example spoken document retrieval
abstract
Recent efforts on the task of spoken document retrieval (SDR) have made use of speech lattices: speech lattices contain information about alternative speech transcription hypotheses other than the 1-best transcripts, and this information can improve retrieval accuracy by overcoming recognition errors present in the 1-best transcription. In this paper, we look at using lattices for the query-by-example spoken document retrieval task - retrieving documents from a speech corpus, where the queries are themselves in the form of complete spoken documents (query exemplars). We extend a previously proposed method for SDR with short queries to the query-by-example task. Specifically, we use a retrieval method based on statistical modeling: we compute expected word counts from document and query lattices, estimate statistical models from these counts, and compute relevance scores as divergences between these models. Experimental results on a speech corpus of conversational English show that the use of statistics from lattices for both documents and query exemplars results in better retrieval accuracy than using only 1-best transcripts for either documents, or queries, or both. In addition, we investigate the effect of stop word removal which further improves retrieval accuracy. To our knowledge, our work is the first to have used a lattice-based approach to query-by-example spoken document retrieval.
Tee Kiah Chia, Khe Chai Sim, Haizhou Li 0001, Hwee Tou Ng
SIGIR4
2007 Domain Adaptation with Active Learning for Word Sense Disambiguation
Yee Seng Chan, Hwee Tou Ng
ACL2
2007 Word Sense Disambiguation Improves Statistical Machine Translation
Yee Seng Chan, Hwee Tou Ng, David Chiang 0001
ACL2
2007 Learning Predictive Structures for Semantic Role Labeling of NomBank
Hwee Tou Ng
ACL2
2007 A Unified Tagging Approach to Text Normalization
Conghui Zhu, Jie Tang 0001, Hang Li 0001, Hwee Tou Ng, Tiejun Zhao
ACL4
2007 A Statistical Language Modeling Approach to Lattice-Based Spoken Document Retrieval
Tee Kiah Chia, Haizhou Li 0001, Hwee Tou Ng
EMNLP-CoNLL3
2007 Identification and Resolution of Chinese Zero Pronouns: A Machine Learning Approach
Shanheng Zhao, Hwee Tou Ng
EMNLP-CoNLL2
2007 One Class per Named Entity: Exploiting Unlabeled Text for Named Entity Recognition
Yingchuan Wong, Hwee Tou Ng
IJCAI2
2006 Estimating Class Priors in Domain Adaptation for Word Sense Disambiguation
abstract
Instances of a word drawn from different domains may have different sense priors (the proportions of the different senses of a word). This in turn affects the accuracy of word sense disambiguation (WSD) systems trained and applied on different domains. This paper presents a method to estimate the sense priors of words drawn from a new domain, and highlights the importance of using well calibrated probabilities when performing these estimations. By using well calibrated probabilities, we are able to estimate the sense priors effectively to achieve significant improvements in WSD accuracy.
Yee Seng Chan, Hwee Tou Ng
ACL2
2006 Semantic Role Labeling of NomBank: A Maximum Entropy Approach
Zheng Ping Jiang, Hwee Tou Ng
EMNLP2
2005 Scaling Up Word Sense Disambiguation via Parallel Texts
Yee Seng Chan, Hwee Tou Ng
AAAI2
2005 Word Sense Disambiguation with Semi-Supervised Learning
Thanh Phong Pham 0001, Hwee Tou Ng, Wee Sun Lee
AAAI2
2005 Word Sense Disambiguation with Distribution Estimation
Yee Seng Chan, Hwee Tou Ng
IJCAI2
2005 Semantic Argument Classification Exploiting Argument Interdependence
Zheng Ping Jiang, Hwee Tou Ng
IJCAI3
2005 A Machine Learning Approach to Identification and Resolution of One-Anaphora
Hwee Tou Ng, Robert Dale, Mary Gardiner
IJCAI1
2004 Mining New Word Translations from Comparable Corpora
Li Shao, Hwee Tou Ng
COLING2
2004 Chinese Part-of-Speech Tagging: One-at-a-Time or All-at-Once? Word-Based or Character-Based?
Hwee Tou Ng, Jin Kiat Low
EMNLP1
2003 Closing the Gap: Learning-Based Information Extraction Rivaling Knowledge-Engineering Methods
abstract
In this paper, we present a learning approach to the scenario template task of information extraction, where information filling one template could come from multiple sentences. When tested on the MUC-4 task, our learning approach achieves accuracy competitive to the best of the MUC-4 systems, which were all built with manually engineered rules. Our analysis reveals that our use of full parsing and state-of-the-art learning algorithms have contributed to the good performance. To our knowledge, this is the first research to have demonstrated that a learning approach to the full-scale information extraction task could achieve performance rivaling that of the knowledge engineering approach.
Hai Leong Chieu, Hwee Tou Ng, Yoong Keok Lee
ACL2
2003 Exploiting Parallel Texts for Word Sense Disambiguation: An Empirical Study
abstract
A central problem of word sense disambiguation (WSD) is the lack of manually sense-tagged data required for supervised learning. In this paper, we evaluate an approach to automatically acquire sense-tagged training data from English-Chinese parallel corpora, which are then used for disambiguating the nouns in the SENSEVAL-2 English lexical sample task. Our investigation reveals that this method of acquiring sense-tagged data is promising. On a subset of the most difficult SENSEVAL-2 nouns, the accuracy difference between the two approaches is only 14.0%, and the difference could narrow further to 6.5% if we disregard the advantage that manually sense-tagged data have in their sense coverage. Our analysis also highlights the importance of the issue of domain dependence in evaluating WSD programs.
Hwee Tou Ng, Yee Seng Chan
ACL1
2003 Named Entity Recognition with a Maximum Entropy Approach
Hai Leong Chieu, Hwee Tou Ng
CoNLL2
2003 Mining topic-specific concepts and definitions on the web
abstract
Traditionally, when one wants to learn about a particular topic, one reads a book or a survey paper. With the rapid expansion of the Web, learning in-depth knowledge about a topic from the Web is becoming increasingly important and popular. This is also due to the Web's convenience and its richness of information. In many cases, learning from the Web may even be essential because in our fast changing world, emerging topics appear constantly and rapidly. There is often not enough time for someone to write a book on such topics. To learn such emerging topics, one can resort to research papers. However, research papers are often hard to understand by non-researchers, and few research papers cover every aspect of the topic. In contrast, many Web pages often contain intuitive descriptions of the topic. To find such Web pages, one typically uses a search engine. However, current search techniques are not designed for in-depth learning. Top ranking pages from a search engine may not contain any description of the topic. Even if they do, the description is usually incomplete since it is unlikely that the owner of the page has good knowledge of every aspect of the topic. In this paper, we attempt a novel and challenging task, mining topic-specific knowledge on the Web. Our goal is to help people learn in-depth knowledge of a topic systematically on the Web. The proposed techniques first identify those sub-topics or salient concepts of the topic, and then find and organize those informative pages, containing definitions and descriptions of the topic and sub-topics, just like those in a book. Experimental results using 28 topics show that the proposed techniques are highly effective.
Bing Liu 0001, Chee Wee Chin, Hwee Tou Ng
WWW3
2002 Teaching a Weaker Classifier: Named Entity Recognition on Upper Case Text
abstract
This paper describes how a machine-learning named entity recognizer (NER) on upper case text can be improved by using a mixed case NER and some unlabeled text. The mixed case NER can be used to tag some unlabeled mixed case text, which are then used as additional training material for the upper case NER. We show that this approach reduces the performance gap between the mixed case NER and the upper case NER substantially, by 39% for MUC-6 and 22% for MUC-7 named entity test data. Our method is thus useful in improving the accuracy of NERs on upper case text, such as transcribed text from automatic speech recognizers where case information is missing.
Hai Leong Chieu, Hwee Tou Ng
ACL2
2002 Named Entity Recognition: A Maximum Entropy Approach Using Global Information
Hai Leong Chieu, Hwee Tou Ng
COLING2
2002 An Empirical Evaluation of Knowledge Sources and Learning Algorithms for Word Sense Disambiguation
abstract
In this paper, we evaluate a variety of knowledge sources and supervised learning algorithms for word sense disambiguation on SENSEVAL-2 and SENSEVAL-1 data. Our knowledge sources include the part-of-speech of neighboring words, single words in the surrounding context, local collocations, and syntactic relations. The learning algorithms evaluated include Support Vector Machines (SVM), Naive Bayes, AdaBoost, and decision tree algorithms. We present empirical results showing the relative contribution of the component knowledge sources and the different learning algorithms. In particular, using all of these knowledge sources and SVM (i.e., a single learning algorithm) achieves accuracy higher than the best official scores on both SENSEVAL-2 and SENSEVAL-1 test data.
Yoong Keok Lee, Hwee Tou Ng
EMNLP2
2002 Refining the Wrapper Approach - Smoothed Error Estimates for Feature Selection
Loo-Nin Teow, Hwee Tou Ng, Eric Yap
ICML3
2002 Bayesian online classifiers for text classification and filtering
abstract
This paper explores the use of Bayesian online classifiers to classify text documents. Empirical results indicate that these classifiers are comparable with the best text classification systems. Furthermore, the online approach offers the advantage of continuous learning in the batch-adaptive text filtering task.
Kian Ming A. Chai, Hai Leong Chieu, Hwee Tou Ng
SIGIR3
2001 Question Answering Using a Large Text Database: A Machine Learning Approach
Hwee Tou Ng, Jennifer Lai-Pheng Kwan, Yiyuan Xia
EMNLP1
2001 A Machine Learning Approach to Coreference Resolution of Noun Phrases
abstract
In this paper, we present a learning approach to coreference resolution of noun phrases in unrestricted text. The approach learns from a small, annotated corpus and the task includes resolving not just a certain type of noun phrase (e.g., pronouns) but rather general noun phrases. It also does not restrict the entity types of the noun phrases; that is, coreference is assigned whether they are of “organization,” “person,” or other types. We evaluate our approach on common data sets (namely, the MUC-6 and MUC-7 coreference corpora) and obtain encouraging results, indicating that on the general noun phrase coreference task, the learning approach holds promise and achieves accuracy comparable to that of nonlearning approaches. Our system is the first learning-based system that offers performance comparable to that of state-of-the-art nonlearning systems on these data sets.
Wee Meng Soon, Hwee Tou Ng, Chung Yong Lim
Comput. Linguistics2
2000 A Machine Learning Approach to Answering Questions for Reading Comprehension Tests
abstract
In this paper, we report results on answering questions for the reading comprehension task, using a machine learning approach. We evaluated our approach on the Remedia data set, a common data set used in several recent papers on the reading comprehension task. Our learning approach achieves accuracy competitive to previous approaches that rely on hand-crafted, deterministic rules and algorithms. To the best of our knowledge, this is the first work that reports that the use of a machine learning approach achieves competitive results on answering questions for reading comprehension tests.
Hwee Tou Ng, Leonghwee Teo, Jennifer Lai-Pheng Kwan
EMNLP1
1999 Learning to Recognize Tables in Free Text
abstract
Many real-world texts contain tables. In order to process these texts correctly and extract the information contained within the tables, it is important to identify the presence and structure of tables. In this paper, we present a new approach that learns to recognize tables in free text, including the boundary, rows and columns of tables. When tested on Wall Street Journal news documents, our learning approach outperforms a deterministic table recognition algorithm that identifies table recognition algorithm that identifies tables based on a fixed set of conditions. Our learning approach is also more flexible and easily adaptable to texts in different domains with different table characteristics.
Hwee Tou Ng, Chung Yong Lim, Jessica Li Teng Koo
ACL1
1999 Corpus-Based Learning for Noun Phrase Coreference Resolution
Wee Meng Soon, Hwee Tou Ng, Chung Yong Lim
EMNLP2
1997 Exemplar-Based Word Sense Disambiguation" Some Recent Improvements
Hwee Tou Ng
EMNLP1
1997 Feature Selection, Perceptron Learning, and a Usability Case Study for Text Categorization
abstract
In this paper, we describe an automated learning approach to text categorization based on perception learning and a new feature selection metric, called correlation coefficient.Our approach has been teated on the standard Reuters text categorization collection.Empirical results indicate that our approach outperforms the best published results on this % uters collection.In particular, our new feature selection method yields comiderable improvement.We also investigate the usability of our automated hxu-n-~approach by actually developing a system that categorizes texts into a tree of categories.We compare tbe accuracy of our learning approach to a rrddmsed, expert system ap preach that uses a text categorization shell built by Cams gie Group.Although our automated learning approach still gives a lower accuracy, by appropriately inmrporating a set of manually chosen worda to use as f~ures, the combined, semi-automated approach yields accuracy close to the * baaed approach.
Hwee Tou Ng, Wei Boon Goh, Kok Leong Low
SIGIR1
1996 Integrating Multiple Knowledge Sources to Disambiguate Word Sense: An Exemplar-Based Approach
abstract
In this paper, we present a new approach for word sense disambiguation (WSD) using an exemplar-based learning algorithm.This approach integrates a diverse set of knowledge sources to disambiguate word sense, including part of speech of neighboring words, morphological form, the unordered set of surrounding words, local collocations, and verb-object syntactic relation.We tested our WSD program, named LEXAS, on both a common data set used in previous work, as well as on a large sense-tagged corpus that we separately constructed.LEXAS achieves a higher accuracy on the common data set, and performs better than the most frequent heuristic on the highly ambiguous words in the large corpus tagged with the refined senses of WoRDNET.
Hwee Tou Ng, Hian Beng Lee
ACL1
1992 Abductive Plan Recognition and Diagnosis: A Comprehensive Empirical Evaluation
Hwee Tou Ng, Raymond J. Mooney
KR1
1991 An Efficient First-Order Horn-Clause Abduction System Based on the ATMS
Hwee Tou Ng, Raymond J. Mooney
AAAI1
1990 On the Role of Coherence in Abductive Explanation
Hwee Tou Ng, Raymond J. Mooney
AAAI1