EDBT 2026 Demo / reviewers in the wild / expert
Thien Huu Nguyen
dblp:17/9407
· DBLP profile ↗
85ranked-venue papers
5as first author
59since 2021 · last 2026
0000-0003-3768-4736ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 75 · 5 first-author · 53 since 2021Graphics, computer vision, multimedia, augmented reality and games · 15 · 1 first-author · 10 since 2021Databases, data management, data science and information retrieval · 11 · 6 since 2021Human-computer interaction and ubiquitous computing · 1Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GloCTM: Cross-Lingual Topic Modeling via a Global Context SpaceabstractCross-lingual topic modeling seeks to uncover coherent and semantically aligned topics across languages—a task central to multilingual understanding. Yet most existing models learn topics in disjoint, language-specific spaces and rely on alignment mechanisms (e.g., bilingual dictionaries) that often fail to capture deep cross-lingual semantics, resulting in loosely connected topic spaces. Moreover, these approaches often overlook the rich semantic signals embedded in multilingual pretrained representations, further limiting their ability to capture fine-grained alignment. We introduce **GloCTM** (**Glo**bal Context Space for **C**ross-Lingual **T**opic **M**odel), a novel framework that enforces cross-lingual topic alignment through a unified semantic space spanning the entire model pipeline. GloCTM constructs enriched input representations by expanding bag-of-words with cross-lingual lexical neighborhoods, and infers topic proportions using both local and global encoders, with their latent representations aligned through internal regularization. At the output level, the global topic-word distribution, defined over the combined vocabulary, structurally synchronizes topic meanings across languages. To further ground topics in deep semantic space, GloCTM incorporates a Centered Kernel Alignment (CKA) loss that aligns the latent topic space with multilingual contextual embeddings. Experiments across multiple benchmarks demonstrate that GloCTM significantly improves topic coherence and cross-lingual alignment, outperforming strong baselines. Nguyen Tien Phat, Ngo Vu Minh, Ngo Van Linh 0001, Nguyen Thi Ngoc Diep, Thien Huu Nguyen |
AAAI | 5 |
| 2026 | TALAS: Teacher-Anchored Layer Alignment with Adaptive Sharpness-Aware Minimization for Embedding DistillationabstractQuoc Phong Dao, Hoang Son Nguyen, Pham Khanh Chi, Linh Ngo Van, Nguyen Thi Ngoc Diep, Thien Huu Nguyen, Trung Le. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Quoc Phong Dao, Hoang Son Nguyen, Pham Khanh Chi, Ngo Van Linh 0001, Nguyen Thi Ngoc Diep, Thien Huu Nguyen, Trung Le 0001 |
ACL (1) | 6 |
| 2026 | Towards Fast and Accurate Modeling for Cross-Lingual Label ProjectionabstractInformation extraction (IE) systems rely on structured data for training, but such annotated data is highly imbalanced across languages, with low-resource languages receiving little attention.Label projection techniques aim to bridge this gap by transferring structured annotations from high-resource to low-resource languages.However, existing methods are either inaccurate or too slow for large-scale use.This work aims to address this problem by developing a more effective method that remains sufficiently efficient for large-scale projection.In particular, we propose to synthesize alignment sequence pairs and fine-tune an encoder model with span alignment objective, while controlling data influence during training.Experimental results across 50+ languages show that our framework consistently outperforms previous state-of-the-art methods while maintaining fast inference speed.In addition, we introduce EXP -the first benchmark for explicit evaluation of label projection, thereby reducing confounders and non-determinism in method assessment. Thang Le, Huy Huu Nguyen, Anh Tuan Luu, Thamar Solorio, Thien Huu Nguyen |
ACL (1) | 5 |
| 2026 | Explainable Disentangled Representation Learning for Generalizable Authorship Attribution in the Era of Generative AIabstractLearning robust representations of authorial style is crucial for authorship attribution and AI-generated text detection.However, existing methods often struggle with content-style entanglement, where models learn spurious correlations between authors' writing styles and topics, leading to poor generalization across domains.To address this challenge, we propose Explainable Authorship Variational Autoencoder (EAVAE), a novel framework that explicitly disentangles style from content through architectural separation-by-design. EAVAE first pretrains style encoders using supervised contrastive learning on diverse authorship data, then finetunes with a Variational Autoencoder (VEA) architecture using separate encoders for style and content representations.Disentanglement is enforced through a novel discriminator that not only distinguishes whether pairs of style/content representations belong to the same or different authors/content sources, but also generates natural language explanation for their decision, simultaneously mitigating confounding information and enhancing interpretability.Extensive experiments demonstrate the effectiveness of EAVAE.On authorship attribution, we achieve state-of-the-art performance on various datasets, including Amazon Reviews, PAN21, and HRS.For AI-generated text detection, EAVAE excels in few-shot learning over the M4 dataset.Code and data repositories are available online 1 2 . Hieu Man, Van-Cuong Pham, Nghia Trung Ngo, Franck Dernoncourt, Thien Huu Nguyen |
ACL (1) | 5 |
| 2026 | Lizard: An Efficient Linearization Framework for Large Language ModelsabstractChien Van Nguyen, Huy Huu Nguyen, Ruiyi Zhang, Hanieh Deilamsalehy, Puneet Mathur, Viet Dac Lai, Haoliang Wang, Jayakumar Subramanian, Ryan A. Rossi, Trung Bui, Nikos Vlassis, Franck Dernoncourt, Thien Huu Nguyen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Chien Van Nguyen, Huy Huu Nguyen, Ruiyi Zhang 0002, Hanieh Deilamsalehy, Puneet Mathur, Viet Dac Lai, Jayakumar Subramanian, Ryan Rossi, Trung Bui, Nikos Vlassis, Franck Dernoncourt, Thien Huu Nguyen |
ACL (1) | 13 |
| 2026 | Octopus: Gated Selective Attention for Memory-Bounded Long-Context Inference in Large Language ModelsabstractChien Van Nguyen, Ryan A. Rossi, Linh Ngo Van, Franck Dernoncourt, Thien Huu Nguyen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Chien Van Nguyen, Ryan Rossi, Ngo Van Linh 0001, Franck Dernoncourt, Thien Huu Nguyen |
ACL (1) | 5 |
| 2026 | CypherSmith: Transforming Text-to-Cypher Generation for LLMs with Synthetic DataabstractKnowledge Graph (KG) retrieval is a promising augmentation to address knowledge gaps and hallucinations in LLMs. As KGs in practice are stored in graph databases (e.g., Wikidata, Freebase), accurate retrieval requires translating natural language questions into structured queries (query generation). A key challenge of query generation is Text-to-Cypher, which generates Cypher queries for property graphs (e.g., Neo4j), a paradigm increasingly adopted in industry for their scalable architectures and expressive schemas. However, compared to other query generation tasks such as Text-to-SQL or Text-to-SPARQL, Text-to-Cypher remains underexplored due to scarce public KGs and datasets. Existing datasets are small, domain-limited, and lack diversity, constraining LLM progress. To address this, we introduce CypherSmith, an instruction-tuning dataset over 12\times larger than prior public Text-to-Cypher datasets, spanning diverse domains to better support LLM fine-tuning. Our key distinction lies in fully leveraging open-source LLMs for large-scale synthetic data generation and introducing a novel likelihood-based filtering technique to ensure high-quality Text-to-Cypher data. Extensive experiments demonstrate the effectiveness of CypherSmith, achieving state-of-the-art LLM performance. Kexuan Sun 0002, Jens-S. Vöckler, Thien Huu Nguyen, Thuy Vu |
ACL (1) | 5 |
| 2026 | mSCoRe: A Multilingual and Scalable Benchmark for Skill-based Commonsense Reasoning
Nghia Trung Ngo, Franck Dernoncourt, Thien Huu Nguyen |
LREC | 3 |
| 2025 | Adaptive Prompting for Continual Relation Extraction: A Within-Task Variance PerspectiveabstractTo address catastrophic forgetting in Continual Relation Extraction (CRE), many current approaches rely on memory buffers to rehearse previously learned knowledge while acquiring new tasks. Recently, prompt-based methods have emerged as potent alternatives to rehearsal-based strategies, demonstrating strong empirical performance. However, upon analyzing existing prompt-based approaches for CRE, we identified several critical limitations, such as inaccurate prompt selection, inadequate mechanisms for mitigating forgetting in shared parameters, and suboptimal handling of cross-task and within-task variances. To overcome these challenges, we draw inspiration from the relationship between prefix tuning and mixture of experts, proposing a novel approach that employs a prompt pool for each task, capturing variations within each task while enhancing cross-task variances. Furthermore, we incorporate a generative model to consolidate prior knowledge within shared parameters, eliminating the need for explicit data storage. Extensive experiments validate the efficacy of our approach, demonstrating superior performance over state-of-the-art prompt-based and rehearsal-free methods in continual relation extraction. Minh Le, Tien Ngoc Luu, An Nguyen The, Thanh-Thien Le, Tung Thanh Nguyen, Ngo Van Linh 0001, Thien Huu Nguyen |
AAAI | 8 |
| 2025 | Few-Shot, No Problem: Descriptive Continual Relation ExtractionabstractFew-shot Continual Relation Extraction is a crucial challenge for enabling AI systems to identify and adapt to evolving relationships in dynamic real-world domains. Traditional memory-based approaches often overfit to limited samples, failing to reinforce old knowledge, with the scarcity of data in few-shot scenarios further exacerbating these issues by hindering effective data augmentation in the latent space. In this paper, we propose a novel retrieval-based solution, starting with a large language model to generate descriptions for each relation. From these descriptions, we introduce a bi-encoder retrieval training paradigm to enrich both sample and class representation learning. Leveraging these enhanced representations, we design a retrieval-based prediction method where each sample "retrieves" the best fitting relation via a reciprocal rank fusion score that integrates both relation description vectors and class prototypes. Extensive experiments on multiple datasets demonstrate that our method significantly advances the state-of-the-art by maintaining robust performance across sequential tasks, effectively addressing catastrophic forgetting. Anh Duc Le, Quyen Tran, Thanh-Thien Le, Ngo Van Linh 0001, Thien Huu Nguyen |
AAAI | 6 |
| 2025 | Mitigating Non-Representative Prototypes and Representation Bias in Few-Shot Continual Relation ExtractionabstractTo address the phenomenon of similar classes, existing methods in few-shot continual relation extraction (FCRE) face two main challenges: non-representative prototypes and representation bias, especially when the number of available samples is limited. In our work, we propose Minion to address these challenges. Firstly, we leverage the General Orthogonal Frame (GOF) structure, based on the concept of Neural Collapse, to create robust class prototypes with clear separation, even between analogous classes. Secondly, we utilize label description representations as global class representatives within the fast-slow contrastive learning paradigm. These representations consistently encapsulate the essential attributes of each relation, acting as global information that helps mitigate overfitting and reduces representation bias caused by the limited local few-shot examples within a class. Extensive experiments on well-known FCRE benchmarks show that our method outperforms state-of-the-art approaches, demonstrating its effectiveness for advancing RE system. Thanh Duc Pham, Ngo Van Linh 0001, Nguyen Thi Ngoc Diep, Sang Dinh, Thien Huu Nguyen |
ACL (1) | 6 |
| 2025 | From Selection to Generation: A Survey of LLM-based Active LearningabstractYu Xia, Subhojyoti Mukherjee, Zhouhang Xie, Junda Wu, Xintong Li, Ryan Aponte, Hanjia Lyu, Joe Barrow, Hongjie Chen, Franck Dernoncourt, Branislav Kveton, Tong Yu, Ruiyi Zhang, Jiuxiang Gu, Nesreen K. Ahmed, Yu Wang, Xiang Chen, Hanieh Deilamsalehy, Sungchul Kim, Zhengmian Hu, Yue Zhao, Nedim Lipka, Seunghyun Yoon, Ting-Hao Kenneth Huang, Zichao Wang, Puneet Mathur, Soumyabrata Pal, Koyel Mukherjee, Zhehao Zhang, Namyong Park, Thien Huu Nguyen, Jiebo Luo, Ryan A. Rossi, Julian McAuley. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yu Xia 0007, Subhojyoti Mukherjee, Zhouhang Xie, Junda Wu, Xintong Li 0001, Ryan Aponte, Hanjia Lyu, Joe Barrow, Hongjie Chen 0003, Franck Dernoncourt, Branislav Kveton, Tong Yu 0001, Ruiyi Zhang 0002, Jiuxiang Gu, Nesreen K. Ahmed, Yu Wang 0160, Xiang Chen 0010, Hanieh Deilamsalehy, Sungchul Kim, Zhengmian Hu, Yue Zhao 0016, Nedim Lipka, Seunghyun Yoon 0002, Ting-Hao 'Kenneth' Huang, Zichao Wang 0001, Puneet Mathur, Soumyabrata Pal, Koyel Mukherjee 0001, Zhehao Zhang 0001, Namyong Park 0001, Thien Huu Nguyen, Jiebo Luo 0001, Ryan Rossi, Julian J. McAuley |
ACL (1) | 31 |
| 2025 | EMO: Embedding Model Distillation via Intra-Model Relation and Optimal Transport AlignmentsabstractKnowledge distillation (KD) is crucial for compressing large text embedding models, but faces challenges when teacher and student models use different tokenizers (Cross-Tokenizer KD -CTKD).Vocabulary mismatches impede the transfer of relational knowledge encoded in deep representations, such as hidden states and attention matrices, which are vital for producing high-quality embeddings.Existing CTKD methods often focus on direct output alignment, neglecting this crucial structural information.We propose a novel framework tailored for CTKD embedding model distillation.We first map tokens one-to-one via Minimum Edit Distance (MinED).Then, we distill intra-model relational knowledge by aligning attention matrix patterns using Centered Kernel Alignment, focusing on the top-m most important tokens of the directly mapped tokens.Simultaneously, we align final hidden states via Optimal Transport with Importance-Scored Mass Assignment, which emphasizes semantically important token representations, based on importance scores derived from attention weights.We evaluate distillation from state-of-the-art embedding models (e.g., LLM2Vec, BGE) to a Bert-base-uncased model on embedding-reliant tasks such as text classification, sentence pair classification, and semantic textual similarity.Our proposed framework significantly outperforms existing CTKD baselines.By preserving attention structure and prioritizing key representations, our approach yields smaller, highfidelity embedding models despite tokenizer differences. Minh-Phuc Truong, Hai An Vu, Tu Vu, Nguyen Thi Ngoc Diep, Ngo Van Linh 0001, Thien Huu Nguyen, Trung Le 0001 |
EMNLP | 6 |
| 2025 | Mutual-pairing Data Augmentation for Fewshot Continual Relation ExtractionabstractNguyen Hoang Anh, Quyen Tran, Thanh Xuan Nguyen, Nguyen Thi Ngoc Diep, Linh Ngo Van, Thien Huu Nguyen, Trung Le. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Nguyen Hoang Anh, Quyen Tran, Nguyen Thi Ngoc Diep, Ngo Van Linh 0001, Thien Huu Nguyen, Trung Le 0001 |
NAACL (Long Papers) | 6 |
| 2025 | Enhancing Discriminative Representation in Similar Relation Clusters for Few-Shot Continual Relation ExtractionabstractAnh Duc Le, Nam Le Hai, Thanh Xuan Nguyen, Linh Ngo Van, Nguyen Thi Ngoc Diep, Sang Dinh, Thien Huu Nguyen. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Anh Duc Le, Ngo Van Linh 0001, Nguyen Thi Ngoc Diep, Sang Dinh, Thien Huu Nguyen |
NAACL (Long Papers) | 7 |
| 2025 | Sharpness-Aware Minimization for Topic Models with High-Quality Document RepresentationsabstractTung Nguyen, Tue Le, Hoang Tran Vuong, Quang Duc Nguyen, Duc Anh Nguyen, Linh Ngo Van, Sang Dinh, Thien Huu Nguyen. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Tue Le, Hoang Tran Vuong, Ngo Van Linh 0001, Sang Dinh, Thien Huu Nguyen |
NAACL (Long Papers) | 8 |
| 2025 | GloCOM: A Short Text Neural Topic Model via Global Clustering ContextabstractQuang Duc Nguyen, Tung Nguyen, Duc Anh Nguyen, Linh Ngo Van, Sang Dinh, Thien Huu Nguyen. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Ngo Van Linh 0001, Sang Dinh, Thien Huu Nguyen |
NAACL (Long Papers) | 6 |
| 2025 | LUSIFER: Language Universal Space Integration for Enhanced Representation in Multilingual Text Embedding ModelsabstractRecent advancements in large language models (LLMs) based embedding models have established new state-of-the-art benchmarks for text embedding tasks, particularly in dense vector-based retrieval. However, these models predominantly focus on English, leaving multilingual embedding capabilities largely unexplored. To address this limitation, we present LUSIFER, a novel zero-shot approach that adapts LLM-based embedding models for multilingual tasks without requiring multilingual supervision. LUSIFER's architecture combines a multilingual encoder, serving as a language-universal learner, with an LLM-based embedding model optimized for embedding-specific tasks. These components are seamlessly integrated through a minimal set of trainable parameters that act as a connector, effectively transferring the multilingual encoder's language understanding capabilities to the specialized embedding model. Additionally, to comprehensively evaluate multilingual embedding performance, we introduce a new benchmark encompassing 5 primary embedding tasks, 123 diverse datasets, and coverage across 14 languages. Extensive experimental results demonstrate that LUSIFER significantly enhances the multilingual performance across various embedding tasks, particularly for medium and low-resource languages, without requiring explicit multilingual training data. The code and dataset for training are available at: https://github.com/hieum98/lusifer Hieu Man, Nghia Trung Ngo, Viet Dac Lai, Ryan Rossi, Franck Dernoncourt, Thien Huu Nguyen |
SIGIR | 6 |
| 2024 | Continual Relation Extraction via Sequential Multi-Task LearningabstractTo build continual relation extraction (CRE) models, those can adapt to an ever-growing ontology of relations, is a cornerstone information extraction task that serves in various dynamic real-world domains. To mitigate catastrophic forgetting in CRE, existing state-of-the-art approaches have effectively utilized rehearsal techniques from continual learning and achieved remarkable success. However, managing multiple objectives associated with memory-based rehearsal remains underexplored, often relying on simple summation and overlooking complex trade-offs. In this paper, we propose Continual Relation Extraction via Sequential Multi-task Learning (CREST), a novel CRE approach built upon a tailored Multi-task Learning framework for continual learning. CREST takes into consideration the disparity in the magnitudes of gradient signals of different objectives, thereby effectively handling the inherent difference between multi-task learning and continual learning. Through extensive experiments on multiple datasets, CREST demonstrates significant improvements in CRE performance as well as superiority over other state-of-the-art Multi-task Learning frameworks, offering a promising solution to the challenges of continual learning in this domain. Thanh-Thien Le, Tung Thanh Nguyen, Ngo Van Linh 0001, Thien Huu Nguyen |
AAAI | 5 |
| 2024 | Mastering Context-to-Label Representation Transformation for Event Causality Identification with Diffusion ModelsabstractTo understand event structures of documents, event causality identification (ECI) emerges as a crucial task, aiming to discern causal relationships among event mentions. The latest approach for ECI has introduced advanced deep learning models where transformer-based encoding models, complemented by enriching components, are typically leveraged to learn effective event context representations for causality prediction. As such, an important step for ECI models is to transform the event context representations into causal label representations to perform logits score computation for training and inference purposes. Within this framework, event context representations might encapsulate numerous complicated and noisy structures due to the potential long context between the input events while causal label representations are intended to capture pure information about the causal relations to facilitate score estimation. Nonetheless, a notable drawback of existing ECI models stems from their reliance on simple feed-forward networks to handle the complex context-to-label representation transformation process, which might require drastic changes in the representations to hinder the learning process. To overcome this issue, our work introduces a novel method for ECI where, instead abrupt transformations, event context representations are gradually updated to achieve effective label representations. This process will be done incrementally to allow filtering of irrelevant structures at varying levels of granularity for causal relations. To realize this, we present a diffusion model to learn gradual representation transition processes between context and causal labels. It operates through a forward pass for causal label representation noising and a reverse pass for reconstructing label representations from random noise. Our experiments on different datasets across multiple languages demonstrate the advantages of the diffusion model with state-of-the-art performance for ECI. Hieu Man, Franck Dernoncourt, Thien Huu Nguyen |
AAAI | 3 |
| 2024 | CAMAL: A Novel Dataset for Multi-label Conversational Argument Move AnalysisabstractUnderstanding the discussion moves that teachers and students use to engage in classroom discussions is important to support pre-service teacher learning and teacher educators. This work introduces a novel conversational multi-label corpus of teaching transcripts collected from a simulated classroom environment for Conversational Argument Move AnaLysis (CAMAL). The dataset offers various argumentation moves used by pre-service teachers and students in mathematics and science classroom discussions. The dataset includes 165 transcripts from these discussions that pre-service elementary teachers facilitated in a simulated classroom environment of five student avatars. The discussion transcripts were annotated by education assessment experts for nine argumentation moves (aka. intents) used by the pre-service teachers and students during the discussions. In this paper, we describe the dataset, our annotation framework, and the models we employed to detect argumentation moves. Our experiments with state-of-the-art models demonstrate the complexity of the CAMAL task presented in the dataset. The result reveals that models that combined CNN and LSTM structures with speaker ID graphs improved the F1-score of our baseline models to detect speakers’ intents by a large margin. Given the complexity of the CAMAL task, it creates research opportunities for future studies. We share the dataset, the source code, and the annotation framework publicly at http://github.com/uonlp/camal-dataset. Viet Dac Lai, Duy Ngoc Pham, Jonathan Steinberg, Jamie Mikeska, Thien Huu Nguyen |
LREC/COLING | 5 |
| 2024 | Hierarchical Selection of Important Context for Generative Event Causality Identification with Optimal TransportsabstractWe study the problem of Event Causality Identification (ECI) that seeks to predict causal relation between event mentions in the text. In contrast to previous classification-based models, a few recent ECI methods have explored generative models to deliver state-of-the-art performance. However, such generative models cannot handle document-level ECI where long context between event mentions must be encoded to secure correct predictions. In addition, previous generative ECI methods tend to rely on external toolkits or human annotation to obtain necessary training signals. To address these limitations, we propose a novel generative framework that leverages Optimal Transport (OT) to automatically select the most important sentences and words from full documents. Specifically, we introduce hierarchical OT alignments between event pairs and the document to extract pertinent contexts. The selected sentences and words are provided as input and output to a T5 encoder-decoder model which is trained to generate both the causal relation label and salient contexts. This allows richer supervision without external tools. We conduct extensive evaluations on different datasets with multiple languages to demonstrate the benefits and state-of-the-art performance of ECI. Hieu Man, Chien Van Nguyen, Nghia Trung Ngo, Ngo Van Linh 0001, Franck Dernoncourt, Thien Huu Nguyen |
LREC/COLING | 6 |
| 2024 | CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 LanguagesabstractExtensive training datasets represent one of the important factors for the impressive learning capabilities of large language models (LLMs). However, these training datasets for current LLMs, especially the recent state-of-the-art models, are often not fully disclosed. Creating training data for high-performing LLMs involves extensive cleaning and deduplication to ensure the necessary level of quality. The lack of transparency for training data has thus hampered research on attributing and addressing hallucination and bias issues in LLMs, hindering replication efforts and further advancements in the community. These challenges become even more pronounced in multilingual learning scenarios, where the available multilingual text datasets are often inadequately collected and cleaned. Consequently, there is a lack of open-source and readily usable dataset to effectively train LLMs in multiple languages. To overcome this issue, we present CulturaX, a substantial multilingual dataset with 6.3 trillion tokens in 167 languages, tailored for LLM development. Our dataset undergoes meticulous cleaning and deduplication through a rigorous pipeline of multiple stages to accomplish the best quality for model training, including language identification, URL-based filtering, metric-based cleaning, document refinement, and data deduplication. CulturaX is released in Hugging Face facilitate research and advancements in multilingual LLMs: https://huggingface.co/datasets/uonlp/CulturaX. Thuat Nguyen, Chien Van Nguyen, Viet Dac Lai, Hieu Man, Nghia Trung Ngo, Franck Dernoncourt, Ryan Rossi, Thien Huu Nguyen |
LREC/COLING | 8 |
| 2024 | BKEE: Pioneering Event Extraction in the Vietnamese LanguageabstractEvent Extraction (EE) is a fundamental task in information extraction, aimed at identifying events and their associated arguments within textual data. It holds significant importance in various applications and serves as a catalyst for the development of related tasks. Despite the availability of numerous datasets and methods for event extraction in various languages, there has been a notable absence of a dedicated dataset for the Vietnamese language. To address this limitation, we propose BKEE, a novel event extraction dataset for Vietnamese. BKEE encompasses over 33 distinct event types and 28 different event argument roles, providing a labeled dataset for entity mentions, event mentions, and event arguments on 1066 documents. Additionally, we establish robust baselines for potential downstream tasks on this dataset, facilitating the analysis of challenges and future development prospects in the field of Vietnamese event extraction. Thi-Nhung Nguyen, Tien-Bang Tran, Trong-Nghia Luu, Thien Huu Nguyen, Kiem-Hieu Nguyen |
LREC/COLING | 4 |
| 2024 | Lifelong Event Detection via Optimal TransportabstractContinual Event Detection (CED) poses a formidable challenge due to the catastrophic forgetting phenomenon, where learning new tasks (with new coming event types) hampers performance on previous ones.In this paper, we introduce a novel approach, Lifelong Event Detection via Optimal Transport (LEDOT), that leverages optimal transport principles to align the optimization of our classification module with the intrinsic nature of each class, as defined by their pre-trained language modeling.Our method integrates replay sets, prototype latent representations, and an innovative Optimal Transport component.Extensive experiments on MAVEN and ACE datasets demonstrate LEDOT's superior performance, consistently outperforming state-of-the-art baselines.The results underscore LEDOT as a pioneering solution in continual event detection, offering a more effective and nuanced approach to addressing catastrophic forgetting in evolving environments. Viet Dao, Van-Cuong Pham, Quyen Tran, Thanh-Thien Le, Ngo Van Linh 0001, Thien Huu Nguyen |
EMNLP | 6 |
| 2024 | Preserving Generalization of Language models in Few-shot Continual Relation ExtractionabstractFew-shot Continual Relations Extraction (FCRE) is an emerging and dynamic area of study where models can sequentially integrate knowledge from new relations with limited labeled data while circumventing catastrophic forgetting and preserving prior knowledge from pre-trained backbones.In this work, we introduce a novel method that leverages oftendiscarded language model heads.By employing these components via a mutual information maximization strategy, our approach helps maintain prior knowledge from the pre-trained backbone and strategically aligns the primary classification head, thereby enhancing model performance.Furthermore, we explore the potential of Large Language Models (LLMs), renowned for their wealth of knowledge, in addressing FCRE challenges.Our comprehensive experimental results underscore the efficacy of the proposed method and offer valuable insights for future work. Quyen Tran, Nguyen Hoang Anh, Trung Le 0001, Ngo Van Linh 0001, Thien Huu Nguyen |
EMNLP | 7 |
| 2024 | Identifying Speakers in Dialogue Transcripts: A Text-based Approach Using Pretrained Language Models
Minh Nguyen 0007, Franck Dernoncourt, Seunghyun Yoon 0002, Hanieh Deilamsalehy, Hao Tan 0002, Ryan Rossi, Quan Hung Tran, Trung Bui, Thien Huu Nguyen |
INTERSPEECH | 9 |
| 2024 | SharpSeq: Empowering Continual Event Detection through Sharpness-Aware Sequential-task LearningabstractThanh-Thien Le, Viet Dao, Linh Nguyen, Thi-Nhung Nguyen, Linh Ngo, Thien Nguyen. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Thanh-Thien Le, Viet Dao, Thi-Nhung Nguyen, Ngo Van Linh 0001, Thien Huu Nguyen |
NAACL-HLT | 6 |
| 2024 | Counterfactual Augmentation for Robust Authorship Representation LearningabstractAuthorship attribution is a task that aims to identify the author of given pieces of writing. Authorship representation learning using neural networks has been shown to work in open-set environment settings with hundreds of thousands of authors. However, the performance of authorship attribution models often degrades significantly when texts are from different domains than the training data. In this work, we propose addressing this issue by adopting a novel causal framework for authorship representation learning. Our key insight is to use causal interventions during training to make models robust to differences in domains. Specifically, we introduce generating style-counterfactual examples by retrieving the most similar content texts by different authors on the same topics/domains. This exposes the model to challenging examples with similar content but distinct styles. Furthermore, we introduce causal masking of topic-indicative words to generate content-counterfactual examples. Content-counterfactuals hide topic content to encourage focusing on writing style. Experiments on three disparate domains - Amazon reviews, fanfiction stories, and Reddit comments - demonstrate that our approach significantly outperforms previous state-of-the-art methods for authorship attribution. Hieu Man, Thien Huu Nguyen |
SIGIR | 2 |
| 2024 | Continual variational dropout: a view of auxiliary local variables in continual learning
Ngo Van Linh 0001, Thien Huu Nguyen, Khoat Than |
Mach. Learn. | 4 |
| 2023 | Hybrid Knowledge Transfer for Improved Cross-Lingual Event Detection via Hierarchical Sample SelectionabstractIn this paper, we address the Event Detection task under a zero-shot cross-lingual setting where a model is trained on a source language but evaluated on a distinct target language for which there is no labeled data available.Most recent efforts in this field follow a direct transfer approach in which the model is trained using language-invariant features and then directly applied to the target language.However, we argue that these methods fail to take advantage of the benefits of the data transfer approach where a cross-lingual model is trained on targetlanguage data and is able to learn task-specific information from syntactical features or wordlabel relations in the target language.As such, we propose a hybrid knowledge-transfer approach that leverages a teacher-student framework where the teacher and student networks are trained following the direct and data transfer approaches, respectively.Our method is complemented by a hierarchical training-sample selection scheme designed to address the issue of noisy labels being generated by the teacher model.Our model achieves state-of-the-art results on 9 morphologically-diverse target languages across 3 distinct datasets, highlighting the importance of exploiting the benefits of hybrid transfer. Luis Guzman-Nateras, Franck Dernoncourt, Thien Huu Nguyen |
ACL (1) | 3 |
| 2023 | Boosting Punctuation Restoration with Data Generation and Reinforcement Learning
Viet Dac Lai, Abel Salinas, Hao Tan 0002, Trung Bui, Quan Tran, Seunghyun Yoon 0002, Hanieh Deilamsalehy, Franck Dernoncourt, Thien Huu Nguyen |
INTERSPEECH | 9 |
| 2023 | Question-Context Alignment and Answer-Context Dependencies for Effective Answer Sentence Selection
Minh Nguyen 0007, Kishan KC, Toàn Quoc Nguyên, Thien Huu Nguyen, Ankit Chadha, Thuy Vu |
INTERSPEECH | 4 |
| 2022 | Selecting Optimal Context Sentences for Event-Event Relation ExtractionabstractUnderstanding events entails recognizing the structural and temporal orders between event mentions to build event structures/graphs for input documents. To achieve this goal, our work addresses the problems of subevent relation extraction (SRE) and temporal event relation extraction (TRE) that aim to predict subevent and temporal relations between two given event mentions/triggers in texts. Recent state-of-the-art methods for such problems have employed transformer-based language models (e.g., BERT) to induce effective contextual representations for input event mention pairs. However, a major limitation of existing transformer-based models for SRE and TRE is that they can only encode input texts of limited length (i.e., up to 512 sub-tokens in BERT), thus unable to effectively capture important context sentences that are farther away in the documents. In this work, we introduce a novel method to better model document-level context with important context sentences for event-event relation extraction. Our method seeks to identify the most important context sentences for a given entity mention pair in a document and pack them into shorter documents to be consume entirely by transformer-based language models for representation learning. The REINFORCE algorithm is employed to train models where novel reward functions are presented to capture model performance, and context-based and knowledge-based similarity between sentences for our problem. Extensive experiments demonstrate the effectiveness of the proposed method with state-of-the-art performance on benchmark datasets. Hieu Man, Nghia Trung Ngo, Ngo Van Linh 0001, Thien Huu Nguyen |
AAAI | 4 |
| 2022 | MECI: A Multilingual Dataset for Event Causality IdentificationabstractEvent Causality Identification (ECI) is the task of detecting causal relations between events mentioned in the text. Although this task has been extensively studied for English materials, it is under-explored for many other languages. A major reason for this issue is the lack of multilingual datasets that provide consistent annotations for event causality relations in multiple non-English languages. To address this issue, we introduce a new multilingual dataset for ECI, called MECI. The dataset employs consistent annotation guidelines for five typologically different languages, i.e., English, Danish, Spanish, Turkish, and Urdu. Our dataset thus enable a new research direction on cross-lingual transfer learning for ECI. Our extensive experiments demonstrate high quality for MECI that can provide ample research challenges and directions for future research. We will publicly release MECI to promote research on multilingual ECI. Viet Dac Lai, Amir Pouran Ben Veyseh, Minh Nguyen 0007, Franck Dernoncourt, Thien Huu Nguyen |
COLING | 5 |
| 2022 | Unsupervised Domain Adaptation for Text Classification via Meta Self-Paced LearningabstractA shift in data distribution can have a significant impact on performance of a text classification model. Recent methods addressing unsupervised domain adaptation for textual tasks typically extracted domain-invariant representations through balancing between multiple objectives to align feature spaces between source and target domains. While effective, these methods induce various new domain-sensitive hyperparameters, thus are impractical as large-scale language models are drastically growing bigger to achieve optimal performance. To this end, we propose to leverage meta-learning framework to train a neural network-based self-paced learning procedure in an end-to-end manner. Our method, called Meta Self-Paced Domain Adaption (MSP-DA), follows a novel but intuitive domain-shift variation of cluster assumption to derive the meta train-test dataset split based on the self-pacing difficulties of source domain’s examples. As a result, MSP-DA effectively leverages self-training and self-tuning domain-specific hyperparameters simultaneously throughout the learning process. Extensive experiments demonstrate our framework substantially improves performance on target domains, surpassing state-of-the-art approaches. Detailed analyses validate our method and provide insight into how each domain affects the learned hyperparameters. Nghia Trung Ngo, Ngo Van Linh 0001, Thien Huu Nguyen |
COLING | 3 |
| 2022 | Event Extraction in Video TranscriptsabstractEvent extraction (EE) is one of the fundamental tasks for information extraction whose goal is to identify mentions of events and their participants in text. Due to its importance, different methods and datasets have been introduced for EE. However, existing EE datasets are limited to formally written documents such as news articles or scientific papers. As such, the challenges of EE in informal and noisy texts are not adequately studied. In particular, video transcripts constitute an important domain that can benefit tremendously from EE systems (e.g., video retrieval), but has not been studied in EE literature due to the lack of necessary datasets. To address this limitation, we propose the first large-scale EE dataset obtained for transcripts of streamed videos on the video hosting platform Behance to promote future research in this area. In addition, we extensively evaluate existing state-of-the-art EE methods on our new dataset. We demonstrate that such systems cannot achieve adequate performance on the proposed dataset, revealing challenges and opportunities for further research effort. Amir Pouran Ben Veyseh, Viet Dac Lai, Franck Dernoncourt, Thien Huu Nguyen |
COLING | 4 |
| 2022 | MACRONYM: A Large-Scale Dataset for Multilingual and Multi-Domain Acronym ExtractionabstractAcronym extraction is the task of identifying acronyms and their expanded forms in texts that is necessary for various NLP applications. Despite major progress for this task in recent years, one limitation of existing AE research is that they are limited to the English language and certain domains (i.e., scientific and biomedical). Challenges of AE in other languages and domains are mainly unexplored. As such, lacking annotated datasets in multiple languages and domains has been a major issue to prevent research in this direction. To address this limitation, we propose a new dataset for multilingual and multi-domain AE. Specifically, 27,200 sentences in 6 different languages and 2 new domains, i.e., legal and scientific, are manually annotated for AE. Our experiments on the dataset show that AE in different languages and learning settings has unique challenges, emphasizing the necessity of further research on multilingual and multi-domain AE. Amir Pouran Ben Veyseh, Nicole Meister, Seunghyun Yoon 0002, Rajiv Jain, Franck Dernoncourt, Thien Huu Nguyen |
COLING | 6 |
| 2022 | Keyphrase Prediction from Video Transcripts: New Dataset and DirectionsabstractKeyphrase Prediction (KP) is an established NLP task, aiming to yield representative phrases to summarize the main content of a given document. Despite major progress in recent years, existing works on KP have mainly focused on formal texts such as scientific papers or weblogs. The challenges of KP in informal-text domains are not yet fully studied. To this end, this work studies new challenges of KP in transcripts of videos, an understudied domain for KP that involves informal texts and non-cohesive presentation styles. A bottleneck for KP research in this domain involves the lack of high-quality and large-scale annotated data that hinders the development of advanced KP models. To address this issue, we introduce a large-scale manually-annotated KP dataset in the domain of live-stream video transcripts obtained by automatic speech recognition tools. Concretely, transcripts of 500+ hours of videos streamed on the behance.net platform are manually labeled with important keyphrases. Our analysis of the dataset reveals the challenging nature of KP in transcripts. Moreover, for the first time in KP, we demonstrate the idea of improving KP for long documents (i.e., transcripts) by feeding models with paragraph-level keyphrases, i.e., hierarchical extraction. To foster future research, we will publicly release the dataset and code. Amir Pouran Ben Veyseh, Quan Hung Tran, Seunghyun Yoon 0002, Varun Manjunatha, Hanieh Deilamsalehy, Rajiv Jain, Trung Bui, Walter Chang, Franck Dernoncourt, Thien Huu Nguyen |
COLING | 10 |
| 2022 | Learning Cross-Task Dependencies for Joint Extraction of Entities, Events, Event Arguments, and RelationsabstractExtracting entities, events, event arguments, and relations (i.e., task instances) from text represents four main challenging tasks in information extraction (IE), which have been solved jointly (JointIE) to boost the overall performance for IE.As such, previous work often leverages two types of dependencies between the tasks, i.e., cross-instance and cross-type dependencies representing relatedness between task instances and correlations between information types of the tasks.However, the crosstask dependencies in prior work are not optimal as they are only designed manually according to some task heuristics.To address this issue, we propose a novel model for JointIE that aims to learn cross-task dependencies from data.In particular, we treat each task instance as a node in a dependency graph where edges between the instances are inferred through information from different layers of a pretrained language model (e.g., BERT).Furthermore, we utilize the Chow-Liu algorithm to learn a dependency tree between information types for JointIE by seeking to approximate the joint distribution of the types from data.Finally, the Chow-Liu dependency tree is used to generate cross-type patterns, serving as anchor knowledge to guide the learning of representations and dependencies between instances for JointIE.Experimental results show that our proposed model significantly outperforms strong JointIE baselines over four datasets with different languages. Minh Nguyen 0007, Bonan Min, Franck Dernoncourt, Thien Huu Nguyen |
EMNLP | 4 |
| 2022 | MEE: A Novel Multilingual Event Extraction DatasetabstractEvent Extraction (EE) is one of the fundamental tasks in Information Extraction (IE) that aims to recognize event mentions and their arguments (i.e., participants) from text.Due to its importance, extensive methods and resources have been developed for Event Extraction.However, one limitation of current research for EE involves the under-exploration for non-English languages in which the lack of high-quality multilingual EE datasets for model training and evaluation has been the main hindrance.To address this limitation, we propose a novel Multilingual Event Extraction dataset (MEE) that provides annotation for more than 50K event mentions in 8 typologically different languages.MEE comprehensively annotates data for entity mentions, event triggers and event arguments.We conduct extensive experiments on the proposed dataset to reveal challenges and opportunities for multilingual EE. Amir Pouran Ben Veyseh, Javid Ebrahimi, Franck Dernoncourt, Thien Huu Nguyen |
EMNLP | 4 |
| 2022 | BehanceCC: A ChitChat Detection Dataset For Livestreaming Video TranscriptsabstractLivestreaming videos have become an effective broadcasting method for both video sharing and educational purposes. However, livestreaming videos contain a considerable amount of off-topic content (i.e., up to 50%) which introduces significant noises and data load to downstream applications. This paper presents BehanceCC, a new human-annotated benchmark dataset for off-topic detection (also called chitchat detection) in livestreaming video transcripts. In addition to describing the challenges of the dataset, our extensive experiments of various baselines reveal the complexity of chitchat detection for livestreaming videos and suggest potential future research directions for this task. The dataset will be made publicly available to foster research in this area. Viet Dac Lai, Amir Pouran Ben Veyseh, Franck Dernoncourt, Thien Huu Nguyen |
LREC | 4 |
| 2022 | BehanceQA: A New Dataset for Identifying Question-Answer Pairs in Video TranscriptsabstractQuestion-Answer (QA) is one of the effective methods for storing knowledge which can be used for future retrieval. As such, identifying mentions of questions and their answers in text is necessary for a knowledge construction and retrieval systems. In the literature, QA identification has been well studied in the NLP community. However, most of the prior works are restricted to formal written documents such as papers or websites. As such, Questions and Answers that are presented in informal/noisy documents have not been adequately studied. One of the domains that can significantly benefit from QA identification is the domain of livestreaming video transcripts that involve abundant QA pairs to provide valuable knowledge for future users and services. Since video transcripts are often transcribed automatically for scale, they are prone to errors. Combined with the informal nature of discussion in a video, prior QA identification systems might not be able to perform well in this domain. To enable comprehensive research in this domain, we present a large-scale QA identification dataset annotated by human over transcripts of 500 hours of streamed videos. We employ Behance.net to collect the videos and their automatically obtained transcripts. Furthermore, we conduct extensive analysis on the annotated dataset to understand the complexity of QA identification for livestreaming video transcripts. Our experiments show that the annotated dataset presents unique challenges for existing methods and more research is necessary to explore more effective methods. The dataset and the models developed in this work will be publicly released for future research. Amir Pouran Ben Veyseh, Viet Dac Lai, Franck Dernoncourt, Thien Huu Nguyen |
LREC | 4 |
| 2022 | Joint Extraction of Entities, Relations, and Events via Modeling Inter-Instance and Inter-Label DependenciesabstractMinh Van Nguyen, Bonan Min, Franck Dernoncourt, Thien Nguyen. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Minh Nguyen 0007, Bonan Min, Franck Dernoncourt, Thien Huu Nguyen |
NAACL-HLT | 4 |
| 2022 | MINION: a Large-Scale and Diverse Dataset for Multilingual Event DetectionabstractAmir Pouran Ben Veyseh, Minh Van Nguyen, Franck Dernoncourt, Thien Nguyen. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Amir Pouran Ben Veyseh, Minh Nguyen 0007, Franck Dernoncourt, Thien Huu Nguyen |
NAACL-HLT | 4 |
| 2021 | Exploiting Document Structures and Cluster Consistencies for Event Coreference ResolutionabstractHieu Minh Tran, Duy Phung, Thien Huu Nguyen. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Hieu Minh Tran, Duy Phung, Thien Huu Nguyen |
ACL/IJCNLP (1) | 3 |
| 2021 | Unleash GPT-2 Power for Event DetectionabstractAmir Pouran Ben Veyseh, Viet Lai, Franck Dernoncourt, Thien Huu Nguyen. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Amir Pouran Ben Veyseh, Viet Dac Lai, Franck Dernoncourt, Thien Huu Nguyen |
ACL/IJCNLP (1) | 4 |
| 2021 | Textual Data Augmentation for Patient Outcomes PredictionabstractDeep learning models have demonstrated superior performance in various healthcare applications. However, the major limitation of these deep models is usually the lack of high-quality training data due to the private and sensitive nature of this field. In this study, we propose a novel textual data augmentation method to generate artificial clinical notes in patients’ Electronic Health Records (EHRs) that can be used as additional training data for patient outcomes prediction. Essentially, we fine-tune the generative language model GPT-2 to synthesize labeled text with the original training data. More specifically, We propose a teacher-student framework where we first pre-train a teacher model on the 0riginal data, and then train a student model on the GPT-augmented data under the guidance of the teacher. We evaluate our method on the most common patient outcome, i.e., the 30-day readmission rate. The experimental results show that deep models can improve their predictive performance with the augmented data, indicating the effectiveness of the proposed architecture. Qiuhao Lu, Dejing Dou, Thien Huu Nguyen |
BIBM | 3 |
| 2021 | Dictionary-Guided Scene Text RecognitionabstractLanguage prior plays an important role in the way humans detect and recognize text in the wild. Current scene text recognition methods do use lexicons to improve recognition performance, but their naive approach of casting the output into a dictionary word based purely on the edit distance has many limitations. In this paper, we present a novel approach to incorporate a dictionary in both the training and inference stage of a scene text recognition system. We use the dictionary to generate a list of possible outcomes and find the one that is most compatible with the visual appearance of the text. The proposed method leads to a robust scene text recognition model, which is better at handling ambiguous cases encountered in the wild, and improves the overall performance of state-of-the-art scene text spotting frameworks. Our work suggests that incorporating language prior is a potential approach to advance scene text detection and recognition methods. Besides, we contribute VinText, a challenging scene text dataset for Vietnamese, where some characters are equivocal in the visual form due to accent symbols. This dataset will serve as a challenging benchmark for measuring the applicability and robustness of scene text detection and recognition algorithms. Code and dataset are available at https://github.com/VinAIResearch/dict-guided. Thu Nguyen 0003, Vinh Tran 0005, Minh-Triet Tran, Thanh Duc Ngo, Thien Huu Nguyen, Minh Hoai |
CVPR | 6 |
| 2021 | Fine-Grained Event Trigger DetectionabstractMost of the previous work on Event Detection (ED) has only considered the datasets with a small number of event types (i.e., up to 38 types).In this work, we present the first study on fine-grained ED (FED) where the evaluation dataset involves much more fine-grained event types (i.e., 449 types).We propose a novel method to transform the Semcor dataset for Word Sense Disambiguation into a large and high-quality dataset for FED.Extensive evaluation of the current ED methods is conducted to demonstrate the challenges of the generated datasets for FED, calling for more research effort in this area. Thien Huu Nguyen |
EACL | 2 |
| 2021 | Learning Prototype Representations Across Few-Shot Tasks for Event DetectionabstractWe address the sampling bias and outlier issues in few-shot learning for event detection, a subtask of information extraction.We propose to model the relations between training tasks in episodic few-shot learning by introducing cross-task prototypes.We further propose to enforce prediction consistency among classifiers across tasks to make the model more robust to outliers.Our extensive experiment shows a consistent improvement on three fewshot learning datasets.The findings suggest that our model is more robust when labeled data of novel event types is limited. Viet Dac Lai, Franck Dernoncourt, Thien Huu Nguyen |
EMNLP (1) | 3 |
| 2021 | Crosslingual Transfer Learning for Relation and Event Extraction via Word Category and Class AlignmentsabstractPrevious work on crosslingual Relation and Event Extraction (REE) suffers from the monolingual bias issue due to the training of models on only the source language data.An approach to overcome this issue is to use unlabeled data in the target language to aid the alignment of crosslingual representations, i.e., via fooling a language discriminator.However, as this approach does not condition on class information, a target language example of a class could be incorrectly aligned to a source language example of a different class.To address this issue, we propose a novel crosslingual alignment method that leverages class information of REE tasks for representation learning.In particular, we propose to learn two versions of representation vectors for each class in an REE task based on either source or target language examples.Representation vectors for corresponding classes will then be aligned to achieve class-aware alignment for crosslingual representations.In addition, we propose to further align representation vectors for languageuniversal word categories (i.e., parts of speech and dependency relations).As such, a novel filtering mechanism is presented to facilitate the learning of word category representations from contextualized representations on input texts based on adversarial learning.We conduct extensive crosslingual experiments with English, Chinese, and Arabic over REE tasks.The results demonstrate the benefits of the proposed method that significantly advances the state-of-the-art performance in these settings. Minh Nguyen 0007, Tuan Ngo Nguyen, Bonan Min, Thien Huu Nguyen |
EMNLP (1) | 4 |
| 2021 | Modeling Document-Level Context for Event Detection via Important Context SelectionabstractEvent Detection (ED) aims to recognize and classify trigger words of events in text.The recent progress has featured advanced transformer-based language models (e.g., BERT) as a critical component in stateof-the-art models for ED.However, the length limit for input texts is a barrier for such ED models as they cannot encode long-range document-level context that has been shown to be beneficial for ED.To address this issue, we propose a novel method to model documentlevel context with BERT for ED that dynamically selects relevant sentences in the document for the event prediction of the target sentence.The target sentence will be then augmented with the selected sentences and consumed entirely by BERT for improved representation learning for ED.To this end, the RE-INFORCE algorithm is employed to train the relevant sentence selection for ED.Several information types are then introduced to form the reward function for the training process, including ED performance, sentence similarity, and discourse relations.Our extensive experiments on multiple benchmark datasets reveal the effectiveness of the proposed model, leading to new state-of-the-art performance. Amir Pouran Ben Veyseh, Minh Nguyen 0007, Nghia Trung Ngo, Bonan Min, Thien Huu Nguyen |
EMNLP (1) | 5 |
| 2021 | Cross-Task Instance Representation Interactions and Label Dependencies for Joint Information Extraction with Graph Convolutional NetworksabstractExisting works on information extraction (IE)have mainly solved the four main tasks separately (entity mention recognition, relation extraction, event trigger detection, and argument extraction), thus failing to benefit from inter-dependencies between tasks.This paper presents a novel deep learning model to simultaneously solve the four tasks of IE in a single model (called FourIE).Compared to few prior work on jointly performing four IE tasks, FourIE features two novel contributions to capture inter-dependencies between tasks.First, at the representation level, we introduce an interaction graph between instances of the four tasks that is used to enrich the prediction representation for one instance with those from related instances of other tasks.Second, at the label level, we propose a dependency graph for the information types in the four IE tasks that captures the connections between the types expressed in an input sentence.A new regularization mechanism is introduced to enforce the consistency between the golden and predicted type dependency graphs to improve representation learning.We show that the proposed model achieves the state-of-the-art performance for joint IE on both monolingual and multilingual learning settings with three different languages. Minh Nguyen 0007, Viet Dac Lai, Thien Huu Nguyen |
NAACL-HLT | 3 |
| 2021 | Graph Convolutional Networks for Event Causality Identification with Rich Document-level StructuresabstractWe study the problem of Event Causality Identification (ECI) to detect causal relation between event mention pairs in text.Although deep learning models have recently shown state-of-the-art performance for ECI, they are limited to the intra-sentence setting where event mention pairs are presented in the same sentences.This work addresses this issue by developing a novel deep learning model for document-level ECI (DECI) to accept intersentence event mention pairs.As such, we propose a graph-based model that constructs interaction graphs to capture relevant connections between important objects for DECI in input documents.Such interaction graphs are then consumed by graph convolutional networks to learn document context-augmented representations for causality prediction between events.Various information sources are introduced to enrich the interaction graphs for DECI, featuring discourse, syntax, and semantic information.Our extensive experiments show that the proposed model achieves state-of-the-art performance on two benchmark datasets. Minh Tran Phu, Thien Huu Nguyen |
NAACL-HLT | 2 |
| 2021 | Inducing Rich Interaction Structures Between Words for Document-Level Event Argument Extraction
Amir Pouran Ben Veyseh, Franck Dernoncourt, Quan Hung Tran, Varun Manjunatha, Rajiv Jain, Doo Soon Kim, Walter Chang, Thien Huu Nguyen |
PAKDD (2) | 9 |
| 2021 | Augmenting Open-Domain Event Detection with Synthetic Data from GPT-2
Amir Pouran Ben Veyseh, Minh Nguyen 0007, Bonan Min, Thien Huu Nguyen |
ECML/PKDD (3) | 4 |
| 2021 | Graph Learning Regularization and Transfer Learning for Few-Shot Event DetectionabstractWe address the poor generalization of few-shot learning models for event detection (ED) using transfer learning and representation regularization. In particular, we propose to transfer knowledge from open-domain word sense disambiguation into few-shot learning models for ED to improve their generalization to new event types. We also propose a novel training signal derived from dependency graphs to regularize the representation learning for ED. Moreover, we evaluate few-shot learning models for ED with a large-scale human-annotated ED dataset to obtain more reliable insights for this problem. Our comprehensive experiments demonstrate that the proposed model outperforms state-of-the-art baseline models in the few-shot learning and supervised learning settings for ED. Code and data splits are available at https://github.com/laiviet/ed-fsl. Viet Dac Lai, Minh Nguyen 0007, Thien Huu Nguyen, Franck Dernoncourt |
SIGIR | 3 |
| 2021 | Predicting Patient Readmission Risk from Medical Text via Knowledge Graph Enhanced Multiview Graph ConvolutionabstractUnplanned intensive care unit (ICU) readmission rate is an important metric for evaluating the quality of hospital care. Efficient and accurate prediction of ICU readmission risk can not only help prevent patients from inappropriate discharge and potential dangers, but also reduce associated costs of healthcare. In this paper, we propose a new method that uses medical text of Electronic Health Records (EHRs) for prediction, which provides an alternative perspective to previous studies that heavily depend on numerical and time-series features of patients. More specifically, we extract discharge summaries of patients from their EHRs, and represent them with multiview graphs enhanced by an external knowledge graph. Graph convolutional networks are then used for representation learning. Experimental results prove the effectiveness of our method, yielding state-of-the-art performance for this task. Qiuhao Lu, Thien Huu Nguyen, Dejing Dou |
SIGIR | 2 |
| 2020 | A Joint Model for Definition Extraction with Syntactic Connection and Semantic ConsistencyabstractDefinition Extraction (DE) is one of the well-known topics in Information Extraction that aims to identify terms and their corresponding definitions in unstructured texts. This task can be formalized either as a sentence classification task (i.e., containing term-definition pairs or not) or a sequential labeling task (i.e., identifying the boundaries of the terms and definitions). The previous works for DE have only focused on one of the two approaches, failing to model the inter-dependencies between the two tasks. In this work, we propose a novel model for DE that simultaneously performs the two tasks in a single framework to benefit from their inter-dependencies. Our model features deep learning architectures to exploit the global structures of the input sentences as well as the semantic consistencies between the terms and the definitions, thereby improving the quality of the representation vectors for DE. Besides the joint inference between sentence classification and sequential labeling, the proposed model is fundamentally different from the prior work for DE in that the prior work has only employed the local structures of the input sentences (i.e., word-to-word relations), and not yet considered the semantic consistencies between terms and definitions. In order to implement these novel ideas, our model presents a multi-task learning framework that employs graph convolutional neural networks and predicts the dependency paths between the terms and the definitions. We also seek to enforce the consistency between the representations of the terms and definitions both globally (i.e., increasing semantic consistency between the representations of the entire sentences and the terms/definitions) and locally (i.e., promoting the similarity between the representations of the terms and the definitions). The extensive experiments on three benchmark datasets demonstrate the effectiveness of our approach.1 Amir Pouran Ben Veyseh, Franck Dernoncourt, Dejing Dou, Thien Huu Nguyen |
AAAI | 4 |
| 2020 | Multi-View Consistency for Relation Extraction via Mutual Information and Structure PredictionabstractRelation Extraction (RE) is one of the fundamental tasks in Information Extraction. The goal of this task is to find the semantic relations between entity mentions in text. It has been shown in many previous work that the structure of the sentences (i.e., dependency trees) can provide important information/features for the RE models. However, the common limitation of the previous work on RE is the reliance on some external parsers to obtain the syntactic trees for the sentence structures. On the one hand, it is not guaranteed that the independent external parsers can offer the optimal sentence structures for RE and the customized structures for RE might help to further improve the performance. On the other hand, the quality of the external parsers might suffer when applied to different domains, thus also affecting the performance of the RE models on such domains. In order to overcome this issue, we introduce a novel method for RE that simultaneously induces the structures and predicts the relations for the input sentences, thus avoiding the external parsers and potentially leading to better sentence structures for RE. Our general strategy to learn the RE-specific structures is to apply two different methods to infer the structures for the input sentences (i.e., two views). We then introduce several mechanisms to encourage the structure and semantic consistencies between these two views so the effective structure and semantic representations for RE can emerge. We perform extensive experiments on the ACE 2005 and SemEval 2010 datasets to demonstrate the advantages of the proposed method, leading to the state-of-the-art performance on such datasets. Amir Pouran Ben Veyseh, Franck Dernoncourt, My T. Thai, Dejing Dou, Thien Huu Nguyen |
AAAI | 5 |
| 2020 | Exploiting the Syntax-Model Consistency for Neural Relation ExtractionabstractThis paper studies the task of Relation Extraction (RE) that aims to identify the semantic relations between two entity mentions in text.In the deep learning models for RE, it has been beneficial to incorporate the syntactic structures from the dependency trees of the input sentences.In such models, the dependency trees are often used to directly structure the network architectures or to obtain the dependency relations between the word pairs to inject the syntactic information into the models via multi-task learning.The major problems with these approaches are the lack of generalization beyond the syntactic structures in the training data or the failure to capture the syntactic importance of the words for RE.In order to overcome these issues, we propose a novel deep learning model for RE that uses the dependency trees to extract the syntax-based importance scores for the words, serving as a tree representation to introduce syntactic information into the models with greater generalization.In particular, we leverage Ordered-Neuron Long-Short Term Memory Networks (ON-LSTM) to infer the model-based importance scores for RE for every word in the sentences that are then regulated to be consistent with the syntax-based scores to enable syntactic information injection.We perform extensive experiments to demonstrate the effectiveness of the proposed method, leading to the state-of-the-art performance on three RE benchmark datasets. Amir Pouran Ben Veyseh, Franck Dernoncourt, Dejing Dou, Thien Huu Nguyen |
ACL | 4 |
| 2020 | Exploiting Node Content for Multiview Graph Convolutional Network and Adversarial RegularizationabstractNetwork representation learning (NRL) is crucial in the area of graph learning.Recently, graph autoencoders and its variants have gained much attention and popularity among various types of node embedding approaches.Most existing graph autoencoder-based methods aim to minimize the reconstruction errors of the input network while not explicitly considering the semantic relatedness between nodes.In this paper, we propose a novel network embedding method which models the consistency across different views of networks.More specifically, we create a second view from the input network which captures the relation between nodes based on node content and enforce the latent representations from the two views to be consistent by incorporating a multiview adversarial regularization module.The experimental studies on benchmark datasets prove the effectiveness of this method, and demonstrate that our method compares favorably with the state-of-the-art algorithms on challenging tasks such as link prediction and node clustering.We also evaluate our method on a real-world application, i.e., 30-day unplanned ICU readmission prediction, and achieve promising results compared with several baseline methods. Qiuhao Lu, Nisansa de Silva, Dejing Dou, Thien Huu Nguyen, Prithviraj Sen, Berthold Reinwald, Yunyao Li 0001 |
COLING | 4 |
| 2020 | What Does This Acronym Mean? Introducing a New Dataset for Acronym Identification and DisambiguationabstractAcronyms are the short forms of phrases that facilitate conveying lengthy sentences in documents and serve as one of the mainstays of writing.Due to their importance, identifying acronyms and corresponding phrases (i.e., acronym identification (AI)) and finding the correct meaning of each acronym (i.e., acronym disambiguation (AD)) are crucial for text understanding.Despite the recent progress on this task, there are some limitations in the existing datasets which hinder further improvement.More specifically, limited size of manually annotated AI datasets or noises in the automatically created acronym identification datasets obstruct designing advanced highperforming acronym identification models.Moreover, the existing datasets are mostly limited to the medical domain and ignore other domains.In order to address these two limitations, we first create a manually annotated large AI dataset for scientific domain.This dataset contains 17,506 sentences which is substantially larger than previous scientific AI datasets.Next, we prepare an AD dataset for scientific domain with 62,441 samples which is significantly larger than previous scientific AD dataset.Our experiments show that the existing state-of-the-art models fall far behind human-level performance on both datasets proposed by this work.In addition, we propose a new deep learning model which utilizes the syntactical structure of the sentence to expand an ambiguous acronym in a sentence.The proposed model outperforms the state-of-the-art models on the new AD dataset, providing a strong baseline for future research on this dataset 1 . Amir Pouran Ben Veyseh, Franck Dernoncourt, Quan Hung Tran, Thien Huu Nguyen |
COLING | 4 |
| 2020 | Event Detection: Gate Diversity and Syntactic Importance Scores for Graph Convolution Neural NetworksabstractRecent studies on event detection (ED) have shown that the syntactic dependency graph can be employed in graph convolution neural networks (GCN) to achieve state-of-the-art performance.However, the computation of the hidden vectors in such graph-based models is agnostic to the trigger candidate words, potentially leaving irrelevant information for the trigger candidate for event prediction.In addition, the current models for ED fail to exploit the overall contextual importance scores of the words, which can be obtained via the dependency tree, to boost the performance.In this study, we propose a novel gating mechanism to filter noisy information in the hidden vectors of the GCN models for ED based on the information from the trigger candidate.We also introduce novel mechanisms to achieve the contextual diversity for the gates and the importance score consistency for the graphs and models in ED.The experiments show that the proposed model achieves state-of-the-art performance on two ED datasets. Viet Dac Lai, Tuan Ngo Nguyen, Thien Huu Nguyen |
EMNLP (1) | 3 |
| 2020 | Introducing a New Dataset for Event Detection in Cybersecurity TextsabstractDetecting cybersecurity events is necessary to keep us informed about the fast growing number of such events reported in text.In this work, we focus on the task of event detection (ED) to identify event trigger words for the cybersecurity domain.In particular, to facilitate the future research, we introduce a new dataset for this problem, characterizing the manual annotation for 30 important cybersecurity event types and a large dataset size to develop deep learning models.Comparing to the prior datasets for this task, our dataset involves more event types and supports the modeling of document-level information to improve the performance.We perform extensive evaluation with the current state-of-the-art methods for ED on the proposed dataset.Our experiments reveal the challenges of cybersecurity ED and present many research opportunities in this area for the future work. Hieu Man Duc Trong, Duc-Trong Le, Amir Pouran Ben Veyseh, Thuat Nguyen, Thien Huu Nguyen |
EMNLP (1) | 5 |
| 2020 | Introducing Syntactic Structures into Target Opinion Word Extraction with Deep LearningabstractTargeted opinion word extraction (TOWE) is a sub-task of aspect based sentiment analysis (ABSA) which aims to find the opinion words for a given aspect-term in a sentence.Despite their success for TOWE, the current deep learning models fail to exploit the syntactic information of the sentences that have been proved to be useful for TOWE in the prior research.In this work, we propose to incorporate the syntactic structures of the sentences into the deep learning models for TOWE, leveraging the syntax-based opinion possibility scores and the syntactic connections between the words.We also introduce a novel regularization technique to improve the performance of the deep learning models based on the representation distinctions between the words in TOWE.The proposed model is extensively analyzed and achieves the state-of-the-art performance on four benchmark datasets. Amir Pouran Ben Veyseh, Nasim Nouri, Franck Dernoncourt, Dejing Dou, Thien Huu Nguyen |
EMNLP (1) | 5 |
| 2020 | Exploiting the Matching Information in the Support Set for Few Shot Event Classification
Viet Dac Lai, Franck Dernoncourt, Thien Huu Nguyen |
PAKDD (2) | 3 |
| 2020 | Learning to Select Important Context Words for Event Detection
Nghia Trung Ngo, Tuan Ngo Nguyen, Thien Huu Nguyen |
PAKDD (2) | 3 |
| 2020 | Bag of biterms modeling for short texts
Anh Phan Tuan, Tran Xuan Bach, Thien Huu Nguyen, Ngo Van Linh 0001, Khoat Than |
Knowl. Inf. Syst. | 3 |
| 2019 | One for All: Neural Joint Modeling of Entities and EventsabstractThe previous work for event extraction has mainly focused on the predictions for event triggers and argument roles, treating entity mentions as being provided by human annotators. This is unrealistic as entity mentions are usually predicted by some existing toolkits whose errors might be propagated to the event trigger and argument role recognition. Few of the recent work has addressed this problem by jointly predicting entity mentions, event triggers and arguments. However, such work is limited to using discrete engineering features to represent contextual information for the individual tasks and their interactions. In this work, we propose a novel model to jointly perform predictions for entity mentions, event triggers and arguments based on the shared hidden representations from deep learning. The experiments demonstrate the benefits of the proposed method, leading to the state-of-the-art performance for event extraction. Trung Minh Nguyen, Thien Huu Nguyen |
AAAI | 2 |
| 2019 | Employing the Correspondence of Relations and Connectives to Identify Implicit Discourse Relations via Label EmbeddingsabstractIt has been shown that implicit connectives can be exploited to improve the performance of the models for implicit discourse relation recognition (IDRR).An important property of the implicit connectives is that they can be accurately mapped into the discourse relations conveying their functions.In this work, we explore this property in a multi-task learning framework for IDRR in which the relations and the connectives are simultaneously predicted, and the mapping is leveraged to transfer knowledge between the two prediction tasks via the embeddings of relations and connectives.We propose several techniques to enable such knowledge transfer that yield the state-of-the-art performance for IDRR on several settings of the benchmark dataset (i.e., the Penn Discourse Treebank dataset). Linh The Nguyen, Ngo Van Linh 0001, Khoat Than, Thien Huu Nguyen |
ACL (1) | 4 |
| 2019 | Graph based Neural Networks for Event Factuality Prediction using Syntactic and Semantic StructuresabstractEvent factuality prediction (EFP) is the task of assessing the degree to which an event mentioned in a sentence has happened.For this task, both syntactic and semantic information are crucial to identify the important context words.The previous work for EFP has only combined these information in a simple way that cannot fully exploit their coordination.In this work, we introduce a novel graph-based neural network for EFP that can integrate the semantic and syntactic information more effectively.Our experiments demonstrate the advantage of the proposed model for EFP. Amir Pouran Ben Veyseh, Thien Huu Nguyen, Dejing Dou |
ACL (1) | 2 |
| 2019 | Rumor detection in social networks via deep contextual modelingabstractFake news and rumors constitute a major problem in social networks recently. Due to the fast information propagation in social networks, it is inefficient to use human labor to detect suspicious news. Automatic rumor detection is thus necessary to prevent devastating effects of rumors on the individuals and society. Previous work has shown that in addition to the content of the news/posts and their contexts (i.e., replies), the relations or connections among those components are important to boost the rumor detection performance. In order to induce such relations between posts and contexts, the prior work has mainly relied on the inherent structures of the social networks (e.g., direct replies), ignoring the potential semantic connections between those objects. In this work, we demonstrate that such semantic relations are also helpful as they can reveal the implicit structures to better capture the patterns in the contexts for rumor detection. We propose to employ the self-attention mechanism in neural text modeling to achieve the semantic structure induction for this problem. In addition, we introduce a novel method to preserve the important information of the main news/posts in the final representations of the entire threads to further improve the performance for rumor detection. Our method matches the main post representations and the thread representations by ensuring that they predict the same latent labels in a multitask learning framework. The extensive experiments demonstrate the effectiveness of the proposed model for rumor detection, yielding the state-of-the-art performance on recent datasets for this problem. Amir Pouran Ben Veyseh, My T. Thai, Thien Huu Nguyen, Dejing Dou |
ASONAM | 3 |
| 2019 | Systematic Generalization: What Is Required and Can It Be Learned?
Dzmitry Bahdanau, Shikhar Murty, Michael Noukhovitch, Thien Huu Nguyen, Harm de Vries, Aaron C. Courville |
ICLR (Poster) | 4 |
| 2019 | BabyAI: A Platform to Study the Sample Efficiency of Grounded Language Learning
Maxime Chevalier-Boisvert, Dzmitry Bahdanau, Salem Lahlou, Lucas Willems, Chitwan Saharia, Thien Huu Nguyen, Yoshua Bengio |
ICLR (Poster) | 6 |
| 2019 | Improving Cross-Domain Performance for Relation Extraction via Dependency Prediction and Information Flow ControlabstractRelation Extraction (RE) is one of the fundamental tasks in Information Extraction and Natural Language Processing. Dependency trees have been shown to be a very useful source of information for this task. The current deep learning models for relation extraction has mainly exploited this dependency information by guiding their computation along the structures of the dependency trees. One potential problem with this approach is it might prevent the models from capturing important context information beyond syntactic structures and cause the poor cross-domain generalization. This paper introduces a novel method to use dependency trees in RE for deep learning models that jointly predicts dependency and semantics relations. We also propose a new mechanism to control the information flow in the model based on the input entity mentions. Our extensive experiments on benchmark datasets show that the proposed model outperforms the existing methods for RE significantly. Amir Pouran Ben Veyseh, Thien Huu Nguyen, Dejing Dou |
IJCAI | 2 |
| 2018 | Graph Convolutional Networks With Argument-Aware Pooling for Event DetectionabstractThe current neural network models for event detection have only considered the sequential representation of sentences. Syntactic representations have not been explored in this area although they provide an effective mechanism to directly link words to their informative context for event detection in the sentences. In this work, we investigate a convolutional neural network based on dependency trees to perform event detection. We propose a novel pooling method that relies on entity mentions to aggregate the convolution vectors. The extensive experiments demonstrate the benefits of the dependency-based convolutional neural networks and the entity mention-based pooling method for event detection. We achieve the state-of-the-art performance on widely used datasets with both perfect and predicted entity mentions. Thien Huu Nguyen, Ralph Grishman |
AAAI | 1 |
| 2018 | Who is Killed by Police: Introducing Supervised Attention for Hierarchical LSTMsabstractFinding names of people killed by police has become increasingly important as police shootings get more and more public attention (police killing detection). Unfortunately, there has been not much work in the literature addressing this problem. The early work in this field (Keith etal., 2017) proposed a distant supervision framework based on Expectation Maximization (EM) to deal with the multiple appearances of the names in documents. However, such EM-based framework cannot take full advantages of deep learning models, necessitating the use of handdesigned features to improve the detection performance. In this work, we present a novel deep learning method to solve the problem of police killing recognition. The proposed method relies on hierarchical LSTMs to model the multiple sentences that contain the person names of interests, and introduce supervised attention mechanisms based on semantical word lists and dependency trees to upweight the important contextual words. Our experiments demonstrate the benefits of the proposed model and yield the state-of-the-art performance for police killing detection. Minh Nguyen 0004, Thien Huu Nguyen |
COLING | 2 |
| 2018 | Similar but not the Same - Word Sense Disambiguation Improves Event Detection via Neural Representation MatchingabstractEvent detection (ED) and word sense disambiguation (WSD) are two similar tasks in that they both involve identifying the classes (i.e.event types or word senses) of some word in a given sentence.It is thus possible to extract the knowledge hidden in the data for WSD, and utilize it to improve the performance on ED.In this work, we propose a method to transfer the knowledge learned on WSD to ED by matching the neural representations learned for the two tasks.Our experiments on two widely used datasets for ED demonstrate the effectiveness of the proposed method. Weiyi Lu, Thien Huu Nguyen |
EMNLP | 2 |
| 2016 | Joint Learning of Local and Global Features for Entity Linking via Neural NetworksabstractPrevious studies have highlighted the necessity for entity linking systems to capture the local entity-mention similarities and the global topical coherence. We introduce a novel framework based on convolutional neural networks and recurrent neural networks to simultaneously model the local and global features for entity linking. The proposed model benefits from the capacity of convolutional neural networks to induce the underlying representations for local contexts and the advantage of recurrent neural networks to adaptively compress variable length sequences of predictions for global constraints. Our evaluation on multiple datasets demonstrates the effectiveness of the model and yields the state-of-the-art performance on such datasets. In addition, we examine the entity linking systems on the domain adaptation setting that further demonstrates the cross-domain robustness of the proposed model. Thien Huu Nguyen, Nicolas R. Fauceglia, Mariano Rodriguez-Muro, Oktie Hassanzadeh, Alfio Massimiliano Gliozzo, Mohammad Sadoghi |
COLING | 1 |
| 2016 | Modeling Skip-Grams for Event Detection with Convolutional Neural NetworksabstractConvolutional neural networks (CNN) have achieved the top performance for event detection due to their capacity to induce the underlying structures of the k-grams in the sentences.However, the current CNN-based event detectors only model the consecutive k-grams and ignore the non-consecutive kgrams that might involve important structures for event detection.In this work, we propose to improve the current CNN models for ED by introducing the non-consecutive convolution.Our systematic evaluation on both the general setting and the domain adaptation setting demonstrates the effectiveness of the nonconsecutive CNN model, leading to the significant performance improvement over the current state-of-the-art systems. Thien Huu Nguyen, Ralph Grishman |
EMNLP | 1 |
| 2016 | Joint Event Extraction via Recurrent Neural NetworksabstractEvent extraction is a particularly challenging problem in information extraction.The stateof-the-art models for this problem have either applied convolutional neural networks in a pipelined framework (Chen et al., 2015) or followed the joint architecture via structured prediction with rich local and global features (Li et al., 2013).The former is able to learn hidden feature representations automatically from data based on the continuous and generalized representations of words.The latter, on the other hand, is capable of mitigating the error propagation problem of the pipelined approach and exploiting the inter-dependencies between event triggers and argument roles via discrete structures.In this work, we propose to do event extraction in a joint framework with bidirectional recurrent neural networks, thereby benefiting from the advantages of the two models as well as addressing issues inherent in the existing approaches.We systematically investigate different memory features for the joint model and demonstrate that the proposed model achieves the state-of-the-art performance on the ACE 2005 dataset. Thien Huu Nguyen, Kyunghyun Cho, Ralph Grishman |
HLT-NAACL | 1 |
| 2015 | Semantic Representations for Domain Adaptation: A Case Study on the Tree Kernel-based Method for Relation ExtractionabstractThien Huu Nguyen, Barbara Plank, Ralph Grishman. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Thien Huu Nguyen, Barbara Plank, Ralph Grishman |
ACL (1) | 1 |
| 2011 | Combining Proper Name-Coreference with Conditional Random Fields for Semi-supervised Named Entity Recognition in Vietnamese Text
Rathany Chan Sam, Huong Thanh Le, Thuy Thanh Nguyen, Thien Huu Nguyen |
PAKDD (1) | 4 |