Ngo Van Linh 0001

dblp:125/3578 · also Linh Ngo Van 0001, Linh Van Ngo 0001 · DBLP profile ↗
← Back
57ranked-venue papers
4as first author
47since 2021 · last 2026
0000-0002-0011-5137ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 49 · 1 first-author · 44 since 2021Databases, data management, data science and information retrieval · 9 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 7 · 7 since 2021Applied, interdisciplinary, general and emerging computing · 1
YearPublicationVenuePosition
2026 CTPD: Cross Tokenizer Preference Distillation
abstract
While knowledge distillation has seen widespread use in pre-training and instruction tuning, its application to aligning language models with human preferences remains underexplored, particularly in the more realistic cross-tokenizer setting. The incompatibility of tokenization schemes between teacher and student models has largely prevented fine-grained, white-box distillation of preference information. To address this gap, we propose Cross-Tokenizer Preference Distillation (CTPD), the first unified framework for transferring human-aligned behavior between models with heterogeneous tokenizers. CTPD introduces three key innovations: (1) Aligned Span Projection, which maps teacher and student tokens to shared character-level spans for precise supervision transfer; (2) a cross-tokenizer adaptation of Token-level Importance Sampling (TIS-DPO) for improved credit assignment; and (3) a Teacher-Anchored Reference, allowing the student to directly leverage the teacher’s preferences in a DPO-style objective. Our theoretical analysis grounds CTPD in importance sampling, and experiments across multiple benchmarks confirm its effectiveness, with significant performance gains over existing methods. These results establish CTPD as a practical and general solution for preference distillation across diverse tokenization schemes, opening the door to more accessible and efficient alignment of language models.
Phi Van Dat, Ngan Nguyen, Ngo Van Linh 0001, Trung Le 0001, Thanh Hong Nguyen
AAAI4
2026 GloCTM: Cross-Lingual Topic Modeling via a Global Context Space
abstract
Cross-lingual topic modeling seeks to uncover coherent and semantically aligned topics across languages—a task central to multilingual understanding. Yet most existing models learn topics in disjoint, language-specific spaces and rely on alignment mechanisms (e.g., bilingual dictionaries) that often fail to capture deep cross-lingual semantics, resulting in loosely connected topic spaces. Moreover, these approaches often overlook the rich semantic signals embedded in multilingual pretrained representations, further limiting their ability to capture fine-grained alignment. We introduce **GloCTM** (**Glo**bal Context Space for **C**ross-Lingual **T**opic **M**odel), a novel framework that enforces cross-lingual topic alignment through a unified semantic space spanning the entire model pipeline. GloCTM constructs enriched input representations by expanding bag-of-words with cross-lingual lexical neighborhoods, and infers topic proportions using both local and global encoders, with their latent representations aligned through internal regularization. At the output level, the global topic-word distribution, defined over the combined vocabulary, structurally synchronizes topic meanings across languages. To further ground topics in deep semantic space, GloCTM incorporates a Centered Kernel Alignment (CKA) loss that aligns the latent topic space with multilingual contextual embeddings. Experiments across multiple benchmarks demonstrate that GloCTM significantly improves topic coherence and cross-lingual alignment, outperforming strong baselines.
Nguyen Tien Phat, Ngo Vu Minh, Ngo Van Linh 0001, Nguyen Thi Ngoc Diep, Thien Huu Nguyen
AAAI3
2026 MCW-KD: Multi-Cost Wasserstein Knowledge Distillation for Large Language Models
Hoang Tran Vuong, Tue Le, Quyen Tran, Ngo Van Linh 0001, Trung Le 0001
AAAI4
2026 MTA: Multi-Granular Trajectory Alignment for Large Language Model Distillation
abstract
Pham Khanh Chi, Quoc Phong Dao, Thuat Nguyen, Linh Ngo Van, Trung Le, Thanh Hong Nguyen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Pham Khanh Chi, Quoc Phong Dao, Thuat Nguyen, Ngo Van Linh 0001, Trung Le 0001, Thanh Hong Nguyen
ACL (1)4
2026 SRA: Span Representation Alignment for Large Language Model Distillation
abstract
Quoc Phong Dao, Hoang Son Nguyen, Pham Khanh Chi, Tung Nguyen, Linh Ngo Van, Nguyen Thi Ngoc Diep, Trung Le. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Quoc Phong Dao, Hoang Son Nguyen, Pham Khanh Chi, Ngo Van Linh 0001, Nguyen Thi Ngoc Diep, Trung Le 0001
ACL (1)5
2026 TALAS: Teacher-Anchored Layer Alignment with Adaptive Sharpness-Aware Minimization for Embedding Distillation
abstract
Quoc Phong Dao, Hoang Son Nguyen, Pham Khanh Chi, Linh Ngo Van, Nguyen Thi Ngoc Diep, Thien Huu Nguyen, Trung Le. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Quoc Phong Dao, Hoang Son Nguyen, Pham Khanh Chi, Ngo Van Linh 0001, Nguyen Thi Ngoc Diep, Thien Huu Nguyen, Trung Le 0001
ACL (1)4
2026 Octopus: Gated Selective Attention for Memory-Bounded Long-Context Inference in Large Language Models
abstract
Chien Van Nguyen, Ryan A. Rossi, Linh Ngo Van, Franck Dernoncourt, Thien Huu Nguyen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Chien Van Nguyen, Ryan Rossi, Ngo Van Linh 0001, Franck Dernoncourt, Thien Huu Nguyen
ACL (1)3
2026 LLM-XTM: Enhancing Cross-Lingual Topic Models with Large Language Models
abstract
Minh Chu Xuan, Tien-Phat Nguyen, Linh Ngo Van, Dinh Viet Sang, Nguyen Thi Ngoc Diep, Trung Le. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Minh Chu Xuan, Tien-Phat Nguyen, Ngo Van Linh 0001, Nguyen Thi Ngoc Diep, Trung Le 0001
ACL (1)3
2026 Global and local context in short text neural topic model
Ngo Van Linh 0001, Anh Nguyen Duc 0002
Artif. Intell.2
2026 WAVE++: Capturing within-task variance for continual relation extraction with adaptive prompting
Bao-Ngoc Dao, Minh Le, Luyen Ngo Dinh, Nam Le 0005, Ngo Van Linh 0001
Neurocomputing6
2026 A multi-objective surrogate framework for training neural topic models
Chi Pham Khanh, Tue Le, Ngo Van Linh 0001
Neurocomputing4
2026 A framework for neural topic modeling using hierarchical clustering and contrastive learning with optimal transport
Chi Pham Khanh, Hoang Tran Vuong, Ngo Van Linh 0001
Neurocomputing4
2026 A framework for cross-lingual topic modeling using global context spaces and unified decoding with semantic knowledge distillation
Tien-Phat Nguyen, Giang Pham Thi Phuong, Thuan Do Phan, Ngo Van Linh 0001
Neurocomputing5
2026 AQS-IDETR: Adaptive Query Selection for Efficient Inference in Real-Time Detection Transformers
Nhat Minh Nguyen Quoc, Hong Dang Nguyen, Ngo Van Linh 0001
Image Vis. Comput.3
2026 Discord: Enhancing knowledge distillation via cross-chain of thought and optimal transport alignment between models with different tokenizers
Anh Duc Le, Tu Vu, Ngo Van Linh 0001
Knowl. Based Syst.4
2026 MoL: Mixture of Layers in Cross-Tokenizer embedding model distillation
Hai An Vu, Minh-Phuc Truong, Tu Vu, Ngo Van Linh 0001
Knowl. Based Syst.4
2026 SAMD: Span-Aware Matryoshka Distillation for Cross-Tokenizer Embedding Models
Thang Duc Tran, Minh-Phuc Truong, Thuan Do Phan, Ngo Van Linh 0001
Mach. Learn.4
2025 Adaptive Prompting for Continual Relation Extraction: A Within-Task Variance Perspective
abstract
To address catastrophic forgetting in Continual Relation Extraction (CRE), many current approaches rely on memory buffers to rehearse previously learned knowledge while acquiring new tasks. Recently, prompt-based methods have emerged as potent alternatives to rehearsal-based strategies, demonstrating strong empirical performance. However, upon analyzing existing prompt-based approaches for CRE, we identified several critical limitations, such as inaccurate prompt selection, inadequate mechanisms for mitigating forgetting in shared parameters, and suboptimal handling of cross-task and within-task variances. To overcome these challenges, we draw inspiration from the relationship between prefix tuning and mixture of experts, proposing a novel approach that employs a prompt pool for each task, capturing variations within each task while enhancing cross-task variances. Furthermore, we incorporate a generative model to consolidate prior knowledge within shared parameters, eliminating the need for explicit data storage. Extensive experiments validate the efficacy of our approach, demonstrating superior performance over state-of-the-art prompt-based and rehearsal-free methods in continual relation extraction.
Minh Le, Tien Ngoc Luu, An Nguyen The, Thanh-Thien Le, Tung Thanh Nguyen, Ngo Van Linh 0001, Thien Huu Nguyen
AAAI7
2025 Few-Shot, No Problem: Descriptive Continual Relation Extraction
abstract
Few-shot Continual Relation Extraction is a crucial challenge for enabling AI systems to identify and adapt to evolving relationships in dynamic real-world domains. Traditional memory-based approaches often overfit to limited samples, failing to reinforce old knowledge, with the scarcity of data in few-shot scenarios further exacerbating these issues by hindering effective data augmentation in the latent space. In this paper, we propose a novel retrieval-based solution, starting with a large language model to generate descriptions for each relation. From these descriptions, we introduce a bi-encoder retrieval training paradigm to enrich both sample and class representation learning. Leveraging these enhanced representations, we design a retrieval-based prediction method where each sample "retrieves" the best fitting relation via a reciprocal rank fusion score that integrates both relation description vectors and class prototypes. Extensive experiments on multiple datasets demonstrate that our method significantly advances the state-of-the-art by maintaining robust performance across sequential tasks, effectively addressing catastrophic forgetting.
Anh Duc Le, Quyen Tran, Thanh-Thien Le, Ngo Van Linh 0001, Thien Huu Nguyen
AAAI5
2025 Mitigating Non-Representative Prototypes and Representation Bias in Few-Shot Continual Relation Extraction
abstract
To address the phenomenon of similar classes, existing methods in few-shot continual relation extraction (FCRE) face two main challenges: non-representative prototypes and representation bias, especially when the number of available samples is limited. In our work, we propose Minion to address these challenges. Firstly, we leverage the General Orthogonal Frame (GOF) structure, based on the concept of Neural Collapse, to create robust class prototypes with clear separation, even between analogous classes. Secondly, we utilize label description representations as global class representatives within the fast-slow contrastive learning paradigm. These representations consistently encapsulate the essential attributes of each relation, acting as global information that helps mitigate overfitting and reduces representation bias caused by the limited local few-shot examples within a class. Extensive experiments on well-known FCRE benchmarks show that our method outperforms state-of-the-art approaches, demonstrating its effectiveness for advancing RE system.
Thanh Duc Pham, Ngo Van Linh 0001, Nguyen Thi Ngoc Diep, Sang Dinh, Thien Huu Nguyen
ACL (1)3
2025 EMO: Embedding Model Distillation via Intra-Model Relation and Optimal Transport Alignments
abstract
Knowledge distillation (KD) is crucial for compressing large text embedding models, but faces challenges when teacher and student models use different tokenizers (Cross-Tokenizer KD -CTKD).Vocabulary mismatches impede the transfer of relational knowledge encoded in deep representations, such as hidden states and attention matrices, which are vital for producing high-quality embeddings.Existing CTKD methods often focus on direct output alignment, neglecting this crucial structural information.We propose a novel framework tailored for CTKD embedding model distillation.We first map tokens one-to-one via Minimum Edit Distance (MinED).Then, we distill intra-model relational knowledge by aligning attention matrix patterns using Centered Kernel Alignment, focusing on the top-m most important tokens of the directly mapped tokens.Simultaneously, we align final hidden states via Optimal Transport with Importance-Scored Mass Assignment, which emphasizes semantically important token representations, based on importance scores derived from attention weights.We evaluate distillation from state-of-the-art embedding models (e.g., LLM2Vec, BGE) to a Bert-base-uncased model on embedding-reliant tasks such as text classification, sentence pair classification, and semantic textual similarity.Our proposed framework significantly outperforms existing CTKD baselines.By preserving attention structure and prioritizing key representations, our approach yields smaller, highfidelity embedding models despite tokenizer differences.
Minh-Phuc Truong, Hai An Vu, Tu Vu, Nguyen Thi Ngoc Diep, Ngo Van Linh 0001, Thien Huu Nguyen, Trung Le 0001
EMNLP5
2025 Mutual-pairing Data Augmentation for Fewshot Continual Relation Extraction
abstract
Nguyen Hoang Anh, Quyen Tran, Thanh Xuan Nguyen, Nguyen Thi Ngoc Diep, Linh Ngo Van, Thien Huu Nguyen, Trung Le. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Nguyen Hoang Anh, Quyen Tran, Nguyen Thi Ngoc Diep, Ngo Van Linh 0001, Thien Huu Nguyen, Trung Le 0001
NAACL (Long Papers)5
2025 Enhancing Discriminative Representation in Similar Relation Clusters for Few-Shot Continual Relation Extraction
abstract
Anh Duc Le, Nam Le Hai, Thanh Xuan Nguyen, Linh Ngo Van, Nguyen Thi Ngoc Diep, Sang Dinh, Thien Huu Nguyen. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Anh Duc Le, Ngo Van Linh 0001, Nguyen Thi Ngoc Diep, Sang Dinh, Thien Huu Nguyen
NAACL (Long Papers)4
2025 Sharpness-Aware Minimization for Topic Models with High-Quality Document Representations
abstract
Tung Nguyen, Tue Le, Hoang Tran Vuong, Quang Duc Nguyen, Duc Anh Nguyen, Linh Ngo Van, Sang Dinh, Thien Huu Nguyen. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Tue Le, Hoang Tran Vuong, Ngo Van Linh 0001, Sang Dinh, Thien Huu Nguyen
NAACL (Long Papers)6
2025 GloCOM: A Short Text Neural Topic Model via Global Clustering Context
abstract
Quang Duc Nguyen, Tung Nguyen, Duc Anh Nguyen, Linh Ngo Van, Sang Dinh, Thien Huu Nguyen. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Ngo Van Linh 0001, Sang Dinh, Thien Huu Nguyen
NAACL (Long Papers)4
2025 Token-Level Self-Play with Importance-Aware Guidance for Large Language Models
abstract
Leveraging the power of Large Language Models (LLMs) through preference optimization is crucial for aligning model outputs with human values. Direct Preference Optimization (DPO) has recently emerged as a simple yet effective method by directly optimizing on preference data without the need for explicit reward models. However, DPO typically relies on human-labeled preference data, which can limit its scalability. Self-Play Fine-Tuning (SPIN) addresses this by allowing models to generate their own rejected samples, reducing the dependence on human annotations. Nevertheless, SPIN uniformly applies learning signals across all tokens, ignoring the fine-grained quality variations within responses. As the model improves, rejected samples increasingly contain high-quality tokens, making the uniform treatment of tokens suboptimal. In this paper, we propose SWIFT (Self-Play Weighted Fine-Tuning), a fine-grained self-refinement method that assigns token-level importance weights estimated from a stronger teacher model. Beyond alignment, we also demonstrate that SWIFT serves as an effective knowledge distillation strategy by using the teacher not for logits matching, but for reward-guided token weighting. Extensive experiments on diverse benchmarks and settings demonstrate that SWIFT consistently surpasses both existing alignment approaches and conventional knowledge distillation methods.
Tue Le, Hoang Tran Vuong, Quyen Tran, Ngo Van Linh 0001, Mehrtash Harandi, Trung Le 0001
NeurIPS4
2025 A Framework for Neural Topic Modeling with Mutual Information and Group Regularization
Ngo Van Linh 0001
Neurocomputing2
2025 Out-of-vocabulary handling and topic quality control strategies in streaming topic models
Ngo Van Linh 0001, Ha Bang Ban, Khoat Than
Neurocomputing3
2025 TopiCOT: Neural topic model aligning with pre-trained clustering and optimal transport
Duy-Tung Pham, Ngo Van Linh 0001
Neurocomputing4
2025 Sub-document neural topic model
Duy-Tung Doan, Huu-Thanh Luong, Ngo Van Linh 0001
Knowl. Based Syst.4
2024 Continual Relation Extraction via Sequential Multi-Task Learning
abstract
To build continual relation extraction (CRE) models, those can adapt to an ever-growing ontology of relations, is a cornerstone information extraction task that serves in various dynamic real-world domains. To mitigate catastrophic forgetting in CRE, existing state-of-the-art approaches have effectively utilized rehearsal techniques from continual learning and achieved remarkable success. However, managing multiple objectives associated with memory-based rehearsal remains underexplored, often relying on simple summation and overlooking complex trade-offs. In this paper, we propose Continual Relation Extraction via Sequential Multi-task Learning (CREST), a novel CRE approach built upon a tailored Multi-task Learning framework for continual learning. CREST takes into consideration the disparity in the magnitudes of gradient signals of different objectives, thereby effectively handling the inherent difference between multi-task learning and continual learning. Through extensive experiments on multiple datasets, CREST demonstrates significant improvements in CRE performance as well as superiority over other state-of-the-art Multi-task Learning frameworks, offering a promising solution to the challenges of continual learning in this domain.
Thanh-Thien Le, Tung Thanh Nguyen, Ngo Van Linh 0001, Thien Huu Nguyen
AAAI4
2024 Hierarchical Selection of Important Context for Generative Event Causality Identification with Optimal Transports
abstract
We study the problem of Event Causality Identification (ECI) that seeks to predict causal relation between event mentions in the text. In contrast to previous classification-based models, a few recent ECI methods have explored generative models to deliver state-of-the-art performance. However, such generative models cannot handle document-level ECI where long context between event mentions must be encoded to secure correct predictions. In addition, previous generative ECI methods tend to rely on external toolkits or human annotation to obtain necessary training signals. To address these limitations, we propose a novel generative framework that leverages Optimal Transport (OT) to automatically select the most important sentences and words from full documents. Specifically, we introduce hierarchical OT alignments between event pairs and the document to extract pertinent contexts. The selected sentences and words are provided as input and output to a T5 encoder-decoder model which is trained to generate both the causal relation label and salient contexts. This allows richer supervision without external tools. We conduct extensive evaluations on different datasets with multiple languages to demonstrate the benefits and state-of-the-art performance of ECI.
Hieu Man, Chien Van Nguyen, Nghia Trung Ngo, Ngo Van Linh 0001, Franck Dernoncourt, Thien Huu Nguyen
LREC/COLING4
2024 Lifelong Event Detection via Optimal Transport
abstract
Continual Event Detection (CED) poses a formidable challenge due to the catastrophic forgetting phenomenon, where learning new tasks (with new coming event types) hampers performance on previous ones.In this paper, we introduce a novel approach, Lifelong Event Detection via Optimal Transport (LEDOT), that leverages optimal transport principles to align the optimization of our classification module with the intrinsic nature of each class, as defined by their pre-trained language modeling.Our method integrates replay sets, prototype latent representations, and an innovative Optimal Transport component.Extensive experiments on MAVEN and ACE datasets demonstrate LEDOT's superior performance, consistently outperforming state-of-the-art baselines.The results underscore LEDOT as a pioneering solution in continual event detection, offering a more effective and nuanced approach to addressing catastrophic forgetting in evolving environments.
Viet Dao, Van-Cuong Pham, Quyen Tran, Thanh-Thien Le, Ngo Van Linh 0001, Thien Huu Nguyen
EMNLP5
2024 Preserving Generalization of Language models in Few-shot Continual Relation Extraction
abstract
Few-shot Continual Relations Extraction (FCRE) is an emerging and dynamic area of study where models can sequentially integrate knowledge from new relations with limited labeled data while circumventing catastrophic forgetting and preserving prior knowledge from pre-trained backbones.In this work, we introduce a novel method that leverages oftendiscarded language model heads.By employing these components via a mutual information maximization strategy, our approach helps maintain prior knowledge from the pre-trained backbone and strategically aligns the primary classification head, thereby enhancing model performance.Furthermore, we explore the potential of Large Language Models (LLMs), renowned for their wealth of knowledge, in addressing FCRE challenges.Our comprehensive experimental results underscore the efficacy of the proposed method and offer valuable insights for future work.
Quyen Tran, Nguyen Hoang Anh, Trung Le 0001, Ngo Van Linh 0001, Thien Huu Nguyen
EMNLP6
2024 SharpSeq: Empowering Continual Event Detection through Sharpness-Aware Sequential-task Learning
abstract
Thanh-Thien Le, Viet Dao, Linh Nguyen, Thi-Nhung Nguyen, Linh Ngo, Thien Nguyen. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Thanh-Thien Le, Viet Dao, Thi-Nhung Nguyen, Ngo Van Linh 0001, Thien Huu Nguyen
NAACL-HLT5
2024 Mixture of Experts Meets Prompt-Based Continual Learning
abstract
Exploiting the power of pre-trained models, prompt-based approaches stand out compared to other continual learning solutions in effectively preventing catastrophic forgetting, even with very few learnable parameters and without the need for a memory buffer. While existing prompt-based continual learning methods excel in leveraging prompts for state-of-the-art performance, they often lack a theoretical explanation for the effectiveness of prompting. This paper conducts a theoretical analysis to unravel how prompts bestow such advantages in continual learning, thus offering a new perspective on prompt design. We first show that the attention block of pre-trained models like Vision Transformers inherently encodes a special mixture of experts architecture, characterized by linear experts and quadratic gating score functions. This realization drives us to provide a novel view on prefix tuning, reframing it as the addition of new task-specific experts, thereby inspiring the design of a novel gating mechanism termed Non-linear Residual Gates (NoRGa). Through the incorporation of non-linear activation and residual connection, NoRGa enhances continual learning performance while preserving parameter efficiency. The effectiveness of NoRGa is substantiated both theoretically and empirically across diverse benchmarks and pretraining paradigms. Our code is publicly available at https://github.com/Minhchuyentoancbn/MoE_PromptCL.
Minh Le, An Nguyen The, Trang Pham, Ngo Van Linh 0001, Nhat Ho
NeurIPS6
2024 Continual variational dropout: a view of auxiliary local variables in continual learning
Ngo Van Linh 0001, Thien Huu Nguyen, Khoat Than
Mach. Learn.3
2023 Dynamic Transformation of Prior Knowledge Into Bayesian Models for Data Streams
abstract
We consider how to effectively use prior knowledge when learning a Bayesian model from streaming environments where the data come endlessly and sequentially. This problem is highly important in the era of data explosion and rich sources of valuable external knowledge such as pre-trained models, ontologies, Wikipedia, etc. We show that some existing approaches can forget any knowledge very fast. We then propose a novel framework that enables to incorporate the prior knowledge of different forms into a base Bayesian model for data streams. Our framework subsumes some existing popular models for time-series/dynamic data. Extensive experiments show that our framework outperforms existing methods with a large margin. In particular, our framework can help Bayesian models generalize well on extremely short text while other methods overfit. An implementation of our framework is available athttp://github.com/bachtranxuan/TPS.
Tran Xuan Bach, Nguyen Duc Anh, Ngo Van Linh 0001, Khoat Than
IEEE Trans. Knowl. Data Eng.3
2022 Selecting Optimal Context Sentences for Event-Event Relation Extraction
abstract
Understanding events entails recognizing the structural and temporal orders between event mentions to build event structures/graphs for input documents. To achieve this goal, our work addresses the problems of subevent relation extraction (SRE) and temporal event relation extraction (TRE) that aim to predict subevent and temporal relations between two given event mentions/triggers in texts. Recent state-of-the-art methods for such problems have employed transformer-based language models (e.g., BERT) to induce effective contextual representations for input event mention pairs. However, a major limitation of existing transformer-based models for SRE and TRE is that they can only encode input texts of limited length (i.e., up to 512 sub-tokens in BERT), thus unable to effectively capture important context sentences that are farther away in the documents. In this work, we introduce a novel method to better model document-level context with important context sentences for event-event relation extraction. Our method seeks to identify the most important context sentences for a given entity mention pair in a document and pack them into shorter documents to be consume entirely by transformer-based language models for representation learning. The REINFORCE algorithm is employed to train models where novel reward functions are presented to capture model performance, and context-based and knowledge-based similarity between sentences for our problem. Extensive experiments demonstrate the effectiveness of the proposed method with state-of-the-art performance on benchmark datasets.
Hieu Man, Nghia Trung Ngo, Ngo Van Linh 0001, Thien Huu Nguyen
AAAI3
2022 Unsupervised Domain Adaptation for Text Classification via Meta Self-Paced Learning
abstract
A shift in data distribution can have a significant impact on performance of a text classification model. Recent methods addressing unsupervised domain adaptation for textual tasks typically extracted domain-invariant representations through balancing between multiple objectives to align feature spaces between source and target domains. While effective, these methods induce various new domain-sensitive hyperparameters, thus are impractical as large-scale language models are drastically growing bigger to achieve optimal performance. To this end, we propose to leverage meta-learning framework to train a neural network-based self-paced learning procedure in an end-to-end manner. Our method, called Meta Self-Paced Domain Adaption (MSP-DA), follows a novel but intuitive domain-shift variation of cluster assumption to derive the meta train-test dataset split based on the self-pacing difficulties of source domain’s examples. As a result, MSP-DA effectively leverages self-training and self-tuning domain-specific hyperparameters simultaneously throughout the learning process. Extensive experiments demonstrate our framework substantially improves performance on target domains, surpassing state-of-the-art approaches. Detailed analyses validate our method and provide insight into how each domain affects the learned hyperparameters.
Nghia Trung Ngo, Ngo Van Linh 0001, Thien Huu Nguyen
COLING2
2022 Auxiliary Local Variables for Improving Regularization/Prior Approach in Continual Learning
Ngo Van Linh 0001, Khoat Than
PAKDD (1)1
2022 Reducing Catastrophic Forgetting in Neural Networks via Gaussian Mixture Approximation
Hoang Phan, Anh Phan Tuan, Ngo Van Linh 0001, Khoat Than
PAKDD (1)4
2022 A graph convolutional topic model for short and noisy text streams
Ngo Van Linh 0001, Tran Xuan Bach, Khoat Than
Neurocomputing1
2022 Balancing stability and plasticity when learning topic models from short and noisy text streams
Trung Mai, Ngo Van Linh 0001, Khoat Than
Neurocomputing4
2022 From implicit to explicit feedback: A deep neural network for modeling sequential behaviours and long-short term preferences of online users
Quyen Tran, Lam Tran, Linh Chu Hai, Ngo Van Linh 0001, Khoat Than
Neurocomputing4
2022 Adaptive infinite dropout for noisy and sparse data streams
Ngo Van Linh 0001, Khoat Than
Mach. Learn.4
2021 Boosting prior knowledge in streaming variational Bayes
Ngo Van Linh 0001, Nguyen Kim Anh, Canh Hao Nguyen, Khoat Than
Neurocomputing2
2020 Bag of biterms modeling for short texts
Anh Phan Tuan, Tran Xuan Bach, Thien Huu Nguyen, Ngo Van Linh 0001, Khoat Than
Knowl. Inf. Syst.4
2019 Employing the Correspondence of Relations and Connectives to Identify Implicit Discourse Relations via Label Embeddings
abstract
It has been shown that implicit connectives can be exploited to improve the performance of the models for implicit discourse relation recognition (IDRR).An important property of the implicit connectives is that they can be accurately mapped into the discourse relations conveying their functions.In this work, we explore this property in a multi-task learning framework for IDRR in which the relations and the connectives are simultaneously predicted, and the mapping is leveraged to transfer knowledge between the two prediction tasks via the embeddings of relations and connectives.We propose several techniques to enable such knowledge transfer that yield the state-of-the-art performance for IDRR on several settings of the benchmark dataset (i.e., the Penn Discourse Treebank dataset).
Linh The Nguyen, Ngo Van Linh 0001, Khoat Than, Thien Huu Nguyen
ACL (1)2
2019 From Implicit to Explicit Feedback: A deep neural network for modeling the sequential behavior of online users
abstract
We demonstrate the advantages of taking into account multiple types of behavior in recommendation systems. Intuitively, each user has to do some \textbf{implicit} actions (e.g., click) before making an \textbf{explicit} decision (e.g., purchase). Previous works showed that implicit and explicit feedback has distinct properties to make a useful recommendation. However, these works exploit implicit and explicit behavior separately and therefore ignore the semantic of interaction between users and items. In this paper, we propose a novel model namely \textit{Implicit to Explicit (ITE)} which directly models the order of user actions. Furthermore, we present an extended version of ITE, namely \textit{Implicit to Explicit with Side information (ITE-Si)}, which incorporates side information to enrich the representations of users and items. The experimental results show that both ITE and ITE-Si outperform existing recommendation systems and also demonstrate the effectiveness of side information in two large scale datasets.
Anh Phan Tuan, Nhat Nguyen Trong, Duong Bui Trong, Ngo Van Linh 0001, Khoat Than
ACML4
2019 Infinite Dropout for training Bayesian models from data streams
abstract
The ability to continuously train Bayesian models in streaming environments is highly important in the era of big data. However, it has to face the famous stability-plasticity dilemma and the problem of noisy and sparse data. We propose a novel and easy-to-implement framework, called Infinite Dropout (iDropout), to address these challenges. iDropout has an easy mechanism to balance between old and new information, which allows models to trade off stability against plasticity. Thanks to the ability to reduce overfitting and the ensemble property of Dropout, our framework obtains better generalization, thus effectively handles undesirable effects of noise and sparsity. Further, iDropout is able to adapt quickly to abnormal changes in data streams. We theoretically analyze the equivalence of Dropout in iDropout to a regularizer, well applied to a much larger context than what was known before. Extensive experiments show that iDropout significantly outperforms the state-of-the-art baselines.
Van-Son Nguyen, Duc-Tung Nguyen, Ngo Van Linh 0001, Khoat Than
IEEE BigData3
2019 Eliminating overfitting of probabilistic topic models on short and noisy text: The role of dropout
Cuong Ha, Van-Dang Tran, Ngo Van Linh 0001, Khoat Than
Int. J. Approx. Reason.3
2018 Collaborative Topic Model for Poisson distributed ratings
Hoa M. Le, Son Ta Cong, Quyen Pham The, Ngo Van Linh 0001, Khoat Than
Int. J. Approx. Reason.4
2017 Keeping Priors in Streaming Bayesian Learning
Ngo Van Linh 0001, Nguyen Kim Anh, Khoat Than
PAKDD (2)2
2017 An effective and interpretable method for document classification
Ngo Van Linh 0001, Nguyen Kim Anh, Khoat Than, Chien Nguyen Dang
Knowl. Inf. Syst.1
2016 Enabling Hierarchical Dirichlet Processes to Work Better for Short Texts at Large Scale
Khai Mai, Sang Mai, Ngo Van Linh 0001, Khoat Than
PAKDD (2)4
2015 Effective and Interpretable Document Classification Using Distinctly Labeled Dirichlet Process Mixture Models of von Mises-Fisher Distributions
Ngo Van Linh 0001, Nguyen Kim Anh, Khoat Than, Nguyen Nguyen Tat
DASFAA (2)1