Chien Van Nguyen

dblp:351/5534 · DBLP profile ↗
← Back
4ranked-venue papers
2as first author
4since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Language models and text generation · 65% Efficient and distributed learning · 22% Deep learning architectures and training · 13%

Topics — the 5 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › language modeling
long-context language modeling
2.022026
Octopus: Gated Selective Attention for Memory-Bounded Long-Context Inference in Large Language Models · ACL (1) 2026
Lizard: An Efficient Linearization Framework for Large Language Models · ACL (1) 2026
Natural language and speech › Language models and text generation › text generation › surface realization
linearization
1.012026
Lizard: An Efficient Linearization Framework for Large Language Models · ACL (1) 2026
Machine learning › Efficient and distributed learning › inference efficiency
memory-efficient inference
1.012026
Octopus: Gated Selective Attention for Memory-Bounded Long-Context Inference in Large Language Models · ACL (1) 2026
Machine learning › Deep learning architectures and training
attention mechanism
0.312026
Octopus: Gated Selective Attention for Memory-Bounded Long-Context Inference in Large Language Models · ACL (1) 2026
Machine learning › Deep learning architectures and training › attention mechanism
selective attention
0.312026
Octopus: Gated Selective Attention for Memory-Bounded Long-Context Inference in Large Language Models · ACL (1) 2026

Methods — techniques the papers use, named apart from their topics

linear attention · 1.0gated attention · 1.0
YearPublicationVenuePosition
2026 Lizard: An Efficient Linearization Framework for Large Language Models
abstract
Chien Van Nguyen, Huy Huu Nguyen, Ruiyi Zhang, Hanieh Deilamsalehy, Puneet Mathur, Viet Dac Lai, Haoliang Wang, Jayakumar Subramanian, Ryan A. Rossi, Trung Bui, Nikos Vlassis, Franck Dernoncourt, Thien Huu Nguyen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Chien Van Nguyen, Huy Huu Nguyen, Ruiyi Zhang 0002, Hanieh Deilamsalehy, Puneet Mathur, Viet Dac Lai, Jayakumar Subramanian, Ryan Rossi, Trung Bui, Nikos Vlassis, Franck Dernoncourt, Thien Huu Nguyen
ACL (1)1
2026 Octopus: Gated Selective Attention for Memory-Bounded Long-Context Inference in Large Language Models
abstract
Chien Van Nguyen, Ryan A. Rossi, Linh Ngo Van, Franck Dernoncourt, Thien Huu Nguyen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Chien Van Nguyen, Ryan Rossi, Ngo Van Linh 0001, Franck Dernoncourt, Thien Huu Nguyen
ACL (1)1
2024 Hierarchical Selection of Important Context for Generative Event Causality Identification with Optimal Transports
abstract
We study the problem of Event Causality Identification (ECI) that seeks to predict causal relation between event mentions in the text. In contrast to previous classification-based models, a few recent ECI methods have explored generative models to deliver state-of-the-art performance. However, such generative models cannot handle document-level ECI where long context between event mentions must be encoded to secure correct predictions. In addition, previous generative ECI methods tend to rely on external toolkits or human annotation to obtain necessary training signals. To address these limitations, we propose a novel generative framework that leverages Optimal Transport (OT) to automatically select the most important sentences and words from full documents. Specifically, we introduce hierarchical OT alignments between event pairs and the document to extract pertinent contexts. The selected sentences and words are provided as input and output to a T5 encoder-decoder model which is trained to generate both the causal relation label and salient contexts. This allows richer supervision without external tools. We conduct extensive evaluations on different datasets with multiple languages to demonstrate the benefits and state-of-the-art performance of ECI.
Hieu Man, Chien Van Nguyen, Nghia Trung Ngo, Ngo Van Linh 0001, Franck Dernoncourt, Thien Huu Nguyen
LREC/COLING2
2024 CulturaX: A Cleaned, Enormous, and Multilingual Dataset for Large Language Models in 167 Languages
abstract
Extensive training datasets represent one of the important factors for the impressive learning capabilities of large language models (LLMs). However, these training datasets for current LLMs, especially the recent state-of-the-art models, are often not fully disclosed. Creating training data for high-performing LLMs involves extensive cleaning and deduplication to ensure the necessary level of quality. The lack of transparency for training data has thus hampered research on attributing and addressing hallucination and bias issues in LLMs, hindering replication efforts and further advancements in the community. These challenges become even more pronounced in multilingual learning scenarios, where the available multilingual text datasets are often inadequately collected and cleaned. Consequently, there is a lack of open-source and readily usable dataset to effectively train LLMs in multiple languages. To overcome this issue, we present CulturaX, a substantial multilingual dataset with 6.3 trillion tokens in 167 languages, tailored for LLM development. Our dataset undergoes meticulous cleaning and deduplication through a rigorous pipeline of multiple stages to accomplish the best quality for model training, including language identification, URL-based filtering, metric-based cleaning, document refinement, and data deduplication. CulturaX is released in Hugging Face facilitate research and advancements in multilingual LLMs: https://huggingface.co/datasets/uonlp/CulturaX.
Thuat Nguyen, Chien Van Nguyen, Viet Dac Lai, Hieu Man, Nghia Trung Ngo, Franck Dernoncourt, Ryan Rossi, Thien Huu Nguyen
LREC/COLING2