Minh Le Nguyen 0001

dblp:n/MinhLeNguyen · also Le Minh Nguyen 0001, Le-Minh Nguyen 0001, Minh-Le Nguyen 0001, Nguyen Le Minh 0001, Nguyen Minh Le 0001 · DBLP profile ↗
← Back
35ranked-venue papers in the field
4as first author
16since 2021 · last 2027
0000-0002-2265-1010ORCID · conflict

Domains — venue-derived; a paper can count in several

Information Retrieval & Web Search · 21 (3 first)Database Systems & Data Management · 8 (1 first)Data Mining & Knowledge Discovery · 6
YearPublicationVenuePosition
2027 M-RAV: Multimodal retrieve-augment-verify framework for boosting zero-shot fact verification system with large language models
Son T. Luu, Trung Vo, Minh Le Nguyen 0001
Inf. Process. Manag.3
2026 CancerRAGent: Evidence-Linked and Safety-Guided Oncology Question Answering
Trung Vo, An Trieu, Yuji Matsumoto 0001, Minh Le Nguyen 0001
ECIR (4)5
2026 Legal Case Entailment via Optimal Transport-Enhanced Retrieval and Large Language Model Ranking
abstract
Legal case entailment embodies a fundamental principle of the legal system, wherein the verdict of historical cases functions as a guiding precedent for subsequent cases sharing analogous factual circumstances. Due to the intricate nature of legal case documents, identifying entailment between legal cases requires considerable time and effort, necessitating a thorough understanding and specialized expertise in legal interpretation and analysis. To accelerate the process of legal case entailment, in this article, we conceptualize this task as a document retrieval problem and propose a two-stage framework focused on entailment information retrieval. Within this framework, we develop a cost-efficient system that utilizes advanced language models for legal case entailment. In the first stage, we present the established ColBERT document retrieval model, augmented with a sparse keyword alignment strategy utilizing the Unbalanced Optimal Transport framework. Our study illustrates that by focusing on the interaction of contextually and semantically similar keyword pairs between the query and the document, the proposed alignment method improves the retrieval capability of ColBERT in the legal domain. For the second stage, we employ a fine-tuned MonoT5 document ranking model to refine the retrieval results and predict entailment instances. Extensive evaluation demonstrates a significant performance improvement of the proposed system compared to previous methods. As an additional study, we benchmark state-of-the-art open source LLMs in legal case entailment to reveal their performance and potential applications. Our findings indicate that while LLMs exhibit sensitivity to prompt formulation, they demonstrate promising zero-shot performance in legal entailment scenarios. To encourage further AI development in the legal domain, we provide the code necessary to reproduce our results ( https://github.com/thanhtcptit/Legal-Case-Entailment-Framework ).
Phuong Minh Nguyen 0001, Minh Le Nguyen 0001
ACM Trans. Knowl. Discov. Data3
2025 Sentiment Analysis of Travel Reviews in Kyoto Using LLMs
Shehan Liyanaarachchi, Tai Dinh, Wuyi Yue, Trung Vo, Minh Le Nguyen 0001
ADMA (3)5
2025 WIP: Iterative Post-training Pruning with Weighted Importance Estimation for Large Language Models
Dinh-Truong Do, Kiyoaki Shirai, Minh Le Nguyen 0001
NLDB (1)3
2025 Improve Smart Contract Vulnerability Explanation with Synthetic Data and Chain-of-Thought Prompting
Minh Le Nguyen 0001, Naoya Inoue
NLDB (1)1
2025 Pretraining Data Exposure in Large Language Models: A Survey of Membership Inference, Data Contamination, and Security Implications
Ziyi Tong, Minh Le Nguyen 0001
NLDB (2)3
2025 Retrieve-Revise-Refine: A novel framework for retrieval of concise entailing legal article set
Phuong Minh Nguyen 0001, Minh Le Nguyen 0001
Inf. Process. Manag.3
2025 K-Bloom: unleashing the power of pre-trained language models in extracting knowledge graph with predefined relations
Trung Vo, Son T. Luu, Minh Le Nguyen 0001
Knowl. Inf. Syst.3
2023 NegT5: A Cross-Task Text-to-Text Framework for Negation in Question Answering
Teeradaj Racharak, Minh Le Nguyen 0001
ACIIDS (2)3
2023 Dual Congestion-Aware Route Planning for Tourists by Multi-agent Reinforcement Learning
Kong Yuntao, Minh Le Nguyen 0001, Qiang Ma 0001
DEXA (2)3
2023 GRAM: Grammar-Based Refined-Label Representing Mechanism in the Hierarchical Semantic Parsing Task
Dinh-Truong Do, Phuong Minh Nguyen 0001, Minh Le Nguyen 0001
NLDB3
2023 ViWiQA: Efficient end-to-end Vietnamese Wikipedia-based Open-domain Question-Answering systems for single-hop and multi-hop questions
Dieu-Hien Nguyen, Nguyen-Khang Le, Minh Le Nguyen 0001
Inf. Process. Manag.3
2022 Learning to Map the GDPR to Logic Representation on DAPRECO-KB
Phuong Minh Nguyen 0001, Thi-Thu-Trang Nguyen, Vu D. Tran, Ha-Thanh Nguyen, Minh Le Nguyen 0001, Ken Satoh
ACIIDS (1)5
2022 Diversity-Oriented Route Planning for Tourists
Wei Kun Kong, Shuyuan Zheng, Minh Le Nguyen 0001, Qiang Ma 0001
DEXA (2)3
2022 Extractive Elementary Discourse Units for Improving Abstractive Summarization
abstract
Abstractive summarization focuses on generating concise and fluent text from an original document while maintaining the original intent and containing the new words that do not appear in the original document. Recent studies point out that rewriting extractive summaries help improve the performance with a more concise and comprehensible output summary, which uses a sentence as a textual unit. However, a single document sentence normally cannot supply sufficient information. In this paper, we apply elementary discourse unit (EDU) as textual unit of content selection. In order to utilize EDU for generating a high quality summary, we propose a novel summarization model that first designs an EDU selector to choose salient content. Then, the generator model rewrites the selected EDUs as the final summary. To determine the relevancy of each EDU on the entire document, we choose to apply group tag embedding, which can establish the connection between summary sentences and relevant EDUs, so that our generator does not only focus on selected EDUs, but also ingest the entire original document. Extensive experiments on the CNN/Daily Mail dataset have demonstrated the effectiveness of our model.
Ye Xiong, Teeradaj Racharak, Minh Le Nguyen 0001
SIGIR3
2019 A Study on Self-attention Mechanism for AMR-to-text Generation
Vu Trong Sinh, Minh Le Nguyen 0001
NLDB2
2019 Web document summarization by exploiting social context with matrix co-factorization
Minh-Tien Nguyen, Tran Viet Cuong, Nguyen Xuan Hoai, Minh Le Nguyen 0001
Inf. Process. Manag.4
2019 Sentence modeling via multiple word embeddings and multi-level comparison for semantic textual similarity
Minh Le Nguyen 0001, Yamasaki Tomohiro, Izuha Tatsuya
Inf. Process. Manag.2
2018 Preface KSE 2016
Ashwin Ittoo, Minh Le Nguyen 0001, Satoshi Tojo
Data Knowl. Eng.2
2018 Automatically classifying source code using tree-based approaches
Anh Viet Phan, Ngoc Phuong Chau, Minh Le Nguyen 0001, Lam Thu Bui
Data Knowl. Eng.3
2018 Multilingual opinion mining on YouTube - A convolutional N-gram BiLSTM word embedding
Minh Le Nguyen 0001
Inf. Process. Manag.2
2018 Exploiting User Posts for Web Document Summarization
abstract
Relevant user posts such as comments or tweets of a Web document provide additional valuable information to enrich the content of this document. When creating user posts, readers tend to borrow salient words or phrases in sentences. This can be considered as word variation. This article proposes a framework that models the word variation aspect to enhance the quality of Web document summarization. Technically, the framework consists of two steps: scoring and selection. In the first step, the social information of a Web document such as user posts is exploited to model intra-relations and inter-relations in lexical and semantic levels. These relations are denoted by a mutual reinforcement similarity graph used to score each sentence and user post. After scoring, summaries are extracted by using a ranking approach or concept-based method formulated in the form of Integer Linear Programming. To confirm the efficiency of our framework, sentence and story highlight extraction tasks were taken as a case study on three datasets in two languages, English and Vietnamese. Experimental results show that: (i) the framework can improve ROUGE-scores compared to state-of-the-art baselines of social context summarization and (ii) the combination of the two relations benefits the sentence extraction of single Web documents.
Minh-Tien Nguyen, Vu D. Tran, Minh Le Nguyen 0001, Xuan-Hieu Phan
ACM Trans. Knowl. Discov. Data3
2017 Summarizing Web Documents Using Sequence Labeling with User-Generated Content and Third-Party Sources
Minh-Tien Nguyen, Vu D. Tran, Chien-Xuan Tran, Minh Le Nguyen 0001
NLDB4
2016 SoLSCSum: A Linked Sentence-Comment Dataset for Social Context Summarization
abstract
This paper presents a dataset named SoLSCSum for social context summarization. The dataset includes 157 open-domain articles along with their comments collected from Yahoo News. The articles and their comments were manually annotated by two annotators to extract standard summaries. The inter-annotator agreement is 74.5% and Cohen's Kappa is 0.5845. To illustrate the potential use of our dataset, a learning to rank model was trained by using a set of local and cross features. Experimental results demonstrate that: (1) our model trained by Ranking SVM obtains significant improvements from 5.5% to 14.8% of ROUGE-1 over state-of-the-art baselines in document summarization and (2) our dataset can be used to train summary methods such as SVM.
Minh-Tien Nguyen, Chien-Xuan Tran, Vu D. Tran, Minh Le Nguyen 0001
CIKM4
2016 SoRTESum: A Social Context Framework for Single-Document Summarization
Minh-Tien Nguyen, Minh Le Nguyen 0001
ECIR2
2014 From Treebank Conversion to Automatic Dependency Parsing for Vietnamese
Dat Quoc Nguyen, Dai Quoc Nguyen, Son Bao Pham, Phuong-Thai Nguyen, Minh Le Nguyen 0001
NLDB5
2014 A semi supervised learning model for mapping sentences to logical forms with ambiguous supervision
Minh Le Nguyen 0001, Akira Shimazu
Data Knowl. Eng.1
2013 EDU-Based Similarity for Paraphrase Identification
Ngo Xuan Bach, Minh Le Nguyen 0001, Akira Shimazu
NLDB2
2012 A Semi Supervised Learning Model for Mapping Sentences to Logical form with Ambiguous Supervision
Minh Le Nguyen 0001, Akira Shimazu
NLDB1
2011 Improving Subtree-Based Question Classification Classifiers with Word-Cluster Models
Minh Le Nguyen 0001, Akira Shimazu
NLDB1
2011 A Hidden Topic-Based Framework toward Building Applications with Short Web Documents
abstract
This paper introduces a hidden topic-based framework for processing short and sparse documents (e.g., search result snippets, product descriptions, book/movie summaries, and advertising messages) on the Web. The framework focuses on solving two main challenges posed by these kinds of documents: 1) data sparseness and 2) synonyms/homonyms. The former leads to the lack of shared words and contexts among documents while the latter are big linguistic obstacles in natural language processing (NLP) and information retrieval (IR). The underlying idea of the framework is that common hidden topics discovered from large external data sets (universal data sets), when included, can make short documents less sparse and more topic-oriented. Furthermore, hidden topics from universal data sets help handle unseen data better. The proposed framework can also be applied for different natural languages and data domains. We carefully evaluated the framework by carrying out two experiments for two important online applications (Web search result classification and matching/ranking for contextual advertising) with large-scale universal data sets and we achieved significant results.
Xuan-Hieu Phan, Cam-Tu Nguyen, Dieu-Thu Le, Minh Le Nguyen 0001, Susumu Horiguchi, Quang-Thuy Ha
IEEE Trans. Knowl. Data Eng.4
2008 Learning to classify short and sparse text & web with hidden topics from large-scale data collections
abstract
This paper presents a general framework for building classifiers that deal with short and sparse text & segments by making the most of hidden topics discovered from large-scale data collections. The main motivation of this work is that many classification tasks working with short segments of text & Web, such as search snippets, forum & chat messages, blog & news feeds, product reviews, and book & movie summaries, fail to achieve high accuracy due to the data sparseness. We, therefore, come up with an idea of gaining external knowledge to make the data more related as well as expand the coverage of classifiers to handle future data better. The underlying idea of the framework is that for each classification task, we collect a large-scale external data collection called universal dataset, and then build a classifier on both a (small) set of labeled training data and a rich set of hidden topics discovered from that data collection. The framework is general enough to be applied to different data domains and genres ranging from search results to medical text. We did a careful evaluation on several hundred megabytes of Wikipedia (30M words) and MEDLINE (18M words) with two tasks: Web search domain disambiguation and disease categorization for medical text, and achieved significant quality enhancement.
Xuan-Hieu Phan, Minh Le Nguyen 0001, Susumu Horiguchi
WWW2
2005 Classification with Maximum Entropy Modeling of Predictive Association Rules
Xuan-Hieu Phan, Minh Le Nguyen 0001, Susumu Horiguchi, Yasushi Inoguchi
ECML2
2005 Improving discriminative sequential learning with rare--but--important associations
abstract
Discriminative sequential learning models like Conditional Random Fields (CRFs) have achieved significant success in several areas such as natural language processing or information extraction. Their key advantage is the ability to capture various non--independent and overlapping features of inputs. However, several unexpected pitfalls have a negative influence on the model's performance; these mainly come from an imbalance among classes/labels, irregular phenomena, and potential ambiguity in the training data. This paper presents a data--driven approach that can deal with such hard--to--predict data instances by discovering and emphasizing rare--but--important associations of statistics hidden in the training data. Mined associations are then incorporated into these models to deal with difficult examples. Experimental results of English phrase chunking and named entity recognition using CRFs show a significant improvement in accuracy. In addition to the technical perspective, our approach also highlights a potential connection between association mining and statistical learning by offering an alternative strategy to enhance learning performance with interesting and useful patterns discovered from large dataset.
Xuan-Hieu Phan, Minh Le Nguyen 0001, Susumu Horiguchi
KDD2