EDBT 2026 Demo / reviewers in the wild / expert
Minh-Tien Nguyen
dblp:119/2163
· DBLP profile ↗
37ranked-venue papers
20as first author
21since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 28 · 14 first-author · 18 since 2021Databases, data management, data science and information retrieval · 8 · 7 first-author · 1 since 2021Systems, architecture and hardware · 3 · 2 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | A Noise-Shaping SAR ADC With Second-Order Input PredictionabstractInternational audience Viet Nguyen-Thien, Chadi Jabbour, Minh-Tien Nguyen, Nicolas Delorme |
ISCAS | 3 |
| 2025 | From Span Extraction to Classification: A Multi-step Framework for Cognitive Distortion Analysis
Manh-Cuong Phan, Thi-Ngoc-Phuong Nguyen, Huu-Loi Le, Huy-The Vu, Hajime Hotta, Minh-Tien Nguyen |
PACLIC | 6 |
| 2025 | Automatic Prompt Selection for Large Language Models
Viet-Tung Do, Xuan-Quang Nguyen, Van-Khanh Hoang, Duy-Hung Nguyen, Shahab Sabahi, Jeff Yang, Hajime Hotta, Minh-Tien Nguyen, Hung Le 0002 |
PAKDD (3) | 8 |
| 2025 | A fine-tuning framework based on question, context, and answer relationships for enhancing legal information retrieval
Nhu Hai Phung, Nguyen Chi Thanh, Minh-Tien Nguyen, Thu Ha Nguyen, Huu-Loi Le, Truong-Phuc Nguyen |
Eng. Appl. Artif. Intell. | 3 |
| 2024 | Aspect-Based Sentiment Analysis of Clothing Reviews in Vietnamese E-commerce
Pham Quoc-Hung, Dinh Van-Dan, Huu-Loi Le, Le Thi-Viet-Huong, Nguyen Thu Ha, Xuan-Hieu Phan, Minh-Tien Nguyen, Pham Ngoc Hung |
PACLIC | 7 |
| 2024 | Improving Speech Recognition with Jargon InjectionabstractThis paper introduces a new method that improves the performance of Automatic speech recognition (ASR) engines, e.g., Whisper in practical cases.Different from prior methods that usually require both speech data and its transcription for decoding, our method only uses jargon as the context for decoding.To do that, the method first represents the jargon in a trie tree structure for efficient storing and traversing.The method next forces the decoding of Whisper to more focus on the jargon by adjusting the probability of generated tokens with the use of the trie tree.To further improve the performance, the method utilizes the prompting method that uses the jargon as the context.Final tokens are generated based on the combination of prompting and decoding.Experimental results on Japanese and English datasets show that the proposed method helps to improve the performance of Whisper, specially for domain-specific data.The method is simple but effective and can be deployed to any encoder-decoder ASR engines in actual cases.The code and data are also accessible. 1 Minh-Tien Nguyen, Dat Phuoc Nguyen, Tuan-Hai Luu, Xuan-Quang Nguyen, Tung-Duong Nguyen, Jeff Yang |
SIGDIAL | 1 |
| 2024 | Improving biomedical Named Entity Recognition with additional external contexts
Bui Duc Tho, Minh-Tien Nguyen, Lin-Lung Ying, Shumpei Inoue, Tri-Thanh Nguyen |
J. Biomed. Informatics | 2 |
| 2024 | Learning to generate text with auxiliary tasks
Pham Quoc-Hung, Minh-Tien Nguyen, Shumpei Inoue, Manh Tran-Tien, Xuan-Hieu Phan |
Knowl. Based Syst. | 2 |
| 2024 | Towards Vietnamese Question and Answer Generation: An Empirical StudyabstractQuestion-answer generation (QAG) is a challenging task that generates both questions and answers from a given input paragraph context. The QAG task has recently achieved promising results thanks to the appearance of large pre-trained language models, yet, QAG models are mainly implemented in common languages, e.g., English. There still remains a gap in domain and language adaptation of these QAG models to low-resource languages such as Vietnamese. To address the gap, this article presents a large-scale and systematic study of QAG in Vietnamese. To do that, we first implement several QAG models by using the common fine-tuning techniques based on powerful pre-trained language models. We next introduce a set of instructions designed for the QAG task. These instructions are used to fine-tuned the pre-trained language and large language models. Extensive experimental results of both automatic and human evaluation on five benchmark machine reading comprehension datasets show two important points. First, the instruction-tuning method has the potential to enhance the performance of QAG models. Second, large language models trained in English need more data for fine-tuning to work well on the downstream QAG tasks of low-resource languages. We also provide a prototype system to demonstrate how our QAG models actually work. The code for fine-tuning QAG models and instructions are also made available. Pham Quoc-Hung, Huu-Loi Le, Dang Nhat Minh, T. Tran Khang, Manh Tran-Tien, Viet-Hung Dang, Huy-The Vu, Minh-Tien Nguyen, Xuan-Hieu Phan |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 8 |
| 2023 | Emotion-Cause Pair Extraction as Question Answering
Huu-Hiep Nguyen, Minh-Tien Nguyen |
ICAART (3) | 2 |
| 2023 | ViPubmedDeBERTa: A Pre-trained Model for Vietnamese Biomedical Text
Manh Tran-Tien, Huu-Loi Le, Dang Nhat Minh, T. Tran Khang, Huy-The Vu, Minh-Tien Nguyen |
PACLIC | 6 |
| 2023 | Label-representative graph convolutional network for multi-label text classification
The H. Vu, Minh-Tien Nguyen, Van-Chien Nguyen, Minh-Hieu Pham, Van-Quyet Nguyen, Van-Hau Nguyen |
Appl. Intell. | 2 |
| 2023 | Gain more with less: Extracting information from business documents with small data
Minh-Tien Nguyen, Nguyen Hong Son, Le Thai Linh |
Expert Syst. Appl. | 1 |
| 2022 | Improving Document Image Understanding with Reinforcement Finetuning
Bao-Sinh Nguyen, Hieu M. Vu, Tuan-Anh D. Nguyen, Minh-Tien Nguyen, Hung Le 0002 |
ICONIP (7) | 5 |
| 2022 | Jointly Learning Span Extraction and Sequence Labeling for Information Extraction from Business DocumentsabstractThis paper introduces a new information extraction model for business documents. Different from prior studies which only base on span extraction or sequence labeling, the model takes into account advantage of both span extraction and sequence labeling. The combination allows the model to deal with long documents with sparse information (the small amount of extracted information). The model is trained end-to-end to jointly optimize the two tasks in a unified manner. Experimental results on four business datasets in English and Japanese show that the model achieves promising results and is significantly faster than the normal span-based extraction method. The code is also available.11https://bit.ly/3iYztCL. It will be available on github. Nguyen Hong Son, Hieu M. Vu, Tuan-Anh D. Nguyen, Minh-Tien Nguyen |
IJCNN | 4 |
| 2022 | Label Correlation Based Graph Convolutional Network for Multi-label Text ClassificationabstractMulti-label text classification aims to assign a set of most relevant labels to a given document. To build such a classifier, apart from demanding an efficient document representation, capturing label information for classification performance improvement is still challenging. In this paper, we propose a novel model based on a graph convolutional network to model label correlation. To do that, we design a correlation matrix from labels in a data-driven way. The learned label correlations are then fused with fine-grained document information extracted by a RoBERTa-based subnet for classification. Furthermore, we introduce a simple mechanism to make the label correlation matrix more effective in propagating information among label nodes. We first normalize the correlation matrix to deal with the highly skewed problem and then filter noisy edges to alleviate the long-tailed distribution problem. Evaluation results show that our model achieves competitive results compared to existing state-of-the-art methods. Ablation studies are also conducted to explore the proposed model's behaviors. Huy-The Vu, Minh-Tien Nguyen, Van-Chien Nguyen, Manh Tran-Tien, Van-Hau Nguyen |
IJCNN | 2 |
| 2022 | Enhance Incomplete Utterance Restoration by Joint Learning Token Extraction and Text GenerationabstractShumpei Inoue, Tsungwei Liu, Son Nguyen, Minh-Tien Nguyen. Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2022. Shumpei Inoue, Tsungwei Liu, Minh-Tien Nguyen |
NAACL-HLT | 4 |
| 2021 | Robust Deep Reinforcement Learning for Extractive Legal Summarization
Duy-Hung Nguyen, Bao-Sinh Nguyen, Nguyen-Viet-Dung Nghiem, Mim Amina Khatun, Minh-Tien Nguyen, Hung Le 0002 |
ICONIP (6) | 6 |
| 2021 | Loss-based Active Learning for Named Entity RecognitionabstractThis paper addresses the practical issue of lacking training data when building named entity recognition (NER) systems. To this aim, we introduce a new active learning method for reducing the number of training samples required by the underlying NER system. Different from prior work that only focuses on training data, we define a new loss function that when estimating loss and uncertainty scores of training samples for selection, it takes also into account the uncertainty of the$K$unlabelled test instances most similar to the unlabelled training instances. Experimental results on both general domain and clinical benchmark datasets show that the proposed active learning method allows to train the NER system with between 5% to 7% less training data compared to state of the art uncertainty sampling methods, while retaining high NER effectiveness. Le Thai Linh, Minh-Tien Nguyen, Guido Zuccon, Gianluca Demartini |
IJCNN | 2 |
| 2021 | Information Extraction of Domain-Specific Business Documents with Limited DataabstractInformation extraction is a key corner-stone in the digitization of office data which requires the conversion of unstructured to structured data. However, in the actual application to business cases, there is a big deadlock to adapt common extraction systems to domain-specific documents due to the limitation of preparation of training data. To overcome this issue, we introduce a model, which employs pre-trained language models with a customized CNN layer for domain adaptation. The model is validated on three Japanese domain-specific and two benchmark machine reading comprehension data sets (SQuADs). Experimental results confirm that our model achieves promising results which are applicable for actual business scenarios. Minh-Tien Nguyen, Le Thai Linh, Nguyen Hong Son, Do Hoang Thai Duong, Bui Cong Minh, Akira Shojiguchi |
IJCNN | 1 |
| 2021 | Transformers-based information extraction with limited data for domain-specific business documents
Minh-Tien Nguyen, Le Thai Linh |
Eng. Appl. Artif. Intell. | 1 |
| 2020 | AURORA: An Information Extraction System of Domain-specific Business Documents with Limited DataabstractInformation extraction is a well-known topic that plays a critical role in many NLP applications as its outputs can be considered as an entrance step for digital transformation. However, there still exist gaps when applying research results to actual business cases. This paper introduces AURORA, an information extraction for domain-specific business documents. The intuition of AURORA is to use transfer learning for extraction. To do that, it utilizes the power of transformers for dealing with the limitation of training data in business cases and stacks additional layers for domain adaptation. We demonstrate AURORA in the context of actual scenarios where users are invited to experience two functions: fine-grained and whole paragraph extraction of Japanese business documents. A video of the system is available at http://y2u.be/xHQpYE41Tqw. Minh-Tien Nguyen, Le Thai Linh, Nguyen Hong Son, Do Hoang Thai Duong, Bui Cong Minh, Nguyen Hai Phong, Nguyen Huu Hiep |
CIKM | 1 |
| 2020 | Sentence Compression as Deletion with Contextual Embeddings
Minh-Tien Nguyen, Bui Cong Minh, Le Thai Linh |
ICCCI | 1 |
| 2020 | Understanding Transformers for Information Extraction with Limited Data
Minh-Tien Nguyen, Nguyen Hong Son, Bui Cong Minh, Do Hoang Thai Duong, Le Thai Linh |
PACLIC | 1 |
| 2019 | Web document summarization by exploiting social context with matrix co-factorization
Minh-Tien Nguyen, Tran Viet Cuong, Nguyen Xuan Hoai, Minh Le Nguyen 0001 |
Inf. Process. Manag. | 1 |
| 2018 | Towards Social Context Summarization with Convolutional Neural Networks
Minh-Tien Nguyen, Vu D. Tran, Viet-Anh Phan, Minh Le Nguyen 0001 |
CICLing (2) | 1 |
| 2018 | TSix: A Human-involved-creation Dataset for Tweet Summarization
Minh-Tien Nguyen, Viet Dac Lai, Minh Le Nguyen 0001 |
LREC | 1 |
| 2018 | Social context summarization using user-generated content and third-party sources
Minh-Tien Nguyen, Vu D. Tran, Minh Le Nguyen 0001 |
Knowl. Based Syst. | 1 |
| 2018 | Exploiting User Posts for Web Document SummarizationabstractRelevant user posts such as comments or tweets of a Web document provide additional valuable information to enrich the content of this document. When creating user posts, readers tend to borrow salient words or phrases in sentences. This can be considered as word variation. This article proposes a framework that models the word variation aspect to enhance the quality of Web document summarization. Technically, the framework consists of two steps: scoring and selection. In the first step, the social information of a Web document such as user posts is exploited to model intra-relations and inter-relations in lexical and semantic levels. These relations are denoted by a mutual reinforcement similarity graph used to score each sentence and user post. After scoring, summaries are extracted by using a ranking approach or concept-based method formulated in the form of Integer Linear Programming. To confirm the efficiency of our framework, sentence and story highlight extraction tasks were taken as a case study on three datasets in two languages, English and Vietnamese. Experimental results show that: (i) the framework can improve ROUGE-scores compared to state-of-the-art baselines of social context summarization and (ii) the combination of the two relations benefits the sentence extraction of single Web documents. Minh-Tien Nguyen, Vu D. Tran, Minh Le Nguyen 0001, Xuan-Hieu Phan |
ACM Trans. Knowl. Discov. Data | 1 |
| 2017 | Summarizing Web Documents Using Sequence Labeling with User-Generated Content and Third-Party Sources
Minh-Tien Nguyen, Vu D. Tran, Chien-Xuan Tran, Minh Le Nguyen 0001 |
NLDB | 1 |
| 2017 | Intra-relation or inter-relation?: Exploiting social information for Web document summarization
Minh-Tien Nguyen, Minh Le Nguyen 0001 |
Expert Syst. Appl. | 1 |
| 2016 | SoLSCSum: A Linked Sentence-Comment Dataset for Social Context SummarizationabstractThis paper presents a dataset named SoLSCSum for social context summarization. The dataset includes 157 open-domain articles along with their comments collected from Yahoo News. The articles and their comments were manually annotated by two annotators to extract standard summaries. The inter-annotator agreement is 74.5% and Cohen's Kappa is 0.5845. To illustrate the potential use of our dataset, a learning to rank model was trained by using a set of local and cross features. Experimental results demonstrate that: (1) our model trained by Ranking SVM obtains significant improvements from 5.5% to 14.8% of ROUGE-1 over state-of-the-art baselines in document summarization and (2) our dataset can be used to train summary methods such as SVM. Minh-Tien Nguyen, Chien-Xuan Tran, Vu D. Tran, Minh Le Nguyen 0001 |
CIKM | 1 |
| 2016 | SoRTESum: A Social Context Framework for Single-Document Summarization
Minh-Tien Nguyen, Minh Le Nguyen 0001 |
ECIR | 1 |
| 2016 | Learning to Summarize Web Documents Using Social InformationabstractThis paper presents a method named SoSVMRank, which integrates the social information of a Web document to generate a high-quality summarization. In order to do that, the summarization was formulated as a learning to rank task, in which the order of a sentence or comment was determined by its informative information. The informative information was measured by a set of local and social features in which the social features were exploited to support the local ones when modeling a sentence or comment. To enrich information, new features were also proposed. After ranking, top m ranked sentences and comments were selected as the summarization. Our method was extensively evaluated on two datasets. Promising results indicate that: (1) by using new features, our method achieves improvements in both ROUGE-1 and ROUGE-2 of the summarization over state-of-the-art baselines and (2) integrating social information benefits the summarization. Minh-Tien Nguyen, Vu D. Tran, Chien-Xuan Tran, Minh Le Nguyen 0001 |
ICTAI | 1 |
| 2016 | A flexible receiver using ΔΣ modulationabstractThis paper presents a reconfigurable 2nd/3rd-order discrete-time direct RF-to-digital ΔΣ receiver architecture for wide frequency range flexible receivers. Using 25% duty-cycle current-driven passive mixer along with RF feedback enables high-Q bandpass filtering and relaxes the linearity requirement on LNTA. Moreover, the passive/active implementation of the loop filter gives a good trade-off between power consumption, linearity and dynamic range. A design example with 10 MHz useful bandwidth and 0.4-4.0 GHz frequency range is conducted to demonstrate the feasibility and the characteristics of this proposed architecture. Minh-Tien Nguyen, Chadi Jabbour, Van-Tam Nguyen 0004 |
ISCAS | 1 |
| 2015 | TSum4act: A Framework for Retrieving and Summarizing Actionable Tweets During a Disaster for Reaction
Minh-Tien Nguyen, Asanobu Kitamoto, Tri-Thanh Nguyen |
PAKDD (2) | 1 |
| 2013 | Direct delta-sigma receiver: Analysis, modelization and simulationabstractThis paper presents a model of direct delta-sigma receiver (DDSR) and methodology for theoretical transfer function (TF). The theoretical analysis is carried out by modeling the key elements of the DDSR, including the N-path filter, down-conversion mixer, baseband delta sigma modulator (DSM) and FIRDACs in order to optimize the loop filter coefficients. The contribution of different noise sources and the impact of the nonlinearities are also analyzed. The obtained simulation results show the accuracy of the proposed models and method. This top-down approach allows the designer to determine the key parameters for DDSR in terms of noise contributions, nonlinearity impacts, filtering effects as well as gain distribution and thus enables the optimization of the whole receiver. Although this methodology was applied to the conventional DDSR, but it can be used for any receiver architecture based on DDSR. Minh-Tien Nguyen, Chadi Jabbour, Cyrius Ouffoue, Rayan Mina, Florent Sibille, Patrick Loumeau, Pascal Triaire, Van Tam Nguyen 0001 |
ISCAS | 1 |