VLDB 2026 Research / reviewers in the wild / expert
Ha-Thanh Nguyen
dblp:158/7894
· DBLP profile ↗
25ranked-venue papers
11as first author
18since 2021 · last 2026
0000-0003-2794-7010ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 16 · 5 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 10 · 6 first-author · 9 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An Empirical Study of Architectural Trade-Offs in Vision-Based Traffic Sign Interpretation Systems
Su Myat Noe, Ha-Thanh Nguyen, May Myo Zin, Ken Satoh |
ICAART (4) | 2 |
| 2026 | BIS Reasoning 1.0: The First Large-Scale Japanese Benchmark for Belief-Inconsistent Syllogistic ReasoningabstractWe present BIS Reasoning 1.0, the first large-scale Japanese dataset of syllogistic reasoning problems explicitly designed to evaluate belief-inconsistent reasoning in large language models (LLMs). Unlike prior resources such as NeuBAROCO and JFLD, which emphasize general or belief-aligned logic, BIS Reasoning 1.0 systematically introduces logically valid yet belief-inconsistent syllogisms to expose belief bias, the tendency to accept believable conclusions irrespective of validity. We benchmark a representative suite of cutting-edge models, including OpenAI GPT-5 variants, GPT-4o, Qwen, and prominent Japanese LLMs, under a uniform, zero-shot protocol. Reasoning-centric models achieve near-perfect accuracy on BIS Reasoning 1.0 (e.g., Qwen3-32B $\approx$99% and GPT-5-mini up to $\approx$99.7%), while GPT-4o attains around 80%. Earlier Japanese-specialized models underperform, often well below 60%, whereas the latest llm-jp-3.1-13b-instruct4 markedly improves to the mid-80% range. These results indicate that robustness to belief-inconsistent inputs is driven more by explicit reasoning optimization than by language specialization or scale alone. Our analysis further shows that even top-tier systems falter when logical validity conflicts with intuitive or factual beliefs, and that performance is sensitive to prompt design and inference-time reasoning effort. We discuss implications for safety-critical domains, including law, healthcare, and scientific literature, where strict logical fidelity must override intuitive belief to ensure reliability. Ha-Thanh Nguyen, Hideyuki Tachibana, Qianying Liu, Su Myat Noe, Koichi Takeda 0003, Sadao Kurohashi |
LREC | 1 |
| 2025 | Detecting Misleading Information with LLMs and Explainable ASPabstractInternational audience Quang-Anh Nguyen, Thu-Trang Pham, Thi-Hai-Yen Vuong, Giang V. Trinh, Ha-Thanh Nguyen |
ICAART (3) | 5 |
| 2025 | DeCoRA: Definition and Context Reasoning in ArgumentationabstractIn the legal field, accurately interpreting and applying legal definitions is crucial yet challenging due to inherent ambiguities. This paper introduces DeCoRA, a novel framework that enhances legal argumentation by incorporating context-based reasoning to address these ambiguities, with a focus on the judge as the central decision maker. Unlike black-box models, such as generative models or outcome prediction systems, which often produce outputs without fully explaining the reasoning behind their conclusions, DeCoRA emphasizes transparency by modeling the judicial decision-making process in a structured and interpretable manner. Our key contributions include: (1) a tree-based knowledge base that organizes legal definitions, highlighting their relationships and effects; (2) a context-aware definition framework enabling judges to interpret definitions considering legal and contextual relevance; and (3) an effective method for handling complex legal scenarios with conflicting or overlapping definitions. Ngoc-Duy Mai, Xuan-Bach Le, Thi-Hai-Yen Vuong, Ha-Thanh Nguyen, Kostas Stathis, Ken Satoh |
ICAIL | 4 |
| 2025 | Uncovering connections: a reference network approach to statute law retrieval
Thi-Hai-Yen Vuong, Hai-Long Nguyen 0001, Tan-Minh Nguyen, Ha-Thanh Nguyen, Minh Le Nguyen 0001, Xuan-Hieu Phan |
Appl. Intell. | 4 |
| 2024 | ConsRAG: Minimize LLM Hallucinations in the Legal DomainabstractRetrieval-Augmented Generation (RAG) systems have shown potential in improving legal question-answering applications. However, they often struggle to provide precise information for legal queries, as broad-topic relevance may not always align with contextual usefulness. To address this challenge, we introduce Constrained Retrieval-Augmented Generation (ConsRAG), a novel approach that employs aspect-based constraints during both retrieval and generation phases. ConsRAG aims to enhance precision and contextual relevance in legal outputs through these constraints and an iterative backtracking mechanism. Our experiments suggest that ConsRAG may offer improvements over existing systems in terms of accuracy and relevance of retrieved documents. This paper presents the framework, implementation, and evaluation of ConsRAG, exploring its potential to enhance the reliability of legal AI applications. Ha-Thanh Nguyen, Ken Satoh |
JURIX | 1 |
| 2023 | How Fine Tuning Affects Contextual Embeddings: A Negative Result Explanation
Ha-Thanh Nguyen, Phuong Minh Nguyen 0001, Minh Le Nguyen 0001, Ken Satoh |
ICAART (3) | 1 |
| 2023 | Improving Translation of Case Descriptions into Logical Fact Formulas using LegalCaseNERabstractThe automated translation of natural language text into structured logical representations is a critical task in various applications, including legal reasoning and decision-making. This paper presents a Name Entity Recognition (NER) based approach for translating the legal case descriptions written in natural language into PROLEG fact formulas. The approach comprises (1) extracting legal entities from the case description using a specialized NER model, namely LegalCaseNER and (2) transforming the extracted entities into PROLEG fact formulas using PROLEG rules. The experimental results demonstrate the efficacy of our proposed approach in accurately extracting relevant entities from legal case descriptions and translating them into the appropriate PROLEG fact formulas. Our approach provides a promising solution for handling complex and diverse case descriptions, enabling their representation in a structured format. This work provides a foundation for future research in the application of logical fact formulas in legal reasoning and decision-making. May Myo Zin, Ha-Thanh Nguyen, Ken Satoh, Saku Sugawara, Fumihito Nishino |
ICAIL | 2 |
| 2023 | LogiLaw Dataset Towards Reinforcement Learning from Logical Feedback (RLLF)abstractLarge Language Models (LLMs) face limitations in logical reasoning, which restrict their applicability in critical domains such as law. Current evaluation methods often lead to inaccurate assessments of LLMs’ capabilities due to their simplicity. This paper presents a refined evaluation method for assessing LLMs’ capability to answer legal questions by eliminating the possibility of obtaining correct responses by chance. Furthermore, we introduce the LogiLaw dataset, which aims to enhance the models’ logical reasoning capacities in general and legal reasoning specifically. By leveraging the refined evaluation technique, the LogiLaw dataset, and the proposed Reinforcement Learning from Logical Feedback (RLLF) approach, our work aims to open new avenues for research to bolster LLMs’ performance in law and other logic-intensive disciplines while addressing the shortcomings of conventional evaluation approaches. Ha-Thanh Nguyen, Wachara Fungwacharakorn, Ken Satoh |
JURIX | 1 |
| 2023 | LawGiBa - Combining GPT, Knowledge Bases, and Logic Programming in a Legal Assistance SystemabstractWe present LawGiBa, a proof-of-concept demonstration system for legal assistance that combines GPT, legal knowledge bases, and Prolog’s logic programming structure to provide explanations for legal queries. This novel combination effectively and feasibly addresses the hallucination issue of large language models (LLMs) in critical domains, such as law. Through this system, we demonstrate how incorporating a legal knowledge base and logical reasoning can enhance the accuracy and reliability of legal advice provided by AI models like GPT. Though our work is primarily a demonstration, it provides a framework to explore how knowledge bases and logic programming structures can be further integrated with generative AI systems, to achieve improved results across various natural languages and legal systems. Ha-Thanh Nguyen, Randy Goebel, Francesca Toni, Kostas Stathis, Ken Satoh |
JURIX | 1 |
| 2023 | Information Extraction from Lengthy Legal Contracts: Leveraging Query-Based Summarization and GPT-3.5abstractIn the legal domain, extracting information from contracts poses significant challenges, primarily due to the scarcity of annotated data. In such situations, leveraging large language models (LLMs), such as the Generative Pretrained Transformer (GPT) models, offers a promising solution. However, the inherent token limitations of these models can be a bottleneck for processing lengthy legal contracts. This paper presents an unsupervised two-step approach to address these challenges. First, we propose a query-based summarization model that extracts sentences pertinent to predefined queries, concisely representing lengthy contracts. This summarization ensures that the core information remains intact while simultaneously addressing the token limitation issue. Subsequently, the generated summary is fed to GPT-3.5 for precise information extraction. Our approach effectively overcomes the challenges of token limitations and zero resources, enabling efficient and scalable information extraction from legal contracts. We compare our results with those obtained from supervised models that have been fine-tuned on domain-specific annotated data. Experimental results demonstrate the remarkable effectiveness of our approach, as it achieves state-of-the-art performance without the need for domain-specific training data. May Myo Zin, Ha-Thanh Nguyen, Ken Satoh, Saku Sugawara, Fumihito Nishino |
JURIX | 2 |
| 2022 | Learning to Map the GDPR to Logic Representation on DAPRECO-KB
Phuong Minh Nguyen 0001, Thi-Thu-Trang Nguyen, Vu D. Tran, Ha-Thanh Nguyen, Minh Le Nguyen 0001, Ken Satoh |
ACIIDS (1) | 4 |
| 2022 | Logical Structure-based Pretrained Models for Legal Text Processing
Ha-Thanh Nguyen, Minh Le Nguyen 0001 |
ICAART (3) | 1 |
| 2022 | A Survey of Pretrained Embeddings for Japanese Legal Representation
Ha-Thanh Nguyen, Minh Le Nguyen 0001, Ken Satoh |
IEA/AIE | 1 |
| 2022 | A Multi-Step Approach in Translating Natural Language into Logical FormulaabstractTranslating often has the meaning of converting from one human language to another. However, in a broader sense, it means transforming a message from one form of communication to another form. Logic is an important form of communication and the ability to translate natural language into logic is important in many different fields, in which logical reasoning and logical arguments are used. In the legal field, for example, judges must often reason from facts and arguments presented in natural language to logical conclusions. In this paper, toward the goal of support for this kind of reasoning with machines, we propose a method for translating natural language into logical representations using a combination of deep learning methods. Our approach contributes methodologies and insights to the development of computational methods for converting natural language into logical representations. Ha-Thanh Nguyen, Wachara Fungwacharakorn, Fumihito Nishino, Ken Satoh |
JURIX | 1 |
| 2022 | An Interactive Natural Language Interface for PROLEGabstractPROLEG is a famous computer program supporting attorneys in the legal inference process. However, the input of this system is expressed in Prolog, which most lawyers are not familiar with. This technical barrier is a serious problem for using PROLEG in the real legal context. A natural language interface is one of the solutions to this problem. We have developed a prototype of such an interface. This paper describes the prototype and its current performance. The prototype translates input facts into Prolog expressions following PROLEG syntax. The system consists of three main modules, (1) natural language perceiver, (2) PROLEG reasoner, and (3) inference explainer. In addition, we analyze the performance of the prototype and identify existing issues and discuss possible solutions. Ha-Thanh Nguyen, Fumihito Nishino, Megumi Fujita, Ken Satoh |
JURIX | 1 |
| 2021 | Few-Shot Tuning Framework for Automated Terms of Service GenerationabstractIn this paper, we introduce BART2S a novel framework based on BART pretrained models to generate terms of service in high quality. The framework contains two parts: a generator finetuned with multiple tasks and a discriminator fine-tuned to distinguish the fair and unfair terms. Besides the novelty in design and the implementation contributions, the proposed framework can support drafting terms of service, a growing need in the digital age. Our proposed approach allows the system to reach a balance between automation and the will expression of the service provider. Through experiments, we demonstrate the effectiveness of the method and discuss potential future directions. Ha-Thanh Nguyen, Kiyoaki Shirai, Minh Le Nguyen 0001 |
JURIX | 1 |
| 2021 | VSEC: Transformer-Based Model for Vietnamese Spelling Correction
Dinh-Truong Do, Ha-Thanh Nguyen, Thang Ngoc Bui, Hieu Dinh Vo |
PRICAI (2) | 2 |
| 2020 | Answering Legal Questions by Learning Neural Attentive Text RepresentationabstractText representation plays a vital role in retrieval-based question answering, especially in the legal domain where documents are usually long and complicated.The better the question and the legal documents are represented, the more accurate they are matched.In this paper, we focus on the task of answering legal questions at the article level.Given a legal question, the goal is to retrieve all the correct and valid legal articles, that can be used as the basic to answer the question.We present a retrieval-based model for the task by learning neural attentive text representation.Our text representation method first leverages convolutional neural networks to extract important information in a question and legal articles.Attention mechanisms are then used to represent the question and articles and select appropriate information to align them in a matching process.Experimental results on an annotated corpus consisting of 5,922 Vietnamese legal questions show that our model outperforms state-of-the-art retrieval-based methods for question answering by large margins in terms of both recall and NDCG. Manh-Kien Phi, Ha-Thanh Nguyen, Ngo Xuan Bach, Vu D. Tran, Minh Le Nguyen 0001, Tu Minh Phuong |
COLING | 2 |
| 2020 | Progressive Training in Recurrent Neural Networks for Chord Progression Modeling
Trung-Kien Vu, Teeradaj Racharak, Satoshi Tojo, Ha-Thanh Nguyen, Minh Le Nguyen 0001 |
ICAART (2) | 4 |
| 2020 | Comparison of Algorithms for Tree-top Detection in Drone Image Mosaics of Japanese Mixed Forests
Yago Diez Donoso, Sarah Kentsch, Maximo Larry Lopez Caceres, Ha-Thanh Nguyen, Daniel Serrano, Ferran Roure |
ICPRAM | 4 |
| 2020 | How State-Of-The-Art Models Can Deal With Long-Form Question Answering
Minh-Quan Bui, Ha-Thanh Nguyen, Minh Le Nguyen 0001 |
PACLIC | 3 |
| 2020 | Latent Topic Refinement based on Distance Metric Learning and Semantics-assisted Non-negative Matrix Factorization
Tran Binh Dang, Ha-Thanh Nguyen, Minh Le Nguyen 0001 |
PACLIC | 2 |
| 2019 | Swarm Filter - A Simple Deep Learning Component Inspired by Swarm ConceptabstractSwarm is a research topic not only of biologists but also for computer scientists for years. With the idea of swarm intelligence in nature, optimal algorithms are proposed to solve different problems. In addition to the proactive aspect, a swarm can provide useful hints for identification problems. There are features that only exist when an individual belongs to a swarm. An idea came to us, deep learning networks have the ability to automatically select features, so they can extract the characteristics of a swarm for identification problems. This is a new idea in the combination of swarm characteristic with deep learning model. The previous studies combined swarm intelligence with neural networks to find the optimal parameters and architecture for the model. When performing our experiments, we were surprised that this simple architecture got a state-of-the-art result. This interesting discovery can be applied to other tasks using deep learning. Ha-Thanh Nguyen, Minh Le Nguyen 0001 |
ICTAI | 1 |
| 2014 | An Architecture for Web Services Mash-Up based on Mobile AgentsabstractIn the present day, there are many useful domain information scattered over the Internet. But domain users need to navigate different websites or information sources to get domain relevant information. Services mash-up can be used to integrate not only web content and also services to adapt the display on users' clients. However, the user's requests for contents generating the UI displays might consider participating in mash-up process as time-consuming, even a trouble of network bottleneck. To address this problem, we proposed an architecture for web service mash-up system based on mobile agents. By taking using the mobile agents, mash-up system can be adapt for current context to provide better visualization support to users. The experiments to support mobile agent base web service mash-up has been implemented. Quang-Dung Vu, Ha-Thanh Nguyen, Danh-Viet Vu, Viet Ha Nguyen 0001, Nobuyasu Nakajima |
iiWAS | 2 |