EDBT 2026 Demo / reviewers in the wild / expert
May Myo Zin
dblp:286/6557
· DBLP profile ↗
9ranked-venue papers
7as first author
9since 2021 · last 2026
0000-0003-1315-7704ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 4 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 6 · 5 first-author · 6 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | An Empirical Study of Architectural Trade-Offs in Vision-Based Traffic Sign Interpretation Systems
Su Myat Noe, Ha-Thanh Nguyen, May Myo Zin, Ken Satoh |
ICAART (4) | 3 |
| 2025 | Data Augmented Pipeline for Legal Information Extraction and ReasoningabstractIn this paper, we propose a pipeline leveraging Large Language Models (LLMs) for data augmentation in Information Extraction tasks within the legal domain. The proposed method is both simple and effective, significantly reducing the manual effort required for data annotation while enhancing the robustness of Information Extraction systems. Furthermore, the method is generalizable, making it applicable to various Natural Language Processing (NLP) tasks beyond the legal domain. Phuong Minh Nguyen 0001, Thanh Ha Nguyen, May Myo Zin, Ken Satoh |
ICAIL | 3 |
| 2025 | Towards Machine-Readable Traffic Laws: Formalizing Traffic Rules into PROLOG Using LLMsabstractEnsuring autonomous vehicles (AVs) adhere to traffic rules is crucial for safety. Formalizing these rules into machine-readable formats offers a consistent, unambiguous foundation for automated reasoning and compliance. However, the formalization process is traditionally manual, resource-intensive, and prone to error. This study explores using large language models (LLMs) to automate the translation of traffic rules into PROLOG, a declarative programming language ideal for encoding logical rules and relationships. The proposed methodology consists of three key phases: extracting traffic rules from diverse textual sources, structuring them into Logical English (LE) for clarity and consistency, and translating them into PROLOG representations using advanced natural language processing (NLP) techniques, including in-context learning and fine-tuning. The experimental results demonstrate the effectiveness of LLMs in automating this process, achieving high accuracy in translation. The findings underscore the potential for scaling this methodology to accommodate broader regulatory frameworks, paving the way for safer, more reliable AV operations in complex and dynamic traffic environments. May Myo Zin, Georg Borges, Ken Satoh, Wachara Fungwacharakorn |
ICAIL | 1 |
| 2025 | From Court Decisions to Guiding Principles: Advancing Complex Legal Summarization with LLMsabstractGuiding principles (Leits´latze) are central to German jurisprudence, capturing the essence of judicial reasoning in concise, doctrinally precise statements. Unlike general case summaries, they distill normative reasoning and key legal holdings rather than recounting factual backgrounds or procedural details. This paper examines how large language models (LLMs) can automatically generate guiding principles from German court decisions, focusing on three dimensions: model choice and adaptation, prompting strategies, and evaluation methods. Comparing GPT-4o with LerLeoLM, a fine-tuned, domain-specific model, we find that GPT-4o outperforms even without fine-tuning, while lightweight tuning on only 100 cases yields the best results. Human–LLM co-designed prompts further enhance quality, surpassing both expert-structured and self-generated prompts. Finally, we show that LLM-based evaluation aligns more closely with expert judgment than traditional metrics, establishing guiding principle generation as a distinct task in Legal NLP. May Myo Zin, Ken Satoh, Georg Borges |
JURIX | 1 |
| 2024 | Leveraging LLM for Identification and Extraction of Normative StatementsabstractThe development of autonomous vehicles (AVs) requires a comprehensive understanding of both explicit and implicit traffic rules to ensure legal compliance and safety. While explicit traffic laws are well-defined in statutes and regulations, implicit rules derived from judicial interpretations and case law are more nuanced and challenging to extract. This research investigates the potential of Large Language Models (LLMs), particularly GPT-4o, in automating the extraction of implicit traffic rules from judicial decisions. By utilizing various prompt engineering techniques, including Standard Prompts, Chain-of-Thought (CoT), Chain-of-Instructions (CoI), and Layer-of-Thought (LoT) prompts, this study aims to assess the effectiveness of GPT-4o in identifying normative content relevant to specific traffic laws. The contributions of this paper include an assessment of LLMs for legal text processing, the automation of implicit rule extraction, and the development of a scalable framework that can continuously update as new legal precedents emerge. The results indicate promising avenues for integrating automated normative extraction in AV systems, improving both the safety and legal compliance of autonomous driving technologies. May Myo Zin, Ken Satoh, Georg Borges |
JURIX | 1 |
| 2023 | Improving Translation of Case Descriptions into Logical Fact Formulas using LegalCaseNERabstractThe automated translation of natural language text into structured logical representations is a critical task in various applications, including legal reasoning and decision-making. This paper presents a Name Entity Recognition (NER) based approach for translating the legal case descriptions written in natural language into PROLEG fact formulas. The approach comprises (1) extracting legal entities from the case description using a specialized NER model, namely LegalCaseNER and (2) transforming the extracted entities into PROLEG fact formulas using PROLEG rules. The experimental results demonstrate the efficacy of our proposed approach in accurately extracting relevant entities from legal case descriptions and translating them into the appropriate PROLEG fact formulas. Our approach provides a promising solution for handling complex and diverse case descriptions, enabling their representation in a structured format. This work provides a foundation for future research in the application of logical fact formulas in legal reasoning and decision-making. May Myo Zin, Ha-Thanh Nguyen, Ken Satoh, Saku Sugawara, Fumihito Nishino |
ICAIL | 1 |
| 2023 | Information Extraction from Lengthy Legal Contracts: Leveraging Query-Based Summarization and GPT-3.5abstractIn the legal domain, extracting information from contracts poses significant challenges, primarily due to the scarcity of annotated data. In such situations, leveraging large language models (LLMs), such as the Generative Pretrained Transformer (GPT) models, offers a promising solution. However, the inherent token limitations of these models can be a bottleneck for processing lengthy legal contracts. This paper presents an unsupervised two-step approach to address these challenges. First, we propose a query-based summarization model that extracts sentences pertinent to predefined queries, concisely representing lengthy contracts. This summarization ensures that the core information remains intact while simultaneously addressing the token limitation issue. Subsequently, the generated summary is fed to GPT-3.5 for precise information extraction. Our approach effectively overcomes the challenges of token limitations and zero resources, enabling efficient and scalable information extraction from legal contracts. We compare our results with those obtained from supervised models that have been fine-tuned on domain-specific annotated data. Experimental results demonstrate the remarkable effectiveness of our approach, as it achieves state-of-the-art performance without the need for domain-specific training data. May Myo Zin, Ha-Thanh Nguyen, Ken Satoh, Saku Sugawara, Fumihito Nishino |
JURIX | 1 |
| 2022 | Expand-Extract: A Parallel Corpus Mining Framework from Comparable Corpora for English-Myanmar Machine TranslationabstractHigh-quality neural machine translation (NMT) systems rely on the availability of large-scale and reliable parallel data. Since Myanmar language is a low-resource language, the parallel corpus of English-Myanmar language pair is sparse in volume. In this paper, we present a simple yet effective framework to create a parallel corpus from the available comparable corpora. Our proposed system first uses self-training and back-translation approaches together with the denoising-based automatic post-editing (DbAPE) system for augmenting synthetic datasets that are used to expand the size of existing comparable corpora. Then, LaBSE-based sentence embeddings and the proposed scoring function are applied to extract parallel sentences from the expanded comparable corpora. The extracted parallel sentences can be used to supplement parallel corpus when training the low-resource English-Myanmar NMT systems. We investigate the effectiveness of our methods by evaluating the NMT systems trained on the concatenation of parallel data created by our framework and an existing dataset. We show that the proposed framework is capable of creating a reliable parallel corpus, and that the created corpus substantially increases translation quality of MT systems trained on the existing parallel data, as measured by automatic evaluation metrics. May Myo Zin, Teeradaj Racharak, Minh Le Nguyen 0001 |
ICTAI | 1 |
| 2021 | Construct-Extract: An Effective Model for Building Bilingual Corpus to Improve English-Myanmar Machine Translation
May Myo Zin, Teeradaj Racharak, Minh Le Nguyen 0001 |
ICAART (2) | 1 |