EDBT 2026 Demo / reviewers in the wild / expert
Mo El-Haj
dblp:355/8324
· DBLP profile ↗
7ranked-venue papers in the field
2as first author
7since 2021 · last 2025
0000-0002-6136-3898ORCID · reported
Domains — venue-derived; a paper can count in several
Big Data, Cloud & Distributed Data Systems · 3 (1 first)Knowledge Engineering, Semantic Web & Information Systems · 3 (1 first)Data Mining & Knowledge Discovery · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | AraFinNews: Arabic Financial Summarisation with Domain-Adapted LLMs
Mo El-Haj, Paul Rayson |
IEEE Big Data | 1 |
| 2025 | Survey on legal information extraction: current status and open challengesabstractAbstract The goal of information extraction is to extract structural knowledge (such as entities, relations and events) from plain and unstructured texts. Information extraction in legal documents has recently gained a lot of attention in the natural language processing (NLP) community due to the high demand for efficient information extraction for legal practitioners and companies. Given that the legal documents are unique and their processing is challenging, there is a pressing need for applications of NLP techniques to tackle these challenges. In this research, we present a survey on the recent advancements in legal information extraction focusing on three tasks: named entity recognition, relationship extraction and event detection. We report language resources and systems in multiple jurisdictions and languages for each task. Based on the thorough review conducted, we identify insights into the techniques employed and promising research directions that merit further exploration in future studies. We maintain a public repository and consistently update related resources at https://github.com/DamithDR/legalinformationextraction . Damith Premasiri, Tharindu Ranasinghe, Ruslan Mitkov, Mo El-Haj, Ingo Frommholz |
Knowl. Inf. Syst. | 4 |
| 2023 | Advancements in Financial Document Structure Extraction: Insights from Five Years of FinTOC (2019-2023)abstractIn this comprehensive paper, we present a detailed overview of the Financial Table Of Content extraction shared task series, FinTOC, conducted over a span of five years from 2019 to 2023. This paper serves as a retrospective analysis of the key developments in the field of financial document structure extraction. The FinTOC series, hosted within the framework of the Financial Narrative Processing (FNP) workshop, has been instrumental in shaping the landscape of Natural Language Processing (NLP) in the financial domain. Our analysis delves into the diverse methodologies proposed by participants across all editions, shedding light on the innovative strategies employed to tackle the intricate challenge of extracting structured information from financial documents. We explore the evolution of techniques, from traditional rule-based approaches to cutting-edge deep learning models, showcasing the dynamic nature of NLP advancements. Furthermore, our study investigates the introduction of multilingual datasets by the organizers, highlighting the importance of cross-lingual analysis in financial document processing. We also examine the contributions made by participants in augmenting the training data with external sources, showcasing the collaborative spirit of the NLP community in enhancing the quality and size of the shared training dataset. Juyeon Kang, Mauli Mehulkumar Patel, Anushka Agrawal, Simhadri Sevitha, Srinivasa Ravi, Sandra Bellato, Anand Kumar Madasamy, Ngawang Dempa Tsang, Mo El-Haj |
IEEE Big Data | 9 |
| 2023 | The Financial Narrative Summarisation Shared Task (FNS 2023)abstractThis paper presents the results and findings of the Financial Narrative Summarisation Shared Task on summarising UK, Greek, and Spanish annual reports. The shared task was organised as part of the 5th Financial Narrative Processing Workshop (FNP 2023). The Financial Narrative summarisation Shared Task (FNS 2023) has been running since 2020 as part of the Financial Narrative Processing (FNP) workshop series [15–20]. The shared task included one main challenge, which is the use of either abstractive or extractive automatic summarisers to summarise long documents in terms of UK, Greek, and Spanish financial annual reports. This shared task is the fourth to target financial documents. The data for the shared task was created and collected from publicly available annual reports published by firms listed on the Stock Exchanges of the UK, Greece, and Spain. A total number of 6 systems from 3 different teams participated in the shared task. Elias Zavitsanos, Aris Kosmopoulos, George Giannakopoulos, Marina Litvak, Blanca Carbajo-Coronado, Antonio Moreno-Sandoval, Mo El-Haj |
IEEE Big Data | 7 |
| 2023 | Unifying Emotion Analysis Datasets using Valence Arousal Dominance (VAD)
Mo El-Haj, Ryutaro Takanami |
LDK | 1 |
| 2023 | Open-Source Thesaurus Development for Under-Resourced Languages: a Welsh Case Study
Nouran Khallaf, Elin Arfon, Mo El-Haj, Jonathan Morris, Dawn Knight, Paul Rayson, Tymaa Hammouda, Mustafa Jarrar |
LDK | 3 |
| 2023 | FinAraT5: A text to text model for financial Arabic text understanding and generation
Nadhem Zmandar, Mo El-Haj, Paul Rayson |
LDK | 2 |