EDBT 2026 Demo / reviewers in the wild / expert
Asep Fajar Firmansyah
dblp:194/8627
· DBLP profile ↗
9ranked-venue papers in the field
3as first author
7since 2021 · last 2026
0000-0001-5775-6966ORCID · corroborated
Domains — venue-derived; a paper can count in several
Information Retrieval & Web Search · 4Knowledge Engineering, Semantic Web & Information Systems · 4 (3 first)Database Systems & Data Management · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | ELEVATE-ID: Extending Large Language Models for End-to-End Entity Linking Evaluation in IndonesianabstractLarge Language Models (LLMs) have demonstrated remarkable performance across a wide range of natural language processing tasks. However, their effectiveness in low-resource languages remains underexplored, particularly in complex tasks such as end-to-end Entity Linking (EL), which requires both mention detection and disambiguation against a knowledge base (KB). In earlier work, we introduced IndEL — the first end-to-end EL benchmark dataset for the Indonesian language — covering both a general domain (news) and a specific domain (religious text from the Indonesian translation of the Quran), and evaluated four traditional end-to-end EL systems on this dataset. In this study, we propose ELEVATE-ID, a comprehensive evaluation framework for assessing LLM performance on end-to-end EL in Indonesian. The framework evaluates LLMs under both zero-shot and fine-tuned conditions, using multilingual and Indonesian monolingual models, with Wikidata as the target KB. Our experiments include performance benchmarking, generalization analysis across domains, and systematic error analysis. Results show that GPT-4 and GPT-3.5 achieve the highest accuracy in zero-shot and fine-tuned settings, respectively. However, even fine-tuned GPT-3.5 underperforms compared to DBpedia Spotlight — the weakest of the traditional model baselines — in the general domain. Interestingly, GPT-3.5 outperforms Babelfy in the specific domain. Generalization analysis indicates that fine-tuned GPT-3.5 adapts more effectively to cross-domain and mixed-domain scenarios. Error analysis uncovers persistent challenges that hinder LLM performance: difficulties with non-complete mentions, acronym disambiguation, and full-name recognition in formal contexts. These issues point to limitations in mention boundary detection and contextual grounding. Indonesian-pretrained LLMs, Komodo and Merak, reveal core weaknesses: template leakage and entity hallucination, respectively—underscoring architectural and training limitations in low-resource end-to-end EL. 1 Ria Hari Gusmita, Asep Fajar Firmansyah, Hamada M. Zahera, Axel-Cyrille Ngonga Ngomo |
Data Knowl. Eng. | 2 |
| 2025 | ANTS: Abstractive Entity Summarization in Knowledge Graphs
Asep Fajar Firmansyah, Hamada M. Zahera, Mohamed Ahmed Sherif, Diego Moussallem, Axel-Cyrille Ngonga Ngomo |
ESWC (1) | 1 |
| 2025 | NL2LS: LLM-based Automatic Linking of Knowledge GraphsabstractIntegrated knowledge graphs form the foundation of numerous data-driven applications, including search engines, conversational agents, and e-commerce solutions. Declarative link discovery frameworks utilize link specifications to define the conditions necessary for establishing a link between knowledge graphs’ resources. Despite domain expertise, defining such link specifications remains challenging due to their intricate syntax, threshold tuning, and the need to precisely express complex linking logic. To address this challenge, we propose NL2LS, a novel language-driven approach that leverages large language models to automatically translate natural language (NL) into link specifications (LSs), enabling domain experts and practitioners to express correct and complex linking rules more effectively. NL2LS employs three distinct training paradigms to handle the complexity of link specifications: zero-shot learning, one-shot learning and supervised fine-tuning. We evaluated NL2LS using different large language model architectures in comparison with a rule-based baseline model on different multi-lingual datasets. Our evaluation using BLEU, METEOR, ChrF++, and TER metrics demonstrates that NL2LS effectively translates natural language into link specifications, lowering the technical barrier and assisting users in specifying link rules more intuitively. Reda Ihtassine, Asep Fajar Firmansyah, Nikit Srivastava, Manzoor Ali, Axel-Cyrille Ngonga Ngomo, Mohamed Ahmed Sherif |
K-CAP | 2 |
| 2024 | ESLM: Improving Entity Summarization by Leveraging Language Models
Asep Fajar Firmansyah, Diego Moussallem, Axel-Cyrille Ngonga Ngomo |
ESWC (1) | 1 |
| 2023 | Explainable Integration of Knowledge Graphs Using Large Language Models
Abdullah Fathi Ahmed, Asep Fajar Firmansyah, Mohamed Ahmed Sherif, Diego Moussallem, Axel-Cyrille Ngonga Ngomo |
NLDB | 2 |
| 2023 | IndQNER: Named Entity Recognition Benchmark Dataset from the Indonesian Translation of the Quran
Ria Hari Gusmita, Asep Fajar Firmansyah, Diego Moussallem, Axel-Cyrille Ngonga Ngomo |
NLDB | 2 |
| 2021 | GATES: Using Graph Attention Networks for Entity SummarizationabstractThe sheer size of modern knowledge graphs has led to increased attention being paid to the entity summarization task. Given a knowledge graph T and an entity e found therein, solutions to entity summarization select a subset of the triples from T which summarize e's concise bound description. Presently, the best performing approaches rely on sequence-to-sequence models to generate entity summaries and use little to none of the structure information of T during the summarization process. We hypothesize that this structure information can be exploited to compute better summaries. To verify our hypothesis, we propose GATES, a new entity summarization approach that combines topological information and knowledge graph embeddings to encode triples. The topological information is encoded by means of a Graph Attention Network. Furthermore, ensemble learning is applied to boost the performance of triple scoring. We evaluate GATES on the DBpedia and LMDB datasets from ESBM (version 1.2), as well as on the FACES datasets. Our results show that GATES outperforms the state-of-the-art approaches on 4 of 6 configuration settings and reaches up to 0.574 F-measure. Pertaining to resulted summaries quality, GATES still underperforms the state of the arts as it obtains the highest score only on 1 of 6 configuration settings at 0.697 NDCG score. An open-source implementation of our approach and of the code necessary to rerun our experiments are available at https://github.com/dice-group/GATES. Asep Fajar Firmansyah, Diego Moussallem, Axel-Cyrille Ngonga Ngomo |
K-CAP | 1 |
| 2018 | Analysis of Study Program Selection Patterns Using FP-Growth and ECLAT MethodsabstractA set of data that correlates form patterns. In determining the patterns of data sets using association rules. The University has a variety of faculties that have several study programs. Many factors that influence prospective students determining the choice of their study program. Researchers consider it important to look for patterns of prospective students in determining their study program choices. With these patterns can help prospective students determining their choices. Forming these patterns uses the association rules with the FP-Growth and ECLAT methods. Researchers conducted a survey to students from various universities on the island of Java, Indonesia. Of the 35 transaction data obtained, the support minimum value is 20% and the confidence value is 45%. There are some association rules obtained, namely 12 rules using FP-Growth algorithm and 9 rules using ECLAT algorithm. Of these rules, interest factors, accreditation, and job prospects are the main factors influencing the choice of study programs. Nurbojatmiko, Eri Rustamaji, Asep Fajar Firmansyah |
iiWAS | 3 |
| 2016 | Generating weighted vector for concepts in indonesian translation of QuranabstractThis paper presents a work in generating Weighted Vector for each Concept in Indonesian Translation of Quran (ITQ). This task is done in aiming to provide a resource needed in implementing a semantic-based question answering system (QAS) for Indonesian ITQ, particularly in retrieving semantically related verses. Semantic approach on QAS employs Ontology concepts of the domain. Since there is no Ontology for ITQ remains, we built one by utilizing the existing Ontology from Quranic Arabic corpus (http://corpus.quran.com/). Furthermore, each leaf concept that enriched by related Quran verse (as its instance) had a representation vector of terms that occur in the corresponding Quran verse to express how strength the concept in relates with verse terms. This vector is assigned with a weight resulted from applying TFIDF method. From 222 leaf concepts in the Ontology, we applied the process only to those that categorized as a member group of Person, Location, and Time named entity. They are 107 in a total. The result shows that the most strength concept in association with verse terms is syaitan which is scored at 0.895 of 1. In overall, 16.82 % concepts had score that more than 0.4, following by 14.95%, 23.36% and 11.21% concepts scored at more than 0.3 ,0.2 and less than 0.1 respectively, and finally the rest ones were the biggest in volume where 33.64% concepts obtained score more than 0.1 and less than 0.2. Syopiansyah Jaya Putra, Khodijah Hulliyah, Nashrul Hakiem, Rayi Pradono Iswara, Asep Fajar Firmansyah |
iiWAS | 5 |