VLDB 2026 Research / reviewers in the wild / expert
Sanju Mishra
dblp:247/4827 · also Sanju Mishra Tiwari, Sanju Tiwari
· DBLP profile ↗
7ranked-venue papers
2as first author
5since 2021 · last 2025
0000-0001-7197-0766ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | EcoRAG: A Multi-hop Economic QA Benchmark for Retrieval Augmented Generation Using Knowledge Graphs
Hanieh Khorashadizadeh, Sanju Mishra, Farah Benamara, Nandana Mihindukulasooriya, Jinghua Groppe, Soror Sahri, Morteza Kamaladdini Ezzabady, Frédéric Ieng, Sven Groppe |
NLDB (2) | 2 |
| 2024 | Scholarly Wikidata: Population and Exploration of Conference Data in Wikidata Using LLMs
Nandana Mihindukulasooriya, Sanju Mishra, Daniil Dobriy, Finn Årup Nielsen, Tek Raj Chhetri, Axel Polleres |
EKAW | 2 |
| 2024 | Attention based hybrid deep learning model for wearable based stress recognition
Ritu Tanwar, Orchid Chetia Phukan, Ghanapriya Singh, Pankaj Kumar Pal, Sanju Mishra |
Eng. Appl. Artif. Intell. | 5 |
| 2023 | Text2KGBench: A Benchmark for Ontology-Driven Knowledge Graph Generation from TextabstractThe recent advances in large language models (LLM) and foundation models with emergent capabilities have been shown to improve the performance of many NLP tasks. LLMs and Knowledge Graphs (KG) can complement each other such that LLMs can be used for KG construction or completion while existing KGs can be used for different tasks such as making LLM outputs explainable or fact-checking in Neuro-Symbolic manner. In this paper, we present Text2KGBench, a benchmark to evaluate the capabilities of language models to generate KGs from natural language text guided by an ontology. Given an input ontology and a set of sentences, the task is to extract facts from the text while complying with the given ontology (concepts, relations, domain/range constraints) and being faithful to the input sentences. We provide two datasets (i) Wikidata-TekGen with 10 ontologies and 13,474 sentences and (ii) DBpedia-WebNLG with 19 ontologies and 4,860 sentences. We define seven evaluation metrics to measure fact extraction performance, ontology conformance, and hallucinations by LLMs. Furthermore, we provide results for two baseline models, Vicuna-13B and Alpaca-LoRA-13B using automatic prompt generation from test cases. The baseline results show that there is room for improvement using both Semantic Web and Natural Language Processing techniques. Resource Type: Evaluation Benchmark Source Repo: https://github.com/cenguix/Text2KGBench DOI: https://doi.org/10.5281/zenodo.7916716 License: Creative Commons Attribution (CC BY 4.0) Nandana Mihindukulasooriya, Sanju Mishra, Carlos F. Enguix, Kusum Lata 0002 |
ISWC | 2 |
| 2021 | Recent trends in knowledge graphs: theory and practice
Sanju Mishra, Fatima N. Al-Aswadi, Devottam Gaurav |
Soft Comput. | 1 |
| 2020 | Machine intelligence-based algorithms for spam filtering on document labeling
Devottam Gaurav, Sanju Mishra, Ayush Goyal, Niketa Gandhi, Ajith Abraham |
Soft Comput. | 2 |
| 2019 | Secure Semantic Smart HealthCare (S3HC)abstractHealthcare is a significant domain having a huge knowledge base, a significant part which comes from medical, diagnostic and imaging devices and sensors.The health status of patients may be monitored and managed remotely by performing reasoning over this knowledge base.Specialists in HealthCare facilities are required to handle large quantity of data generated and make decisions.However, the heterogeneous and complex nature and the huge amount of data generated; the way it is represented and presented; and the security challenges may overburden the core abilities of thinking and reasoning of even highly skilled and knowledgeable experts putting the lives of patients at risk.The situation may become even worse when data is coming from various healthcare devices and sensors which are themselves characterized by a number of representation and serialization formats.To address the various challenges in healthcare, this paper tries to represent and hence exchange the data collected by healthcare devices meaningfully and securely.This allows all healthcare devices to operate in conjunction with each other facilitating deeper insights and enabling generation of intelligent recommendations. Sanju Mishra, Sarika Jain 0001, Ajith Abraham, Smita Shandilya |
J. Web Eng. | 1 |