Steven J. Lynden

dblp:l/SJLynden · DBLP profile ↗
← Back
28ranked-venue papers
8as first author
17since 2021 · last 2026
0000-0001-6642-6934ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Databases, data management, data science and information retrieval · 16 · 3 first-author · 13 since 2021Artificial intelligence and machine learning · 15 · 2 first-author · 12 since 2021Systems, architecture and hardware · 2 · 1 first-authorComputer networks · 1 · 1 first-authorSoftware engineering, systems software and programming languages · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Knowledge Graph Adapter-Based Augmentation Testbed for Large Language Models
Ushtar Ali, Steven J. Lynden, Akiyoshi Matono, Toshiyuki Amagasa
DEXA (1)2
2026 ACLBot: A Knowledge Graph-Driven Assistant for ACL Anthology Research
Jan Buchmann, Steven J. Lynden, Kristiina Jokinen
LREC2
2026 When structure predicts hallucination: Aligning LLMs with knowledge graph features
abstract
Large Language Models (LLMs) have demonstrated remarkable factual accuracy in producing human-like and AI-generated texts across a wide range of natural language tasks, including question answering. Despite these advances, their tendency to hallucinate and produce fabricated, false or incorrect responses is a persistent limitation. This limitation undermines their reliability and remains a critical challenge, especially in areas where high precision and trustworthiness are required. To address this challenge, we investigate whether the features derived from Knowledge Graphs (KGs) align with the accuracy of answers produced by the LLMs. In particular, we focus on entropy-based KG features, which capture diversity and uncertainty within structured knowledge. By analyzing the correlation between the entropy-based KG features and the accuracy of LLM responses, we are able to identify “blind spots” where LLMs are prone to hallucination. This provides insights not only into when an LLM is correct, but also into the conditions under which it fails. We present results across several datasets, including two developed for this study, demonstrating that entropy-based KG features can effectively align with the accuracy of LLM responses. Motivated by these findings, we propose a probing strategy for assessing LLM accuracy by focusing on areas where LLM accuracy is weak. The experimental results confirm that KG features can guide the probing effectively, highlighting the importance of using structured features from KGs in building more reliable and hallucination-free AI based systems.
Ushtar Ali, Steven J. Lynden, Akiyoshi Matono, Toshiyuki Amagasa
Data Knowl. Eng.2
2025 Entropy-Guided Probing for Predicting LLM Hallucinations with Knowledge Graph Features
Ushtar Ali, Steven J. Lynden, Akiyoshi Matono, Toshiyuki Amagasa
DEXA (1)2
2025 Action Sequence Analysis Using Temporal Commonsense Knowledge
Steven J. Lynden, Kyoung-Sook Kim 0001, Akiyoshi Matono, Hai-Tao Yu 0003, Xin Liu 0020
PAKDD (6)1
2025 AssistEM: Domain Instruction Tuning for Enhanced Entity Matching
John Bosco Mugeni, Steven J. Lynden, Toshiyuki Amagasa, Akiyoshi Matono
PAKDD (5)2
2025 How Useful Is Graph Pooling for Node-Level Tasks?
Yijun Duan, Xin Liu 0020, Steven J. Lynden, Akiyoshi Matono, Qiang Ma 0001
ECML/PKDD (3)3
2025 Estimating the plausibility of commonsense statements by novelly fusing large language model and graph neural network
Hai-Tao Yu 0003, Yijun Duan, Xin Liu 0020, Steven J. Lynden, Kyoung-Sook Kim 0001, Akiyoshi Matono, Adam Jatowt
Inf. Process. Manag.6
2025 Implicit knowledge-augmented prompting for commonsense explanation generation
abstract
Abstract Commonsense explanation generation refers to reasoning and explaining why a commonsense statement contradicts commonsense knowledge, such as why the statement “My dad grew volleyballs in his garden” is nonsensical. While such reasoning is trivial for humans, it remains a challenge for AI systems. Despite their notable performance in tasks like text generation and reasoning, large language models (LLMs) often fall short of consistently generating coherent and accurate commonsense explanations. To bridge this gap, we propose a novel Two-stage Identification and Prompting (TIP) framework for enhancing LLMs’ ability to handle the task of commonsense explanation generation. Specifically, in the first stage, TIP identifies the nonsensical concept in the given statement, pinpointing the specific element that contradicts commonsense knowledge. In the second stage, TIP generates implicit knowledge based on the identified nonsensical concept and then leverages this implicit knowledge to guide the adopted LLMs in generating explanations. In order to demonstrate the effectiveness of the proposed TIP framework for commonsense explanation generation, we conducted extensive experiments based on the ComVE dataset and a newly constructed CSE dataset, where a variety of LLMs are evaluated. The experimental results show that TIP consistently outperforms all baseline methods across multiple metrics, demonstrating its effectiveness in improving LLMs’ commonsense reasoning and explanation generation capabilities.
Hai-Tao Yu 0003, Xin Liu 0020, Adam Jatowt, Kyoung-Sook Kim 0001, Steven J. Lynden, Akiyoshi Matono
Knowl. Inf. Syst.7
2024 MultiMatch: Low-Resource Generalized Entity Matching Using Task-Conditioned Hyperadapters in Multitask Learning
John Bosco Mugeni, Steven J. Lynden, Toshiyuki Amagasa, Akiyoshi Matono
DaWaK2
2024 Inexact Graph Representation Learning
abstract
Graph is the universal language for modeling data across various domains. In recent years, graph representation learning has achieved outstanding performance in a series of computational tasks on graphs, such as graph node classification and community detection, etc. However, a commonly hidden assumption underlying these tasks is that label information corresponds to nodes in the training set on a one-to-one basis. In real-world scenarios, node label information may be uncertain and concealed within higher-order graph structural labels. In this paper, we define a set of arbitrary nodes on the graph as a bag and assume that label information corresponds to bags rather than individual nodes. Furthermore, bag labels are generated based on node-level labels through some hidden mechanism, thus inexactly encompassing node label information1. Therefore, we propose for the first time a novel and widely applicable task: learning the latent representation of bags on the graph and predicting their labels. For this task, we propose a hierarchical model that is highly interpretable and scalable, incorporating information propagation among nodes, information aggregation from nodes to bags and inter-bag relationship prediction. Experiments on diverse standard datasets demonstrate that our proposed model shows higher classification accuracy compared to strong baselines. Our research introduces a new perspective to the field of graph representation learning.
Yijun Duan, Xin Liu 0020, Adam Jatowt, Hai-Tao Yu 0003, Steven J. Lynden, Akiyoshi Matono
IJCNN5
2024 Semi-supervised Named Entity Recognition for Low-Resource Languages Using Dual PLMs
Mehari Yohannes Hailemariam, Steven J. Lynden, Toshiyuki Amagasa, Akiyoshi Matono
NLDB (1)2
2023 Commonsense Temporal Action Knowledge (CoTAK) Dataset
Steven J. Lynden, Mehari Yohannes Hailemariam, Kyoung-Sook Kim 0001, Adam Jatowt, Akiyoshi Matono, Hai-Tao Yu 0003, Xin Liu 0020, Yijun Duan
CIKM1
2023 What Wikipedia Misses About Yuriko Nakamura? Predicting Missing Biography Content by Learning Latent Life Patterns
abstract
Action-related KnowledGe (AKG) is important for facilitating deeper understanding of people’s life patterns, objectives and motivations. In this study, we present a novel framework for automatically predicting missing human biography records in Wikipedia by generating such knowledge. The generation method, which is based on a neural network matrix factorization model, is capable of encoding action semantics from diverse perspectives and discovering latent inter-action relations. By correctly predicting missing information and correcting errors, our work can effectively improve the quality of data about the behavioral records of historical figures in the knowledge base (e.g., biographies in Wikipedia), thus contributing to the understanding and study of human actions by the general public on the one hand, and can be considered as a new paradigm for managing action-related knowledge in digital libraries on the other. Extensive experiments demonstrate that the AKG we generate can capture well missing or “forgotten” human biography related information in Wikipedia.
Yijun Duan, Xin Liu 0020, Adam Jatowt, Chenyi Zhuang, Hai-Tao Yu 0003, Steven J. Lynden, Kyoung-Sook Kim 0001, Akiyoshi Matono
ECAI6
2023 AdapterEM: Pre-trained Language Model Adaptation for Generalized Entity Matching using Adapter-tuning
abstract
Entity Matching (EM) involves identifying different data representations referring to the same entity from multiple data sources and is typically formulated as a binary classification problem. It is a challenging problem in data integration due to the heterogeneity of data representations. State-of-the-art solutions have adopted NLP techniques based on pre-trained language models (PrLMs) via the fine-tuning paradigm, however, sequential fine-tuning of overparameterized PrLMs can lead to catastrophic forgetting, especially in low-resource scenarios. In this study, we propose a parameter-efficient paradigm for fine-tuning PrLMs based on adapters, small neural networks encapsulated between layers of a PrLM, by optimizing only the adapter and classifier weights while the PrLMs parameters are frozen. Adapter-based methods have been successfully applied to multilingual speech problems achieving promising results, however, the effectiveness of these methods when applied to EM is not yet well understood, particularly for generalized EM with heterogeneous data. Furthermore, we explore using (i) pre-trained adapters and (ii) invertible adapters to capture token-level language representations and demonstrate their benefits for transfer learning on the generalized EM benchmark. Our results show that our solution achieves comparable or superior performance to full-scale PrLM fine-tuning and prompt-tuning baselines while utilizing a significantly smaller computational footprint of the PrLM parameters.
John Bosco Mugeni, Steven J. Lynden, Toshiyuki Amagasa, Akiyoshi Matono
IDEAS2
2022 Dual Cost-sensitive Graph Convolutional Network
abstract
In graph node classification tasks, traditional graph neural network (GNN) models assume that different types of misclassification have equal loss and thus seek to maximize the posterior probability of sample nodes under labeled classes. However, the graph data in realistic scenarios tend to follow unbalanced long-tail class distributions, making it difficult for GNN to accurately represent the minority class samples because of the overfitting tendency to the majority class features. To address this problem, in this paper we propose a novel GNN model, named Dual Cost-sensitive Graph Convolutional Network (DCSGCN). The DCSGCN is a two-tower model containing two sub-networks that compute the posterior probability and the misclassification cost separately. It uses the cost as complementary information in classification to correct the posterior probability under the minimal risk perspective. Furthermore, we propose a series of new methods to compute node cost labels based on graph topological information and node class distribution. Extensive experiments lead to the observation that DCSGCN outperforms the state-of-the-art model on diverse real-world imbalanced graphs.
Yijun Duan, Xin Liu 0020, Adam Jatowt, Hai-Tao Yu 0003, Steven J. Lynden, Kyoung-Sook Kim 0001, Akiyoshi Matono
IJCNN5
2022 Anonymity can Help Minority: A Novel Synthetic Data Over-Sampling Strategy on Multi-label Graphs
Yijun Duan, Xin Liu 0020, Adam Jatowt, Hai-Tao Yu 0003, Steven J. Lynden, Kyoung-Sook Kim 0001, Akiyoshi Matono
ECML/PKDD (2)5
2018 Network Embedding Based on a Quasi-Local Similarity Measure
Xin Liu 0020, Natthawut Kertkeidkachorn, Tsuyoshi Murata, Kyoung-Sook Kim 0001, Julien Leblay, Steven J. Lynden
PRICAI (1)6
2017 Exploring the Veracity of Online Claims with BackDrop
abstract
Using the Web to assess the validity of claims presents many challenges. Whether the data comes from social networks or established media outlets, individual or institutional data publishers, one has to deal with scale and heterogeneity, as well as with incomplete, imprecise and sometimes outright false information. All of these are closely studied issues. Yet in many situations, the claims under scrutiny, and the data itself, have some inherent context-dependency making them impossible to completely disprove, or evaluate through a simple (e.g. scalar) measure. While data models used on the Web typically deal with universal knowledge, we believe the time has come to put context, such as time or provenance, at the forefront and watch knowledge through multiple lenses. We present BackDrop, an application that enables annotating knowledge and ontologies found online to explore how the veracity of claims varies with context. BackDrop comes in the form of a Web interface, in which users can interactively populate and annotate knowledge bases, and explore under which circumstances certain claims are more or less credible.
Julien Leblay, Steven J. Lynden
CIKM3
2011 Dynamic Data Redistribution for MapReduce Joins
abstract
MapReduce has become a popular method for data processing, in particular for large scale datasets, due to its accessibility as a scalable yet convenient programming paradigm. Data processing tasks often involve joins, and the repartition and fragment-replicate joins are two widely-used join algorithms utilised within the MapReduce framework. This paper presents a multi-join supporting tuple redistribution, building on both the repartition and fragment-replicate joins. Hadoop is used to demonstrate how reduce tasks may improve performance by passing intermediate results to other reduce tasks that are better able to process them using Apache ZooKeeper as a means of communication and data transfer. A performance analysis is presented showing the technique has the potential to reduce response times when processing multiple joins in single MapReduce jobs.
Steven J. Lynden, Yusuke Tanimura, Isao Kojima, Akiyoshi Matono
CloudCom1
2010 ADERIS: Adaptively Integrating RDF Data from SPARQL Endpoints
Steven J. Lynden, Isao Kojima, Akiyoshi Matono, Yusuke Tanimura
DASFAA (2)1
2010 Adaptive join processing in pipelined plans
abstract
In adaptive query processing, the way in which a query is evaluated is changed in the light of feedback obtained from the environment during query evaluation. Such feedback may, for example, establish that misleading selectivity estimates were used when the query was compiled, leading to the optimizer choosing an inappropriate join order or unsuitable join algorithms. This paper describes how joins can be reordered, and the join algorithms used replaced, while they are being evaluated in pipelined plans. Where joins are reordered and/or replaced during their evaluation, the approach avoids duplicating work that has already been carried out, by resuming from where the previous plan left off. The approach has been evaluated empirically, and shown to be effective for improving query performance in the light of misleading selectivity estimates.
Kwanchai Eurviriyanukul, Norman W. Paton, Alvaro A. A. Fernandes, Steven J. Lynden
EDBT4
2009 The design and implementation of OGSA-DQP: A service-based distributed query processor
Steven J. Lynden, Arijit Mukherjee, Alastair C. Hume, Alvaro A. A. Fernandes, Norman W. Paton, Rizos Sakellariou, Paul Watson 0001
Future Gener. Comput. Syst.1
2004 Dealing with Complex Networks of Process Interactions: A Security Measure
abstract
The majority of faults and consequent errors and failures in computer systems stem from the complexity of the system itself according to B. Schneier (2000). Yet complexity as a non-functional property is largely disregarded from initial development phases or at best is considered "manageable" by the process life cycle. Although at large the property is considered as having being conquered, valid sources inform us as stated in W. B. Kernighan (1987) that complexity as a property - borne by the system or emergent - is the cause of most failures. Furthermore the second most frequent case of failures is that of interaction with the system. In this paper we are looking at one aspect of process interactions. We are presenting a method for modelling and analyzing information system process interactions for the purpose of enhancing security. The paper introduces our work in the area of complex systems and gives a general introduction to our modelling methodology and its usage.
Panayiotis Periorellis, Olusola C. Idowu, Steven J. Lynden, Malcolm P. Young, Peter Andras 0001
ICECCS3
2002 Coordinating Learning Agents via Utility Assignment
Steven J. Lynden, Omer F. Rana
IDEAL1
2002 Coordinated Learning to Support Resource Management in Computational Grids
abstract
Managing resources in large scale distributed systems is an important concern for both peer-2-peer and computational grid systems, and is a complex and time sensitive process. Although existing peer-2-peer systems are divided into those that support computation (CPU) sharing or data sharing, users in a computational grid generally need to share both. Identifying which resources to select is important to guarantee reasonable execution time and cost to a given user or group of users. We provide first a description of a framework to support learning in the context a community of interacting peers, and show how this can be used to support resource sharing in computational grids.
Steven J. Lynden, Omer F. Rana
Peer-to-Peer Computing1
2000 Emergent coordination for distributed information management
abstract
With an increase in the complexity of information systems, decentralised management techniques are becoming important. Such information systems will generally involve various computational and data nodes, managed by various organisations with differing policies and needs. Furthermore, in the absence of a centralised manager, co-ordination and interaction policies between such nodes become crucial. Aspects of co-ordination can range from data formats describing shared information to convergence time to a desired goal. Each entity within such a networked environment operates with incomplete knowledge about the state of others, and makes use of an asynchronous mode of operation. An overview of techniques for supporting coordination in such an environment is first provided, followed by a proposed technique that makes use of Q-learning and data sharing between a collection of nodes, each of which is modelled as a neural network. We propose a decentralised technique for managing a cluster of nodes, where emergence is used to converge towards a 'useful' result.
Steven J. Lynden, Omer F. Rana, Steve Margetts, Antonia J. Jones
CEC1
2000 PaDDMAS: Parallel and Distributed Data Mining Application Suite
abstract
Discovering complex associations, anomalies and patterns in distributed data sets is gaining popularity in a range of scientific, medical and business applications. Various algorithms are employed to perform data analysis within a domain, and range from statistical to machine learning and AI based techniques. Several issues need to be addressed however to scale such approaches to large data sets, particularly when these are applied to data distributed at various sites. As new analysis techniques are identified, the core tool set must enable easy integration of such analytical components. Similarly, results from an analysis engines must be sharable, to enable storage, visualisation or further analysis of results. We describe the architecture of PaDDMAS, a component based system for developing distributed data mining applications. PaDDMAS provides a tool set for combining pre-developed or custom components using a dataflow approach, with components performing analysis, data extraction or data management and translation. Each component is wrapped as a Java/CORBA object, and has an interface defined in XML. Components can be serial or parallel objects, and may be binary or contain a more complex internal structure. We demonstrate a prototype using a neural network analysis algorithm.
Omer F. Rana, David W. Walker, Maozhen Li 0001, Steven J. Lynden, Mike Ward
IPDPS4