VLDB 2026 Research / reviewers in the wild / expert
Sean O'Brien
dblp:11/10553
· DBLP profile ↗
7ranked-venue papers
1as first author
7since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 4 since 2021Artificial intelligence and machine learning · 3 · 3 since 2021Software engineering, systems software and programming languages · 3 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | The Wave Platform a LCNC Platform for Communities of Research
Tiziana Margaria, Sean O'Brien, Marco Krumrey, Daniel Sami Mitwalli, Salim Saay, Sebastian Teumert |
COMPSAC | 2 |
| 2025 | Specialized Training in LCNC and AI: from the Pedagogical Concept to the ExperienceabstractTrain the trainer is a pedagogical model widely used across a range of disciplines and fields. Originating in non- governmental organisations, the model was developed in a bid to reduce costs associated with training while also increasing the capacity of staff to effectively disseminate knowledge within organisations. While its use is now widespread, there is little consensus in the literature regarding its efficacy. This paper describes a train-the-trainer programme that was designed to enhance early-career researchers’ ability to deliver training based on their research. The programme was delivered to a pilot group who then delivered training to external stakeholders on two occasions. The train-the-trainer programme was then evaluated from the perspectives of the trained early-career researchers along with the participants who subsequently took part in workshops delivered by this cohort. Results indicate that the programme was largely successful in improving the early-career researchers’ confidence with respect to training delivery, while also creating a positive learning experience for the participants. The paper concludes with a discussion around the future plans for the development of the train-the-trainer programme based on the results from this pilot implementation. Sean O'Brien, Colm Brandon, Daniel Busch, Marco Krumrey, Daniel Sami Mitwalli, Sebastian Teumert, Tiziana Margaria |
COMPSAC | 1 |
| 2025 | Diagrams as Visual Knowledge Communication Tools in Interdisciplinary Postgraduate EducationabstractThis research investigates the use of diverse visual knowledge communication tools in the multidisciplinary training produced in the BC4ECO project. During this project, teaching and learning content was developed to enable postgraduate learners to utilise blockchain and Distributed Ledger Technology (DLT) to solve complex, real-world problems related to environmental sustainability. We examine how visual modelling techniques enhance comprehension across disciplines and facilitate interdisciplinary and transdisciplinary collaboration, particularly for learners who may not have prior programming or computer science knowledge. Findings suggest that using these tools leads to enhanced collaboration between computer scientists and non-computer scientists and aids in the facilitation of joint problem solving between seemingly disparate disciplines. Additionally, the study highlights how the use of tools for established software engineering modelling, like Visual Paradigm for UML, bridges the gap between theoretical knowledge and practical implementation, strengthening problem solving skills and improving software development education. Salim Saay, Sean O'Brien, Amalia de Götzen, Viktoria Voronova, Tiziana Margaria |
COMPSAC | 2 |
| 2025 | NovelHopQA: Diagnosing Multi-Hop Reasoning Failures in Long Narrative ContextsabstractCurrent large language models (LLMs) struggle to answer questions that span tens of thousands of tokens, especially when multi-hop reasoning is involved.While prior benchmarks explore long-context comprehension or multihop reasoning in isolation, none jointly vary context length and reasoning depth in natural narrative settings.We introduce NOVEL-HOPQA, the first benchmark to evaluate 1-4 hop QA over 64k-128k-token excerpts from 83 full-length public-domain novels.A keywordguided pipeline builds hop-separated chains grounded in coherent storylines.We evaluate seven state-of-the-art (SOTA) models and apply oracle-context filtering to ensure all questions are genuinely answerable.Human annotators validate both alignment and hop depth.We additionally present retrieval-augmented generation (RAG) evaluations to test model performance when only selected passages are provided instead of the full context.We noticed consistent accuracy drops with increased hops and context length, even in frontier models-revealing that sheer scale does not guarantee robust reasoning.Our failure mode analysis highlights common breakdowns, such as missed final-hop integration and long-range drift.NOVELHOPQA offers a controlled diagnostic setting to test multi-hop reasoning at scale.All code and datasets are available at: https://novelhopqa.github.io. Abhay Gupta, Kevin Zhu, Vasu Sharma, Sean O'Brien, Michael Lu |
EMNLP | 4 |
| 2025 | Self-Updatable Large Language Models by Integrating Context into Model ParametersabstractDespite significant advancements in large language models (LLMs), the rapid and frequent integration of small-scale experiences, such as interactions with sur- rounding objects, remains a substantial challenge. Two critical factors in assimilating these experiences are (1) **Efficacy**: the ability to accurately remember recent events; (2) **Retention**: the capacity to recall long-past experiences. Current methods either embed experiences within model parameters using continual learning, model editing, or knowledge distillation techniques, which often struggle with rapid updates and complex interactions, or rely on external storage to achieve long-term retention, thereby increasing storage requirements. In this paper, we propose **SELF-PARAM** (Self-Updatable Large Language Models with Parameter Integration). SELF-PARAM requires no extra parameters while ensuring near-optimal efficacy and long-term retention. Our method employs a training objective that minimizes the Kullback-Leibler (KL) divergence between the predictions of an original model (with access to contextual information) and a target model (without such access). By generating diverse question-answer pairs related to the knowledge and minimizing the KL divergence across this dataset, we update the target model to internalize the knowledge seamlessly within its parameters. Evaluations on question-answering and conversational recommendation tasks demonstrate that SELF-PARAM significantly outperforms existing methods, even when accounting for non-zero storage requirements. This advancement paves the way for more efficient and scalable integration of experiences in large language models by embedding knowledge directly into model parameters. Yu Wang 0170, Xinshuang Liu, Xiusi Chen, Sean O'Brien, Junda Wu, Julian J. McAuley |
ICLR | 4 |
| 2025 | Disentangling Likes and Dislikes in Personalized Generative Explainable RecommendationabstractRecent research on explainable recommendation generally frames the task as a standard text generation problem, and evaluates models simply based on the textual similarity between the predicted and ground-truth explanations. However, this approach fails to consider one crucial aspect of the systems: whether their outputs accurately reflect the users' (post-purchase) sentiments, i.e., whether and why they would like and/or dislike the recommended items. To shed light on this issue, we introduce new datasets and evaluation methods that focus on the users' sentiments. Specifically, we construct the datasets by explicitly extracting users' positive and negative opinions from their post-purchase reviews using an LLM, and propose to evaluate systems based on whether the generated explanations 1) align well with the users' sentiments, and 2) accurately identify both positive and negative opinions of users on the target items. We benchmark several recent models on our datasets and demonstrate that achieving strong performance on existing metrics does not ensure that the generated explanations align well with the users' sentiments. Lastly, we find that existing models can provide more sentiment-aware explanations when the users' (predicted) ratings for the target items are directly fed into the models as input. The datasets and benchmark implementation are available at: https://github.com/jchanxtarov/sent_xrec. Ryotaro Shimizu, Takashi Wada 0001, Yu Wang 0170, Johannes Kruse 0002, Sean O'Brien, Sai Htaung Kham, Linxin Song, Yuya Yoshikawa, Yuki Saito 0002, Fugee Tsung, Masayuki Goto, Julian J. McAuley |
WWW | 5 |
| 2024 | Linear Layer Extrapolation for Fine-Grained Emotion ClassificationabstractCertain abilities of Transformer-based language models consistently emerge in their later layers.Previous research has leveraged this phenomenon to improve factual accuracy through self-contrast, penalizing early-exit predictions based on the premise that later-layer updates are more factually reliable than earlierlayer associations.We observe a similar pattern for fine-grained emotion classification in text, demonstrating that self-contrast can enhance encoder-based text classifiers.Additionally, we reinterpret self-contrast as a form of linear extrapolation, which motivates a refined approach that dynamically adjusts the contrastive strength based on the selected intermediate layer.Experiments across multiple models and emotion classification datasets show that our method outperforms standard classification techniques in fine-grained emotion classification tasks. Mayukh Sharma, Sean O'Brien, Julian J. McAuley |
EMNLP | 2 |