VLDB 2026 Research / reviewers in the wild / expert
Lei Shi 0032
dblp:29/563-32
· DBLP profile ↗
12ranked-venue papers
1as first author
12since 2021 · last 2025
0000-0003-1628-1559ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 6 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 3 · 3 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ActionDiffusion: An Action-Aware Diffusion Model for Procedure Planning in Instructional VideosabstractWe present ActionDiffusion - a novel diffusion model for procedure planning in instructional videos that is the first to take temporal inter-dependencies between actions into account. Our approach is in stark contrast to existing methods that fail to exploit the rich information content available in the particular order in which actions are performed. Our method unifies the learning of temporal dependencies between actions and denoising of the action plan in the diffusion process by projecting the action in-formation into the noise space. This is achieved 1) by adding action embeddings in the noise masks in the noise-adding phase and 2) by introducing an attention mecha-nism in the noise prediction network to learn the corre-lations between different action steps. We report exten-sive experiments on three instructional video benchmark datasets (CrossTask, Coin, and NIV) and show that our method outperforms previous state-of-the-art methods on all metrics on CrossTask and NIV and all metrics except accuracy on Coin dataset. We show that by adding action embeddings into the noise mask the diffusion model can better learn action temporal dependencies and increase the performances on procedure planning. Codes are available at https://www.collaborative-ai.org/publications/shi25_wacv/ Lei Shi 0032, Paul C. Bürkner, Andreas Bulling |
WACV | 1 |
| 2024 | Neural Reasoning about Agents' Goals, Preferences, and ActionsabstractWe propose the Intuitive Reasoning Network (IRENE) - a novel neural model for intuitive psychological reasoning about agents' goals, preferences, and actions that can generalise previous experiences to new situations. IRENE combines a graph neural network for learning agent and world state representations with a transformer to encode the task context. When evaluated on the challenging Baby Intuitions Benchmark, IRENE achieves new state-of-the-art performance on three out of its five tasks - with up to 48.9% improvement. In contrast to existing methods, IRENE is able to bind preferences to specific agents, to better distinguish between rational and irrational agents, and to better understand the role of blocking obstacles. We also investigate, for the first time, the influence of the training tasks on test performance. Our analyses demonstrate the effectiveness of IRENE in combining prior knowledge gained during training for unseen evaluation tasks. Matteo Bortoletto, Lei Shi 0032, Andreas Bulling |
AAAI | 2 |
| 2024 | Limits of Theory of Mind Modelling in Dialogue-Based Collaborative Plan AcquisitionabstractMatteo Bortoletto, Constantin Ruhdorfer, Adnen Abdessaied, Lei Shi, Andreas Bulling. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Matteo Bortoletto, Constantin Ruhdorfer, Adnen Abdessaied, Lei Shi 0032, Andreas Bulling |
ACL (1) | 4 |
| 2024 | VSA4VQA: Scaling A Vector Symbolic Architecture To Visual Question Answering on Natural Images
Anna Penzkofer, Lei Shi 0032, Andreas Bulling |
CogSci | 2 |
| 2024 | Explicit Modelling of Theory of Mind for Belief Prediction in Nonverbal Social InteractionsabstractWe propose MToMnet – a Theory of Mind (ToM) neural network for predicting beliefs and their dynamics during human social interactions from multimodal input. ToM is key for effective nonverbal human communication and collaboration, yet existing methods for belief modelling have not included explicit ToM modelling or have typically been limited to one or two modalities. MToMnet encodes contextual cues (scene videos and object locations) and integrates them with person-specific cues (human gaze and body language) in a separate MindNet for each person. Inspired by prior research on social cognition and computational ToM, we propose three different MToMnet variants: two involving the fusion of latent representations and one involving the re-ranking of classification scores. We evaluate our approach on two challenging real-world datasets, one focusing on belief prediction while the other examining belief dynamics prediction. Our results demonstrate that MToMnet surpasses existing methods by a large margin while at the same time requiring a significantly smaller number of parameters. Taken together, our method opens up a highly promising direction for future work on artificial intelligent systems that can robustly predict human beliefs from their non-verbal behaviour and, as such, more effectively collaborate with humans. Matteo Bortoletto, Constantin Ruhdorfer, Lei Shi 0032, Andreas Bulling |
ECAI | 3 |
| 2024 | Multi-modal Video Dialog State Tracking in the Wild
Adnen Abdessaied, Lei Shi 0032, Andreas Bulling |
ECCV (57) | 2 |
| 2024 | Explaining Disagreement in Visual Question Answering Using Eye TrackingabstractWhen presented with the same question about an image, human annotators often give valid but disagreeing answers indicating that their reasoning was different. Such differences are lost in a single ground truth label used to train and evaluate visual question answering (VQA) methods. In this work, we explore whether visual attention maps, created using stationary eye tracking, provide insight into the reasoning underlying disagreement in VQA. We first manually inspect attention maps in the recent VQA-MHUG dataset and find cases in which attention differs consistently for disagreeing answers. We further evaluate the suitability of four different similarity metrics to detect attention differences matching the disagreement. We show that attention maps plausibly surface differences in reasoning underlying one type of disagreement, and that the metrics complementarily detect them. Taken together, our results represent an important first step to leverage eye-tracking to explain disagreement in VQA. Susanne Hindennach, Lei Shi 0032, Andreas Bulling |
ETRA | 2 |
| 2024 | VD-GR: Boosting Visual Dialog with Cascaded Spatial-Temporal Multi-Modal GRaphsabstractWe propose $\mathbb{V}\mathbb{D}{\text{ - }}\mathbb{G}\mathbb{R}$ – a novel visual dialog model that combines pre-trained language models (LMs) with graph neural networks (GNNs). Prior works mainly focused on one class of models at the expense of the other, thus missing out on the opportunity of combining their respective benefits. At the core of $\mathbb{V}\mathbb{D}{\text{ - }}\mathbb{G}\mathbb{R}$ is a novel integration mechanism that alternates between spatial-temporal multi-modal GNNs and BERT layers, and that covers three distinct contributions: First, we use multi-modal GNNs to process the features of each modality (image, question, and dialog history) and exploit their local structures before performing BERT global attention. Second, we propose hub-nodes that link to all other nodes within one modality graph, allowing the model to propagate information from one GNN (modality) to the other in a cascaded manner. Third, we augment the BERT hidden states with fine-grained multi-modal GNN features before passing them to the next $\mathbb{V}\mathbb{D}{\text{ - }}\mathbb{G}\mathbb{R}$ layer. Evaluations on VisDial v1.0, VisDial v0.9, VisDialConv, and VisPro show that $\mathbb{V}\mathbb{D}{\text{ - }}\mathbb{G}\mathbb{R}$ achieves new state-of-the-art results across all four datasets. Adnen Abdessaied, Lei Shi 0032, Andreas Bulling |
WACV | 2 |
| 2024 | Mindful Explanations: Prevalence and Impact of Mind Attribution in XAI ResearchabstractWhen users perceive AI systems as mindful, independent agents, they hold them responsible instead of the AI experts who created and designed these systems. So far, it has not been studied whether explanations support this shift in responsibility through the use of mind-attributing verbs like "to think". To better understand the prevalence of mind-attributing explanations we analyse AI explanations in 3,533 explainable AI (XAI) research articles from the Semantic Scholar Open Research Corpus (S2ORC). Using methods from semantic shift detection, we identify three dominant types of mind attribution: (1) metaphorical (e.g. "to learn" or "to predict"), (2) awareness (e.g. "to consider"), and (3) agency (e.g. "to make decisions"). We then analyse the impact of mind-attributing explanations on awareness and responsibility in a vignette-based experiment with 199 participants. We find that participants who were given a mind-attributing explanation were more likely to rate the AI system as aware of the harm it caused. Moreover, the mind-attributing explanation had a responsibility-concealing effect: Considering the AI experts' involvement lead to reduced ratings of AI responsibility for participants who were given a non-mind-attributing or no explanation. In contrast, participants who read the mind-attributing explanation still held the AI system responsible despite considering the AI experts' involvement. Taken together, our work underlines the need to carefully phrase explanations about AI systems in scientific writing to reduce mind attribution and clearly communicate human responsibility. Susanne Hindennach, Lei Shi 0032, Filip Miletic 0002, Andreas Bulling |
Proc. ACM Hum. Comput. Interact. | 2 |
| 2023 | Improving neural saliency prediction with a cognitive model of human visual attention
Ekta Sood, Lei Shi 0032, Matteo Bortoletto, Yao Wang 0018, Philipp Müller 0001, Andreas Bulling |
CogSci | 2 |
| 2023 | Exploring Natural Language Processing Methods for Interactive Behaviour Modelling
Matteo Bortoletto, Zhiming Hu 0003, Lei Shi 0032, Mihai Bâce, Andreas Bulling |
INTERACT (3) | 4 |
| 2022 | Comparison of Deep Learning Models in Position Based Visual ServoingabstractIn this paper a Position Based Visual Servoing (PBVS) algorithm using deep neural network is applied to a UR10 cobot. A pre-trained Convolutional Neural Network (CNN) will be re-purposed and fine-tuned in an offline stage. In order implement and validate the CNN for a visual servoing application, a dataset was created in a simulated environment by moving a simulated UR10 robot to various positions and capturing an image with the corresponding relative pose to the target object. The obtained dataset was validated with ground truth dataset collected using the real robot. To control the motion of the cobot (simulated/real) a meta-operating system and a vision based control law was designed in ROS. The visual servoing task is defined as a repositioning task whereby the performance of the visual servoing is evaluated by the convergence to one desired pose from various, arbitrarily selected starting poses. Three different network architectures were implemented and tested. The obtained results reveal that all network architecture can be successfully applied to visual servoing systems. Cosmin Copot, Lei Shi 0032, Elke Smet, Clara M. Ionescu, Steve Vanlanduit |
ETFA | 2 |