VLDB 2026 Research / reviewers in the wild / expert
Stuart James
dblp:19/10673
· DBLP profile ↗
31ranked-venue papers
2as first author
19since 2021 · last 2025
0000-0002-2649-2133ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 22 · 2 first-author · 13 since 2021Artificial intelligence and machine learning · 12 · 1 first-author · 8 since 2021Human-computer interaction and ubiquitous computing · 10 · 7 since 2021Systems, architecture and hardware · 2 · 1 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | ReassembleNet: Learnable Keypoints and Diffusion for 2D Fresco ReconstructionabstractThe task of reassembly is a significant challenge across multiple domains, including archaeology, genomics, and molecular docking, requiring the precise placement and orientation of elements to reconstruct an original structure. In this work, we address key limitations in state-of-the-art Deep Learning methods for reassembly, namely i) scalability; ii) multimodality; and iii) real-world applicability: beyond square or simple geometric shapes, realistic and complex erosion, or other real-world problems. We propose ReassembleNet, a method that reduces complexity by representing each input piece as a set of contour keypoints and learning to select the most informative ones by Graph Neural Networks pooling inspired techniques. ReassembleNet effectively lowers computational complexity while enabling the integration of features from multiple modalities, including both geometric and texture data. Further enhanced through pretraining on a semi-synthetic dataset. We then apply diffusion-based pose estimation to recover the original structure. We improve on prior methods by 57% and 87% for RMSE Rotation and Translation, respectively. Adeela Islam, Stefano Fiorini, Stuart James, Pietro Morerio, Alessio Del Bue |
ICCV | 3 |
| 2025 | GANzzle++: Generative approaches for jigsaw puzzle solving as local to global assignment in latent spatial representationsabstractJigsaw puzzles are a popular and enjoyable pastime that humans can easily solve, even with many pieces. However, solving a jigsaw is a combinatorial problem, and the space of possible solutions is exponential in the number of pieces, intractable for pairwise solutions. In contrast to the classical pairwise local matching of pieces based on edge heuristics, we estimate an approximate solution image, i.e., a mental image , of the puzzle and exploit it to guide the placement of pieces as a piece-to-global assignment problem. Therefore, from unordered pieces, we consider conditioned generation approaches, including Generative Adversarial Networks (GAN) models, Slot Attention (SA) and Vision Transformers (ViT), to recover the solution image. Given the generated solution representation, we cast the jigsaw solving as a 1-to-1 assignment matching problem using Hungarian attention, which places pieces in corresponding positions in the global solution estimate. Results show that the newly proposed GANzzle-SA and GANzzle-VIT benefit from the early fusion strategy where pieces are jointly compressed and gathered for global structure recovery. A single deep learning model generalizes to puzzles of different sizes and improves the performances by a large margin. Evaluated on PuzzleCelebA and PuzzleWikiArts, our approaches bridge the gap of deep learning strategies with respect to optimization-based classic puzzle solvers. • We present new generative modules for estimating the jigsaw solution image. • We show the effect of estimating the target image for placement of pieces. • We evaluate on open datasets showing a large margin improvement. Davide Talon, Alessio Del Bue, Stuart James |
Pattern Recognit. Lett. | 3 |
| 2024 | PRAGO: Differentiable Multi-View Pose Optimization From Objectness DetectionsabstractRobustly estimating camera poses from a set of images is a fundamental task which remains challenging for differentiable methods, especially in the case of small and sparse camera pose graphs. To overcome this challenge, we propose Pose-refined Rotation Averaging Graph Optimization (PRAGO). From a set of objectness detections on unordered images, our method reconstructs the rotational pose, and in turn, the absolute pose, in a differentiable manner benefiting from the optimization of a sequence of geometrical tasks. We show how our objectness pose-refinement module in PRAGO is able to refine the inherent ambiguities in pairwise relative pose estimation without removing edges and avoiding making early decisions on the viability of graph edges. PRAGO then refines the absolute rotations through iterative graph construction, reweighting the graph edges to compute the final rotational pose, which can be converted into absolute poses using translation averaging. We show that PRAGO is able to outperform non-differentiable solvers on small and sparse scenes extracted from 7-Scenes achieving a relative improvement of 21% for rotations while achieving similar translation estimates. Matteo Taiana, Matteo Toso, Stuart James, Alessio Del Bue |
3DV | 3 |
| 2024 | 6DGS: 6D Pose Estimation from a Single Image and a 3D Gaussian Splatting Model
Matteo Bortolon, Theodore Tsesmelis, Stuart James, Fabio Poiesi, Alessio Del Bue |
ECCV (52) | 3 |
| 2024 | Interactive Digital Storytelling Navigating the Inherent Currents of the Diasporic Mind
Valentina Nisi, Paulo Bala, Miguel Pessoa, Stuart James, Nuno Nunes 0001 |
ICIDS (1) | 4 |
| 2024 | IFFNeRF: Initialisation Free and Fast 6DoF pose estimation from a single image and a NeRF modelabstractWe introduce IFFNeRF to estimate the six degrees-of-freedom (6DoF) camera pose of a given image, building on the Neural Radiance Fields (NeRF) formulation. IFFNeRF is specifically designed to operate in real-time and eliminates the need for an initial pose guess that is proximate to the sought solution. IFFNeRF utilizes the Metropolis-Hasting algorithm to sample surface points from within the NeRF model. From these sampled points, we cast rays and deduce the color for each ray through pixel-level view synthesis. The camera pose can then be estimated as the solution to a Least Squares problem by selecting correspondences between the query image and the resulting bundle. We facilitate this process through a learned attention mechanism, bridging the query image embedding with the embedding of parameterized rays, thereby matching rays pertinent to the image. Through synthetic and real evaluation settings, we show that our method can improve the angular and translation error accuracy by 80.1% and 67.3%, respectively, compared to iNeRF while performing at 34fps on consumer hardware and not requiring the initial pose guess. Project page: https://mbortolon97.github.io/frenerf/ Matteo Bortolon, Theodore Tsesmelis, Stuart James, Fabio Poiesi, Alessio Del Bue |
ICRA | 3 |
| 2024 | ArtAI4DS: AI Art and Its Empowering Role in Digital Storytelling
Teresa Fernandes, Valentina Nisi, Nuno Nunes 0001, Stuart James |
ICEC | 4 |
| 2024 | Re-assembling the past: The RePAIR dataset and benchmark for real world 2D and 3D puzzle solvingabstractThis paper proposes the RePAIR dataset that represents a challenging benchmark to test modern computational and data driven methods for puzzle-solving and reassembly tasks. Our dataset has unique properties that are uncommon to current benchmarks for 2D and 3D puzzle solving. The fragments and fractures are realistic, caused by a collapse of a fresco during a World War II bombing at the Pompeii archaeological park. The fragments are also eroded and have missing pieces with irregular shapes and different dimensions, challenging further the reassembly algorithms. The dataset is multi-modal providing high resolution images with characteristic pictorial elements, detailed 3D scans of the fragments and meta-data annotated by the archaeologists. Ground truth has been generated through several years of unceasing fieldwork, including the excavation and cleaning of each fragment, followed by manual puzzle solving by archaeologists of a subset of approx. 1000 pieces among the 16000 available. After digitizing all the fragments in 3D, a benchmark was prepared to challenge current reassembly and puzzle-solving methods that often solve more simplistic synthetic scenarios. The tested baselines show that there clearly exists a gap to fill in solving this computationally complex problem. Theodore Tsesmelis, Luca Palmieri 0002, Marina Khoroshiltseva, Adeela Islam, Gur Elkin, Ofir Itzhak Shahar, Gianluca Scarpellini, Stefano Fiorini, Yaniv Ohayon, Nadav Alali, Sinem Aslan, Pietro Morerio, Sebastiano Vascon, Elena Gravina, Maria Cristina Napolitano, Giuseppe Scarpati, Gabriel Zuchtriegel, Alexandra Spühler, Michel E. Fuchs, Stuart James, Ohad Ben-Shahar, Marcello Pelillo, Alessio Del Bue |
NeurIPS | 20 |
| 2024 | Positional diffusion: Graph-based diffusion models for set orderingabstractPositional reasoning is the process of ordering an unsorted set of parts into a consistent structure. To address this problem, we present Positional Diffusion , a plug-and-play graph formulation with Diffusion Probabilistic Models. Using a diffusion process, we add Gaussian noise to the set elements’ position and map them to a random position in a continuous space. Positional Diffusion learns to reverse the noising process and recover the original positions through an Attention-based Graph Neural Network. To evaluate our method, we conduct extensive experiments on three different tasks and seven datasets, comparing our approach against the state-of-the-art methods for visual puzzle-solving, sentence ordering, and room arrangement, demonstrating that our method outperforms long-lasting research on puzzle solving with up to + 17 % compared to the second-best deep learning method, and performs on par against the state-of-the-art methods on sentence ordering and room rearrangement. Our work highlights the suitability of diffusion models for ordering problems and proposes a novel formulation and method for solving various ordering tasks. We release our code at https://github.com/IIT-PAVIS/Positional_Diffusion . • The article presents a novel method for Ordering Elements of a Set in 1D and 2D space. • We propose a task-agnostic method, Positional Diffusion for different ordering tasks • Our approach combines Graph Neural Networks with Diffusion Probabilistic Models. • Without any task-specific modes, our method can outperform task-specific approaches. • We test our approach on Sentence ordering, Visual Puzzles, and Furniture Arrangement. Francesco Giuliari, Gianluca Scarpellini, Stefano Fiorini, Stuart James, Pietro Morerio, Yiming Wang 0002, Alessio Del Bue |
Pattern Recognit. Lett. | 4 |
| 2023 | Inclusive Digital Storytelling: Artificial Intelligence and Augmented Reality to Re-centre Stories from the Margins
Valentina Nisi, Stuart James, Paulo Bala, Alessio Del Bue, Nuno Nunes 0001 |
ICIDS (1) | 2 |
| 2023 | "Connected to the people": Social Inclusion & Cohesion in Action through a Cultural Heritage Digital ToolabstractCurrent cultural policies are evolving from social inclusion (removing barriers and promoting equality for participation in culture) to social cohesion (fostering solid bonds between groups despite their differences). Digital interventions can create spaces that promote social inclusion and cohesion. In this paper, we report on the design and evaluation of a cultural heritage and digital storytelling application supporting a participatory approach to culture and hosting society. We evaluate our intervention in three marginalized communities with different social-cultural contexts: migrant women in Barcelona, a community living in a priority neighbourhood in Paris and second and third-generation migrants in Lisbon. Through an analysis of their application use, our findings point at their needs and desires, highlighting how the app can support social inclusion as the first step towards cohesion, but that these are heterogeneous concepts susceptible to nuanced appropriations by the different communities. Valentina Nisi, Paulo Bala, Vanessa Cesário, Stuart James, Alessio Del Bue, Nuno Nunes 0001 |
Proc. ACM Hum. Comput. Interact. | 4 |
| 2023 | Locality-aware subgraphs for inductive link prediction in knowledge graphsabstractRecent methods for inductive reasoning on Knowledge Graphs (KGs) transform the link prediction problem into a graph classification task. They first extract a subgraph around each target link based on the k-hop neighborhood of the target entities, encode the subgraphs using a Graph Neural Network (GNN), then learn a function that maps subgraph structural patterns to link existence. Although these methods have witnessed great successes, increasing k often leads to an exponential expansion of the neighborhood, thereby degrading the GNN expressivity due to oversmoothing. In this paper, we formulate the subgraph extraction as a local clustering procedure that aims at sampling tightly-related subgraphs around the target links, based on a personalized PageRank (PPR) approach. Empirically, on three real-world KGs, we show that reasoning over subgraphs extracted by PPR-based local clustering can lead to a more accurate link prediction model than relying on neighbors within fixed hop distances. Furthermore, we investigate graph properties such as average clustering coefficient and node degree, and show that there is a relation between these and the performance of subgraph-based link prediction. Hebatallah A. Mohamed Hassan 0001, Diego Pilutti, Stuart James, Alessio Del Bue, Marcello Pelillo, Sebastiano Vascon |
Pattern Recognit. Lett. | 3 |
| 2022 | PoserNet: Refining Relative Camera Poses Exploiting Object Detections
Matteo Taiana, Matteo Toso, Stuart James, Alessio Del Bue |
ECCV (33) | 3 |
| 2022 | Writing with (Digital) Scissors: Designing a Text Editing Tool for Assisted Storytelling Using Crowd-Generated Content
Paulo Bala, Stuart James, Alessio Del Bue, Valentina Nisi |
ICIDS | 2 |
| 2022 | Ganzzle: Reframing Jigsaw Puzzle Solving as a Retrieval Task using a Generative Mental ImageabstractPuzzle solving is a combinatorial challenge due to the difficulty of matching adjacent pieces. Instead, we infer a mental image from all pieces, which a given piece can then be matched against avoiding the combinatorial explosion. Exploiting advancements in Generative Adversarial methods, we learn how to reconstruct the image given a set of unordered pieces, allowing the model to learn a joint embedding space to match an encoding of each piece to the cropped layer of the generator. Therefore we frame the problem as a R@1 retrieval task, and then solve the linear assignment using differentiable Hungarian attention, making the process end-to-end. In doing so our model is puzzle size agnostic, in contrast to prior deep learning methods which are single size. We evaluate on two new large-scale datasets, where our model is on par with deep learning methods, while generalizing to multiple puzzle sizes. Davide Talon, Alessio Del Bue, Stuart James |
ICIP | 3 |
| 2021 | Lights: Light Specularity Dataset For Specular Detection In Multi-ViewabstractSpecular highlights are commonplace in images, however, methods for detecting them and removing the phenomenon are particularly challenging. A reason for this is the difficulty in creating a dataset for training or evaluation, as in the real world, we lack the necessary control over the environment. Therefore, we propose a novel physically-based rendered LIGHT Specularity (LIGHTS) Dataset for the evaluation of the specular highlight detection task. Our dataset consists of 18 high-quality architectural scenes, where each scene is rendered with multiple views. In total, the dataset contains 2, 603 views with an average of 145 views per scene. Additionally, we propose a simple aggregation based method for specular highlight detection that outperforms prior work by 3.6% in two orders of magnitude less time on our dataset. Mohamed Dahy Elkhouly, Theodore Tsesmelis, Alessio Del Bue, Stuart James |
ICIP | 4 |
| 2021 | Amnesia in the Atlantic: An AI Driven Serious Game on Marine Biodiversity
Mara Dionisio, Valentina Nisi, Jin Xin, Paulo Bala, Stuart James, Nuno Nunes 0001 |
ICEC | 5 |
| 2021 | Mixing Modalities of 3D Sketching and Speech for Interactive Model Retrieval in Virtual RealityabstractSketch and speech are intuitive interaction methods that convey complementary information and have been independently used for 3D model retrieval in virtual environments. While sketch has been shown to be an effective retrieval method, not all collections are easily navigable using this modality alone. We design a new challenging database for sketch comprised of 3D chairs where each of the components (arms, legs, seat, back) are independently colored. To overcome this, we implement a multimodal interface for querying 3D model databases within a virtual environment. We base the sketch on the state-of-the-art for 3D Sketch Retrieval, and use a Wizard-of-Oz style experiment to process the voice input. In this way, we avoid the complexities of natural language processing which frequently requires fine-tuning to be robust. We conduct two user studies and show that hybrid search strategies emerge from the combination of interactions, fostering the advantages provided by both modalities. Daniele Giunchi, Alejandro Sztrajman, Stuart James, Anthony Steed |
IMX | 3 |
| 2021 | Perceived Realism of Pedestrian Crowds Trajectories in VRabstractCrowd simulation algorithms play an essential role in populating Virtual Reality (VR) environments with multiple autonomous humanoid agents. The generation of plausible trajectories can be a significant computational cost for real-time graphics engines, especially in untethered and mobile devices such as portable VR devices. Previous research explores the plausibility and realism of crowd simulations on desktop computers but fails to account the impact it has on immersion. This study explores how the realism of crowd trajectories affects the perceived immersion in VR. We do so by running a psychophysical experiment in which participants rate the realism of real/synthetic trajectories data, showing similar level of perceived realism. Daniele Giunchi, Riccardo Bovo, Panayiotis Charalambous, Fotis Liarokapis, Alastair Shipman, Stuart James, Anthony Steed, Thomas Heinis |
VRST | 6 |
| 2020 | Machine Learning for Cultural Heritage: A SurveyabstractThe application of Machine Learning (ML) to Cultural Heritage (CH) has evolved since basic statistical approaches such as Linear Regression to complex Deep Learning models. The question remains how much of this actively improves on the underlying algorithm versus using it within a ‘black box’ setting. We survey across ML and CH literature to identify the theoretical changes which contribute to the algorithm and in turn them suitable for CH applications. Alternatively, and most commonly, when there are no changes, we review the CH applications, features and pre/post-processing which make the algorithm suitable for its use. We analyse the dominant divides within ML, Supervised, Semi-supervised and Unsupervised, and reflect on a variety of algorithms that have been extensively used. From such an analysis, we give a critical look at the use of ML in CH and consider why CH has only limited adoption of ML. Marco Fiorucci, Marina Khoroshiltseva, Massimiliano Pontil, Arianna Traviglia, Alessio Del Bue, Stuart James |
Pattern Recognit. Lett. | 6 |
| 2018 | Visual Graphs from Motion (VGfM): Scene Understanding with Object Geometry Reasoning
Paul Gay, Stuart James, Alessio Del Bue |
ACCV (3) | 2 |
| 2018 | Multi-view Aggregation for Color Naming with Shadow Detection and RemovalabstractThis paper presents a set of methods for classifying the color attribute of objects when multiple images of the same objects are available. This problem is more complex than the single image estimation since varying environmental effects, such as, shadows or specularities from light sources, can result in poor accuracy. These depend primarily on the camera positions and the material type of the objects. Single image techniques focus on improving the discrimination of between colors, whereas in multi-view systems additional information is available but should be utilized wisely. To this end, we propose three methods to aggregate image pixel information in multi-view that boost the performance of color name classification. Moreover, we study the effect of shadows by employing automatic shadow detection and correction techniques on the color naming problem. We tested our proposals on a new multi-view color names dataset (M3DCN) which contain indoor and outdoor objects. The experimental evaluation shows that one out of the three presented aggregation methods is very efficient and it achieves the highest accuracy in term of classification results. Also, we experimentally show that addressing visual outliers like shadow in multi-view images improves the performance of the color attribute decision process. Mohamed Dahy Elkhouly, Stuart James, Alessio Del Bue |
IPAS | 2 |
| 2018 | Model Retrieval by 3D Sketching in Immersive Virtual RealityabstractWe describe a novel method for searching 3D model collections using free-form sketches within a virtual environment as queries. As opposed to traditional Sketch Retrieval, our queries are drawn directly onto an example model. Using immersive virtual reality the user can express their query through a sketch that demonstrates the desired structure, color and texture. Unlike previous sketch-based retrieval methods, users remain immersed within the environment without relying on textual queries or 2D projections which can disconnect the user from the environment. We show how a convolutional neural network (CNN) can create multi-view representations of colored 3D sketches. Using such a descriptor representation, our system is able to rapidly retrieve models and in this way, we provide the user with an interactive method of navigating large object datasets. Through a preliminary user study we demonstrate that by using our VR 3D model retrieval system, users can perform quick and intuitive search. Using our system users can rapidly populate a virtual environment with specific models from a very large database, and thus the technique has the potential to be broadly applicable in immersive editing systems. Daniele Giunchi, Stuart James, Anthony Steed |
VR | 2 |
| 2017 | Texture Stationarization: Turning Photos into Tileable TexturesabstractTexture synthesis has grown into a mature field in computer graphics, allowing the synthesis of naturalistic textures and images from photographic exemplars. Surprisingly little work, however, has been dedicated to synthesizing tileable textures, that is, textures that when laid out in a regular grid of tiles form a homogeneous appearance suitable for use in memory-sensitive real-time graphics applications. One of the key challenges in doing so is that most natural input exemplars exhibit uneven spatial variations that, when tiled, show as repetitive patterns. We propose an approach to synthesize tileable textures while enforcing stationarity properties that effectively mask repetitions while maintaining the unique characteristics of the exemplar. We explore a number of alternative measures for texture stationarity and show how each measure can be integrated into a standard texture synthesis method (PatchMatch) to enforce stationarity at user-controlled scales. We demonstrate the efficacy of our approach using a database of 118 exemplar images, both from publicly available sources as well as new ones captured under uncontrolled conditions, and we quantitatively analyze alternative stationarity measures for their robustness across many test runs using different random seeds. In conclusion, we suggest a novel synthesis approach that employs local histogram matching to reliably turn input photographs of natural surfaces into tiles well suited for artifact-free tiling. Joep Moritz, Stuart James, Tom S. F. Haines, Tobias Ritschel 0001, Tim Weyrich |
Comput. Graph. Forum | 2 |
| 2016 | Evolutionary data purification for social media classificationabstractWe present a novel algorithm for the semantic labeling of photographs shared via social media. Such imagery is diverse, exhibiting high intra-class variation that demands large training data volumes to learn representative classifiers. Unfortunately image annotation at scale is noisy resulting in errors in the training corpus that confound classifier accuracy. We show how evolutionary algorithms may be applied to select a 'purified' subset of the training corpus to optimize classifier performance. We demonstrate our approach over a variety of image descriptors (including deeply learned features) and support vector machines. Stuart James, John P. Collomosse |
ICPR | 1 |
| 2014 | Admixed portrait: reflections on being online as a new parentabstractThis Pictorial documents the process of designing a device as an intervention within a field study of new parents. The device was deployed in participating parents' homes to invite reflection on their everyday experiences of portraying self and others through social media in their transition to parenthood. Diego Trujillo-Pisanty, Abigail Durrant, Sarah Martindale, Stuart James, John P. Collomosse |
Conference on Designing Interactive Systems | 4 |
| 2014 | A particle filtering approach to salient video object localizationabstractWe describe a novel fully automatic algorithm for identifying salient objects in video based on their motion. Spatially coherent clusters of optical flow vectors are sampled to generate estimates of affine motion parameters local to super-pixels identified within each frame. These estimates, combined with spatial data, form coherent point distributions in a 5D solution space corresponding to objects or parts there-of. These distributions are temporally denoised using a particle filtering approach, and clustered to estimate the position and motion parameters of salient moving objects in the clip. We demonstrate localization of salient object/s in a variety of clips exhibiting moving and cluttered backgrounds. Charles Gray, Stuart James, John P. Collomosse, Paul Asente |
ICIP | 2 |
| 2014 | ReEnact: Sketch based Choreographic Design from Archival Dance FootageabstractWe describe a novel system for synthesising video choreography using sketched visual storyboards comprising human poses (stick men) and action labels. First, we describe an algorithm for searching archival dance footage using sketched pose. We match using an implicit representation of pose parsed from a mix of challenging low and high fidelity footage. In a training pre-process we learn a mapping between a set of exemplar sketches and corresponding pose representations parsed from the video, which are generalized at query-time to enable retrieval over previously unseen frames, and over additional unseen videos. Second, we describe how a storyboard of sketched poses, interspersed with labels indicating connecting actions, may be used to drive the synthesis of novel video choreography from the archival footage. Stuart James, Manuel J. Fonseca, John P. Collomosse |
ICMR | 1 |
| 2013 | Markov random fields for sketch based video retrievalabstractWe describe a new system for searching video databases using free-hand sketched queries. Our query sketches depict both object appearance and motion, and are annotated with keywords that indicate the semantic category of each object. We parse space-time volumes from video to form graph representation, which we match to sketches under a Markov Random Field (MRF) optimization. The MRF energy function is used to rank videos for relevance and contains unary, pairwise and higher-order potentials that reflect the colour, shape, motion and type of sketched objects. We evaluate performance over a dataset of 500 sports footage clips. Rui Hu 0007, Stuart James, Tinghuai Wang, John P. Collomosse |
ICMR | 2 |
| 2012 | Annotated Free-Hand Sketches for Video Retrieval Using Object Semantics and Motion
Rui Hu 0007, Stuart James, John P. Collomosse |
MMM | 2 |
| 2012 | Skeletons from sketches of dancing posesabstractThe contribution of this paper is a sketch parser able to recognize the several components of a skeleton described using the drawing of a stick-man. We describe the sketch parser in detail, and briefly outline how it is applied to form the front-end of a sketch based retrieval system capable of searching for human poses in archival dance footage. Manuel J. Fonseca, Stuart James, John P. Collomosse |
VL/HCC | 2 |