VLDB 2026 Research / reviewers in the wild / expert
Stuart W. Perry
dblp:50/6281
· DBLP profile ↗
15ranked-venue papers
4as first author
7since 2021 · last 2026
0000-0002-2794-3178ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 11 · 3 first-author · 5 since 2021Artificial intelligence and machine learning · 4 · 1 first-author · 3 since 2021Applied, interdisciplinary, general and emerging computing · 1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | StructGS: Adaptive Spherical Harmonics and Rendering Enhancements for Superior 3D Gaussian SplattingabstractRecent advancements in 3D reconstruction coupled with neural rendering techniques have greatly improved the creation of photo-realistic 3D scenes, influencing both academic research and industry applications. The technique of 3D Gaussian Splatting and its variants incorporate the strengths of both primitive-based and volumetric representations, achieving superior rendering quality. While 3D Geometric Scattering (3DGS) and its variants have advanced the field of 3D representation, they fall short in capturing the stochastic properties of non-local structural information during the training process. Additionally, the initialisation of spherical functions in 3DGS-based methods often fails to engage higher-order terms in early training rounds, leading to unnecessary computational overhead as training progresses. Furthermore, current 3DGS-based approaches require training on higher resolution images to render higher resolution outputs, significantly increasing memory demands and prolonging training durations. We introduce StructGS, a framework that enhances 3D Gaussian Splatting (3DGS) for improved novel-view synthesis in 3D reconstruction. StructGS innovatively incorporates a patch-based SSIM loss, dynamic spherical harmonics initialisation and a Multi-scale Residual Network (MSRN) to address the above-mentioned limitations, respectively. Our framework significantly reduces computational redundancy, enhances detail capture and supports high-resolution rendering from low-resolution inputs. Experimentally, StructGS demonstrates superior performance over state-of-the-art (SOTA) models, achieving higher quality and more detailed renderings with fewer artifacts. (The link to the code will be made available after publication.). Zexu Huang, Min Xu 0001, Stuart W. Perry |
IEEE Trans. Multim. | 3 |
| 2025 | RLD-GS: Reinforcement Learning-Driven Gaussian Splatting for High-Fidelity Neural Rendering
Zexu Huang, Min Xu 0001, Stuart W. Perry |
ICONIP (2) | 3 |
| 2025 | ReviveDiff: A Universal Diffusion Model for Restoring Images in Adverse Weather ConditionsabstractImages captured in challenging environments-such as nighttime, smoke, rainy weather, and underwater-often suffer from significant degradation, resulting in a substantial loss of visual quality. The effective restoration of these degraded images is critical for the subsequent vision tasks. While many existing approaches have successfully incorporated specific priors for individual tasks, these tailored solutions limit their applicability to other degradations. In this work, we propose a universal network architecture, dubbed "ReviveDiff", which can address various degradations and restore images to their original quality by enhancing and restoring their details. Our approach is inspired by the observation that, unlike degradation caused by movement or electronic issues, quality degradation under adverse conditions primarily stems from natural media (such as fog, water, and low luminance), which generally preserves the original structures of objects. To restore the quality of such images, we leveraged the latest advancements in diffusion models and developed ReviveDiff to restore image quality from both macro and micro levels across some key factors determining image quality, such as sharpness, distortion, noise level, dynamic range, and color accuracy. We rigorously evaluated ReviveDiff on seven benchmark datasets covering five types of degrading conditions: Rainy, Underwater, Low-light, Smoke, and Nighttime Hazy. Our experimental results demonstrate that ReviveDiff outperforms the state-of-the-art methods both quantitatively and visually. Wenfeng Huang, Guoan Xu, Wenjing Jia, Stuart W. Perry, Guangwei Gao |
IEEE Trans. Image Process. | 4 |
| 2023 | JPEG Pleno Call for Proposals Responses Quality AssessmentabstractIn this paper, the quality evaluation of the responses to the Call for Proposals (CfP) of JPEG Pleno Point Cloud Coding is presented. Three responses to the CfP were evaluated together with the state of the art anchor codecs G-PCC and VPCC from MPEG. The JPEG committee selected a set of eight point clouds that were encoded at different pre-established bitrates. For the subjective evaluation of the responses to the CfP, a set of video sequences were created where the reference and distorted decoded point clouds were rotated about their axes side by side. Furthermore, the objective quality metrics PCQM, PSNR D1, PSNR D2, PSNR Y and PSNR YUV were computed, and compared with the subjective evaluation results. This study revealed that the deep learning solutions outperformed G-PCC but were still below the performance of V-PCC regarding color representation. PCQM showed the best performance in predicting the compression quality. João Prazeres, António M. G. Pinheiro, Luís Alberto da Silva Cruz, Stuart W. Perry |
ICASSP | 5 |
| 2021 | Single-View 3D Object Reconstruction From Shape Priors in MemoryabstractExisting methods for single-view 3D object reconstruction directly learn to transform image features into 3D representations. However, these methods are vulnerable to images containing noisy backgrounds and heavy occlusions because the extracted image features do not contain enough information to reconstruct high-quality 3D shapes. Humans routinely use incomplete or noisy visual cues from an image to retrieve similar 3D shapes from their memory and reconstruct the 3D shape of an object. Inspired by this, we propose a novel method, named Mem3D, that explicitly constructs shape priors to supplement the missing information in the image. Specifically, the shape priors are in the forms of "image-voxel" pairs in the memory network, which is stored by a well-designed writing strategy during training. We also propose a voxel triplet loss function that helps to retrieve the precise 3D shapes that are highly related to the input image from shape priors. The LSTM-based shape encoder is introduced to extract information from the retrieved 3D shapes, which are useful in recovering the 3D shape of an object that is heavily occluded or in complex environments. Experimental results demonstrate that Mem3D significantly improves reconstruction quality and performs favorably against state-of-the-art methods on the ShapeNet and Pix3D datasets. Shuo Yang 0006, Min Xu 0001, Haozhe Xie, Stuart W. Perry, Jiahao Xia 0001 |
CVPR | 4 |
| 2021 | Comparison of Remote Subjective Assessment Strategies in the Context of the JPEG Pleno Point Cloud ActivityabstractIn this work we compare two different options to perform on-line subjective quality assessment experiments in the context of the Call for Evidence on JPEG Pleno Point Cloud Coding. A deep-learning based point cloud codec submitted to the Call was tested against current MPEG point cloud compression methods. The first option is based on participants downloading the entire set of stimuli and running a set of scripts in MATLAB to perform the experiment. The second option involves the participants accessing a server on the web and viewing and judging the stimuli using a web browser. Quality scores compiled using both methods were compared showing strong correlation. A second analysis compared the quality scores with those obtained in a prior laboratory-based study using higher resolution screens. The entire study also brought to light each option’s unique advantages and disadvantages that make each one better suited to specific types of subjective evaluation contexts and situations. Stuart W. Perry, Luís Alberto da Silva Cruz, Emil Dumic, Nhung Hong Thi Nguyen, António M. G. Pinheiro, Evangelos Alexiou |
MMSP | 1 |
| 2021 | Two-stage convolutional neural network for road crack detection and segmentationabstractAutomatic detection of road cracks is an important task to support road inspection for transport infrastructure. Various methods have been proposed for road crack detection and segmentation, however, there is no established method for handling real road images that are noisy and of low quality. In this paper, a new method utilising a two-stage convolutional neural network (CNN) is proposed for road crack detection and segmentation in images at the pixel level. Our novel contribution is a framework where the first stage serves to remove noise or artifacts and isolate the potential cracks to a small area, and the second stage is able to learn the context of cracks in the detected area. This is hence more effective than learning over the entire original noisy image. Extensive experiments on real datasets including public sources and our collected dataset have been conducted. The experimental results show that the two-stage CNN model outperformed existing approaches, especially for noisy, low-resolution images, and imbalanced datasets. Our approach achieves an F1-measure of over 0.91 on three datasets. Nhung Hong Thi Nguyen, Stuart W. Perry, Don Bone, Thi Thuy Nguyen |
Expert Syst. Appl. | 2 |
| 2020 | Quality Evaluation Of Static Point Clouds Encoded Using MPEG CodecsabstractThis paper presents a quality evaluation study of point cloud codecs that have been recently standardised by the MPEG committee. In particular, a subjective experiment to assess their performance in terms of bitrate against visual quality is designed and realized in four independent laboratories. The experimental setup of each laboratory varies; yet, the obtained subjective scores exhibit high inter laboratory correlation, confirming that the adopted assessment protocol is robust to equipment selection and viewing conditions, ensuring reliability and facilitating repeatability. Our study confirms the superior compression performance of the MPEG V-PCC, when compared to MPEG G-PCC, in the case of static contents. Finally, results from a benchmark of the most popular objective quality metrics using the obtained subjective scores as ground truth, reveal that the point2plane with mean square error is the most accurate quality predictor, closely followed by the point2point also using mean square error as distance measure. Stuart W. Perry, Huy Phi Cong, Luís Alberto da Silva Cruz, João Prazeres, Manuela Pereira, António M. G. Pinheiro, Emil Dumic, Evangelos Alexiou, Touradj Ebrahimi |
ICIP | 1 |
| 2020 | Recall What You See Continually Using GridLSTM in Image CaptioningabstractThe goal of image captioning is to automatically describe an image with a sentence, and the task has attracted research attention from both the computer vision and natural-language processing research communities. The existing encoder-decoder model and its variants, which are the most popular models for image captioning, use the image features in three ways: first, they inject the encoded image features into the decoder only once at the initial step, which does not enable the rich image content to be explored sufficiently while gradually generating a text caption; second, they concatenate the encoded image features with text as extra inputs at every step, which introduces unnecessary noise; and, third, they using an attention mechanism, which increases the computational complexity due to the introduction of extra neural nets to identify the attention regions. Different from the existing methods, in this paper, we propose a novel network, Recall Network, for generating captions that are consistent with the images. The recall network selectively involves the visual features by using a GridLSTM and, thus, is able to recall image contents while generating each word. By importing the visual information as the latent memory along the depth dimension LSTM, the decoder is able to admit the visual features dynamically through the inherent LSTM structure without adding any extra neural nets or parameters. The Recall Network efficiently prevents the decoder from deviating from the original image content. To verify the efficiency of our model, we conducted exhaustive experiments on full and dense image captioning. The experimental results clearly demonstrate that our recall network outperforms the conventional encoder-decoder model by a large margin and that it performs comparably to the state-of-the-art methods. Lingxiang Wu, Min Xu 0001, Jinqiao Wang, Stuart W. Perry |
IEEE Trans. Multim. | 4 |
| 2019 | Airborne Object Detection Using Hyperspectral Imaging: Deep Learning Review
Thuy T. Pham, Madhumita A. Takalkar, Min Xu 0001, Dinh Thai Hoang, H. A. Truong, Eryk Dutkiewicz, Stuart W. Perry |
ICCSA (1) | 7 |
| 2018 | A novel spatial pooling method for 3D mesh quality assessment based on percentile weighting strategy
Xiang Feng 0007, Stuart W. Perry, Song Zhu |
Comput. Graph. | 4 |
| 2018 | A new mesh visual quality metric using saliency weighting-based pooling strategy
Xiang Feng 0007, Stuart W. Perry, Song Zhu, Zexin Liu |
Graph. Model. | 4 |
| 2000 | Weight assignment for adaptive image restoration by neural networksabstractThis paper presents a scheme for adaptively training the weights, in terms of varying the regularization parameter, in a neural network for the restoration of digital images. The flexibility of neural-network-based image restoration algorithms easily allow the variation of restoration parameters such as blur statistics and regularization value spatially and temporally within the image. This paper focuses on spatial variation of the regularization parameter.We first show that the previously proposed neural-network method based on gradient descent can only find suboptimal solutions, and then introduce a regional processing approach based on local statistics. A method is presented to vary the regularization parameter spatially. This method is applied to a number of images degraded by various levels of noise, and the results are examined. The method is also applied to an image degraded by spatially variant blur. In all cases, the proposed method provides visually satisfactory results in an efficient way. Stuart W. Perry, Ling Guan |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 1998 | Neural vision system and applications in image processing and analysisabstractWe present a computer vision system based on an integrated neural network architecture. In the low level vision subsystem, a network of networks-a biologically inspired network is used to recursively perform filtering, segmentation and edge detection; in the intermediate level and the high level, hierarchically structured arrays of self-organizing tree maps-extension of the popular self-organizing map are utilized to carry out image/feature analysis. The system has been applied to solve a number of real world problems. Some interesting and encouraging results are reported. Ling Guan, Stuart W. Perry, Raniero Romagnoli, Hau-San Wong, Haosong Kong |
ICASSP | 2 |
| 1998 | Perception based adaptive image restorationabstractThis paper presents an image restoration technique which uses a cost function based on a novel image error measure. The cost function presented here takes into account local statistical information of the image when performing restoration. It is shown that this technique compares favourably with other techniques, especially when applied to colour images. Stuart W. Perry, Ling Guan |
ICASSP | 1 |