EDBT 2026 Demo / reviewers in the wild / expert
Gregory Slabaugh
dblp:s/GregoryGSlabaugh · also Greg Slabaugh, Gregory G. Slabaugh
· DBLP profile ↗
74ranked-venue papers
9as first author
25since 2021 · last 2026
0000-0003-4060-5226ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 52 · 8 first-author · 16 since 2021Artificial intelligence and machine learning · 40 · 4 first-author · 14 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Human-computer interaction and ubiquitous computing · 1Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Deep learning models to map osteocyte networks from confocal microscopy can successfully distinguish between young and aged boneabstractOsteocytes, the most abundant and mechanosensitive cells in bone tissue, play a pivotal role in bone homeostasis and mechano-responsiveness, orchestrating the delicate balance between bone formation and resorption under daily activity. Studying osteocyte connectivity and understanding their intricate arrangement within the lacunar canalicular network is essential for unravelling bone physiology, which is significantly disrupted during ageing. Much work has been carried out to investigate this relationship, often involving high resolution microscopy of discrete fragments of this network, alongside advanced computational modelling of individual cells. However, traditional methods of segmenting and measuring osteocyte connectomics are time-consuming and labour-intensive, often hindered by human subjectivity and limited throughput. In this study, we explored the application of deep learning and computer vision techniques to automate the segmentation and measurement of osteocyte connectomics, enabling more efficient and accurate analysis. For this specific application, once trained, the analysis was completed within 10 seconds, compared to manual segmentation time of 130 hours. We compared a number of state-of-the-art computer vision models (U-Nets and Vision Transformers) to successfully segment the osteocyte network, finding that an Attention U-Net model can accurately segment and measure 81.8% of osteocytes and 42.1% of dendritic processes, when compared to manual labelling. While further development is required, we demonstrated that this degree of accuracy is already sufficient to distinguish between bones of young (2-month-old) and aged (36-month-old) mice, as well as partially capturing the degeneration induced by genetic modification of osteocytes. Comparison of the model predictions with manual measurements showed no significant difference, indicating that, with additional training, such deep learning algorithms could be trained to human-level accuracy when measuring the osteocyte network. By harnessing the power of these advanced technologies, further developments will likely shed light on the complexities of osteocyte networks with ever-increasing efficiency. Simon D. Vetter, Charles A. Schurman, Tamara Alliston, Gregory Slabaugh, Stefaan W. Verbruggen |
PLoS Comput. Biol. | 4 |
| 2026 | Adapter-RL: Adaptation of Any Agent Using Reinforcement LearningabstractThis study introduces Adapter-RL, a novel architecture aimed at improving the performance of existing agents in reinforcement learning tasks. The approach integrates human-knowledge-based systems with deep reinforcement learning, combining the interpretability and rule-based logic of the former with the adaptive learning capabilities of the latter. A crucial aspect of this method is the use of “adapters”—concise modules integrated with a base-agent, designed to adjust the policy for specific tasks. The Adapter-RL framework comprises a base-agent responsible for initial decision-making and an adapter module that refines these decisions to meet task-specific requirements. The adapter facilitates efficient training, reduces parameter requirements, and mitigates catastrophic forgetting, enhancing overall performance and adaptability. This architecture enables agents to be fine-tuned effectively, allowing them to adapt to complex tasks with rapidly changing or uncertain conditions. The research demonstrates the efficacy of Adapter-RL through experiments in microRTS, a challenging real-time strategy game. The results demonstrate that Adapter-RL significantly accelerates the training process and outperforms base-agents across various tasks, highlighting its efficiency and robustness. In addition, the study investigates the temperature coefficient tradeoff in adapter training, finding that optimal performance is achievable within a broad range of coefficients. This underscores the stability of the method. The Adapter-RL method enables the specialization of base AI for specific characters or scenarios. Yizhao Jin, Gregory Slabaugh, Simon M. Lucas |
IEEE Trans. Games | 2 |
| 2025 | BioX-CPath: Biologically-driven Explainable Diagnostics for Multistain IHC Computational PathologyabstractThe development of biologically interpretable and explainable models remains a key challenge in computational pathology, particularly for multistain immunohistochemistry (IHC) analysis. We present BioX-CPath, an explainable graph neural network architecture for whole slide image (WSI) classification that leverages both spatial and semantic features across multiple stains. At its core, BioX-CPath introduces a novel Stain-Aware Attention Pooling (SAAP) module that generates biologically meaningful, stain-aware patient embeddings. Our approach achieves state-of-the-art performance on both Rheumatoid Arthritis and Sjogren’s Disease multistain datasets. Beyond performance metrics, BioX-CPath provides interpretable insights through stain attention scores, entropy measures, and stain interaction scores, that permit measuring model alignment with known pathological mechanisms. This biological grounding, combined with strong classification performance, makes BioX-CPath particularly suitable for clinical applications where interpretability is key. Source code and documentation can be found at: https://github.com/AmayaGS/BioX-CPath. Amaya Gallagher-Syed, Henry Senior, Omnia Alwazzan, Elena Pontarini, Michele Bombardieri, Costantino Pitzalis, Myles J. Lewis, Michael R. Barnes, Luca Rossi 0011, Gregory Slabaugh |
CVPR | 10 |
| 2025 | STaR: Seamless Spatial-Temporal Aware Motion Retargeting with Penetration and Consistency ConstraintsabstractMotion retargeting seeks to faithfully replicate the spatio-temporal motion characteristics of a source character onto a target character with a different body shape. Apart from motion semantics preservation, ensuring geometric plausibility and maintaining temporal consistency are also crucial for effective motion retargeting. However, many existing methods prioritize either geometric plausibility or temporal consistency. Neglecting geometric plausibility results in interpenetration while neglecting temporal consistency leads to motion jitter. In this paper, we propose a novel sequence-to-sequence model for seamless Spatial-Temporal aware motion Retargeting (STaR), with penetration and consistency constraints. STaR consists of two modules: (1) a spatial module that incorporates dense shape representation and a novel limb penetration constraint to ensure geometric plausibility while preserving motion semantics, and (2) a temporal module that utilizes a temporal transformer and a novel temporal consistency constraint to predict the entire motion sequence at once while enforcing multi-level trajectory smoothness. The seamless combination of the two modules helps us achieve a good balance between the semantic, geometric, and temporal targets. Extensive experiments on the Mixamo and ScanRet datasets demonstrate that our method produces plausible and coherent motions while significantly reducing interpenetration rates compared with other approaches. Code page: https://github.com/XiaohangYang829/STaR. Xiaohang Yang, Gregory Slabaugh, Shanxin Yuan |
ICCV | 4 |
| 2025 | XFMamba: Cross-Fusion Mamba for Multi-view Medical Image Classification
Xiaoyu Zheng 0001, Xu Chen 0030, Shaogang Gong, Xavier Griffin, Gregory Slabaugh |
MICCAI (1) | 5 |
| 2025 | Robust Photo-Realistic Hand Gesture Generation: from Single View to Multiple ViewabstractHigh-fidelity hand gesture generation represents a significant challenge in human-centric generation tasks. Existing methods typically employ a single-view mesh-rendered image prior to enhancing gesture generation quality. However, the spatial complexity of hand gestures and the inherent limitations of single-view rendering make it difficult to capture complete gesture information, particularly when fingers are occluded. The fundamental contradiction lies in the loss of 3D topological relationships through 2D projection and the incomplete spatial coverage inherent to single-view representations. Diverging from single-view prior approaches, we propose a multi-view prior framework, named Multi-Modal UNet-based Feature Encoder (MUFEN), to guide diffusion models in learning comprehensive 3D hand information. Specifically, we extend conventional front-view rendering to include rear, left, right, top, and bottom perspectives, selecting the most information-rich view combination as training priors to address occlusion. This multi-view prior with a dedicated dual stream encoder significantly improves the model's understanding of complete hand features. Furthermore, we design a bounding box feature fusion module, which can fuse the gesture localization features and multi-modal features to enhance the location-awareness of the MUFEN features to the gesture-related features. Experiments demonstrate that our method achieves state-of-the-art performance in both quantitative metrics and qualitative evaluations. The source code is available at https://github.com/fuqifan/MUFEN. Qifan Fu, Xu Chen 0030, Muhammad Asad 0001, Shanxin Yuan, Changjae Oh, Gregory Slabaugh |
ACM Multimedia | 6 |
| 2025 | ViDAR: Video Diffusion-Aware 4D Reconstruction From Monocular InputsabstractDynamic Novel View Synthesis aims to generate photorealistic views of moving subjects from arbitrary viewpoints. This task is particularly challenging when relying on monocular video, where disentangling structure from motion is ill-posed and supervision is scarce. We introduce Video Diffusion-Aware Reconstruction (ViDAR), a novel 4D reconstruction framework that leverages personalised diffusion models to synthesise a pseudo multi-view supervision signal for training a Gaussian splatting representation. By conditioning on scene-specific features, ViDAR recovers fine-grained appearance details while mitigating artefacts introduced by monocular ambiguity. To address the spatio-temporal inconsistency of diffusion-based supervision, we propose a diffusion-aware loss function and a camera pose optimisation strategy that aligns synthetic views with the underlying scene geometry. Experiments on DyCheck, a challenging benchmark with extreme viewpoint variation, show that ViDAR outperforms all state-of-the-art baselines in visual quality and geometric consistency. We further highlight ViDAR’s strong improvement over baselines on dynamic regions and provide a new benchmark to compare performance in reconstructing motion-rich parts of the scene. Michal Nazarczuk, Sibi Catley-Chandar, Thomas Tanay, Zhensong Zhang, Gregory Slabaugh, Eduardo Pérez-Pellitero |
NeurIPS | 5 |
| 2025 | CAMS: Convolution and Attention-Free Mamba-Based Cardiac Image SegmentationabstractConvolutional Neural Networks (CNNs) and Transformer-based self-attention models have become the standard for medical image segmentation. This paper demonstrates that convolution and self-attention, while widely used, are not the only effective methods for segmentation. Breaking with convention, we present a Convolution and self-Attention-free Mamba-based seman-tic Segmentation Network named CAMS-Net. Specifically, we design Mamba-based Channel Aggregator and Spatial Aggregator, which are applied independently in each encoder-decoder stage. The Channel Aggregator extracts information across different channels, and the Spatial Ag-gregator learns features across different spatial locations. We also propose a Linearly Interconnected Factorized Mamba (LIFM) block to reduce the computational complexity of a Mamba block and to enhance its decision function by introducing a non-linearity between two factor-ized Mamba blocks. Our model outperforms the existing state-of-the-art CNN, self-attention, and Mamba-based methods on CMR and M&Ms-2 Cardiac segmentation datasets, showing how this innovative, convolution, and self-attention-free method can inspire further research beyond CNN and Transformer paradigms, achieving linear complexity and reducing the number of parameters. Source code and pre-trained models are available at: https://github.com/kabbas570/CAMS-Net. Muhammad Asad 0001, Martin Benning, Caroline H. Roney, Gregory Slabaugh |
WACV | 5 |
| 2025 | Compositional Segmentation of Cardiac Images Leveraging MetadataabstractCardiac image segmentation is essential for automated cardiac function assessment and monitoring of changes in cardiac structures over time. Inspired by coarse-to-fine approaches in image analysis, we propose a novel multitask compositional segmentation approach that can simultaneously localize the heart in a cardiac image and perform part-based segmentation of different regions of interest. We demonstrate that this compositional approach achieves better results than direct segmentation of the anatomies. Further, we propose a novel Cross-Modal Feature Integration (CMFI) module to leverage the meta-data related to cardiac imaging collected during image acquisition. We perform experiments on two different modalities, MRI and ultrasound, using public datasets, Multi-Disease, Multi-View, and Multi-Centre (M &Ms-2) and Multi-structure Ultrasound Segmentation (CAMUS) data, to showcase the efficiency of the proposed compositional segmentation method and Cross-Modal Feature Integration module incorporating metadata within the proposed compositional segmentation network. The source code is available: https://github.com/kabbas570/CompSeg-MetaData. Muhammad Asad 0001, Martin Benning, Caroline H. Roney, Gregory Slabaugh |
WACV | 5 |
| 2025 | Partial Advantage Estimator for Proximal Policy OptimizationabstractThis paper proposes an innovative approach to the Generalized Advantage Estimator (GAE) to address the bias-variance trade-off in truncated roll-outs during reinforcement learning. In typical GAE implementations, the k-step advantage is estimated using a lambda-weighted average, until the terminal state. While this method provides constant bias-variance properties at any time step, it often necessitates truncated roll-outs with shorter horizons for faster learning and policy updates within a single episode. This study highlights an unexplored issue: the bias-variance properties differ for small versus considerable time steps within truncated roll-outs. Specifically, smaller time steps may have a significant bias, prompting a need for their increase. The proposed solution involves a partial GAE update, calculating the advantage estimates for all time steps but updating the policy only for a specified range. To prevent data wastage, the data from this range is retained for further processing and policy parameter updates. This partial GAE approach, despite the increased memory requirements, promises enhanced computation speed and optimal data utilization. Empirical validation was conducted on four MuJoCo tasks and microRTS. The results show a performance improvement trend with the partial GAE estimator, outperforming regular GAE in task completion speed in microRTS. These findings offer a promising direction for improving policy update efficiency in reinforcement learning. Yizhao Jin, Xiulei Song, Gregory Slabaugh, Simon M. Lucas |
IEEE Trans. Games | 3 |
| 2025 | Graph neural networks in vision-language image understanding: a surveyabstractAbstract 2D image understanding is a complex problem within computer vision, but it holds the key to providing human-level scene comprehension. It goes further than identifying the objects in an image, and instead, it attempts to understand the scene. Solutions to this problem form the underpinning of a range of tasks, including image captioning, visual question answering (VQA), and image retrieval. Graphs provide a natural way to represent the relational arrangement between objects in an image, and thus, in recent years graph neural networks (GNNs) have become a standard component of many 2D image understanding pipelines, becoming a core architectural component, especially in the VQA group of tasks. In this survey, we review this rapidly evolving field and we provide a taxonomy of graph types used in 2D image understanding approaches, a comprehensive list of the GNN models used in this domain, and a roadmap of future potential developments. To the best of our knowledge, this is the first comprehensive survey that covers image captioning, visual question answering, and image retrieval techniques that focus on using GNNs as the main part of their architecture. Henry Senior, Gregory Slabaugh, Shanxin Yuan, Luca Rossi 0011 |
Vis. Comput. | 2 |
| 2024 | RoGUENeRF: A Robust Geometry-Consistent Universal Enhancer for NeRF
Sibi Catley-Chandar, Richard Shaw, Gregory Slabaugh, Eduardo Pérez-Pellitero |
ECCV (12) | 3 |
| 2024 | RAVE: Residual Vector Embedding for CLIP-Guided Backlit Image Enhancement
Tatiana Gaintseva, Martin Benning, Gregory Slabaugh |
ECCV (79) | 3 |
| 2024 | GoLDFormer: A global-local deformable window transformer for efficient image restoration
Bolun Zheng, Chenggang Yan 0001, Zunjie Zhu, Tingyu Wang 0002, Gregory Slabaugh, Shanxin Yuan |
J. Vis. Commun. Image Represent. | 6 |
| 2024 | Non-local degradation modeling for spatially adaptive single image super-resolution
Qianyu Zhang 0002, Bolun Zheng, Zongpeng Li, Yu Liu 0005, Zunjie Zhu, Gregory Slabaugh, Shanxin Yuan |
Neural Networks | 6 |
| 2023 | Improving Dynamic HDR Imaging with Fusion TransformerabstractReconstructing a High Dynamic Range (HDR) image from several Low Dynamic Range (LDR) images with different exposures is a challenging task, especially in the presence of camera and object motion. Though existing models using convolutional neural networks (CNNs) have made great progress, challenges still exist, e.g., ghosting artifacts. Transformers, originating from the field of natural language processing, have shown success in computer vision tasks, due to their ability to address a large receptive field even within a single layer. In this paper, we propose a transformer model for HDR imaging. Our pipeline includes three steps: alignment, fusion, and reconstruction. The key component is the HDR transformer module. Through experiments and ablation studies, we demonstrate that our model outperforms the state-of-the-art by large margins on several popular public datasets. Rufeng Chen, Bolun Zheng, Chenggang Yan 0001, Gregory Slabaugh, Shanxin Yuan |
AAAI | 6 |
| 2023 | Multi-Stain Self-Attention Graph Multiple Instance Learning Pipeline for Histopathology Whole Slide Images
Amaya Gallagher-Syed, Luca Rossi 0011, Felice Rivellese, Costantino Pitzalis, Myles J. Lewis, Michael R. Barnes, Gregory Slabaugh |
BMVC | 7 |
| 2023 | Joint Dense-Point Representation for Contour-Aware Graph Segmentation
Kit Mills Bransby, Gregory Slabaugh, Christos V. Bourantas, Qianni Zhang |
MICCAI (3) | 2 |
| 2023 | Diagnosing and Preventing Instabilities in Recurrent Video ProcessingabstractRecurrent models are a popular choice for video enhancement tasks such as video denoising or super-resolution. In this work, we focus on their stability as dynamical systems and show that they tend to fail catastrophically at inference time on long video sequences. To address this issue, we (1) introduce a diagnostic tool which produces input sequences optimized to trigger instabilities and that can be interpreted as visualizations of temporal receptive fields, and (2) propose two approaches to enforce the stability of a model during training: constraining the spectral norm or constraining the stable rank of its convolutional layers. We then introduce Stable Rank Normalization for Convolutional layers (SRN-C), a new algorithm that enforces these constraints. Our experimental results suggest that SRN-C successfully enforces stablility in recurrent video processing models without a significant performance loss. Thomas Tanay, Aivar Sootla, Matteo Maggioni, Puneet K. Dokania, Philip Torr 0001, Ales Leonardis, Gregory Slabaugh |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2022 | DomainPlus: Cross Transform Domain Learning towards High Dynamic Range ImagingabstractHigh dynamic range (HDR) imaging by combining multiple low dynamic range (LDR) images of different exposures provides a promising way to produce high quality photographs. However, the misalignment between the input images leads to ghosting artifacts in the reconstructed HDR image. In this paper, we propose a cross-transform domain neural network for efficient HDR imaging. Our approach consists of two modules: a merging module and a restoration module. For the merging module, we propose a Multiscale Attention with Fronted Fusion (MAFF) mechanism to achieve coarse-to-fine spatial fusion. For the restoration module, we propose fronted Discrete Wavelet Transform (DWT) and Discrete Cosine Transform (DCT)-based learnable bandpass filters to formulate a cross-transform domain learning block, dubbed DomainPlus Block (DPB) for effective ghosting removal. Our ablation study and comprehensive experiments show that DomainPlus outperforms the existing state-of-the-art on several datasets. Bolun Zheng, Xiaokai Pan, Xiaofei Zhou 0003, Gregory Slabaugh, Chenggang Yan 0001, Shanxin Yuan |
ACM Multimedia | 5 |
| 2022 | Bidirectional difference locating and semantic consistency reasoning for change captioningabstractChange captioning is an emerging task to describe the changes between a pair of images. The difficulty in this task is to discover the differences between the two images. Recently, some methods have been proposed to address this problem. However, they all employ unidirectional difference localization to identify the changes. This can lead to ambiguity about the nature of the changes. Instead, we propose a framework with bidirectional difference localization and semantic consistency reasoning to describe the image changes. First, we locate the changes in the two images by capturing bidirectional differences. Then we design a decoder with spatial-channel attention to generate the change caption. Finally, we introduce semantic consistency reasoning to constrain our bidirectional difference localization module and spatial-channel attention module. Extensive experiments on three public data sets show that the performance of our proposed model outperforms the state-of-the-art change captioning models by a large margin. Yaoqi Sun, Liang Li 0003, Tongyv Lu, Bolun Zheng, Chenggang Yan 0001, Yongjun Bao, Guiguang Ding, Gregory Slabaugh |
Int. J. Intell. Syst. | 10 |
| 2022 | A Continual Learning Survey: Defying Forgetting in Classification TasksabstractArtificial neural networks thrive in solving the classification problem for a particular rigid task, acquiring knowledge through generalized learning behaviour from a distinct training phase. The resulting network resembles a static entity of knowledge, with endeavours to extend this knowledge without targeting the original task resulting in a catastrophic forgetting. Continual learning shifts this paradigm towards networks that can continually accumulate knowledge over different tasks without the need to retrain from scratch. We focus on task incremental classification, where tasks arrive sequentially and are delineated by clear boundaries. Our main contributions concern: (1) a taxonomy and extensive overview of the state-of-the-art; (2) a novel framework to continually determine the stability-plasticity trade-off of the continual learner; (3) a comprehensive experimental comparison of 11 state-of-the-art continual learning methods; and (4) baselines. We empirically scrutinize method strengths and weaknesses on three benchmarks, considering Tiny Imagenet and large-scale unbalanced iNaturalist and a sequence of recognition datasets. We study the influence of model capacity, weight decay and dropout regularization, and the order in which the tasks are presented, and qualitatively compare methods in terms of required memory, computation time, and storage. Matthias De Lange, Rahaf Aljundi, Marc Masana, Sarah Parisot, Xu Jia 0012, Ales Leonardis, Gregory Slabaugh, Tinne Tuytelaars |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2022 | Learning Frequency Domain Priors for Image DemoireingabstractImage demoireing is a multi-faceted image restoration task involving both moire pattern removal and color restoration. In this paper, we raise a general degradation model to describe an image contaminated by moire patterns, and propose a novel multi-scale bandpass convolutional neural network (MBCNN) for single image demoireing. For moire pattern removal, we propose a multi-block-size learnable bandpass filters (M-LBFs), based on a block-wise frequency domain transform, to learn the frequency domain priors of moire patterns. We also introduce a new loss function named Dilated Advanced Sobel loss (D-ASL) to better sense the frequency information. For color restoration, we propose a two-step tone mapping strategy, which first applies a global tone mapping to correct for a global color shift, and then performs local fine tuning of the color per pixel. To determine the most appropriate frequency domain transform, we investigate several transforms including DCT, DFT, DWT, learnable non-linear transform and learnable orthogonal transform. We finally adopt the DCT. Our basic model won the AIM2019 demoireing challenge. Experimental results on three public datasets show that our method outperforms state-of-the-art methods by a large margin. Bolun Zheng, Shanxin Yuan, Chenggang Yan 0001, Xiang Tian 0002, Jiyong Zhang 0001, Yaoqi Sun, Lin Liu 0016, Ales Leonardis, Gregory Slabaugh |
IEEE Trans. Pattern Anal. Mach. Intell. | 9 |
| 2022 | CBREN: Convolutional Neural Networks for Constant Bit Rate Video Quality EnhancementabstractConstant bit rate (CBR) videos are widely used in streaming playback applications. However, the image quality of the CBR video is often unstable, especially for scenes with large motion. To this end, we design a new model to represent the distortion of High Efficiency Video Coding (HEVC) constant bit rate video, and propose a neural network for a constant bit rate video quality enhancement (CBREN). We propose a dual-domain restoration module (DRM) to jointly learn the prior knowledge in the pixel domain and the frequency domain. To address the degradation resulting from compression, we propose a two-step quantization degradation estimation strategy. The Inverse DCT (IDCT) Translation Unit (ITU) is used to constrain the quantization table of the constant bit rate video to a suitable range, and the Dynamic Alpha Unit (DAU) is used to fine-tune the quantization table according to the content of each frame. In order to effectively reduce the block distortion of different sizes produced in the compression process, we adopt a multi-scale network. Extensive experiments show that our approach can greatly enhance the quality of CBR compressed video. Moreover, our method can also be applied to constant quantization parameter (CQP) video enhancement tasks, and is certainly superior to existing methods. Hengrun Zhao, Bolun Zheng, Shanxin Yuan, Chenggang Yan 0001, Liang Li 0003, Gregory Slabaugh |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2022 | FlexHDR: Modeling Alignment and Exposure Uncertainties for Flexible HDR ImagingabstractHigh dynamic range (HDR) imaging is of fundamental importance in modern digital photography pipelines and used to produce a high-quality photograph with well exposed regions despite varying illumination across the image. This is typically achieved by merging multiple low dynamic range (LDR) images taken at different exposures. However, over-exposed regions and misalignment errors due to poorly compensated motion result in artefacts such as ghosting. In this paper, we present a new HDR imaging technique that specifically models alignment and exposure uncertainties to produce high quality HDR results. We introduce a strategy that learns to jointly align and assess the alignment and exposure reliability using an HDR-aware, uncertainty-driven attention map that robustly merges the frames into a single high quality HDR image. Further, we introduce a progressive, multi-stage image fusion approach that can flexibly merge any number of LDR images in a permutation-invariant manner. Experimental results show our method can produce better quality HDR images with up to 1.1dB PSNR improvement to the state-of-the-art, and subjective improvements in terms of better detail, colours, and fewer artefacts. Sibi Catley-Chandar, Thomas Tanay, Lucas Vandroux, Ales Leonardis, Gregory Slabaugh, Eduardo Pérez-Pellitero |
IEEE Trans. Image Process. | 5 |
| 2020 | TESA: Tensor Element Self-Attention via MatricizationabstractRepresentation learning is a fundamental part of modern computer vision, where abstract representations of data are encoded as tensors optimized to solve problems like image segmentation and inpainting. Recently, self-attention in the form of Non-Local Block has emerged as a powerful technique to enrich features, by capturing complex interdependencies in feature tensors. However, standard self-attention approaches leverage only spatial relationships, drawing similarities between vectors and overlooking correlations between channels. In this paper, we introduce a new method, called Tensor Element Self-Attention (TESA) that generalizes such work to capture interdependencies along all dimensions of the tensor using matricization. An order R tensor produces R results, one for each dimension. The results are then fused to produce an enriched output which encapsulates similarity among tensor elements. Additionally, we analyze self-attention mathematically, providing new perspectives on how it adjusts the singular values of the input feature tensor. With these new insights, we present experimental results demonstrating how TESA can benefit diverse problems including classification and instance segmentation. By simply adding a TESA module to existing networks, we substantially improve competitive baselines and set new state-of-the-art results for image inpainting on Celeb and low light raw-to-rgb image translation on SID. Francesca Babiloni, Ioannis Marras, Gregory Slabaugh, Stefanos Zafeiriou |
CVPR | 3 |
| 2020 | Video Super-Resolution With Temporal Group AttentionabstractVideo super-resolution, which aims at producing a high-resolution video from its corresponding low-resolution version, has recently drawn increasing attention. In this work, we propose a novel method that can effectively incorporate temporal information in a hierarchical way. The input sequence is divided into several groups, with each one corresponding to a kind of frame rate. These groups provide complementary information to recover missing details in the reference frame, which is further integrated with an attention module and a deep intra-group fusion module. In addition, a fast spatial alignment is proposed to handle videos with large motion. Extensive results demonstrate the capability of the proposed model in handling videos with various motion. It achieves favorable performance against state-of-the-art methods on several benchmark datasets. Takashi Isobe, Songjiang Li, Xu Jia 0012, Shanxin Yuan, Gregory Slabaugh, Chunjing Xu, Yali Li 0001, Shengjin Wang, Qi Tian 0001 |
CVPR | 5 |
| 2020 | A Multi-Hypothesis Approach to Color ConstancyabstractContemporary approaches frame the color constancy problem as learning camera specific illuminant mappings. While high accuracy can be achieved on camera specific data, these models depend on camera spectral sensitivity and typically exhibit poor generalisation to new devices. Additionally, regression methods produce point estimates that do not explicitly account for potential ambiguities among plausible illuminant solutions, due to the ill-posed nature of the problem. We propose a Bayesian framework that naturally handles color constancy ambiguity via a multi-hypothesis strategy. Firstly, we select a set of candidate scene illuminants in a data-driven fashion and apply them to a target image to generate a set of corrected images. Secondly, we estimate, for each corrected image, the likelihood of the light source being achromatic using a camera-agnostic CNN. Finally, our method explicitly learns a final illumination estimate from the generated posterior probability distribution. Our likelihood estimator learns to answer a camera-agnostic question and thus enables effective multi-camera training by disentangling illuminant estimation from the supervised learning task. We extensively evaluate our proposed approach and additionally set a benchmark for novel sensor generalisation without re-training. Our method provides state-of-the-art accuracy on multiple public datasets (up to 11% median angular error improvement) while maintaining real-time execution. Daniel Hernández Juárez, Sarah Parisot, Benjamin Busam, Ales Leonardis, Gregory Slabaugh, Steven McDonagh 0001 |
CVPR | 5 |
| 2020 | Unsupervised Model Personalization While Preserving Privacy and Scalability: An Open ProblemabstractThis work investigates the task of unsupervised model personalization, adapted to continually evolving, unlabeled local user images. We consider the practical scenario where a high capacity server interacts with a myriad of resource-limited edge devices, imposing strong requirements on scalability and local data privacy. We aim to address this challenge within the continual learning paradigm and provide a novel Dual User-Adaptation framework (DUA) to explore the problem. This framework flexibly disentangles user-adaptation into model personalization on the server and local data regularization on the user device, with desirable properties regarding scalability and privacy constraints. First, on the server, we introduce incremental learning of task-specific expert models, subsequently aggregated using a concealed unsupervised user prior. Aggregation avoids retraining, whereas the user prior conceals sensitive raw user data, and grants unsupervised adaptation. Second, local user-adaptation incorporates a domain adaptation point of view, adapting regularizing batch normalization parameters to the user data. We explore various empirical user configurations with different priors in categories and a tenfold of transforms for MIT Indoor Scene recognition, and classify numbers in a combined MNIST and SVHN setup. Extensive experiments yield promising results for data-driven local adaptation and elicit user priors for server adaptation to depend on the model rather than user data. Hence, although user-adaptation remains a challenging open problem, the DUA framework formalizes a principled foundation for personalizing both on server and user device, while maintaining privacy and scalability. Matthias De Lange, Xu Jia 0012, Sarah Parisot, Ales Leonardis, Gregory Slabaugh, Tinne Tuytelaars |
CVPR | 5 |
| 2020 | DeepLPF: Deep Local Parametric Filters for Image EnhancementabstractDigital artists often improve the aesthetic quality of digital photographs through manual retouching. Beyond global adjustments, professional image editing programs provide local adjustment tools operating on specific parts of an image. Options include parametric (graduated, radial filters) and unconstrained brush tools. These highly expressive tools enable a diverse set of local image enhancements. However, their use can be time consuming, and requires artistic capability. State-of-the-art automated image enhancement approaches typically focus on learning pixel-level or global enhancements. The former can be noisy and lack interpretability, while the latter can fail to capture fine-grained adjustments. In this paper, we introduce a novel approach to automatically enhance images using learned spatially local filters of three different types (Elliptical Filter, Graduated Filter, Polynomial Filter). We introduce a deep neural network, dubbed Deep Local Parametric Filters (DeepLPF), which regresses the parameters of these spatially localized filters that are then automatically applied to enhance the image. DeepLPF provides a natural form of model regularization and enables interpretable, intuitive adjustments that lead to visually pleasing results. We report on multiple benchmarks and show that DeepLPF produces state-of-the-art performance on two variants of the MIT-Adobe 5k dataset, often using a fraction of the parameters required for competing methods. Sean Moran, Pierre Marza, Steven McDonagh 0001, Sarah Parisot, Gregory Slabaugh |
CVPR | 5 |
| 2020 | Image Demoireing with Learnable Bandpass FiltersabstractImage demoireing is a multi-faceted image restoration task involving both texture and color restoration. In this paper, we propose a novel multiscale bandpass convolutional neural network (MBCNN) to address this problem. As an end-to-end solution, MBCNN respectively solves the two sub-problems. For texture restoration, we propose a learnable bandpass filter (LBF) to learn the frequency prior for moire texture removal. For color restoration, we propose a two-step tone mapping strategy, which first applies a global tone mapping to correct for a global color shift, then performs local fine tuning of the color per pixel. Through an ablation study, we demonstrate the effectiveness of the different components of MBCNN. Experimental results on two public datasets show that our method outperforms state-of-the-art methods by a large margin (more than 2dB in terms of PSNR). Bolun Zheng, Shanxin Yuan, Gregory Slabaugh, Ales Leonardis |
CVPR | 3 |
| 2020 | More Classifiers, Less Forgetting: A Generic Multi-classifier Paradigm for Incremental Learning
Yu Liu 0012, Sarah Parisot, Gregory Slabaugh, Xu Jia 0012, Ales Leonardis, Tinne Tuytelaars |
ECCV (26) | 3 |
| 2020 | Wavelet-Based Dual-Branch Network for Image Demoiréing
Lin Liu 0016, Jianzhuang Liu, Shanxin Yuan, Gregory Slabaugh, Ales Leonardis, Wengang Zhou 0001, Qi Tian 0001 |
ECCV (13) | 4 |
| 2020 | Reconstructing the Noise Variance Manifold for Image Denoising
Ioannis Marras, Grigorios Chrysos 0002, Ioannis Alexiou, Gregory Slabaugh, Stefanos Zafeiriou |
ECCV (9) | 4 |
| 2020 | Low Light Video Enhancement Using Synthetic Data Produced with an Intermediate Domain Mapping
Danai Triantafyllidou, Sean Moran, Steven McDonagh 0001, Sarah Parisot, Gregory Slabaugh |
ECCV (13) | 5 |
| 2020 | PROPEL: Probabilistic Parametric Regression Loss for Convolutional Neural NetworksabstractIn recent years, Convolutional Neural Networks (CNNs) have enabled significant advancements to the state-of-the-art in computer vision. For classification tasks, CNNs have widely employed probabilistic output and have shown the significance of providing additional confidence for predictions. However, such probabilistic methodologies are not widely applicable for addressing regression problems using CNNs, as regression involves learning unconstrained continuous and, in many cases, multi-variate target variables. We propose a PRObabilistic Parametric rEgression Loss (PROPEL) that facilitates CNNs to learn parameters of probability distributions for addressing probabilistic regression problems. PROPEL is fully differentiable and, hence, can be easily incorporated for end - to-end training of existing CNN regression architectures using existing optimization algorithms. The proposed method is flexible as it enables learning complex unconstrained probabilities while being generalizable to higher dimensional multi-variate regression problems. We utilize a PROPEL-based CNN to address the problem of learning hand and head orientation from uncalibrated color images. Our experimental validation and comparison with existing CNN regression loss functions show that PROPEL improves the accuracy of a CNN by enabling probabilistic regression, while significantly reducing required model parameters by 10×, resulting in improved generalization as compared to the existing state-of-the-art. Muhammad Asad 0001, Rilwan Remilekun Basaru, S. M. Masudur Rahman Al-Arif, Gregory Slabaugh |
ICPR | 4 |
| 2020 | CURL: Neural Curve Layers for Global Image EnhancementabstractWe present a novel approach to adjust global image properties such as colour, saturation, and luminance using human-interpretable image enhancement curves, inspired by the Photoshop curves tool. Our method, dubbed neural CURve Layers (CURL), is designed as a multi-colour space neural retouching block trained jointly in three different colour spaces (HSV, CIELab, RGB) guided by a novel multi-colour space loss. The curves are fully differentiable and are trained end-to-end for different computer vision problems including photo enhancement (RGB-to-RGB) and as part of the image signal processing pipeline for image formation (RAW-to-RGB). To demonstrate the effectiveness of CURL we combine this global image transformation block with a pixel-level (local) image multi-scale encoder-decoder backbone network. In an extensive experimental evaluation we show that CURL produces state-of-the-art image quality versus recently proposed deep learning approaches in both objective and perceptual metrics, setting new state-of-the-art performance on multiple public datasets. Our code is publicly available at: https://github.com/sjmoran/CURL. Sean Moran, Steven McDonagh 0001, Gregory Slabaugh |
ICPR | 3 |
| 2020 | Self-Adaptively Learning to Demoiré from Focused and Defocused Image PairsabstractMoiré artifacts are common in digital photography, resulting from the interference between high-frequency scene content and the color filter array of the camera. Existing deep learning-based demoiréing methods trained on large scale datasets are limited in handling various complex moiré patterns, and mainly focus on demoiréing of photos taken of digital displays. Moreover, obtaining moiré-free ground-truth in natural scenes is difficult but needed for training. In this paper, we propose a self-adaptive learning method for demoiréing a high-frequency image, with the help of an additional defocused moiré-free blur image. Given an image degraded with moiré artifacts and a moiré-free blur image, our network predicts a moiré-free clean image and a blur kernel with a self-adaptive strategy that does not require an explicit training stage, instead performing test-time adaptation. Our model has two sub-networks and works iteratively. During each iteration, one sub-network takes the moiré image as input, removing moiré patterns and restoring image details, and the other sub-network estimates the blur kernel from the blur image. The two sub-networks are jointly optimized. Extensive experiments demonstrate that our method outperforms state-of-the-art methods and can produce high-quality demoiréd results. It can generalize well to the task of removing moiré artifacts caused by display screens. In addition, we build a new moiré dataset, including images with screen and texture moiré artifacts. As far as we know, this is the first dataset with real texture moiré patterns. Lin Liu 0016, Shanxin Yuan, Jianzhuang Liu, Liping Bao, Gregory Slabaugh, Qi Tian 0001 |
NeurIPS | 5 |
| 2018 | SPNet: Shape Prediction Using a Fully Convolutional Neural Network
S. M. Masudur Rahman Al-Arif, Karen Knapp, Gregory Slabaugh |
MICCAI (1) | 3 |
| 2018 | Data-driven recovery of hand depth using CRRF on stereo imagesabstractHand pose is emerging as an important interface for human–computer interaction. This study presents a data‐driven method to estimate a high‐quality depth map of a hand from a stereoscopic camera input by introducing a novel superpixel‐based regression framework that takes advantage of the smoothness of the depth surface of the hand. To this end, the authors introduce conditional regressive random forest (CRRF), a method that combines a conditional random field (CRF) and an RRF to model the mapping from a stereo red, green and blue image pair to a depth image. The RRF provides a unary term that adaptively selects different stereo‐matching measures as it implicitly determines matching pixels in a coarse‐to‐fine manner. While the RRF makes depth prediction for each superpixel independently, the CRF unifies the prediction of depth by modelling pairwise interactions between adjacent superpixels. Experimental results show that CRRF can generate a depth image more accurately than the leading contemporary techniques using an inexpensive stereo camera. Rilwan Remilekun Basaru, Christopher Child, Eduardo Alonso 0001, Gregory Slabaugh |
IET Comput. Vis. | 4 |
| 2018 | DAGAN: Deep De-Aliasing Generative Adversarial Networks for Fast Compressed Sensing MRI ReconstructionabstractCompressed sensing magnetic resonance imaging (CS-MRI) enables fast acquisition, which is highly desirable for numerous clinical applications. This can not only reduce the scanning cost and ease patient burden, but also potentially reduce motion artefacts and the effect of contrast washout, thus yielding better image quality. Different from parallel imaging-based fast MRI, which utilizes multiple coils to simultaneously receive MR signals, CS-MRI breaks the Nyquist-Shannon sampling barrier to reconstruct MRI images with much less required raw data. This paper provides a deep learning-based strategy for reconstruction of CS-MRI, and bridges a substantial gap between conventional non-learning methods working only on data from a single image, and prior knowledge from large training data sets. In particular, a novel conditional Generative Adversarial Networks-based model (DAGAN)-based model is proposed to reconstruct CS-MRI. In our DAGAN architecture, we have designed a refinement learning method to stabilize our U-Net based generator, which provides an end-to-end network to reduce aliasing artefacts. To better preserve texture and edges in the reconstruction, we have coupled the adversarial loss with an innovative content loss. In addition, we incorporate frequency-domain information to enforce similarity in both the image and frequency domains. We have performed comprehensive comparison studies with both conventional CS-MRI reconstruction methods and newly investigated deep learning approaches. Compared with these methods, our DAGAN method provides superior reconstruction with preserved perceptual image details. Furthermore, each image is reconstructed in about 5 ms, which is suitable for real-time processing. Guang Yang 0006, Simiao Yu, Hao Dong 0003, Gregory Slabaugh, Pier Luigi Dragotti, Xujiong Ye, Fangde Liu, Simon R. Arridge, Jennifer Keegan, Yike Guo, David N. Firmin |
IEEE Trans. Medical Imaging | 4 |
| 2017 | Probabilistic Spatial Regression using a Deep Fully Convolutional Neural Network
S. M. Masudur Rahman Al-Arif, Karen Knapp, Gregory Slabaugh |
BMVC | 3 |
| 2017 | SPORE: Staged Probabilistic Regression for Hand Orientation Inference
Muhammad Asad 0001, Gregory Slabaugh |
Comput. Vis. Image Underst. | 2 |
| 2017 | Patch-based corner detection for cervical vertebrae in X-ray images
S. M. Masudur Rahman Al-Arif, Muhammad Asad 0001, Michael Gundry, Karen Knapp, Gregory Slabaugh |
Signal Process. Image Commun. | 5 |
| 2016 | The Need for Knowledge Extraction: Understanding Harmful Gambling Behavior with Neural NetworksabstractResponsible gambling is a field of study that involves supporting gamblers so as to reduce the harm that their gambling activity might cause. Recently in the literature, machine learning algorithms have been introduced as a way to predict potentially harmful gambling based on patterns of gambling behavior, such as trends in amounts wagered and the time spent gambling. In this paper, neural network models are analyzed to help predict the outcome of a partial proxy for harmful gambling behavior: when a gambler “self-excludes”, requesting a gambling operator to prevent them from accessing gambling opportunities. Drawing on survey and interview insights from industry and public officials as to the importance of interpretability, a variant of the knowledge extraction algorithm TREPAN is proposed which can produce compact, human-readable logic rules efficiently, given a neural network trained on gambling data. To the best of our knowledge, this paper reports the first industrial-strength application of knowledge extraction from neural networks, which otherwise are black-boxes unable to provide the explanatory insights which are crucially required in this area of application. We show that through knowledge extraction one can explore and validate the kinds of behavioral and demographic profiles that best predict self-exclusion, while developing a machine learning approach with greater potential for adoption by industry and treatment providers. Experimental results reported in this paper indicate that the rules extracted can achieve high fidelity to the trained neural network while maintaining competitive accuracy and providing useful insight to domain experts in responsible gambling. Christian Percy, Artur S. d'Avila Garcez, Simo Dragicevic, Manoel V. M. França, Gregory Slabaugh, Tillman Weyde |
ECAI | 5 |
| 2016 | Super-Resolved Enhancement of a Single Image and Its Application in Cardiac MRI
Guang Yang 0006, Xujiong Ye, Gregory Slabaugh, Jennifer Keegan, Raad Mohiaddin, David N. Firmin |
ICISP | 3 |
| 2015 | Generatinga 3D hand model from frontal color and range scansabstractRealistic 3D modeling of human hand anatomy has a number of important applications, including real-time tracking, pose estimation, and human-computer interaction. However the use of RGB-D sensors to accurately capture the full 3D shape of a hand is limited by self-occlusions, relatively smaller size of the hand and the requirement to capture multiple images. In this paper, we propose a method for generating a detailed, realistic hand model from a single frontal range scan and registered color image. In essence, our method converts this 2.5D data into a fully 3D model. The proposed approach extracts joint locations from the color image using a fingertip and interfinger region detector with a Naive Bayes probabilistic model. Direct correspondence between these joint locations in the range scan and a synthetic hand model are used to perform rigid registration, followed by a thin-plate-spline deformation that non-rigidly registers a synthetic model. This reconstructed model maintains similar geometric properties as the range scan, but also includes the back side of the hand. Experimental results demonstrate the promise of the method to produce detailed and realistic 3D hand geometry. Muhammad Asad 0001, Enguerrand Gentet, Rilwan Remilekun Basaru, Gregory Slabaugh |
ICIP | 4 |
| 2013 | Endoluminal surface registration for CT colonography using haustral fold matchingabstractComputed Tomographic (CT) colonography is a technique used for the detection of bowel cancer or potentially precancerous polyps. The procedure is performed routinely with the patient both prone and supine to differentiate fixed colonic pathology from mobile faecal residue. Matching corresponding locations is difficult and time consuming for radiologists due to colonic deformations that occur during patient repositioning. We propose a novel method to establish correspondence between the two acquisitions automatically. The problem is first simplified by detecting haustral folds using a graph cut method applied to a curvature-based metric applied to a surface mesh generated from segmentation of the colonic lumen. A virtual camera is used to create a set of images that provide a metric for matching pairs of folds between the prone and supine acquisitions. Image patches are generated at the fold positions using depth map renderings of the endoluminal surface and optimised by performing a virtual camera registration over a restricted set of degrees of freedom. The intensity difference between image pairs, along with additional neighbourhood information to enforce geometric constraints over a 2D parameterisation of the 3D space, are used as unary and pair-wise costs respectively, and included in a Markov Random Field (MRF) model to estimate the maximum a posteriori fold labelling assignment. The method achieved fold matching accuracy of 96.0% and 96.1% in patient cases with and without local colonic collapse. Moreover, it improved upon an existing surface-based registration algorithm by providing an initialisation. The set of landmark correspondences is used to non-rigidly transform a 2D source image derived from a conformal mapping process on the 3D endoluminal surface mesh. This achieves full surface correspondence between prone and supine views and can be further refined with an intensity based registration showing a statistically significant improvement (p<0.001), and decreasing mean error from 11.9 mm to 6.0 mm measured at 1743 reference points from 17 CTC datasets. Thomas Hampshire, Holger Roth, Emma Helbren, Andrew Plumb, Darren Boone, Gregory Slabaugh, Steve Halligan, David J. Hawkes |
Medical Image Anal. | 6 |
| 2013 | Erosion band signatures for spatial extraction of features
Eduard Vazquez, Xiaoyun Yang, Gregory Slabaugh |
Mach. Vis. Appl. | 3 |
| 2011 | Automatic Prone to Supine Haustral Fold Matching in CT Colonography Using a Markov Random Field Model
Thomas Hampshire, Holger Roth, Mingxing Hu, Darren Boone, Gregory Slabaugh, Shonit Punwani, Steve Halligan, David J. Hawkes |
MICCAI (1) | 5 |
| 2011 | Generating shapes by analogies: An application to hearing aid design
Gozde Unal, Delphine Nain, Gregory Slabaugh, Tong Fang |
Comput. Aided Des. | 3 |
| 2011 | Automatic Detection of Bridge Deck Condition From Ground Penetrating Radar ImagesabstractAccurate assessment of the quality of concrete bridge decks and identification of corrosion induced delamination lead to economic management of bridge decks. It has been demonstrated that ground penetrating radar (GPR) can be successfully used for such purposes. The growing demand on GPR has brought into the challenge of developing automatic processes necessary to produce a final accurate interpretation. However, there have been few publications targeting at automatic detection of bridge deck delamination from GPR data. This paper proposes a novel method using partial differential equations to detect rebar (or steel-bar) mat signatures of concrete bridges from GPR data so that the delamination within the bridge deck can be effectively located. The proposed algorithm was tested on both synthetic and real GPR images and the experimental results have demonstrated its accuracy and reliability, even for diminished image contrast and low signal-to-noise ratio. Therefore, an accurate deterioration map of the bridge deck can be generated automatically. Zhe Wendy Wang, MengChu Zhou, Gregory Slabaugh, Jiefu Zhai, Tong Fang |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2010 | Automatic Detection of Anatomical Features on 3D Ear Impressions for Canonical Representation
Sajjad Baloch, Rupen Melkisetoglu, Simon Flöry, Sergei Azernikov, Gregory Slabaugh, Alexander Zouhar, Tong Fang |
MICCAI (3) | 5 |
| 2010 | Establishing Spatial Correspondence between the Inner Colon Surfaces from Prone and Supine CT Colonography
Holger Roth, Jamie McClelland, Marc Modat, Darren Boone, Mingxing Hu, Sébastien Ourselin, Gregory Slabaugh, Steve Halligan, David J. Hawkes |
MICCAI (3) | 7 |
| 2010 | 3D ball skinning using PDEs for generation of smooth tubular surfaces
Gregory Slabaugh, Brian Whited, Jarek Rossignac, Tong Fang, Gozde Unal |
Comput. Aided Des. | 1 |
| 2009 | Fast pseudo-enhancement correction in CT colonography using linear shift-invariant filtersabstractThis paper presents a novel method to approximate shift-variant Gaussian filtering of an image using a set of shift-invariant Gaussian filters. This approximation affords filtering of the image using fast convolution techniques that rely on the FFT, while achieving a result that closely matches the shift-variant result. We demonstrate the method in a CT colonography application that reduces the pseudo-enhancement effect, which is a local brightening artifact in CT imaging that can result from the use of oral contrast agents. Experimental results demonstrate the effectiveness of the method and emphasize its computational efficiency. Richard Boyes, Gregory Slabaugh, Gareth Beddoe |
ICIP | 2 |
| 2008 | Variational Skinning of an Ordered Set of Discrete 2D Balls
Gregory Slabaugh, Gozde Unal, Tong Fang, Jarek Rossignac, Brian Whited |
GMP | 1 |
| 2008 | Customized Design of Hearing Aids Using Statistical Shape Learning
Gozde Unal, Delphine Nain, Gregory Slabaugh, Tong Fang |
MICCAI (1) | 3 |
| 2008 | Shape-Driven Segmentation of the Arterial Wall in Intravascular Ultrasound ImagesabstractSegmentation of arterial wall boundaries from intravascular images is an important problem for many applications in the study of plaque characteristics, mechanical properties of the arterial wall, its 3-D reconstruction, and its measurements such as lumen size, lumen radius, and wall radius. We present a shape-driven approach to segmentation of the arterial wall from intravascular ultrasound images in the rectangular domain. In a properly built shape space using training data, we constrain the lumen and media-adventitia contours to a smooth, closed geometry, which increases the segmentation quality without any tradeoff with a regularizer term. In addition to a shape prior, we utilize an intensity prior through a nonparametric probability-density-based image energy, with global image measurements rather than pointwise measurements used in previous methods. Furthermore, a detection step is included to address the challenges introduced to the segmentation process by side branches and calcifications. All these features greatly enhance our segmentation method. The tests of our algorithm on a large dataset demonstrate the effectiveness of our approach. Gozde Unal, Susann Bucher, Stephane G. Carlier, Gregory Slabaugh, Tong Fang, Kaoru Tanaka |
IEEE Trans. Inf. Technol. Biomed. | 4 |
| 2008 | Guest Editorial Introduction to the Special Section on Computer Vision for Intravascular and Intracardiac ImagingabstractThe ten papers in this special section focus on computer vision for intravascular and intracardiac imaging. The papers are summarized here. Gozde Unal, Gregory Slabaugh, Ioannis A. Kakadiaris, Allen R. Tannenbaum |
IEEE Trans. Inf. Technol. Biomed. | 2 |
| 2007 | A Variational Approach to the Evolution of Radial Basis Functions for Image SegmentationabstractIn this paper we derive differential equations for evolving radial basis functions (RBFs) to solve segmentation problems. The differential equations result from applying variational calculus to energy functionals designed for image segmentation. Our methodology supports evolution of all parameters of each RBF, including its position, weight, orientation, and anisotropy, if present. Our framework is general and can be applied to numerous RBF interpolants. The resulting approach retains some of the ideal features of implicit active contours, like topological adaptivity, while requiring low storage overhead due to the sparsity of our representation, which is an unstructured list of RBFs. We present the theory behind our technique and demonstrate its usefulness for image segmentation. Gregory Slabaugh, Huong Quynh Dinh, Gozde Unal |
CVPR | 1 |
| 2007 | Variational Guidewire Tracking Using Phase Congruency
Gregory Slabaugh, Koon Kong, Gozde Unal, Tong Fang |
MICCAI (2) | 1 |
| 2007 | An information-theoretic detector based scheme for registration of speckled medical imagesabstractSeveral studies dealt with medical ultrasound registration. Their similarity metrics relied on pixel-to-pixel intensity comparisons. Hence, they are not well suited to the case of speckled images. To better handle the speckle noise, our previous work proposed an information-theoretic feature detector-based registration approach. This work aims to extend it to the cases where the image speckle model is Rayleigh or normalized Fisher-Tippett distributed. Using speckle modeling based on these distributions, a speckle-speci.c informationtheoretic feature detector is constructed and applied to provide feature images. Those feature images are then registered using differential equations, the solution of which provides a transformation to bring the images into alignment. Compared to standard gradient-based techniques, the experimental results demonstrate the effectiveness of our method, particularly for low contrast ultrasound images. Zhe Wendy Wang, Gregory Slabaugh, Gozde Unal, MengChu Zhou, Tong Fang |
SMC | 2 |
| 2007 | A Variational Approach to Problems in Calibration of Multiple CamerasabstractThis paper addresses the problem of calibrating camera parameters using variational methods. One problem addressed is the severe lens distortion in low-cost cameras. For many computer vision algorithms aiming at reconstructing reliable representations of 3D scenes, the camera distortion effects will lead to inaccurate 3D reconstructions and geometrical measurements if not accounted for. A second problem is the color calibration problem caused by variations in camera responses that result in different color measurements and affects the algorithms that depend on these measurements. We also address the extrinsic camera calibration that estimates relative poses and orientations of multiple cameras in the system and the intrinsic camera calibration that estimates focal lengths and the skew parameters of the cameras. To address these calibration problems, we present multiview stereo techniques based on variational methods that utilize partial and ordinary differential equations. Our approach can also be considered as a coordinated refinement of camera calibration parameters. To reduce computational complexity of such algorithms, we utilize prior knowledge on the calibration object, making a piecewise smooth surface assumption, and evolve the pose, orientation, and scale parameters of such a 3D model object without requiring a 2D feature extraction from camera views. We derive the evolution equations for the distortion coefficients, the color calibration parameters, the extrinsic and intrinsic parameters of the cameras, and present experimental results. Gozde Unal, Anthony J. Yezzi, Stefano Soatto, Gregory Slabaugh |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2006 | Ultrasound-Specific Segmentation via Decorrelation and Statistical Region-Based Active ContoursabstractSegmentation of ultrasound images is often a very challenging task due to speckle noise that contaminates the image. It is well known that speckle noise exhibits an asymmetric distribution as well as significant spatial correlation. Since these attributes can be difficult to model, many previous ultrasound segmentation methods oversimplify the problem by assuming that the noise is white and/or Gaussian, resulting in generic approaches that are actually more suitable to MR and X-ray segmentation than ultrasound. Unlike these methods, in this paper we present an ultrasound-specific segmentation approach that first decorrelates the image, and then performs segmentation on the whitened result using statistical region-based active contours. In particular, we design a gradient ascent flow that evolves the active contours to maximize a log likelihood functional based on the Fisher-Tippett distribution. We present experimental results that demonstrate the effectiveness of our method. Gregory Slabaugh, Gozde Unal, Tong Fang, Michael Wels |
CVPR (1) | 1 |
| 2006 | Semi-Automatic 3-D Segmentation of Anatomical Structures of Brain MRI Volumes using Graph CutsabstractWe present a semi-automatic segmentation technique of the anatomical structures of the brain: cerebrum, cerebellum, and brain stem. The method uses graph cuts segmentation with an anatomic template for initialization. First, a skull stripping procedure is applied to remove non-brain tissues. Then, the segmentation is done hierarchically by first, extracting first the cerebrum from the brain, and then from the remaining volume the cerebellum and the brain stem are separated. This method is fast and can separate different anatomical structures of the brain in spite of weak boundaries. We describe our approach and present experimental results demonstrating its usefulness. Huy-Nam Doan, Gregory Slabaugh, Gozde Unal, Tong Fang |
ICIP | 2 |
| 2006 | Semi-Automatic Lymph Node Segmentation in LN-MRIabstractAccurate staging of nodal cancer still relies on surgical exploration because many primary malignancies spread via lymphatic dissemination. The purpose of this study was to utilize nanoparticle-enhanced lymphotropic magnetic resonance imaging (LN-MRI) to explore semi-automated noninvasive nodal cancer staging. We present a joint image segmentation and registration approach, which makes use of the problem specific information to increase the robustness of the algorithm to noise and weak contrast often observed in medical imaging applications. The effectiveness of the approach is demonstrated with a given lymph node segmentation problem in post-contrast pelvic MRI sequences. Gozde Unal, Gregory Slabaugh, Andreas Ess, Anthony J. Yezzi, Tong Fang, Jason Tyan, Martin Requardt, Robert Krieg, Ravi T. Seethamraju, Mukesh Harisinghani, Ralph Weissleder |
ICIP | 2 |
| 2005 | Active Polyhedron: Surface Evolution Theory Applied to Deformable MeshesabstractThis paper presents a novel 3D deformable surface that we call an active polyhedron. Rooted in surface evolution theory, an active polyhedron is a polyhedral surface whose vertices deform to minimize a regional and/or boundary-based energy functional. Unlike continuous active surface models, the vertex motion of an active polyhedron is computed by integrating speed terms over polygonal faces of the surface. The resulting ordinary differential equations (ODEs) provide improved robustness to noise and allow for larger time steps compared to continuous active surfaces implemented with level set methods. We describe an electrostatic regularization technique that achieves global regularization while better preserving sharper local features. Experimental results demonstrate the effectiveness of an active polyhedron in solving segmentation problems as well as surface reconstruction from unorganized points. Gregory Slabaugh, Gozde Unal |
CVPR (2) | 1 |
| 2005 | Coupled PDEs for Non-Rigid Registration and SegmentationabstractIn this paper we present coupled partial differential equations (PDEs) for the problem of joint segmentation and registration. The registration component of the method estimates a deformation field between boundaries of two structures. The desired coupling comes from two PDEs that estimate a common surface through segmentation and its non-rigid registration with a target image. The solutions of these two PDEs both decrease the total energy of the surface, and therefore aid each other in finding a locally optimal solution. Our technique differs from recently popular joint segmentation and registration algorithms, all of which assume a rigid transformation among shapes. We present both the theory and results that demonstrate the effectiveness of the approach. Gozde Unal, Gregory Slabaugh |
CVPR (1) | 2 |
| 2005 | Graph cuts segmentation using an elliptical shape priorabstractWe present a graph cuts-based image segmentation technique that incorporates an elliptical shape prior. Inclusion of this shape constraint restricts the solution space of the segmentation result, increasing robustness to misleading information that results from noise, weak boundaries, and clutter. We argue that combining a shape prior with a graph cuts method suggests an iterative approach that updates an intermediate result to the desired solution. We first present the details of our method and then demonstrate its effectiveness in segmenting vessels and lymph nodes from pelvic magnetic resonance images, as well as human faces. Gregory Slabaugh, Gozde Unal |
ICIP (2) | 1 |
| 2004 | Methods for Volumetric Reconstruction of Visual Scenes
Gregory Slabaugh, W. Bruce Culbertson, Thomas Malzbender, Mark R. Stevens, Ronald W. Schafer |
Int. J. Comput. Vis. | 1 |
| 2002 | Multi-resolution space carving using level set methodsabstractWe present a multi-resolution space carving algorithm that reconstructs a 3D model of a visual scene photographed by a calibrated digital camera placed at multiple viewpoints. Our approach employs a level set framework for reconstructing the scene. Unlike most standard space carving approaches, our level set approach produces a smooth reconstruction composed of manifold surfaces. Our method outputs a polygonal model, instead of a collection of voxels. We texture-map the reconstructed geometry using the photographs, and then render the model to produce photo-realistic new views of the scene. Gregory Slabaugh, Ronald W. Schafer, Mathieu Hans |
ICIP (2) | 1 |
| 2002 | Reconstructing Surfaces by Volumetric Regularization Using Radial Basis FunctionsabstractWe present a new method of surface reconstruction that generates smooth and seamless models from sparse, noisy, nonuniform, and low resolution range data. Data acquisition techniques from computer vision, such as stereo range images and space carving, produce 3D point sets that are imprecise and nonuniform when compared to laser or optical range scanners. Traditional reconstruction algorithms designed for dense and precise data do not produce smooth reconstructions when applied to vision-based data sets. Our method constructs a 3D implicit surface, formulated as a sum of weighted radial basis functions. We achieve three primary advantages over existing algorithms: (1) the implicit functions we construct estimate the surface well in regions where there is little data, (2) the reconstructed surface is insensitive to noise in data acquisition because we can allow the surface to approximate, rather than exactly interpolate, the data, and (3) the reconstructed surface is locally detailed, yet globally smooth, because we use radial basis functions that achieve multiple orders of smoothness. Huong Quynh Dinh, Greg Turk, Gregory Slabaugh |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2001 | Reconstructing Surfaces Using Anisotropic Basis FunctionsabstractPoint sets obtained from computer vision techniques are often noisy and non-uniform. We present a new method of surface reconstruction that can handle such data sets using anisotropic basis functions. Our reconstruction algorithm draws upon the work in variational implicit surfaces for constructing smooth and seamless 3D surfaces. Implicit functions are often formulated as a sum of weighted basis functions that are radially symmetric. Using radially symmetric basis functions inherently assumes, however that the surface to be reconstructed is, everywhere, locally symmetric. Such an assumption is true only at planar regions, and hence, reconstruction using isotropic basis is insufficient to recover objects that exhibit sharp features. We preserve sharp features using anisotropic basis that allow the surface to vary locally. The reconstructed surface is sharper along edges and at corner points. We determine the direction of anisotropy at a point by performing principal component analysis of the data points in a small neighborhood. The resulting field of principle directions across the surface is smoothed through tensor filtering. We have applied the anisotropic basis functions to reconstruct surfaces from noisy synthetic 3D data and from real range data obtained from space carving. Huong Quynh Dinh, Greg Turk, Gregory Slabaugh |
ICCV | 3 |