Nantheera Anantrasirichai

dblp:98/5858 · DBLP profile ↗
← Back
54ranked-venue papers
30as first author
26since 2021 · last 2026
0000-0002-2122-5781ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 42 · 23 first-author · 19 since 2021Artificial intelligence and machine learning · 8 · 5 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 4 since 2021Systems, architecture and hardware · 3 · 3 since 2021
YearPublicationVenuePosition
2026 Splatography: Sparse Multi-View Dynamic Gaussian Splatting for Filmmaking Challenges
abstract
Deformable Gaussian Splatting (GS) accomplishes photorealistic dynamic 3-D reconstruction from dense multi-view video (MVV) by learning to deform a canonical GS representation. However, in filmmaking, tight budgets can result in sparse camera configurations, which limits state-of-the-art (SotA) methods when capturing complex dynamic features. To address this issue, we introduce an approach that splits the canonical Gaussians and deformation field into foreground and background components using a sparse set of masks for frames at$t=0$. Each representation is separately trained on different loss functions during canonical pre-training. Then, during dynamic training, different parameters are modeled for each deformation field following common filmmaking practices. The foreground stage contains diverse dynamic features so changes in color, position and rotation are learned. While, the background containing film-crew and equipment, is typically dimmer and less dynamic so only changes in point position are learned. Experiments on 3-D and 2.5-D entertainment datasets show that our method produces SotA qualitative and quantitative results; up to 3 PSNR higher with half the model size on 3-D scenes. Unlike the SotA and without the need for dense mask supervision, our method also produces segmented dynamic reconstructions including transparent and dynamic textures. Code and video comparisons are available online: http://bit.ly/4oqzZrO
Adrian Azzarelli, Nantheera Anantrasirichai, David Bull 0001
3DV2
2026 Bayesian Neural Networks for One-to-Many Mapping in Image Enhancement
abstract
In image enhancement tasks, such as low-light and underwater image enhancement, a degraded image can correspond to multiple plausible target images due to dynamic photography conditions. This naturally results in a one-to-many mapping problem. To address this, we propose a Bayesian Enhancement Model (BEM) that incorporates Bayesian Neural Networks (BNNs) to capture data uncertainty and produce diverse outputs. To enable fast inference, we introduce a BNN-DNN framework: a BNN is first employed to model the one-to-many mapping in a low-dimensional space, followed by a Deterministic Neural Network (DNN) that refines fine-grained image details. Extensive experiments on multiple low-light and underwater image enhancement benchmarks demonstrate the effectiveness of our method.
Guoxi Huang, Ruirui Lin, Zipeng Qi, David Bull 0001, Nantheera Anantrasirichai
AAAI6
2026 RMFAT: Recurrent Multi-scale Feature Atmospheric Turbulence Mitigator
abstract
Atmospheric turbulence (AT) severely degrades video quality by introducing distortions such as geometric warping, blur, and temporal flickering, posing significant challenges to both visual clarity and temporal consistency. Current state-of-the-art methods are based on transformer, 3D architectures and require multi-frame input, but their large computational cost and memory usage limit real-time deployment. In this work, we propose RMFAT, Recurrent Multi-scale Feature Atmospheric Turbulence Mitigator, designed for efficient and temporally consistent video restoration under AT conditions. RMFAT adopts a lightweight recurrent framework that restores each frame using only two inputs at a time, significantly reducing temporal window size and computational burden. It further integrates multi-scale feature encoding and decoding with temporal warping modules at both encoder and decoder stages to enhance spatial detail and temporal coherence. Extensive experiments conducted on synthetic and real-world atmospheric turbulence datasets demonstrate that RMFAT not only outperforms existing methods in terms of clarity restoration (with nearly a 9% improvement in SSIM) but also achieves significantly improved inference speed (achieving a more than fourfold reduction), making it particularly suitable for real-time atmospheric turbulence suppression tasks.
Nantheera Anantrasirichai
AAAI2
2026 GFix: Perceptually Enhanced Gaussian Splatting Video Compression
abstract
3D Gaussian Splatting (3DGS) enhances 3D scene reconstruction through explicit representation and fast rendering, demonstrating potential benefits for various low-level vision tasks, including video compression. However, existing 3DGS-based video codecs generally exhibit more noticeable visual artifacts and relatively low compression ratios. In this paper, we specifically target the perceptual enhancement of 3DGS-based video compression, based on the assumption that artifacts from 3DGS rendering and quantization resemble noisy latents sampled during diffusion training. Building on this premise, we propose a content-adaptive framework, GFix, comprising a streamlined, single-step diffusion model that serves as an off-the-shelf neural enhancer. Moreover, to increase compression efficiency, We propose a modulated LoRA scheme that freezes the low-rank decompositions and modulates the intermediate hidden states, thereby achieving efficient adaptation of the diffusion backbone with highly compressible updates. Experimental results show that GFix delivers strong perceptual quality enhancement, outperforming GSVC with up to 72.1% BD-rate savings in LPIPS and 21.4% in FID.
Siyue Teng, Ge Gao 0005, Duolikun Danier, Yuxuan Jiang 0015, Fan Zhang 0017, Nantheera Anantrasirichai, Thomas Davis, Zoe Liu, David Bull 0001
ISCAS6
2026 DMAT: An End-to-End Framework for Joint Atmospheric Turbulence Mitigation and Object Detection
abstract
Atmospheric Turbulence (AT) degrades the clarity and accuracy of surveillance imagery, posing challenges not only for visualization quality but also for object classification and scene tracking. Deep learning-based methods have been proposed to improve visual quality, but spatio-temporal distortions remain a significant issue. Although deep learning-based object detection performs well under normal conditions, it struggles to operate effectively on sequences distorted by atmospheric turbulence. In this paper, we propose a novel framework that learns to compensate for distorted features while simultaneously improving visualization and object detection. This end-to-end training strategy leverages and exchanges knowledge of low-level distorted features in the AT mitigator with semantic features extracted in the object detector. Specifically, in the AT mitigator a 3D Mamba-based structure is used to handle the spatio-temporal displacements and blurring caused by turbulence. Optimization is achieved through back-propagation in both the AT mitigator and object detector. Our proposed DMAT outperforms state-of-the-art AT mitigation and object detection systems up to a 15% improvement on datasets corrupted by generated turbulence. The code is available at https://github.com/pui-nantheera/DMAT and datasets are available at https://zenodo.org/records/17673509.
Paul R. Hill, Alin Achim, David Bull 0001, Nantheera Anantrasirichai
WACV5
2026 ViVo: A Dataset for Human Volumetric Video Reconstruction and Compression
Adrian Azzarelli, Ge Gao 0005, Ho Man Kwan, Fan Zhang 0017, Nantheera Anantrasirichai, Oliver Moolan-Feroze, David Bull 0001
IEEE Trans. Circuits Syst. Video Technol.5
2025 MAMAT: 3D Mamba-Based Atmospheric Turbulence Removal and its Object Detection Capability
abstract
Restoration and enhancement are essential for improving the quality of videos captured under atmospheric turbulence conditions, aiding visualization, object detection, classification, and tracking in surveillance systems. In this paper, we introduce a novel Mamba-Based method, the 3D Mamba-Based Atmospheric Turbulence Removal (MAMAT), which employs a dual-module strategy to mitigate these distortions. The first module utilizes deformable 3D convolutions for non-rigid registration to minimize spatial shifts, while the second module enhances contrast and detail. Leveraging the advanced capabilities of the 3D Mamba architecture, experimental results demonstrate that MAMAT outperforms state-of-the-art learning-based methods, achieving up to a 3% improvement in visual quality and a 15% boost in object detection. It not only enhances visualization but also significantly improves object detection accuracy, bridging the gap between visual restoration and the effectiveness of surveillance applications. The code is available at https://github.com/csprh/MAMAT.
Paul R. Hill, Nantheera Anantrasirichai
AVSS3
2025 Multi-Scale Denoising in the Feature Space for Low-Light Instance Segmentation
abstract
Instance segmentation for low-light imagery remains largely unexplored due to the challenges imposed by such conditions, for example shot noise due to low photon count, color distortions and reduced contrast. In this paper, we propose an end-to-end solution to address this challenging task. Our proposed method implements weighted non-local blocks (wNLB) in the feature extractor. This integration enables an inherent denoising process at the feature level. As a result, our method eliminates the need for aligned ground truth images during training, thus supporting training on real-world low-light datasets. We introduce additional learnable weights at each layer in order to enhance the network’s adaptability to real-world noise characteristics, which affect different feature scales in different ways. Experimental results on several object detectors show that the proposed method outperforms the pre-trained networks with an Average Precision (AP) improvement of at least +7.6, with the introduction of wNLB further enhancing AP by upto +1.3.
Joanne Lin, Nantheera Anantrasirichai, David Bull 0001
ICASSP2
2025 BVI-CR: A Multi-View Human Dataset for Volumetric Video Compression
abstract
The advances in immersive technologies and 3D reconstruction have enabled the creation of digital replicas of real-world objects and environments with fine details. These processes generate vast amounts of 3D data, requiring more efficient compression methods to satisfy the memory and bandwidth constraints associated with data storage and transmission. However, the development and validation of efficient 3D data compression methods are constrained by the lack of comprehensive and high-quality volumetric video datasets, which typically require much more effort to acquire and consume increased resources compared to 2D image and video databases. To bridge this gap, we present an open multi-view volumetric human dataset, denoted BVI-CR, which contains 18 multi-view RGB-D captures and their corresponding textured polygonal meshes, depicting a range of diverse human actions. Each video sequence contains 10 views in 1080p resolution with durations between 10-15 seconds at 30FPS. Using BVI-CR, we benchmarked three conventional and neural coordinate-based multi-view video compression methods, following the MPEG MIV Common Test Conditions, and reported their rate quality performance based on various quality metrics. The results show the great potential of neural representation based methods in volumetric video compression compared to conventional video coding methods (with an up to 38% average coding gain in PSNR). This dataset provides a development and validation platform for a variety of tasks including volumetric reconstruction, compression, and quality assessment. The database will be shared publicly at https://github.com/fan-aaron-zhang/bvi-cr.
Ge Gao 0005, Adrian Azzarelli, Ho Man Kwan, Nantheera Anantrasirichai, Fan Zhang 0017, Will Andrew, Oliver Moolan-Feroze, David Bull 0001
ISCAS4
2025 AquaNeRF: Neural Radiance Fields in Underwater Media with Distractor Removal
abstract
Neural radiance field (NeRF) research has made significant progress in modeling static video content captured in the wild. However, current models and rendering processes rarely consider scenes captured underwater, which are useful for studying and filming ocean life. They fail to address visual artifacts unique to underwater scenes, such as moving fish and suspended particles. This paper introduces a novel NeRF renderer and optimization scheme for an implicit MLP-based NeRF model. Our renderer reduces the influence of floaters and moving objects that interfere with static objects of interest by estimating a single surface per ray. We use a Gaussian weight function with a small offset to ensure that the transmittance of the surrounding media remains constant. Additionally, we enhance our model with a depth-based scaling function to upscale gradients for near-camera volumes. Overall, our method outperforms the baseline Nerfacto by approximately 7.5% and SeaThru-NeRF by 6.2% in terms of PSNR. Subjective evaluation also shows a significant reduction of artifacts while preserving details of static targets and background compared to the state of the arts.
Luca Gough, Adrian Azzarelli, Fan Zhang 0017, Nantheera Anantrasirichai
ISCAS4
2025 Guiding WaveMamba with Frequency Maps for Image Debanding
Xinyi Wang 0011, Smaranda Tasmoc, Nantheera Anantrasirichai, Angeliki V. Katsenou
PCS3
2025 UW-GS: Distractor-Aware 3D Gaussian Splatting for Enhanced Underwater Scene Reconstruction
abstract
3D Gaussian splatting (3DGS) offers the capability to achieve real-time high quality 3D scene rendering. However, 3DGS assumes that the scene is in a clear medium environment and struggles to generate satisfactory representations in underwater scenes, where light absorption and scattering are prevalent and moving objects are involved. To overcome these, we introduce a novel Gaussian Splatting-based method, UW-GS, designed specifically for underwater applications. It introduces a color appearance that models distance-dependent color variation, employs a new physics-based density control strategy to enhance clarity for distant objects, and uses a binary motion mask to handle dynamic content. Optimized with a well-designed loss function supporting for scattering media and strengthened by pseudo-depth maps, UW-GS outperforms existing methods with PSNR gains up to 1.26dB. To fully verify the effectiveness of the model, we also developed a new underwater dataset, S-UW, with dynamic object masks. The code of UW-GS and S-UW will be available at https://github.com/WangHaoran16/UW-GS.
Nantheera Anantrasirichai, Fan Zhang 0017, David Bull 0001
WACV2
2024 A Spatio-Temporal Aligned SUNet Model For Low-Light Video Enhancement
abstract
Distortions caused by low-light conditions are not only visually unpleasant but also degrade the performance of computer vision tasks. The restoration and enhancement have proven to be highly beneficial. However, there are only a limited number of enhancement methods explicitly designed for videos acquired in low-light conditions. We propose a Spatio-Temporal Aligned SUNet (STA-SUNet) model using a Swin Transformer as a backbone to capture low light video features and exploit their spatio-temporal correlations. The STA-SUNet model is trained on a novel, fully registered dataset (BVI), which comprises dynamic scenes captured under varying light conditions. It is further analysed comparatively against various other models over three test datasets. The model demonstrates superior adaptivity across all datasets, obtaining the highest PSNR and SSIM values. It is particularly effective in extreme low-light conditions, yielding fairly good visualisation results.
Ruirui Lin, Nantheera Anantrasirichai, Alexandra Malyugina, David Bull 0001
ICIP2
2024 Anomaly Detection for the Identification of Volcanic Unrest in Satellite Imagery
abstract
Satellite images have the potential to detect volcanic deformation prior to eruptions, but while a vast number of images are routinely acquired, only a small percentage contain volcanic deformation events. Manual inspection could miss these anomalies, and an automatic system modelled with supervised learning requires suitably labelled datasets. To tackle these issues, this paper explores the use of unsupervised deep learning on satellite data for the purpose of identifying volcanic deformation as anomalies. Our detector is based on Patch Distribution Modeling (PaDiM), and the detection performance is enhanced with a weighted distance, assigning greater importance to features from deeper layers. Additionally, we propose a preprocessing approach to handle noisy and incomplete data points. The final framework was tested with five volcanoes, which have different deformation characteristics and its performance was compared against the supervised learning method for volcanic deformation detection.
Robert Gabriel Popescu, Nantheera Anantrasirichai, Juliet Biggs
ICIP2
2024 TaGAT: Topology-Aware Graph Attention Network for Multi-modal Retinal Image Fusion
Xin Tian 0009, Nantheera Anantrasirichai, Lindsay Nicholson, Alin Achim
MICCAI (1)2
2023 ST-MFNET Mini: Knowledge Distillation-Driven Frame Interpolation
abstract
Currently, one of the major challenges in deep learning-based video frame interpolation (VFI) is the large model size and high computational complexity associated with many high performance VFI approaches. In this paper, we present a distillation-based two-stage workflow for obtaining compressed VFI models which perform competitively compared to the state of the art, but with significantly reduced model size and complexity. Specifically, an optimization-based network pruning method is applied to a state of the art frame interpolation model, ST-MFNet, which suffers from large model size. The resulting network architecture achieves a 91% reduction in parameter numbers and a 35% increase in speed. The performance of the new network is further enhanced through a teacher-student knowledge distillation training process using a Laplacian distillation loss. The final low complexity model, ST-MFNet Mini, achieves a comparable performance to most existing high-complexity VFI methods, only outperformed by the original ST-MFNet. Our source code is available at https://github.com/crispianm/ST-MFNet-Mini
Crispian Morris, Duolikun Danier, Fan Zhang 0017, Nantheera Anantrasirichai, David Bull 0001
ICIP4
2023 Atmospheric turbulence removal with complex-valued convolutional neural network
abstract
Atmospheric turbulence distorts visual imagery and is always problematic for information interpretation by both human and machine. Most well-developed approaches to remove atmospheric turbulence distortion are model-based. However, these methods require high computation and large memory making real-time operation infeasible. Deep learning-based approaches have hence gained more attention but currently work efficiently only on static scenes. This paper presents a novel learning-based framework offering short temporal spanning to support dynamic scenes. We exploit complex-valued convolutions as phase information, altered by atmospheric turbulence, is captured better than using ordinary real-valued convolutions. Two concatenated modules are proposed. The first module aims to remove geometric distortions and, if enough memory, the second module is applied to refine micro details of the videos. Experimental results show that our proposed framework efficiently mitigates the atmospheric turbulence distortion and significantly outperforms existing methods.
Nantheera Anantrasirichai
Pattern Recognit. Lett.1
2023 A topological loss function for image Denoising on a new BVI-lowlight dataset
abstract
Although image denoising algorithms have attracted significant research attention, surprisingly few have been proposed for, or evaluated on, noise from imagery acquired under real low-light conditions. Moreover, noise characteristics are often assumed to be spatially invariant, leading to edges and textures being distorted after denoising. Here, we introduce a novel topological loss function which is based on persistent homology. The method performs in the space of image patches, where topological invariants are calculated and represented in persistent diagrams. The loss function is a combination of ℓ1 or ℓ2 losses with the new persistence-based topological loss. We compare its performance across popular denoising architectures and loss functions, training the networks on our new comprehensive dataset of natural images captured in low-light conditions – BVI-LOWLIGHT. Analysis reveals that this approach outperforms existing methods, adapting well to complex structures and suppressing common artifacts.
Alexandra Malyugina, Nantheera Anantrasirichai, David Bull 0001
Signal Process.2
2022 DeTurb: Atmospheric Turbulence Mitigation with Deformable 3D Convolutions and 3D Swin Transformers
Zhicheng Zou, Nantheera Anantrasirichai
ACCV (4)2
2022 ICIP 2022 Challenge on Parasitic Egg Detection and Classification in Microscopic Images: Dataset, Methods and Results
abstract
Manual examination of faecal smear samples to identify the existence of parasitic eggs is very time-consuming and can only be done by specialists. Therefore, an automated system is required to tackle this problem since it can relate to serious intestinal parasitic infections. This paper reviews the ICIP 2022 Challenge on parasitic egg detection and classification in microscopic images. We describe a new dataset for this application, which is the largest dataset of its kind. The methods used by participants in the challenge are summarised and discussed along with their results.
Nantheera Anantrasirichai, Thanarat H. Chalidabhongse, Duangdao Palasuwan, Korranat Naruenatthanaset, Thananop Kobchaisawat, Nuntiporn Nunthanasup, Kanyarat Boonpeng, Xudong Ma, Alin Achim
ICIP1
2022 Unsupervised Image Fusion Using Deep Image Priors
abstract
A significant number of researchers have applied deep learning methods to image fusion. However, most works require a large amount of training data or depend on pre-trained models or frameworks to capture features from source images. This is inevitably hampered by a shortage of training data or a mismatch between the framework and the actual problem. Deep Image Prior (DIP) has been introduced to exploit convolutional neural networks’ ability to synthesize the ‘prior’ in the input image. However, the original design of DIP is hard to be generalized to multi-image processing problems, particularly for image fusion. Therefore, we propose a new image fusion technique that extends DIP to fusion tasks formulated as inverse problems. Additionally, we apply a multichannel approach to enhance DIP’s effect further. The evaluation is conducted with several commonly used image fusion assessment metrics. The results are compared with state-of-the-art image fusion methods. Our method outperforms these techniques for a range of metrics. In particular, it is shown to provide the best objective results for most metrics when applied to medical images.
Xudong Ma, Paul R. Hill, Nantheera Anantrasirichai, Alin Achim
ICIP3
2022 Optimal Transport-Based Graph Matching for 3D Retinal Oct Image Registration
abstract
Registration of longitudinal optical coherence tomography (OCT) images assists disease monitoring and is essential in image fusion applications. Mouse retinal OCT images are often collected for longitudinal study of eye disease models such as uveitis, but their quality is often poor compared with human imaging. This paper presents a novel but efficient framework involving an optimal transport based graph matching (OT-GM) method for 3D mouse OCT image registration. We first perform registration of fundus-like images obtained by projecting all b-scans of a volume on a plane orthogonal to them, hereafter referred to as the x-y plane. We introduce Adaptive Weighted Vessel Graph Descriptors (AWVGD) and 3D Cube Descriptors (CD) to identify the correspondence between nodes of graphs extracted from segmented vessels within the OCT projection images. The AWVGD comprises scaling, translation and rotation, which are computationally efficient, whereas CD exploits 3D spatial and frequency domain information. The OT-GM method subsequently performs the correct alignment in the x-y plane. Finally, registration along the direction orthogonal to the x-y plane (the z-direction) is guided by the segmentation of two important anatomical features peculiar to mouse b-scans, the Internal Limiting Membrane (ILM) and the hyaloid remnant (HR). Both subjective and objective evaluation results demonstrate that our framework outperforms other well-established methods on mouse OCT images within a reasonable execution time.
Xin Tian 0009, Nantheera Anantrasirichai, Lindsay Nicholson, Alin Achim
ICIP2
2022 Self-Supervised Contrastive Learning for Volcanic Unrest Detection
abstract
Ground deformation measured from interferometric synthetic aperture radar (InSAR) data is considered a sign of volcanic unrest, statistically linked to a volcanic eruption. Recent studies have shown the potential of using Sentinel-1 InSAR data and supervised deep learning (DL) methods for the detection of volcanic deformation signals, toward global volcanic hazard mitigation. However, detection accuracy is compromised from the lack of labeled data and class imbalance. To overcome this, synthetic data are typically used for fine-tuning DL models pretrained on the ImageNet dataset. This approach suffers from poor generalization on real InSAR data. This letter proposes the use of self-supervised contrastive learning to learn quality visual representations hidden in unlabeled InSAR data. Our approach, based on the SimCLR framework, provides a solution that does not require a specialized architecture nor a large labeled or synthetic dataset. We show that our self-supervised pipeline achieves higher accuracy with respect to the state-of-the-art methods and shows excellent generalization even for out-of-distribution test data. Finally, we showcase the effectiveness of our approach for detecting the unrest episodes preceding the recent Icelandic Fagradalsfjall volcanic eruption.
Nikolaos-Ioannis Bountos, Ioannis Papoutsis, Dimitrios Michail 0001, Nantheera Anantrasirichai
IEEE Geosci. Remote. Sens. Lett.4
2021 Contextual Colorization and Denoising for Low-Light Ultra High Resolution Sequences
abstract
Low-light image sequences generally suffer from spatiotemporal incoherent noise, flicker and blurring of moving objects. These artefacts significantly reduce visual quality and, in most cases, post-processing is needed in order to generate acceptable quality. Most state-of-the-art enhancement methods based on machine learning require ground truth data but this is not usually available for naturally captured low light sequences. We tackle these problems with an unpaired learning method that offers simultaneous colorization and denoising. Our approach is an adaptation of the CycleGAN structure. To overcome the excessive memory limitations associated with ultra high resolution content, we propose a multiscale patch-based framework, capturing both local and contextual features. Additionally, an adaptive temporal smoothing technique is employed to remove flickering artefacts. Experimental results show that our method outperforms existing approaches in terms of subjective quality and that it is robust to variations in brightness levels and noise.
Nantheera Anantrasirichai, David Bull 0001
ICIP1
2021 Detecting Ground Deformation in the Built Environment Using Sparse Satellite InSAR Data With a Convolutional Neural Network
abstract
The large volumes of Sentinel-1 data produced over Europe are being used to develop pan-national ground motion services. However, simple analysis techniques like thresholding cannot detect and classify complex deformation signals reliably making providing usable information to a broad range of nonexpert stakeholders a challenge. Here, we explore the applicability of deep learning approaches by adapting a pretrained convolutional neural network (CNN) to detect deformation in a national-scale velocity field. For our proof-of-concept, we focus on the U.K. where previously identified deformation is associated with coal-mining, ground water withdrawal, landslides, and tunneling. The sparsity of measurement points and the presence of spike noise make this a challenging application for deep learning networks, which involve calculations of the spatial convolution between images. Moreover, insufficient ground truth data exist to construct a balanced training data set, and the deformation signals are slower and more localized than in previous applications. We propose three enhancement methods to tackle these problems: 1) spatial interpolation with modified matrix completion; 2) a synthetic training data set based on the characteristics of the real U.K. velocity map; and 3) enhanced overwrapping techniques. Using velocity maps spanning 2015-2019, our framework detects several areas of coal mining subsidence, uplift due to dewatering, slate quarries, landslides, and tunnel engineering works. The results demonstrate the potential applicability of the proposed framework to the development of automated ground motion analysis systems.
Nantheera Anantrasirichai, Juliet Biggs, Krisztina Kelevitz, Zahra Sadeghi, Tim J. Wright, Alin Achim, David Bull 0001
IEEE Trans. Geosci. Remote. Sens.1
2021 River Planform Extraction From High-Resolution SAR Images via Generalized Gamma Distribution Superpixel Classification
abstract
The extraction of river planforms from remotely sensed satellite images is a task of crucial importance to many applications such as land planning, water resource monitoring, or flood prediction. In this article, we present a novel framework for the extraction of rivers from synthetic aperture radar (SAR) images, based on superpixel segmentation and subsequent classification. Superpixel segmentation is achieved by a modeling of the image pixels' amplitudes and spatial coordinates as a finite mixture model, where the generalized Gamma distribution is used to model accurately a variety of high-resolution SAR scenes. A number of features describing image texture and statistics are extracted on a superpixel level, facilitating the identification of river superpixels-planforms are then extracted by unsupervised, agglomerative clustering, thus eliminating the need for labeled training data. We present the results of our proposed method on the ICEYE-X2 and SENTINEL-1 SAR data, demonstrating its ability to produce pixel-accurate river masks.
Odysseas A. Pappas, Nantheera Anantrasirichai, Alin Achim, Byron A. Adams
IEEE Trans. Geosci. Remote. Sens.2
2020 Image fusion via sparse regularization with non-convex penalties
Nantheera Anantrasirichai, Rencheng Zheng, Ivan W. Selesnick, Alin Achim
Pattern Recognit. Lett.1
2020 Corrigendum to "Image fusion via sparse regularization with non-convex penalties" Pattern Recognition Letters Volume 131, March 2020, Pages 355-360
Nantheera Anantrasirichai, Rencheng Zheng, Ivan W. Selesnick, Alin Achim
Pattern Recognit. Lett.1
2019 Defectnet: Multi-Class Fault Detection on Highly-Imbalanced Datasets
abstract
As a data-driven method, the performance of deep convolutional neural networks (CNN) relies heavily on training data. The prediction results of traditional networks give a bias toward larger classes, which tend to be the background in the semantic segmentation task. This becomes a major problem for fault detection, where the targets appear very small on the images and vary in both types and sizes. In this paper we propose a new network architecture, DefectNet, that offers multi-class (including but not limited to) defect detection on highly-imbalanced datasets. DefectNet consists of two parallel paths, which are a fully convolutional network and a dilated convolutional network to detect large and small objects respectively. We propose a hybrid loss maximising the usefulness of a dice loss and a cross entropy loss, and we also employ the leaky rectified linear unit (ReLU) to deal with rare occurrence of some targets in training batches. The prediction results show that our DefectNet outperforms state-of-the-art networks for detecting multi-class defects with the average accuracy improvement of approximately 10% on a wind turbine.
Nantheera Anantrasirichai, David Bull 0001
ICIP1
2018 Atmospheric Turbulence Mitigation for Sequences with Moving Objects Using Recursive Image Fusion
abstract
This paper describes a new method for mitigating the effects of atmospheric distortion on observed sequences that include large moving objects. In order to provide accurate detail from objects behind the distorting layer, we solve the space-variant distortion problem using recursive image fusion based on the Dual Tree Complex Wavelet Transform (DT-CWT). The moving objects are detected and tracked using the improved Gaussian mixture models (GMM) and Kalman filtering. New fusion rules are introduced which work on the magnitudes and angles of the DT-CWT coefficients independently to achieve a sharp image and to reduce atmospheric distortion, respectively. The subjective results show that the proposed method achieves better video quality than other existing methods with competitive speed.
Nantheera Anantrasirichai, Alin Achim, David Bull 0001
ICIP1
2018 Fixation Prediction and Visual Priority Maps for Biped Locomotion
abstract
This paper presents an analysis of the low-level features and key spatial points used by humans during locomotion over diverse types of terrain. Although, a number of methods for creating saliency maps and task-dependent approaches have been proposed to estimate the areas of an image that attract human attention, none of these can straightforwardly be applied to sequences captured during locomotion, which contain dynamic content derived from a moving viewpoint. We used a novel learning-based method for creating a visual priority map informed by human eye tracking data. Our proposed priority map is created based on two fixation types: first exploiting the observation that humans search for safe foot placement and second that they observe the edges of a path as a guide to safe traversal of the terrain. Texture features and the difference between them, observed at the region around an eye position, are employed within a support vector machine to create a visual priority map for biped locomotion. The results show that our proposed method outperforms the state-of-the-art, particularly for more complex terrains, where achieving smooth locomotion needs more attention on the traversing path.
Nantheera Anantrasirichai, Katherine A. J. Daniels, Jeremy F. Burn, Iain D. Gilchrist, David Bull 0001
IEEE Trans. Cybern.1
2017 Line detection in speckle images using Radon transform and ℓ1 regularization
abstract
Boundaries and lines in medical images are important structures as they can delineate between tissue types, organs, and membranes. Although, a number of image enhancement and segmentation methods have been proposed to detect lines, none of these have considered line artefacts, which are more difficult to visualise as they are not physical structures, yet are still meaningful for clinical interpretation. This paper presents a novel method to restore lines, including line artefacts, in speckle images. We address this as a sparse estimation problem using a convex optimisation technique based on a Radon transform and sparsity regularisation (ℓ1norm). This problem divides into subproblems which are solved using the alternating direction method of multipliers, thereby achieving line detection and deconvolution simultaneously. The results for both simulated and in vivo ultrasound images show that the proposed method outperforms existing methods, in particular for detecting B-lines in lung ultrasound images, where the performance can be improved by up to 30 %.
Nantheera Anantrasirichai, Marco Allinovi, Wesley Hayes, David Bull 0001, Alin Achim
ICASSP1
2017 Line Detection as an Inverse Problem: Application to Lung Ultrasound Imaging
abstract
This paper presents a novel method for line restoration in speckle images. We address this as a sparse estimation problem using both convex and non-convex optimization techniques based on the Radon transform and sparsity regularization. This breaks into subproblems, which are solved using the alternating direction method of multipliers, thereby achieving line detection and deconvolution simultaneously. We include an additional deblurring step in the Radon domain via a total variation blind deconvolution to enhance line visualization and to improve line recognition. We evaluate our approach on a real clinical application: the identification of B-lines in lung ultrasound images. Thus, an automatic B-line identification method is proposed, using a simple local maxima technique in the Radon transform domain, associated with known clinical definitions of line artefacts. Using all initially detected lines as a starting point, our approach then differentiates between B-lines and other lines of no clinical significance, including Z-lines and A-lines. We evaluated our techniques using as ground truth lines identified visually by clinical experts. The proposed approach achieves the best B-line detection performance as measured by the F score when a non-convex [Formula: see text] regularization is employed for both line detection and deconvolution. The F scores as well as the receiver operating characteristic (ROC) curves show that the proposed approach outperforms the state-of-the-art methods with improvements in B-line detection performance of 54%, 40%, and 33% for [Formula: see text], [Formula: see text], and [Formula: see text], respectively, and of 24% based on ROC curve evaluations.
Nantheera Anantrasirichai, Wesley Hayes, Marco Allinovi, David Bull 0001, Alin Achim
IEEE Trans. Medical Imaging1
2016 Visual salience and priority estimation for locomotion using a deep convolutional neural network
abstract
This paper presents a novel method of salience and priority estimation for the human visual system during locomotion. This visual information contains dynamic content derived from a moving viewpoint. The priority map, ranking key areas on the image, is created from probabilities of gaze fixations, merged from bottom-up features and top-down control on the locomotion. Two deep convolutional neural networks (CNNs), inspired by models of the primate visual system, are employed to capture local salience features and compute probabilities. The first network operates through the foveal and peripheral areas around the eye positions. The second network obtains the importance of fixated points that have long durations or multiple visits, of which such areas need more times to process or to recheck to ensure smooth locomotion. The results show that our proposed method outperforms the state-of-the-art by up to 30 %, computed from average of four well known metrics for saliency estimation.
Nantheera Anantrasirichai, Iain D. Gilchrist, David Bull 0001
ICIP1
2016 Fixation identification for low-sample-rate mobile eye trackers
abstract
This paper presents a novel method of fixation identification for mobile eye trackers. The most significant benefit of our method over the state-of-the-art is that it achieves high accuracy for low-sample-rate devices worn during locomotion. This in turn delivers higher quality datasets for further use in human behaviour research, robotics and the development of guidance aids for the visually impaired. The proposed method employs temporal characteristics of the eye positions combined with statistical visual features extracted using a deep convolutional neural network, inspired by models of the primate visual system, through the fovea and peripheral areas around the eye positions. The results show that the proposed method outperforms existing methods by up to 16 % in terms of classification accuracy.
Nantheera Anantrasirichai, Iain D. Gilchrist, David Bull 0001
ICIP1
2015 Robust texture features based on undecimated dual-tree complex wavelets and local magnitude binary patterns
abstract
Image degradation due to illumination change, blur and noise can have a significant influence on classification performance, and yet no descriptors that perform well under these conditions exist. We propose a novel method for obtaining texture features, robust to these distortions, based on an undecimated dual-tree complex wavelet transform (UDT-CWT)1. As the UDT-CWT provides a local spatial relationship between scales, we can straightforwardly create bit-planes of the images representing local phases of wavelet coefficients. Magnitudes of the UDT-CWT are captured via a local binary pattern (LBP), after discarding some of the finest scales that are most affected by the blur and noise. A histogram of the resulting binary code words then forms the features used in texture classification. Results show that our approach outperforms existing methods, that claim to be invariant to feature degradations.
Nantheera Anantrasirichai, Jeremy F. Burn, David Bull 0001
ICIP1
2015 Undecimated Dual-Tree Complex Wavelet Transforms
Paul R. Hill, Nantheera Anantrasirichai, Alin Achim, Mohammed E. Al-Mualla, David Bull 0001
Signal Process. Image Commun.2
2015 Terrain Classification From Body-Mounted Cameras During Human Locomotion
abstract
This paper presents a novel algorithm for terrain type classification based on monocular video captured from the viewpoint of human locomotion. A texture-based algorithm is developed to classify the path ahead into multiple groups that can be used to support terrain classification. Gait is taken into account in two ways. Firstly, for key frame selection, when regions with homogeneous texture characteristics are updated, the frequency variations of the textured surface are analyzed and used to adaptively define filter coefficients. Secondly, it is incorporated in the parameter estimation process where probabilities of path consistency are employed to improve terrain-type estimation. When tested with multiple classes that directly affect mobility-a hard surface, a soft surface, and an unwalkable area-our proposed method outperforms existing methods by up to 16%, and also provides improved robustness.
Nantheera Anantrasirichai, Jeremy F. Burn, David Bull 0001
IEEE Trans. Cybern.1
2014 Orientation estimation for planar textured surfaces based on complex wavelets
abstract
The gradient of a road or terrain influences the appropriate speed and power of a vehicle traversing it. Therefore, gradient prediction is necessary if autonomous vehicles are to optimise their locomotion. This paper presents a novel texture-based method for estimating the orientation of planar surfaces under the basic assumption of homogeneity. Based on a backprojection technique, we propose a new method for measuring texture consistency in the Dual-Tree Complex Wavelet Transform (DT-CWT) domain. Texture histograms computed with a new anti-aliasing compensation approach are employed to find the rotation of a planar surface. The proposed method performs well with various types of textured surfaces and outperforms other existing methods with significantly reduced computational complexity up to 35%.
Nantheera Anantrasirichai, Jeremy F. Burn, David Bull 0001
ICIP1
2014 Robust texture features for blurred images using Undecimated Dual-Tree Complex Wavelets
abstract
This paper presents a new descriptor for texture classification. The descriptor is rotationally invariant and blur insensitive, which provides great benefits for various applications that suffer from out-of-focus content or involve fast moving or shaking cameras. We employ an Undecimated Dual-Tree Complex Wavelet Transform (UDT-CWT) [1] to extract texture features. As the UDT-CWT fully provides local spatial relationship between scales and subband orientations, we can straightforwardly create bit-planes of the images representing local phases of wavelet coefficients. We also discard some of the finest decomposition levels where are most affected by the blur. A histogram of the resulting code words is created and used as features in texture classification. Experimental results show that our approach outperforms existing methods by up to 40% for synthetic blurs and up to 30% for natural video content due to camera motion when walking.
Nantheera Anantrasirichai, Jeremy F. Burn, David Bull 0001
ICIP1
2013 Projective image restoration using sparsity regularization
abstract
This paper presents a method of image restoration for projective ground images which lie on a projection orthogonal to the camera axis. The ground images are initially transformed using homography, and then the proposed image restoration is applied. The process is performed in the dual-tree complex wavelet transform domain in conjunction with L0 reweighting and L2 minimisation (L0RL2) employed to solve this ill-posed problem. We also propose instant estimation of a blur kernel arising from the projective transform and the subsequent interpolation of sparse data. Subjective results show significant improvement of image quality. Furthermore, classification of surface type at various distances (evaluated using a support vector machine classifier) is also improved for the images restored using our proposed algorithm.
Nantheera Anantrasirichai, Jeremy F. Burn, David Bull 0001
ICIP1
2013 Adaptive-weighted bilateral filtering for optical coherence tomography
abstract
This paper presents an image enhancement method for retinal optical coherence tomography (OCT) images. Raw OCT images contain a large amount of speckle which causes images to be grainy and very low contrast. The raw OCT images thus need to be processed before any clinical interpretation is made. We propose a novel method to remove speckle, while preserving useful information contained in each retinal layer. The process starts with multi-scale despeckling based on a dual-tree complex wavelet transform (DT-CWT). Then, we further enhance the OCT image through a smoothing process that uses a novel adaptive-weighted bilateral filter (AWBF). This offers the desirable property of preserving texture within the OCT images. Glaucoma classification results confirm that our method can significantly enhance the clinical usefulness of OCT images.
Nantheera Anantrasirichai, Lindsay Nicholson, James E. Morgan, Irina Erchova, Alin Achim
ICIP1
2013 Curvelet domain image fusion of OCT and fundus imagery using convolution of Meridian distributions
abstract
This paper presents a novel statistical model based method aimed at fusing Optical Coherence Tomography and Fundus Photographic imagery of the eye. The presented method utilises the Discrete Curvelet Transform to decompose the images into sub-band coefficients. The Meridian distribution, a specialized case of the generalized Cauchy distribution, is used to model the curvelet decomposition coefficients. The convolution of the input image distributions is used as a probabilistic prior for modelling the fused image coefficients. Experimental results show this method to provide very high-quality fusion results.
Odysseas A. Pappas, Nantheera Anantrasirichai, Lindsay Nicholson, James E. Morgan, Irina Erchova, Alin Achim
ICIP2
2013 Atmospheric Turbulence Mitigation Using Complex Wavelet-Based Fusion
abstract
Restoring a scene distorted by atmospheric turbulence is a challenging problem in video surveillance. The effect, caused by random, spatially varying, perturbations, makes a model-based solution difficult and in most cases, impractical. In this paper, we propose a novel method for mitigating the effects of atmospheric distortion on observed images, particularly airborne turbulence which can severely degrade a region of interest (ROI). In order to extract accurate detail about objects behind the distorting layer, a simple and efficient frame selection method is proposed to select informative ROIs only from good-quality frames. The ROIs in each frame are then registered to further reduce offsets and distortions. We solve the space-varying distortion problem using region-level fusion based on the dual tree complex wavelet transform. Finally, contrast enhancement is applied. We further propose a learning-based metric specifically for image quality assessment in the presence of atmospheric distortion. This is capable of estimating quality in both full- and no-reference scenarios. The proposed method is shown to significantly outperform existing methods, providing enhanced situational awareness in a range of surveillance scenarios.
Nantheera Anantrasirichai, Alin Achim, Nick G. Kingsbury, David Bull 0001
IEEE Trans. Image Process.1
2012 Mitigating the effects of atmospheric distortion using DT-CWT fusion
abstract
This paper describes a new method for mitigating the effects of atmospheric distortion on observed images, particularly airborne turbulence which degrades a region of interest (ROI). In order to provide accurate detail from objects behind the distorting layer, a simple and efficient frame selection method is proposed to pick informative ROIs from only good-quality frames. We solve the space-variant distortion problem using region-based fusion based on the Dual Tree Complex Wavelet Transform (DT-CWT). We also propose an object alignment method for pre-processing the ROI since this can exhibit significant offsets and distortions between frames. Simple haze removal is used as the final step. The proposed method performs very well with atmospherically distorted videos and outperforms other existing methods.
Nantheera Anantrasirichai, Alin Achim, David Bull 0001, Nick G. Kingsbury
ICIP1
2011 Colour volumetric compression for realistic view synthesis applications
Nantheera Anantrasirichai, Cedric Nishan Canagarajah, David W. Redmill, Akbar Sheikh Akbari, David Bull 0001
Multim. Tools Appl.1
2010 Spatiotemporal super-resolution for low bitrate H.264 video
abstract
Super-resolution and frame interpolation enhance low resolution low-framerate videos. Such techniques are especially important for limited bandwidth communications. This paper proposes a novel technique to up-scale videos compressed with H.264 at low bit-rate both in spatial and temporal dimensions. A quantisation noise model is used in the super-resolution estimator, designed for low bitrate video, and a weighting map for decreasing inaccuracy of motion estimation are proposed. Results show improvement both in rate-distortion and perceived image quality.
Nantheera Anantrasirichai, Cedric Nishan Canagarajah
ICIP1
2010 In-Band Disparity Compensation for Multiview Image Compression and View Synthesis
abstract
This paper presents a novel framework to achieve scalable multiview image compression and view synthesis. The open-loop wavelet-lifting scheme for geometric filtering has been exploited to achieve signal-to-noise ratio scalability and view-type scalability (mono, stereo, or multiview). Spatial scalability is achieved by employing in-band prediction which removes correlations among subbands (level-by-level) via shift-invariant references obtained by overcomplete discrete wavelet transforms. We propose a novel in-band disparity compensated view filtering approach, akin to motion compensated temporal filtering, for achieving a scalable multiview codec. In our codec, hybrid prediction is proposed to deal with occlusions, and a novel cost function in dynamic programming (DP) for disparity estimation is introduced to improve view synthesis quality. Experiments show comparable results at full resolution and significant improvements at coarser resolutions, compared to a conventional spatial prediction scheme. View synthesis efficiency is extensively improved by utilizing disparity estimation from the proposed DP approach.
Nantheera Anantrasirichai, Cedric Nishan Canagarajah, David W. Redmill, David Bull 0001
IEEE Trans. Circuits Syst. Video Technol.1
2009 Enhanced spatially interleaved DVC using diversity and selective feedback
abstract
Systems with cheap/simple/power efficient encoders but complex decoders make applications such as low cost, low power remote sensors practical. Bandwidth considerations however are still an issue and compression efficiency has to remain high. In this paper, we present a distributed video codec (DVC) that we are developing with the aim of achieving such a low power paradigm at the cost of only a small compression performance deficit relative to the current state of the art, H.264. The proposed system employs spatial interleaving of KEY and Wyner-Ziv data which allows efficient side information (SI) generation through block-based error concealment, a Gray code that increases the accuracy of bit probability estimation, and a diversity scheme that produces more reliable results by exploiting multiple SI generated data. Simulation results show an improvement of the proposed scheme over H.264 intra coding of up to 1.5 dB. We additionally propose two mechanisms for selective parity bit feedback requests that can further reduce the WZ bitrate by up to 15%.
Nantheera Anantrasirichai, Dimitris Agrafiotis, David Bull 0001
ICASSP1
2008 A concealment based approach to distributed video coding
abstract
This paper presents a concealment based approach to distributed video coding that uses hybrid key/WZ frames via an FMO type interleaving of macroblocks. Our motivation stems from a previous work of ours that showed promising results relative to the more common approach of splitting the sequence in key and WZ frames. In this paper, we extend our previous scheme to the case of I-B-P frame structures and transform domain DVC. We additionally introduce a number of enhancements at the decoder including use of spatio-temporal concealment for generating the side information on a MB basis, mode selection for switching between the two concealment approaches and for deciding how the correlation noise is estimated, local (MB wise) correlation noise estimation and modified B frame quantisation. The results presented indicate considerable improvement (up to 30%) compared to corresponding frame extrapolation and frame interpolation schemes.
Nantheera Anantrasirichai, Dimitris Agrafiotis, David Bull 0001
ICIP1
2007 Colour Volumetric Compression for Realistic View Synthesis Applications
abstract
The colour volumetric data which is constructed from a set of multi-view images is capable of providing realistic immersive experience. However it is not widely applicable due to its manifold increase in bandwidth. This paper presents a novel framework to achieve scalable volumetric compression. Based on wavelet transformation, data rearrangement algorithm is proposed to compact volumetric data leading to high efficiency of transformation. The colour data is also rearranged by using the characteristics of human eye sensitivity. Moreover, the pre-processing for adaptive resolution is proposed in this paper. The low resolution overcomes the limitation of the data transmission at low bit rate, whilst the fine resolution improves the synthesised images' quality. The results of our proposed schemes show significant improvement of the compression performance over the traditional 3D coding.
Nantheera Anantrasirichai, Cedric Nishan Canagarajah, David W. Redmill, David Bull 0001
ICASSP (1)1
2006 Dynamic Programming for Multi-View Disparity/Depth Estimation
abstract
A novel algorithm for disparity/depth estimation from multi-view images is presented. A dynamic programming approach with window-based correlation and a novel cost function is proposed. The smoothness of disparity/depth map is embedded in dynamic programming approach, whilst the window-based correlation increases reliability. The enhancement methods are included, i.e. adaptive window size and shiftable window are used to increase reliability in homogenous areas and to increase sharpness at object boundaries. First, the algorithms estimate depth maps along a single camera axis. The algorithms exploits then combines the depth estimates from different axis to derive a suitable depth map for multi-view images. The proposed scheme outperforms existing approaches in parallel and in the non-parallel camera configurations
Nantheera Anantrasirichai, Cedric Nishan Canagarajah, David W. Redmill, David Bull 0001
ICASSP (2)1
2006 Volumetric Representation for Sparse Multi-Views
abstract
In this paper, we propose a novel volumetric representation for a sparse set of calibrated multi-view images of a non-Lambertian scene. The depth map of each reference view is registered into a volume and a simple algorithm to shape the volume is introduced. Particular colours are defined for each voxel to render a smooth and realistic image. Synthesized results demonstrate the good performance with the proposed scheme, both for parallel-camera and non-parallel-camera geometries.
Nantheera Anantrasirichai, Cedric Nishan Canagarajah, David W. Redmill, David Bull 0001
ICIP1
2005 Multi-view image coding with wavelet lifting and in-band disparity compensation
abstract
This paper presents a novel framework to achieve scalable multi-view image coding. As open loop operation, the wavelet lifting scheme for geometric filtering has been exploited to overcome the limitation of SNR scalability and to attain view scalability. The essential key for achieving the spatial scalability is the in-band prediction. It removes correlations among subbands level-by-level via shift-invariant references obtained by overcomplete discrete wavelet transforms (ODWT). Additionally, the proposed disparity compensated view filtering is allowed to exploit the different filters and estimation parameters for each resolution level. The experiments show comparable results at full resolution and the significant improvement at coarser resolution over the conventional spatial prediction scheme.
Nantheera Anantrasirichai, Cedric Nishan Canagarajah, David Bull 0001
ICIP (3)1