VLDB 2026 Research / reviewers in the wild / expert
Nantheera Anantrasirichai
dblp:98/5858
· DBLP profile ↗
54ranked-venue papers
30as first author
26since 2021 · last 2026
0000-0002-2122-5781ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Graphics, computer vision, multimedia, augmented reality and games · 42 · 23 first-author · 19 since 2021Artificial intelligence and machine learning · 8 · 5 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 4 since 2021Systems, architecture and hardware · 3 · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Splatography: Sparse Multi-View Dynamic Gaussian Splatting for Filmmaking ChallengesabstractDeformable Gaussian Splatting (GS) accomplishes photorealistic dynamic 3-D reconstruction from dense multi-view video (MVV) by learning to deform a canonical GS representation. However, in filmmaking, tight budgets can result in sparse camera configurations, which limits state-of-the-art (SotA) methods when capturing complex dynamic features. To address this issue, we introduce an approach that splits the canonical Gaussians and deformation field into foreground and background components using a sparse set of masks for frames at$t=0$. Each representation is separately trained on different loss functions during canonical pre-training. Then, during dynamic training, different parameters are modeled for each deformation field following common filmmaking practices. The foreground stage contains diverse dynamic features so changes in color, position and rotation are learned. While, the background containing film-crew and equipment, is typically dimmer and less dynamic so only changes in point position are learned. Experiments on 3-D and 2.5-D entertainment datasets show that our method produces SotA qualitative and quantitative results; up to 3 PSNR higher with half the model size on 3-D scenes. Unlike the SotA and without the need for dense mask supervision, our method also produces segmented dynamic reconstructions including transparent and dynamic textures. Code and video comparisons are available online: http://bit.ly/4oqzZrO Adrian Azzarelli, Nantheera Anantrasirichai, David Bull 0001 |
3DV | 2 |
| 2026 | Bayesian Neural Networks for One-to-Many Mapping in Image EnhancementabstractIn image enhancement tasks, such as low-light and underwater image enhancement, a degraded image can correspond to multiple plausible target images due to dynamic photography conditions. This naturally results in a one-to-many mapping problem. To address this, we propose a Bayesian Enhancement Model (BEM) that incorporates Bayesian Neural Networks (BNNs) to capture data uncertainty and produce diverse outputs. To enable fast inference, we introduce a BNN-DNN framework: a BNN is first employed to model the one-to-many mapping in a low-dimensional space, followed by a Deterministic Neural Network (DNN) that refines fine-grained image details. Extensive experiments on multiple low-light and underwater image enhancement benchmarks demonstrate the effectiveness of our method. Guoxi Huang, Ruirui Lin, Zipeng Qi, David Bull 0001, Nantheera Anantrasirichai |
AAAI | 6 |
| 2026 | RMFAT: Recurrent Multi-scale Feature Atmospheric Turbulence MitigatorabstractAtmospheric turbulence (AT) severely degrades video quality by introducing distortions such as geometric warping, blur, and temporal flickering, posing significant challenges to both visual clarity and temporal consistency. Current state-of-the-art methods are based on transformer, 3D architectures and require multi-frame input, but their large computational cost and memory usage limit real-time deployment. In this work, we propose RMFAT, Recurrent Multi-scale Feature Atmospheric Turbulence Mitigator, designed for efficient and temporally consistent video restoration under AT conditions. RMFAT adopts a lightweight recurrent framework that restores each frame using only two inputs at a time, significantly reducing temporal window size and computational burden. It further integrates multi-scale feature encoding and decoding with temporal warping modules at both encoder and decoder stages to enhance spatial detail and temporal coherence. Extensive experiments conducted on synthetic and real-world atmospheric turbulence datasets demonstrate that RMFAT not only outperforms existing methods in terms of clarity restoration (with nearly a 9% improvement in SSIM) but also achieves significantly improved inference speed (achieving a more than fourfold reduction), making it particularly suitable for real-time atmospheric turbulence suppression tasks. Nantheera Anantrasirichai |
AAAI | 2 |
| 2026 | GFix: Perceptually Enhanced Gaussian Splatting Video Compressionabstract3D Gaussian Splatting (3DGS) enhances 3D scene reconstruction through explicit representation and fast rendering, demonstrating potential benefits for various low-level vision tasks, including video compression. However, existing 3DGS-based video codecs generally exhibit more noticeable visual artifacts and relatively low compression ratios. In this paper, we specifically target the perceptual enhancement of 3DGS-based video compression, based on the assumption that artifacts from 3DGS rendering and quantization resemble noisy latents sampled during diffusion training. Building on this premise, we propose a content-adaptive framework, GFix, comprising a streamlined, single-step diffusion model that serves as an off-the-shelf neural enhancer. Moreover, to increase compression efficiency, We propose a modulated LoRA scheme that freezes the low-rank decompositions and modulates the intermediate hidden states, thereby achieving efficient adaptation of the diffusion backbone with highly compressible updates. Experimental results show that GFix delivers strong perceptual quality enhancement, outperforming GSVC with up to 72.1% BD-rate savings in LPIPS and 21.4% in FID. Siyue Teng, Ge Gao 0005, Duolikun Danier, Yuxuan Jiang 0015, Fan Zhang 0017, Nantheera Anantrasirichai, Thomas Davis, Zoe Liu, David Bull 0001 |
ISCAS | 6 |
| 2026 | DMAT: An End-to-End Framework for Joint Atmospheric Turbulence Mitigation and Object DetectionabstractAtmospheric Turbulence (AT) degrades the clarity and accuracy of surveillance imagery, posing challenges not only for visualization quality but also for object classification and scene tracking. Deep learning-based methods have been proposed to improve visual quality, but spatio-temporal distortions remain a significant issue. Although deep learning-based object detection performs well under normal conditions, it struggles to operate effectively on sequences distorted by atmospheric turbulence. In this paper, we propose a novel framework that learns to compensate for distorted features while simultaneously improving visualization and object detection. This end-to-end training strategy leverages and exchanges knowledge of low-level distorted features in the AT mitigator with semantic features extracted in the object detector. Specifically, in the AT mitigator a 3D Mamba-based structure is used to handle the spatio-temporal displacements and blurring caused by turbulence. Optimization is achieved through back-propagation in both the AT mitigator and object detector. Our proposed DMAT outperforms state-of-the-art AT mitigation and object detection systems up to a 15% improvement on datasets corrupted by generated turbulence. The code is available at https://github.com/pui-nantheera/DMAT and datasets are available at https://zenodo.org/records/17673509. Paul R. Hill, Alin Achim, David Bull 0001, Nantheera Anantrasirichai |
WACV | 5 |
| 2026 | ViVo: A Dataset for Human Volumetric Video Reconstruction and Compression
Adrian Azzarelli, Ge Gao 0005, Ho Man Kwan, Fan Zhang 0017, Nantheera Anantrasirichai, Oliver Moolan-Feroze, David Bull 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2025 | MAMAT: 3D Mamba-Based Atmospheric Turbulence Removal and its Object Detection CapabilityabstractRestoration and enhancement are essential for improving the quality of videos captured under atmospheric turbulence conditions, aiding visualization, object detection, classification, and tracking in surveillance systems. In this paper, we introduce a novel Mamba-Based method, the 3D Mamba-Based Atmospheric Turbulence Removal (MAMAT), which employs a dual-module strategy to mitigate these distortions. The first module utilizes deformable 3D convolutions for non-rigid registration to minimize spatial shifts, while the second module enhances contrast and detail. Leveraging the advanced capabilities of the 3D Mamba architecture, experimental results demonstrate that MAMAT outperforms state-of-the-art learning-based methods, achieving up to a 3% improvement in visual quality and a 15% boost in object detection. It not only enhances visualization but also significantly improves object detection accuracy, bridging the gap between visual restoration and the effectiveness of surveillance applications. The code is available at https://github.com/csprh/MAMAT. Paul R. Hill, Nantheera Anantrasirichai |
AVSS | 3 |
| 2025 | Multi-Scale Denoising in the Feature Space for Low-Light Instance SegmentationabstractInstance segmentation for low-light imagery remains largely unexplored due to the challenges imposed by such conditions, for example shot noise due to low photon count, color distortions and reduced contrast. In this paper, we propose an end-to-end solution to address this challenging task. Our proposed method implements weighted non-local blocks (wNLB) in the feature extractor. This integration enables an inherent denoising process at the feature level. As a result, our method eliminates the need for aligned ground truth images during training, thus supporting training on real-world low-light datasets. We introduce additional learnable weights at each layer in order to enhance the network’s adaptability to real-world noise characteristics, which affect different feature scales in different ways. Experimental results on several object detectors show that the proposed method outperforms the pre-trained networks with an Average Precision (AP) improvement of at least +7.6, with the introduction of wNLB further enhancing AP by upto +1.3. Joanne Lin, Nantheera Anantrasirichai, David Bull 0001 |
ICASSP | 2 |
| 2025 | BVI-CR: A Multi-View Human Dataset for Volumetric Video CompressionabstractThe advances in immersive technologies and 3D reconstruction have enabled the creation of digital replicas of real-world objects and environments with fine details. These processes generate vast amounts of 3D data, requiring more efficient compression methods to satisfy the memory and bandwidth constraints associated with data storage and transmission. However, the development and validation of efficient 3D data compression methods are constrained by the lack of comprehensive and high-quality volumetric video datasets, which typically require much more effort to acquire and consume increased resources compared to 2D image and video databases. To bridge this gap, we present an open multi-view volumetric human dataset, denoted BVI-CR, which contains 18 multi-view RGB-D captures and their corresponding textured polygonal meshes, depicting a range of diverse human actions. Each video sequence contains 10 views in 1080p resolution with durations between 10-15 seconds at 30FPS. Using BVI-CR, we benchmarked three conventional and neural coordinate-based multi-view video compression methods, following the MPEG MIV Common Test Conditions, and reported their rate quality performance based on various quality metrics. The results show the great potential of neural representation based methods in volumetric video compression compared to conventional video coding methods (with an up to 38% average coding gain in PSNR). This dataset provides a development and validation platform for a variety of tasks including volumetric reconstruction, compression, and quality assessment. The database will be shared publicly at https://github.com/fan-aaron-zhang/bvi-cr. Ge Gao 0005, Adrian Azzarelli, Ho Man Kwan, Nantheera Anantrasirichai, Fan Zhang 0017, Will Andrew, Oliver Moolan-Feroze, David Bull 0001 |
ISCAS | 4 |
| 2025 | AquaNeRF: Neural Radiance Fields in Underwater Media with Distractor RemovalabstractNeural radiance field (NeRF) research has made significant progress in modeling static video content captured in the wild. However, current models and rendering processes rarely consider scenes captured underwater, which are useful for studying and filming ocean life. They fail to address visual artifacts unique to underwater scenes, such as moving fish and suspended particles. This paper introduces a novel NeRF renderer and optimization scheme for an implicit MLP-based NeRF model. Our renderer reduces the influence of floaters and moving objects that interfere with static objects of interest by estimating a single surface per ray. We use a Gaussian weight function with a small offset to ensure that the transmittance of the surrounding media remains constant. Additionally, we enhance our model with a depth-based scaling function to upscale gradients for near-camera volumes. Overall, our method outperforms the baseline Nerfacto by approximately 7.5% and SeaThru-NeRF by 6.2% in terms of PSNR. Subjective evaluation also shows a significant reduction of artifacts while preserving details of static targets and background compared to the state of the arts. Luca Gough, Adrian Azzarelli, Fan Zhang 0017, Nantheera Anantrasirichai |
ISCAS | 4 |
| 2025 | Guiding WaveMamba with Frequency Maps for Image Debanding
Xinyi Wang 0011, Smaranda Tasmoc, Nantheera Anantrasirichai, Angeliki V. Katsenou |
PCS | 3 |
| 2025 | UW-GS: Distractor-Aware 3D Gaussian Splatting for Enhanced Underwater Scene Reconstructionabstract3D Gaussian splatting (3DGS) offers the capability to achieve real-time high quality 3D scene rendering. However, 3DGS assumes that the scene is in a clear medium environment and struggles to generate satisfactory representations in underwater scenes, where light absorption and scattering are prevalent and moving objects are involved. To overcome these, we introduce a novel Gaussian Splatting-based method, UW-GS, designed specifically for underwater applications. It introduces a color appearance that models distance-dependent color variation, employs a new physics-based density control strategy to enhance clarity for distant objects, and uses a binary motion mask to handle dynamic content. Optimized with a well-designed loss function supporting for scattering media and strengthened by pseudo-depth maps, UW-GS outperforms existing methods with PSNR gains up to 1.26dB. To fully verify the effectiveness of the model, we also developed a new underwater dataset, S-UW, with dynamic object masks. The code of UW-GS and S-UW will be available at https://github.com/WangHaoran16/UW-GS. Nantheera Anantrasirichai, Fan Zhang 0017, David Bull 0001 |
WACV | 2 |
| 2024 | A Spatio-Temporal Aligned SUNet Model For Low-Light Video EnhancementabstractDistortions caused by low-light conditions are not only visually unpleasant but also degrade the performance of computer vision tasks. The restoration and enhancement have proven to be highly beneficial. However, there are only a limited number of enhancement methods explicitly designed for videos acquired in low-light conditions. We propose a Spatio-Temporal Aligned SUNet (STA-SUNet) model using a Swin Transformer as a backbone to capture low light video features and exploit their spatio-temporal correlations. The STA-SUNet model is trained on a novel, fully registered dataset (BVI), which comprises dynamic scenes captured under varying light conditions. It is further analysed comparatively against various other models over three test datasets. The model demonstrates superior adaptivity across all datasets, obtaining the highest PSNR and SSIM values. It is particularly effective in extreme low-light conditions, yielding fairly good visualisation results. Ruirui Lin, Nantheera Anantrasirichai, Alexandra Malyugina, David Bull 0001 |
ICIP | 2 |
| 2024 | Anomaly Detection for the Identification of Volcanic Unrest in Satellite ImageryabstractSatellite images have the potential to detect volcanic deformation prior to eruptions, but while a vast number of images are routinely acquired, only a small percentage contain volcanic deformation events. Manual inspection could miss these anomalies, and an automatic system modelled with supervised learning requires suitably labelled datasets. To tackle these issues, this paper explores the use of unsupervised deep learning on satellite data for the purpose of identifying volcanic deformation as anomalies. Our detector is based on Patch Distribution Modeling (PaDiM), and the detection performance is enhanced with a weighted distance, assigning greater importance to features from deeper layers. Additionally, we propose a preprocessing approach to handle noisy and incomplete data points. The final framework was tested with five volcanoes, which have different deformation characteristics and its performance was compared against the supervised learning method for volcanic deformation detection. Robert Gabriel Popescu, Nantheera Anantrasirichai, Juliet Biggs |
ICIP | 2 |
| 2024 | TaGAT: Topology-Aware Graph Attention Network for Multi-modal Retinal Image Fusion
Xin Tian 0009, Nantheera Anantrasirichai, Lindsay Nicholson, Alin Achim |
MICCAI (1) | 2 |
| 2023 | ST-MFNET Mini: Knowledge Distillation-Driven Frame InterpolationabstractCurrently, one of the major challenges in deep learning-based video frame interpolation (VFI) is the large model size and high computational complexity associated with many high performance VFI approaches. In this paper, we present a distillation-based two-stage workflow for obtaining compressed VFI models which perform competitively compared to the state of the art, but with significantly reduced model size and complexity. Specifically, an optimization-based network pruning method is applied to a state of the art frame interpolation model, ST-MFNet, which suffers from large model size. The resulting network architecture achieves a 91% reduction in parameter numbers and a 35% increase in speed. The performance of the new network is further enhanced through a teacher-student knowledge distillation training process using a Laplacian distillation loss. The final low complexity model, ST-MFNet Mini, achieves a comparable performance to most existing high-complexity VFI methods, only outperformed by the original ST-MFNet. Our source code is available at https://github.com/crispianm/ST-MFNet-Mini Crispian Morris, Duolikun Danier, Fan Zhang 0017, Nantheera Anantrasirichai, David Bull 0001 |
ICIP | 4 |
| 2023 | Atmospheric turbulence removal with complex-valued convolutional neural networkabstractAtmospheric turbulence distorts visual imagery and is always problematic for information interpretation by both human and machine. Most well-developed approaches to remove atmospheric turbulence distortion are model-based. However, these methods require high computation and large memory making real-time operation infeasible. Deep learning-based approaches have hence gained more attention but currently work efficiently only on static scenes. This paper presents a novel learning-based framework offering short temporal spanning to support dynamic scenes. We exploit complex-valued convolutions as phase information, altered by atmospheric turbulence, is captured better than using ordinary real-valued convolutions. Two concatenated modules are proposed. The first module aims to remove geometric distortions and, if enough memory, the second module is applied to refine micro details of the videos. Experimental results show that our proposed framework efficiently mitigates the atmospheric turbulence distortion and significantly outperforms existing methods. Nantheera Anantrasirichai |
Pattern Recognit. Lett. | 1 |
| 2023 | A topological loss function for image Denoising on a new BVI-lowlight datasetabstractAlthough image denoising algorithms have attracted significant research attention, surprisingly few have been proposed for, or evaluated on, noise from imagery acquired under real low-light conditions. Moreover, noise characteristics are often assumed to be spatially invariant, leading to edges and textures being distorted after denoising. Here, we introduce a novel topological loss function which is based on persistent homology. The method performs in the space of image patches, where topological invariants are calculated and represented in persistent diagrams. The loss function is a combination of ℓ1 or ℓ2 losses with the new persistence-based topological loss. We compare its performance across popular denoising architectures and loss functions, training the networks on our new comprehensive dataset of natural images captured in low-light conditions – BVI-LOWLIGHT. Analysis reveals that this approach outperforms existing methods, adapting well to complex structures and suppressing common artifacts. Alexandra Malyugina, Nantheera Anantrasirichai, David Bull 0001 |
Signal Process. | 2 |
| 2022 | DeTurb: Atmospheric Turbulence Mitigation with Deformable 3D Convolutions and 3D Swin Transformers
Zhicheng Zou, Nantheera Anantrasirichai |
ACCV (4) | 2 |
| 2022 | ICIP 2022 Challenge on Parasitic Egg Detection and Classification in Microscopic Images: Dataset, Methods and ResultsabstractManual examination of faecal smear samples to identify the existence of parasitic eggs is very time-consuming and can only be done by specialists. Therefore, an automated system is required to tackle this problem since it can relate to serious intestinal parasitic infections. This paper reviews the ICIP 2022 Challenge on parasitic egg detection and classification in microscopic images. We describe a new dataset for this application, which is the largest dataset of its kind. The methods used by participants in the challenge are summarised and discussed along with their results. Nantheera Anantrasirichai, Thanarat H. Chalidabhongse, Duangdao Palasuwan, Korranat Naruenatthanaset, Thananop Kobchaisawat, Nuntiporn Nunthanasup, Kanyarat Boonpeng, Xudong Ma, Alin Achim |
ICIP | 1 |
| 2022 | Unsupervised Image Fusion Using Deep Image PriorsabstractA significant number of researchers have applied deep learning methods to image fusion. However, most works require a large amount of training data or depend on pre-trained models or frameworks to capture features from source images. This is inevitably hampered by a shortage of training data or a mismatch between the framework and the actual problem. Deep Image Prior (DIP) has been introduced to exploit convolutional neural networks’ ability to synthesize the ‘prior’ in the input image. However, the original design of DIP is hard to be generalized to multi-image processing problems, particularly for image fusion. Therefore, we propose a new image fusion technique that extends DIP to fusion tasks formulated as inverse problems. Additionally, we apply a multichannel approach to enhance DIP’s effect further. The evaluation is conducted with several commonly used image fusion assessment metrics. The results are compared with state-of-the-art image fusion methods. Our method outperforms these techniques for a range of metrics. In particular, it is shown to provide the best objective results for most metrics when applied to medical images. Xudong Ma, Paul R. Hill, Nantheera Anantrasirichai, Alin Achim |
ICIP | 3 |
| 2022 | Optimal Transport-Based Graph Matching for 3D Retinal Oct Image RegistrationabstractRegistration of longitudinal optical coherence tomography (OCT) images assists disease monitoring and is essential in image fusion applications. Mouse retinal OCT images are often collected for longitudinal study of eye disease models such as uveitis, but their quality is often poor compared with human imaging. This paper presents a novel but efficient framework involving an optimal transport based graph matching (OT-GM) method for 3D mouse OCT image registration. We first perform registration of fundus-like images obtained by projecting all b-scans of a volume on a plane orthogonal to them, hereafter referred to as the x-y plane. We introduce Adaptive Weighted Vessel Graph Descriptors (AWVGD) and 3D Cube Descriptors (CD) to identify the correspondence between nodes of graphs extracted from segmented vessels within the OCT projection images. The AWVGD comprises scaling, translation and rotation, which are computationally efficient, whereas CD exploits 3D spatial and frequency domain information. The OT-GM method subsequently performs the correct alignment in the x-y plane. Finally, registration along the direction orthogonal to the x-y plane (the z-direction) is guided by the segmentation of two important anatomical features peculiar to mouse b-scans, the Internal Limiting Membrane (ILM) and the hyaloid remnant (HR). Both subjective and objective evaluation results demonstrate that our framework outperforms other well-established methods on mouse OCT images within a reasonable execution time. Xin Tian 0009, Nantheera Anantrasirichai, Lindsay Nicholson, Alin Achim |
ICIP | 2 |
| 2022 | Self-Supervised Contrastive Learning for Volcanic Unrest DetectionabstractGround deformation measured from interferometric synthetic aperture radar (InSAR) data is considered a sign of volcanic unrest, statistically linked to a volcanic eruption. Recent studies have shown the potential of using Sentinel-1 InSAR data and supervised deep learning (DL) methods for the detection of volcanic deformation signals, toward global volcanic hazard mitigation. However, detection accuracy is compromised from the lack of labeled data and class imbalance. To overcome this, synthetic data are typically used for fine-tuning DL models pretrained on the ImageNet dataset. This approach suffers from poor generalization on real InSAR data. This letter proposes the use of self-supervised contrastive learning to learn quality visual representations hidden in unlabeled InSAR data. Our approach, based on the SimCLR framework, provides a solution that does not require a specialized architecture nor a large labeled or synthetic dataset. We show that our self-supervised pipeline achieves higher accuracy with respect to the state-of-the-art methods and shows excellent generalization even for out-of-distribution test data. Finally, we showcase the effectiveness of our approach for detecting the unrest episodes preceding the recent Icelandic Fagradalsfjall volcanic eruption. Nikolaos-Ioannis Bountos, Ioannis Papoutsis, Dimitrios Michail 0001, Nantheera Anantrasirichai |
IEEE Geosci. Remote. Sens. Lett. | 4 |
| 2021 | Contextual Colorization and Denoising for Low-Light Ultra High Resolution SequencesabstractLow-light image sequences generally suffer from spatiotemporal incoherent noise, flicker and blurring of moving objects. These artefacts significantly reduce visual quality and, in most cases, post-processing is needed in order to generate acceptable quality. Most state-of-the-art enhancement methods based on machine learning require ground truth data but this is not usually available for naturally captured low light sequences. We tackle these problems with an unpaired learning method that offers simultaneous colorization and denoising. Our approach is an adaptation of the CycleGAN structure. To overcome the excessive memory limitations associated with ultra high resolution content, we propose a multiscale patch-based framework, capturing both local and contextual features. Additionally, an adaptive temporal smoothing technique is employed to remove flickering artefacts. Experimental results show that our method outperforms existing approaches in terms of subjective quality and that it is robust to variations in brightness levels and noise. Nantheera Anantrasirichai, David Bull 0001 |
ICIP | 1 |
| 2021 | Detecting Ground Deformation in the Built Environment Using Sparse Satellite InSAR Data With a Convolutional Neural NetworkabstractThe large volumes of Sentinel-1 data produced over Europe are being used to develop pan-national ground motion services. However, simple analysis techniques like thresholding cannot detect and classify complex deformation signals reliably making providing usable information to a broad range of nonexpert stakeholders a challenge. Here, we explore the applicability of deep learning approaches by adapting a pretrained convolutional neural network (CNN) to detect deformation in a national-scale velocity field. For our proof-of-concept, we focus on the U.K. where previously identified deformation is associated with coal-mining, ground water withdrawal, landslides, and tunneling. The sparsity of measurement points and the presence of spike noise make this a challenging application for deep learning networks, which involve calculations of the spatial convolution between images. Moreover, insufficient ground truth data exist to construct a balanced training data set, and the deformation signals are slower and more localized than in previous applications. We propose three enhancement methods to tackle these problems: 1) spatial interpolation with modified matrix completion; 2) a synthetic training data set based on the characteristics of the real U.K. velocity map; and 3) enhanced overwrapping techniques. Using velocity maps spanning 2015-2019, our framework detects several areas of coal mining subsidence, uplift due to dewatering, slate quarries, landslides, and tunnel engineering works. The results demonstrate the potential applicability of the proposed framework to the development of automated ground motion analysis systems. Nantheera Anantrasirichai, Juliet Biggs, Krisztina Kelevitz, Zahra Sadeghi, Tim J. Wright, Alin Achim, David Bull 0001 |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | River Planform Extraction From High-Resolution SAR Images via Generalized Gamma Distribution Superpixel ClassificationabstractThe extraction of river planforms from remotely sensed satellite images is a task of crucial importance to many applications such as land planning, water resource monitoring, or flood prediction. In this article, we present a novel framework for the extraction of rivers from synthetic aperture radar (SAR) images, based on superpixel segmentation and subsequent classification. Superpixel segmentation is achieved by a modeling of the image pixels' amplitudes and spatial coordinates as a finite mixture model, where the generalized Gamma distribution is used to model accurately a variety of high-resolution SAR scenes. A number of features describing image texture and statistics are extracted on a superpixel level, facilitating the identification of river superpixels-planforms are then extracted by unsupervised, agglomerative clustering, thus eliminating the need for labeled training data. We present the results of our proposed method on the ICEYE-X2 and SENTINEL-1 SAR data, demonstrating its ability to produce pixel-accurate river masks. Odysseas A. Pappas, Nantheera Anantrasirichai, Alin Achim, Byron A. Adams |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2020 | Image fusion via sparse regularization with non-convex penalties
Nantheera Anantrasirichai, Rencheng Zheng, Ivan W. Selesnick, Alin Achim |
Pattern Recognit. Lett. | 1 |
| 2020 | Corrigendum to "Image fusion via sparse regularization with non-convex penalties" Pattern Recognition Letters Volume 131, March 2020, Pages 355-360
Nantheera Anantrasirichai, Rencheng Zheng, Ivan W. Selesnick, Alin Achim |
Pattern Recognit. Lett. | 1 |
| 2019 | Defectnet: Multi-Class Fault Detection on Highly-Imbalanced DatasetsabstractAs a data-driven method, the performance of deep convolutional neural networks (CNN) relies heavily on training data. The prediction results of traditional networks give a bias toward larger classes, which tend to be the background in the semantic segmentation task. This becomes a major problem for fault detection, where the targets appear very small on the images and vary in both types and sizes. In this paper we propose a new network architecture, DefectNet, that offers multi-class (including but not limited to) defect detection on highly-imbalanced datasets. DefectNet consists of two parallel paths, which are a fully convolutional network and a dilated convolutional network to detect large and small objects respectively. We propose a hybrid loss maximising the usefulness of a dice loss and a cross entropy loss, and we also employ the leaky rectified linear unit (ReLU) to deal with rare occurrence of some targets in training batches. The prediction results show that our DefectNet outperforms state-of-the-art networks for detecting multi-class defects with the average accuracy improvement of approximately 10% on a wind turbine. Nantheera Anantrasirichai, David Bull 0001 |
ICIP | 1 |
| 2018 | Atmospheric Turbulence Mitigation for Sequences with Moving Objects Using Recursive Image FusionabstractThis paper describes a new method for mitigating the effects of atmospheric distortion on observed sequences that include large moving objects. In order to provide accurate detail from objects behind the distorting layer, we solve the space-variant distortion problem using recursive image fusion based on the Dual Tree Complex Wavelet Transform (DT-CWT). The moving objects are detected and tracked using the improved Gaussian mixture models (GMM) and Kalman filtering. New fusion rules are introduced which work on the magnitudes and angles of the DT-CWT coefficients independently to achieve a sharp image and to reduce atmospheric distortion, respectively. The subjective results show that the proposed method achieves better video quality than other existing methods with competitive speed. Nantheera Anantrasirichai, Alin Achim, David Bull 0001 |
ICIP | 1 |
| 2018 | Fixation Prediction and Visual Priority Maps for Biped LocomotionabstractThis paper presents an analysis of the low-level features and key spatial points used by humans during locomotion over diverse types of terrain. Although, a number of methods for creating saliency maps and task-dependent approaches have been proposed to estimate the areas of an image that attract human attention, none of these can straightforwardly be applied to sequences captured during locomotion, which contain dynamic content derived from a moving viewpoint. We used a novel learning-based method for creating a visual priority map informed by human eye tracking data. Our proposed priority map is created based on two fixation types: first exploiting the observation that humans search for safe foot placement and second that they observe the edges of a path as a guide to safe traversal of the terrain. Texture features and the difference between them, observed at the region around an eye position, are employed within a support vector machine to create a visual priority map for biped locomotion. The results show that our proposed method outperforms the state-of-the-art, particularly for more complex terrains, where achieving smooth locomotion needs more attention on the traversing path. Nantheera Anantrasirichai, Katherine A. J. Daniels, Jeremy F. Burn, Iain D. Gilchrist, David Bull 0001 |
IEEE Trans. Cybern. | 1 |
| 2017 | Line detection in speckle images using Radon transform and ℓ1 regularizationabstractBoundaries and lines in medical images are important structures as they can delineate between tissue types, organs, and membranes. Although, a number of image enhancement and segmentation methods have been proposed to detect lines, none of these have considered line artefacts, which are more difficult to visualise as they are not physical structures, yet are still meaningful for clinical interpretation. This paper presents a novel method to restore lines, including line artefacts, in speckle images. We address this as a sparse estimation problem using a convex optimisation technique based on a Radon transform and sparsity regularisation (ℓ1norm). This problem divides into subproblems which are solved using the alternating direction method of multipliers, thereby achieving line detection and deconvolution simultaneously. The results for both simulated and in vivo ultrasound images show that the proposed method outperforms existing methods, in particular for detecting B-lines in lung ultrasound images, where the performance can be improved by up to 30 %. Nantheera Anantrasirichai, Marco Allinovi, Wesley Hayes, David Bull 0001, Alin Achim |
ICASSP | 1 |
| 2017 | Line Detection as an Inverse Problem: Application to Lung Ultrasound ImagingabstractThis paper presents a novel method for line restoration in speckle images. We address this as a sparse estimation problem using both convex and non-convex optimization techniques based on the Radon transform and sparsity regularization. This breaks into subproblems, which are solved using the alternating direction method of multipliers, thereby achieving line detection and deconvolution simultaneously. We include an additional deblurring step in the Radon domain via a total variation blind deconvolution to enhance line visualization and to improve line recognition. We evaluate our approach on a real clinical application: the identification of B-lines in lung ultrasound images. Thus, an automatic B-line identification method is proposed, using a simple local maxima technique in the Radon transform domain, associated with known clinical definitions of line artefacts. Using all initially detected lines as a starting point, our approach then differentiates between B-lines and other lines of no clinical significance, including Z-lines and A-lines. We evaluated our techniques using as ground truth lines identified visually by clinical experts. The proposed approach achieves the best B-line detection performance as measured by the F score when a non-convex [Formula: see text] regularization is employed for both line detection and deconvolution. The F scores as well as the receiver operating characteristic (ROC) curves show that the proposed approach outperforms the state-of-the-art methods with improvements in B-line detection performance of 54%, 40%, and 33% for [Formula: see text], [Formula: see text], and [Formula: see text], respectively, and of 24% based on ROC curve evaluations. Nantheera Anantrasirichai, Wesley Hayes, Marco Allinovi, David Bull 0001, Alin Achim |
IEEE Trans. Medical Imaging | 1 |
| 2016 | Visual salience and priority estimation for locomotion using a deep convolutional neural networkabstractThis paper presents a novel method of salience and priority estimation for the human visual system during locomotion. This visual information contains dynamic content derived from a moving viewpoint. The priority map, ranking key areas on the image, is created from probabilities of gaze fixations, merged from bottom-up features and top-down control on the locomotion. Two deep convolutional neural networks (CNNs), inspired by models of the primate visual system, are employed to capture local salience features and compute probabilities. The first network operates through the foveal and peripheral areas around the eye positions. The second network obtains the importance of fixated points that have long durations or multiple visits, of which such areas need more times to process or to recheck to ensure smooth locomotion. The results show that our proposed method outperforms the state-of-the-art by up to 30 %, computed from average of four well known metrics for saliency estimation. Nantheera Anantrasirichai, Iain D. Gilchrist, David Bull 0001 |
ICIP | 1 |
| 2016 | Fixation identification for low-sample-rate mobile eye trackersabstractThis paper presents a novel method of fixation identification for mobile eye trackers. The most significant benefit of our method over the state-of-the-art is that it achieves high accuracy for low-sample-rate devices worn during locomotion. This in turn delivers higher quality datasets for further use in human behaviour research, robotics and the development of guidance aids for the visually impaired. The proposed method employs temporal characteristics of the eye positions combined with statistical visual features extracted using a deep convolutional neural network, inspired by models of the primate visual system, through the fovea and peripheral areas around the eye positions. The results show that the proposed method outperforms existing methods by up to 16 % in terms of classification accuracy. Nantheera Anantrasirichai, Iain D. Gilchrist, David Bull 0001 |
ICIP | 1 |
| 2015 | Robust texture features based on undecimated dual-tree complex wavelets and local magnitude binary patternsabstractImage degradation due to illumination change, blur and noise can have a significant influence on classification performance, and yet no descriptors that perform well under these conditions exist. We propose a novel method for obtaining texture features, robust to these distortions, based on an undecimated dual-tree complex wavelet transform (UDT-CWT)1. As the UDT-CWT provides a local spatial relationship between scales, we can straightforwardly create bit-planes of the images representing local phases of wavelet coefficients. Magnitudes of the UDT-CWT are captured via a local binary pattern (LBP), after discarding some of the finest scales that are most affected by the blur and noise. A histogram of the resulting binary code words then forms the features used in texture classification. Results show that our approach outperforms existing methods, that claim to be invariant to feature degradations. Nantheera Anantrasirichai, Jeremy F. Burn, David Bull 0001 |
ICIP | 1 |
| 2015 | Undecimated Dual-Tree Complex Wavelet Transforms
Paul R. Hill, Nantheera Anantrasirichai, Alin Achim, Mohammed E. Al-Mualla, David Bull 0001 |
Signal Process. Image Commun. | 2 |
| 2015 | Terrain Classification From Body-Mounted Cameras During Human LocomotionabstractThis paper presents a novel algorithm for terrain type classification based on monocular video captured from the viewpoint of human locomotion. A texture-based algorithm is developed to classify the path ahead into multiple groups that can be used to support terrain classification. Gait is taken into account in two ways. Firstly, for key frame selection, when regions with homogeneous texture characteristics are updated, the frequency variations of the textured surface are analyzed and used to adaptively define filter coefficients. Secondly, it is incorporated in the parameter estimation process where probabilities of path consistency are employed to improve terrain-type estimation. When tested with multiple classes that directly affect mobility-a hard surface, a soft surface, and an unwalkable area-our proposed method outperforms existing methods by up to 16%, and also provides improved robustness. Nantheera Anantrasirichai, Jeremy F. Burn, David Bull 0001 |
IEEE Trans. Cybern. | 1 |
| 2014 | Orientation estimation for planar textured surfaces based on complex waveletsabstractThe gradient of a road or terrain influences the appropriate speed and power of a vehicle traversing it. Therefore, gradient prediction is necessary if autonomous vehicles are to optimise their locomotion. This paper presents a novel texture-based method for estimating the orientation of planar surfaces under the basic assumption of homogeneity. Based on a backprojection technique, we propose a new method for measuring texture consistency in the Dual-Tree Complex Wavelet Transform (DT-CWT) domain. Texture histograms computed with a new anti-aliasing compensation approach are employed to find the rotation of a planar surface. The proposed method performs well with various types of textured surfaces and outperforms other existing methods with significantly reduced computational complexity up to 35%. Nantheera Anantrasirichai, Jeremy F. Burn, David Bull 0001 |
ICIP | 1 |
| 2014 | Robust texture features for blurred images using Undecimated Dual-Tree Complex WaveletsabstractThis paper presents a new descriptor for texture classification. The descriptor is rotationally invariant and blur insensitive, which provides great benefits for various applications that suffer from out-of-focus content or involve fast moving or shaking cameras. We employ an Undecimated Dual-Tree Complex Wavelet Transform (UDT-CWT) [1] to extract texture features. As the UDT-CWT fully provides local spatial relationship between scales and subband orientations, we can straightforwardly create bit-planes of the images representing local phases of wavelet coefficients. We also discard some of the finest decomposition levels where are most affected by the blur. A histogram of the resulting code words is created and used as features in texture classification. Experimental results show that our approach outperforms existing methods by up to 40% for synthetic blurs and up to 30% for natural video content due to camera motion when walking. Nantheera Anantrasirichai, Jeremy F. Burn, David Bull 0001 |
ICIP | 1 |
| 2013 | Projective image restoration using sparsity regularizationabstractThis paper presents a method of image restoration for projective ground images which lie on a projection orthogonal to the camera axis. The ground images are initially transformed using homography, and then the proposed image restoration is applied. The process is performed in the dual-tree complex wavelet transform domain in conjunction with L0 reweighting and L2 minimisation (L0RL2) employed to solve this ill-posed problem. We also propose instant estimation of a blur kernel arising from the projective transform and the subsequent interpolation of sparse data. Subjective results show significant improvement of image quality. Furthermore, classification of surface type at various distances (evaluated using a support vector machine classifier) is also improved for the images restored using our proposed algorithm. Nantheera Anantrasirichai, Jeremy F. Burn, David Bull 0001 |
ICIP | 1 |
| 2013 | Adaptive-weighted bilateral filtering for optical coherence tomographyabstractThis paper presents an image enhancement method for retinal optical coherence tomography (OCT) images. Raw OCT images contain a large amount of speckle which causes images to be grainy and very low contrast. The raw OCT images thus need to be processed before any clinical interpretation is made. We propose a novel method to remove speckle, while preserving useful information contained in each retinal layer. The process starts with multi-scale despeckling based on a dual-tree complex wavelet transform (DT-CWT). Then, we further enhance the OCT image through a smoothing process that uses a novel adaptive-weighted bilateral filter (AWBF). This offers the desirable property of preserving texture within the OCT images. Glaucoma classification results confirm that our method can significantly enhance the clinical usefulness of OCT images. Nantheera Anantrasirichai, Lindsay Nicholson, James E. Morgan, Irina Erchova, Alin Achim |
ICIP | 1 |
| 2013 | Curvelet domain image fusion of OCT and fundus imagery using convolution of Meridian distributionsabstractThis paper presents a novel statistical model based method aimed at fusing Optical Coherence Tomography and Fundus Photographic imagery of the eye. The presented method utilises the Discrete Curvelet Transform to decompose the images into sub-band coefficients. The Meridian distribution, a specialized case of the generalized Cauchy distribution, is used to model the curvelet decomposition coefficients. The convolution of the input image distributions is used as a probabilistic prior for modelling the fused image coefficients. Experimental results show this method to provide very high-quality fusion results. Odysseas A. Pappas, Nantheera Anantrasirichai, Lindsay Nicholson, James E. Morgan, Irina Erchova, Alin Achim |
ICIP | 2 |
| 2013 | Atmospheric Turbulence Mitigation Using Complex Wavelet-Based FusionabstractRestoring a scene distorted by atmospheric turbulence is a challenging problem in video surveillance. The effect, caused by random, spatially varying, perturbations, makes a model-based solution difficult and in most cases, impractical. In this paper, we propose a novel method for mitigating the effects of atmospheric distortion on observed images, particularly airborne turbulence which can severely degrade a region of interest (ROI). In order to extract accurate detail about objects behind the distorting layer, a simple and efficient frame selection method is proposed to select informative ROIs only from good-quality frames. The ROIs in each frame are then registered to further reduce offsets and distortions. We solve the space-varying distortion problem using region-level fusion based on the dual tree complex wavelet transform. Finally, contrast enhancement is applied. We further propose a learning-based metric specifically for image quality assessment in the presence of atmospheric distortion. This is capable of estimating quality in both full- and no-reference scenarios. The proposed method is shown to significantly outperform existing methods, providing enhanced situational awareness in a range of surveillance scenarios. Nantheera Anantrasirichai, Alin Achim, Nick G. Kingsbury, David Bull 0001 |
IEEE Trans. Image Process. | 1 |
| 2012 | Mitigating the effects of atmospheric distortion using DT-CWT fusionabstractThis paper describes a new method for mitigating the effects of atmospheric distortion on observed images, particularly airborne turbulence which degrades a region of interest (ROI). In order to provide accurate detail from objects behind the distorting layer, a simple and efficient frame selection method is proposed to pick informative ROIs from only good-quality frames. We solve the space-variant distortion problem using region-based fusion based on the Dual Tree Complex Wavelet Transform (DT-CWT). We also propose an object alignment method for pre-processing the ROI since this can exhibit significant offsets and distortions between frames. Simple haze removal is used as the final step. The proposed method performs very well with atmospherically distorted videos and outperforms other existing methods. Nantheera Anantrasirichai, Alin Achim, David Bull 0001, Nick G. Kingsbury |
ICIP | 1 |
| 2011 | Colour volumetric compression for realistic view synthesis applications
Nantheera Anantrasirichai, Cedric Nishan Canagarajah, David W. Redmill, Akbar Sheikh Akbari, David Bull 0001 |
Multim. Tools Appl. | 1 |
| 2010 | Spatiotemporal super-resolution for low bitrate H.264 videoabstractSuper-resolution and frame interpolation enhance low resolution low-framerate videos. Such techniques are especially important for limited bandwidth communications. This paper proposes a novel technique to up-scale videos compressed with H.264 at low bit-rate both in spatial and temporal dimensions. A quantisation noise model is used in the super-resolution estimator, designed for low bitrate video, and a weighting map for decreasing inaccuracy of motion estimation are proposed. Results show improvement both in rate-distortion and perceived image quality. Nantheera Anantrasirichai, Cedric Nishan Canagarajah |
ICIP | 1 |
| 2010 | In-Band Disparity Compensation for Multiview Image Compression and View SynthesisabstractThis paper presents a novel framework to achieve scalable multiview image compression and view synthesis. The open-loop wavelet-lifting scheme for geometric filtering has been exploited to achieve signal-to-noise ratio scalability and view-type scalability (mono, stereo, or multiview). Spatial scalability is achieved by employing in-band prediction which removes correlations among subbands (level-by-level) via shift-invariant references obtained by overcomplete discrete wavelet transforms. We propose a novel in-band disparity compensated view filtering approach, akin to motion compensated temporal filtering, for achieving a scalable multiview codec. In our codec, hybrid prediction is proposed to deal with occlusions, and a novel cost function in dynamic programming (DP) for disparity estimation is introduced to improve view synthesis quality. Experiments show comparable results at full resolution and significant improvements at coarser resolutions, compared to a conventional spatial prediction scheme. View synthesis efficiency is extensively improved by utilizing disparity estimation from the proposed DP approach. Nantheera Anantrasirichai, Cedric Nishan Canagarajah, David W. Redmill, David Bull 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 1 |
| 2009 | Enhanced spatially interleaved DVC using diversity and selective feedbackabstractSystems with cheap/simple/power efficient encoders but complex decoders make applications such as low cost, low power remote sensors practical. Bandwidth considerations however are still an issue and compression efficiency has to remain high. In this paper, we present a distributed video codec (DVC) that we are developing with the aim of achieving such a low power paradigm at the cost of only a small compression performance deficit relative to the current state of the art, H.264. The proposed system employs spatial interleaving of KEY and Wyner-Ziv data which allows efficient side information (SI) generation through block-based error concealment, a Gray code that increases the accuracy of bit probability estimation, and a diversity scheme that produces more reliable results by exploiting multiple SI generated data. Simulation results show an improvement of the proposed scheme over H.264 intra coding of up to 1.5 dB. We additionally propose two mechanisms for selective parity bit feedback requests that can further reduce the WZ bitrate by up to 15%. Nantheera Anantrasirichai, Dimitris Agrafiotis, David Bull 0001 |
ICASSP | 1 |
| 2008 | A concealment based approach to distributed video codingabstractThis paper presents a concealment based approach to distributed video coding that uses hybrid key/WZ frames via an FMO type interleaving of macroblocks. Our motivation stems from a previous work of ours that showed promising results relative to the more common approach of splitting the sequence in key and WZ frames. In this paper, we extend our previous scheme to the case of I-B-P frame structures and transform domain DVC. We additionally introduce a number of enhancements at the decoder including use of spatio-temporal concealment for generating the side information on a MB basis, mode selection for switching between the two concealment approaches and for deciding how the correlation noise is estimated, local (MB wise) correlation noise estimation and modified B frame quantisation. The results presented indicate considerable improvement (up to 30%) compared to corresponding frame extrapolation and frame interpolation schemes. Nantheera Anantrasirichai, Dimitris Agrafiotis, David Bull 0001 |
ICIP | 1 |
| 2007 | Colour Volumetric Compression for Realistic View Synthesis ApplicationsabstractThe colour volumetric data which is constructed from a set of multi-view images is capable of providing realistic immersive experience. However it is not widely applicable due to its manifold increase in bandwidth. This paper presents a novel framework to achieve scalable volumetric compression. Based on wavelet transformation, data rearrangement algorithm is proposed to compact volumetric data leading to high efficiency of transformation. The colour data is also rearranged by using the characteristics of human eye sensitivity. Moreover, the pre-processing for adaptive resolution is proposed in this paper. The low resolution overcomes the limitation of the data transmission at low bit rate, whilst the fine resolution improves the synthesised images' quality. The results of our proposed schemes show significant improvement of the compression performance over the traditional 3D coding. Nantheera Anantrasirichai, Cedric Nishan Canagarajah, David W. Redmill, David Bull 0001 |
ICASSP (1) | 1 |
| 2006 | Dynamic Programming for Multi-View Disparity/Depth EstimationabstractA novel algorithm for disparity/depth estimation from multi-view images is presented. A dynamic programming approach with window-based correlation and a novel cost function is proposed. The smoothness of disparity/depth map is embedded in dynamic programming approach, whilst the window-based correlation increases reliability. The enhancement methods are included, i.e. adaptive window size and shiftable window are used to increase reliability in homogenous areas and to increase sharpness at object boundaries. First, the algorithms estimate depth maps along a single camera axis. The algorithms exploits then combines the depth estimates from different axis to derive a suitable depth map for multi-view images. The proposed scheme outperforms existing approaches in parallel and in the non-parallel camera configurations Nantheera Anantrasirichai, Cedric Nishan Canagarajah, David W. Redmill, David Bull 0001 |
ICASSP (2) | 1 |
| 2006 | Volumetric Representation for Sparse Multi-ViewsabstractIn this paper, we propose a novel volumetric representation for a sparse set of calibrated multi-view images of a non-Lambertian scene. The depth map of each reference view is registered into a volume and a simple algorithm to shape the volume is introduced. Particular colours are defined for each voxel to render a smooth and realistic image. Synthesized results demonstrate the good performance with the proposed scheme, both for parallel-camera and non-parallel-camera geometries. Nantheera Anantrasirichai, Cedric Nishan Canagarajah, David W. Redmill, David Bull 0001 |
ICIP | 1 |
| 2005 | Multi-view image coding with wavelet lifting and in-band disparity compensationabstractThis paper presents a novel framework to achieve scalable multi-view image coding. As open loop operation, the wavelet lifting scheme for geometric filtering has been exploited to overcome the limitation of SNR scalability and to attain view scalability. The essential key for achieving the spatial scalability is the in-band prediction. It removes correlations among subbands level-by-level via shift-invariant references obtained by overcomplete discrete wavelet transforms (ODWT). Additionally, the proposed disparity compensated view filtering is allowed to exploit the different filters and estimation parameters for each resolution level. The experiments show comparable results at full resolution and the significant improvement at coarser resolution over the conventional spatial prediction scheme. Nantheera Anantrasirichai, Cedric Nishan Canagarajah, David Bull 0001 |
ICIP (3) | 1 |