Xiaoyin Xu

dblp:35/46 · DBLP profile ↗
← Back
39ranked-venue papers
9as first author
20since 2021 · last 2026
0000-0003-0813-7979ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 19 · 3 first-author · 11 since 2021Artificial intelligence and machine learning · 16 · 1 first-author · 13 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 5 first-author · 4 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
YearPublicationVenuePosition
2026 Spherical Geometry Diffusion: Generating High-quality 3D Face Geometry via Sphere-anchored Representations
abstract
A fundamental challenge in text-to-3D face generation is achieving high-quality geometry. The core difficulty lies in the arbitrary and intricate distribution of vertices in 3D space, making it challenging for existing models to establish clean connectivity and resulting in suboptimal geometry. To address this, our core insight is to simplify the underlying geometric structure by constraining the distribution onto a simple and regular manifold, a topological sphere. Building on this, we first propose the Spherical Geometry Representation, a novel face representation that anchors geometric signals to uniform spherical coordinates. This guarantees a regular point distribution, from which the mesh connectivity can be robustly reconstructed. Critically, this canonical sphere can be seamlessly unwrapped into a 2D map, creating a perfect synergy with powerful 2D generative models. We then introduce Spherical Geometry Diffusion, a conditional diffusion framework built upon this 2D map. It enables diverse and controllable generation by jointly modeling geometry and texture, where the geometry explicitly conditions the texture synthesis process. Our method's effectiveness is demonstrated through its success in a wide range of tasks: text-to-3D generation, face reconstruction, and text-based 3D editing. Extensive experiments show that our approach substantially outperforms existing methods in geometric quality, textual fidelity, and inference efficiency.
Yunhong Lu, Wenzhe Qian, Xiaoyin Xu, David Gu
AAAI6
2026 Rethinking Personalized T2I Diffusion Models from the Perspective of Redundancy
Xierui Wang, Bohan Lei, Xiaoyin Xu, Fei Wu 0001, Min Zhang 0069
Int. J. Comput. Vis.3
2026 Image compression using optimal transport mapping based on ranking visual saliency
Dongsheng An, Xianfeng Gu, Xiaoyin Xu, Min Zhang 0069
Pattern Recognit.4
2025 SOP-Guided Co-Enhanced CLIP for Action Recognition in Logistics Warehouse Packaging
Wenting Qi, Xiaoyin Xu, Hengle Qin
IEEE Big Data5
2025 InPO: Inversion Preference Optimization with Reparametrized DDIM for Efficient Diffusion Model Alignment
abstract
Without using explicit reward, direct preference optimization (DPO) employs paired human preference data to fine-tune generative models, a method that has garnered considerable attention in large language models (LLMs). However, exploration of aligning text-to-image (T2I) diffusion models with human preferences remains limited. In comparison to supervised fine-tuning, existing methods that align diffusion model suffer from low training efficiency and subpar generation quality due to the long Markov chain process and the intractability of the reverse process. To address these limitations, we introduce DDIM-InPO, an efficient method for direct preference alignment of diffusion models. Our approach conceptualizes diffusion model as a single-step generative model, allowing us to fine-tune the outputs of specific latent variables selectively. In order to accomplish this objective, we first assign implicit rewards to any latent variable directly via a reparameterization technique. Then we construct an Inversion technique to estimate appropriate latent variables for preference optimization. This modification process enables the diffusion model to only fine-tune the outputs of latent variables that have a strong correlation with the preference dataset. Experimental results indicate that our DDIM-InPO achieves state-of-the-art performance with just 400 steps of fine-tuning, surpassing all preference aligning baselines for T2I diffusion models in human preference evaluation tasks.
Yunhong Lu, Hengyuan Cao, Xierui Wang, Xiaoyin Xu
CVPR5
2025 Smoothed Preference Optimization via ReNoise Inversion for Aligning Diffusion Models with Varied Human Preferences
abstract
Direct Preference Optimization (DPO) aligns text-to-image (T2I) generation models with human preferences using pairwise preference data. Although substantial resources are expended in collecting and labeling datasets, a critical aspect is often neglected: *preferences vary across individuals and should be represented with more granularity.* To address this, we propose SmPO-Diffusion, a novel method for modeling preference distributions to improve the DPO objective, along with a numerical upper bound estimation for the diffusion optimization objective. First, we introduce a smoothed preference distribution to replace the original binary distribution. We employ a reward model to simulate human preferences and apply preference likelihood averaging to improve the DPO loss, such that the loss function approaches zero when preferences are similar. Furthermore, we utilize an inversion technique to simulate the trajectory preference distribution of the diffusion model, enabling more accurate alignment with the optimization objective. Our approach effectively mitigates issues of excessive optimization and objective misalignment present in existing methods through straightforward modifications. Experimental results demonstrate that our method achieves state-of-the-art performance in preference evaluation tasks, surpassing baselines across various metrics, while reducing the training costs.
Yunhong Lu, Hengyuan Cao, Xiaoyin Xu, Min Zhang 0069
ICML4
2025 tanh As a robust feature scaling method in training deep learning models with imbalanced data
Aijia Yang, Huai Chen, Taihao Li, Shupeng Liu, Xiaoyin Xu
Pattern Recognit.6
2024 An Optimal Transport-Based Method For Medical Image Generation
abstract
Optimal transport (OT) is an effective technique for mapping between two distributions. Introducing a definition of distance between the distributions, OT provides the best calculable mapping plan, especially from an irregular distribution to a regular one. For this reason, OT has found wide applications, including deep learning, to mitigate some inherent risks in existing techniques. For example, generative adversarial networks (GANs) are widely used to generate medical images but sometimes experience mode collapse, a phenomenon that can be avoided by using OT. We use OT to map latent space distribution in a GAN model to a reasonable Gaussian distribution and generate more convincing and complete medical images to help train medical AI models. Our method consists of three steps. First, we train an auto-encoder as a basic structure of our model. Second, we introduce OT to the trained auto-encoder, which aligns the latent space distribution with a Gaussian distribution. This mapping helps to enforce a more realistic distribution of the generated images, enhancing their clinical relevance. Third, we construct a GAN-type model by training the decoder part of the auto-encoder and discriminator networks. We conducted a comprehensive experimental evaluation on a diverse medical dataset by comparing the performance of our proposed method to other state-of-the-art models. Results show that our method performs relatively better than other models by FID scores and provides more detailed information.
Bohan Lei, Yueting Zhuang, Xiaoyin Xu
ICIP3
2024 Domain Knowledge Enhanced Vision-Language Pretrained Model for Dynamic Facial Expression Recognition
abstract
Dynamic facial expression recognition (DFER) is a rapidly developing field that focuses on recognizing facial expressions in video sequences. However, the complex temporal modeling caused by noisy frames, along with the limited training data significantly hinder the further development of DFER. Previous efforts in this domain have been limited as they tackled these issues separately. Inspired by recent advances of pretrained vision-language models (e.g., CLIP), we propose to leverage it to jointly address the two limitations in DFER. Since the raw CLIP model lacks the ability to model temporal relationships and determine the optimal task-related textual prompts, we utilize DFER-specific domain knowledge, including characteristics of temporal correlations and relationships between facial behavior descriptions at different levels, to guide the adaptation of CLIP to DFER. Specifically, we propose enhancements to CLIP's visual encoder through the design of a hierarchical video encoder that captures both short- and long-term temporal correlations in DFER. Meanwhile, we align facial expressions with action units through prior knowledge to construct semantically rich textual prompts, which are further enhanced with visual contents. Furthermore, we introduce a class-aware consistency regularization mechanism that adaptively filters out noisy frames, bolstering the model's robustness against interference. Extensive experiments on three in-the-wild dynamic facial expression datasets demonstrate that our method outperforms the state-of-the-art DFER approaches. The code is available at https://github.com/liliupeng28/DK-CLIP.
Liupeng Li, Yuhua Zheng, Shupeng Liu, Xiaoyin Xu, Taihao Li
ACM Multimedia4
2024 Use estimated signal and noise to adjust step size for image restoration
Shupeng Liu, Taihao Li, Huai Chen, Xiaoyin Xu
Pattern Recognit. Lett.5
2024 Design of a differentiable L-1 norm for pattern recognition and machine learning
Min Zhang 0069, Taihao Li, Shupeng Liu, Xianfeng Gu, Xiaoyin Xu
Pattern Recognit. Lett.7
2024 Autoencoder-based conditional optimal transport generative adversarial network for medical image generation
abstract
Recently, there has been a significant surge of interest in medical image generation. In this study, we developed a model known as AE-COT-GAN (autoencoder-based conditional optimal transport generative adversarial network) to generate medical images that belong to specific categories. The primary objective of our research is to address the prevalent challenges often encountered during the training of generative adversarial networks (GANs), including issues such as mode collapse and mode mixing. The training process of our model encompasses three fundamental components. First, we employ an autoencoder model to obtain a low-dimensional manifold representation of real images. Second, we apply extended semi-discrete optimal transport to map Gaussian noise distribution to the latent space distribution and obtain corresponding labels effectively. This procedure leads to the generation of new latent codes with known labels. Finally, we integrate a GAN to train the decoder further to generate medical images. To evaluate the performance of the AE-COT-GAN model, we conducted experiments on two medical image datasets, namely DermaMNIST and BloodMNIST. The model’s performance was compared with state-of-the-art generative models. Results show that the AE-COT-GAN model had excellent performance in generating medical images. Moreover, it effectively addressed the common issues associated with traditional GANs.
Jun Wang 0039, Bohan Lei, Xiaoyin Xu, Xianfeng Gu, Min Zhang 0069
Vis. Informatics4
2023 Volumetric Optimal Transportation by Fast Fourier Transform
Na Lei, Dongsheng An, Min Zhang 0069, Xiaoyin Xu, Xianfeng Gu
ICLR4
2023 Multi-modal Semi-supervised Evidential Recycle Framework for Alzheimer's Disease Classification
Yingjie Feng, Wei Chen 0130, Xianfeng Gu, Xiaoyin Xu, Min Zhang 0069
MICCAI (1)4
2023 TransLiver: A Hybrid Transformer Model for Multi-phase Liver Lesion Classification
Xierui Wang, Hanning Ying, Xiaoyin Xu, Xiujun Cai
MICCAI (2)3
2022 Efficient Optimal Transport Algorithm by Accelerated Gradient Descent
abstract
Optimal transport (OT) plays an essential role in various areas like machine learning and deep learning. However, computing discrete optimal transport plan for large scale problems with adequate accuracy and efficiency is still highly challenging. Recently, methods based on the Sinkhorn algorithm add an entropy regularizer to the prime problem and get a trade off between efficiency and accuracy. In this paper, we propose a novel algorithm to further improve the efficiency and accuracy based on Nesterov's smoothing technique. Basically, the non-smooth c-transform of the Kantorovich potential is approximated by the smooth Log-Sum-Exp function, which finally smooths the original non-smooth Kantorovich dual functional. The smooth Kantorovich functional can be optimized by the fast proximal gradient algorithm (FISTA) efficiently. Theoretically, the computational complexity of the proposed method is lower than current estimation of the Sinkhorn algorithm in terms of the precision. Empirically, compared with the Sinkhorn algorithm, our experimental results demonstrate that the proposed method achieves faster convergence and better accuracy with the same parameter.
Dongsheng An, Na Lei, Xiaoyin Xu, Xianfeng Gu
AAAI3
2022 Image Compression Based on Importance Using Optimal Mass Transportation Map
abstract
Demand for efficient image transmission and storage is increasing rapidly because of the continuing growth of multimedia technology and VR and AR applications. In this paper, we proposed an image compression method based on the recognition of importance of regions in images. As not all the information in an image is equally useful, we can identify important regions in an image for high fidelity compression and accept a comparatively more lossy compression about less important regions of the image. First, we segment images to two parts, namely, foreground and background, where the foreground represents the more important component and the background is of less importance. Second, we apply optimal mass transportation mapping in a GAN (generative adversarial network) framework to both the foreground and background to magnify the foreground and shrink the background while keeping the shape and total image area unchanged. As a result, in the processed image, the ratio of foreground to background is larger than the corrresponding ratio in the original image. This ratio is controllable in our process, giving users the ability to control the degree of compression. The GAN-processed image is then used for compression. To restore the image, we apply a GAN model to the compressed image and recover the ratio of foreground and background using an optimal mass transportation map. Test results show that our method is highly effective in reconstructing detail of important components in compressed images while achieving a high compression ratio.
Dongsheng An, Yingjie Feng, Xianfeng Gu, Xiaoyin Xu, Min Zhang 0069
ICIP5
2022 End-to-End Evidential-Efficient Net for Radiomics Analysis of Brain MRI to Predict Oncogene Expression and Overall Survival
Yingjie Feng, Jun Wang 0039, Dongsheng An, Xianfeng Gu, Xiaoyin Xu, Min Zhang 0069
MICCAI (3)5
2022 A new framework of designing iterative techniques for image deblurring
Min Zhang 0069, Geoffrey S. Young, Yanmei Tie, Xianfeng Gu, Xiaoyin Xu
Pattern Recognit.5
2021 Cortical Surface Shape Analysis Based on Alexandrov Polyhedra
abstract
Shape analysis has been playing an important role in early diagnosis and prognosis of neurodegenerative diseases such as Alzheimer's diseases (AD). However, obtaining effective shape representations remains challenging. This paper proposes to use the Alexandrov polyhedra as surface-based shape signatures for cortical morphometry analysis. Given a closed genus-0 surface, its Alexandrov polyhedron is a convex representation that encodes its intrinsic geometry information. We propose to compute the polyhedra via a novel spherical optimal transport (OT) computation. In our experiments, we observe that the Alexandrov polyhedra of cortical surfaces between pathology-confirmed AD and cognitively unimpaired individuals are significantly different. Moreover, we propose a visualization method by comparing local geometry differences across cortical surfaces. We show that the proposed method is effective in pinpointing regional cortical structural changes impacted by AD.
Min Zhang 0069, Na Lei, Xiaoyin Xu, Yalin Wang 0001, Xianfeng Gu
ICCV6
2019 A new design in iterative image deblurring for improved robustness and performance
abstract
In many applications, image deblurring is a pre-requisite to improve the sharpness of an image before it can be further processed. Iterative methods are widely used for deblurring images but care must be taken to ensure that the iterative process is robust, meaning that the process does not diverge and reaches the solution reasonably fast, two goals that sometimes compete against each other. In practice, it remains challenging to choose parameters for the iterative process to be robust. We propose a new approach consisting of relaxed initialization and pixel-wise updates of the step size for iterative methods to achieve robustness. The first novel design of the approach is to modify the initialization of existing iterative methods to stop a noise term from being propagated throughout the iterative process. The second novel design is the introduction of a vectorized step size that is adaptively determined through the iteration to achieve higher stability and accuracy in the whole iterative process. The vectorized step size aims to update each pixel of an image individually, instead of updating all the pixels by the same factor. In this work, we implemented the above designs based on the Landweber method to test and demonstrate the new approach. Test results showed that the new approach can deblur images from noisy observations and achieve a low mean squared error with a more robust performance.
Taihao Li, Huai Chen, Shupeng Liu, Shunren Xia, Xinhua Cao, Geoffrey S. Young, Xiaoyin Xu
Pattern Recognit.8
2018 Using feature points and angles between them to recognise facial expression by a neural network approach
abstract
In this study, the authors propose a neural network (NN) method that uses feature points and the angles formed between the points to recognise facial expressions. Accurate facial expression recognition is an important part of affective computing with many practical applications. Yet, achieving acceptable levels of facial recognition accuracy has proven difficult. Feature points and the distances between the points are used to model basic expressions in NN‐based approaches, but, in some cases, they cannot generate satisfactory performance. They expand on the characterisation of facial expression by considering the angles formed between feature points to augment the amount of information that is sent to the NNs. Furthermore, to circumvent a common challenge in facial expressions recognition, which is the difficulty of differentiating among several expressions, they designed a post‐processing step to assess the output of the NN against a threshold. The whole method makes a decision only when the output of the NN exceeds the threshold. Otherwise, the frame under consideration is assigned to a ‘no decision’ class. They tested our method on the widely used facial expression CK + database and found that it can achieve good accuracy.
Taihao Li, Cuifen Du, Tuya Naren, Shupeng Liu, Jianshe Zhou, Xiaoyin Xu
IET Image Process.7
2017 Joint volumetric extraction and enhancement of vasculature from low-SNR 3-D fluorescence microscopy images
Sepideh Almasi, Ayal Ben-Zvi, Baptiste Lacoste, Chenghua Gu, Eric L. Miller 0001, Xiaoyin Xu
Pattern Recognit.6
2015 A novel method for identifying a graph-based representation of 3-D microvascular networks from fluorescence microscopy image stacks
Sepideh Almasi, Xiaoyin Xu, Ayal Ben-Zvi, Baptiste Lacoste, Chenghua Gu, Eric L. Miller 0001
Medical Image Anal.2
2014 A New Iterative Triclass Thresholding Technique in Image Segmentation
abstract
We present a new method in image segmentation that is based on Otsu's method but iteratively searches for subregions of the image for segmentation, instead of treating the full image as a whole region for processing. The iterative method starts with Otsu's threshold and computes the mean values of the two classes as separated by the threshold. Based on the Otsu's threshold and the two mean values, the method separates the image into three classes instead of two as the standard Otsu's method does. The first two classes are determined as the foreground and background and they will not be processed further. The third class is denoted as a to-be-determined (TBD) region that is processed at next iteration. At the succeeding iteration, Otsu's method is applied on the TBD region to calculate a new threshold and two class means and the TBD region is again separated into three classes, namely, foreground, background, and a new TBD region, which by definition is smaller than the previous TBD regions. Then, the new TBD region is processed in the similar manner. The process stops when the Otsu's thresholds calculated between two iterations is less than a preset threshold. Then, all the intermediate foreground and background regions are, respectively, combined to create the final segmentation result. Tests on synthetic and real images showed that the new iterative method can achieve better performance than the standard Otsu's method in many challenging cases, such as identifying weak objects and revealing fine structures of complex objects while the added computational cost is minimal.
Hongmin Cai, Xinhua Cao, Weiming Xia, Xiaoyin Xu
IEEE Trans. Image Process.5
2011 Effective image noise removal based on difference eigenvalue
abstract
Preservation of fine feature of an image is essential during the process of noise removal, especially via some types of smoothing such as using diffusion process-based methods to enhance images. In this paper, we present a new edge indicator called difference eigenvalue to measure image gradient magnitude in the diffusion process. Based on the eigenvalues of the Hessian matrix, the difference eigenvalue manifest itself in terms of structural information of an image. We adapt the new edge indicator to a diffusion model to achieve a better balance between noise removal and detail preservation. Experiments on both synthetic and real images show that the new model can obtain good results and outperforms existing methods.
Haiying Tian, Hongmin Cai, Jian-Huang Lai, Xiaoyin Xu
ICIP4
2009 Robust 3D reconstruction and identification of dendritic spines from optical microscopy imaging
Firdaus Janoos, Kishore Mosaliganti, Xiaoyin Xu, Raghu Machiraju, Kun Huang 0001, Stephen T. C. Wong
Medical Image Anal.3
2009 A method based on rank-ordered filter to detect edges in cellular image
Xiaoyin Xu, Yaming Wang
Pattern Recognit. Lett.1
2008 Classification and Uncertainty Visualization of Dendritic Spines from Optical Microscopy Imaging
abstract
Abstract Neuronal dendrites and their spines affect the connectivity of neural networks, and play a significant role in many neurological conditions. Neuronal function is observed to be closely correlated with the appearance, disappearance and morphology of the spines. Automatic 3‐D reconstruction of neurons from light microscopy images, followed by the identification, classification and visualization of dendritic spines is therefore essential for studying neuronal physiology and biophysical properties. In this paper, we present a method to reconstruct dendrites using a surface representation of the dendrite. The 1‐D skeleton of the dendritic surface is then extracted by a medial geodesic function that is robust and topologically correct. This is followed by a Bayesian identification and classification of the spines. The dendrite and spines are visualized in a manner that displays the spines' types and the inherent uncertainty in identification and classification. We also describe a user study conducted to validate the accuracy of the classification and the efficacy of the visualization.
Firdaus Janoos, Boonthanome Nouanesengsy, Xiaoyin Xu, Raghu Machiraju, Stephen T. C. Wong
Comput. Graph. Forum3
2008 Using nonlinear diffusion and mean shift to detect and connect cross-sections of axons in 3D optical microscopy images
Hongmin Cai, Xiaoyin Xu, Ju Lu, Jeff Lichtman, Siu-Pang Yung, Stephen T. C. Wong
Medical Image Anal.2
2007 A New Nonlinear Diffusion Method to Improve Image Quality
abstract
We propose a nonlinear diffusion method based on the gradient vector field construction to remove noises in image while preserving fine details. The blocky effect and over-smoothing, as usually seen in images processed by diffusion operators, are greatly reduced by our method. Results obtained from various images, including synthetic and magnetic resonance imaging (MRI), are used to demonstrate the performance of our new method. Comparing it with other diffusion methods, we find it obtains better performance in terms of removing noises without destroying detail features of images.
Xiaoyin Xu, Hongmin Cai, Siu-Pang Yung, Stephen T. C. Wong
ICIP (1)2
2004 Adaptive two-pass rank order filter to remove impulse noise in highly corrupted images
abstract
In this paper, we present an adaptive two-pass rank order filter to remove impulse noise in highly corrupted images. When the noise ratio is high, rank order filters, such as the median filter for example, can produce unsatisfactory results. Better results can be obtained by applying the filter twice, which we call two-pass filtering. To further improve the performance, we develop an adaptive two-pass rank order filter. Between the passes of filtering, an adaptive process is used to detect irregularities in the spatial distribution of the estimated impulse noise. The adaptive process then selectively replaces some pixels changed by the first pass of filtering with their original observed pixel values. These pixels are then kept unchanged during the second filtering. In combination, the adaptive process and the second filter eliminate more impulse noise and restore some pixels that are mistakenly altered by the first filtering. As a final result, the reconstructed image maintains a higher degree of fidelity and has a smaller amount of noise. The idea of adaptive two-pass processing can be applied to many rank order filters, such as a center-weighted median filter (CWMF), adaptive CWMF, lower-upper-middle filter, and soft-decision rank-order-mean filter. Results from computer simulations are used to demonstrate the performance of this type of adaptation using a number of basic rank order filters.
Xiaoyin Xu, Eric L. Miller 0001, Dongbin Chen, Mansoor Sarhadi
IEEE Trans. Image Process.1
2003 Minimum entropy regularization in frequency-wavenumber migration to localize subsurface objects
abstract
Optimized versions of frequency-wavenumber (F-K) migration methods are introduced to better focus ground-penetrating radar (GPR) data in applications of shallow subsurface object localization, e.g., landmine remediation. Migration methods are based on the wave equation and operate by backpropagating the received data into the earth so as to localize buried objects. Traditional F-K migration is based on an underlying assumption that the wavefields propagate in a homogeneous medium. The presence of a rough air-ground interface in the GPR case degrades the localization ability. To overcome this problem in the context of the F-K algorithm, we introduce lateral variations in the velocity of waves in the medium. An optimization approach is employed to choose that velocity function that results in a well-focused image where an entropy-like criterion is used to quantify the notion of focus. Extension of the basic method to lossy medium is also described. The utility of these techniques is demonstrated using field data from a number of GPR systems.
Xiaoyin Xu, Eric L. Miller 0001, Carey M. Rappaport
IEEE Trans. Geosci. Remote. Sens.1
2002 Adaptive two-pass median filter to remove impulsive noise
abstract
In this paper, we present an adaptive two-pass median filter to remove impulsive noise. In two-pass median filtering, an image contaminated by impulsive noise is processed by a median filter twice. Median filtering is a non-reversible process, i.e., useful information discarded by the filter cannot be recovered. This behavior becomes more apparent in two-pass median filtering. To correct this problem, between the two filtering processes we introduce an adaptive process to selectively replace some pixels by their original values based on the spatial distribution of estimated impulsive noise. Compared with standard median filtering and two-pass median filtering, better results are obtained in terms of visual appreciation and mean squared error. We use examples to demonstrate the performance of the method.
Xiaoyin Xu, Eric L. Miller 0001
ICIP (1)1
2002 Adaptive difference of Gaussians to improve subsurface object detection using GPR imagery
abstract
We present an adaptive difference of Gaussians (ADOG) method to improve landmine detection from ground penetrating radar (GPR) imagery. GPR is widely used for a variety of subsurface sensing problems including landmine detection and localization. In most all applications, specular reflection from the air-ground interface is the most significant source of interference often eclipsing the returns from the buried object. From an image processing viewpoint, the specular reflection constitutes a sharp edge and the landmine reflected signals are more smooth and of low spatial frequency. As an edge-preserving method, the ADOG keeps the specular reflection intact and removes the object reflected signals. Then by subtracting the ADOG output from the original image, we eliminate the specular reflection and obtain an image with enhanced object reflected signals. Using the difference image, we can obtain better result in detecting buried objects. Results from applying our algorithm on a landmine detection application are used to demonstrate the performance of the method.
Xiaoyin Xu, Eric L. Miller 0001
ICIP (2)1
2002 Optimization of migration method to locate buried object in lossy medium
abstract
We present an optimized frequency-wavenumber (F-K) migration method to localize buried objects such as landmines in lossy medium. F-K migration has been proposed to find the location of a buried object using ground penetrating radar (GPR) data. This approach makes use of a wave equation in the Fourier domain to back-propagate the received wavefield. For GPR applications however, standard F-K migration assumes that the ground surface is flat and the medium is loss-free which are not true in reality. When implemented in the Fourier domain, the wave equation becomes the Helmholtz equation. It is then straightforward to incorporate a complex index of refraction in the Helmholtz equation to describe wave phenomenon in lossy medium. We generalize F-K migration to the case of rough ground surface and lossy medium. In the framework of Tikhonov regularization, we develop an algorithm that optimally alters the wave propagation velocity and the complex index of refraction to take into account of the ground roughness and lossy medium. In the process of searching the optimal velocity and complex index of refraction, the algorithm is constrained to produce an image of minimum entropy. By minimizing the entropy of the resulting image, better results are obtained in terms of enhanced mainlobe, suppressed sidelobes, and reduced noise. We use examples from field data to demonstrate the performance of our method.
Xiaoyin Xu, Eric L. Miller 0001
IGARSS1
2002 Adaptive difference of Gaussians to improve subsurface imagery
abstract
In detection of landmines using ground penetrating radar (GPR), the most significant interference is the specular reflection. Compared with the specular reflection, landmine scattered signals are of small amplitude and difficult to observe. Better detection results can be obtained if the specular reflection can be well separated from the landmine scattered signals. Difference of Gaussians (DOG) is an operation that generates a sharp image from an original image, i.e. in the case of a GPR image, it keeps the specular reflection intact. Therefore by subtracting the DOG output from an original GPR image, we are able to remove the specular reflection and enhance the landmine scattered signals. The DOG takes the difference between two Gaussian curves of zero means and different standard deviations and convolves with the original image. One advantage of the DOG is that the two standard deviations can be chosen properly to suit different applications. We develop an adaptive DOG (ADOG) to process GPR images to improve detection of buried landmines. In the ADOG, the two standard deviations of Gaussians are computed adaptively as a GPR image is scanned from the top to the bottom row. At each row, two windows are used to calculate the two standard deviations and the current row is convolved with the DOG. The output has an unchanged specular reflection while there is a reduced landmine scattered signal. The final image is obtained by subtracting the ADOG output from the original image to remove the specular reflection. The landmine scattered signal is greatly enhanced, allowing more accurate detection.
Xiaoyin Xu, Eric L. Miller 0001
IGARSS1
2002 On the use of contrast stretch and adaptive filter to enhance ground penetrating radar imagery
abstract
We propose using an adaptive filter to enhance ground penetrating radar (GPR) images. It is well known that GPR images are usually dominated by the specular ground reflection. The specular reflection makes the object scattered signals difficult to observe, so it first must be removed before more refined detection and classification processing can be employed. To remove the specular reflection, the biggest challenge is that, due to ground roughness, the reflection cannot be satisfactorily subtracted by some simple methods such as a moving-average filter. Using contrast stretch we can enhance the object reflect signal and then use standard background removal method to eliminate most of the specular reflection. During the contrast stretch, the GPR images may have "streaky" artifacts because of the background removal. To overcome this side-effect, we apply an adaptive filter to remove "streaky" artifacts. Using field data, we show that images of higher quality can be obtained by our method.
Xiaoyin Xu, Eric L. Miller 0001
IGARSS1
2002 Statistical method to detect subsurface objects using array ground-penetrating radar data
abstract
We introduce a combination of high-dimensional analysis of variance (HANOVA) and sequential probability ratio test (SPRT) to detect buried objects from an array ground-penetrating radar (GPR) surveying a region of interest in a progressive manner. Using HANOVA, we exploit the transient characteristic of GPR signals in the time domain to extract information about buried objects at fixed positions of the array. Based on the output of the HANOVA, the SPRT is employed to make detection decisions recursively as the array moves downtrack. The method is on-line implementable and of low computational complexity. Our approach is validated using field-data from two quite different GPR sensing systems designed for landmine detection applications.
Xiaoyin Xu, Eric L. Miller 0001, Carey M. Rappaport, Gary D. Sower
IEEE Trans. Geosci. Remote. Sens.1