EDBT 2026 Demo / reviewers in the wild / expert
Changming Sun
dblp:s/SMSun
· DBLP profile ↗
104ranked-venue papers
15as first author
43since 2021 · last 2026
0000-0001-5943-1989ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 66 · 13 first-author · 23 since 2021Graphics, computer vision, multimedia, augmented reality and games · 36 · 3 first-author · 11 since 2021Applied, interdisciplinary, general and emerging computing · 18 · 14 since 2021Security and privacy · 2Databases, data management, data science and information retrieval · 2 · 1 first-authorHuman-computer interaction and ubiquitous computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Adaptive feature selection-based feature reconstruction network for few-shot learning
Yaohui An, Tao Lei 0003, Junpo Yang, Zicheng Pan, Yongsheng Gao 0001, Changming Sun |
Pattern Recognit. | 9 |
| 2026 | Unrectified-stereo: A new paradigm for stereo matching without epipolar rectification
Xiucai Zhang, Jun Lin 0003, Changming Sun, Wenqi Ma, Yihan Bai, Yuhai Wang, Huanyu Zhao, Yang Liu 0333 |
Pattern Recognit. | 4 |
| 2026 | KG-CMI: Knowledge Graph Enhanced Cross-Mamba Interaction for Medical Visual Question Answering
Xianyao Zheng, Hui Cui 0002, Changming Sun, Xiangyu Li 0004, Ran Su, Leyi Wei, Qiangguo Jin |
IEEE Trans. Ind. Informatics | 4 |
| 2026 | PGST: A Prototype-Guided Parameter-Efficient Network for Spatial Transcriptomics PredictionabstractSpatial transcriptomics (ST) aims to decode spatially resolved gene expression patterns while preserving tissue morphology. Current methods tend to use lower-cost deep learning approaches for gene expression prediction, yet face severe challenges. First, existing methods fail to give sufficient consideration to the spatial specificity of positional encoding inherent in ST; second, they neglect to leverage spatially coherent co-expression patterns across different domains; third, their reliance on linearly weighted aggregation induces vulnerability to noise and distribution shifts; and finally, these architectures exhibit limited parameter efficiency. To address these issues, we introduce prototype-guided network for spatial transcriptomics (PGST), which includes four parts: (1) oriented signal propagation through polar embedding strategy for spatial transcriptomics (PEST); (2) prototype-guided aggregation for global co-feature preservation; (3) global consistency enforcement via shared decoder with reconstruction loss; and (4) lightweight architectural design. Our framework integrates contrastive learning with graph neural networks to balance local-global spatial dependencies and cross-modal consistency. Experimental results on multiple datasets from ST demonstrate the superior performance of our PGST model than existing methods. Yuan He 0016, Kaimiao Hu, Changming Sun, Leyi Wei, Ran Su |
IEEE J. Biomed. Health Informatics | 3 |
| 2025 | ADSA-Net: Addressing Intra- and Inter-Class Variabilities for Severity Assessment of Atopic DermatitisabstractAtopic dermatitis (AD) is a chronic inflammatory skin disorder characterized by recurrent itching, erythema, dryness, and eczematous lesions. Automated AD severity assessment is crucial for cost-effective and precision clinical decision-making but remains challenging. This is due to the subtle contrast variations between key dermatological signs and significant variations in lesion sizes across patients and disease stages. To address these issues, we propose ADSA-Net, which is designed to handle both intra- and inter-class variabilities. ADSA-Net first extracts multi-scale texture-aware features to effectively model variations in lesion size and texture. It then leverages contrastive learning to enhance intra- and inter-class differentiation, strengthening model's discriminatory ability for samples that are difficult to distinguish. Finally, ADSA-Net refines the learning process by leveraging a dynamic feature pool of correctly classified samples to guide the calibration of misclassified instances, enhancing overall accuracy. We further establish a dataset for AD severity assessment. Comprehensive experiments on this dataset show that ADSA-Net significantly outperforms existing state-of-the-art methods. Qiangguo Jin, Xurong Chen, Hui Cui 0002, Changming Sun, Youpeng Deng, Cong Cong 0001, Yuqi Fang, Ran Su, Leyi Wei |
BIBM | 4 |
| 2025 | Temperature Inversion guided Attention for Infrared Small Target DetectionabstractTargets in infrared images typically lack distinct texture and color features, posing significant challenges for infrared small target detection. Although existing methods have made some progress, they still have limitations in terms of interpretability and are unable to provide sufficient theoretical grounds to explain their effectiveness. Moreover, these methods generally overlook the temperature information in the imaging process, failing to fully exploit the subtle thermal differences in infrared images, which consequently limits detection accuracy. To this end, we propose a novel temperature-aware Transformer network (TATransNet). By introducing temperature information, it allows the model to be more focused on the target area. Specifically, inspired by the temperature inversion technique in remote sensing, we propose a learnable temperature inversion attention module (TIAM), which simulates the inverse process of infrared imaging to perceive the temperature distribution across feature maps of different scales, thereby enhancing the distinguishability between foreground and background. We further integrate TIAM with Swin Transformer to construct a Transformer-CNN hybrid encoder unit, termed temperature-aware Transformer block (TATB). By stacking multiple TATBs, both global context and local temperature features can be captured simultaneously, thereby strengthening the representation of dim and small targets effectively. Experimental results on five benchmark datasets, i.e., Small-ExtIRShip, Small-SSDD, NUAA-SIRST, IRSTD-1K, and IHAST demonstrate that our method not only significantly improves detection accuracy and reduces false alarm rates, but also demonstrates strong generalization capability, enabling effective transfer to other approaches with different types of encoders. Ting Zhang 0012, Zhaoying Liu, Bo Liu 0011, Changming Sun |
IJCNN | 5 |
| 2025 | Iterative clustering algorithm G-DESC-E and pan-cancer key gene analysis based on single-cell sequencing dataabstractSingle-cell sequencing technology has profoundly revolutionized the field of cancer genomics, enabling researchers to explore gene expression profiles at the resolution of individual cells. Despite its extensive applications in the study of cancer gene states, pan-cancer analyses remain relatively underexplored. In this study, we propose the G-DESC-E algorithm, which effectively distinguishes dimensionality-reduced data through a grid-based approach, filters out outliers during the preprocessing phase, and employs the Louvain algorithm for prescreening cluster centroids as initial clusters. We construct an objective function by integrating label entropy with the Kullback-Leibler divergence formula, achieving final clustering results through iterative optimization. Our findings demonstrate the effectiveness of the G-DESC-E algorithm in enhancing clustering accuracy. By applying our methodology to real-world datasets, we illustrate its capability to identify critical transcriptional features associated with distinct cancer subtypes. Coupled with clustering visualization and gene ontology analysis, we identify over thirty genes potentially related to cancer occurrence and progression. The algorithm and research framework presented in this study pave the way for new directions in clinical research by applying single-cell sequencing technology to the analysis of key genes within the realm of pan-cancer analysis for the first time. This approach offers valuable insights that can inform further clinical investigations. Ke Wu 0023, Changming Sun, Jie Geng 0001, Leyi Wei, Ran Su |
Briefings Bioinform. | 2 |
| 2025 | Synergizing multimodal data and fingerprint space exploration for mechanism of action predictionabstractMOTIVATION: Effective computational methods for predicting the mechanism of action (MoA) of compounds are essential in drug discovery. Current MoA prediction models mainly utilize the structural information of compounds. However, high-throughput screening technologies have generated more targeted cell perturbation data for MoA prediction, a factor frequently disregarded by the majority of current approaches. Moreover, exploring the commonalities and specificities among different fingerprint representations remains challenging. RESULTS: In this paper, we propose IFMoAP, a model integrating cell perturbation image and fingerprint data for MoA prediction. Firstly, we modify the Res-Net to accommodate the feature extraction of five-channel cell perturbation images and establish a granularity-level attention mechanism to combine coarse- and fine-grained features. To learn both common and specific fingerprint features, we introduce an FP-CS module, projecting four fingerprint embeddings into distinct spaces and incorporating two loss functions for effective learning. Finally, we construct two independent classifiers based on image and fingerprint features for prediction and for weighting the two prediction scores. Experimental results demonstrate that our model achieves highest accuracy of 0.941 when using multimodal data. The comparison with other methods and explorations further highlights the superiority of our proposed model and the complementary characteristics of multimodal data. AVAILABILITY AND IMPLEMENTATION: The source code is available at https://github.com/ s1mplehu/IFMoAP. The raw image data of Cell Painting can be accessed from Figshare (https://doi.org/10.17044/scilifelab.21378906). Kaimiao Hu, Jianguo Wei, Changming Sun, Jie Geng 0001, Leyi Wei, Ran Su |
Bioinform. | 3 |
| 2025 | Iterative pseudo-labeling based adaptive copy-paste supervision for semi-supervised tumor segmentation
Qiangguo Jin, Hui Cui 0002, Changming Sun, Yimiao He, Ping Xuan, Cong Cong 0001, Leyi Wei, Ran Su |
Knowl. Based Syst. | 4 |
| 2025 | Efficient transformer with compressed-attention for stereo image super-resolutionabstractWhile self-attention mechanisms in transformers exhibit superior performance in image super-resolution tasks, improving efficiency remains a challenge. To enhance the efficiency of self-attention mechanisms for stereo image super-resolution, we propose an efficient transformer with compressed attention for stereo image super-resolution (ETCASSR). Specifically, we propose a simple yet effective compressed-attention mechanism that organizes channels from partial to full for attention operations. Using this mechanism, we develop a compressed window-based self-attention block and a compressed transposed self-attention block, enabling efficient intra-view feature extraction. To further enrich feature representation, we introduce a spatial local feature branch and a channel global feature branch to complement these two blocks. Furthermore, a compressed cross-attention block for cross-view feature extraction is designed by extending the compressed-attention mechanism. Combining these blocks, ETCASSR achieves state-of-the-art performance on stereo image super-resolution while maintaining low computational complexity and fast running speed. Additionally, we introduce ETCASR for single-image super-resolution by omitting the cross-view components from ETCASSR, also achieving superior performance with high efficiency. The proposed transformers offer significant potential applications in other vision tasks. Source code will be released at https://github.com/jianwensong/ETCASSR . Jianwen Song, Arcot Sowmya, Changming Sun |
Knowl. Based Syst. | 4 |
| 2025 | Efficient frequency feature aggregation transformer for image super-resolutionabstractAlthough vision transformers have shown remarkable performance in image super-resolution tasks, the key component, i.e., the self-attention mechanism, suffers from insufficient high-frequency information extraction capability and high computational costs, hindering further advancement. To address these limitations, we propose an efficient frequency feature aggregation transformer for single image super-resolution (EFATSR). Specifically, a frequency self-attention aggregation block is proposed to enhance the extraction of high-frequency information. This block incorporates a frequency spatial feature aggregation branch to supplement high-frequency feature extraction for a self-attention branch, enabling the model to capture high-frequency information more effectively. Additionally, a frequency channel-spatial aggregation block is proposed to extract channel and spatial features in the frequency domain, enhancing the efficiency of deep feature extraction. Extensive experiments on single image super-resolution demonstrate that EFATSR achieves state-of-the-art performance while maintaining low computational complexity. Furthermore, we extend EFATSR for stereo image super-resolution by incorporating a multi-head parallax-attention block, forming EFATSSR, which also shows remarkable performance and high efficiency. Source code is avaliable at https://github.com/jianwensong/EFATSR . Jianwen Song, Arcot Sowmya, Changming Sun |
Pattern Recognit. | 3 |
| 2025 | An Image Terrain Map Model for Texture FilteringabstractThe purpose of texture measurement is to describe and quantify the texture features of pixels in an image. The accuracy of texture measurement plays a crucial role in determining the effectiveness of texture filtering. However, current texture measurement methods face challenges in achieving accurate texture measurement results, particularly for multi-scale texture measurements. This limitation often leads to unsatisfactory texture filtering results, particularly with image details and high-contrast textures. We find that when moving the texture measurement regions for pixels near texture edges further away from the texture edge and keeping the texture measurement regions for pixels far from texture edges unchanged results in an improved accuracy of texture measurement. Based on this observation, we propose a novel texture measurement approach that employs a circular neighborhood with a variable radius as the texture measurement region for each pixel. Furthermore, we proposed an image terrain map model based on a one-pixel texture edge to obtain optimal parameters for texture measurement regions. This model significantly enhances the accuracy of texture measurement at any scale in an image. The experimental results show that the texture filtering method based on our image terrain map model is significantly better than existing methods in terms of edge-preservation, small-structure preservation, and high-contrast texture filtering. Additionally, we presented some applications of the image terrain map model in other areas of image processing to demonstrate its versatility. Yiyao Fan, Jun Lin 0003, Changming Sun, Tianhao Wang 0009, Yuehan Qi, Yang Liu 0333 |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2025 | TPNET: A Time-Sensitive Small Sample Multimodal Network for Cardiotoxicity Risk PredictionabstractCancer therapy-related cardiac dysfunction (CTRCD) is a potential complication associated with cancer treatment, particularly in patients with breast cancer, requiring monitoring of cardiac health during the treatment process. Tissue Doppler imaging (TDI) is a remarkable technique that can provide a comprehensive reflection of the left ventricle's physiological status. We hypothesized that the combination of TDI features with deep learning techniques could be utilized to predict CTRCD. To evaluate the hypothesis, we developed a temporal-multimodal pattern network for efficient training (TPNET) model to predict the incidence of CTRCD over a 24-month period based on TDI, function, and clinical data from 270 patients. Our model achieved an area under curve (AUC) of 0.83 and sensitivity of 0.88, demonstrating greater robustness compared to other existing visual models. To further translate our model's findings into practical applications, we utilized the integrated gradients (IG) attribution to perform a detailed evaluation of all the features. This analysis has identified key pathogenic signs that may have remained unnoticed, providing a viable option for implementing our model in preoperative breast cancer patients. Additionally, our findings demonstrate the potential of TPNET in discovering new causative agents for CTRCD. Yuan He 0016, Fengyun Zhang, Kaimiao Hu, Changming Sun, Jie Geng 0001, Ning Ren, Ran Su |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | BFGTP: A BERT-Guided Two-Stage Molecular Representation Learning Framework for Toxicity PredictionabstractAccurate prediction of molecular toxicity is vital for drug development. Most mainstream methods rely on fingerprints or graph-based feature extraction, the emergence of large language models (LLMs) offers new prospects for molecular representation learning in toxicity prediction. Although several studies attempt to leverage LLMs to integrate molecular sequence data for pretraining molecular representations, certain limitations remain. Current LLM-based approaches usually utilize solely on class embedding features, overlooking the rich information in sequence embedding. Moreover, integrating pre-trained molecular representations with multi-modal molecular data may further enhance performance in toxicity prediction. To address these challenges, we propose BFGTP, a BERT-guided two-stage molecular representation learning framework for toxicity prediction. Firstly, we design independent encoders for molecular descriptions of three modalities, where the fingerprint encoder with dual level attention mechanisms effectively integrates multi-category fingerprints. Then, the two-stage guide strategy is introduced to fully utilize the prior knowledge of LLMs, employing contrastive learning to align and fuse the tri-modal representations and knowledge distillation to align predicted value distributions. BFGTP ultimately combines fingerprint and graph representations to predict molecular toxicity. Experiments on seven toxicity datasets show that BFGTP outperforms baselines, achieving the highest AUC on five datasets and the best average performance across five evaluation metrics. Ablation studies, t-SNE visualization and case study confirm the effectiveness of BFGTP's components and its ability to capture meaningful molecular representations. Kaimiao Hu, Yuan He 0016, Jianguo Wei, Changming Sun, Jie Geng 0001, Leyi Wei, Ran Su |
IEEE J. Biomed. Health Informatics | 4 |
| 2025 | Enhancing Infrared Small Target Detection: A Saliency-Guided Multi-Task Learning ApproachabstractObject detection in infrared images poses a considerable challenge due to its small-scale targets, low contrast and poor signal-to-clutter ratio, often resulting in a high false alarm rate. To improve the detection accuracy on infrared small targets, we introduce Light-SGMTLM, a lightweight and saliency-guided multi-task learning model. This model integrates saliency detection into the YOLOv5x framework through a parallel multi-task learning structure and employs a joint loss function during training. Such integration significantly alleviates the impact of complex backgrounds and improves the precision of small target localization. Moreover, we have developed a streamlined module, termed SIWD, to create a more agile backbone, which establishes an optimal balance between precision and efficiency, making the model more suitable for situations with limited computational resources. Comprehensive comparative experiments were conducted on six infrared small target datasets, namely, Small-ExtIRShip, Small-SSDD, IHAST, NUAA-SIRST, IRSTD-1k, and IRDST, and we assessed the model’s performance against ten leading target detection models, such as YOLOv7, YOLOv8, DINO, and Relation-DETR. The findings reveal that our method’s unique joint learning architecture, combining saliency and object detection tasks, significantly improves accuracy for infrared small target detection. Notably, it achieved impressive mean average precision (mAP) values of 92.60% and 75.71% on the NUAA-SIRST and IRSTD-1k datasets, respectively. Zhaoying Liu, Junran He, Ting Zhang 0012, Sadaqat ur Rehman, Mohammad Saraee, Changming Sun |
IEEE Trans. Intell. Transp. Syst. | 7 |
| 2024 | Location Embedding Based Pairwise Distance Learning for Fine-Grained Diagnosis of Urinary Stones
Qiangguo Jin, Jiapeng Huang, Changming Sun, Hui Cui 0002, Ping Xuan, Ran Su, Leyi Wei, Yu-Jie Wu, Chia-An Wu, Henry Been-Lirn Duh, Yueh-Hsun Lu |
MICCAI (11) | 3 |
| 2024 | Inter- and intra-uncertainty based feature aggregation model for semi-supervised histopathology image segmentation
Qiangguo Jin, Hui Cui 0002, Changming Sun, Jiangbin Zheng 0001, Leilei Cao, Leyi Wei, Ran Su |
Expert Syst. Appl. | 3 |
| 2024 | Few-Shot Stereo Matching with High Domain Adaptability Based on Adaptive Recursive Network
Rongcheng Wu, Zhidong Li, Jianlong Zhou, Fang Chen 0001, Changming Sun |
Int. J. Comput. Vis. | 7 |
| 2024 | Efficient masked feature and group attention network for stereo image super-resolutionabstractCurrent stereo image super-resolution methods do not fully exploit cross-view and intra-view information, resulting in limited performance. While vision transformers have shown great potential in super-resolution, their application in stereo image super-resolution is hindered by high computational demands and insufficient channel interaction. This paper introduces an efficient masked feature and group attention network for stereo image super-resolution (EMGSSR) designed to integrate the strengths of transformers into stereo super-resolution while addressing their inherent limitations. Specifically, an efficient masked feature block is proposed to extract local features from critical areas within images, guided by sparse masks. A group-weighted cross-attention module consisting of group-weighted cross-view feature interactions along epipolar lines is proposed to fully extract cross-view information from stereo images. Additionally, a group-weighted self-attention module consisting of group-weighted self-attention feature extractions with different local windows is proposed to effectively extract intra-view information from stereo images. Experimental results demonstrate that the proposed EMGSSR outperforms state-of-the-art methods at relatively low computational costs. The proposed EMGSSR offers a robust solution that effectively extracts cross-view and intra-view information for stereo image super-resolution, bringing a promising direction for future research in high-fidelity stereo image super-resolution. Source codes will be released at https://github.com/jianwensong/EMGSSR . Jianwen Song, Arcot Sowmya, Jien Kato, Changming Sun |
Image Vis. Comput. | 4 |
| 2024 | Re-abstraction and perturbing support pair network for few-shot fine-grained image classificationabstractThe goal of few-shot fine-grained image classification (FSFGIC) is to distinguish subordinate-level categories with subtle visual differences such as the species of bird and models of car with only a few samples. In this work, we argue that a designed network that has the ability to better distinguish feature descriptors of different categories will effectively improve the performance of FSFGIC. We propose a re-abstraction and perturbing support pair network (RaPSPNet) for FSFGIC. Specifically, we first design a feature re-abstraction embedding (FRaE) module which can not only effectively amplify the difference between the feature information from different categories but also better extract the feature information from images. Furthermore, a novel perturbing support pair (PSP) based similarity measure module is designed which evaluates the relationships of feature information among a query image and two different categories of support images (a support pair) at the same time for guiding the designed FRaE module to find salient feature information from the same category of query and support images and find distinguishable feature information from the different categories of query and support images. Extensive experiments on FSFGIC tasks demonstrate the superiority of the proposed methods over state-of-the-art benchmarks. Yali Zhao, Yongsheng Gao 0001, Changming Sun |
Pattern Recognit. | 4 |
| 2024 | Efficient Hybrid Feature Interaction Network for Stereo Image Super-ResolutionabstractIt is very challenging to fully use cross-view information for stereo image super-resolution. Previous methods using pixel-based parallax-attention mechanisms do not consider neighborhood pixels. Also, they typically use convolutions for basic feature extraction, which may not be as effective as modern self-attention mechanisms in transformers. To address these limitations, we propose an efficient hybrid feature interaction network for stereo image super-resolution. Specifically, we propose a shifted cross-view interaction block that integrates neighborhood pixels and imposes constraints on the disparity range during cross-view interactions. In addition, we propose a hybrid feature interaction block consisting of local and global interaction branches for extracting intra-view features efficiently. In this block, we propose a design that incorporates lightweight attention connections and a partial downsampling operation to enhance spatial and channel feature interaction with high efficiency. Additionally, a dilated efficient channel attention mechanism is proposed to obtain cross-channel interactions within features. Experimental results evaluated on various metrics (PSNR, SSIM, and LPIPS) demonstrate that the proposed method achieves state-of-the-art stereo image super-resolution performance at relatively low computational cost. Moreover, the super-resolution images obtained by the proposed method achieve the smallest stereo matching errors compared to other methods. Source code will be publicly available athttps://github.com/jianwensong/EHFSSR. Jianwen Song, Arcot Sowmya, Changming Sun |
IEEE Trans. Multim. | 3 |
| 2023 | Shape-aware contrastive deep supervision for esophageal tumor segmentation from CT scansabstractAccurate tumor segmentation is crucial for esophageal cancer radiotherapy treatment planning. The low contrast among the esophagus, tumors, and surrounding tissues, and irregular tumor shapes limit the performance of automatic segmentation methods. In this paper, we aim to exploit the irregular shapes of tumors to facilitate accurate segmentation. We propose a simple and pluggable shape-aware contrastive deep supervision network (SCDSNet) with shape-aware regularization and voxel-to-voxel contrastive deep supervision. Specifically, the shape-aware regularization with an uncertainty minimization strategy encourages the precise predictions of an additional shape-aware head. The voxel-to-voxel contrastive deep supervision enhances the multi-scale shape-tumor contrast for better voxel-to-voxel prediction of shapes. The proposed method is simple and highly pluggable, which can easily be extended to other frameworks. Further, we establish a large in-house dataset on esophageal cancer to validate the effectiveness of our proposed method. The quantitative and qualitative experimental results demonstrate the effectiveness of SCDSNet on the esophageal cancer dataset. Qiangguo Jin, Hui Cui 0002, Changming Sun, Jiapeng Huang, Ping Xuan, Yiyue Xu, Leilei Cao, Leyi Wei, Ran Su |
BIBM | 3 |
| 2023 | Learning Partial Correlation based Deep Visual Representation for Image ClassificationabstractVisual representation based on covariance matrix has demonstrates its efficacy for image classification by characterising the pairwise correlation of different channels in convolutional feature maps. However, pairwise correlation will become misleading once there is another channel correlating with both channels of interest, resulting in the “confounding” effect. For this case, “partial correlation” which removes the confounding effect shall be estimated instead. Nevertheless, reliably estimating partial correlation requires to solve a symmetric positive definite matrix optimisation, known as sparse inverse covariance estimation (SICE). How to incorporate this process into CNN remains an open issue. In this work, we formulate SICE as a novel structured layer of CNN. To ensure end-to-end trainability, we develop an iterative method to solve the above matrix optimisation during forward and backward propagation steps. Our work obtains a partial correlation based deep visual representation and mitigates the small sample problem often encountered by covariance matrix estimation in CNN. Computationally, our model can be effectively trained with GPU and works well with a large number of channels of advanced CNNs. Experiments show the efficacy and superior classification performance of our deep visual representation compared to covariance matrix based counterparts. Saimunur Rahman, Piotr Koniusz, Lei Wang 0001, Luping Zhou, Peyman Moghadam, Changming Sun |
CVPR | 6 |
| 2023 | Calibrating a Deep Neural Network with Its PredecessorsabstractConfidence calibration - the process to calibrate the output probability distribution of neural networks - is essential for safety-critical applications of such networks. Recent works verify the link between mis-calibration and overfitting. However, early stopping, as a well-known technique to mitigate overfitting, fails to calibrate networks. In this work, we study the limitions of early stopping and comprehensively analyze the overfitting problem of a network considering each individual block. We then propose a novel regularization method, predecessor combination search (PCS), to improve calibration by searching a combination of best-fitting block predecessors, where block predecessors are the corresponding network blocks with weight parameters from earlier training stages. PCS achieves the state-of-the-art calibration performance on multiple datasets and architectures. In addition, PCS improves model robustness under dataset distribution shift. Supplementary material and code are available at https://github.com/Linwei94/PCS Linwei Tao, Minjing Dong, Daochang Liu, Changming Sun, Chang Xu 0002 |
IJCAI | 4 |
| 2023 | Multi-modality Contrastive Learning for Sarcopenia Screening from Hip X-rays and Clinical Information
Qiangguo Jin, Changjiang Zou, Hui Cui 0002, Changming Sun, Shu-Wei Huang, Yi-Jie Kuo, Ping Xuan, Leilei Cao, Ran Su, Leyi Wei, Henry Been-Lirn Duh, Yu-Pin Chen |
MICCAI (6) | 4 |
| 2023 | Swin-ResUNet+: An edge enhancement module for road extraction from remote sensing images
Yingshan Jing, Ting Zhang 0012, Zhaoying Liu, Yuewu Hou, Changming Sun |
Comput. Vis. Image Underst. | 5 |
| 2023 | ECFRNet: Effective corner feature representations network for image corner detection
Junfeng Jing, Chao Liu 0033, Yongsheng Gao 0001, Changming Sun |
Expert Syst. Appl. | 5 |
| 2023 | Image Feature Information Extraction for Interest Point Detection: A Comprehensive ReviewabstractInterest point detection is one of the most fundamental and critical problems in computer vision and image processing. In this paper, we carry out a comprehensive review on image feature information (IFI) extraction techniques for interest point detection. To systematically introduce how the existing interest point detection methods extract IFI from an input image, we propose a taxonomy of the IFI extraction techniques for interest point detection. According to this taxonomy, we discuss different types of IFI extraction techniques for interest point detection. Furthermore, we identify the main unresolved issues related to the existing IFI extraction techniques for interest point detection and any interest point detection methods that have not been discussed before. The existing popular datasets and evaluation standards are provided and the performances for fifteen state-of-the-art approaches are evaluated and discussed. Moreover, future research directions on IFI extraction techniques for interest point detection are elaborated. Junfeng Jing, Tian Gao 0004, Yongsheng Gao 0001, Changming Sun |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | Image Intensity Variation Information for Interest Point DetectionabstractInterest point detection methods are gaining more attention and are widely applied in computer vision tasks such as image retrieval and 3D reconstruction. However, there still exist two main problems to be solved: (1) from the perspective of mathematical representations, the differences among edges, corners, and blobs have not been convincingly explained and the relationships among the amplitude response, scale factor, and filtering orientation for interest points have not been thoroughly explained; (2) the existing design mechanism for interest point detection does not show how to accurately obtain intensity variation information on corners and blobs. In this paper, the first- and second-order Gaussian directional derivative representations of a step edge, four common genres of corners, an anisotropic-type blob, and an isotropic-type blob are analyzed and derived. Multiple interest point characteristics are discovered. The characteristics for interest points that we obtained help us describe the differences among edges, corners, and blobs, explain why the existing interest point detection methods with multiple scales cannot properly obtain interest points from images, and present novel corner and blob detection methods. Extensive experiments demonstrate the superiority of our proposed methods in terms of detection performance, robustness to affine transformations, noise, image matching, and 3D reconstruction. Changming Sun, Yongsheng Gao 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Super-Resolution Phase Retrieval Network for Single-Pattern Structured Light 3D ImagingabstractStructured light 3D imaging is often used for obtaining accurate 3D information via phase retrieval. Single-pattern structured light 3D imaging is much faster than multi-pattern versions. Current phase retrieval methods for single-pattern structured light 3D imaging are however not accurate enough. Besides, the projector resolution in a structured light 3D imaging system is expensive to improve due to hardware costs. To address the issues of low accuracy and low resolution of single-pattern structured light 3D imaging, this work proposes a super-resolution phase retrieval network (SRPRNet). Specifically, a phase-shifting module is proposed to extract multi-scale features with different phase shifts, and a refinement and super-resolution module is proposed to obtain refined and super-resolution phase components. After phase demodulation and unwrapping, high-resolution absolute phase is obtained. A sine shifting loss and a cosine shifting loss are also introduced to form the regularization term of the loss function. As far as can be ascertained, the proposed SRPRNet is the first network for super-resolution phase retrieval by using a single pattern, and it can also be used for standard-resolution phase retrieval. Experimental results on three datasets show that SRPRNet achieves state-of-the-art performance on $1\times $ , $2\times $ , and $4\times $ super-resolution phase retrieval tasks. Jianwen Song, Kai Liu 0012, Arcot Sowmya, Changming Sun |
IEEE Trans. Image Process. | 4 |
| 2022 | Semi-supervised Histological Image Segmentation via Hierarchical Consistency Enforcement
Qiangguo Jin, Hui Cui 0002, Changming Sun, Jiangbin Zheng 0001, Leyi Wei, Zhenyu Fang, Zhaopeng Meng, Ran Su |
MICCAI (2) | 3 |
| 2022 | EOCSA: Predicting prognosis of Epithelial ovarian cancer with whole slide histopathological images
Tianling Liu, Ran Su, Changming Sun, Xiu-Ting Li, Leyi Wei |
Expert Syst. Appl. | 3 |
| 2022 | Recent advances on image edge detection: A comprehensive review
Junfeng Jing, Shenjuan Liu, Gang Wang 0031, Changming Sun |
Neurocomputing | 5 |
| 2022 | Complex shearlets and rotary phase congruence tensor for corner detectionabstractCorner detection algorithms based on multi-scale analysis attract more attention due to their promising performance. However, they only consider amplitude information, neglect phase information and partially utilize multi-scale decomposition coefficients to detect corners. This limits their detection accuracy, repeatability and localization ability. This paper describes a new multi-scale analysis based corner detector. To overcome the problems of bilateral margin responses, edge extension and lack of phase information in traditional shearlets, a novel complex shearlet transform is proposed to better localize distributed discontinuities and especially to extract phase information from geometrical features. Moreover, a new rotary phase congruence tensor is proposed to utilize all amplitude and phase information for corner detection. Its tolerances to noise and ability for corner localization are improved further by screening and normalizing the amplitude information. Experimental results demonstrate that the localization ability and detection accuracy of the proposed method are superior to current detectors, and its repeatability is generally higher than current detectors and recent machine learning based interest point detectors. Changming Sun, Arcot Sowmya |
Pattern Recognit. | 2 |
| 2022 | Improving Monocular Visual Odometry Using Learned DepthabstractMonocular visual odometry (VO) is an important task in robotics and computer vision. Thus far, how to build accurate and robust monocular VO systems that can work well in diverse scenarios remains largely unsolved. In this article, we propose a framework to exploit monocular depth estimation for improving VO. The core of our framework is a monocular depth estimation module with a strong generalization capability for diverse scenes. It consists of two separate working modes to assist the localization and mapping. With a single monocular image input, the depth estimation module predicts a relative depth to help the localization module on improving the accuracy. With a sparse depth map and an RGB image input, the depth estimation module can generate accurate scale-consistent depth for dense mapping. Compared with current learning-based VO methods, our method demonstrates a stronger generalization ability to diverse scenes. More significantly, our framework is able to boost the performances of existing geometry-based VO methods by a large margin. Libo Sun 0002, Wei Yin 0006, Enze Xie, Zhengrong Li, Changming Sun, Chunhua Shen |
IEEE Trans. Robotics | 5 |
| 2021 | A Novel Decision Mechanism for Image Edge Detection
Junfeng Jing, Shenjuan Liu, Chao Liu 0033, Tian Gao 0004, Changming Sun |
ICIC (1) | 6 |
| 2021 | Disentangled Representation Learning and Enhancement Network for Single Image De-RainingabstractIn this paper, we present a disentangled representation learning and enhancement network (DRLE-Net) to address the challenging single image de-raining problems, i.e., raindrop and rain streak removal. Specifically, the DRLE-Net is formulated as a multi-task learning framework, and an elegant knowledge transfer strategy is designed to train the encoder of DRLE-Net to embed a rainy image into two separated latent spaces representing the task (clean image reconstruction in this paper) relevant and irrelevant variations respectively, such that only the essential task-relevant factors will be used by the decoder of DRLE-Net to generate high-quality de-raining results. Furthermore, visual attention information is modeled and fed into the disentangled representation learning network to enhance the task-relevant factor learning. To facilitate the optimization of the hierarchical network, a new adversarial loss formulation is proposed and used together with the reconstruction loss to train the proposed DRLE-Net. Extensive experiments are carried out for removing raindrops or rainstreaks from both synthetic and real rainy images, and DRLE-Net is demonstrated to produce significantly better results than state-of-the-art models. Guoqing Wang 0001, Changming Sun, Xing Xu 0001, Jingjing Li 0001, Zheng Wang 0044, Zeyu Ma 0002 |
ACM Multimedia | 2 |
| 2021 | Domain adaptation based self-correction model for COVID-19 infection segmentation in CT images
Qiangguo Jin, Hui Cui 0002, Changming Sun, Zhaopeng Meng, Leyi Wei, Ran Su |
Expert Syst. Appl. | 3 |
| 2021 | Context-Enhanced Representation Learning for Single Image Deraining
Guoqing Wang 0001, Changming Sun, Arcot Sowmya |
Int. J. Comput. Vis. | 2 |
| 2021 | Free-form tumor synthesis in computed tomography images via richer generative adversarial network
Qiangguo Jin, Hui Cui 0002, Changming Sun, Zhaopeng Meng, Ran Su |
Knowl. Based Syst. | 3 |
| 2021 | Corner Detection Using Second-Order Generalized Gaussian Directional Derivative RepresentationsabstractCorner detection is a critical component of many image analysis and image understanding tasks, such as object recognition and image matching. Our research indicates that existing corner detection algorithms cannot properly depict the difference between edges and corners and this results in wrong corner detections. In this paper, the capability of second-order generalized (isotropic and anisotropic) Gaussian directional derivative filters to suppress Gaussian noise is evaluated. The second-order generalized Gaussian directional derivative representations of step edge, L-type corner, Y- or T-type corner, X-type corner, and star-type corner are investigated and obtained. A number of properties for edges and corners are discovered which enable us to propose a new image corner detection method. Finally, the criteria on detection accuracy and average repeatability under affine image transformation, JPEG compression, and noise degradation, and the criteria on region repeatability are used to evaluate the proposed detector against nine state-of-the-art methods. The experimental results show that our proposed detector outperforms all the other tested detectors. Changming Sun |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2021 | Attentive Feature Refinement Network for Single Rainy Image RestorationabstractDespite the fact that great progress has been made on single image deraining tasks, it is still challenging for existing models to produce satisfactory results directly, and it often requires a single or multiple refinement stages to gradually improve the quality. However, in this paper, we demonstrate that existing image-level refinement with a stage-independent learning design is problematic with the side effect of over/under-deraining. To resolve this issue, we for the first time propose the mechanism of learning to carry out refinement on the unsatisfactory features, and propose a novel attentive feature refinement (AFR) module. Specifically, AFR is designed as a two-branched network for simultaneous rain-distribution-aware attention map learning and attention guided hierarchy-preserving feature refinement. Guided by task-specific attention, coarse features are progressively refined to better model the diversified rainy effects. By using a separable convolution as the basic component, our AFR module introduces little computation overhead and can be readily integrated into most rainy-to-clean image translation networks for achieving better deraining results. By incorporating a series of AFR modules into a general encoder-decoder network, AFR-Net is constructed for deraining and it achieves new state-of-the-art results on both synthetic and real images. Furthermore, by using AFR-Net as a teacher model, we explore the use of knowledge distillation to successfully learn a student model that is also able to achieve state-of-the-art results but with a much faster inference speed (i.e., it only takes 0.08 second to process a 512×512 rainy image). Code and pre-trained models are available at 〈 https://github.com/RobinCSIRO/AFR-Net 〉 . Guoqing Wang 0001, Changming Sun, Arcot Sowmya |
IEEE Trans. Image Process. | 2 |
| 2021 | Robust Tensor Decomposition for Image Representation Based on Generalized CorrentropyabstractTraditional tensor decomposition methods, e.g., two dimensional principal component analysis and two dimensional singular value decomposition, that minimize mean square errors, are sensitive to outliers. To overcome this problem, in this paper we propose a new robust tensor decomposition method using generalized correntropy criterion (Corr-Tensor). A Lagrange multiplier method is used to effectively optimize the generalized correntropy objective function in an iterative manner. The Corr-Tensor can effectively improve the robustness of tensor decomposition with the existence of outliers without introducing any extra computational cost. Experimental results demonstrated that the proposed method significantly reduces the reconstruction error on face reconstruction and improves the accuracies on handwritten digit recognition and facial image clustering. Miaohua Zhang, Yongsheng Gao 0001, Changming Sun, Michael Blumenstein |
IEEE Trans. Image Process. | 3 |
| 2020 | Instance-Aware Embedding for Point Cloud Instance Segmentation
Tong He 0001, Yifan Liu 0001, Chunhua Shen, Changming Sun |
ECCV (30) | 5 |
| 2020 | ReDro: Efficiently Learning Large-Sized SPD Visual Representation
Saimunur Rahman, Lei Wang 0001, Changming Sun, Luping Zhou |
ECCV (15) | 3 |
| 2020 | Corner Detection Using Multi-directional Structure Tensor with Multiple Scales
Changming Sun |
Int. J. Comput. Vis. | 2 |
| 2020 | Fusing convolutional neural network features with hand-crafted features for osteoporosis diagnoses
Ran Su, Tianling Liu, Changming Sun, Qiangguo Jin, Rachid Jennane, Leyi Wei |
Neurocomputing | 3 |
| 2020 | Deep learning based HEp-2 image classification: A comprehensive review
Saimunur Rahman, Lei Wang 0001, Changming Sun, Luping Zhou |
Medical Image Anal. | 3 |
| 2020 | Corner detection based on shearlet transform and multi-directional structure tensor
Changming Sun, Arcot Sowmya |
Pattern Recognit. | 3 |
| 2020 | A robust matching pursuit algorithm using information theoretic learning
Miaohua Zhang, Yongsheng Gao 0001, Changming Sun, Michael Blumenstein |
Pattern Recognit. | 3 |
| 2020 | Multi-Weighted Co-Occurrence Descriptor Encoding for Vein RecognitionabstractDespite being highly secure, vein recognition suffers from the high inter-class similarity and intra-class variation resulting from the uncontrolled image capture, making the design of discriminative and robust representation very important. The recent success of convolutional neural network (CNN) for various image understanding tasks makes it a promising method for feature extraction. However, limited variability in small-scale datasets leads to systems derived from the direct training or fine-tuning not transferable and unreliable for practical biometric applications. This motivates the design of a multi-weighted co-occurrence descriptor encoding (MWCDE) model for vein recognition. Instead of directly conducting a feed-forward operation with a pre-trained CNN for obtaining the semantic features from the fully connected layers, co-occurrence features among convolutional filters are modeled first in MWCDE by a simple convolution between an indicator filter in a higher layer with a to-be-reweighted filter in a lower layer, and a redundancy-driven indicator filter selection algorithm is designed for filtering out some ambiguous representations. Second, another hard feature weighting strategy with a binary masking scheme is proposed for discarding noisy background and feature redundancy. The selected high-order descriptors are then embedded and aggregated into the compact feature vectors with a saliency driven spatial weighted Fisher vector algorithm, followed by the introduction of a generalized support vector machine for recognition. Extensive experiments with three benchmark vein datasets demonstrate that the proposed framework can achieve state-of-the-art results, and an additional experiment with the PolyU multispectral palmprint database illustrates its generalization ability. Code is available at (https://github.com/RobinCSIRO/MWCDE-for-Vein-Recognition). Guoqing Wang 0001, Changming Sun, Arcot Sowmya |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2020 | Learning a Compact Vein Discrimination Model With GANerated SamplesabstractDespite the great success achieved by convolutional neural networks (CNNs) in various image understanding tasks, it is still difficult for CNNs to be applied to vein recognition tasks due to the problems of insufficient training datasets, intra-class variations, and inter-class similarities. Besides, due to the essential requirement on the storage of millions of parameters for CNN, it is challenging to use a CNN for designing a vein-based embedded person identification system. In this paper, these two problems are addressed by learning a discriminative and compact vein recognition model. For the first problem, a hierarchical generative adversarial network (HGAN) consisting of a constrained CNN and a CycleGAN is proposed for data augmentation. Two similarity losses are defined for estimating the self-similarity and inter-class dissimilarity, and a CycleGAN model is properly trained with these two losses for better task-specific training sample generation. After obtaining a baseline vein recognition model fine-tuned on the augmented datasets, the existence of parameter redundancy in the over-parameterized network motivates the proposal of model compression by way of filter pruning and low rank approximation, thus making the compressed model more suitable for deployment on embedded systems. Through the vein recognition experiments with two different datasets and an additional palmprint recognition experiment, the proposed algorithms are shown to yield a highly compact model while keeping the accuracy acceptable for application. Guoqing Wang 0001, Changming Sun, Arcot Sowmya |
IEEE Trans. Inf. Forensics Secur. | 2 |
| 2020 | Cascaded Attention Guidance Network for Single Rainy Image RestorationabstractRestoring a rainy image with raindrops or rainstreaks of varying scales, directions, and densities is an extremely challenging task. Recent approaches attempt to leverage the rain distribution (e.g., location) as prior to generate satisfactory results. However, concatenation of a single distribution map with the rainy image or with intermediate feature maps is too simplistic to fully exploit the advantages of such priors. To further explore this valuable information, an advanced cascaded attention guidance network, dubbed as CAG-Net, is formulated and designed as a three-stage model. In the first stage, a multitask learning network is constructed for producing the attention map and coarse de-raining results simultaneously. Subsequently, the coarse results and the rain distribution map are concatenated and fed to the second stage for results refinement. In this stage, the attention map generation network from the first stage is used to formulate a novel semantic consistency loss for better detail recovery. In the third stage, a novel pyramidal "whereand- how" learning mechanism is formulated. At each pyramid level, a two-branch network is designed to take the features from previous stages as inputs to generate better attention-guidance features and de-raining features, which are then combined via a gating scheme to produce the final de-raining results. Moreover, the uncertainty maps are also generated in this stage for more accurate pixel-wise loss calculation. Extensive experiments are carried out for removing raindrops or rainstreaks from both synthetic and real rainy images, and CAG-Net is demonstrated to produce significantly better results than state-of-the-art models. Code will be publicly available after paper acceptance. Guoqing Wang 0001, Changming Sun, Arcot Sowmya |
IEEE Trans. Image Process. | 2 |
| 2019 | Knowledge Adaptation for Efficient Semantic SegmentationabstractBoth accuracy and efficiency are of significant importance to the task of semantic segmentation. Existing deep FCNs suffer from heavy computations due to a series of high-resolution feature maps for preserving the detailed knowledge in dense estimation. Although reducing the feature map resolution (i.e., applying a large overall stride) via subsampling operations (e.g., polling and convolution striding) can instantly increase the efficiency, it dramatically decreases the estimation accuracy. To tackle this dilemma, we propose a knowledge distillation method tailored for semantic segmentation to improve the performance of the compact FCNs with large overall stride. To handle the inconsistency between the features of the student and teacher network, we optimize the feature similarity in a transferred latent domain formulated by utilizing a pre-trained autoencoder. Moreover, an affinity distillation module is proposed to capture the long-range dependency by calculating the non local interactions across the whole image. To validate the effectiveness of our proposed method, extensive experiments have been conducted on three popular benchmarks: Pascal VOC, Cityscapes and Pascal Context. Built upon a highly competitive baseline, our proposed method can improve the performance of a student network by 2.5% (mIOU boosts from 70.2 to 72.7 on the cityscapes test set) and can train a better compact model with only 8% float operations (FLOPS) of a model that achieves comparable performances. Tong He 0001, Chunhua Shen, Zhi Tian, Dong Gong, Changming Sun, Youliang Yan |
CVPR | 5 |
| 2019 | ERL-Net: Entangled Representation Learning for Single Image De-RainingabstractDespite the significant progress achieved in image de-raining by training an encoder-decoder network within the image-to-image translation formulation, blurry results with missing details indicate the deficiency of the existing models. By interpreting the de-raining encoder-decoder network as a conditional generator, within which the decoder acts as a generator conditioned on the embedding learned by the encoder, the unsatisfactory output can be attributed to the low-quality embedding learned by the encoder. In this paper, we hypothesize that there exists an inherent mapping between the low-quality embedding to a latent optimal one, with which the generator (decoder) can produce much better results. To improve the de-raining results significantly over existing models, we propose to learn this mapping by formulating a residual learning branch, that is capable of adaptively adding residuals to the original low-quality embedding in a representation entanglement manner. Using an embedding learned this way, the decoder is able to generate much more satisfactory de-raining results with better detail recovery and rain artefacts removal, providing new state-of-the-art results on four benchmark datasets with considerable improvement (i.e., on the challenging Rain100H data, an improvement of 4.19dB on PSNR and 5% on SSIM is obtained). The entanglement can be easily adopted into any encoder-decoder based image restoration networks. Besides, we propose a series of evaluation metrics to investigate the specific contribution of the proposed entangled representation learning mechanism. Codes are available at 〈https://github.com/RobinCSIRO/ERL-Net-for-Single-Image-Deraining〉. Guoqing Wang 0001, Changming Sun, Arcot Sowmya |
ICCV | 2 |
| 2019 | Robust Sparse Learning Based on Kernel Non-Second Order MinimizationabstractPartial occlusions in face images pose a great problem for most face recognition algorithms due to the fact that most of these algorithms mainly focus on solving a second order loss function, e.g., mean square error (MSE), which will magnify the effect from occlusion parts. In this paper, we proposed a kernel non-second order loss function for sparse representation (KNS-SR) to recognize or restore partially occluded facial images, which both take the advantages of the correntropy and the non-second order statistics measurement. The resulted framework is more accurate than the MSE-based ones in locating and eliminating outliers information. Experimental results from image reconstruction and recognition tasks on publicly available databases show that the proposed method achieves better performances compared with existing methods. Miaohua Zhang, Yongsheng Gao 0001, Changming Sun, Michael Blumenstein |
ICIP | 3 |
| 2019 | Kernel Mean P Power Error Loss for Robust Two-Dimensional Singular Value DecompositionabstractTraditional matrix-based dimensional reduction methods, e.g., two-dimensional principal component analysis (2DPCA) and two-dimensional singular value decomposition (2DSVD), minimize mean square errors (MSE), which is sensitive to outliers. To overcome this problem, in this paper we propose a new robust 2DSVD method based on the kernel mean p power error loss (KMPE-2DSVD). Different from the MSE and the correntropy based ones which are second order statistics based measurements, the KMPE-2DSVD is based on the non-second order statistics in the kernel space, and thus is more flexible in controlling the representation error. Experimental results show that the proposed method significantly improves the accuracy of facial image clustering. Miaohua Zhang, Yongsheng Gao 0001, Changming Sun, Michael Blumenstein |
ICIP | 3 |
| 2019 | Chord Bunch Walks for Recognizing Naturally Self-Overlapped and Compound LeavesabstractEffectively describing and recognizing leaf shapes under arbitrary variations, particularly from a large database, remains an unsolved problem. In this research, we attempted a new strategy of describing leaf shapes by walking and measuring along a bunch of chords that pass through the shape. A novel chord bunch walks (CBW) descriptor is developed through the chord walking behavior that effectively integrates the shape image function over the walked chord to reflect both the contour features and the inner properties of the shape. For each contour point, the chord bunch groups multiple pairs of chords to build a hierarchical framework for a coarse-to-fine description that can effectively characterize not only the subtle differences among leaf margin patterns but also the interior part of the shape contour formed inside a self-overlapped or compound leaf. Instead of using optimal correspondence based matching, a Log-Min distance that encourages one-to-one correspondences is proposed for efficient and effective CBW matching. The proposed CBW shape analysis method is invariant to rotation, scaling, translation, and mirror transforms. Five experiments, including image retrieval of compound leaves, image retrieval of naturally self-overlapped leaves, and retrieval of mixed leaves on three large scale datasets, are conducted. The proposed method achieved large accuracy increases with low computational costs over the state-of-the-art benchmarks, which indicates the research potential along this direction. Bin Wang 0041, Yongsheng Gao 0001, Changming Sun, Michael Blumenstein, John La Salle |
IEEE Trans. Image Process. | 3 |
| 2019 | Discrete Curvature Representations for Noise Robust Image Corner DetectionabstractImage corner detection is very important in the fields of image analysis and computer vision. Curvature calculation techniques are used in many contour-based corner detectors. We identify that existing calculation of curvature is sensitive to local variation and noise in the discrete domain and does not perform well when corners are closely located. In this paper, discrete curvature representations of single and double corner models are investigated and obtained. A number of model properties have been discovered, which help us detect corners on contours. It is shown that the proposed method has a high corner resolution (the ability to accurately detect neighboring corners), and a corresponding corner resolution constant is also derived. Meanwhile, this method is less sensitive to any local variations and noise on the contour; and false corner detection is less likely to occur. The proposed detector is compared with seven state-of-the-art detectors. Three test images with ground truths are used to assess the detection capability and localization accuracy of these methods in cases with noise-free and different noise levels; 24 images with various scenes without ground truths are used to evaluate their repeatability under affine transformation, JPEG compression, and noise degradations. The experimental results show that our proposed detector attains a better overall performance. Changming Sun, Toby P. Breckon, Naif Alshammari |
IEEE Trans. Image Process. | 2 |
| 2019 | Cell Segmentation Based on FOPSO Combined With Shape Information Improved Intuitionistic FCMabstractFuzzy c-means (FCM) clustering algorithms have been proved to be effective image segmentation techniques. However, FCM clustering algorithms are sensitive to noises and initialization. They cannot effectively segment cell images with inhomogeneous gray value distributions and complex touching cells. Aiming to overcome these disadvantages, this paper proposes a cell image segmentation algorithm using fractional-order velocity based particle swarm optimization (FOPSO) combined with shape information improved intuitionistic FCM (SI-IFCM) clustering. Iterations are carried out between FOPSO and SI-IFCM to achieve final cell segmentation. Experimental results demonstrate that the proposed algorithm has advantages on cell image segmentation, with the highest recall (90.25%) and lowest false discovery rate (0.28%) compared with the state-of-the-art algorithms. Xiangzhi Bai, Chuxiong Sun, Changming Sun |
IEEE J. Biomed. Health Informatics | 3 |
| 2019 | Design of a Clinical Decision Support System for Predicting Erectile Dysfunction in Men Using NHIRD DatasetabstractErectile dysfunction (ED) affects millions of men worldwide. Men with ED generally complain failure to attain or maintain an adequate erection during sexual activity. The prevalence of ED is strongly correlated with age, affecting about 40% of men at age 40 and nearly 70% at age 70. A variety of chronic diseases, including diabetes, ischemic heart disease, congestive heart failure, hypertension, depression, chronic renal failure, obstructive sleep apnea, prostate disease, gout, and sleep disorder, were reported to be associated with ED. In this study, data retrieved from a subset of the National Health Insurance Research Database of Taiwan were used for designing the clinical decision support system (CDSS) for predicting ED incidences in men. The positive cases were male patients aged 20-65 who were diagnosed with ED between January 2000 and December 2010 confirmed by at least three outpatient visits or at least one inpatient visit, while the negative cases were randomly selected from the database without a history of ED and were frequency (1:1), age, and index year matched with the ED patients. Data of a total of 2832 ED patients and 2832 non-ED patients, each consisting of 41 features including index age, 10 comorbidities, and 30 other comorbidity-related variables, were retrieved for designing the predictive models. Integrated genetic algorithm and support vector machine was adopted to design the CDSSs with two experiments of independent training and testing (ITT) conducted to verify their effectiveness. In the 1st ITT experiment, data extracted from January 2000 till December 2005 (61.51%, 1742 positive cases and 1742 negative cases) were used for training and validating and the data retrieved from January 2006 till December 2010 were used for testing (38.49%), whereas in the 2nd ITT experiment, data in the training set (77.78%) were extracted from January 2000 till Deceber 2007 and those in the testing set (22.22%) were retrieved afterward. Tenfold cross validation and three different objective functions were adopted for obtaining the optimal models with best predictive performance in the training phase. The testing results show that the CDSSs achieved a predictive performance with accuracy, sensitivity, specificity, g-mean, and area under ROC curve of 74.72%-76.65%, 72.33%-83.76%, 69.54%-77.10%, 0.7468-0.7632, and 0.766-0.817, respectively. In conclusion, the CDSSs designed based on cost-sensitive objective functions as well as salient comorbidity-related features achieve satisfactory predictive performance for predicting ED incidences. Yung-fu Chen, Chih-Sheng Lin, Chun-Fu Hong, Dah-Jye Lee, Changming Sun, Hsuan-Hung Lin |
IEEE J. Biomed. Health Informatics | 5 |
| 2018 | An End-to-End TextSpotter With Explicit Alignment and AttentionabstractText detection and recognition in natural images have long been considered as two separate tasks that are processed sequentially. Jointly training two tasks is non-trivial due to significant differences in learning difficulties and convergence rates. In this work, we present a conceptually simple yet efficient framework that simultaneously processes the two tasks in a united framework. Our main contributions are three-fold: (1) we propose a novel text-alignment layer that allows it to precisely compute convolutional features of a text instance in arbitrary orientation, which is the key to boost the performance; (2) a character attention mechanism is introduced by using character spatial information as explicit supervision, leading to large improvements in recognition; (3) two technologies, together with a new RNN branch for word recognition, are integrated seamlessly into a single model which is end-to-end trainable. This allows the two tasks to work collaboratively by sharing convolutional features, which is critical to identify challenging text instances. Our model obtains impressive results in end-to-end recognition on the ICDAR 2015 [19], significantly advancing the most recent results [2], with improvements of F-measure from (0.54, 0.51, 0.47) to (0.82, 0.77, 0.63), by using a strong, weak and generic lexicon respectively. Thanks to joint training, our method can also serve as a good detector by achieving a new state-of-the-art detection performance on related benchmarks. Code is available at https://github.com/tonghe90/textspotter. Tong He 0001, Zhi Tian, Chunhua Shen, Yu Qiao 0001, Changming Sun |
CVPR | 6 |
| 2018 | Matching Pursuit Based on Kernel Non-Second Order MinimizationabstractThe orthogonal matching pursuit (OMP) is an important sparse approximation algorithm to recover sparse signals from compressed measurements. However, most MP algorithms are based on the mean square error(MSE) to minimize the recovery error, which is suboptimal when there are outliers. In this paper, we present a new robust OMP algorithm based on kernel non-second order statistics (KNS-OMP), which not only takes advantages of the outlier resistance ability of correntropy but also further extends the second order statistics based correntropy to a non-second order similarity measurement to improve its robustness. The resulted framework is more accurate than the second order ones in reducing the effect of outliers. Experimental results on synthetic and real data show that the proposed method achieves better performances compared with existing methods. Miaohua Zhang, Yongsheng Gao 0001, Changming Sun, Michael Blumenstein |
ICIP | 3 |
| 2017 | Can Walking and Measuring Along Chord Bunches Better Describe Leaf Shapes?abstractEffectively describing and recognizing leaf shapes under arbitrary deformations, particularly from a large database, remains an unsolved problem. In this research, we attempted a new strategy of describing shape by walking along a bunch of chords that pass through the shape to measure the regions trespassed. A novel chord bunch walks (CBW) descriptor is developed through the chord walking that effectively integrates the shape image function over the walked chord to reflect the contour features and the inner properties of the shape. For each contour point, the chord bunch groups multiple pairs of chord walks to build a hierarchical framework for a coarse-to-fine description. The proposed CBW descriptor is invariant to rotation, scaling, translation, and mirror transforms. Instead of using the expensive optimal correspondence based matching, an improved Hausdorff distance encoded correspondence information is proposed for efficient yet effective shape matching. In experimental studies, the proposed method obtained substantially higher accuracies with low computational cost over the benchmarks, which indicates the research potential along this direction. Bin Wang 0041, Yongsheng Gao 0001, Changming Sun, Michael Blumenstein, John La Salle |
CVPR | 3 |
| 2017 | Image dehazing using adaptive bi-channel priors on superpixels
Changming Sun, Yu Zhao 0004, Li Yang 0003 |
Comput. Vis. Image Underst. | 2 |
| 2017 | Fog Density Estimation and Image Defogging Based on Surrogate Modeling for Optical DepthabstractIn order to estimate fog density correctly and to remove fog from foggy images appropriately, a surrogate model for optical depth is presented in this paper. We comprehensively investigate various fog-relevant features and propose a novel feature based on the hue, saturation, and value color space, which correlate well with the perception of fog density. We use a surrogate-based method to learn a refined polynomial regression model for optical depth with informative fog-relevant features, such as dark-channel, saturation-value, and chroma, which are selected on the basis of sensitivity analysis. Based on the obtained accurate surrogate model for optical depth, an effective method for fog density estimation and image defogging is proposed. The effectiveness of our proposed method is verified quantitatively and qualitatively by the experimental results on both synthetic and real-world foggy images. Changming Sun, Yu Zhao 0004, Li Yang 0003 |
IEEE Trans. Image Process. | 2 |
| 2016 | Robust tensor factorization using maximum correntropy criterionabstractTraditional tensor decomposition methods, e.g., two dimensional principle component analysis (2DPCA) and two dimensional singular value decomposition (2DSVD), minimize mean square errors (MSE) and are sensitive to outliers. In this paper, we propose a new robust tensor factorization method using maximum correntropy criterion (MCC) to improve the robustness of traditional tensor decomposition methods. A half-quadratic optimization algorithm is adopted to effectively optimize the correntropy objective function in an iterative manner. It can effectively improve the robustness of a tensor decomposition method to outliers without introducing any extra computational cost. Experimental results demonstrated that the proposed method significantly reduces the reconstruction error on face reconstruction and improves the accuracy rate on handwritten digit recognition. Miaohua Zhang, Yongsheng Gao 0001, Changming Sun, John La Salle, Junli Liang |
ICPR | 3 |
| 2016 | Orientation-guided geodesic weighting for PatchMatch-based stereo matching
Changming Sun, Xiao Tan 0001, Li Yang 0003 |
Inf. Sci. | 2 |
| 2016 | Infrared ship target segmentation through integration of multiple feature maps
Zhaoying Liu, Xiangzhi Bai, Changming Sun, Fugen Zhou |
Image Vis. Comput. | 3 |
| 2016 | Guided image completion by confidence propagation
Xiao Tan 0001, Changming Sun, Kwan-Yee Kenneth Wong, Tuan D. Pham |
Pattern Recognit. | 2 |
| 2016 | Stereo matching based on multi-direction polynomial model
Xiao Tan 0001, Changming Sun, Tuan D. Pham |
Signal Process. Image Commun. | 2 |
| 2016 | Edge-Aware Filtering with Local Polynomial Approximation and Rectangle-Based WeightingabstractThis paper presents a novel method for performing guided image filtering using local polynomial approximation (LPA) with range guidance. In our method, the LPA is introduced into a multipoint framework for reliable model regression and better preservation on image spatial variation which usually contains the essential information in the input image. In addition, we develop a weighting scheme which has the spatial flexibility during the filtering process. All components in our method are efficiently implemented and a constant computation complexity is achieved. Compared with conventional filtering methods, our method provides clearer boundaries and performs especially better in recovering spatial variation from noisy images. We conduct a number of experiments for different applications: depth image upsampling, joint image denoising, details enhancement, and image abstraction. Both quantitative and qualitative comparisons demonstrate that our method outperforms state-of-the-art methods. Xiao Tan 0001, Changming Sun, Tuan D. Pham |
IEEE Trans. Cybern. | 2 |
| 2015 | Feature matching in stereo images encouraging uniform spatial distribution
Xiao Tan 0001, Changming Sun, Xavier Sirault, Robert Furbank, Tuan D. Pham |
Pattern Recognit. | 2 |
| 2014 | Multipoint Filtering with Local Polynomial Approximation and Range GuidanceabstractThis paper presents a novel guided image filtering method using multipoint local polynomial approximation (LPA) with range guidance. In our method, the LPA is extended from a pointwise model into a multipoint model for reliable filtering and better preserving image spatial variation which usually contains the essential information in the input image. In addition, we develop a scheme with constant computational complexity (invariant to the size of filtering kernel) for generating a spatial adaptive support region around a point. By using the hybrid of the local polynomial model and color/intensity based range guidance, the proposed method not only preserves edges but also does a much better job in preserving spatial variation than existing popular filtering methods. Our method proves to be effective in a number of applications: depth image upsampling, joint image denoising, details enhancement, and image abstraction. Experimental results show that our method produces better results than state-of-the-art methods and it is also computationally efficient. Xiao Tan 0001, Changming Sun, Tuan D. Pham |
CVPR | 2 |
| 2014 | Soft Cost Aggregation with Multi-resolution Fusion
Xiao Tan 0001, Changming Sun, Dadong Wang, Yi Guo 0001, Tuan D. Pham |
ECCV (5) | 2 |
| 2014 | Iterative infrared ship target segmentation based on multiple features
Zhaoying Liu, Fugen Zhou, Xiangzhi Bai, Changming Sun |
Pattern Recognit. | 5 |
| 2014 | A new method for linear feature and junction enhancement in 2D images based on morphological operation, oriented anisotropic Gaussian function and Hessian information
Ran Su, Changming Sun, Chao Zhang 0011, Tuan D. Pham |
Pattern Recognit. | 2 |
| 2014 | Stereo matching using cost volume watershed and region merging
Xiao Tan 0001, Changming Sun, Xavier Sirault, Robert Furbank, Tuan D. Pham |
Signal Process. Image Commun. | 2 |
| 2014 | Evaluation and Comparison of Current Fetal Ultrasound Image Segmentation Methods for Biometric Measurements: A Grand ChallengeabstractThis paper presents the evaluation results of the methods submitted to Challenge US: Biometric Measurements from Fetal Ultrasound Images, a segmentation challenge held at the IEEE International Symposium on Biomedical Imaging 2012. The challenge was set to compare and evaluate current fetal ultrasound image segmentation methods. It consisted of automatically segmenting fetal anatomical structures to measure standard obstetric biometric parameters, from 2D fetal ultrasound images taken on fetuses at different gestational ages (21 weeks, 28 weeks, and 33 weeks) and with varying image quality to reflect data encountered in real clinical environments. Four independent sub-challenges were proposed, according to the objects of interest measured in clinical practice: abdomen, head, femur, and whole fetus. Five teams participated in the head sub-challenge and two teams in the femur sub-challenge, including one team who tackled both. Nobody attempted the abdomen and whole fetus sub-challenges. The challenge goals were two-fold and the participants were asked to submit the segmentation results as well as the measurements derived from the segmented objects. Extensive quantitative (region-based, distance-based, and Bland-Altman measurements) and qualitative evaluation was performed to compare the results from a representative selection of current methods submitted to the challenge. Several experts (three for the head sub-challenge and two for the femur sub-challenge), with different degrees of expertise, manually delineated the objects of interest to define the ground truth used within the evaluation framework. For the head sub-challenge, several groups produced results that could be potentially used in clinical settings, with comparable performance to manual delineations. The femur sub-challenge had inferior performance to the head sub-challenge due to the fact that it is a harder segmentation problem and that the techniques presented relied more on the femur's appearance. Sylvia Rueda, Sana Fathima, Caroline L. Knight, Mohammad Yaqub, Aris T. Papageorghiou, Bahbibi Rahmatullah, Alessandro Foi, Matteo Maggioni, Antonietta Pepe, Jussi Tohka, Richard V. Stebbing, John McManigle, Anca Ciurte, Xavier Bresson, Meritxell Bach Cuadra, Changming Sun, Gennady V. Ponomarev, Mikhail S. Gelfand, Marat D. Kazanov, Ching-Wei Wang, Hsiang-Chou Chen, Chun-Wei Peng, Chu-Mei Hung, J. Alison Noble |
IEEE Trans. Medical Imaging | 16 |
| 2013 | Trinocular stereo image rectification in closed-form only using fundamental matricesabstractTrinocular stereo image rectification is a process to transform a set of three images into a new set so that the epipolar lines in the transformed images have the same direction as the image row or column and matching epipolar lines in different images have the same row or column indices to enable an efficient and reliable dense stereo matching. In this paper we propose a new closed-form method to rectify three-view stereo images just using fundamental matrices. The algorithm involves only direct and purely geometric transformation processes. No iteration or optimization process is involved in our method. Real images have been used for testing purposes, and accurate results have been obtained using our algorithm. Changming Sun |
ICIP | 1 |
| 2012 | Cross Image Inference Scheme for Stereo Matching
Xiao Tan 0001, Changming Sun, Xavier Sirault, Robert Furbank, Tuan D. Pham |
ACCV (4) | 2 |
| 2012 | Embedded Voxel Colouring with Adaptive Threshold Selection Using Globally Minimal Surfaces
Carlos Leung, Ben Appleton, Mitchell Buckley, Changming Sun |
Int. J. Comput. Vis. | 4 |
| 2012 | Junction detection for linear structures based on Hessian, correlation and shape information
Ran Su, Changming Sun, Tuan D. Pham |
Pattern Recognit. | 2 |
| 2010 | Automated Feature Weighting in Fuzzy Declustering-based Vector QuantizationabstractFeature weighting plays an important role in improving the performance of clustering technique. We propose an automated feature weighting in fuzzy declustering-based vector quantization (FDVQ), namely AFDVQ algorithm, for enhancing effectiveness and efficiency in classification. The proposed AFDVQ imposes weights on the modified fuzzy c-means (FCM) so that it can automatically calculate feature weights based on their degrees of importance rather than treating them equally. Moreover, the extension of FDVQ and AFDVQ algorithms based on generalized improved fuzzy partitions (GIFP), known as GIFP-FDVQ and GIFP-AFDVQ respectively, are proposed. The experimental results on real data (original and noisy data) and modified data (biased and noisy-biased data) have demonstrated that the proposed algorithms outperformed standard algorithms in classifying clusters especially for biased data. Theam Foo Ng, Tuan D. Pham, Changming Sun |
ICPR | 3 |
| 2010 | Automated Detection of Nucleoplasmic Bridges for DNA Damage Scoring in Binucleated CellsabstractQuantification of DNA damage, which may be caused by radiation or exposure to chemicals, is very important and can be very time consuming and subject to variability if carried out visually. The quantification of scoring DNA damage includes biomarkers such as micronuclei, nucleoplasmic bridges, and nuclear buds as scored in cytokinesis-blocked binucleated cells. In this paper, we present a new algorithm based on a shortest path technique that enables us to detect the nucleoplasmic bridges joining two nuclei in cell images of binucleated cells. The effectiveness of our algorithm is illustrated using a set of cell images. We believe that this is the first time that a feasible automated nucleoplasmic bridge detection system has been reported. Changming Sun, Pascal Vallotton, Michael Fenech, Philip Thomas |
ICPR | 1 |
| 2009 | Splitting touching cells based on concave points and ellipse fitting
Xiangzhi Bai, Changming Sun, Fugen Zhou |
Pattern Recognit. | 2 |
| 2009 | Membrane boundary extraction using circular multiple paths
Changming Sun, Pascal Vallotton, Dadong Wang, Jamie Lopez, Yvonne Ng, David E. James |
Pattern Recognit. | 1 |
| 2008 | Iterated dynamic programming and quadtree subregioning for fast stereo matching
Carlos Leung, Ben Appleton, Changming Sun |
Image Vis. Comput. | 3 |
| 2008 | Boundary extraction of linear features using dual paths through gradient profiles
Ryan Lagerstrom, Changming Sun, Pascal Vallotton |
Pattern Recognit. Lett. | 2 |
| 2007 | Thickness measurement and crease detection of wheat grains using stereo vision
Changming Sun, Mark Berman 0002, David Coward, Brian Osborne |
Pattern Recognit. Lett. | 1 |
| 2006 | Moving average algorithms for diamond, hexagon, and general polygonal shaped window operations
Changming Sun |
Pattern Recognit. Lett. | 1 |
| 2005 | Multiple Paths Extraction in Images Using a Constrained Expanded TrellisabstractSingle shortest path extraction algorithms have been used in a number of areas such as network flow and image analysis. In image analysis, shortest path techniques can be used for object boundary detection, crack detection, or stereo disparity estimation. Sometimes one needs to find multiple paths as opposed to a single path in a network or an image where the paths must satisfy certain constraints. In this paper, we propose a new algorithm to extract multiple paths simultaneously within an image using a constrained expanded trellis (CET) for feature extraction and object segmentation. We also give a number of application examples for our multiple paths extraction algorithm. Changming Sun, Ben Appleton |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2004 | Fast Stereo Matching by Iterated Dynamic Programming and Quadtree SubregioningabstractThe application of energy minimisation methods for stereo matching has been demonstrated to produce high quality disparity maps. However the majority of these methods are known to be computationally expensive, requiring minutes or even hours of computation. We propose a fast minimisation scheme that produces strongly competitive results for significantly reduced computation, requiring only a few seconds of computation. In this paper, we present our iterated dynamic programming algorithm along with a quadtree subregioning process for fast stereo matching. 1 Carlos Leung, Ben Appleton, Changming Sun |
BMVC | 3 |
| 2004 | Fast panoramic stereo matching using cylindrical maximum surfacesabstractThis paper presents a fast panoramic stereo matching algorithm using a cylindrical maximum surface technique. The disparity for a pair of panoramic images is found in a cylindrical shaped correlation coefficient volume by obtaining the maximum surface rather than simply choosing a position that gives the maximum correlation coefficient value. The use of our cylindrical maximum surface technique ensures that the disparities obtained at the left and the right columns of the panoramic stereo images are properly constrained. Typical running time for a pair of 1324 x 120 images is about 0.33 s on a 1.7-GHz PC. A variety of real images have been tested, and good results have been obtained. Changming Sun, Shmuel Peleg |
IEEE Trans. Syst. Man Cybern. Part B | 1 |
| 2003 | Uncalibrated three-view image rectification
Changming Sun |
Image Vis. Comput. | 1 |
| 2003 | Circular shortest paths by branch and bound
Ben Appleton, Changming Sun |
Pattern Recognit. | 2 |
| 2003 | Circular shortest path in images
Changming Sun, Stefano Pallottino |
Pattern Recognit. | 1 |
| 2002 | Fast Stereo Matching Using Rectangular Subregioning and 3D Maximum-Surface Techniques
Changming Sun |
Int. J. Comput. Vis. | 1 |
| 2002 | Fast optical flow using 3D shortest path techniques
Changming Sun |
Image Vis. Comput. | 1 |
| 2000 | Algorithm Performance ContestabstractThis contest involved the running and evaluation of computer vision and pattern recognition techniques on different data sets with known groundwidth. The contest included three areas; binary shape recognition, symbol recognition and image flow estimation. A package was made available for each area. Each package contained either real images with manual groundtruth or programs to generate data sets of ideal as well as noisy images with known groundtruth. They also contained programs to evaluate the results of an algorithm according to the given groundtruth. These evaluation criteria included the generation of confusion matrices, computation of the misdetection and false alarm rates and other performance measures suitable for the problems. The paper summarizes the data generation for each area and experimental results for a total of six participating algorithms. Selim Aksoy, Michael L. Schauf, Mingzhou Song 0001, Yalin Wang 0001, Robert M. Haralick, Jim R. Parker, Juraj Pivovarov, Dominik Royko, Changming Sun, Gunnar Farnebäck |
ICPR | 10 |
| 1997 | Skew and Slant Correction for Document Images Using Gradient DirectionabstractA fast algorithm is presented for skew and slant correction in printed document images. The algorithm employs only the gradient information. The skew angle is obtained by searching for a peak in the histogram of the gradient orientation of the input grey-level image. The skewness of the document is corrected by a rotation at such an angle. The slant of characters can also be detected using the same technique, and can be corrected by a shear operation. A second method for character slant correction by fitting parallelograms to the connected components is also described. Document images with different contents (tables, figures, and photos) have been tested for skew correction and the algorithm gives accurate results on all the test images, and the algorithm is very easy to implement. Changming Sun, Deyi Si |
ICDAR | 1 |
| 1997 | 3D Symmetry Detection Using The Extended Gaussian ImageabstractSymmetry detection is important in the area of computer vision. A 3D symmetry detection algorithm is presented in this paper. The symmetry detection problem is converted to the correlation of the Gaussian image. Once the Gaussian image of the object has been obtained, the algorithm is independent of the input format. The algorithm can handle different kinds of images or objects. Simulated and real images have been tested in a variety of formats, and the results show that the symmetry can be determined using the Gaussian image. Changming Sun, Jamie Sherrah |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 1995 | Symmetry detection using gradient information
Changming Sun |
Pattern Recognit. Lett. | 1 |
| 1994 | Robust Estimation for Motion ParametersabstractThe performance of least squares method ca'n be improved by changing the error metric so that points which lie far from the bulk of data do not influence the final value—that is to reject the outliers. In this paper, the combination of two robust estimators are used to obtain motion parameters of a camera from matched image features. Results obtained show that the robust estimator has the ability to remove gross errors and mismatch points automatically. Therefore the existence of a small amount of outliers or mismatches will not affect the final results. Different robust estimators can be used for the purpose of parameters estimation. In our case the Huber and Tukey's estimators are used which allows ten per cent mismatches. Median absolute estimator can also be used which can allow as high as fifty per cent outliers in the whole corresponding points. But one of the disadvantage is that it will take much more time. 1 Changming Sun |
BMVC | 1 |