Sutong Wang

dblp:256/0200 · DBLP profile ↗
← Back
12ranked-venue papers
3as first author
12since 2021 · last 2026
0000-0001-6603-6047ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Applied, interdisciplinary, general and emerging computing · 7 · 7 since 2021Artificial intelligence and machine learning · 3 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021
YearPublicationVenuePosition
2026 SMR-agents: Synergistic medical reasoning agents for zero-shot medical visual question answering with MLLMs
Dujuan Wang, T. C. E. Cheng, Sutong Wang, Youhua (Frank) Chen, Yunqiang Yin
Inf. Process. Manag.3
2025 An explainable lesion detection transformer model for medical imaging diagnosis decision support: Design science research
Sutong Wang, Dujuan Wang, T. C. E. Cheng
Decis. Support Syst.3
2024 Sea-Land Segmentation of Polarimetric SAR Images Based on Thresholding USING Polarimetric Likelihood Ratio
abstract
A sea-land segmentation method in polarimetric SAR images is proposed based on thresholding using polarimetric likelihood ratio in this article. Firstly, the volume scattering power is extracted based on three-component decomposition. Then, a typical sampling region of the sea is extracted based on the mean and variance of the region to determine the range interval of the threshold. Finally, the optimal threshold such that the polarimetric likelihood ratio between the sea and land region reaches the maximum is searched in the interval. Experiments of different sensors of data in San Francisco including AirSAR, RADARSAT-2, TerraSAR-X and Gaofen-3 show that both the precision and recall rates of the proposed method are higher than 90%, significantly surpassing those of the comparison methods.
Sutong Wang, Jian Yang 0011
IGARSS2
2024 Coastline Detection in Polarimetric SAR Images Based on Freeman Decomposition and Three-Region Markov Random Field Segmentation
abstract
The difficulty of coastline detection in polarimetric SAR images is that the coast zone includes both the sea noised by the sidelobe echo of strong scattering targets and intertidal zones with varying water content. Methods based solely on the statistical distribution of quad-polarization coherent matrix or single-channel scattering intensity would result in incorrect segmentation. In this paper, a coastline detection method in polarimetric SAR images based on three-region Markov random field (MRF) segmentation with embedded components obtained by Freeman decomposition is proposed. Firstly, the powers of Freeman three-component decomposition of the land cover are obtained. Then, the 3 × 3 scattering coherent matrix is decomposed into two parts of volume scattering power and non-volume scattering 2×2 coherent matrix based on the reflective symmetry of natural land covers. A joint probability distribution of the two parts is constructed and embedded into the three-region MRF model to segment the image into sea, strong scattering, and other regions. The coastline is finally extracted by determining the boundary of the segmented sea region. To improve the accuracy and robustness of the initial segmentation, a threshold segmentation method by sampling typical regions is proposed to segment the double-bounce and volume scattering powers. Experimental results using RADARSAT-2 data from regions in Dalian and Boao of China, and Singapore demonstrate that the proposed method accurately detects coastlines in different scenes, with an average offset of less than 2 pixels. The performance is significantly superior to coastline detection methods based on two-region thresholding segmentation, MRF segmentation, and level set segmentation.
Jian Yang 0011, Sutong Wang
IEEE Geosci. Remote. Sens. Lett.4
2024 Context Feature-Based Bridge Detection for Polarimetric SAR Images
abstract
A bridge detection method in polarimetric SAR images based on the context feature of the bridge is proposed in this article. Using the context feature that bridge crossing over water, water and land are segmented using the Markov random field (MRF) segmentation method based on the initialization segmentation of the edge energy of the volume scattering component first. The candidate bridge pixels surrounded by water are then detected by scanning the jetties in multidirection. Using the geometric feature that bridge appears as a long strip, an approximate rectangle detection criterion is proposed for the initial extraction of the regions of interest (ROIs) of bridges first. A directional entropy (DE) parameter used for the determination of the main direction and boundary of the extracted long strip region is then proposed to refine the ROIs. Using the context feature that both upper and lower ends of a bridge are large land masses and both left and right sides are large water areas, bridges are finally recognized by measuring the proportion of land pixels at both ends and water pixels at both sides. The results obtained by data of AirSAR in San Francisco of USA, RADARSAT-2 in Fujian, Boao and Lingshui of China demonstrate the effectiveness of the proposed method. Compared with the spatial-based bridge detection method, the proposed method can effectively reduce false alarms and missed detections, and the intersection over union (IoU) index is improved by at least 5%.
Sutong Wang, Jian Yang 0011
IEEE Geosci. Remote. Sens. Lett.2
2024 Thresholding of Polarimetric SAR Images of Coastal Zones Based on Three-Component Decomposition and Likelihood Ratio
abstract
To address the problem of threshold segmentation of polarimetric images in coastal zones when the intensity is not bimodal under complex coastal environments, a thresholding method based on three-component decomposition and likelihood ratio is proposed in this article. First, the double-bounce and volume scattering powers are extracted by three-component decomposition. Then, a three-step thresholding diagram is carried out in the two scattering powers. The threshold base point is determined by sampling a typical window regions with the minimum or maximum average power and variance first. The threshold interval is then determined by adding or subtracting a fluctuation range to the base point. The threshold point with the maximum likelihood ratio is then searched within the interval. The final two-region or three-region segmentations of coastal images are carried out by segmenting the volume scattering power to obtain the low scattering region and segmenting the double-bounce scattering power to obtain the strong scattering region. The experimental results of RADARSAT-2 Singapore, AirSAR San Francisco, as well as TerraSAR-X Singapore data, demonstrate that the proposed method outperforms the traditional thresholding method. It effectively reduces the wrong segmentation in complex coastal scenes, showcasing an average segmentation precision exceeding 85% and an average intersection over union rate closing to 0.75. Code is available at https://github.com/polSAR-lab/polSAR-thresholding.
Sutong Wang, Jian Yang 0011
IEEE Geosci. Remote. Sens. Lett.2
2023 Context Feature Based Bridge Detection for Polarimetric SAR Images
abstract
It is difficult to extract accurate geometric features for bridge detection in SAR images. A bridge detection method in polarimetric SAR images based on context feature of the bridge is proposed in this paper. Using the context feature that bridge crossing over water, water and land are segmented using the Markov random field segmentation method by extracting the surface, double-bounce, and volume scattering powers of the terrain first. The candidate bridge pixels are then detected by the proposed multi-direction jetty scanning method. Using the context feature that bridge appears as a long strip and both ends are large land regions, an approximate rectangle detection method is proposed to extract the regions of interest (ROIs) of bridges first. The false alarms generated by dykes and dams are then discriminated by the context features. The bridge detection results using AirSAR in San Francisco, RADARSAT-2 in Fujian and TerraSAR-X in Singapore demonstrate the effectiveness of the proposed method. The detection rate is high and the false alarm rate is low.
Sutong Wang, Jian Yang 0011
IGARSS2
2023 Explainable Multitask Shapley Explanation Networks for Real-Time Polyp Diagnosis in Videos
abstract
Colorectal cancer is mostly caused by colorectal polyps, which can be prevented through polyp diagnosis using colonoscopy. The current computer-aided decision-making methods suffer from a variety of drawbacks, including inaccurate polyp classification, poor real-time performance, and poor interpretability. To address these issues, we propose an explainable multitask Shapley explanation networks (EMSEN) that can perform real-time explainable multitasks such as polyp detection and classification in colonoscopy videos. The EMSEN accepts two multimodal inputs of different light sources, and outputs the polyp location, classification type, and diagnosis results according to the real-time colonoscopy video, where efficient channel attention (ECA) Mechanism-based network and Shapley explanation networks (ShapNet) are designed to improve the feature extraction performance and model interpretability, respectively. Extensive experiment studies are conducted to verify the efficiency and effectiveness of the proposed method by comparing with the experts and state-of-the-art methods. The results demonstrate that the developed method performs the best, which achieves competitive diagnosis performance.
Dujuan Wang, Sutong Wang, Yunqiang Yin
IEEE Trans. Ind. Informatics3
2023 Interpretable Multi-Modal Stacking-Based Ensemble Learning Method for Real Estate Appraisal
abstract
With the development of online real estate trading platforms, multi-modal housing trading data, including structural information, location, and interior image data, are being accumulated. The accurate appraisal of real estate makes sense for government officials, urban policymakers, real estate sellers, and personal purchasers. In this study, we propose an interpretable multi-modal stacking-based ensemble learning (IMSEL) method that deals with various modalities for real estate appraisals. We crawl the structural and image data of real estate in Chengdu city, China from the nation's largest real estate transaction platform with the location information, including public services, within 2 km of the real estate using Baidu map. We then compare the predictive results from IMSEL with those from previous state-of-art methods in the literature in terms of the root mean square error, mean absolute percentage error, mean absolute error, and coefficient of determination (R2). The comparison results show that IMSEL outperformed the other methods. We verified the improvement of introducing a data transformation strategy and deep visual features through a 10-fold cross-validation. We also discuss the managerial implications of our research findings.
Sutong Wang, Yunqiang Yin, Dujuan Wang, T. C. E. Cheng, Yanzhang Wang
IEEE Trans. Multim.1
2022 Interpretability-Based Multimodal Convolutional Neural Networks for Skin Lesion Diagnosis
abstract
Skin lesion diagnosis is a key step for skin cancer screening, which requires high accuracy and interpretability. Though many computer-aided methods, especially deep learning methods, have made remarkable achievements in skin lesion diagnosis, their generalization and interpretability are still a challenge. To solve this issue, we propose an interpretability-based multimodal convolutional neural network (IM-CNN), which is a multiclass classification model with skin lesion images and metadata of patients as input for skin lesion diagnosis. The structure of IM-CNN consists of three main paths to deal with metadata, features extracted from segmented skin lesion with domain knowledge, and skin lesion images, respectively. We add interpretable visual modules to provide explanations for both images and metadata. In addition to area under the ROC curve (AUC), sensitivity, and specificity, we introduce a new indicator, an AUC curve with a sensitivity larger than 80% (AUC_SEN_80) for performance evaluation. Extensive experimental studies are conducted on the popular HAM10000 dataset, and the results indicate that the proposed model has overwhelming advantages compared with popular deep learning models, such as DenseNet, ResNet, and other state-of-the-art models for melanoma diagnosis. The proposed multimodal model also achieves on average 72% and 21% improvement in terms of sensitivity and AUC_SEN_80, respectively, compared with the single-modal model. The visual explanations can also help gain trust from dermatologists and realize man-machine collaborations, effectively reducing the limitation of black-box models in supporting medical decision making.
Sutong Wang, Yunqiang Yin, Dujuan Wang, Yanzhang Wang, Yaochu Jin
IEEE Trans. Cybern.1
2022 A Clustering-Based Optimization Method for the Driving Cycle Construction: A Case Study in Fuzhou and Putian, China
abstract
Driving cycle is a crucial topic for the auto industry. It is developed to provide a quantitative measure on the fuel consumption and emission of a vehicle. In recent years, massive amount of driving data has been collected but has not yet been commonly used for the evaluation of driving cycle. We believe the collection of such data and the advancement in analytics models may provide a fresh perspective for the construction of driving cycle. Therefore, we propose a novel clustering-based optimization method for the construction of driving cycles. We employ the principal component analysis and spectral clustering algorithms to eliminate redundant features and analyze data structure. We further develop an adaptive optimization algorithm to select the appropriate kinematic segments to form a representative driving cycle. To demonstrate the effectiveness of our method, we compare our performance against the baselines including the New European Driving Cycle (NEDC), Federal Test Procedure (FTP), and Markov chain-based methods. The model performance is evaluated with real driving data from two cities in Fujian, China. Our proposed method is shown to be superior to all baselines. In addition, based on our optimized driving cycle, we can also estimate the fuel consumption to evaluate its energy economy. To sum up, this study offers a novel methodology to establish the driving cycle based on real and localized traffic data, where the constructed driving cycle can further be used for the development of energy economy and emission control.
Huaxin Qiu 0002, Shaoze Cui, Sutong Wang, Yanzhang Wang, Mengling Feng
IEEE Trans. Intell. Transp. Syst.3
2021 An interpretable deep neural network for colorectal polyp diagnosis under colonoscopy
Sutong Wang, Yunqiang Yin, Dujuan Wang, Zehui Lv, Yanzhang Wang, Yaochu Jin
Knowl. Based Syst.1