Ming Wu 0001

dblp:32/2572-1 · DBLP profile ↗
← Back
40ranked-venue papers
0as first author
34since 2021 · last 2026
0000-0001-8390-5398ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 25 · 20 since 2021Artificial intelligence and machine learning · 15 · 13 since 2021Applied, interdisciplinary, general and emerging computing · 12 · 12 since 2021Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1
YearPublicationVenuePosition
2026 From Sight to Insight: Enhancing Confusable Structure Segmentation via Vision-Language Mutual Prompting
abstract
Confusable structure segmentation (CSS) is a type of semantic segmentation applied in remote sensing sea fog detection, medical image segmentation, camouflaged object detection, etc. Structural similarity and visual ambiguity are two critical issues in CSS that pose difficulties in distinguishing foreground objects from the background. Current methods focus primarily on enhancing visual representations and do not often incorporate multimodal information, which leads to performance bottlenecks. Inspired by recent achievements in vision-language models, we proposeVision-LanguageMutualPrompting (VLMP), a novel and unified language-guided framework that leverages text prompts to enhance CSS. Specifically, VLMP consists of vision-to-language prompting and language-to-vision prompting, which bidirectionally model the interactions between visual and linguistic features, thereby facilitating cross-modal complementary information flow. To prevent the predominance of one modality over another, we design a feature integration modulator that modulates and balances feature weights for adaptive multimodal fusion. Our framework is designed to be modular and flexible, allowing for integration with any backbone, including CNNs and transformers. We evaluate VLMP with three diverse datasets: SFDD-H8, QaTa-COV19, and CAMO-COD10K. Extensive experiments demonstrate the effectiveness and superiority of the proposed framework over those of state-of-the-art methods across these datasets. This shift from basicsightto deeperinsightin CSS through vision-language integration represents a significant advancement in the field.
Yihao Zuo, Mengqiu Xu, Kaixin Chen 0001, Ming Wu 0001, Zhanyu Ma, Jun Guo 0002
IEEE Trans. Multim.5
2025 GTR: General Handwritten Lines Text Recognition Dataset
abstract
In recent years, as printed content recognition and understanding have improved, the more challenging task of handwritten content recognition and understanding has come into focus. Due to the diversity of targets and the heterogeneity of patterns, previous work has typically designed task-specific architectures and objectives for individual tasks, which has inadvertently led to model isolation and complex workflows. However, in real-world handwriting, especially in notes from many common disciplines, mixed objectives are more commonly encountered scenarios. To this end, we propose a new paradigm for precise multi-type handwritten content recognition in scenarios involving a mix of line-level handwritten texts in Chinese, Latin, Arabic, formulas, sheet music, etc., to enhance understanding in complex scenarios. For the first time, we present a multi-type mixed line-level handwritten text dataset GTR-D to assist existing research. At the same time, we propose GTR, a General-purpose end2end handwritten Text Recognition model with strong general capability, which shows good performance in each scenario. Extensive experiments on multiple standard benchmarks demonstrate that, despite its simple design, the proposed network achieves state-of-the-art (SOTA) or highly competitive performance across 5 handwritten content recognition tasks and 9 datasets. It is worth noting that our method is significantly superior to GTP4V in this field. Our data can be accessed at https://github.com/JXXXX001/GTR-D.
Haizhao Sun, Ming Wu 0001
CBMI4
2025 MFogHub: Bridging Multi-Regional and Multi-Satellite Data for Global Marine Fog Detection and Forecasting
abstract
Deep learning approaches for marine fog detection and forecasting have outperformed traditional methods, demonstrating significant scientific and practical importance. However, the limited availability of open-source datasets remains a major challenge. Existing datasets, often focused on a single region or satellite, restrict the ability to evaluate model performance across diverse conditions and hinder the exploration of intrinsic marine fog characteristics. To address these limitations, we introduce MFogHub, the first multi-regional and multi-satellite dataset to integrate annotated marine fog observations from 15 coastal fog-prone regions and six geostationary satellites, comprising over 68,000 high-resolution samples. By encompassing diverse regions and satellite perspectives, MFogHub facilitates rigorous evaluation of both detection and forecasting methods under varying conditions. Extensive experiments with 16 baseline models demonstrate that MFogHub can reveal generalization fluctuations due to regional and satellite discrepancy, while also serving as a valuable resource for the development of targeted and scalable fog prediction techniques. Through MFogHub, we aim to advance both the practical monitoring and scientific understanding of marine fog dynamics on a global scale. The dataset and code are at https://github.com/kaka0910/MFogHub.
Mengqiu Xu, Kaixin Chen 0001, Heng Guo 0003, Ming Wu 0001, Jun Guo 0002
CVPR5
2025 Enhancing Multi-task Learning Capability of Medical Generalist Foundation Model via Image-centric Multi-annotation Data
abstract
The emergence of medical generalist foundation models has revolutionized conventional task-specific model development paradigms, aiming to better handle multiple tasks through joint training on large-scale medical datasets. However, recent advances prioritize simple data scaling or architectural component enhancement, while neglecting to re-examine multi-task learning from a data-centric perspective. Critically, simply aggregating existing data resources leads to decentralized image-task alignment, which fails to cultivate comprehensive image understanding or align with clinical needs for multi-dimensional image interpretation. In this paper, we introduce the image-centric multi-annotation X-ray dataset (IMAX), the first attempt to enhance the multi-task learning capabilities of medical multi-modal large language models (MLLMs) from the data construction level. To be specific, IMAX is featured from the following attributes: 1) High-quality data curation. A comprehensive collection of more than 354K entries applicable to seven different medical tasks. 2) Image-centric dense annotation. Each X-ray image is associated with an average of 4.10 tasks and 7.46 training entries, ensuring multi-task representation richness per image. Compared to the general decentralized multi-annotation X-ray dataset (DMAX), IMAX consistently demonstrates significant multi-task average performance gains ranging from 3.20% to 21.05% across seven open-source state-of-the-art medical MLLMs. Moreover, we investigate differences in statistical patterns exhibited by IMAX and DMAX training processes, exploring potential correlations between optimization dynamics and multi-task performance. Finally, leveraging the core concept of IMAX data construction, we propose an optimized DMAX-based training strategy to alleviate the dilemma of obtaining high-quality IMAX data in practical scenarios. Related resources are available at https://github.com/MSIIP/IMAX.
Fanbin Mo, Yiming Shi, Ming Wu 0001, Miao Li 0003, Ji Wu 0002
ACM Multimedia6
2025 LF-SAM: Prompting for Land Fog Recognition With Ground Observation Station Data
abstract
The utilization of ground observation stations for land fog monitoring is constrained by station distribution, which results in incomplete and nonuniform coverage of actual conditions. This problem can be overcome by using geostationary meteorological satellite multichannel observation images and semantic segmentation to recognize fog areas. Existing methods, however, rely on a large number of annotated satellite images and fail to make full use of other auxiliary information, such as the ground observation station data. The segment anything model (SAM) proposed by Meta AI can use prompts to guide segmentation, making it possible to use data from multiple sources to undertake land fog recognition tasks. In this letter, we proposed LF-SAM, a method designed to automatically generate point prompts derived from ground observation station data and extract high-frequency features, thereby aiding in the segmentation of satellite images. We created a dedicated dataset named FY-OBS for land fog recognition, including 300 training sets and 60 test sets, and validated our method on it. Our method reduces the amount of fully annotated data required and achieves significantly better performance than prompt-free methods. The mean intersection over union (mIoU) and average accuracy (aAcc) of LF-SAM reach 73.99% and 95.24%, which were increased by 3.05% and 0.93%, respectively, compared with the method without prompts.
Ming Wu 0001, Mengqiu Xu, Sundingkai Su
IEEE Geosci. Remote. Sens. Lett.2
2025 Enhanced Seafog Detection Using Deep Learning and Contrastive Learning: Integrating Spectral and Motion Features
abstract
Seafog detection is a complex challenge in meteorology, primarily due to the extensive distribution of seafog and the limited availability of maritime observation stations. Remote sensing meteorological satellites, with their wide-area observation capabilities and abundant data resources, hold significant potential for seafog detection. However, current seafog detection approaches—whether based on traditional thresholding techniques or deep learning methods—do not fully exploit the multichannel information provided by remote sensing data, nor do they effectively capture the physical motion differences between seafog and other cloud types. Deep learning has demonstrated notable performance in seafog detection tasks, yet traditional threshold-based methods for seafog detection still hold the advantage of strong physical interpretability. Furthermore, challenges such as fuzzy seafog boundaries and overlap with clouds lead to discontinuous and incomplete detection regions, posing significant challenges for practical applications. To address these issues, we propose a deep learning approach that integrates the spectral and motion physical characteristics of seafog and various cloud types, enhancing the feature distinctions between seafog and other clouds to enable more accurate and interpretable feature extraction. To further optimize feature extraction, we propose a contrastive learning mechanism for seafog, aimed at amplifying the distinctions between seafog and other cloud types while simultaneously reducing intraclass variability within seafog across different conditions. Finally, we propose a probabilistic mask representation, which effectively mitigates the issue of regional discontinuities in detection, thereby enhancing the overall performance of the detection process. Our extensive experiments on the FY-4A satellite dataset demonstrate a CSI of 64.98%, which is 7.41% higher than the current best-performing method.
Ming Wu 0001, Luming Xiao, Mengqiu Xu
IEEE Trans. Geosci. Remote. Sens.2
2025 GenP2C: A General Pipeline for Road Extraction From Partial Road Maps
abstract
We propose a pipeline for extracting complete road networks from partial road maps. The pipeline is architecturally simpler, more powerful, more computationally efficient, and better suited for general road extraction models than the existing method like P2CNet. Inspired by the philosophy of conditional control in diffusion models from AIGC research, we introduce anEfficient partialMap Encoder (EME) that extracts multi-level partial Map embeddings to guide the road extraction model in learning road topology. To interact the EME with the main (image) encoder, we propose a Self-Activation Addition (SAA) module to fuse the Map embedding with the Image embedding, where SAA only occupies a few portion of whole computational budget and conducts a level-to-level fusion to effectively prompt the road extraction model with discriminative road information. The resulting multi-level Image-Map embeddings are fed into our newly developed hierarchical decoder, theMulti-Road-Scale Attention Network (Mrs.Former), which is based on the varying-window multi-scale attention mechanism. Overall, our pipeline, termedGeneral Partial roadtoComplete road (GenP2C), uses SAA to connect EME with the image encoder, producing Image-Map embeddings for the decoder (Mrs.Former), and its learning objective is identical to general road extraction models. Compared with P2CNet, GenP2C consistently improves performance by an average 1.7% road IoU on challenging datasets such as DeepGlobe, SpaceNet, and OSM, while using only 72% of P2CNet’s computation cost. More prominently, EME and SAA can serve as plug-and-play modules applicable to other road extraction models that aim to extract complete roads from partial ones. GenP2C demonstrates strong compatibility with general road extraction models, offering 0.8%-1.2% improvements in road IoU across five prevalent road extraction models on the P2C road extraction with cheap computational overhead. GenP2C also demonstrates strong generalization to multi-spectral satellite images and spatial misalignments between partial maps and imagery by continuously achieving better performance than P2CNet.
Kaiwen Jing, Yihao Zuo, Ming Wu 0001
IEEE Trans. Geosci. Remote. Sens.4
2025 Class-Customized Domain Adaptation: Unlock Each Customer-Specific Class With Single Annotation
abstract
Model customization mitigates the issues of inadequate performance, resource wastage, and privacy risks associated with using general-purpose models in specialized domains and well-defined tasks. However, achieving customization at a low annotation cost still poses a challenge. Existing domain adaptation research has addressed cases where all customized classes are present in the labeled database, yet scenarios involving customer-specific classes are still unresolved. Therefore, this paper proposes a novel Class-Customized Domain Adaptation (CCDA) method, addressing the latter scenario with just one additional annotation for each customer-specific class. CCDA adopts the classic adaptation training framework and comprises two innovative techniques. Firstly, to ensure the shared class knowledge from the database and the private class knowledge from additional annotations are transferred and propagated to the correct regions within the target domain, we design the partial-feature alignment strategy, based on the mechanical properties of feature alignment. Second, we propose soft-balanced sampling to tackle the long-tail distribution problem in labeled data, preventing the model from overfitting to the labeled samples of customer-specific classes. The effectiveness of CCDA has been validated across 48 tasks simulated on domain adaptation benchmarks and two real-world customization scenarios, consistently showing excellent performance. Additionally, extensive analytical experiments illustrate the contributions of two innovative techniques. The code is available at https://github.com/CHEN-kx/ClassCustomizedDA.
Kaixin Chen 0001, Huiying Chang, Mengqiu Xu, Ruoyi Du, Ming Wu 0001, Zhanyu Ma
IEEE Trans. Image Process.5
2024 Privileged Prior Information Distillation for Image Matting
abstract
Performance of trimap-free image matting methods is limited when trying to decouple the deterministic and undetermined regions, especially in the scenes where foregrounds are semantically ambiguous, chromaless, or high transmittance. In this paper, we propose a novel framework named Privileged Prior Information Distillation for Image Matting (PPID-IM) that can effectively transfer privileged prior environment-aware information to improve the performance of trimap-free students in solving hard foregrounds. The prior information of trimap regulates only the teacher model during the training stage, while not being fed into the student network during actual inference. To achieve effective privileged cross-modality (i.e. trimap and RGB) information distillation, we introduce a Cross-Level Semantic Distillation (CLSD) module that reinforces the students with more knowledgeable semantic representations and environment-aware information. We also propose an Attention-Guided Local Distillation module that efficiently transfers privileged local attributes from the trimap-based teacher to trimap-free students for the guidance of local-region optimization. Extensive experiments demonstrate the effectiveness and superiority of our PPID on image matting. The code will be released soon.
Jiake Xie, Bo Xu 0031, Cheng Lu 0006, Han Huang 0005, Ming Wu 0001
AAAI7
2024 A Multimodal Network on Handwritten Chinese Character Error Correction
Haizhao Sun, Ming Wu 0001
BMVC5
2024 Multi-Scale Representations by Varying Window Attention for Semantic Segmentation
abstract
Multi-scale learning is central to semantic segmentation. We visualize the effective receptive field (ERF) of canonical multi-scale representations and point out two risks learning them: \textit{scale inadequacy} and \textit{field inactivation}. A novel multi-scale learner, \textbf{varying window attention} (VWA), is presented to address these issues. VWA leverages the local window attention (LWA) and disentangles LWA into the query window and context window, allowing the context's scale to vary for the query to learn representations at multiple scales. However, varying the context to large-scale windows (enlarging ratio $R$) can significantly increase the memory footprint and computation cost ($R^2$ times larger than LWA). We propose a simple but professional re-scaling strategy to zero the extra induced cost without compromising performance. Consequently, VWA uses the same cost as LWA to overcome the receptive limitation of the local window. Furthermore, depending on VWA and employing various MLPs, we introduce a multi-scale decoder (MSD), \textbf{VWFormer}, to improve multi-scale representations for semantic segmentation. VWFormer achieves efficiency competitive with the most compute-friendly MSDs, like FPN and MLP decoder, but performs much better than any MSDs. For instance, using nearly half of UPerNet's computation, VWFormer outperforms it by $1.0\%-2.5\%$ mIoU on ADE20K. At little extra overhead, $\sim 10$G FLOPs, Mask2Former armed with VWFormer improves by $1.0\%-1.3\%$.
Ming Wu 0001
ICLR2
2024 MMSISP: A Satellite Image Sequence Prediction Network with Multi-factor Decoupling and Multi-modal Fusion
Fanbin Mo, Ming Wu 0001
ICPR (22)3
2024 Multi-Scale Receptive Field Rectification of Remote Sensing Images
abstract
Remote sensing semantic segmentation refers to the pixel-level classification of high-resolution imagery obtained through remote sensing technology. In the era of deep learning, U-shaped network structures have been gaining popularity. These networks adopt a multi-level backbone to obtain multi-scale features with multiple receptive fields of different scales, which can make the network hierarchically understand contextual information. As the backbone goes deeper, the network is believed to obtain a larger receptive field. However, an in-depth effective receptive-field analysis reveals that such enlargement is unclear and the deepest receptive field is still local. Therefore, this paper conducts studies on the U-net structure and leverages dilated convolution to rectify the receptive field of multi-scale features. The effectiveness of the rectification is evaluated on the LoveDA benchmark and the rectified receptive fields are compared comprehensively with the non-rectified receptive fields.
Ming Wu 0001
IGARSS3
2024 CLIP-SP: Vision-language model with adaptive prompting for scene parsing
abstract
We present a novel framework, CLIP-SP, and a novel adaptive prompt method to leverage pre-trained knowledge from CLIP for scene parsing. Our approach addresses the limitations of DenseCLIP, which demonstrates the superior image segmentation provided by CLIP pre-trained models over ImageNet pre-trained models, but struggles with rough pixel-text score maps for complex scene parsing. We argue that, as they contain all textual information in a dataset, the pixel-text score maps, i.e., dense prompts, are inevitably mixed with noise. To overcome this challenge, we propose a two-step method. Firstly, we extract visual and language features and perform multi-label classification to identify the most likely categories in the input images. Secondly, based on the top-k categories and confidence scores, our method generates scene tokens which can be treated as adaptive prompts for implicit modeling of scenes, and incorporates them into the visual features fed into the decoder for segmentation. Our method imposes a constraint on prompts and suppresses the probability of irrelevant categories appearing in the scene parsing results. Our method achieves competitive performance, limited by the available visual-language pre-trained models. Our CLIP-SP performs 1.14% better (in terms of mIoU) than DenseCLIP on ADE20K, using a ResNet-50 backbone.
Jiaao Li, Ming Wu 0001
Comput. Vis. Media3
2024 Relation fusion propagation network for transductive few-shot learning
Hongyu Hao, Weichao Ge, Ming Wu 0001, Jun Guo 0002
Pattern Recognit.5
2023 SwiftAvatar: Efficient Auto-Creation of Parameterized Stylized Character on Arbitrary Avatar Engines
abstract
The creation of a parameterized stylized character involves careful selection of numerous parameters, also known as the "avatar vectors" that can be interpreted by the avatar engine. Existing unsupervised avatar vector estimation methods that auto-create avatars for users, however, often fail to work because of the domain gap between realistic faces and stylized avatar images. To this end, we propose SwiftAvatar, a novel avatar auto-creation framework that is evidently superior to previous works. SwiftAvatar introduces dual-domain generators to create pairs of realistic faces and avatar images using shared latent codes. The latent codes can then be bridged with the avatar vectors as pairs, by performing GAN inversion on the avatar images rendered from the engine using avatar vectors. Through this way, we are able to synthesize paired data in high-quality as many as possible, consisting of avatar vectors and their corresponding realistic faces. We also propose semantic augmentation to improve the diversity of synthesis. Finally, a light-weight avatar vector estimator is trained on the synthetic pairs to implement efficient auto-creation. Our experiments demonstrate the effectiveness and efficiency of SwiftAvatar on two different avatar engines. The superiority and advantageous flexibility of SwiftAvatar are also verified in both subjective and objective evaluations.
Shizun Wang, Weihong Zeng, Ming Wu 0001, Yunzhao Zeng
AAAI7
2023 Weather2K: A Multivariate Spatio-Temporal Benchmark Dataset for Meteorological Forecasting Based on Real-Time Observation Data from Ground Weather Stations
abstract
Weather forecasting is one of the cornerstones of meteorological work. In this paper, we present a new benchmark dataset named Weather2K, which aims to make up for the deficiencies of existing weather forecasting datasets in terms of real-time, reliability, and diversity, as well as the key bottleneck of data quality. To be specific, our Weather2K is featured from the following aspects: 1) Reliable and real-time data. The data is hourly collected from 2,130 ground weather stations covering an area of 6 million square kilometers. 2) Multivariate meteorological variables. 20 meteorological factors and 3 constants for position information are provided with a length of 40,896 time steps. 3) Applicable to diverse tasks. We conduct a set of baseline tests on time series forecasting and spatio-temporal forecasting. To the best of our knowledge, our Weather2K is the first attempt to tackle weather forecasting task by taking full advantage of the strengths of observation data from ground weather stations. Based on Weather2K, we further propose Meteorological Factors based Multi-Graph Convolution Network (MFMGCN), which can effectively construct the intrinsic correlation among geographic locations based on meteorological factors. Sufficient experiments show that MFMGCN improves both the forecasting performance and temporal robustness. We hope our Weather2K can significantly motivate researchers to develop efficient and accurate algorithms to advance the task of weather forecasting. The dataset can be available at https://github.com/bycnfz/weather2k/.
Yutong Xiong, Ming Wu 0001, Gaozhen Nie
AISTATS3
2023 E2SAM: A Pipeline for Efficiently Extending SAM's Capability on Cross-Modality Data via Knowledge Inheritance
Sundingkai Su, Mengqiu Xu, Kaixin Chen 0001, Ming Wu 0001
BMVC4
2023 Ariadne's Thread: Using Text Prompts to Improve Segmentation of Infected Areas from Chest X-ray Images
Mengqiu Xu, Kongming Liang, Kaixin Chen 0001, Ming Wu 0001
MICCAI (4)5
2023 Change-Aware Network for Damaged Roads Recognition and Assessment Based on Multi-temporal Remote Sensing Imageries
Ming Wu 0001, Binzhu Xie
PRCV (4)2
2023 Weakly Supervised Sea Fog Detection in Remote Sensing Images via Prototype Learning
abstract
Sea fog detection is a challenging and significant task in the field of remote sensing. Deep learning-based methods have shown promising potential, but require a large amount of pixel-level labeled data that are time-consuming and labor-intensive to acquire. To scale up the dataset and overcome the limitations of pixel-level annotation, we attempt to explore the existing knowledge from historical statistics for label efficient sea fog detection. In this paper, we propose an image-level Weakly Supervised Sea Fog Detection Dataset (WS-SFDD) and a novel weakly supervised sea fog detection framework via prototype learning, named ProCAM. According to the sea fog events recorded by the Marine Weather Review published quarterly by the National Meteorological Center of China, we collect the sea fog images from Himawari-8 satellite data and obtain free image-level labels to construct the dataset. However, with image-level annotations, existing weakly supervised semantic segmentation methods mainly rely on class activation maps (CAMs) and have limitations when applied to such a specific scenario: 1) the pseudo labels mainly cover the most discriminative part of object regions that are incomplete; 2) the background is complex with varying atmospheric conditions and it is difficult to distinguish sea fog from low clouds due to their high similarity in spectral characteristics; 3) the co-occurring context like ‘sea’ distracts the model and thus degrades the performance. To address the above issues, in our proposed ProCAM, we first design a prototype re-activation (PRA) module that reactivates self-similar sea fog regions by pixel-to-prototype feature matching to improve the robustness and completeness of CAMs. Then, we develop a pixel-to-prototype contrastive (PPC) learning method to increase the distance between sea fog and background in the embedding space for learning more discriminative dense features. Finally, a self-augmented regularization (SAR) strategy is presented to decouple sea fog from its co-occurring context and thus avoid background interference. Extensive experiments on the WS-SFDD dataset demonstrate our proposed method ProCAM achieves superior performance with an F1-score of 77.59% and a critical success index of 63.39%. To the best of our knowledge, this is the first work to perform image-level weakly supervised sea fog detection in remote sensing images. The dataset and code are available at https://github.com/yixianghuang/ProCAM.
Ming Wu 0001, Xin Jiang 0036, Jiaao Li, Mengqiu Xu, Jun Guo 0002
IEEE Trans. Geosci. Remote. Sens.2
2022 Dive into Plane: Lightweight & Modular Linear Projection Cross-dimensional Network for Retinal Vessel Segmentation in OCTA Images
abstract
Optical coherence tomography angiography (OCTA) is a new non-invasive imaging technique that can generate high-resolution volumetric blood flow information in a few seconds for retinal vessel imaging. However, retinal vessel segmentation from OCTA images is a challenging cross-dimensional task which requires efficiently acquiring spatial feature from voxels and projecting it into a plane. Considering that using 3D convolution to extract features causes exponential computational resource consumption, we simplify the design of the projection module and improve the decoding capability of the segmentation network. In this paper, we propose a lightweight Linear Projection Module (LPM) for dimensionality reduction and transformation of spatial features, and a general Multiple Attention Decoder (MADecoder) for fusing and enhancing the representation of multi-level semantic information. Our proposed cross-dimensional end-to-end segmentation network achieves SOTA results on the OCTA-500 dataset. Compared with previous method, our method reduces GPU memory usage by 55.18% and arithmetic complexity by 95.29%, dice score improved by 0.25% while reducing the number of parameters by 22.25%.
Mengqiu Xu, Ming Wu 0001
BIBM3
2022 PaRK-Detect: Towards Efficient Multi-Task Satellite Imagery Road Extraction via Patch-Wise Keypoints Detection
Shenwei Xie, Wanfeng Zheng, Zhenglin Xian, Junli Yang, Ming Wu 0001
BMVC6
2022 Pointshift: Point-Wise Shift MLP for Pixel-Level Cloud Type Classification in Meteorological Satellite Imagery
abstract
The deep neural network has recently achieved promising results on cloud type classification, which gets rid of the hand-crafted features and plays an essential role in climate change analysis. Previous methods perform context reasoning with single-scale representation at the centre of the network, which is challenging to capture sufficient contextual information. In this paper, we propose a point-wise shift multi-layer perceptron (MLP) for pixel-level cloud type classification, termed PointShift, which effectively models point-wise and multi-scale neighbour information. We design a shift operation to compose a multi-scale receptive field in a non-parametric manner. Besides, we introduce split attention to improve the interaction of feature channels. Extensive experiments on the Himawari-8 image dataset demonstrate that our proposed architecture achieves the best mIoU of 71.06% and a competitive trade-off between efficiency and performance.
Zhaoqing Wang, Xin Jiang 0036, Ming Wu 0001, Jun Guo 0002
IGARSS4
2022 PickDet: A Detection Framework for Aerial-view Scene
abstract
Detecting objects in the aerial-view scene is challenging for the objects usually have small scales relative to the image, making it hard to achieve high accuracy in full-image detection. Slice detection tries to overcome this by cutting the full image into slices before detecting them, but objects are sparsely distributed and usually clustered in local areas, a large number of background areas without objects can be ignored to improve detection efficiency. In this paper, we present PickDet, a framework for efficient and effective object detection in the aerial-view scene, which only chooses slices containing objects to conduct detection. The key components of PickDet include a lightweight convolutional network (PickNet), a screening strategy (SoftPick), and fine-tuned detectors. Given slices of aerial-view images, PickNet first outputs the probability of object existence. Then SoftPick conducts a double-threshold screening strategy to pick the slices which contain objects. Finally, all picked slices are fed into the detector in parallel and full-image detection is used as an auxiliary mean. Compared with previous methods, PickDet achieves higher accuracy and more efficiency in the aerial-view scene. We evaluate PickDet on Visdrone and Oiltank datasets, experiments show that PickDet can result in up to 28.0% AP improvement compared to full-image detection, and can result in up to 2.9% AP increase and up to 5 times inference speedup compared to slice detection.
Shizun Wang, Ming Wu 0001
VCIP4
2022 Identify, Guess and Reconstruct: Three Principles for Cloud Removal Task
abstract
Remote sensing images serve a significant role in earth observation to tackle climate change and post-disaster reconstruction concerns. However, optical images are obscured by clouds or haze, preventing precise earth observation; hence, cloud removal has been a hot topic among concerned scholars. The objective of this article is to make cloud removal more efficient and explicable by proposing three principles: identifying clouds, guessing objects beneath the clouds, and reconstructing the cloudy area. In addition, a modified dual contrastive learning Generative Adversarial Network is proposed based on these three principles by adding cloud detection and weight sharing strategy to obtain cloud semantics. In particular, we align two datasets by forming a quaternary sample pair that includes not only optical pictures and SAR images, but also region information for a more precise reconstruction. Our experiment results on the integrated dataset reveal the superiority of proposed method over previous cloud removal methods and the effectiveness of added modules through ablation experiments, with PSNR and SSIM values of 26.2 and 0.728, respectively.
Sibo Wu, Mengqiu Xu, Ming Wu 0001
VCIP3
2022 Annotating Only at Definite Pixels: A Novel Weakly Supervised Semantic Segmentation Method for Sea Fog Recognition
abstract
Sea fog recognition is a challenging and significant semantic segmentation task in remote sensing images. The fully supervised learning method relies on the pixel-level label, which is labor-intensive and time-consuming. Moreover, it is impossible to accurately annotate all pixels of the sea fog region due to the limited ability of the human eye to distinguish between low clouds and sea fog. In this paper, we propose a novel approach of point-based annotation for weakly supervised semantic segmentation with the auxiliary information of International Comprehensive Ocean-Atmosphere Data Set (ICOADS) visibility data. It only needs several definite points for both foreground and background, which significantly reduces the annotation cost of manpower. We conduct extensive experiments on Himawari-8 satellite remote sensing images to demonstrate the effectiveness of our annotation method. The mean intersection over union (mIoU) and overall recognition accuracy of our annotation method reach 82.72% and 95.18 %, respectively. Compared with the fully supervised learning method, the accuracy and the recognition rate of sea fog area are improved with a maximum increase of 7.69% and 9.69 %, respectively.
Mengqiu Xu, Ming Wu 0001
VCIP3
2022 A Correlation Context-Driven Method for Sea Fog Detection in Meteorological Satellite Imagery
abstract
Sea fog detection is a challenging and essential issue in satellite remote sensing. Although conventional threshold methods and deep learning methods can achieve pixel-level classification, it is difficult to distinguish ambiguous boundaries and thin structures from the background. Considering the correlations between neighbor pixels and the affinities between superpixels, a correlation context-driven method for sea fog detection is proposed in this letter, which mainly consists of a two-stage superpixel-based fully convolutional network (SFCNet), named SFCNet. A fully connected Conditional Random Field (CRF) is utilized to model the dependencies between pixels. To alleviate the problem of high cloud occlusion, an attentive Generative Adversarial Network (GAN) is implemented for image enhancement by exploiting contextual information. Experimental results demonstrate that our proposed method achieves 91.65% mIoU and obtains more refined segmentation results, performing well in detecting fogs in small, broken bits and weak contrast thin structures, as well as detects more obscured parts.
Ming Wu 0001, Jun Guo 0002, Mengqiu Xu
IEEE Geosci. Remote. Sens. Lett.2
2021 SamplingAug: On the Importance of Patch Sampling Augmentation for Single Image Super-Resolution
Shizun Wang, Ming Lu 0002, Kaixin Chen 0001, Jiaming Liu 0003, Xiaoqi Li 0009, Ming Wu 0001
BMVC7
2021 Overfitting the Data: Compact Neural Video Delivery via Content-aware Feature Modulation
abstract
Internet video delivery has undergone a tremendous explosion of growth over the past few years. However, the quality of video delivery system greatly depends on the Internet bandwidth. Deep Neural Networks (DNNs) are utilized to improve the quality of video delivery recently. These methods divide a video into chunks, and stream LR video chunks and corresponding content-aware models to the client. The client runs the inference of models to super-resolve the LR chunks. Consequently, a large number of models are streamed in order to deliver a video. In this paper, we first carefully study the relation between models of different chunks, then we tactfully design a joint training framework along with the Content-aware Feature Modulation (CaFM) layer to compress these models for neural video delivery. With our method, each video chunk only requires less than 1% of original parameters to be streamed, achieving even better SR performance. We conduct extensive experiments across various SR backbones, video time length, and scaling factors to demonstrate the advantages of our method. Besides, our method can be also viewed as a new approach of video coding. Our primary experiments achieve better video quality compared with the commercial H.264 and H.265 standard under the same storage cost, showing the great potential of the proposed method. Code is available at: https://github.com/Neural-video-delivery/ CaFM-Pytorch-ICCV2021
Jiaming Liu 0003, Ming Lu 0002, Kaixin Chen 0001, Xiaoqi Li 0009, Shizun Wang, Zhaoqing Wang, Enhua Wu, Yurong Chen 0001, Ming Wu 0001
ICCV10
2021 Damaged Road Extraction Based on Simulated Post-Disaster Remote Sensing Images
abstract
Damaged road extraction is a challenging task in the field of remote sensing. Some existing methods include the step to extract road from pre- and post-disaster remote sensing images of the same area. In practice, it often occurs that one of these two images is missing. To solve this problem, we use CoCosNet, the model for exemplar-based image translation, to translate pre-disaster images to simulated post-disaster ones. Then we use D-LinkNet, the state-of-the-art method in road extraction, to extract road from the pre- and post-disaster images of the same area. We extract damaged road area by comparing pre-disaster road masks with post-disaster ones and output the damage level by calculating the proportion of the damaged road area. Finally, we evaluate the damaged road extraction accuracy. Experimental results on simulated post-disaster images prove the effectiveness of the simulation method and the framework for damaged road extraction and damage level evaluation.
Yansong Huang, Haocai Wei, Junli Yang, Ming Wu 0001
IGARSS4
2021 Vecnet: A Spectral and Multi-Scale Spatial Fusion Deep Network for Pixel-Level Cloud Type Classification in Himawari-8 Imagery
abstract
Classifying cloud type is essential for analyzing the Earth's radiation budget and global climate change. However, cloud type classification is a challenging task as it requires domain knowledge while extracting features, various sets of parameters obtained from satellites, thresholds determined, and human intervention for analysis. This study proposes a deep learning-based pixel-level cloud type classification method, VecNet, to alleviate the dependency of the domain knowledge. VecNet mainly utilizes 1 × 1 convolution to learn robust local spectral information to generate fine-grained predictions, and designs a semantic pyramid module to capture multi-scale features to accurately classify different-morphological clouds. We conduct extensive experiments on all year Himawari-8 satellite images (6622 images), and VectNet achieved a competitive result of 67.9% mIOU with real-time inference speed.
Zhaoqing Wang, Zhanbei Cui, Ming Wu 0001, Mingming Gong, Tongliang Liu
IGARSS4
2021 Did-Linknet: Polishing D-Block with Dense Connection and Iterative Fusion for Road Extraction
abstract
Since the D-LinkNet won the first place in CVPR2018 Deep-Globe Challenge, the deep neural network built on encoder-decoder structure has become dominant in pixel-wise road extraction [1], [2]. The primary reason for D-LinkNet's success is the creative design of D-Block, which extracts features by dilated convolutions with receptive fields of progressively growing sizes and fuses the features with element-wise addition. Although D-Block promotes road connectivity significantly’ D-LinkNet still struggles in detecting road in regions with compact road topologies and indistinguishable ground objects from the real road. Therefore, in this paper, we obligate in augmenting the extraction and fusion of feature-maps in D-Block through two architectural amendments and upgrade D-Block to DID-Block. The first amendment is introduced to maximize the information flow between any layers in D-Block, which we name dense connection. The second one, iterative fusion, is proposed to aggregate representations learned by each layer with combining the layers' outputs iteratively. To construct the DID-LinkNet, we replace the D-Block in D-LinkNet with DID-Block. We demonstrate the profit of these two structural modifications separately on DeepGlobe2018 dataset, and the experimental results show that the proposed DID-LinkNet enjoys a further road connectivity gains.
Junli Yang, Ming Wu 0001
IGARSS4
2021 AP-CNN: Weakly Supervised Attention Pyramid Convolutional Neural Network for Fine-Grained Visual Classification
abstract
Classifying the sub-categories of an object from the same super-category (e.g., bird species and cars) in fine-grained visual classification (FGVC) highly relies on discriminative feature representation and accurate region localization. Existing approaches mainly focus on distilling information from high-level features. In this article, by contrast, we show that by integrating low-level information (e.g., color, edge junctions, texture patterns), performance can be improved with enhanced feature representation and accurately located discriminative regions. Our solution, named Attention Pyramid Convolutional Neural Network (AP-CNN), consists of 1) a dual pathway hierarchy structure with a top-down feature pathway and a bottom-up attention pathway, hence learning both high-level semantic and low-level detailed feature representation, and 2) an ROI-guided refinement strategy with ROI-guided dropblock and ROI-guided zoom-in operation, which refines features with discriminative local regions enhanced and background noises eliminated. The proposed AP-CNN can be trained end-to-end, without the need of any additional bounding box/part annotation. Extensive experiments on three popularly tested FGVC datasets (CUB-200-2011, Stanford Cars, and FGVC-Aircraft) demonstrate that our approach achieves state-of-the-art performance. Models and code are available at https://github.com/PRIS-CV/AP-CNN_Pytorch-master.
Zhanyu Ma, Shaoguo Wen, Jiyang Xie 0001, Dongliang Chang, Zhongwei Si, Ming Wu 0001, Haibin Ling
IEEE Trans. Image Process.7
2020 GINet: Graph Interaction Network for Scene Parsing
Yu Zhu 0006, Ming Wu 0001, Zhanyu Ma, Guodong Guo
ECCV (17)5
2020 Handwritten Style Recognition for Chinese Characters on HCL2020 Dataset
Peiyi Hu, Mengqiu Xu, Ming Wu 0001
PRCV (2)3
2020 The Devil is in the Channels: Mutual-Channel Loss for Fine-Grained Image Classification
abstract
The key to solving fine-grained image categorization is finding discriminate and local regions that correspond to subtle visual traits. Great strides have been made, with complex networks designed specifically to learn part-level discriminate feature representations. In this paper, we show that it is possible to cultivate subtle details without the need for overly complicated network designs or training mechanisms - a single loss is all it takes. The main trick lies with how we delve into individual feature channels early on, as opposed to the convention of starting from a consolidated feature map. The proposed loss function, termed as mutual-channel loss (MC-Loss), consists of two channel-specific components: a discriminality component and a diversity component. The discriminality component forces all feature channels belonging to the same class to be discriminative, through a novel channel-wise attention mechanism. The diversity component additionally constraints channels so that they become mutually exclusive across the spatial dimension. The end result is therefore a set of feature channels, each of which reflects different locally discriminative regions for a specific class. The MC-Loss can be trained end-to-end, without the need for any bounding-box/part annotations, and yields highly discriminative regions during inference. Experimental results show our MC-Loss when implemented on top of common base networks can achieve state-of-the-art performance on all four fine-grained categorization datasets (CUB-Birds, FGVC-Aircraft, Flowers-102, and Stanford Cars). Ablative studies further demonstrate the superiority of the MC-Loss when compared with other recently proposed general-purpose losses for visual classification, on two different base networks.
Dongliang Chang, Jiyang Xie 0001, Ayan Kumar Bhunia, Zhanyu Ma, Ming Wu 0001, Jun Guo 0002, Yi-Zhe Song
IEEE Trans. Image Process.7
2019 Multi-scale Deep Residual Network for Satellite Image Super-Resolution Reconstruction
Ming Wu 0001
PRCV (3)3
2017 Sketch-based cross-domain image retrieval via heterogeneous network
abstract
The development of image diversity has led multiple fields' application of cross-domain image retrieval. In this paper, we propose a heterogeneous dual network (two different networks) based on sketches and images. End-to-end cross-domain image retrieval is realized by limiting the similarity of features extracted from the two networks through the method of combining the contrastive loss with the triplet ranking loss. We also study how the order of drawings affects the sketch retrieval. Compared to the Siamese network, we avoid edge extraction for preprocessing, and the effect of fine-grained retrieval is further enhanced.
Ming Wu 0001
VCIP3
2014 Influence inflation in online social networks
abstract
Online marketing exploits social influence to trigger chain-like cascades. However, recent practices actively employ agents to collaboratively inflate the spreading of influences. Through supporting structures, they help each other with false feedback and signals to attract other users in the spreading process and thus alter the spontaneous social dynamics. In this paper, we proposed a modeling framework to explain the mechanism of such operations and characterize the spreading dynamics. Model analytics and numerical simulations both showed a lifting in overall spreading influence. As empirical evidence, experiments on a large Weibo network revealed well-structured advertising groups that prominently amplified the influences of promoted commercials via meticulous cooperation in a core-peripheral structure. The inflation effect also brings new considerations into influence maximization problems. Based on our models, we solved the problem of maximizing inflated influence by optimizing the selection of agents under KKT conditions and their supporting structure using its submodular property.
Jianjun Xie, Ming Wu 0001, Yun Huang 0001
ASONAM3