VLDB 2026 Research / reviewers in the wild / expert
Ming-Wen Shao
dblp:97/6241 · also Mingwen Shao
· DBLP profile ↗
134ranked-venue papers
36as first author
110since 2021 · last 2026
0000-0001-7323-5896ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 88 · 24 first-author · 67 since 2021Graphics, computer vision, multimedia, augmented reality and games · 33 · 7 first-author · 33 since 2021Applied, interdisciplinary, general and emerging computing · 9 · 2 first-author · 9 since 2021Systems, architecture and hardware · 8 · 8 since 2021Databases, data management, data science and information retrieval · 6 · 3 first-author · 2 since 2021Theory of computation · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Anti-Avatar: Protect Against Unauthorized 3D Head Avatar Generation via Dual-Space DivergenceabstractHead avatar generation is facilitated to construct high-fidelity 3D virtual personas from a single portrait, but it also raises the risk of unauthorized personal avatars generation. Recent 2D portrait protection methods actively prevent malicious image generation by perturbing the identity features. However, there are two key limitations when directly applied to prevent 3D head avatar generation: 1) These methods neglect the inherent 3D geometric structure of portrait, thus failing to disrupt the modeling of 3D shapes or poses. 2) They focus only on identity offset and are unable to interfere with the overall appearance, resulting in excessive preservation of facial characteristics. To overcomes these limitations, we propose a 3D defense framework termed Anti-Avatar, tailored to protect against unauthorized 3D head avatar generation from a single portrait. Specifically, Anti-Avatar consists of two key designs: Geometric Disruption and Perceptual Confusion. The former disrupts the precise reconstruction of 3D structure by interfering with the estimation of geometric parameters, thus affecting the structural accuracy of the 3D avatar. Collaboratively, the latter confuses image features by dispersing attention distribution, thereby hindering the effective perception of portrait appearance. Benefiting from the above dual-space divergence in geometry and perception, the avatars generated by our protected portraits exhibit substantial discrepancies from the originals. Extensive experiments show that our Anti-Avatar outperforms 2D methods in protection performance and effectively resists reconstruction and manipulation by state-of-the-art 3D head avatar generation methods. Lingzhuang Meng, Ming-Wen Shao, Yuanjian Qiao 0001, Jie Zhang 0133 |
AAAI | 2 |
| 2026 | Optimization-based fast single-image three-dimensional clothed human reconstruction via Gaussian Splatting
Ming-Wen Shao |
Eng. Appl. Artif. Intell. | 2 |
| 2026 | 3D adversarial objects generation for wider-view face recognition attacks
Lingzhuang Meng, Ming-Wen Shao, Yuanjian Qiao 0001, Yecong Wan |
Knowl. Based Syst. | 2 |
| 2026 | ExposureGS: Illumination-aware Gaussian splatting for sparse-view 3D exposure correction
Yuanjian Qiao 0001, Ming-Wen Shao, Lingzhuang Meng, Yecong Wan |
Knowl. Based Syst. | 2 |
| 2026 | FreeMD: Training-free multi-domain text-to-image generation with any control
Ming-Wen Shao, Chang Liu 0115, Lingzhuang Meng, Yecong Wan, Zhengyi Gong |
Neural Networks | 1 |
| 2026 | FedCAD: Cross-modal semantic alignment and distillation for cross-domain heterogeneous federated learning
Huijun Yuan, Ming-Wen Shao |
Neural Networks | 3 |
| 2026 | Gaussian splitting attack: Gaussian splatting-based multi-view 3D adversarial attack
Lingzhuang Meng, Ming-Wen Shao, Yuanjian Qiao 0001 |
Pattern Recognit. | 2 |
| 2026 | Separating anything from image in context
Yecong Wan, Ming-Wen Shao, Yuanshuo Cheng, Deyu Meng, Wangmeng Zuo |
Pattern Recognit. | 2 |
| 2026 | Removing Multiple Hybrid Adverse Weather in Video via a Unified ModelabstractVideos captured under real-world adverse weather conditions typically suffer from uncertain hybrid weather efforts. However, existing algorithms can only remove one type of weather degradation at a time and deal with different weather conditions with separate models, thus may fail to handle real-world stochastic hybrid scenarios. Besides, the model training is also infeasible due to the lack of paired video data to characterize the coexistence of multiple weather. To ameliorate the aforementioned issue, we propose a novel unified model, dubbed UniWRV, to remove multiple heterogeneous video weather degradations in an all-in-one fashion. Specifically, to tackle degenerate spatial feature heterogeneity, we propose a tailored weather prior guided module that queries exclusive priors for different instances as prompts to steer spatial feature characterization. To tackle degenerate temporal feature heterogeneity, we propose a dynamic routing aggregation module that can automatically select optimal fusion paths for different instances to dynamically integrate temporal features. Furthermore, we propose a real-world adaptation training scheme that leverages CLIP priors to provide semantic supervision for unlabeled real-world weather-degraded videos, thereby enabling the model to better cope with the diverse and complex real-world weather conditions. Additionally, we managed to construct a new synthetic video dataset, termed HWVideo, for learning and benchmarking multiple hybrid adverse weather removal, which contains 15 hybrid weather conditions with a total of 1500 adverse-weather/clean paired video clips. Real-world hybrid weather videos are also collected to facilitate model generalizability. Comprehensive experiments demonstrate that our UniWRV exhibits robust and superior adaptation capability in multiple heterogeneous degradations learning scenarios, including various generic video restoration tasks beyond weather removal. Yecong Wan, Ming-Wen Shao, Yuanshuo Cheng, Shuigen Wang |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2025 | Physical-aware Neural Radiance Fields for Efficient Exposure CorrectionabstractNeural Radiance Fields (NeRF) has achieved remarkable success in synthesizing impressive novel views. However, existing methods usually fail to handle scenes with adverse lighting conditions caused by external time variations and different camera settings, leading to poor visual quality. To address this challenge, we propose a physical-aware NeRF for efficient exposure correction, named PHY-NeRF. Specifically, we design Adaptive Lighting Particles inspired by the theory of light scattering and absorption, which can adjust the illumination intensity during volume rendering. Subsequently, we can handle scenes with different lighting conditions by jointly optimizing camera parameters and these lighting particles. Moreover, to promote natural brightness transitions, we devise a global illumination consistency module to control the lighting intensity across views at the feature level while completing more details. Benefiting from the above designs, our PHY-NeRF can tackle arbitrary low-light or overexposed scenes in an unsupervised manner. Extensive experiments show that our PHY-NeRF achieves state-of-the-art results in addressing adverse lighting problems while ensuring high rendering efficiency. Ming-Wen Shao, Yuanjian Qiao 0001 |
AAAI | 2 |
| 2025 | RestorGS: Depth-aware Gaussian Splatting for Efficient 3D Scene Restorationabstract3D Gaussian Splatting (3DGS) has recently achieved remarkable progress in novel view synthesis. However, existing methods rely heavily on high-quality data for rendering and struggle to handle degraded scenes with multi-view inconsistency, leading to inferior rendering quality. To address this challenge, we propose a novel Depth-aware Gaussian Splatting for efficient 3D scene Restoration, called RestorGS, which flexibly restores multiple degraded scenes using a unified framework. Specifically, RestorGS consists of two core designs: Appearance Decoupling and Depth-Guided Modeling. The former exploits appearance learning over spherical harmonics to decouple clear and degraded Gaussian, thus separating the clear views from the degraded ones. Collaboratively, the latter leverages the depth information to guide the degradation modeling, thereby facilitating the decoupling process. Benefiting from the above optimization strategy, our method achieves high-quality restoration while enabling real-time rendering speed. Extensive experiments show that our RestorGS outperforms existing methods significantly in underwater, nighttime, and hazy scenes. Yuanjian Qiao 0001, Ming-Wen Shao, Lingzhuang Meng |
CVPR | 2 |
| 2025 | S2Gaussian: Sparse-View Super-Resolution 3D Gaussian SplattingabstractIn this paper, we aim ambitiously for a realistic yet challenging problem, namely, how to reconstruct high-quality 3D scenes from sparse low-resolution views that simultaneously suffer from deficient perspectives and clarity. Whereas existing methods only deal with either sparse views or low-resolution observations, they fail to handle such hybrid and complicated scenarios. To this end, we propose a novel Sparse-view Super-resolution 3D Gaussian Splatting framework, dubbed S2Gaussian, that can reconstruct structure-accurate and detail-faithful 3D scenes with only sparse and low-resolution views. The S2Gaussian operates in a two-stage fashion. In the first stage, we initially optimize a low-resolution Gaussian representation with depth regularization and densify it to initialize the high-resolution Gaussians through a tailored Gaussian Shuffle Split operation. In the second stage, we refine the high-resolution Gaussians with the super-resolved images generated from both original sparse views and pseudo-views rendered by the low-resolution Gaussians. In which a customized blur-free inconsistency modeling scheme and a 3D robust optimization strategy are elaborately designed to mitigate multi-view inconsistency and eliminate erroneous updates caused by imperfect supervision. Extensive experiments demonstrate superior results and in particular establishing new state-of-the-art performances with more consistent geometry and finer details. Project Page https://jeasco.github.io/S2Gaussian/. Yecong Wan, Ming-Wen Shao, Yuanshuo Cheng, Wangmeng Zuo |
CVPR | 2 |
| 2025 | SUV: Suppressing Undesired Video Content via Semantic Modulation Based on Text Embeddings
Ming-Wen Shao, Lingzhuang Meng, Chang Liu 0115, Yecong Wan |
ICCV | 2 |
| 2025 | Wave-Mambaad: Wavelet-Driven State Space Model for Multi-Class Unsupervised Anomaly Detection
Ming-Wen Shao |
ICCV | 2 |
| 2025 | MoE-based Mamba for Multi-scene Universal Remote Sensing Semantic SegmentationabstractRemote sensing semantic segmentation (RSSS) aims to achieve pixel-level classification of remote sensing imagery for land cover identification. However, most existing RSSS methods are tailored for single-scene tasks and lack generalizability across diverse scenes. Extending these models to multi-scene tasks often results in decreased accuracy, increased training time and computational demands. In this paper, we propose MoE-SegMamba, a universal model for efficient multi-scene RSSS. Specifically, we propose an efficient encoder based on the novel TMoESSM Block, which includes a 2D Selective Scan module for capturing global information and a Task-aware Mixture-of-Experts (TMoE) Block to address multi-scene segmentation challenges. To mitigate task interference, we introduce the MoE Guidance Instructions (MGI) module, which generates task-related instructions to assist the TMoE Block in reducing spatial-dimension interference and support the Instruction Channel Gating Module (ICGM) in mitigating channel-dimension interference. Experimental results demonstrate that our proposed MoE-SegMamba achieves State-of-the-Art performance across four semantic segmentation scenes. The code is available at https://github.com/quanquans931225/MoE-SegMamba. Jie Zhang 0133, Ming-Wen Shao, Xiaodong Tan 0002, Xiangyong Cao |
ICME | 2 |
| 2025 | Indirect Alignment and Relationship Preservation for Domain GeneralizationabstractDomain generalization (DG) aims to train models on multiple source domains to generalize effectively to unseen target domains, addressing performance degradation caused by domain shifts. Many existing methods rely on direct feature alignment, which disrupts natural sequence relationships, causes misalignment and feature distortion, and leads to overfitting, especially with significant domain gaps. To tackle these issues, we propose a novel DG approach with two key modules: the Sample Difference Keeping (SDK) module, which preserves natural sequence relationships to enhance feature diversity and separability, and the Sample Consistency Alignment (SCA) module, which achieves indirect alignment by modeling inter-class and inter-domain relationship consistencies. This approach mitigates overfitting and misalignment, ensuring adaptability to significant domain gaps. Extensive experiments demonstrate that our framework consistently outperforms state-of-the-art methods. Zixiong Li, Ming-Wen Shao |
IJCAI | 4 |
| 2025 | DEGauss: Defending Against Malicious 3D Editing for Gaussian Splattingabstract3D editing with Gaussian splatting is exciting in creating realistic content, but it also poses abuse risks for generating malicious 3D content. Existing 2D defense approaches mainly focus on adding perturbations to single image to resist malicious image editing. However, there remain two limitations when applied directly to 3D scenes: (1) These methods fail to reflect 3D spatial correlations, thus protecting ineffectively under multiple viewpoints. (2) Such pixel-level perturbation is easily eliminated during the iterations of 3D editing, leading to failure of protection. To address the above issues, we propose a novel Defense framework against malicious 3D Editing for Gaussian splatting (DEGauss) for robustly disrupting the trajectory of 3D editing in multi-views. Specifically, to enable the effectiveness of perturbation across various views, we devise a view-focal gradient fusion mechanism that dynamically emphasizes the contributions of the most challenging views to adaptively optimize 3D perturbations. Furthermore, we design a dual discrepancy optimization strategy that both maximize the semantic deviation and the edit direction deviation of the guidance conditions to stably disrupt the editing trajectory. Benefiting from the collaborative designs, our method achieves effective resistance to 3D editing from various views while preserving photorealistic rendering quality. Extensive experiments demonstrate that our DEGauss not only performs excellent defense in different scenes, but also exhibits strong generalization across various state-of-the-art 3D editing pipelines. Lingzhuang Meng, Ming-Wen Shao, Yuanjian Qiao 0001 |
NeurIPS | 2 |
| 2025 | Prompt-guided and degradation prior supervised transformer for adverse weather image restoration
Ming-Wen Shao, Lingzhuang Meng, Yuanjian Qiao 0001, Zhiyuan Bao |
Appl. Intell. | 2 |
| 2025 | FCAT-Diff: Flexible and Consistent Appearance Transfer Based on Training-free Diffusion Model
Zhengyi Gong, Ming-Wen Shao, Chang Liu 0115, Huan Liu 0012 |
Comput. Graph. | 2 |
| 2025 | MirrorDiff: Prompt redescription for zero-shot grounded text-to-image generation with attention modulation
Chang Liu 0115, Ming-Wen Shao, Zhengyi Gong, Lingzhuang Meng |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | Adaptive prompt guided unified image restoration with latent diffusion model
Ming-Wen Shao, Yecong Wan, Yuanjian Qiao 0001, Changzhong Wang |
Eng. Appl. Artif. Intell. | 2 |
| 2025 | Adaptive generative knowledge distillation for dense object detection with reused detectorabstractKnowledge distillation has shown its potential in the field of computer vision, significantly reducing model parameters and memory usage with minimal impact on model performance. Although current mainstream methods face issues related to inconsistencies between distillation targets and real targets, as well as insufficient learning of teacher features. To address these issues, this paper presents a novel knowledge distillation framework, termed ADKD, designed to mitigate the limitations of current methodologies. It consists of two modules: AGD (Adaptive Generative Distillation) and RHD (Reused Head Distillation). Through AGD, the teacher guides the learning of student network, enabling student to achieve stronger representational power. Meanwhile, RHD effectively addresses the discrepancies between real targets and distillation targets. By masking and reusing feature maps and utilizing the teacher detection head, a more effective object detection model is obtained. The proposed approach is straightforward to implement and effective, having undergone extensive experimentation on datasets PASCAL VOC and MS COCO to demonstrate its efficacy. The results indicate that our method outperforms other knowledge distillation techniques. Shuaiqing Wang, Ming-Wen Shao |
Intell. Data Anal. | 4 |
| 2025 | Training-free prior guided diffusion model for zero-reference low-light image enhancement
Kai Shang 0001, Ming-Wen Shao, Chao Wang 0102, Yuanjian Qiao 0001, Yecong Wan |
Neurocomputing | 2 |
| 2025 | Pool-mamba: Pooling state space model for low-light image enhancement
Ming-Wen Shao |
Neurocomputing | 2 |
| 2025 | Efficient spatio-temporal modeling and text-enhanced prototype for few-shot action recognition
Ming-Wen Shao |
Neurocomputing | 3 |
| 2025 | Controlling vision-language model for enhancing image restoration
Ming-Wen Shao, Qiwang Li, Lingzhuang Meng, Yecong Wan |
Image Vis. Comput. | 1 |
| 2025 | DFDW: Distribution-aware Filter and Dynamic Weight for open-mixed-domain Test-time adaptation
Ming-Wen Shao, Xun Shao, Lingzhuang Meng |
Image Vis. Comput. | 1 |
| 2025 | BM-Edit: Background retention and motion consistency for zero-shot video editing
Ming-Wen Shao, Yecong Wan, Yuanshuo Cheng, Lingzhuang Meng |
Knowl. Based Syst. | 2 |
| 2025 | Dual-level semantic collaboration and inference network for medical image report generation
Junsan Zhang, Yuxue Liu, Ming-Wen Shao, Chenglizhao Chen, Zixuan Wang 0012, Yao Wan 0001, Philip S. Yu |
Knowl. Based Syst. | 3 |
| 2025 | CLIP-SDMG:CLIP knowledge distillation based on semantic decoupling and mask generationabstractContrastive language–image pretraining (CLIP) achieves cross-modal semantic alignment via image text contrastive learning and delivers remarkable performance in zero-shot image text retrieval and image classification tasks. However, its massive parameter size restricts practical deployment. Moreover, current CLIP distillation methods rely on response distillation, which causes a knowledge capacity barrier and loss of fine-grained features in the student model, eventually affecting task accuracy. To address these issues, we propose a semantic feature distillation framework, CLIP-SDMG. We design a progressive global semantic loss that enables the student model to gradually understand the process by which the teacher model generates complete image and text responses along with the teacher’s underlying reasoning mechanism; through an exponential decay scheduling, a smooth transition is achieved for global semantic learning. Considering that global semantics are updated through interactions with local semantics, we further propose a collaborative local semantic distillation strategy. For shallow local semantics, a dynamic attention balance mechanism is adopted, wherein the student initially relies on the teacher’s attention to focus on key local semantic features and gradually transitions to independent exploration. For deep local semantics, we construct a dual-path mask generation strategy for the vision and text modalities: the vision branch integrates squeeze-and- excitation attention with adaptive residual reconstruction to enhance visual semantic representation, while the text branch incorporates masked language modeling to improve contextual reasoning. We propose a semantic feature distillation framework for CLIP that enables the student model to receive guidance from the teacher at a global semantic level, thereby learning to comprehend the process and underlying reasoning by which the teacher generates the final responses. To further enhance the student’s ability to capture fine-grained features, we introduce a local semantic collaborative distillation strategy that effectively alleviates the knowledge capacity barrier and mitigates the loss of fine-grained features in the student model. As a result, the student CLIP notably outperforms baseline methods in zero-shot cross-modal image–text retrieval and zero-shot image classification tasks. Zhicheng Si, Ming-Wen Shao |
Knowl. Based Syst. | 3 |
| 2025 | PDbDa: Prompt-tuned dual-branch framework for unsupervised domain adaptationabstract• PDbDa integrates prompt learning and dual-branch design for robust domain adaptation. • The foundation branch use hierarchical prompts design and prompt knowledge constraints to enhance class discriminability. • The adaptation branch uses feature libraries and domain-aware feature tuning to align source and target domains. • PDbDa consistently outperforms state-of-the-art UDA and prompt-tuning methods on three benchmarks. Large-scale vision-language models demonstrate excellent performance on downstream tasks. However, they face challenges in unsupervised domain adaptation, particularly when domain shifts and semantic loss occur. Although existing prompt learning methods can decouple task-specific semantics from general knowledge, they face two main issues: (i) ineffective coordination between task-specific and general knowledge, leading to poor class discriminability, and (ii) inability to adequately address the domain shift. To address these challenges, we propose PDbDa, a prompt-tuned dual-branch framework for unsupervised domain adaptation that jointly optimizes learnable prompt vectors. PDbDa introduces two key innovations: (i) a foundation branch with a prompt knowledge constraint to regularize task-specific and pretrained knowledge, addressing class discriminability, and (ii) an adaptation branch with a domain-aware feature tuning block to align source and target domain features, mitigating the domain shift. These two branches function synergistically, improving model accuracy and training efficiency. The experimental results on the Office-Home, VisDA-2017, and DomainNet datasets indicate that PDbDa outperforms conventional prompt tuning and UDA methods by an average of 2.2 %-3 %. Yurui Zhao, Ming-Wen Shao |
Knowl. Based Syst. | 3 |
| 2025 | Do-DETR: enhancing DETR training convergence with integrated denoising and RoI mechanism
Ming-Wen Shao |
Multim. Syst. | 4 |
| 2025 | DiffRA: universal restorative adversarial attack based on diffusion model
Ming-Wen Shao, Lingzhuang Meng, Huan Liu 0012, Xiaodong Tan 0002 |
Multim. Syst. | 1 |
| 2025 | Meta-prompt tuning for low-resource visual question answering
Ming-Wen Shao, Lingzhuang Meng, Xun Shao |
Multim. Syst. | 1 |
| 2025 | DS-Diff: a dual-stage network with degradation-aware and semantic-aware for adverse weather removal based on diffusion models
Ming-Wen Shao |
Multim. Syst. | 3 |
| 2025 | Cross-domain attention-guided domain adaptive method for image real rain removal
Yuexian Liu, Ming-Wen Shao, Yuanshuo Cheng, Yecong Wan, Minggui Han |
Multim. Tools Appl. | 2 |
| 2025 | Adaptive pseudo-label threshold for source-free domain adaptation
Ming-Wen Shao, Lixu Zhang |
Neural Comput. Appl. | 1 |
| 2025 | Degradation-Guided cross-consistent deep unfolding network for video restoration under diverse weathers
Yuanshuo Cheng, Ming-Wen Shao, Yecong Wan, Yuanjian Qiao 0001, Wangmeng Zuo, Deyu Meng |
Neural Networks | 2 |
| 2025 | SeBIR: Semantic-guided burst image restoration
Huan Liu 0012, Ming-Wen Shao, Yecong Wan, Yuexian Liu, Kai Shang 0001 |
Neural Networks | 2 |
| 2025 | Learning physical-aware diffusion priors for zero-shot restoration of scattering-affected images
Yuanjian Qiao 0001, Ming-Wen Shao, Lingzhuang Meng, Wangmeng Zuo |
Pattern Recognit. | 2 |
| 2025 | Learning to Restore Arbitrary Hybrid adverse weather Conditions in one go
Yecong Wan, Ming-Wen Shao, Yuanshuo Cheng, Yuexian Liu, Zhiyuan Bao |
Pattern Recognit. | 2 |
| 2025 | Adaptive Fuzzy Degradation Perception Based on CLIP Prior for All-in-One Image RestorationabstractDespite substantial progress, the existing all-in-one image restoration methods still lack the ability to adaptively sense and accurately represent degradation information, thus hindering the enhancement of restoration performance. In addition, due to the large uncertainty and fuzziness of the data distribution in real scenarios compared to the training data, the model's generalization ability is often limited. To address the above issues, we propose a novel adaptive fuzzy degradation perception approach based on fuzzy theory that includes two tactics: 1) Fuzzy Degradation Perceiver (FDP); and 2) Test-time Self-supervised Prompt Fine-tuning (TSPF). On the one hand, we introduce the FDP, which leverages the rich visual language prior knowledge in CLIP to learn the prompt representations of different degradations. These prompts are regarded as semantic representations of various degradation fuzzy sets, achieving adaptive degradation perception by computing the degrees of membership between input images and the fuzzy sets. On the other hand, we propose the TSPF strategy, which is capable of self-supervised optimization of degraded fuzzy sets according to real-world scenarios during testing. This strategy improves the model's ability to perceive and represent the degraded information in data with real-world distributions. Thanks to the above key strategies, our method significantly improves degradation perception capability and image restoration quality while exhibiting excellent generalization in complex real-world scenarios. Extensive experiments on multiple benchmark datasets confirm that our approach achieves state-of-the-art performance in all-in-one image restoration. Ming-Wen Shao, Yuexian Liu, Yuanshuo Cheng, Yecong Wan, Changzhong Wang |
IEEE Trans. Fuzzy Syst. | 1 |
| 2025 | PromptSeg: Prompt for Universal Remote Sensing Semantic Segmentation
Jie Zhang 0133, Ming-Wen Shao, Lingzhuang Meng, Xiangyong Cao, Shuigen Wang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2025 | LSA-MEP: layer-wise sparsity allocation multi-metric evaluation pruning
Quanyi Guo, Ming-Wen Shao |
J. Supercomput. | 3 |
| 2025 | Multi-level navigation network: advancing fine-grained visual classification
Ming-Wen Shao |
J. Supercomput. | 3 |
| 2025 | MAQT: multi-scale attention and query-optimized transformer for end-to-end pose estimation
Cuiping Wang, Ming-Wen Shao |
J. Supercomput. | 3 |
| 2025 | Multi-cue SORT: integrating weak cues with appearance and motion for multi-object tracking
Mingchen Xu, Ming-Wen Shao |
J. Supercomput. | 4 |
| 2025 | ReDepthNet: a radar and camera depth estimation model based on semantic segmentation mask region alignment
Ming-Wen Shao |
J. Supercomput. | 4 |
| 2025 | MEKF: long-tailed visual recognition via multiple experts with knowledge fusion
Chenghao Ji, Ming-Wen Shao |
J. Supercomput. | 3 |
| 2025 | An efficient video transformer network with token discard and keyframe enhancement for action recognition
Zuosui Yang, Ming-Wen Shao |
J. Supercomput. | 3 |
| 2025 | Latent Code Augmentation Based on Stable Diffusion for Data-Free Substitute AttacksabstractSince the training data of the target model is not available in the black-box substitute attack, most recent schemes utilize generative adversarial networks (GANs) to generate data for training the substitute model. However, these GANs-based schemes suffer from low training efficiency as the generator needs to be retrained for each target model during the substitute training process, as well as low generation quality. To overcome these limitations, we consider utilizing the diffusion model (DM) to generate data and propose a novel data-free substitute attack scheme based on stable diffusion (SD) to improve the efficiency and accuracy of substitute training. Despite the data generated by the SD exhibited high quality, it presented a different distribution of domains and a large variation of positive and negative samples for the target model. For this problem, we propose latent code augmentation (LCA) to facilitate SD in generating data that aligns with the data distribution of the target model. Specifically, we augment the latent codes of the inferred member data with LCA and use them as guidance for SD. With the guidance of LCA, the data generated by the SD not only meets the discriminative criteria of the target model but also exhibits high diversity. By utilizing this data, it is possible to train the substitute model that closely resembles the target model more efficiently. Extensive experiments demonstrate that our LCA achieves higher attack success rates (ASRs) and requires fewer query budgets compared to GANs-based schemes for different target models. Our codes are available at https://github.com/LzhMeng/LCA. Ming-Wen Shao, Lingzhuang Meng, Yuanjian Qiao 0001, Lixu Zhang, Wangmeng Zuo |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2025 | Frequency-Aware Uncertainty Gaussian Splatting for Dynamic Scene Reconstructionabstract3D Gaussian splatting has recently achieved remarkable progress in dynamic scene reconstruction. However, there remain two practical challenges: (1) Existing methods typically employ a strict point-wise deformation structure to model dynamic attributes, while neglecting the uncertain motion correlation in local space, leading to inferior adaptability to complex scenes. (2) The inherent low-frequency bias properties of Gaussians often lead to blurring artifacts due to the insufficient high-frequency learning of variable motions. To address these challenges, we propose a novel Frequency-aware Uncertainty Gaussian Splatting, termed FUGS, for adaptively reconstructing dynamic scenes in the Fourier space. Specifically, we design an Uncertainty-aware Deformation Model (UDM) that explicitly models motion attributes using learnable uncertainty relations with neighboring Gaussian points. Such a paradigm is capable of facilitating temporal and spatial motion correlation learning, thereby enabling flexible Gaussian deformations. Subsequently, a Dynamic Spectrum Regularization (DSR) is developed to perform coarse-to-fine Gaussian densification through low-to-high frequency filtering. By weighting the gradient with frequency distance, the Gaussian attribute is adaptively adjusted according to the scene complexity. Benefiting from the flexible optimization, our method achieves high-fidelity reconstruction of complex scenes while enjoying real-time rendering. Extensive experiments on synthetic and real-world datasets show that our FUGS exhibits significant superiority over state-of-the-art methods. The code will be available at https://github.com/KevinJoee/GS. Ming-Wen Shao, Yuanjian Qiao 0001, Kai Zhang 0029, Lingzhuang Meng |
IEEE Trans. Vis. Comput. Graph. | 1 |
| 2025 | Contrastive local constraint for irregular image reconstruction and editability
Qiwang Li, Ming-Wen Shao, Fukang Liu, Yuanjian Qiao 0001 |
Vis. Comput. | 2 |
| 2025 | Interacthand: Robust 3D Hand Mesh Reconstruction via Interaction-Aware Segmentation and Refinement
Ming-Wen Shao, Xiaolin Lu |
Vis. Comput. | 2 |
| 2024 | Multi-Domain Multi-Scale Diffusion Model for Low-Light Image EnhancementabstractDiffusion models have achieved remarkable progress in low-light image enhancement. However, there remain two practical limitations: (1) existing methods mainly focus on the spatial domain for the diffusion process, while neglecting the essential features in the frequency domain; (2) conventional patch-based sampling strategy inevitably leads to severe checkerboard artifacts due to the uneven overlapping. To address these limitations in one go, we propose a Multi-Domain Multi-Scale (MDMS) diffusion model for low-light image enhancement. In particular, we introduce a spatial-frequency fusion module to seamlessly integrates spatial and frequency information. By leveraging the Multi-Domain Learning (MDL) paradigm, our proposed model is endowed with the capability to adaptively facilitate noise distribution learning, thereby enhancing the quality of the generated images. Meanwhile, we propose a Multi-Scale Sampling (MSS) strategy that follows a divide-ensemble manner by merging the restored patches under different resolutions. Such a multi-scale learning paradigm explicitly derives patch information from different granularities, thus leading to smoother boundaries. Furthermore, we empirically adopt the Bright Channel Prior (BCP) which indicates natural statistical regularity as an additional restoration guidance. Experimental results on LOL and LOLv2 datasets demonstrate that our method achieves state-of-the-art performance for the low-light image enhancement task. Codes are available at https://github.com/Oliiveralien/MDMS. Kai Shang 0001, Ming-Wen Shao, Chao Wang 0102, Yuanshuo Cheng, Shuigen Wang |
AAAI | 2 |
| 2024 | Inter-Class Topology Alignment for Efficient Black-Box Substitute Attacks
Lingzhuang Meng, Ming-Wen Shao, Yuanjian Qiao 0001 |
ECCV (34) | 2 |
| 2024 | Frequency-aware network for low-light image enhancement
Kai Shang 0001, Ming-Wen Shao, Yuanjian Qiao 0001, Huan Liu 0012 |
Comput. Graph. | 2 |
| 2024 | Pseudo initialization based Few-Shot Class Incremental Learning
Ming-Wen Shao, Xinkai Zhuang, Lixu Zhang, Wangmeng Zuo |
Comput. Vis. Image Underst. | 1 |
| 2024 | When guided diffusion model meets zero-shot image super-resolution
Huan Liu 0012, Ming-Wen Shao, Kai Shang 0001, Yuanjian Qiao 0001, Shuigen Wang |
Eng. Appl. Artif. Intell. | 2 |
| 2024 | Mining positive and negative rules via one-sided fuzzy three-way concept lattices
Ming-Wen Shao, Ju-Sheng Mi, Weizhi Wu 0001 |
Fuzzy Sets Syst. | 2 |
| 2024 | Distance metric-based learning for long-tail object detection
Ming-Wen Shao, Zilu Peng |
Image Vis. Comput. | 1 |
| 2024 | RDM-IR: Task-adaptive deep unfolding network for All-In-One image restoration
Yuanshuo Cheng, Ming-Wen Shao, Yecong Wan, Chao Wang 0102 |
Knowl. Based Syst. | 2 |
| 2024 | PMFN-SSL: Self-supervised learning-based progressive multimodal fusion network for cancer diagnosis and prognosis
Hudan Pan, Yong Liang 0001, Ming-Wen Shao, Shengli Xie 0001, Shanghui Lu, Shuilin Liao |
Knowl. Based Syst. | 4 |
| 2024 | Image all-in-one adverse weather removal via dynamic model weights generation
Yecong Wan, Ming-Wen Shao, Yuanshuo Cheng, Wangmeng Zuo |
Knowl. Based Syst. | 2 |
| 2024 | Mutually guided learning of global semantics and local representations for image restoration
Yuanshuo Cheng, Ming-Wen Shao, Yecong Wan |
Multim. Tools Appl. | 2 |
| 2024 | Fuzzy-based cross-image pixel contrastive learning for compact medical image segmentation
Yecong Wan, Ming-Wen Shao, Yuanshuo Cheng, Weiping Ding 0001 |
Multim. Tools Appl. | 2 |
| 2024 | A dual progressive strategy for long-tailed visual recognition
Guoqing Cao, Ming-Wen Shao |
Mach. Vis. Appl. | 3 |
| 2024 | Learning Depth-Density Priors for Fourier-Based Unpaired Image RestorationabstractDeep learning-based image restoration methods trained on synthetic datasets have witnessed notable progress, but suffer from significant performance drops on real-world images due to huge domain shifts. To alleviate this issue, some recent methods strive to improve the generalization ability of models with unpaired training. However, these solutions typically handle each problem individually and ignore the shared physical properties of different harsh scenarios, i.e., heavy rain, hazy and low-light images degrade more densely with increasing scene depth. Such limitations make them generalize poorly to real-world images. In this paper, we propose a novel Physically Oriented Generative Adversarial Network (POGAN) for unpaired image restoration with depth-density priors. Specifically, our POGAN consists of two core designs: Physical Restoration Network (PRNet) and Degradation Rendering Network (DRNet). The former focuses on estimating the physical components related to the depth and density distribution for restoration, while the latter re-renders degradation effects guided by the estimated depth information. To further facilitate learning the above physical prior, we design a Spatial-Frequency Interaction Residual block (SFIR), which efficiently learns global frequency information and local spatial features in an interactive manner. Extensive experiments on synthetic and real-world datasets demonstrate the superiority of our method in heavy rain, haze, and low-light scenarios. Yuanjian Qiao 0001, Ming-Wen Shao, Leiquan Wang, Wangmeng Zuo |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Multidimensional Dynamic Pruning: Exploring Spatial and Channel Fuzzy SparsityabstractDynamic pruning is an effective model compression method to reduce the computational cost of networks. However, existing dynamic pruning methods are limited to pruning along a single dimension (channel, spatial or depth), which cannot maximally excavate the redundancy of the network. Meanwhile, most of the current state-of-the-arts usually implement dynamic pruning via masked-out partial channels and pixels for training, while failing to accelerate the inference speed. To tackle these limitations, we propose a novel fuzzy-based Multi-Dimensional Dynamic Pruning (MDDP) paradigm to dynamically compress neural networks along both the channel and spatial dimensions. Specifically, we design a multi-dimensional fuzzy-mask block to simultaneously learn which spatial positions or channels are redundant and need to be pruned. Then, the Gumbel-Softmax trick combined with a sparsity loss is introduced to train these mask modules in an end-to-end manner. During the testing stage, we convert features and convolution kernels into two matrices respectively, and then implement sparse convolution through matrix multiplication to accelerate the network inference. Extensive experiments demonstrate that our method outperforms existing methods in terms of accuracy and computational cost. For instance, on the CIFAR-10 dataset, our method prunes 68% FLOPs of ResNet-56 with only a 0.07% Top-1 accuracy drop Ming-Wen Shao, Jiandong Kuang, Chao Wang 0102, Wangmeng Zuo, Guoyin Wang 0001 |
IEEE Trans. Fuzzy Syst. | 1 |
| 2024 | Low-Rank Prompt-Guided Transformer for Hyperspectral Image DenoisingabstractHyperspectral image (HSI) denoising is an essential preprocessing step for downstream applications. Although vision transformer (ViT)-based approaches show impressive denoising performance through self-similarity modeling, these methods still fail to exploit spatial and spectral correlations while ensuring flexibility and efficacy. To address this issue, we propose a hyperspectral denoising transformer using low-rank prompt (HyLoRa), simultaneously taking the spatial self-similarity and spectral low-rank property into account for HSI denoising. Specifically, to fully utilize intrinsic similarity in spatial domain, we perform cross-shaped window-based spatial self-attention for effectively modeling local and global similarity. Moreover, to exploit low-rank inductive bias, we integrate a low-rank prompt module into attention calculation for counting corrected low-dimensional vectors from a large collection of HSIs. This helps to better refine underlying noise-free structure representations. Compared to existing works, powerful capabilities for modeling spatial and spectral correlations can be built to correct low-rank representation in the feature space. Extensive experiments on both simulated and real remote sensing noise demonstrate that our HyLoRa consistently surpasses the state-of-the-art methods. Xiaodong Tan 0002, Ming-Wen Shao, Yuanjian Qiao 0001, Tiyao Liu, Xiangyong Cao |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Summator-Subtractor Network: Modeling Spatial and Channel Differences for Change DetectionabstractThe field of remote sensing (RS) image change detection (CD) has made significant progress, largely due to the powerful feature representation abilities of deep learning. However, traditional methods have not fully exploited the valuable information in differences. These methods often treat deep models as tools to extract features from individual images, which limits their ability to effectively describe differences. Additionally, many approaches tend to focus on spatial differences, while neglecting variations in the channel dimension. In this study, we introduce a novel Summator–Subtractor network for CD (${S}^{2}$CD), which adeptly captures subtle differences within both the spatial and channel aspects of bi-temporal images. The initial spatial and channel differences are derived through summation and subtraction operations on the bi-temporal images. The summator computes initial channel variations, while the subtractor captures initial spatial disparities. Transformers are then used to pull out meaningful differences in both spatial and channel patterns, allowing for a more nuanced understanding than methods relying solely on features from individual images. Finally, a heterogeneous modulation block integrates channel and spatial difference features, thus amplifying overall differences. Through extensive experimentation on four widely acknowledged CD benchmark datasets, our proposed${S}^{2}$CD method outperforms existing techniques, showcasing its superior performance and promising potential. The codes of this work will be available for the sake of reproducibility at:https://github.com/qianday/SSCD-CD. Leiquan Wang, Ye Fang, Chunlei Wu, Mingming Xu 0001, Ming-Wen Shao |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2024 | Boundary-Aware Spatial and Frequency Dual-Domain Transformer for Remote Sensing Urban Images SegmentationabstractSemantic segmentation of remote sensing (RS) images refers to labeling each pixel with a class to identify objects or land cover types. Existing mainstream spatial-domain semantic segmentation methods are mainly categorized into convolutional neural network (CNN)-based and vision transformer (ViT)-based approaches. The former excels at capturing local features, while the latter is adept at extracting global features. Several recent approaches consider combining CNN and ViT to efficiently capture local and global features. However, these approaches still struggle to capture complete features of the RS images, resulting in inaccurate segmentation. To address this issue, we introduce the fast Fourier transform (FFT), which transforms images into the frequency domain for feature extraction, acquiring the image-size receptive field that can complement spatial-domain methods. Based on this, we propose a boundary-aware spatial and frequency dual-domain transformer, termed dual-domain transformer. Specifically, our dual-domain transformer incorporates a dual-domain mixer (DualM), where the spatial-domain branch combines depthwise convolution and the attention mechanism to extract local and global features effectively, while the frequency-domain branch uses FFT to extract image-size features. The two branches complement each other, enabling a more comprehensive feature extraction of RS images. Meanwhile, a boundary-guided training strategy utilizing a boundary-aware module (BAM) is devised to constrain the model extract and predict boundary detail texture, which is an auxiliary task. In addition, the decoder incorporates a scale-feature fusion module (SFM) for adaptive information fusion between the encoder and decoder. Comprehensive experiments on the Zeebrugge and ISPRS datasets, including Vaihingen and Potsdam, showcase that the dual-domain transformer significantly outperforms state-of-the-art (SOTA) methods. Jie Zhang 0133, Ming-Wen Shao, Yecong Wan, Lingzhuang Meng, Xiangyong Cao, Shuigen Wang |
IEEE Trans. Geosci. Remote. Sens. | 2 |
| 2024 | Advancing Few-Shot Black-Box Attack With Alternating TrainingabstractConvolutional neural networks (CNNs) are known to be vulnerable to adversarial examples even in black-box scenarios, posing a significant threat to their reliability and security. Most existing black-box attack methods primarily focus on data-free scenarios, which often require a large number of queries and yield low attack success rates. But in practical applications, it is feasible to collect a small amount of data associated with the target network. In light of this, in this article, we propose an advancing few-shot black-box attack with alternating training scheme using few data and alternating training to improve the efficiency and attack success rate. Specifically, we propose an alternating training approach consisting of two parts, both aimed at optimizing the substitute network, which alternate and reinforce each other, leading to a significant reduction in the query budget required for a successful attack. In addition, we propose an image degradation (ID) module that expands the data volume diversity through ID techniques to mitigate the problem of generator overfitting. Furthermore, we design a model specific adapter to enable the substitute networks to dynamically adjust the parameters for different target networks. Extensive experiments demonstrate the efficacy of our approach in significantly reducing the query budget while achieving higher attack success rates compared to state-of-the-art competitors. Lingzhuang Meng, Ming-Wen Shao, Yuanjian Qiao 0001, Zhaofei Xu |
IEEE Trans. Reliab. | 2 |
| 2024 | Image-to-image translation using an offset-based multi-scale codes GAN encoder
Ming-Wen Shao, Shunhang Li |
Vis. Comput. | 2 |
| 2024 | Frequency domain-enhanced transformer for single image deraining
Ming-Wen Shao, Zhiyuan Bao, Yuanjian Qiao 0001, Yecong Wan |
Vis. Comput. | 1 |
| 2023 | An Efficient Frequency Domain Separation Network for Paired and Unpaired Image Super-ResolutionabstractAlthough existing super-resolution (SR) techniques have made great progress, they are often tailored for either paired or unpaired scenery, thus may result in poor migration ability. In this work, we propose a generalized Frequency Domain Separation Network (FDSNet) for both paired and unpaired SR settings. Firstly, through statistical analysis, we found that real-world low-resolution (LR) images and high-resolution (HR) images differ greatly in high frequencies but less in low frequencies. Inspired by this, we perform high and low-frequency separation of LR images and guide our model to reconstruct the HR contents in the different frequency domains. Then, according to the varying attention on frequencies of traditional CNN and Transformer models, we design a parallel pipeline: LFNet based on Transformer for low-frequency feature extraction, and HFNet based on CNN for high frequencies. In LFNet, to further alleviate the high complexity and data dependency of Transformer, Simplified Multi-head Self Attention (SMSA) is proposed at a low computational cost. And original MLP is replaced by our Spatial Enhancement MLP (SEMLP) to take full advantage of local spatial contexts. Finally, to further facilitate frequency separation and learning, a Frequency attention block is designed to impose guidance on high frequencies. Experiments indicate that our FDSNet achieves promising performance in terms of quantitative and qualitative evaluations while enjoying a faster speed and much fewer parameters. Huan Liu 0012, Ming-Wen Shao, Yuanjian Qiao 0001, Fukang Liu |
IJCNN | 2 |
| 2023 | MSLANet: multi-scale long attention network for skin lesion classification
Yecong Wan, Yuanshuo Cheng, Ming-Wen Shao |
Appl. Intell. | 3 |
| 2023 | High-fidelity GAN inversion by frequency domain guidance
Fukang Liu, Ming-Wen Shao, Lixu Zhang |
Comput. Graph. | 2 |
| 2023 | Mutual channel prior guided dual-domain interaction network for single image raindrop removal
Yuanjian Qiao 0001, Ming-Wen Shao, Huan Liu 0012, Kai Shang 0001 |
Comput. Graph. | 2 |
| 2023 | Hairstyle transfer via manipulating decoupled latent codes of StyleGAN2
Ming-Wen Shao, Fukang Liu, Yuanjian Qiao 0001 |
Comput. Graph. | 1 |
| 2023 | Progressive convolutional transformer for image restoration
Yecong Wan, Ming-Wen Shao, Yuanshuo Cheng, Deyu Meng, Wangmeng Zuo |
Eng. Appl. Artif. Intell. | 2 |
| 2023 | Meta-BN Net for few-shot learning
Ming-Wen Shao, Xinkai Zhuang |
Frontiers Comput. Sci. | 2 |
| 2023 | Graph neural networks induced by concept lattices for classification
Ming-Wen Shao, Weizhi Wu 0001, Huan Liu 0012 |
Int. J. Approx. Reason. | 1 |
| 2023 | Multi-domain clustering pruning: Exploring space and frequency similarity based on GAN
Junsan Zhang, Yeqi Feng, Chao Wang 0093, Ming-Wen Shao, Jian Wang 0010 |
Neurocomputing | 4 |
| 2023 | A novel fuzzy hierarchical fusion attention convolution neural network for medical image super-resolution reconstruction
Changzhong Wang, Ming-Wen Shao |
Inf. Sci. | 3 |
| 2023 | Uncertainty-guided hierarchical frequency domain Transformer for image restoration
Ming-Wen Shao, Yuanjian Qiao 0001, Deyu Meng, Wangmeng Zuo |
Knowl. Based Syst. | 1 |
| 2023 | Two-stage structural information enhancement for source-free domain adaptation
Ming-Wen Shao, Lixu Zhang, Zhiyuan Bao |
Mach. Vis. Appl. | 2 |
| 2023 | Improving knowledge distillation via pseudo-multi-teacher network
Shunhang Li, Ming-Wen Shao, Xinkai Zhuang |
Mach. Vis. Appl. | 2 |
| 2023 | Adaptive one-stage generative adversarial network for unpaired image super-resolution
Ming-Wen Shao, Huan Liu 0012, Jianxin Yang, Feilong Cao |
Neural Comput. Appl. | 1 |
| 2023 | FDDN: frequency-guided network for single image dehazing
Haozhen Shen, Chao Wang 0102, Liang-Jian Deng, Liangtian He, Ming-Wen Shao, Deyu Meng |
Neural Comput. Appl. | 6 |
| 2023 | Image Super-Resolution Using a Simple Transformer Without Pretraining
Huan Liu 0012, Ming-Wen Shao, Chao Wang 0102, Feilong Cao |
Neural Process. Lett. | 2 |
| 2023 | Adversarial-Based Ensemble Feature Knowledge Distillation
Ming-Wen Shao, Shunhang Li, Zilu Peng, Yuantao Sun |
Neural Process. Lett. | 1 |
| 2023 | Global-local transformer for single-image rain removal
Yecong Wan, Ming-Wen Shao, Zhi-Yuan Bao, Yuanshuo Cheng |
Pattern Anal. Appl. | 2 |
| 2023 | Unpaired image super-resolution using a lightweight invertible neural network
Huan Liu 0012, Ming-Wen Shao, Yuanjian Qiao 0001, Yecong Wan, Deyu Meng |
Pattern Recognit. | 2 |
| 2023 | Deep Fuzzy Clustering Transformer: Learning the General Property of Corruptions for Degradation-Agnostic Multitask Image RestorationabstractFor the sake of eliminating multiple degradations, most existing multitask image restoration methods prefer to learn the properties of each degradation type, which is often accompanied by a bloated model size and a heavy learning burden. To tackle the aforementioned issues, in this article, we propose to treat multiple degradations uniformly to achieve degradation type-agnostic multitask image restoration. We observe that the degradations in different spatial locations are always morphologically similar while the background sceneries vary greatly. In accordance with the aforementioned observation, we decouple the degradation features and the background features by an efficient fuzzy clustering method. The degradation features contain all the diverse degradation information, while the images are recovered from the decoupled background features. In practice, we discover a uniformity between the fuzzy C-means algorithm and cross attention and propose a deep fuzzy clustering transformer to achieve degradation type-agnostic background extraction via feature map clustering based on spatial distribution characteristics. Furthermore, to capture the spatial distribution properties of an image, an efficient global attention tree (GAT) is devised to provide a global spatial receptive field for the clustering process. By virtue of the quadtree structure, the proposed GATs enable more efficient global modeling than existing methods. Our experimental analysis showed that the proposed method outperformed the state-of-the-art models in terms of both efficiency and performance. Yuanshuo Cheng, Ming-Wen Shao, Yecong Wan, Yue-Xian Liu, Huan Liu 0012, Deyu Meng |
IEEE Trans. Fuzzy Syst. | 2 |
| 2023 | Eliminating Spatial Correlations of Anomaly: Corner-Visible Network for Unsupervised Hyperspectral Anomaly DetectionabstractHyperspectral anomaly detection (HAD) is crucial for identifying and analyzing abnormal objects in various domains. While existing methods have shown promising results by designing detection methods tailored to specific anomaly characteristics, there is a need for a highly versatile approach that can effectively handle anomalies, particularly those with large spatial sizes. In this article, we propose an end-to-end corner-visible network (CVNet) for unsupervised HAD. Specifically, we introduce a corner-visible convolution that leverages the statistical dependencies of the background within the receptive field while eliminating spatial correlations with potential anomalies for background generation. To address the grid effect caused by the corner-visible convolution, a background smoothing module is employed by using the conventional convolution. Furthermore, adaptive mean-squared error (MSE) and structural similarity index (SSIM) losses are employed to suppress anomaly reconstruction, resulting in a reliable reconstructed background map. Anomalies are identified through the residual of the original hyperspectral image (HSI) and the reconstructed background. Extensive experiments conducted on three publicly datasets demonstrate the effectiveness of our proposed method in handling different types of anomalies. The state-of-the-art performance showcases the versatility and applicability of CVNet in HAD. The codes of this work will be available for the sake of reproducibility athttps://github.com/Cloudynewbee/CVNet-HAD. Leiquan Wang, Chunlei Wu, Mingming Xu 0001, Ming-Wen Shao |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2023 | FGPGAN: a finer-grained CNN pruning via generative adversarial network
Shaoshuai Han, Ming-Wen Shao |
J. Supercomput. | 3 |
| 2023 | Inverse transformation sampling-based attentive cutout for fine-grained visual recognition
Yaojin Lin, Meiyan Xu, Ming-Wen Shao, Junfeng Yao |
Vis. Comput. | 4 |
| 2023 | Branch aware assignment for object detection
Ming-Wen Shao, Bingbing Fan |
Vis. Comput. | 1 |
| 2022 | From the whole to detail: Progressively sampling discriminative parts for fine-grained recognition
Yaojin Lin, Shengyu Chen, Zhichun Zeng, Ming-Wen Shao, Shaozi Li |
Knowl. Based Syst. | 5 |
| 2022 | Image rain removal and illumination enhancement done in one go
Yecong Wan, Yuanshuo Cheng, Ming-Wen Shao, Jordi Gonzàlez 0001 |
Knowl. Based Syst. | 3 |
| 2022 | Context-Based Multiscale Unified Network for Missing Data Reconstruction in Remote Sensing ImagesabstractMissing data reconstruction is a classical yet challenging problem in remote sensing images. Most current methods based on traditional convolutional neural network require supplementary data and can only handle one specific task. To address these limitations, we propose a novel generative adversarial network-based missing data reconstruction method in this letter, which is capable of various reconstruction tasks given only single source data as input. Two auxiliary patch-based discriminators are deployed to impose additional constraints on the local and global regions, respectively. In order to better fit the nature of remote sensing images, we introduce special convolutions and attention mechanism in a two-stage generator, thereby benefiting the tradeoff between accuracy and efficiency. Combining with perceptual and multiscale adversarial losses, the proposed model can produce coherent structure with better details. Qualitative and quantitative experiments demonstrate the uncompromising performance of the proposed model against multisource methods in generating visually plausible reconstruction results. Moreover, further exploration shows a promising way for the proposed model to utilize spatio-spectral-temporal information. The codes and models are available athttps://github.com/Oliiveralien/Inpainting-on-RSI. Ming-Wen Shao, Chao Wang 0102, Tianjun Wu, Deyu Meng, Jiancheng Luo |
IEEE Geosci. Remote. Sens. Lett. | 1 |
| 2022 | Dual-Pyramidal Image Inpainting With Dynamic NormalizationabstractDeep autoencoder-based approaches have achieved significant improvements on restoring damaged images, yet they still suffer from artifacts due to the inadequate representation and inaccurate regularization of existing features. In this paper, we propose a dual-pyramidal inpainting framework called DPNet to address these two limitations, which seamlessly integrates sufficient feature learning and dynamic regularization within an autoencoder network. Specifically, to exhaustively extract multi-scale features, we adopt layer-wise pyramidal convolution in encoder, which provides an arbitrary combination pool of various receptive fields. Subsequently, to tackle the patch deterioration problem in previous cross-scale non-local schemes, we further propose a Pyramidal Attention Mechanism (PAM) in decoder to acquire finer patches directly from learned layers. Mutually benefited with pyramidal features extraction in encoder, the dissemination space for non-local pixels in our PAM is notably enlarged to pyramidal level, thus significantly benefiting the feature representation. Moreover, to avoid the mask error accumulation in existing works, a dynamic normalization mechanism utilizing the spatial mask information updated in encoder is introduced, which further ensures the feature integrity and consistency. Such a dual-pyramidal structure along with dynamic normalization significantly improve the inpainting quality, outperforming existing competitors. Comprehensive experiments conducted on three benchmark datasets demonstrate that our DPNet performs favorably against the state-of-the-arts. Chao Wang 0102, Ming-Wen Shao, Deyu Meng, Wangmeng Zuo |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2022 | Efficient Pyramidal GAN for Versatile Missing Data Reconstruction in Remote Sensing ImagesabstractMissing data reconstruction is a classical yet challenging problem in remote sensing image processing due to the complex atmospheric environment and variability of satellite sensors. Most of the contemporary reconstruction methods either handle only one specific task or require supplementary data, while the single-input for multi-task reconstruction has not been explored yet. In this paper we propose a novel Generative Adversarial Network-based unified framework for missing remote sensing image reconstruction, which is capable of various reconstruction tasks given only single source data as input. Specifically, we first propose a Mask Extraction Network (MEN) to obtain a united soft mask, which represents the intrinsic prior under various scenarios and indicates not only location but context information. The versatility of mask extraction enables the multi-task reconstruction of remote sensing images. Besides, we propose a Unified Inpainting Network (UIN) to repair diverse degraded images. Being specifically tailored for remote sensing images, Dilated pyramidal convolutions (DPC) and an Attention Fusion Mechanism (AFM) are introduced to further improve the feature extraction ability and thus exhaustly leveraging the single-input information. Extensive experiments demonstrate the uncompromising performance of the proposed method against state-of-the-art multi-input methods on diverse missing restoration. Moreover, further exploration shows the potential of the proposed method to utilize joint spatio-spectral-temporal information, which is evaluated to outperform existing competitors on remote sense images. Ming-Wen Shao, Chao Wang 0102, Wangmeng Zuo, Deyu Meng |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2021 | Selective generative adversarial network for raindrop removal from a single image
Ming-Wen Shao, Hong Wang 0021, Deyu Meng |
Neurocomputing | 1 |
| 2021 | Target attack on biomedical image segmentation model based on multi-scale gradients
Ming-Wen Shao, Gaozhi Zhang, Wangmeng Zuo, Deyu Meng |
Inf. Sci. | 1 |
| 2021 | IIT-GAT: Instance-level image transformation via unsupervised generative attention networks with disentangled representations
Ming-Wen Shao, Youcai Zhang, Wangmeng Zuo, Deyu Meng |
Knowl. Based Syst. | 1 |
| 2021 | DMDIT: Diverse multi-domain image-to-image translation
Ming-Wen Shao, Youcai Zhang, Huan Liu 0012, Chao Wang 0102, Xun Shao |
Knowl. Based Syst. | 1 |
| 2021 | EAA-Net: A novel edge assisted attention network for single image dehazing
Chao Wang 0008, Haozhen Shen, Ming-Wen Shao, Chuan-Sheng Yang, Jiancheng Luo, Liang-Jian Deng |
Knowl. Based Syst. | 4 |
| 2021 | Uncertainty Guided Multi-Scale Attention Network for Raindrop Removal From a Single ImageabstractRaindrops adhered to a glass window or camera lens appear in various blurring degrees and resolutions due to the difference in the degrees of raindrops aggregation. The removal of raindrops from a rainy image remains a challenging task because of the density and diversity of raindrops. The abundant location and blur level information are strong prior guide to the task of raindrop removal. However, existing methods use a binary mask to locate and estimate the raindrop with the value 1 (adhesion of raindrops) and 0 (no adhesion), which ignores the diversity of raindrops. Meanwhile, it is noticed that different scale versions of a rainy image have similar raindrop patterns, which makes it possible to employ such complementary information to represent raindrops. In this work, we first propose a soft mask with the value in [-1,1] indicating the blurring level of the raindrops on the background, and explore the positive effect of the blur degree attribute of raindrops on the task of raindrop removal. Secondly, we explore the multi-scale fusion representation for raindrops based on the deep features of the input multi-scale images. The framework is termed uncertainty guided multi-scale attention network (UMAN). Specifically, we construct a multi-scale pyramid structure and introduce an iterative mechanism to extract blur-level information about raindrops to guide the removal of raindrops at different scales. We further introduce the attention mechanism to fuse the input image with the blur-level information, which will highlight raindrop information and reduce the effects of redundant noise. Our proposed method is extensively evaluated on several benchmark datasets and obtains convincing results. Ming-Wen Shao, Deyu Meng, Wangmeng Zuo |
IEEE Trans. Image Process. | 1 |
| 2020 | Multi-Attention Generative Adversarial Network for image captioning
Leiquan Wang, Haiwen Cao, Ming-Wen Shao, Chunlei Wu |
Neurocomputing | 4 |
| 2020 | Knowledge reduction methods of covering approximate spaces based on concept lattice
Ming-Wen Shao, Weizhi Wu 0001, Xizhao Wang, Changzhong Wang |
Knowl. Based Syst. | 1 |
| 2020 | Multi-scale generative adversarial inpainting network based on cross-layer attention transfer mechanism
Ming-Wen Shao, Wangmeng Zuo, Deyu Meng |
Knowl. Based Syst. | 1 |
| 2020 | Feature Selection Based on Neighborhood Self-InformationabstractThe concept of dependency in a neighborhood rough set model is an important evaluation function for the feature selection. This function considers only the classification information contained in the lower approximation of the decision while ignoring the upper approximation. In this paper, we construct a class of uncertainty measures: decision self-information for the feature selection. These measures take into account the uncertainty information in the lower and the upper approximations. The relationships between these measures and their properties are discussed in detail. It is proven that the fourth measure, called relative neighborhood self-information, is better for feature selection than the other measures, because not only does it consider both the lower and the upper approximations but also the change of its magnitude is largest with the variation of feature subsets. This helps to facilitate the selection of optimal feature subsets. Finally, a greedy algorithm for feature selection has been designed and a series of numerical experiments was carried out to verify the effectiveness of the proposed algorithm. The experimental results show that the proposed algorithm often chooses fewer features and improves the classification accuracy in most cases. Changzhong Wang, Yang Huang 0009, Ming-Wen Shao, Qinghua Hu, Degang Chen 0002 |
IEEE Trans. Cybern. | 3 |
| 2020 | Fuzzy Rough Attribute Reduction for Categorical DataabstractClassical rough set theory is considered a useful tool for dealing with the uncertainty of categorical data. The major deficiency of this model is that the classical rough set model is sensitive to noise in classification learning due to the stringent condition of equivalence relation. Thus, a class of fuzzy similarity relations was introduced to describe the similarity between samples with categorical attributes. However, these kinds of similarity relations also have deficiencies when they are used in fuzzy rough computation. In this article, we propose a new fuzzy-rough-set model for categorical data by introducing a variable parameter to control the similarity of samples. This model employs the iterative computation strategy to define fuzzy rough approximations and dependence functions. It is proved that the proposed rough dependence function is monotonic. Finally, the proposed model is applied to the attribute reduction of categorical data. The experimental results indicate that the proposed model is more effective for categorical data than some existing algorithms. Changzhong Wang, Ming-Wen Shao, Degang Chen 0002 |
IEEE Trans. Fuzzy Syst. | 3 |
| 2019 | Uncertainty measures for general fuzzy relations
Changzhong Wang, Yang Huang 0009, Ming-Wen Shao, Degang Chen 0002 |
Fuzzy Sets Syst. | 3 |
| 2019 | Attribute reduction based on k-nearest neighborhood rough sets
Changzhong Wang, Yunpeng Shi, Xiaodong Fan, Ming-Wen Shao |
Int. J. Approx. Reason. | 4 |
| 2019 | Fuzzy rough set-based attribute reduction using distance measures
Changzhong Wang, Yang Huang 0009, Ming-Wen Shao, Xiaodong Fan |
Knowl. Based Syst. | 3 |
| 2019 | Homomorphism between ordered decision systems
Changzhong Wang, Yang Huang 0009, Xiaodong Fan, Ming-Wen Shao |
Soft Comput. | 4 |
| 2017 | Attribute reduction in generalized one-sided formal contexts
Ming-Wen Shao, Kewen Li 0002 |
Inf. Sci. | 1 |
| 2017 | A unified information measure for general binary relations
Changzhong Wang, Qiang He 0003, Ming-Wen Shao, Qinghua Hu |
Knowl. Based Syst. | 3 |
| 2017 | A Fitting Model for Feature Selection With Fuzzy Rough SetsabstractA fuzzy rough set is an important rough set model used for feature selection. It uses the fuzzy rough dependency as a criterion for feature selection. However, this model can merely maintain a maximal dependency function. It does not fit a given dataset well and cannot ideally describe the differences in sample classification. Therefore, in this study, we introduce a new model for handling this problem. First, we define the fuzzy decision of a sample using the concept of fuzzy neighborhood. Then, a parameterized fuzzy relation is introduced to characterize the fuzzy information granules, using which the fuzzy lower and upper approximations of a decision are reconstructed and a new fuzzy rough set model is introduced. This can guarantee that the membership degree of a sample to its own category reaches the maximal value. Furthermore, this approach can fit a given dataset and effectively prevents samples from being misclassified. Finally, we define the significance measure of a candidate attribute and design a greedy forward algorithm for feature selection. Twelve datasets selected from public data sources are used to compare the proposed algorithm with certain existing algorithms, and the experimental results show that the proposed reduction algorithm is more effective than classical fuzzy rough sets, especially for those datasets for which different categories exhibit a large degree of overlap. Changzhong Wang, Yali Qi, Ming-Wen Shao, Qinghua Hu, Degang Chen 0002, Yaojin Lin |
IEEE Trans. Fuzzy Syst. | 3 |
| 2016 | Axiomatic characterizations of (S, T)-fuzzy rough approximation operators
Weizhi Wu 0001, You-Hong Xu, Ming-Wen Shao, Guoyin Wang 0001 |
Inf. Sci. | 3 |
| 2016 | Granular reducts of formal fuzzy contexts
Ming-Wen Shao, Yee Leung, Xizhao Wang, Weizhi Wu 0001 |
Knowl. Based Syst. | 1 |
| 2016 | Feature subset selection based on fuzzy neighborhood rough sets
Changzhong Wang, Ming-Wen Shao, Qiang He 0003, Yali Qi |
Knowl. Based Syst. | 2 |
| 2015 | Knowledge reduction in formal fuzzy contexts
Ming-Wen Shao, Hong-Zhi Yang, Weizhi Wu 0001 |
Knowl. Based Syst. | 1 |
| 2014 | Rule acquisition and complexity reduction in formal decision contexts
Ming-Wen Shao, Yee Leung, Weizhi Wu 0001 |
Int. J. Approx. Reason. | 1 |
| 2014 | Relations between granular reduct and dominance reduct in formal contexts
Ming-Wen Shao, Yee Leung |
Knowl. Based Syst. | 1 |
| 2013 | Vector-based Attribute Reduction Method for Formal ContextsabstractAttribute reduction is one basic issue in knowledge discovery of information systems. In this paper, based on the object oriented concept lattice and classical concept lattice, the approach of attribute reduction for formal contexts is investigated. We consider attribute reduction and attribute characteristics from the perspective of linear dependence of vectors. We first introduce the notion of context matrix and the operations of corresponding column vectors, then present some judgment theorems of attribute reduction for formal contexts. Furthermore, we propose a new method to reducing formal context and show corresponding reduction algorithms. Compared with previous reduction approaches which employ discernibility matrix and discernibility function to determine all reducts, the proposed approach is more simpler and easier to implement. Ming-Wen Shao, Min Liu 0013 |
Fundam. Informaticae | 1 |
| 2013 | Generalized fuzzy rough approximation operators determined by fuzzy implicators
Weizhi Wu 0001, Yee Leung, Ming-Wen Shao |
Int. J. Approx. Reason. | 3 |
| 2011 | Rule acquisition and attribute reduction in real decision formal contexts
Hong-Zhi Yang, Yee Leung, Ming-Wen Shao |
Soft Comput. | 3 |
| 2007 | Set approximations in fuzzy formal concept analysis
Ming-Wen Shao, Min Liu 0013, Wen-Xiu Zhang |
Fuzzy Sets Syst. | 1 |
| 2006 | Knowledge Reduction Based on Evidence Reasoning Theory in Ordered Information Systems
Weihua Xu 0003, Ming-Wen Shao, Wen-Xiu Zhang |
KSEM | 2 |
| 2005 | Dominance relation and rules in an incomplete ordered information systemabstractRough sets theory has proved to be a useful mathematical tool for classification and prediction. However, as many real-world problems deal with ordering objects instead of classifying objects, one of the extensions of the classical rough sets approach is the dominance-based rough sets approach, which is mainly based on substitution of the indiscernibility relation by a dominance relation. In this article, we present a dominance-based rough sets approach to reasoning in incomplete ordered information systems. The approach shows how to find decision rules directly from an incomplete ordered decision table. We propose a reduction of knowledge that eliminates only that information that is not essential from the point of view of the ordering of objects or decision rules. © 2005 Wiley Periodicals, Inc. Int J Int Syst 20: 13–27, 2005. Ming-Wen Shao, Wen-Xiu Zhang |
Int. J. Intell. Syst. | 1 |