Simiao Wang

dblp:219/5763 · DBLP profile ↗
← Back
24ranked-venue papers
4as first author
21since 2021 · last 2026
0009-0001-4210-7960ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 14 · 2 first-author · 13 since 2021Artificial intelligence and machine learning · 6 · 1 first-author · 6 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Databases, data management, data science and information retrieval · 2 · 1 since 2021
YearPublicationVenuePosition
2026 MambaGen: Efficient visual representation learning for automatic radiology report generation
Xiaodi Hou 0001, Xiaobo Li 0007, Simiao Wang, Mingyu Lu, Hongfei Lin, Yi-Jia Zhang 0001
Expert Syst. Appl.3
2026 Latent diffusion-augmented cross-modal representation learning for radiology report generation
Xiaodi Hou 0001, Xiaobo Li 0007, Simiao Wang, Mingyu Lu, Hongfei Lin, Yi-Jia Zhang 0001
Inf. Process. Manag.3
2026 Debiased medication recommendation through fusing frequent pattern and temporal medical records
Xiaobo Li 0007, Xiaodi Hou 0001, Simiao Wang, Shilong Wang 0004, Xiaokun Zhang 0001, Yi-Jia Zhang 0001
Neural Networks3
2026 Frequency-Enhanced Feature Pyramid Network With Global Saliency Kernel Module for Infrared Small Target Detection
Simiao Wang, Yunan Liu 0001, Mingyu Lu
IEEE Signal Process. Lett.1
2025 Weakly Supervised Object Detection Framework based on Classification-Localization Consistency
abstract
The inconsistency between classification and localization brings a challenge in object detection. Fully supervised object detection (FSOD) benefits from bounding-box regression networks to alleviate that, which is absent in weakly supervised object detection (WSOD). Consequently, there is a significant performance gap between the two paradigms. To bridge the performance and technical gaps between WSOD and FSOD, this paper proposes a novel weakly supervised object detection framework based on classification-localization consistency. We propose Max Score Pooling (MSP), which compels features relevant to classification to align with features relevant to localization, thereby achieving consistency between classification and localization. Additionally, we propose a Proposal Fusion Mechanism (PFM) to generate pseudo-supervision for training the bounding box regression network, further reducing the impact of classification-localization inconsistency. Extensive experiments are conducted on PASCAL VOC 2007, PASCAL VOC 2012 and MS COCO 2017 datasets, demonstrating our framework’s superior performance.
Yihuan Zhu, Simiao Wang, Mingyu Lu, Zhengxing Sun
ICME2
2025 RRG-Mamba: Efficient Radiology Report Generation with State Space Model
abstract
Recent advancements in radiology report generation have utilized deep neural networks such as CNNs and Transformers, achieving notable improvements in generating accurate and detailed reports. However, their practical adoption is hindered by the challenge of balancing global dependency modeling with computational efficiency. The state space model, particularly its enhanced variant Mamba, offers promising linear-complexity solutions for long-range dependency modeling. Despite its strengths, Mamba’s fixed positional encoding limits its ability to effectively capture complex spatial dependencies. To address this gap, we propose RRG-Mamba, an advanced framework for efficient radiology report generation. Within the RRGMamba, we enhance the vanilla Mamba by integrating rotary position encoding (RoPE), enabling dynamic modeling of relative positional information in visual feature sequences. Furthermore, we design a global dependency learning module to optimize long-range visual feature sequence modeling. Extensive experiments on publicly available datasets, including IU X-Ray and MIMIC-CXR, demonstrate that RRG-Mamba achieves a 3.7% improvement in BLEU-4 score over existing models, along with significant gains in computational and memory efficiency. Our code is available at https://github.com/Eleanorhxd/RRG-Mamba.
Xiaodi Hou 0001, Xiaobo Li 0007, Mingyu Lu, Simiao Wang, Yi-Jia Zhang 0001
IJCAI4
2025 Mask-Guided Cross-Modality Fusion Network for Visible-Infrared Vehicle Detection
abstract
Drone-based vehicle detection is crucial for intelligent traffic management. However, current methods relying solely on single visible or infrared modalities struggle with precision and robustness, especially in adverse weather conditions. The effective integration of cross-modal information to enhance vehicle detection still poses significant challenges. In this letter, we propose a masked-guided cross-modality fusion method, called MCMF, for robust and accurate visible-infrared vehicle detection. Firstly, we construct a framework consisting of three branches, with two dedicated to the visible and infrared modalities respectively, and another tailored for the fused multi-modal. Secondly, we introduce a Location-Sensitive Masked AutoEncoder (LMAE) for intermediate-level feature fusion. Specifically, our LMAE utilizes masks to cover intermediate-level features of one modality based on the prediction hierarchy of another modality, and then distills cross-modality guidance information through regularization constraints. This strategy, through a self-learning paradigm, effectively preserves the useful information from both modalities while eliminating redundant information from each. Finally, the fused features are input into an uncertainty-based detection head to generate predictions for bounding boxes of vehicles. When evaluated on the DroneVehicle dataset, our MCIF reaches 71.42% w.r..t. mAP, outperforming an established baseline method by 7.42%. Ablation studies further demonstrate the effectiveness of our LMAE for visible-infrared fusion.
Lingyun Tian, Zilong Deng, Simiao Wang
IEEE Signal Process. Lett.5
2024 In-WSOD: Integrality Weakly Supervised Object Detection with Classification and Localization Consistency
Yihuan Zhu, Simiao Wang, Zhengxing Sun
ICONIP (7)2
2024 Meta-Learning Based Knowledge Distillation for Domain Adaptive Nighttime Segmentation
Simiao Wang, Yunan Liu 0001, Mingyu Lu
PRCV (2)3
2024 Latent domain knowledge distillation for nighttime semantic segmentation
Yunan Liu 0001, Simiao Wang, Chunpeng Wang 0001, Mingyu Lu
Eng. Appl. Artif. Intell.2
2024 Image encryption algorithm using multi-base diffusion and a new four-dimensional chaotic system
Simiao Wang, Baichao Sun, Baoxiang Du
Multim. Tools Appl.1
2024 Prior based Pyramid Residual Clique Network for human body image super-resolution
Simiao Wang, Mingyu Lu, Jinguang Sun
Pattern Recognit.1
2024 Complementary Masked-Guided Meta-Learning for Domain Adaptive Nighttime Segmentation
abstract
Semantic segmentation in nighttime scenes presents a significant challenge in autonomous driving. Unsupervised domain adaptation (UDA) offers an effective solution by learning domain-invariant features to transfer models from the source domain (daytime scenes) to the target domain (nighttime scenes). Many methods introduce a latent domain to reduce the difficulty of UDA. However, they often build only a single adaptation pair of “latent-to-target”, which limits the effectiveness of knowledge transfer across different domains. In this letter, we propose a Masked Guided Meta-Learning (MGML) framework for domain-adaptive nighttime semantic segmentation. Within the MGML framework, we explore two key issues: how to generate the latent domain, and how to leverage the latent domain to assist meta-learning in reducing domain discrepancy. For the first issue, we employ the fast Fourier transform along with a complementary masking strategy to generate masked latent images that resemble the target scenes in the latent domain without adding to the training burden. For the second issue, we nest a mask-based consistency constraint within a bi-level meta-learning framework, enabling cross-domain knowledge acquired from the pair of “source-to-latent” to enhance the “latent-to-target” adaptation. Experiments on benchmark datasets demonstrate that our MGML achieves state-of-the-art performance, demonstrating the effectiveness of our approach in nighttime semantic segmentation.
Ruiying Chen 0001, Yuming Bo, Panlong Wu, Simiao Wang, Yunan Liu 0001
IEEE Signal Process. Lett.4
2024 Hierarchical Noise-Tolerant Meta-Learning With Noisy Labels
abstract
Due to the detrimental impact of noisy labels on the generalization of deep neural networks, learning with noisy labels has become an important task in modern deep learning applications. Many previous efforts have mitigated this problem by either removing noisy samples or correcting labels. In this letter, we address this issue from a new perspective and empirically find that models trained with both clean and mislabeled samples exhibit distinguishable activation feature distributions. Building on this observation, we propose a novel meta-learning approach called the Hierarchical Noise-tolerant Meta-Learning (HNML) method, which involves a bi-level optimization comprising meta-training and meta-testing. In the meta-training stage, we incorporate consistency loss at the output prediction hierarchy to facilitate model adaptation to dynamically changing label noise. In the meta-testing stage, we extract activation feature distributions using class activation maps and propose a new mask-guided self-learning method to correct biases in the foreground regions. Through the bi-level optimization of HNML, we ensure that the model generates discriminative feature representations that are insensitive to noisy labels. When evaluated on both synthetic and real-world noisy datasets, our HNML method achieves significant improvements over previous state-of-the-art methods.
Jian Wang 0078, Yuntai Yang, Renlong Wang, Simiao Wang
IEEE Signal Process. Lett.5
2024 iPCa-Former: A Multi-Task Transformer Framework for Perceiving Incidental Prostate Cancer
abstract
Despite significant progress in medical image analysis using deep learning, predicting incidental prostate cancer (iPCa) remains challenging due to subtle differences in multiparametric magnetic resonance imaging (mpMRI) and a lower incidence rate. To address these challenges, we propose iPCa-Former, a transformer-based framework designed to enhance iPCa prediction within prostate mpMRI slices. Firstly, built on an encoder-decoder architecture, our iPCa-Former facilitates the simultaneous optimization of two tasks through mutual learning: prostate transition zone segmentation and iPCa prediction. Secondly, we introduce a joint optimization function that combines focal loss and boundary-based mutual information (BMI) loss, effectively addressing the imbalance of positive and negative samples in classification and the challenge posed by a small proportion of the foreground region in segmentation. Moreover, we construct an iPCa mpMRI dataset comprising 10,276 prostate mpMRI slices from 485 patients clinically diagnosed with benign prostatic hyperplasia, however, 27 out of these patients are identified as iPCa. When evaluated on this benchmark dataset, our iPCa-Former outperforms state-of-the-art methods, demonstrating the superior performance of our approach.
Xianwei Pan, Simiao Wang, Yunan Liu 0001, Lijie Wen 0002, Mingyu Lu
IEEE Signal Process. Lett.2
2024 Dual-Task Mutual Learning With QPHFM Watermarking for Deepfake Detection
abstract
Deepfake technology has rapidly evolved and emerged in recent years, posing significant threats to individuals' reputations and security. Although passive detection methods can achieve reasonable accuracy, they still lack proactive defense mechanisms. To address this issue, this letter proposes a proactive detection framework that combines Quaternion Polar Harmonic Fourier Moments (QPHFMs) with Dual-Task Mutual Learning (DTML) framework. Firstly, watermark information is embedded into QPHFMs, ensuring high imperceptibility while enhancing robustness against common attacks. Secondly, DTML is introduced, where the knowledge distilled from watermark detection can facilitate more accurate deepfake detection. Experimental results on benchmark datasets demonstrate that our method surpasses state-of-the-art techniques, delivering exceptional performance in watermark robustness and imperceptibility while simultaneously accomplishing accurate deepfake detection.
Chunpeng Wang 0001, Chaoyi Shi, Simiao Wang, Bin Ma 0003
IEEE Signal Process. Lett.3
2024 Intermediate Domain-Based Meta Learning Framework for Adaptive Object Detection
abstract
Deep learning based object detection methods have made significant progress in recent years. However, these methods often suffer from a substantial performance drop when domain shifts occur, making it difficult to generalize a source domain trained object detector to a new target domain. To address this problem, we propose an Online Meta Learning Framework (OMLF) for unsupervised domain adaptive object detection. In our proposed framework, we adopt the Polar Harmonic Fourier Moment (PHFM) to generate target-like intermediate data. The purpose is to construct a two-pair framework that learns meta knowledge (i.e. model initial parameters) from the pair of “source-to-intermediate” to assist another pair of “intermediate-to-target”. Moreover, the optimizing process requires a heavy computational load due to triggering higher-order gradients. To alleviate this problem, we introduce a shortest-path update strategy that accelerates optimization. When evaluated on several benchmark adaptation scenarios (i.e. normal-to-foggy weather, cross cameras, synthetic-to-real, and real-to-artistic), our OMLF achieves state-of-the-art results, demonstrating its effectiveness.
Yihuan Zhu, Yunan Liu 0001, Chunpeng Wang 0001, Simiao Wang, Mingyu Lu
IEEE Trans. Circuits Syst. Video Technol.4
2024 A Lightweight Network With Latent Representations for UAV Thermal Image Super-Resolution
abstract
While thermal imaging technology on unmanned aerial vehicles (UAVs) has made significant progress, the widespread issue of insufficient resolution poses a serious challenge to comprehending the content of thermal images. Moreover, deploying super-resolution (SR) models on resource-limited UAVs presents considerable difficulties. In an effort to address these challenges, we propose a lightweight thermal image super-resolution (LTSR) model that efficiently extracts multiscale features and learns latent representations. First, we construct a multiscale knowledge distillation (MSKD) network to extract discriminative features from low-resolution (LR) inputs. To achieve this, we use convolution with varying dilation rates to extract features from diverse receptive fields and compress these features through knowledge distillation. Second, to effectively establish continuous relationships among features in the latent space, we develop a forward Markovian restoration process involving multiple diffusion iterations. In each iteration, we integrate the lightweight MSKD network and latent neural representation into a unified end-to-end framework. When evaluated on the challenging benchmark dataset, our method not only has fewer parameters but also outperforms state-of-the-art methods in SR accuracy. Extensive ablation analysis validates the effectiveness of each component in our LTSR.
Tong Liu 0032, Yunan Liu 0001, Simiao Wang, Xinjun Zhang, Jinguang Sun
IEEE Trans. Geosci. Remote. Sens.5
2024 Mask-Guided Mamba Fusion for Drone-Based Visible-Infrared Vehicle Detection
abstract
Drone-based vehicle detection is a critical task within intelligent transportation systems. The existing methods that rely solely on single visible or infrared modalities often struggle to achieve both precise and robust detection. Effectively integrating cross-modal information to assist in vehicle detection remains a significant challenge. In this article, we propose a mask-guided Mamba fusion (MGMF) method for visible-infrared vehicle detection in aerial scenes. The proposed MGMF framework consists of two key components: the masked regularization constraint module (MRCM) and the state-space fusion module (SSFM). First, in MAEM, we use candidate regions from one modality to cover corresponding regions of intermediate-level features from another modality, while a regularization constraint extracts cross-modal guidance. This design allows cross-modal features focused on vehicle areas to be extracted from both modalities for fusion. Second, in SSFM, we propose mapping cross-modal features into a shared hidden state for interaction. This reduces disparities between the cross-modal features and enhances the representation, enabling better perception of intermodal correlations. When evaluated on the DroneVehicle dataset, our MGMF achieves an 80.24% with respect to mAP, establishing a new benchmark for state-of-the-art performance. Ablation studies further demonstrate the effectiveness of our MAEM and SSFM in enhancing visible-infrared fusion for vehicle detection.
Simiao Wang, Chunpeng Wang 0001, Chaoyi Shi, Yunan Liu 0001, Mingyu Lu
IEEE Trans. Geosci. Remote. Sens.1
2023 A bit plane image encryption algorithm based on compound chaos
Simiao Wang, Baoxiang Du
Multim. Tools Appl.2
2021 Medical image super-resolution via deep residual neural network in the shearlet domain
Chunpeng Wang 0001, Simiao Wang, Qi Li 0029, Bin Ma 0003, Jian Li 0034, Meihong Yang, Yun Q. Shi 0001
Multim. Tools Appl.2
2020 Distributed Pregel-based provenance-aware regular path query processing on RDF knowledge graphs
Xin Wang 0030, Simiao Wang, Yueqi Xin, Yajun Yang, Jianxin Li 0001, Xiaofei Wang 0001
World Wide Web2
2019 Transform Domain Based Medical Image Super-resolution via Deep Multi-scale Network
abstract
This paper proposes a new medical image super-resolution (SR) network, namely deep multi-scale network (DMSN), in the uniform discrete curvelet transform (UDCT) domain. DMSN is made up of a set of cascaded multi-scale fushion (MSF) blocks. In each MSF block, we use convolution kernels of different sizes to adaptively detect the local multi-scale feature, and then local residual learning (LRL) is used to learn effective feature from preceding MSF block and current multi-scale features. After obtaining multi-scale features of different MSF block, we use global feature fusion (GFF) to jointly and adaptively learn global hierarchical features in a holistic manner. Finally, compared with other prediction methods in spatial domain, we applied DMSN in UDCT domain, which enables a better representation of global topological structure and local texture detail of HR images. DM-SN shows superior performance over other state-of-the-art medical image SR methods.
Chunpeng Wang 0001, Simiao Wang, Bin Ma 0003, Jian Li 0034, Xiangjun Dong 0001
ICASSP2
2018 Distributed Efficient Provenance-Aware Regular Path Queries on Large RDF Graphs
Yueqi Xin, Xin Wang 0030, Di Jin 0001, Simiao Wang
DASFAA (1)4