Mingjie Li 0006

dblp:48/10103-6 · DBLP profile ↗
← Back
23ranked-venue papers
8as first author
20since 2021 · last 2026
0000-0001-6096-9858ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 11 · 5 first-author · 11 since 2021Graphics, computer vision, multimedia, augmented reality and games · 10 · 3 first-author · 9 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 2 first-author · 3 since 2021Databases, data management, data science and information retrieval · 2 · 1 first-author · 2 since 2021
YearPublicationVenuePosition
2026 Advancing In-Context Learning for Efficient and Stable Medical Report Generation
abstract
Vision-language models (VLMs) have shown strong generalization across multimodal tasks, but adapting them to medical report generation (MRG) often demands extensive paired image-text data that are limited due to data privacy and annotation cost. In-context learning (ICL) offers a promising training-free alternative, yet standard ICL approaches rely on long demonstration prompts that are computationally inefficient and often yield inconsistent or clinically inaccurate descriptions. To address these challenges, we propose Principal In-Context Vectors (PCVs), a compact latent-guidance framework that distills multimodal demonstrations into stable semantic representations. By extracting hidden states from auto-regressive VLMs and applying principal component analysis (PCA), we identify robust semantic directions that remain stable under input perturbations. These PCVs are then injected into new queries to steer generation toward accurate and clinically meaningful outputs without any model tuning. Extensive experiments on four MRG benchmark datasets show that our approach can enhance both zero-shot and fully supervised generation quality across diverse settings, including cross-center, cross-disease, and longitudinal scenarios. This work provides a lightweight and scalable approach to adapt pre-trained VLMs for practical clinical deployment.
Mingjie Li 0006, Zeyi Shi, Mingfei Han 0002, Lina Yao 0001, Zhihui Li 0001, Xiaojun Chang, Kilian M. Pohl, Md Tauhidul Islam, Lei Xing 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 Distilling Object Detectors via Monte Carlo Dropout
abstract
Knowledge distillation (KD) has become a fundamental technique for model compression in object detection tasks. The data noise and training randomness may cause the knowledge of the teacher model to be unreliable, referred to as knowledge uncertainty. Existing methods neglect this uncertainty, potentially hindering the student's capacity to capture and understand latent "dark knowledge". In this work, we introduce a novel strategy that explicitly incorporates knowledge uncertainty, named Uncertainty-Driven Knowledge Extraction and Transfer (UET). Given the unknown, high-dimensional nature of the knowledge distribution, we employ Monte Carlo dropout to effectively estimate the teacher's uncertainty. Leveraging information theory, we combine uncertainty with deterministic knowledge, enabling the student to benefit from both precision and diversity. UET is a plug-and-play method that integrates seamlessly with existing distillation techniques. We validate our approach through comprehensive experiments across various distillation strategies, detectors, and backbones. Specifically, UET achieves state-of-the-art results, with a ResNet50-based GFL detector obtaining 44.1% mAP on the COCO dataset-surpassing baseline performance by 3.9%.
Junfei Yi, Hui Zhang 0023, Jianxu Mao, Tengfei Liu 0005, Mingjie Li 0006, Sihao Lin, Hanyu Gu, Zhihui Li 0001, Xiaojun Chang, Yaonan Wang 0001
IEEE Trans. Pattern Anal. Mach. Intell.5
2026 Knowledge-guided multi-modality transformer for multi-label genetic mutation prediction
Gexin Huang, Chenfei Wu, Mingjie Li 0006, Xiaojun Chang, Ying Sun 0001, Lei Xing 0001, Xiaodan Liang, Liang Lin 0004
Pattern Recognit.3
2026 Mitigating Data Redundancy to Revitalize Transformer-Based Long-Term Time Series Forecasting System
abstract
Long-term time series forecasting (LTSF) is fundamental to various real-world applications, where Transformer-based models have become the dominant framework due to their ability to capture long-range dependencies. However, these models often experience overfitting due to data redundancy in rolling forecasting settings, limiting their generalization ability particularly evident in longer sequences with highly similar adjacent data. In this work, we introduce CLMFormer, a novel framework that mitigates redundancy through curriculum learning and a memory-driven decoder. Specifically, we progressively introduce Bernoulli noise to the training samples, which effectively breaks the high similarity between adjacent data points. This curriculum-driven noise introduction aids the memory-driven decoder by supplying more diverse and representative training data, enhancing the decoder’s ability to model seasonal tendencies and dependencies in the time series data. To further enhance forecasting accuracy, we introduce a memory-driven decoder. This component enables the model to capture seasonal tendencies and dependencies in the time series data and leverages temporal relationships to facilitate the forecasting process. Extensive experiments on six real-world LTSF benchmarks show that CLMFormer consistently improves Transformer-based models by up to 30%, demonstrating its effectiveness in long-horizon forecasting.
Mingjie Li 0006, Guangsi Shi, Mingfei Han 0002, Lina Yao 0001, Xiaojun Chang, Ling Chen 0006
ACM Trans. Intell. Syst. Technol.1
2026 Learning Causality-Aware Exploration with Transformers for Goal-Oriented Navigation
abstract
Navigation is a fundamental task in the research of Embodied AI, and recent advances in machine learning algorithms have garnered growing interest in developing versatile Embodied AI systems. However, current research in this domain reveals opportunities for improvement. First, the direct application of RNNs and Transformers often overlooks the distinct characteristics of navigation tasks compared to traditional sequential data modeling. These methods are inherently designed to capture long-term dependencies, which are relatively weak in navigation scenarios, potentially limiting their performance in such tasks. Second, the reliance on task-specific configurations, such as pre-trained modules and dataset-specific logic, compromises the generalizability of these methods. We address these constraints by initially exploring the unique differences between Navigation tasks and other sequential data tasks through the lens of Causality, presenting a causal framework to elucidate the inadequacies of conventional sequential methods for Navigation. By leveraging this causal perspective, we propose Causality-Aware Transformer (CAT) Networks for Navigation, featuring a Causal Understanding Module to enhance the model’s Environmental Understanding capability. Meanwhile, our method is devoid of task-specific inductive biases and can be trained in an End-to-End manner, which enhances the method’s generalizability across various contexts. Empirical evaluations demonstrate that our methodology consistently surpasses benchmark performances across a spectrum of settings, tasks, and simulation environments, specifically, in Object Navigation within RoboTHOR, Objective Navigation, Point Navigation in Habitat, and R2R Navigation. Extensive ablation studies reveal that the performance gains can be attributed to the Causal Understanding Module, which demonstrates effectiveness and efficiency in both Reinforcement Learning and Supervised Learning settings. Additionally, further analysis highlights the robustness of our method, demonstrating its capacity to consistently perform well across diverse experimental settings and varying conditions. This robustness underscores the adaptability and generalizability of our approach, reinforcing its potential for application across a wide range of tasks.
Ruoyu Wang 0038, Tong Yu 0001, Mingjie Li 0006, Yuanjiang Cao, Yao Liu 0017, Lina Yao 0001
ACM Trans. Intell. Syst. Technol.3
2026 Integrating Anatomical Priors Into a Causal Diffusion Model
abstract
3D brain MRI studies often examine subtle morphometric differences between cohorts that are hard to detect visually. Given the high cost of MRI acquisition, these studies could greatly benefit from image syntheses, particularly counterfactual image generation, as has been the case for applications in computer vision. However, counterfactual models struggle to produce anatomically plausible MRIs due to a lack of explicit inductive biases to preserve fine-grained anatomical details. This shortcoming arises from the training of models that optimize overall image appearance (e.g., via cross-entropy) rather than preserving subtle, yet medically relevant, local variations across subjects. To preserve subtle variations, we propose to explicitly integrate anatomical constraints at the voxel level as priors into a generative diffusion framework. Termed Probabilistic Causal Graph Model (PCGM), the approach captures anatomical constraints via a probabilistic graph module and translates those constraints into spatial binary masks of regions where subtle variations occur. The masks (encoded by a 3D ControlNet) constrain a novel counterfactual denoising UNet, whose encodings are then transferred into high-quality brain MRIs via our 3D diffusion decoder. Extensive experiments across multiple datasets demonstrate that PCGM generates structural brain MRIs of higher quality than several baseline approaches. Furthermore, we show, for the first time, that brain measurements extracted from counterfactuals (generated by PCGM) replicate the subtle effects of a disease on cortical brain regions previously reported in the neuroscience literature. This achievement is an important milestone in the use of synthetic MRIs in studies investigating subtle morphological differences. The codes are available at https://github.com/AndyCA111/PCGM.
Binxu Li, Wei Peng 0009, Mingjie Li 0006, Ehsan Adeli-Mosabbeb, Kilian M. Pohl
IEEE Trans. Medical Imaging3
2025 HC-LLM: Historical-Constrained Large Language Models for Radiology Report Generation
abstract
Radiology report generation (RRG) models typically focus on individual exams, often overlooking the integration of historical visual or textual data, which is crucial for patient follow-ups. Traditional methods usually struggle with long sequence dependencies when incorporating historical information, but large language models (LLMs) excel at in-context learning, making them well-suited for analyzing longitudinal medical data. In light of this, we propose a novel Historical-Constrained Large Language Models (HC-LLM) framework for RRG, empowering LLMs with longitudinal report generation capabilities by constraining the consistency and differences between longitudinal images and their corresponding reports. Specifically, our approach extracts both time-shared and time-specific features from longitudinal chest X-rays and diagnostic reports to capture disease progression. Then, we ensure consistent representation by applying intra-modality similarity constraints and aligning various features across modalities with multimodal contrastive and structural constraints. These combined constraints effectively guide the LLMs in generating diagnostic reports that accurately reflect the progression of the disease, achieving state-of-the-art results on the Longitudinal-MIMIC dataset. Notably, our approach performs well even without historical data during testing and can be easily adapted to other multimodal large models, enhancing its versatility.
Tengfei Liu 0005, Jiapu Wang, Yongli Hu, Mingjie Li 0006, Junfei Yi, Xiaojun Chang, Junbin Gao
AAAI4
2025 Towards Interpretable Counterfactual Generation via Multimodal Autoregression
Chenglong Ma 0002, Yuanfeng Ji, Jin Ye 0002, Lu Zhang 0060, Tianbin Li, Mingjie Li 0006, Junjun He, Hongming Shan
MICCAI (2)7
2025 BossNAS Family: Block-Wisely Self-Supervised Neural Architecture Search
abstract
Recent advances in hand-crafted neural architectures for visual recognition underscore the pressing need to explore architecture designs comprising diverse building blocks. Concurrently, neural architecture search (NAS) methods have gained traction as a means to alleviate human efforts. Nevertheless, the question of whether NAS methods can efficiently and effectively manage diversified search spaces featuring disparate candidates, such as Convolutional Neural Networks (CNNs) and transformers, remains an open question. In this work, we introduce a novel unsupervised NAS approach called BossNAS (Block-wisely Self-supervised Neural Architecture Search), which aims to address the problem of inaccurate predictive architecture ranking caused by a large weight-sharing space while mitigating potential ranking issue caused by biased supervision. To achieve this, we factorize the search space into blocks and introduce a novel self-supervised training scheme called Ensemble Bootstrapping, to train each block separately in an unsupervised manner. In the search phase, we propose an unsupervised Population-Centric Search, optimizing the candidate architecture towards the population center. Additionally, we enhance our NAS method by integrating masked image modeling and present BossNAS++ to overcome the lack of dense supervision in our block-wise self-supervised NAS. In BossNAS++, we introduce the training technique named Masked Ensemble Bootstrapping for block-wise supernet, accompanied by a Masked Population-Centric Search scheme to promote fairer architecture selection. Our family of models, discovered through BossNAS and BossNAS++, delivers impressive results across various search spaces and datasets. Our transformer model discovered by BossNAS++ attains a remarkable accuracy of 83.2% on ImageNet with only 10.5B MAdds, surpassing DeiT-B by 1.4% while maintaining a lower computation cost. Moreover, our approach excels in architecture rating accuracy, achieving Spearman correlations of 0.78 and 0.76 on the canonical MBConv search space with ImageNet and the NATS-Bench size search space with CIFAR-100, respectively, outperforming state-of-the-art NAS methods.
Sihao Lin, Tang Tao, Guangrun Wang, Mingjie Li 0006, Xiaodan Liang, Xiaojun Chang
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 Tackling Real-World Complexity: Hierarchical Modeling and Dynamic Prompting for Multimodal Long Document Classification
abstract
With the rapid growth of internet content, multimodal long document data has become increasingly prominent, drawing significant attention from researchers. However, most existing methods primarily focus on scenarios where all modalities are present, often overlooking more challenging and realistic cases involving missing image modality. To address this limitation, we propose a robust multimodal long document classification (MLDC) framework that integrates hierarchical modeling and dynamic prompting to handle complex multimodal long document data. Our approach begins by leveraging hierarchical modeling combined with an Adaptive Correlation Multimodal Transformer (ACMT) to effectively capture relationships between text and images at both section and sentence levels. We also introduce a Dynamic Prompt Generation (DPG) module at both levels to enhance the model’s robustness in handling missing image data. By evaluating sample uncertainty, the DPG module dynamically adjusts both the number of prompts and the prompts themselves, allowing the model to better adapt to the varying needs of different samples. Finally, a Hierarchical Heterogeneous Graph (HHG) is introduced to enhance feature interactions across levels, further improving the coherence and accuracy of the model. Extensive experiments on four multi-modal long document datasets demonstrate that our model shows superior performance compared to existing state-of-the-art MLDC classification methods in various conditions.
Tengfei Liu 0005, Yongli Hu, Mingjie Li 0006, Junfei Yi, Xiaojun Chang, Junbin Gao
IEEE Trans. Circuits Syst. Video Technol.3
2025 FMSD: Focal Multi-Scale Shape-Feature Distillation Network for Small Fasteners Detection in Electric Power Scene
abstract
In the electric power scene, fasteners play a pivotal role in securing and connecting electrical equipment, with small fastener detection (SFD) being crucial for ensuring operational stability. Despite the replacement of manual inspection methods by non-destructive techniques employing deep learning, these approaches often demand substantial computational resources and involve numerous parameters. While knowledge distillation (KD) can be a viable solution, existing KD methods may often fail to achieve satisfactory performance when dealing with small object presentation and little inter-class variability in SFD tasks. To alleviate this, we propose a Focal Multi-scale Shape-feature Distillation Network (FMSD) to achieve efficient and precise fastener detection in electric power scenarios. Specifically, we propose a novel Multi-Scale Shape-Aware Feature Aggregation module (MSFA) to augment the network's perception of object shape and scale during the KD process. Additionally, we propose a Contour-Guided Distillation (CGD) module to optimize the transfer of the extracted shape-sensitive knowledge between the teacher and student models. Through a series of experiments compared with existing state-of-the-art (SOTA) methods, our method demonstrates superior performance over existing SOTA techniques, both efficiently and effectively. Furthermore, validation on publicly available power scene datasets confirms the generalizability and adaptability of our proposed FMSD across various settings.
Junfei Yi, Jianxu Mao, Hui Zhang 0023, Mingjie Li 0006, Kai Zeng 0010, Mingtao Feng, Xiaojun Chang, Yaonan Wang 0001
IEEE Trans. Circuits Syst. Video Technol.4
2024 Contrastive Learning with Counterfactual Explanations for Radiology Report Generation
Mingjie Li 0006, Haokun Lin, Xiaodan Liang, Ling Chen 0006, Abdulmotaleb El Saddik, Xiaojun Chang
ECCV (43)1
2024 In-Context Learning for Zero-shot Medical Report Generation
Mingjie Li 0006, Ling Chen 0006, Xiaojun Chang, Lina Yao 0001
ACM Multimedia2
2024 DCCN: A dual-cross contrastive neural network for 3D point cloud representation learning
Guangsi Shi, Zexing Zhao, Mingjie Li 0006, Xiaojun Gao, Xiaoli Yan
Expert Syst. Appl.4
2024 Progressive Frame-Proposal Mining for Weakly Supervised Video Object Detection
abstract
In this paper, we focus on the weakly supervised video object detection problem, where each training video is only tagged with object labels, without any bounding box annotations of objects. To effectively train object detectors from such weakly-annotated videos, we propose a Progressive Frame-Proposal Mining (PFPM) framework by exploiting discriminative proposals in a coarse-to-fine manner. First, we design a flexible Multi-Level Selection (MLS) scheme, with explicit guidance of video tags. By selecting object-relevant frames and mining important proposals from these frames, the proposed MLS can effectively reduce frame redundancy as well as improve proposal effectiveness to boost weakly-supervised detectors. Moreover, we develop a novel Holistic-View Refinement (HVR) scheme, which can globally evaluate importance of proposals among frames, and thus correctly refine pseudo ground truth boxes for training video detectors in a self-supervised manner. Finally, we evaluate the proposed PFPM on a large-scale benchmark for video object detection, on ImageNet VID, under the setting of weak annotations. The experimental results demonstrate that our PFPM significantly outperforms the state-of-the-art weakly-supervised detectors.
Mingfei Han 0002, Yali Wang 0001, Mingjie Li 0006, Xiaojun Chang, Yi Yang 0001, Yu Qiao 0001
IEEE Trans. Image Process.3
2023 Dynamic Graph Enhanced Contrastive Learning for Chest X-Ray Report Generation
abstract
Automatic radiology reporting has great clinical potential to relieve radiologists from heavy workloads and improve diagnosis interpretation. Recently, researchers have enhanced data-driven neural networks with medical knowledge graphs to eliminate the severe visual and textual bias in this task. The structures of such graphs are exploited by using the clinical dependencies formed by the disease topic tags via general knowledge and usually do not update during the training process. Consequently, the fixed graphs can not guarantee the most appropriate scope of knowledge and limit the effectiveness. To address the limitation, we propose a knowledge graph with Dynamic structure and nodes to facilitate chest X-ray report generation with Contrastive Learning, named DCL. In detail, the fundamental structure of our graph is pre-constructed from general knowledge. Then we explore specific knowledge extracted from the retrieved reports to add additional nodes or redefine their relations in a bottom-up manner. Each image feature is integrated with its very own updated graph before being fed into the decoder module for report generation. Finally, this paper introduces Image-Report Contrastive and Image-Report Matching losses to better represent visual features and textual information. Evaluated on IU-Xray and MIMIC-CXR datasets, our DCL outperforms previous state-of-the-art models on these two benchmarks.
Mingjie Li 0006, Bingqian Lin, Zicong Chen, Haokun Lin, Xiaodan Liang, Xiaojun Chang
CVPR1
2023 Mask Propagation for Efficient Video Semantic Segmentation
abstract
Video Semantic Segmentation (VSS) involves assigning a semantic label to each pixel in a video sequence. Prior work in this field has demonstrated promising results by extending image semantic segmentation models to exploit temporal relationships across video frames; however, these approaches often incur significant computational costs. In this paper, we propose an efficient mask propagation framework for VSS, called MPVSS. Our approach first employs a strong query-based image segmentor on sparse key frames to generate accurate binary masks and class predictions. We then design a flow estimation module utilizing the learned queries to generate a set of segment-aware flow maps, each associated with a mask prediction from the key frame. Finally, the mask-flow pairs are warped to serve as the mask predictions for the non-key frames. By reusing predictions from key frames, we circumvent the need to process a large volume of video frames individually with resource-intensive segmentors, alleviating temporal redundancy and significantly reducing computational costs. Extensive experiments on VSPW and Cityscapes demonstrate that our mask propagation framework achieves SOTA accuracy and efficiency trade-offs. For instance, our best model with Swin-L backbone outperforms the SOTA MRCFA using MiT-B5 by 4.0% mIoU, requiring only 26% FLOPs on the VSPW dataset. Moreover, our framework reduces up to 4× FLOPs compared to the per-frame Mask2Former baseline with only up to 2% mIoU degradation on the Cityscapes validation set. Code is available at https://github.com/ziplab/MPVSS.
Yuetian Weng, Mingfei Han 0002, Haoyu He 0001, Mingjie Li 0006, Lina Yao 0001, Xiaojun Chang, Bohan Zhuang
NeurIPS4
2023 Video Pivoting Unsupervised Multi-Modal Machine Translation
abstract
The main challenge in the field of unsupervised machine translation (UMT) is to associate source-target sentences in the latent space. As people who speak different languages share biologically similar visual systems, various unsupervised multi-modal machine translation (UMMT) models have been proposed to improve the performances of UMT by employing visual contents in natural images to facilitate alignment. Commonly, relation information is the important semantic in a sentence. Compared with images, videos can better present the interactions between objects and the ways in which an object transforms over time. However, current state-of-the-art methods only explore scene-level or object-level information from images without explicitly modeling objects relation; thus, they are sensitive to spurious correlations, which poses a new challenge for UMMT models. In this paper, we employ a spatial-temporal graph obtained from videos to exploit object interactions in space and time for disambiguation purposes and to promote latent space alignment in UMMT. Our model employs multi-modal back-translation and features pseudo-visual pivoting, in which we learn a shared multilingual visual-semantic embedding space and incorporate visually pivoted captioning as additional weak supervision. Experimental results on the VATEX Translation 2020 and HowToWorld datasets validate the translation capabilities of our model on both sentence-level and word-level and generalizes well when videos are not available during the testing phase.
Mingjie Li 0006, Po-Yao Huang 0001, Xiaojun Chang, Junjie Hu 0001, Yi Yang 0001, Alex Hauptmann 0001
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Auxiliary signal-guided knowledge encoder-decoder for medical report generation
abstract
Medical reports have significant clinical value to radiologists and specialists, especially during a pandemic like COVID. However, beyond the common difficulties faced in the natural image captioning, medical report generation specifically requires the model to describe a medical image with a fine-grained and semantic-coherence paragraph that should satisfy both medical commonsense and logic. Previous works generally extract the global image features and attempt to generate a paragraph that is similar to referenced reports; however, this approach has two limitations. Firstly, the regions of primary interest to radiologists are usually located in a small area of the global image, meaning that the remainder parts of the image could be considered as irrelevant noise in the training procedure. Secondly, there are many similar sentences used in each medical report to describe the normal regions of the image, which causes serious data bias. This deviation is likely to teach models to generate these inessential sentences on a regular basis. To address these problems, we propose an Auxiliary Signal-Guided Knowledge Encoder-Decoder (ASGK) to mimic radiologists' working patterns. Specifically, the auxiliary patches are explored to expand the widely used visual patch features before fed to the Transformer encoder, while the external linguistic signals help the decoder better master prior knowledge during the pre-training process. Our approach performs well on common benchmarks, including CX-CHR, IU X-Ray, and COVID-19 CT Report dataset (COV-CTR), demonstrating combining auxiliary signals with transformer architecture can bring a significant improvement in terms of medical report generation. The experimental results confirm that auxiliary signals driven Transformer-based models are with solid capabilities to outperform previous approaches on both medical terminology classification and paragraph generation metrics.
Mingjie Li 0006, Xiaojun Chang, Xiaodan Liang
World Wide Web (WWW)1
2022 Cross-modal Clinical Graph Transformer for Ophthalmic Report Generation
abstract
Automatic generation of ophthalmic reports using datadriven neural networks has great potential in clinical practice. When writing a report, ophthalmologists make inferences with prior clinical knowledge. This knowledge has been neglected in prior medical report generation methods. To endow models with the capability of incorporating expert knowledge, we propose a Cross-modal clinical Graph Transformer (CGT) for ophthalmic report generation (ORG), in which clinical relation triples are injected into the visual features as prior knowledge to drive the decoding procedure. However, two major common Knowledge Noise (KN) issues may affect models' effectiveness. 1) Existing general biomedical knowledge bases such as the UMLS may not align meaningfully to the specific context and language of the report, limiting their utility for knowledge injection. 2) Incorporating too much knowledge may divert the visual features from their correct meaning. To overcome these limitations, we design an automatic information extraction scheme based on natural language processing to obtain clinical entities and relations directly from in-domain training reports. Given a set of ophthalmic images, our CGT first restores a sub-graph from the clinical graph and injects the restored triples into visual features. Then visible matrix is employed during the encoding procedure to limit the impact of knowledge. Finally, reports are predicted by the encoded cross-modal features via a Transformer decoder. Extensive experiments on the large-scale FFA-IR benchmark demonstrate that the proposed CGT is able to outperform previous benchmark methods and achieve state-of-the-art performances.
Mingjie Li 0006, Wenjia Cai, Karin Verspoor, Shirui Pan, Xiaodan Liang, Xiaojun Chang
CVPR1
2020 GIS-Supervised Building Extraction With Label Noise-Adaptive Fully Convolutional Neural Network
abstract
Automatic building extraction from aerial or satellite images is a dense pixel prediction task for many applications. It demands a large number of clean label data to train a deep neural network for building extraction. But it is labor expensive to collect such pixel-wise annotated data manually. Fortunately, the building footprint data of geographic information system (GIS) maps provide a cheap way of generating building label data, but these labels are imperfect due to misalignment between the GIS maps and images. In this letter, we consider the task of learning a deep neural network to label images pixel-wise from such noisy label data for building extraction. To this end, we propose a general label noise-adaptive (NA) neural network framework consisting of a base network followed by an additional probability transition modular (PTM) which is introduced to capture the relationship between the true label and the noisy label. The parameters of the PTM can be estimated as part of the training process of the whole network by the off-the-shelf backpropagation algorithm. We conduct experiments on real-world data set to demonstrate that our proposed PTM can better handle noisy labels and improve the performance of convolutional neural networks (CNNs) trained on the noisy label data generated by GIS maps for building extraction. The experimental results indicate that being armed with our proposed PTM for fully CNN, it provides a promising solution to reduce manual annotation effort for the labor-expensive object extraction tasks from remote sensing images.
Zenghui Zhang, Weiwei Guo, Mingjie Li 0006, Wenxian Yu
IEEE Geosci. Remote. Sens. Lett.3
2019 Knowledge driven temporal activity localization
Zhihui Li 0001, ZongYuan Ge, Mingjie Li 0006
J. Vis. Commun. Image Represent.4
2018 Rotated Region Based Fully Convolutional Network for Ship Detection
abstract
Ship detection from high-resolution optical remote sensing images has been a prevalent domain in recent years. Unlike objects in natural images, ships of interest can be anywhere in optical remote sensing images with multi-scale and multi-oriented which makes it more different to be detected. In this paper, we propose a novel method based on the fully convolutional network to detect ships. Our method has three important components: 1) we design a network merging different levels of feature map to fuse multi-scale information. Determining the existence of large ship require features from deep layers in the network, while predicting rotated bounding box enclosing small ships need shallow layers information; 2) The network can be trained end-to-end to generate score maps which indicates the confidence score for the ship region of interest in pixel-wise level through all locations and scaled of an image; 3) We design a rotated bounding box regression model to localize the ships. The experimental results on our dataset collected from Google Earth has demonstrated our proposed method achieves promising performance on ship detection in terms of both efficiency and accuracy in high-resolution optical remote sensing images.
Mingjie Li 0006, Weiwei Guo, Zenghui Zhang, Wenxian Yu, Tao Zhang 0027
IGARSS1