Mingmin Chi

dblp:03/2079 · DBLP profile ↗
← Back
60ranked-venue papers
13as first author
29since 2021 · last 2026
0000-0003-2650-4146ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 30 · 27 since 2021Artificial intelligence and machine learning · 20 · 3 first-author · 10 since 2021Applied, interdisciplinary, general and emerging computing · 20 · 10 first-author · 1 since 2021Databases, data management, data science and information retrieval · 5 · 3 first-author
YearPublicationVenuePosition
2026 Real-IAD Variety: Pushing Industrial Anomaly Detection Dataset to a Modern Era
Wenbing Zhu, Chengjie Wang 0001, Bin-Bin Gao, Jiangning Zhang, Guannan Jiang, Jie Hu 0021, Zhenye Gan, Ziqing Zhou, Jianghui Zhang, Linjie Cheng, Yurui Pan, Mingmin Chi, Lizhuang Ma
Pattern Recognit.14
2025 MM-Tracker: Motion Mamba for UAV-platform Multiple Object Tracking
abstract
Multiple object tracking (MOT) from unmanned aerial vehicle (UAV) platforms requires efficient motion modeling. This is because UAV-MOT faces both local object motion and global camera motion. Motion blur also increases the difficulty of detecting large moving objects. Previous UAV motion modeling approaches either focus only on local motion or ignore motion blurring effects, thus limiting their tracking performance and speed. To address these issues, we propose the Motion Mamba Module, which explores both local and global motion features through cross-correlation and bi-directional Mamba Modules for better motion modeling. To address the detection difficulties caused by motion blur, we also design motion margin loss to effectively improve the detection accuracy of motion blurred objects. Based on the Motion Mamba module and motion margin loss, our proposed MM-Tracker surpasses the state-of-the-art in two widely open-source UAV-MOT datasets.
Mufeng Yao, Jinlong Peng, Qingdong He, Mingmin Chi, Jón Atli Benediktsson
AAAI6
2025 Dual-Interrelated Diffusion Model for Few-Shot Anomaly Image Generation
abstract
The performance of anomaly inspection in industrial manufacturing is constrained by the scarcity of anomaly data. To overcome this challenge, researchers have started employing anomaly generation approaches to augment the anomaly dataset. However, existing anomaly generation methods suffer from limited diversity in the generated anomalies and struggle to achieve a seamless blending of this anomaly with the original image. Moreover, the generated mask is usually not aligned with the generated anomaly. In this paper, we overcome these challenges from a new perspective, simultaneously generating a pair of the overall image and the corresponding anomaly part. We propose DualAnoDiff, a novel diffusion-based few-shot anomaly image generation model, which can generate diverse and realistic anomaly images by using a dual-interrelated diffusion model, where one of them is employed to generate the whole image while the other one generates the anomaly part. Moreover, we extract background and shape information to mitigate the distortion and blurriness phenomenon in few-shot image generation. Extensive experiments demonstrate the superiority of our proposed model over state-of-the-art methods in terms of diversity, realism and the accuracy of mask. Overall, our approach significantly improves the performance of downstream anomaly inspection tasks, including anomaly detection, anomaly localization, and anomaly classification tasks. Code will be made available.
Jinlong Peng, Qingdong He, Jiafu Wu, Wenbing Zhu, Mingmin Chi, Jun Liu 0116, Yabiao Wang
CVPR9
2025 OSV: One Step is Enough for High-Quality Image to Video Generation
abstract
Video diffusion models have shown great potential in generating high-quality videos, making them an increasingly popular focus. However, their inherent iterative nature leads to substantial computational and time costs. Although techniques such as consistency distillation and adversarial training have been employed to accelerate video diffusion by reducing inference steps, these methods often simply transfer the generation approaches from Image diffusion models to video diffusion models. As a result, these methods frequently fall short in terms of both performance and training stability. In this work, we introduce a two-stage training framework that effectively combines consistency distillation with adversarial training to address these challenges. Additionally, we propose a novel video discriminator design, which eliminates the need for decoding the video latents and improves the final performance. Our model is capable of producing high-quality videos in merely one-step, with the flexibility to perform multi-step refinement for further performance enhancement. Our quantitative evaluation on the OpenVid-1M benchmark shows that our model significantly outperforms existing methods. Notably, our 1-step performance (FVD 171.15) exceeds the 8-step performance of the consistency distillation based method, AnimateLCM (FVD 184.79), and approaches the 25-step performance of advanced Stable Video Diffusion (FVD 156.94).
Xiaofeng Mao, Zhengkai Jiang 0001, Fu-Yun Wang, Jiangning Zhang, Mingmin Chi, Yabiao Wang, Wenhan Luo
CVPR6
2025 Real-IAD D3: A Real-World 2D/Pseudo-3D/3D Dataset for Industrial Anomaly Detection
abstract
The increasing complexity of industrial anomaly detection (IAD) has positioned multimodal detection methods as a focal area of machine vision research. However, dedicated multimodal datasets specifically tailored for IAD remain limited. Pioneering datasets like MVTec 3D have laid essential groundwork in multimodal IAD by incorporating RGB+3D data, but still face challenges in bridging the gap with real industrial environments due to limitations in scale and resolution. To address these challenges, we introduce Real-IAD D3, a high-precision multimodal dataset that uniquely incorporates an additional pseudo-3D modality generated through photometric stereo, alongside high-resolution RGB images and micrometer-level 3D point clouds. Real-IAD D3features finer defects, diverse anomalies, and greater scale across 20 categories, providing a challenging benchmark for multimodal IAD Additionally, we introduce an effective approach that integrates RGB, point cloud, and pseudo-3D depth information to leverage the complementary strengths of each modality, enhancing detection performance. Our experiments highlight the importance of these modalities in boosting detection robustness and overall IAD performance. The dataset and code are publicly accessible for research purposes at https://realiad4ad.github.io/Real-IAD_D3.
Wenbing Zhu, Ziqing Zhou, Chengjie Wang 0001, Yurui Pan, Ruoyi Zhang, Zhuhao Chen, Linjie Cheng, Bin-Bin Gao, Jiangning Zhang, Zhenye Gan, Yuxie Wang, Shuguang Qian, Mingmin Chi, Lizhuang Ma
CVPR15
2025 Mamba-YOLO-World: Marrying YOLO-World with Mamba for Open-Vocabulary Detection
abstract
Open-vocabulary detection (OVD) aims to detect objects beyond a predefined set of categories. As a pioneering model incorporating the YOLO series into OVD, YOLO-World is well-suited for scenarios prioritizing speed and efficiency. However, its performance is hindered by its neck feature fusion mechanism, which causes the quadratic complexity and the limited guided receptive fields. To address these limitations, we present Mamba-YOLO-World, a novel YOLO-based OVD model employing the proposed MambaFusion Path Aggregation Network (MambaFusion-PAN) as its neck architecture. Specifically, we introduce an innovative State Space Model-based feature fusion mechanism consisting of a Parallel-Guided Selective Scan algorithm and a Serial-Guided Selective Scan algorithm with linear complexity and globally guided receptive fields. It leverages multi-modal input sequences and mamba hidden states to guide the selective scanning process. Experiments demonstrate that our model outperforms the original YOLO-World on the COCO and LVIS benchmarks in both zero-shot and fine-tuning settings while maintaining comparable parameters and FLOPs. Additionally, it surpasses existing state-of-the-art OVD methods with fewer parameters and FLOPs.
Qingdong He, Jinlong Peng, Mingmin Chi, Yabiao Wang
ICASSP5
2025 MambaRF: A Bi-directional Mamba Structure for Radio Frequency Signal Classification of Unmanned Aerial Vehicle
abstract
The rapid development of Unmanned Aerial Vehicle (UAV) technology has facilitated the widespread use of UAVs in daily life, this advancement brings huge regulatory demand for UAVs. Automatic identification of UAVs using radio frequency (RF) signals can effectively reduce regulatory costs. Previous studies first extracts spectrogram features from raw RF signals using short-time Fourier transform, and then use neural networks to classify the spectrograms. These studies are mainly based on local convolution, ignoring the long-range dependence of the spectrograms, and thus are limited in terms of UAV classification accuracy. In addition, these works also focus only on classifying UAV types without identifying flight states (e.g., hovering, flying, and switch on). To this end, we propose MambaRF, which employs a bi-directional selective scanning mechanism to capture long-range information of spectrograms from two different directions. MambaRF also jointly classifies UAV types and flight states by using two decoupled classification heads, which has not been addressed in previous studies. Experiments on two publicly available datasets show that our proposed MambaRF outperforms VMamba and other previous studies in both type classification and flight state classification.
Mufeng Yao, Lexu Xie, Mingmin Chi
ICASSP4
2025 Unicombine: Unified Multi-Conditional Combination with Diffusion Transformer
abstract
With the rapid development of diffusion models in image generation, the demand for more powerful and flexible controllable frameworks is increasing. Although existing methods can guide generation beyond text prompts, the challenge of effectively combining multiple conditional inputs while maintaining consistency with all of them remains unsolved. To address this, we introduce UniCombine, a DiT-based multi-conditional controllable generative framework capable of handling any combination of conditions, including but not limited to text prompts, spatial maps, and subject images. Specifically, we introduce a novel Conditional MMDiT Attention mechanism and incorporate a trainable LoRA module to build both the training-free and training-based versions. Additionally, we propose a new pipeline to construct SubjectSpatial200K, the first dataset designed for multi-conditional generative tasks covering both the subject-driven and spatially-aligned conditions. Extensive experimental results on multi-conditional generation demonstrate the outstanding universality and powerful capability of our approach with state-of-the-art performance.
Jinlong Peng, Qingdong He, Jiafu Wu, Xiaobin Hu, Yanjie Pan, Zhenye Gan, Mingmin Chi, Yabiao Wang
ICCV10
2025 GraphTEN: Graph Enhanced Texture Encoding Network
abstract
Texture recognition is a fundamental problem in computer vision and pattern recognition. Recent progress leverages feature aggregation into discriminative descriptions based on convolutional neural networks (CNNs). However, modeling non-local context relations through visual primitives remains challenging due to the variability and randomness of texture primitives in spatial distributions. In this paper, we propose a graph-enhanced texture encoding network (GraphTEN) designed to capture both local and global features of texture primitives. GraphTEN models global associations through fully connected graphs and captures cross-scale dependencies of texture primitives via bipartite graphs. Additionally, we introduce a patch encoding module that utilizes a codebook to achieve an orderless representation of texture by encoding multi-scale patch features into a unified feature space. The proposed GraphTEN achieves superior performance compared to state-of-the-art methods across five publicly available datasets.
Mufeng Yao, Jianghui Zhang, Mingmin Chi, Jiang Tao
ICME6
2025 A Streamlined System for Multimodal Industrial Anomaly Detection via 2D and 3D Feature Fusion
abstract
We demonstrate an end-to-end system for real-time, multimodal industrial anomaly detection (IAD), built upon a custom hardware platform for synchronized 2D and 3D data acquisition. Our core contribution is a novel cross-modal residual mechanism that identifies defects by quantifying predictive errors between visual and geometric feature spaces. Instead of traditional concatenation, our dual-stream architecture mutually predicts features across modalities, leveraging the prediction residual's magnitude as a direct and robust anomaly indicator. The entire system achieves sub-second inference from acquisition to decision, enabled by efficient depth map analysis that circumvents the complexity of direct point cloud processing, offering a deployable solution for high-speed inspection.
Wenbing Zhu, Mingmin Chi, Bo Peng 0032
ACM Multimedia2
2024 Solving Spectrum Unmixing as a Multi-Task Bayesian Inverse Problem with Latent Factors for Endmember Variability
abstract
With the increasing customization of spectrometers, spectral unmixing has become a widely used technique in fields such as remote sensing, textiles, and environmental protection. However, endmember variability is a common issue for unmixing, where changes in lighting, atmospheric, temporal conditions, or the intrinsic spectral characteristics of materials, can all result in variations in the measured spectrum. Recent studies have employed deep neural networks to tackle endmember variability. However, these approaches rely on generic networks to implicitly resolve the issue, which struggles with the ill-posed nature and lack of effective convergence constraints for endmember variability. This paper proposes a streamlined multi-task learning model to rectify this problem, incorporating abundance regression and multi-label classification with Unmixing as a Bayesian Inverse Problem, denoted as BIPU. To address the issue of the ill-posed nature, the uncertainty of unmixing is quantified and minimized through the Laplace approximation in a Bayesian inverse solver. In addition, to improve convergence under the influence of endmember variability, the paper introduces two types of constraints. The first separates background factors of variants from the initial factors for each endmember, while the second identifies and eliminates the influence of non-existent endmembers via multi-label classification during convergence. The effectiveness of this model is demonstrated not only on a self-collected near-infrared spectral textile dataset (FENIR), but also on three commonly used remote sensing hyperspectral image datasets, where it achieves state-of-the-art unmixing performance and exhibits strong generalization capabilities.
Mingmin Chi, Xuan Zang
AAAI2
2024 Learning Unified Reference Representation for Unsupervised Multi-class Anomaly Detection
Liren He, Zhengkai Jiang 0001, Jinlong Peng, Wenbing Zhu, Liang Liu 0007, Qiangang Du, Xiaobin Hu, Mingmin Chi, Yabiao Wang, Chengjie Wang 0001
ECCV (67)8
2024 TransAVS: End-to-End Audio-Visual Segmentation with Transformer
abstract
Audio-Visual Segmentation (AVS) is a challenging task, which aims to segment sounding objects in video frames by exploring audio signals. Generally AVS faces two key challenges: (1) Audio signals inherently exhibit a high degree of information density, as sounds produced by multiple objects are entangled within the same audio stream; (2) Objects of the same category tend to produce similar audio signals, making it difficult to distinguish between them and thus leading to unclear segmentation results. Toward this end, we propose TransAVS, the first Transformer-based end-to-end framework for AVS task. Specifically, TransAVS disentangles the audio stream as audio queries, which will interact with images and decode into segmentation masks with full transformer architectures. This scheme not only promotes comprehensive audio-image communication but also explicitly excavates instance cues encapsulated in the scene. Meanwhile, to encourage these audio queries to capture distinctive sounding objects instead of degrading to be homogeneous, we devise two self-supervised loss functions at both query and mask levels, allowing the model to capture distinctive features within similar audio data and achieve more precise segmentation. Our experiments demonstrate that TransAVS achieves state-of-the-art results on the AVSBench dataset, highlighting its effectiveness in bridging the gap between audio and visual modalities.
Yuhang Ling, Yuxi Li 0009, Zhenye Gan, Jiangning Zhang, Mingmin Chi, Yabiao Wang
ICASSP5
2024 Learning Hybrid Negative Probability Model for Weakly-Supervised Whole Slide Image Recognition
abstract
Classifying an entire Whole Slide Image (WSI) in a single forward pass is challenging due to its vast resolution. Consequently, current effort on WSI classification resorts to multiple instance learning (MIL), using patch-wise instances to predict categories under image-wise supervision. However, recent MIL methods usually follow implicit instance selection strategy and ignore the effect from inherent patch category imbalances. In a statistical sense, negative patches dominate in WSIs and provide sufficient samples for accurate density estimation. Therefore, in this paper, we learn from anomaly detection and propose a deep MIL framework which learns a hybrid negative probability model to bootstrap discovery of potential positive lesion. We associate attention-based MIL approach with a regularization loss function to explicitly improve patch selection process in positive images. Experiments conducted on benchmarks of WSI recognition demonstrate that our method brings significant improvement to classic attention-based MIL baseline and achieves state-of-the-art performance.
Yining Qiu, Yuxi Li 0009, Jiafu Wu, Zhenye Gan, Mingmin Chi, Yabiao Wang, Chengjie Wang 0001
ICASSP5
2024 Search for Gravitational Wave Probes - A Self-Supervised Learning for Pulsars Based on Signal Contexts
abstract
The recent successful detection of gravitational waves (GWs) at nanohertz based on pulsar timing arrays has underscored the growing significance of searching for new pulsars, which serve as valuable probes for GWs. However, one of the challenges in this endeavor is the lack of labeled data, which can lead to overfitting and poor generalization in supervised deep neural networks. In this paper, we propose a self-supervised pretext task based on signal con-texts to obtain discriminative radio signal representation. Specially, signal attentions are designed to enhance pulse signals within time-phase or frequency-phase images whenever a pulsar is detected. To validate our proposed model, we conducted experiments using the FAST public dataset with a significant improvement in recall and AUC compared to existing single and multimodal deep models with different attentions. As a result, we searched more than 30 new pulsars.
Xiaofeng Cheng, Yuhang Ling, Mingmin Chi, Zhongyi Sun 0002, Yabiao Wang
ICASSP6
2024 MDT-A2G: Exploring Masked Diffusion Transformers for Co-Speech Gesture Generation
abstract
Recent advancements in the field of Diffusion Transformers have substantially improved the generation of high-quality 2D images, 3D videos, and 3D shapes. However, the effectiveness of the Transformer architecture in the domain of co-speech gesture generation remains relatively unexplored, as prior methodologies have predominantly employed the Convolutional Neural Network (CNNs) or simple a few transformer layers. In an attempt to bridge this research gap, we introduce a novel Masked Diffusion Transformer for co-speech gesture generation, referred to as MDT-A2G, which directly implements the denoising process on gesture sequences. To enhance the contextual reasoning capability of temporally aligned speech-driven gestures, we incorporate a novel Masked Diffusion Transformer. This model employs a mask modeling scheme specifically designed to strengthen temporal relation learning among sequence gestures, thereby expediting the learning process and leading to coherent and realistic motions. Apart from audio, Our MDT-A2G model also integrates multi-modal information, encompassing text, emotion, and identity. Furthermore, we propose an efficient inference strategy that diminishes the denoising computation by leveraging previously calculated results, thereby achieving a speedup with negligible performance degradation. Experimental results demonstrate that MDT-A2G excels in gesture generation, boasting a learning speed that is over 6× faster than traditional diffusion transformers and an inference speed that is 5.7× than the standard diffusion model. Our code is available at https://xiaofenmao.github.io/web-project/MDT-A2G/
Xiaofeng Mao, Zhengkai Jiang 0001, Chencan Fu, Jiangning Zhang, Jiafu Wu, Yabiao Wang, Chengjie Wang 0001, Wei Li 0281, Mingmin Chi
ACM Multimedia10
2024 MDR: Multi-stage Decoupled Relational Knowledge Distillation with Adaptive Stage Selection
abstract
The effectiveness of contrastive-learning-based Knowledge Distillation (KD) has sparked renewed interest in relational distillation, but these methods typically focus on angle-wise information from the penultimate layer. We show that exploiting relational information derived from intermediate layers further improves the effectiveness of distillation. We also find that adding distance-wise relational information to contrastive-learning-based methods negatively impacts distillation quality, revealing an implicit contention between angle-wise and distance-wise attributes. Therefore, we propose a Multi-stage Decoupled Relational (MDR) KD framework equipped with an adaptive stage selection to identify the stages that maximize the efficacy of transferring the relational knowledge. MDR framework decouples angle-wise and distance-wise information to resolve their conflicts while still preserving complete relational knowledge, thereby resulting in an elevated transferring efficiency and distillation quality. To evaluate the proposed method, we conduct extensive experiments on multiple image benchmarks i.e. CIFAR100, ImageNet and Pascal VOC, covering various tasks i.e. classification, few-shot learning, transfer learning and object detection. Our method exhibits superior performance under diverse scenarios, surpassing the state of the art by an average improvement of 1.22% on CIFAR-100 across extensively utilized teacher-student network pairs.
Jiaqi Wang 0009, Mingmin Chi
ACM Multimedia3
2023 Hierarchical Multi-Task Learning for Fabric Component Analysis Based on NIR Spectral Signals
abstract
Near Infrared (NIR) Spectral signal has been successfully applied to fabric component analysis (FCA), which is used to identify the category of the textile (defined as a classification task) and its corresponding content for that category (defined as a regression problem). Unlike conventional classification tasks, the prediction results usually contain more than two types of materials, i.e., classes. In addition, the spectral curves belonging to the same parent fiber are similar and thus lead to the problem that the intra-class variance is usually larger than the inter-class variance. The paper proposes a hierarchical architecture for Multi-Label Classification (MLC) and Multi-Output Regression (MOR) to simultaneously identify fabric classes and their contents, i.e., fabric component analysis. In addition, two constraints are applied to the loss function for final classification and regression results. Experimental results conducted on the FEAT-NIR dataset show that the proposed method successfully obtains the best performance compared to the baselines.
Joseph Kim, Mingmin Chi, Gaoqi Xu
ICASSP3
2023 A Method of Constructing and Automatically Labeling Radio Frequency Signal Training Dataset for UAV
abstract
The problem of signal detection and classification of multiple UAVs can be solved using object detection techniques in computer vision. However, this requires collecting and labeling a large amount of reliable raw data. Since the UAV signal dataset cannot be directly applied to object detection, we propose a method using time-frequency domain filtering and automatic labeling to construct a large-scale time-frequency spectrogram dataset. Experimental results show that the average recognition accuracies of image transmission signals and remote control signals under interference conditions are 97% and 82%, respectively, while the average errors of signal parameters are 0.93% and 5.57%.
Ruipeng Ma, Zheng Si, Mingmin Chi
ICASSP4
2023 A Multi-Signal Perception Network for Textile Composition Identification
abstract
Textile composition identification (TCI) is an essential basic link in the textile industry. Methods based on computer vision or near-infrared (NIR) signal processing have shown potential for the nondestructive TCI task. However, these methods ignore that the integration of NIR signals and visual information may help the model learn a better representation through information complementarity. This paper propose a Multi-Signal Perception Network (MSPNet) for nondestructive textile composition identification, allowing the model to benefit from the advantages of multimodal data. Firstly, a two-way feature extraction network is used to obtain multi-modal features. After that, we propose a multimodal signal fusion module to control the aggregation granularity among multimodal data. Specifically, the target areas of the image are perceived by a target area perception module (TAP). Then a bi-gated aggregation (Bi-GFA) is designed to capture consistent semantic information from signal to image and image to signal. The quantitative and qualitative results of the proposed MSPNet are significantly improved compared to both single and multimodal approaches.
Liren He, Mingmin Chi
ICASSP4
2023 Bipartite Graph Convolutional Networks with Adversarial Domain Transfer
abstract
Bipartite graphs have been widely used in many applications such as recommender systems, search engines and so on. Recent works consider bipartite graphs as homogeneous graphs and apply graph convolution networks for link prediction or node classification. However, in bipartite graphs, there are two types of nodes which are from different domains such as users and items in recommender systems, and cannot be in the same embedding space. In this paper, we proposed a novel graph convolution operation to propagate in bipartite graph with less spatial and temporal complexities, and two mapping functions with adversarial constraints to transfer features between two domains. Experimental results show that the proposed model achieves the improved performance on the tasks of link prediction and recommendation in real-world scenarios.
Xuan Zang, Mingmin Chi
ICASSP5
2023 Align, Perturb and Decouple: Toward Better Leverage of Difference Information for RSI Change Detection
abstract
Change detection is a widely adopted technique in remote sense imagery (RSI) analysis in the discovery of long-term geomorphic evolution. To highlight the areas of semantic changes, previous effort mostly pays attention to learning representative feature descriptors of a single image, while the difference information is either modeled with simple difference operations or implicitly embedded via feature interactions. Nevertheless, such difference modeling can be noisy since it suffers from non-semantic changes and lacks explicit guidance from image content or context. In this paper, we revisit the importance of feature difference for change detection in RSI, and propose a series of operations to fully exploit the difference information: Alignment, Perturbation and Decoupling (APD). Firstly, alignment leverages contextual similarity to compensate for the non-semantic difference in feature space. Next, a difference module trained with semantic-wise perturbation is adopted to learn more generalized change estimators, which reversely bootstraps feature extraction and prediction. Finally, a decoupled dual-decoder structure is designed to predict semantic changes in both content-aware and content-agnostic manners. Extensive experiments are conducted on benchmarks of LEVIR-CD, WHU-CD and DSIFN-CD, demonstrating our proposed operations bring significant improvement and achieve competitive results under similar comparative conditions. Code is available at https://github.com/wangsp1999/CD-Research/tree/main/openAPD
Supeng Wang, Yuxi Li 0009, Mingmin Chi, Yabiao Wang, Chengjie Wang 0001, Wenbing Zhu
IJCAI4
2023 PVG: Progressive Vision Graph for Vision Recognition
abstract
Convolution-based and Transformer-based vision backbone networks process images into the grid or sequence structures, respectively, which are inflexible for capturing irregular objects. Though Vision GNN (ViG) adopts graph-level features for complex images, it has some issues, such as inaccurate neighbor node selection, expensive node information aggregation calculation, and over-smoothing in the deep layers. To address the above problems, we propose a Progressive Vision Graph (PVG) architecture for vision recognition task. Compared with previous works, PVG contains three main components: 1) Progressively Separated Graph Construction (PSGC) to introduce second-order similarity by gradually increasing the channel of the global graph branch and decreasing the channel of local branch as the layer deepens; 2) Neighbor nodes information aggregation and update module by using Max pooling and mathematical Expectation (MaxE) to aggregate rich neighbor information; 3) Graph error Linear Unit (GraphLU) to enhance low-value information in a relaxed form to reduce the compression of image detail information for alleviating the over-smoothing. Extensive experiments on mainstream benchmarks demonstrate the superiority of PVG over state-of-the-art methods, e.g., our PVG-S obtains 83.0% Top-1 accuracy on ImageNet-1K that surpasses GNN-based ViG-S by +0.9↑ with the parameters reduced by 18.5%, while the largest PVG-B obtains 84.2% that has +0.5↑ improvement than ViG-B. Furthermore, our PVG-S obtains +1.3↑ box AP and +0.4↑ mask AP gains than ViG-S on COCO dataset.
Jiafu Wu, Jian Li 0062, Jiangning Zhang, Boshen Zhang, Mingmin Chi, Yabiao Wang, Chengjie Wang 0001
ACM Multimedia5
2023 FOLT: Fast Multiple Object Tracking from UAV-captured Videos Based on Optical Flow
abstract
Multiple object tracking (MOT) has been successfully investigated in computer vision. However, MOT for the videos captured by unmanned aerial vehicles (UAV) is still challenging due to small object size, blurred object appearance, and very large and/or irregular motion in both ground objects and UAV platforms. In this paper, we propose FOLT to mitigate these problems and reach fast and accurate MOT in UAV view. Aiming at speed-accuracy trade-off, FOLT adopts a modern detector and light-weight optical flow extractor to extract object detection features and motion features at a minimum cost. Given the extracted flow, the flow-guided feature augmentation is designed to augment the object detection feature based on its optical flow, which improves the detection of small objects. Then the flow-guided motion prediction is also proposed to predict the object's position in the next frame, which improves the tracking performance of objects with very large displacements between adjacent frames. Finally, the tracker matches the detected objects and predicted objects using a spatially matching scheme to generate tracks for every object. Experiments on Visdrone and UAVDT datasets show that our proposed model can successfully track small objects with large and irregular motion and outperform existing state-of-the-art methods in UAV-MOT tasks.
Mufeng Yao, Jiaqi Wang 0009, Jinlong Peng, Mingmin Chi
ACM Multimedia4
2022 Learning Distinctive Margin toward Active Domain Adaptation
abstract
Despite plenty of efforts focusing on improving the domain adaptation ability (DA) under unsupervised or few-shot semi-supervised settings, recently the solution of active learning started to attract more attention due to its suitability in transferring model in a more practical way with limited annotation resource on target data. Nevertheless, most active learning methods are not inherently designed to handle domain gap between data distribution, on the other hand, some active domain adaptation methods (ADA) usually requires complicated query functions, which is vulnerable to overfitting. In this work, we propose a concise but effective ADA method called Select-by-Distinctive-Margin (SDM), which consists of a maximum margin loss and a margin sampling algorithm for data selection. We provide theoretical analysis to show that SDM works like a Support Vector Machine, storing hard examples around decision boundaries and exploiting them to find informative and transferable data. In addition, we propose two variants of our method, one is designed to adaptively adjust the gradient from margin loss, the other boosts the selectivity of margin sampling by taking the gradient direction into account. We benchmark SDM with standard active learning setting, demonstrating our algorithm achieves competitive results with good data scalability. Code is available at https://github.com/TencentYoutuResearch/ActiveLearning-SDM
Yuxi Li 0009, Yabiao Wang, Zekun Luo, Zhenye Gan, Zhongyi Sun 0002, Mingmin Chi, Chengjie Wang 0001
CVPR7
2022 Hierarchical Signal Fusion Network for Pulsar Detection with Phase-Correlation and Signal Attentions
abstract
The discovery of pulsars is of great importance to human understanding of the universe. Deep learning has exploited to find pulsars based on radio astronomical folded data, which includes time-phase and frequency-phase images and dispersion curve (DM). In this paper, a hierarchical signal fusion network with phase-correlation and signal attentions are pro-posed. Specially, signal attentions are designed to reinforce the pulse signals of time-phase or frequency-phase images if a pulsar appears. After that, pulse signal can be reinforced in the same phase where pulses exist in both images. Accordingly, the pulsar search network is implemented by hierarchical data fusion. The first layer is performed at the feature level through phase correlation attention. The second layer of data fusion done at the decision level is performed by calculating a mapping function weighted by the values provided by the phase-correlation attention and filtered out by the DM peak features for the final discrimination . The proposed model is validated on the FAST public dataset with a significant improvement in recall and AUC compared to existing single and multimodal deep models with different attentions.
Huajian Wu, Mingmin Chi
ICASSP2
2022 A Progressive and Multi-Prior-Guided Network for Image Inpainting
abstract
Deep learning techniques have recently made considerable progress in image inpainting by introducing prior knowledge, e.g. texture and structure. However, the existing methods still suffer from artefacts such as distorted texture and abrupt colors due to insufficient consideration of correlation between the prior visual features. In this paper, we propose a novel progressive and multi-prior-guided network (PMPN) for image inpainting inspired by the human painting process, which first constructs the sketch of painting art, then generates the corresponding textures finally fills the colors to the appropriate locations. In particular, to model the global multi-scale contexts during the reconstruction process, we design a bi-directional cross-stage perception module that captures spatial information across branches and stages and guides the model to synthesize a natural and consistent texture. Our proposed PMPN network is evaluated on three publicly available datasets, outperforming the current state-of-the-art models.
Yining Qiu, Liren He, Mingmin Chi
ICME5
2022 Image-Signal Correlation Network for Textile Fiber Identification
abstract
Identifying fiber compositions is an important aspect of the textile industry. In recent decades, near-infrared spectroscopy has shown its potential in the automatic detection of fiber components. However, for plant fibers such as cotton and linen, the chemical compositions are the same and thus the absorption spectra are very similar, leading to the problem of "different materials with the same spectrum, whereas the same material with different spectrums" and it is difficult using a single mode of NIR signals to capture the effective features to distinguish these fibers. To solve this problem, textile experts under a microscope measure the cross-sectional or longitudinal characteristics of fibers to determine fiber contents with a destructive way. In this paper, we construct the first NIR signal-microscope image textile fiber composition dataset (NIRITFC). Based on the NIRITFC dataset, we propose an image-signal correlation network (ISiC-Net) and design image-signal correlation perception and image-signal correlation attention modules, respectively, to effectively integrate the visual features (esp. local texture details of fibers) with the finer absorption spectrum information of the NIR signal to capture the deep abstract features of bimodal data for nondestructive textile fiber identification. To better learn the spectral characteristics of the fiber components, the endmember vectors of the corresponding fibers are generated by embedding encoding, and the reconstruction loss is designed to guide the model to reconstruct the NIR signals of the corresponding fiber components by a nonlinear mapping. The quantitative and qualitative results are significantly improved compared to both single and bimodal approaches, indicating the great potential of combining microscopic images and NIR signals for textile fiber composition identification.
Liren He, Yining Qiu, Mingmin Chi
ACM Multimedia5
2022 Non-IID federated learning via random exchange of local feature maps for textile IIoT secure computing
Mingmin Chi
Sci. China Inf. Sci.2
2019 Relation Parsing Neural Network for Human-Object Interaction Detection
abstract
Human-Object Interaction Detection devotes to infer a tripletbetween human and objects. In this paper, we propose a novel model, i.e., Relation Parsing Neural Network (RPNN), to detect human-object interactions. Specifically, the network is represented by two graphs, i.e., Object-Bodypart Graph and Human-Bodypart Graph. Here, the Object-Bodypart Graph dynamically captures the relationship between body parts and the surrounding objects. The Human-Bodypart Graph infers the relationship between human and body parts, and assembles body part contexts to predict actions. These two graphs are associated through an action passing mechanism. The proposed RPNN model is able to implicitly parse a pairwise relation in two graphs without supervised labels. Experiments conducted on V-COCO and HICO-DET datasets confirm the effectiveness of the proposed RPNN network which significantly outperforms state-of-the-art methods.
Penghao Zhou, Mingmin Chi
ICCV2
2019 Similarity join on time series by utilizing a dynamic segmentation index
Zhongsheng Li, Peng Wang 0027, Yang Wang 0041, Wei Wang 0009, Ningting Pan, Mingmin Chi
Knowl. Inf. Syst.8
2018 Empowering Dynamic Task-Based Applications with Agile Virtual Infrastructure Programmability
abstract
The IaaS (Infrastructure-as-a-Service) offered by Clouds provides applications with the capability of customizing VMs and configuring their network. Compared to traditional service-based IaaS applications such as persistent web services, most task-based applications have a relatively short duration but are triggered on demand. A typical way to support such kinds of application is to provision a shared and fixed virtual infrastructure based on pre-estimated size in advance, and then perform all the processing tasks. However, due to unpredictable workloads, this solution can lead to either cost inefficiency caused by over-provisioning, or failure to deliver the performance required by applications. CloudsStorm is a dynamic control framework proposed to provide applications with agile programmability and flexibility in controlling the virtual infrastructure. With its front end, applications can design their networked infrastructure and program that infrastructure with our interpreted infrastructure code language. With the back-end engine, the infrastructure code can be executed to provision the networked infrastructure, deploy and execute the application to obtain results, and release resources. Moreover, we adopt multi-threading to support parallel operation. Finally, we conduct experiments in an assumed scenario to demonstrate functionalities of CloudsStorm. The evaluation results prove CloudsStorm is efficient for task-based applications that need to exploit Clouds but reduce the monetary cost.
Huan Zhou 0006, Yang Hu 0013, Jinshu Su, Mingmin Chi, Cees T. A. M. de Laat, Zhiming Zhao
IEEE CLOUD4
2018 Modeling and Evaluating MID1 ICAL Pipeline on Spark
Zhongsheng Li, Wei Wang 0009, Fengbin Qi, Mingmin Chi
DASFAA (2)6
2018 Classifying High Resolution Remote Sensing Images by Fine-Tuned VGG Deep Networks
abstract
Deep convolutional networks perform well in remote sensing (RS) image classification. Usually, it is difficult to obtain a large number of labeled samples in remote sensing classification tasks. Traditionally, the acquisition of remote sensing images is quite different from the photos provided by digital cameras. However, the imaging system for high resolution (HR) RS images (often with RGB 3 channels) is similar to those provided by digital cameras. In the paper, a transfer learning algorithm based on deep neural networks is proposed to attack the problem of lacking labeled RS samples, in particular on the context of pre-trained deep convolutional networks, i.e., VGGNet. Here, the VGGNet is trained on labeled multimedia images provided by “ImageNet Large Scale Visual Recognition Challenge” (ILSVRC). In the proposed strategy, the VGGNet is adopted as a base classifier, and then labeled RS data samples are exploited to fine-tune higher hidden layers in the 16-layer VGG deep neural networks by the back-propagation algorithm. The proposed method is denoted as RS-VGGNet. The proposed RS-VGGNet is validated by real HR remote sensing images, which were acquired from the National Agriculture Imagery Program(NAIP) dataset. Experimental results show that the RS-VGGNet can achieve a higher accuracy compared to the original VGGNet and shallow machine learning methods. And the proposed RS-VGGNet significantly reduces training times and computing burden as well.
Mingmin Chi, Yiqing Qin
IGARSS2
2018 Classification of High Resolution Urban Remote Sensing Images Using Deep Networks by Integration of Social Media Photos
abstract
In recent decades, it is easy to obtain remote sensing images which have been successfully applied to various applications, such as urban planning, hazard monitoring, etc. In particular, high resolution (HR) remote sensing (RS) images can better monitor our living environment from a broader spatial perspective. However, raw remote sensing images provide no labeling information to train a classifier, which usually is exploited to generate remote sensing maps. Based on our previous work, in the paper, an automatic classification system is proposed to classify high resolution urban RS images using deep neural networks, in particular, convolutional neural networks and fully convolutional networks. The labeling information is assigned on the context of both social media photos and HR remote sensing images by significantly reducing the cost of manual labeling without the necessity of remote sensing experts. The experiments carried out on high resolution remote sensing images acquired in the city Frankfurt taken by the Jilin-1 satellites confirm the effectiveness of the proposed strategy compared to the state of the art.
Yiqing Qin, Mingmin Chi, Yijian Zeng, Zhiming Zhao
IGARSS2
2017 Hierarchical Parameter Sharing in Recursive Neural Networks with Long Short-Term Memory
Fengyu Li, Mingmin Chi, Junyu Niu
ICONIP (2)2
2017 Oil Spill Detection via Multitemporal Optical Remote Sensing Images: A Change Detection Perspective
abstract
Oil spill monitoring in optical remote sensing (RS) images is a challenging task due to the complexity of target discrimination in an oil spill scenario. Differently from traditional oil spill detection methods that are mainly carried out in a monotemporal image, in this letter, a novel solution is given in a multitemporal domain by investigating potential capability of change detection (CD) techniques, and it mainly contributes to an unsupervised, semiautomatic, and efficient approach. It opens a new perspective for solving an oil spill detection problem. In particular, a coarse-to-fine multitemporal change analysis procedure is designed to investigate the spectral–temporal variation of change targets present in the scenario. Changes relevant and irrelevant to suspected oil spills are identified and discriminated according to a binary and a multiple CD process, respectively. The proposed approach provides a quick yet effective oil spill detection solution, which is valuable and important in practical applications. The proposed method was validated on two real multitemporal RS data sets presenting the oil spill event in northern Gulf of Mexico in 2010. Experimental results confirmed its effectiveness.
Sicong Liu 0001, Mingmin Chi, Yangxiu Zou, Alim Samat, Jón Atli Benediktsson, Antonio Plaza
IEEE Geosci. Remote. Sens. Lett.2
2017 A Novel Methodology to Label Urban Remote Sensing Images Based on Location-Based Social Media Photos
abstract
With the rapid development of the internet and popularization of intelligent mobile devices, social media is evolving fast and contains rich spatial information, such as geolocated posts, tweets, photos, video, and audio. Those location-based social media data have offered new opportunities for hazards and disaster identification or tracking, recommendations for locations, friends or tags, pay-per-click advertising, etc. Meanwhile, a massive amount of remote sensing (RS) data can be easily acquired in both high temporal and spatial resolution with a multiple satellite system, if RS maps can be provided, to possibly enable the monitoring of our location-based living environments with some devices like charge-coupled device (CCD) cameras but on a much larger scale. To generate the classification maps, usually, labeled RS image pixels should be provided by RS experts to train a classification system. Traditionally, labeled samples are obtained according to ground surveys, image photo interpretation or a combination of the aforementioned strategies. All the strategies should be taken care of by domain experts, in a means which is costly, time consuming, and sometimes of a low quality due to reasons such as photo interpretation based on RS images only. These practices and constraints make it more challenging to classify land-cover RS images using big RS data. In this paper, a new methodology is proposed to classify urban RS images by exploiting the semantics of location-based social media photos (SMPs). To validate the effectiveness of this methodology, an automatic classification system is developed based on RS images as well as SMPs via big data analysis techniques including active learning, crowdsourcing, shallow machine learning, and deep learning. As the labels of RS training data are given by ordinary people with a crowdsourcing technique, the developed system is named Crowd4RS. The quantitative and qualitative experiments confirm the effectiveness of the proposed Crowd4RS system as well as the proposed methodology for automatically generating RS image maps in terms of classification results based on big RS data made up of multispectral RS images in a high spatial resolution and a large amount of photos from social media sites, such as Flickr and Panoramio.
Mingmin Chi, Zhongyi Sun 0002, Yiqing Qin, Jinsheng Shen, Jón Atli Benediktsson
Proc. IEEE1
2016 A multitemporal change detection solution to oil spill monitoring
abstract
This paper develops a novel oil spill detection approach by using the multitemporal optical remote sensing images. Differently from the traditional oil spill detection methods that mainly carried out on a monotemporal image, the proposed approach opens a new perspective to solve the considered oil spill detection problem in a multitemporal domain by investigating the potential capability of change detection (CD) techniques. A coarse to fine multitemporal change analysis is defined to analyze the spectral-temporal variation of change targets that present in the oil spill scenario. Suspected oil spills and non-relevant changes are identified and discriminated according to a multiple-change detection in the proposed technique. The proposed approach provides a quick, yet effective oil spill detection solution in an unsupervised way, which is valuable and important in practical oil spill detection applications. Experimental results obtained on real HJ-1 satellite images presenting the oil spill event in northern Gulf of Mexico in 2010 confirmed the effectiveness of the proposed method.
Sicong Liu 0001, Mingmin Chi, Yangxiu Zou, Alim Samat
IGARSS2
2016 Computational Efficiency Active Learning for classification of hyperspectral images
abstract
Active learning usually is conducted in an iterative way. In the paper, a Computational Efficiency Active Learning (CEAL) algorithm is proposed to address this problem based on diversity measurement for classification of hyperspectral images. In particular, each unlabeled sample is pre-assigned a group label, which can be carried out by such as a clustering algorithm. After that, candidate patterns are selected from each group to satisfy the diversity assumption in each round. The proposed CEAL algorithm is validated by real hyperspectral images. Experimental results show that the proposed CEAL algorithm can obtain not only high classification accuracies but also yield a two to four order of magnitude increase in computational efficiency.
Zhongyi Sun 0002, Mingmin Chi, Jón Atli Benediktsson
IGARSS2
2016 Big Data for Remote Sensing: Challenges and Opportunities
abstract
Every day a large number of Earth observation (EO) spaceborne and airborne sensors from many different countries provide a massive amount of remotely sensed data. Those data are used for different applications, such as natural hazard monitoring, global climate change, urban planning, etc. The applications are data driven and mostly interdisciplinary. Based on this it can truly be stated that we are now living in the age of big remote sensing data. Furthermore, these data are becoming an economic asset and a new important resource in many applications. In this paper, we specifically analyze the challenges and opportunities that big data bring in the context of remote sensing applications. Our focus is to analyze what exactly does big data mean in remote sensing applications and how can big data provide added value in this context. Furthermore, this paper describes the most challenging issues in managing, processing, and efficient exploitation of big data for remote sensing problems. In order to illustrate the aforementioned aspects, two case studies discussing the use of big data in remote sensing are demonstrated. In the first test case, big data are used to automatically detect marine oil spills using a large archive of remote sensing data. In the second test case, content-based information retrieval is performed using high-performance computing (HPC) to extract information from a large database of remote sensing images, collected after the terrorist attack to the World Trade Center in New York City. Both cases are used to illustrate the significant challenges and opportunities brought by the use of big data in remote sensing applications.
Mingmin Chi, Antonio Plaza, Jón Atli Benediktsson, Zhongyi Sun 0002, Jinsheng Shen, Yangyong Zhu
Proc. IEEE1
2012 Input-output-consistent domain adaptation algorithm for remote sensing data classification
abstract
A domain adaptation problem is dealt with where the marginal probability in a target domain is different from but correlated to the one in the source domain but the classification tasks are the same. This problem occurs frequently in classification of remote sensing data, e.g., when data are collected in the same area but at different dates or when data are acquired by the same sensor with the same class label set but in different locations. Traditional learning machines cannot deal with this problem in a satisfactory manner. In this paper, we propose a rationale input-output-consistency where samples in the same cluster and defined by spectral signatures (input space) should have the same class label (output space) if they are accurately classified. With the rationale, samples of high confidence in the target domain are selected to define a new prediction function. Since two domains that are related can have different distributions, the data in the source domain which cannot adapt to the distribution in the target domain are deleted from the training data set. Therefore, the proposed algorithm is denoted as input-consistent-output domain adaptation (iCODA) and works in an iterative way. After the selection of highly-confident target samples and the deletion of source data, a new training data set is used to define a new prediction model. The proposed iCODA algorithm was evaluated on EO-1 hyperspectral data sets from Botswana. Experimental results demonstrate much better classification accuracies when compared to a traditionally used supervised classifier.
Mingmin Chi, Jiangfeng Bao, Xintao Chen, Jón Atli Benediktsson
IGARSS1
2012 Construction of Chinese A-shares Network Using Latent Dirichlet Allocation
abstract
Currently, there are more than 2,400 stocks in Chinese A-shares market and there is almost one IPO share coming into the emerging market per day. The rapid growth of Chinese stock market makes investors difficult to manage portfolio. In the paper, a Chinese A-shares Network (CAN) is constructed using a topic model (i.e., Latent Dirichlet Allocation) to efficiently divide all the A-shares to individual sectors in a probabilistic way in terms of the Business Scope Descriptions (BSD) of the listed companies until December 31, 2011. In the meanwhile, a novel visualization profile is proposed to friendly show stock-sectors relationships. Experimental results validate the effectiveness of the CAN system: the stocks in the same ``sector" defined by the CAN have higher pair wise correlations than those by the experts.
Mingmin Chi, Huijun He, Jiangfeng Bao, Yangyong Zhu
Web Intelligence1
2011 Scalable semi-supervised classification of hyperspectral remote sensing data with spectral and spatial information
abstract
Semi-supervised learning using both labeled and unlabeled data is usually adopted to design a high-accuracy and robust classification system on small-size remote sensing training data set. As suggested in the machine learning literature, the larger amount of unlabeled patterns are used, the better classification accuracies can be obtained. Nevertheless, most recently proposed semi-supervised algorithms are unable to handle a large amount of unlabeled samples. In the paper, we present a scalable semi-supervised learning algorithm by using whole hyperspectral remote sensing image. In particular, both spectral features and spatial information of a remote sensing image are adopted for the scalable semi-supervised learning. The accuracy and the reliability of the proposed algorithm have been evaluated on the ROSIS university hyperspectral remote sensing image. The accuracies are better or comparable when compared to the supervised state-of-the-art algorithms on both small-size and the original training sets.
Mingmin Chi, Jiangfeng Bao, Jón Atli Benediktsson
IGARSS1
2010 Mixture model label propagation
abstract
Usually, we can use a classification or clustering machine learning algorithm to manage knowledge and information retrieval. If we have a small size of known information with a large scale of unknown data, a semi-supervised learning (SSL) algorithm is often preferred. Under the cluster or manifold assumption, usually, the larger amount of unlabeled data are used for learning, the bigger gains of the SSL approaches are achieved. In the paper, we adopt the graph-based SSL algorithm to solve the problem. However the graph-based SSL algorithms are unable to be learnt with large-scale unlabeled samples and originally can only work in a transductive setting. In the paper, we propose a scalable graph-based SSL algorithm to attack the problems aforementioned by Gaussian mixture model label propagation. Experiments conducted on the real dataset illustrate the effectiveness of the proposed algorithm.
Mingmin Chi, Xisheng He, Shipeng Yu
CIKM1
2009 Temporal context as cortical spatial codes
abstract
It is largely unknown how the brain deals with time. The new field of research on autonomous development must enable machines to develop intelligent behaviors that respond not only to spatial features, but also temporal features. Hidden Markov Model (HMM) has a probability based mechanism to deal with time warping, but no effective online method exists that can deal with general temporal structure and temporal abstraction. By online, we mean that the agent must respond to spatial and temporal context immediately while the sensory stream flows in. By general temporal context, we mean various desirable temporal subsets, such as deletion (e.g., stop words) and variable temporal lengths (e.g., beyond bigrams and trigrams). By temporal abstraction, we mean using abstract meaning of context, instead of concrete forms. This paper proposes a brain inspired online scheme for making sequential decisions based on general temporal context. By sequential decisions, the action from the network depends on not only inputs and outputs but also emergent internal context states. In our neuromorphic scheme, the internal states are not predefined symbols, but distributed context depending on the internal attention. Our complexity analysis shows how this scheme greatly reduces the exponential time complexity O(2t) of all the possible number of contexts of length t down to linear time complexity O(cnt), where n is the number of neurons in the network and c is the average number of synapses of each neuron. In this paper, we concentrate on processing sequential text inputs by an online agent network under motor-supervised learning.
Juyang Weng, Mingmin Chi, Xiangyang Xue 0001
IJCNN3
2009 Web image retrieval reranking with multi-view clustering
abstract
General image retrieval is often carried out by a text-based search engine, such as Google Image Search. In this case, natural language queries are used as input to the search engine. Usually, the user queries are quite ambiguous and the returned results are not well-organized as the ranking often done by the popularity of an image. In order to address these problems, we propose to use both textual and visual contents of retrieved images to reRank web retrieved results. In particular, a machine learning technique, a multi-view clustering algorithm is proposed to reorganize the original results provided by the text-based search engine. Preliminary results validate the effectiveness of the proposed framework.
Mingmin Chi, Peiwu Zhang, Yingbin Zhao, Rui Feng 0001, Xiangyang Xue 0001
WWW1
2009 Ensemble Classification Algorithm for Hyperspectral Remote Sensing Data
abstract
In real applications, it is difficult to obtain a sufficient number of training samples in supervised classification of hyperspectral remote sensing images. Furthermore, the training samples may not represent the real distribution of the whole space. To attack these problems, an ensemble algorithm which combines generative (mixture of Gaussians) and discriminative (support cluster machine) models for classification is proposed. Experimental results carried out on hyperspectral data set collected by the reflective optics system imaging spectrometer sensor, validates the effectiveness of the proposed approach.
Mingmin Chi, Qian Kun, Jón Atli Benediktsson
IEEE Geosci. Remote. Sens. Lett.1
2008 MTForest: Ensemble Decision Trees based on Multi-Task Learning
abstract
Many ensemble methods, such as Bagging, Boosting, Random Forest, etc, have been proposed and widely used in real world applications. Some of them are better than others on noise-free data while some of them are better than others on noisy data. But in reality, ensemble methods that can consistently gain good performance in situations with or without noise are more desirable. In this paper, we propose a new method namely MTForest, to ensemble decision tree learning algorihms by enumerating each input attribute as extra task to introduce different additional inductive bias to generate diverse yet accurate component decision tree learning algorithms in the ensemble. The experimental results show that in situations without classification noise, MTForest is comparable to Boosting and Random Forest and significantly better than Bagging, while in situations with classification noise, MTForest is significantly better than Boosting and Random Forest and is slightly better than Bagging. So MTForest is a good choice for ensemble decision tree learning algorithms in situations with or without noise. We conduct the experiments on the basis of 36 widely used UCI data sets that cover a wide range of domains and data characteristics and run all the algorithms within the Weka platform.
Liang Zhang 0019, Mingmin Chi, Jiankui Guo
ECAI3
2008 Cluster-Based Ensemble Classification for Hyperspectral Remote Sensing Images
abstract
Hyperspectral remote sensing images play a very important role in the discrimination of spectrally similar land-cover classes. In order to obtain a reliable classifier, a larger amount of representative training samples are necessary compared to multi-spectral remote sensing data. In real applications, it is difficult to obtain a sufficient number of training samples for supervised learning. Besides, the training samples may not represent the real distribution of the whole space. To attack the quality problems of training samples, we proposed a Cluster-based ENsemble Algorithm (CENA) for the classification of hyperspectral remote sensing images. Data set collected from ROSIS university validates the effectiveness of the proposed approach.
Mingmin Chi, Qun Qian, Jón Atli Benediktsson
IGARSS (1)1
2007 Efficient Feature Extraction for Image Classification
abstract
In many image classification applications, input feature space is often high-dimensional and dimensionality reduction is necessary to alleviate the curse of dimensionality or to reduce the cost of computation. In this paper, we extract discriminant features for image classification by learning a low-dimensional embedding from finite labeled samples. In the new feature space, intra-class compactness and extra-class separability are achieved simultaneously. Target dimensionality of the embedding is selected by spectral analysis. Our method is designed suitable for data with both uni- and multi-modal class distributions. We also develop its two-dimensional variant which makes use of the matrix representation of images. Experimental results on three real image datasets demonstrate the efficacy of our method compared to the state of the art.
Wei Zhang 0016, Xiangyang Xue 0001, Zichen Sun, Yue-Fei Guo, Mingmin Chi, Hong Lu 0001
ICCV5
2007 Support cluster machine
abstract
For large-scale classification problems, the training samples can be clustered beforehand as a downsampling pre-process, and then only the obtained clusters are used for training. Motivated by such assumption, we proposed a classification algorithm, Support Cluster Machine (SCM), within the learning framework introduced by Vapnik. For the SCM, a compatible kernel is adopted such that a similarity measure can be handled not only between clusters in the training phase but also between a cluster and a vector in the testing phase. We also proved that the SCM is a general extension of the SVM with the RBF kernel. The experimental results confirm that the SCM is very effective for largescale classification problems due to significantly reduced computational costs for both training and testing and comparable classification accuracies. As a by-product, it provides a promising approach to dealing with privacy-preserving data mining problems.
Bin Li 0015, Mingmin Chi, Jianping Fan 0001, Xiangyang Xue 0001
ICML2
2007 Classification of hyperspectral data by continuation semi-supervised SVM
abstract
This paper presents a semi-supervised technique for the solution of ill-posed classification problems in remote sensing applications. The proposed technique is based on semisupervised support vector machines (S3VMs) implemented in the primal formulation of the learning problem. In particular, a global optimization algorithm, based on the continuation method, is adopted in the learning phase of the classifier according to an iterative learning procedure. The use of this algorithm can result in a better approximation to the global minimum of the associated cost function. Experimental results, obtained on hyperspectral remote sensing images, point out the advantages and the limitation of the proposed continuation S3VM (cS3VMs) with respect to other implementations of S3VMs.
Mingmin Chi, Lorenzo Bruzzone
IGARSS1
2007 Integration of field work and hyperspectral data for oil and gas exploration
abstract
Hydrocarbon microseepage theory establishes a cause-and-effect relation between oil and gas reservoirs and special surface anomalies, mainly including surface hydrocarbons and related mineral alterations. Thus, diagnostic spectral features of hydrocarbons and mineral alterations are capable of providing reliable evidences for the oil and gas exploration. In the paper, the integrated practical system of exploring for oil, gas is introduced by using reflectance spectroscopy and hydrocarbon microseepage theory. This system is applied to processing and analyzing not only hyperspectral remote sensing data but also the data provided by field work. Furthermore, great efforts are focused on spectral model of hydrocarbon microseepage and Hyperion data classification algorithms. In our work, two exploration targets of natural gas are identified from the study area which covers 2100 km2. All the two exploration targets have been proven industrial reserves by China National Petroleum Corporation (CNPC) in July 2006.
Daqi Xu, Guoqiang Ni, Mingmin Chi
IGARSS5
2007 Semisupervised Classification of Hyperspectral Images by SVMs Optimized in the Primal
abstract
This paper addresses classification of hyperspectral remote sensing images with kernel-based methods defined in the framework of semisupervised support vector machines (S3VMs). In particular, we analyzed the critical problem of the nonconvexity of the cost function associated with the learning phase of S3VMs by considering different (S3VMs) techniques that solve optimization directly in the primal formulation of the objective function. As the nonconvex cost function can be characterized by many local minima, different optimization techniques may lead to different classification results. Here, we present two implementations, which are based on different rationales and optimization methods. The presented techniques are compared with S3VMs implemented in the dual formulation in the context of classification of real hyperspectral remote sensing images. Experimental results point out the effectiveness of the techniques based on the optimization of the primal formulation, which provided higher accuracy and better generalization ability than the S3VMs optimized in the dual formulation
Mingmin Chi, Lorenzo Bruzzone
IEEE Trans. Geosci. Remote. Sens.1
2006 A continuation method for semi-supervised SVMs
abstract
Semi-Supervised Support Vector Machines (S3VMs) are an appealing method for using unlabeled data in classification: their objective function favors decision boundaries which do not cut clusters. However their main problem is that the optimization problem is non-convex and has many local minima, which often results in suboptimal performances. In this paper we propose to use a global optimization technique known as continuation to alleviate this problem. Compared to other algorithms minimizing the same objective function, our continuation method often leads to lower test errors.
Olivier Chapelle, Mingmin Chi, Alexander Zien
ICML2
2006 An ensemble-driven k-NN approach to ill-posed classification problems
Mingmin Chi, Lorenzo Bruzzone
Pattern Recognit. Lett.1
2006 A Novel Transductive SVM for Semisupervised Classification of Remote-Sensing Images
abstract
This paper introduces a semisupervised classification method that exploits both labeled and unlabeled samples for addressing ill-posed problems with support vector machines (SVMs). The method is based on recent developments in statistical learning theory concerning transductive inference and in particular transductive SVMs (TSVMs). TSVMs exploit specific iterative algorithms which gradually search a reliable separating hyperplane (in the kernel space) with a transductive process that incorporates both labeled and unlabeled samples in the training phase. Based on an analysis of the properties of the TSVMs presented in the literature, a novel modified TSVM classifier designed for addressing ill-posed remote-sensing problems is proposed. In particular, the proposed technique: 1) is based on a novel transductive procedure that exploits a weighting strategy for unlabeled patterns, based on a time-dependent criterion; 2) is able to mitigate the effects of suboptimal model selection (which is unavoidable in the presence of small-size training sets); and 3) can address multiclass cases. Experimental results confirm the effectiveness of the proposed method on a set of ill-posed remote-sensing classification problems representing different operative conditions
Lorenzo Bruzzone, Mingmin Chi, Mattia Marconcini
IEEE Trans. Geosci. Remote. Sens.2
2005 Transductive SVMs for semisupervised classification of hyperspectral data
abstract
This paper presents transductive support vector machines (TSVMs) for the semisupervised classification of hyperspectral remote sensing images. On the basis of the analysis of TSVMs recently introduced in the machine learning literature and of the properties of hyperspectral classification problems, a specific TSVM algorithm is proposed to alleviate the Hughes phenomenon in a nonparametric and kernel-based classification framework. The extension of the proposed technique to multiclass cases is also discussed. Experimental results obtained on a real hyperspectral image point out that when small-size training data are available, the proposed TSVMs outperform standard inductive support vector machines (ISVMs).
Lorenzo Bruzzone, Mingmin Chi, Mattia Marconcini
IGARSS2
2005 A semilabeled-sample-driven bagging technique for ill-posed classification problems
abstract
In this letter, a semilabeled-sample-driven bootstrap aggregating (bagging) technique based on a co-inference (inductive and transductive) framework is proposed for addressing ill-posed classification problems. The novelties of the proposed technique lie in: 1) the definition of a general classification strategy for ill-posed problems by the joint use of training and semilabeled samples (i.e., original unlabeled samples labeled by the classification process); and 2) the design of an effective bagging method (driven by semilabeled samples) for a proper exploitation of different classifiers based on bootstrapped hybrid training sets. Although the proposed technique is general and can be applied to any classification algorithm, in this letter multilayer perceptron neural networks (MLPs) are used to develop the basic classifier of the proposed architecture. In this context, a novel cost function for the training of MLPs is defined, which properly considers the contribution of semilabeled samples in the learning of each member of the ensemble. The experimental results, which are obtained on different ill-posed classification problems, confirm the effectiveness of the proposed technique.
Mingmin Chi, Lorenzo Bruzzone
IEEE Geosci. Remote. Sens. Lett.1