VLDB 2026 Research / reviewers in the wild / expert
Qiguang Miao
dblp:03/4610
· DBLP profile ↗
194ranked-venue papers
11as first author
134since 2021 · last 2027
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 105 · 5 first-author · 76 since 2021Graphics, computer vision, multimedia, augmented reality and games · 84 · 3 first-author · 57 since 2021Applied, interdisciplinary, general and emerging computing · 15 · 2 first-author · 10 since 2021Databases, data management, data science and information retrieval · 11 · 1 first-author · 6 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Systems, architecture and hardware · 1 · 1 since 2021Computer networks · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2027 | IPRec: A multimodal recommendation model with item-specific features and progressive knowledge distillation
Junmei Feng, Yaomin Zhao, Qiguang Miao, Zixiang Lu, Zhaoqiang Xia |
Expert Syst. Appl. | 4 |
| 2026 | PriorRG: Prior-Guided Contrastive Pre-training and Coarse-to-Fine Decoding for Chest X-ray Report GenerationabstractChest X-ray report generation aims to reduce radiologists' workload by automatically producing high-quality preliminary reports. A critical yet underexplored aspect of this task is the effective use of patient-specific prior knowledge---including clinical context (e.g., symptoms, medical history) and the most recent prior image---which radiologists routinely rely on for diagnostic reasoning. Most existing methods generate reports from single images, neglecting this essential prior information and thus failing to capture diagnostic intent or disease progression. To bridge this gap, we propose PriorRG, a novel chest X-ray report generation framework that emulates real-world clinical workflows via a two-stage training pipeline. In Stage 1, we introduce a prior-guided contrastive pre-training scheme that leverages clinical context to guide spatiotemporal feature extraction, allowing the model to align more closely with the intrinsic spatiotemporal semantics in radiology reports. In Stage 2, we present a prior-aware coarse-to-fine decoding for report generation that progressively integrates patient-specific prior knowledge with the vision encoder's hidden states. This decoding allows the model to align with diagnostic focus and track disease progression, thereby enhancing the clinical accuracy and fluency of the generated reports. Extensive experiments on MIMIC-CXR and MIMIC-ABN datasets demonstrate that PriorRG outperforms state-of-the-art methods, achieving a 3.6% BLEU-4 and 3.8% F1 score improvement on MIMIC-CXR, and a 5.9% BLEU-1 gain on MIMIC-ABN. Kang Liu 0025, Zhuoqi Ma, Zikang Fang, Yunan Li 0001, Kun Xie 0011, Qiguang Miao |
AAAI | 6 |
| 2026 | Patient-specific multimodal learning with multi-view contrastive alignment for chest X-ray report generationabstractMOTIVATION: Radiology reports play a pivotal role in guiding treatment planning and enabling effective doctor-patient communication. However, their manual composition imposes a substantial workload on radiologists. Although automatic radiology report generation has emerged as a promising alternative, existing approaches predominantly rely on single-view chest X-rays and fail to adequately leverage patient-specific context, thereby limiting diagnostic accuracy. RESULTS: To address this challenge, we propose EVOKE, a novel chest X-ray report generation framework that incorporates multi-view contrastive learning and patient-specific knowledge. Specifically, we introduce a multi-view contrastive learning method that captures semantic correspondences both among multi-view radiographs within a study and between these radiographs and their associated report, thereby improving visual representation learning. We further present a knowledge-guided report generation module that integrates available patient-specific knowledge (i.e., indication, which includes symptom descriptions) to facilitate the generation of accurate and coherent radiology reports. To support research in multi-view report generation, we construct Multiview CXR and Two-view CXR datasets using publicly available sources. Our proposed EVOKE surpasses recent state-of-the-art methods across multiple datasets, achieving a 2.9% F1 RadGraph improvement on MIMIC-CXR, a 5.0% BLEU-1 improvement on MIMIC-ABN, a 1.5% BLEU-4 improvement on Multi-view CXR, and an 8.2% F1,mic-14 CheXbert improvement on Two-view CXR. AVAILABILITY: Code is publicly available at https://github.com/mk-runner/EVOKE, with an archived release available on Zenodo (doi:10.5281/zenodo.21000219). Qiguang Miao, Kang Liu 0025, Zhuoqi Ma, Yunan Li 0001, Xiaolu Kang, Ruixuan Liu, Kun Xie 0011 |
Bioinform. | 1 |
| 2026 | Enhancing person-job fit through multi-temporal career trajectory modeling
Junmei Feng, Shuchun Li, Qiguang Miao, Zhaoqiang Xia |
Expert Syst. Appl. | 4 |
| 2026 | Factual serialization enhancement: A key innovation for chest X-ray report generation
Kang Liu 0025, Zhuoqi Ma, Zhicheng Jiao, Xiaolu Kang, Qiguang Miao, Kun Xie 0011 |
Expert Syst. Appl. | 6 |
| 2026 | Weakly semantic-guided skeleton feature distillation for human action recognition
Ruyi Liu 0001, Qiguang Miao, Wentian Xin, Xiangzeng Liu, Long Li 0005 |
Expert Syst. Appl. | 4 |
| 2026 | Holistic co-speech motion generation via cross-gated attention and cross-limb interaction
Zixiang Lu, Zhixiang Sheng, Zhitong He, Yunan Li 0001, Qiguang Miao |
Expert Syst. Appl. | 6 |
| 2026 | DESformer: Disentangling emotion and style for co-speech body-motion synthesis
Zhitong He, Zixiang Lu, Qiguang Miao, Yunan Li 0001, Kun Xie 0011, Bujia Tian |
Neurocomputing | 4 |
| 2026 | Top-k uniform projection for training-free policy fusion in sequential decision making
Bocheng Zhao, Wucheng Wang, Wenxing Zhang, Qiguang Miao |
Neurocomputing | 6 |
| 2026 | Label acceptance based label propagation algorithm for community detection
Xunlian Wu, Jingqi Hu, Yining Quan, Qiguang Miao, Peng Gang Sun |
Inf. Process. Manag. | 6 |
| 2026 | LRAR: Luminance-ranking autoregressive for low-light image enhancement
Yuntai Liao, Zongfang Ma, Wen Lu 0004, Luze Jia, Qiguang Miao |
Inf. Sci. | 6 |
| 2026 | A two-stage sign language generation framework with self-supervised latent representation learning
Qiguang Miao, Guanwen Feng, Junwei Jing, Yilin Zhang 0007, Yunan Li 0001, Chi-Man Pun |
Knowl. Based Syst. | 1 |
| 2026 | No blind alignment but generation: A different view of continuous sign language recognition based on diffusion
Xi Geng, Yunan Li 0001, Zhuoqi Ma, Zixiang Lu, Qiguang Miao |
Pattern Recognit. | 5 |
| 2026 | Query expansion with topic-aware in-context learning and vocabulary projection for open-domain dense retrieval
Ronghan Li, Mingze Cui, Benben Wang, Yu Wang 0313, Qiguang Miao |
Pattern Recognit. | 5 |
| 2026 | Open-set domain adaptation via unknown sample exploration for hyperspectral image classification
Cheng Shi 0002, Qiguang Miao, Zhiyong Lv |
Pattern Recognit. | 3 |
| 2026 | ITDRCNN: An Interactive Task-Decoupled RCNN for enhanced object detection
Shuai Wu 0001, Hang Wei 0005, Yining Quan, Yong Xu 0001, Chunwei Tian, Qiguang Miao |
Pattern Recognit. | 6 |
| 2026 | LES-Talker: Fine-Grained Emotion Editing for Talking Head Generation in Linear Emotion SpaceabstractWhile existing one-shot talking head generation models have achieved progress in coarse-grained emotion editing, there is still a lack of fine-grained emotion editing models with high interpretability. We argue that for an approach to be considered fine-grained, it needs to provide clear definitions and sufficiently detailed differentiation. We present LES-Talker, a novel one-shot talking head generation model with high interpretability, to achieve fine-grained emotion editing across emotion types, emotion levels, and facial units. We propose a Linear Emotion Space (LES) definition based on Facial Action Units to characterize emotion transformations as vector transformations. We design the Cross-Dimension Attention Net (CDAN) to deeply mine the correlation between LES representation and 3D model representation. Through mining multiple relationships across different feature and structure dimensions, we enable LES representation to guide the controllable deformation of 3D model. In order to adapt the multimodal data with deviations to the LES and enhance visual quality, we utilize specialized network design and training strategies. Experiments show that our method provides high visual quality along with multilevel and inter pretable fine-grained emotion editing, outperforming mainstream methods. Project page: https://peterfanfan.github.io/LES-Talker/ Guanwen Feng, Zhihao Qian 0001, Yunan Li 0001, Qiguang Miao, Chi-Man Pun |
IEEE Trans. Affect. Comput. | 5 |
| 2026 | Reliable-Teacher: Uncertainty-Guided Collaborative Learning for Nighttime Object DetectionabstractNighttime object detection presents significant challenges due to the scarcity of large-scale, high-quality annotations across diverse nighttime scenarios. To circumvent the need for manual nighttime image annotation, researchers have explored Unsupervised Domain Adaptive Object Detection (UDA-OD), which transfers knowledge from labeled daytime datasets to unlabeled nighttime data through pseudo-labeling. While existing approaches have shown promising results, their effectiveness remains limited by the low quality of pseudo labels, restricting model adaptation to nighttime conditions. To address these limitations, we propose Reliable-Teacher, a novel mutual-learning framework that comprehensively leverages target domain knowledge through Uncertainty-Guided Collaborative Learning. Specifically, our approach consists of three key components: 1) A Collaborative Pseudo-Label Construction module that intelligently integrates reliable Teacher-generated pseudo-labels into Student proposals, significantly enhancing pseudo-label quality; 2) An Uncertainty-Guided Consistency Reasoning module that enforces inter-category consistency between Teacher and Student predictions at both anchor and bounding box levels; 3) A Reliability-Weighted Classification Loss that minimizes the influence of unreliable predictions to further enhance uncertainty-guided learning. Extensive experiments demonstrate that Reliable-Teacher significantly outperforms state-of-the-art methods, achieving performance gain of up to 3.1%, 2.2% and 1.7% mAP on BDD100K [1], SHIFT [2], and VisDrone [3] benchmarks, respectively. Upon acceptance, our code will be released to facilitate further research in this domain. Wenjing Jia, Jiaqi Xiao, Jinchang Ren, Di Yuan 0002, Qiguang Miao, Xiangjian He |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | Rule-Semantic Generative Calibration Blur Detection for UAV Imagery
Yihan Wen, Zhuo Zhang 0020, Xianping Ma, Peipei Zhu, Jinglei Li, Guanchong Niu, Qiguang Miao |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2026 | EmoSpeaker: One-Shot Fine-Grained Emotion-Controlled Talking Face GenerationabstractImplementing fine-grained emotion control is crucial for emotion generation tasks because it enhances the expressive capability of the generative model, allowing it to accurately and comprehensively capture and express various nuanced emotional states, thereby improving the emotional quality and personalization of generated content. Generating fine-grained facial animations that accurately portray emotional expressions using only a portrait and an audio recording presents a challenge. In order to address this challenge, we propose a visual attribute-guided audio decoupler. This enables the obtention of content vectors solely related to the audio content, enhancing the stability of subsequent lip movement coefficient predictions. To achieve more precise emotional expression, we introduce a fine-grained emotion coefficient prediction module. Additionally, we propose an emotion intensity control method using a fine-grained emotion matrix. Through these, effective control over emotional expression in the generated videos and finer classification of emotion intensity are accomplished. Subsequently, a series of 3DMM coefficient generation networks are designed to predict 3D coefficients, followed by the utilization of a rendering network to generate the final video. Our experimental results demonstrate that our proposed method, EmoSpeaker, outperforms existing emotional talking face generation methods in terms of expression variation and lip synchronization. Project page:https://peterfanfan.github.io/EmoSpeaker/ Guanwen Feng, Yunan Li 0001, Chaoneng Li, Zhihao Qian 0001, Qiguang Miao, Chi-Man Pun |
IEEE Trans. Multim. | 7 |
| 2026 | A Systematic Review of Skeleton-Based Action Recognition: Methods, Challenges, and Future DirectionsabstractHuman action recognition (HAR), which aims to recognize and understand individual actions and intentions, has rapidly become a research hotspot in computer vision. Compared with other data modalities, skeleton data offers more efficient node semantics and more coherent spatio-temporal motion patterns, effectively reducing the impact of lighting and background changes. In recent years, many researchers have focused on skeleton-based action recognition methods and have made significant progress. However, we believe that the current skeleton-based action recognition methods still face three major challenges: 1) reducing reliance on expensive labeled data while maintaining model performance; 2) enabling the model to understand and recognize new behavior classes with a limited number of samples; and 3) addressing the challenges posed by the lack of skeleton information in single-modality spatio-temporal motion representation learning. Based on these challenges, we conduct a comprehensive review of the existing skeleton-based action recognition methods. Additionally, we provide an extensive review and analysis of publicly available action recognition datasets. This review aims to offer researchers a comprehensive perspective, stimulate more innovative ideas, and promote the application and breakthrough of skeleton action recognition in a wider range of computer vision tasks. Ruyi Liu 0001, Yuzhi Hu, Wentian Xin, Qiguang Miao, Shuai Wu 0001, Long Li 0005 |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2026 | Boosting Semi-Supervised Medical Image Segmentation Through Inter-Instance Information ComplementarityabstractThe acquisition of expert-annotated data remains a critical bottleneck for medical image segmentation, thereby constraining the clinical applicability of highly accurate models. Crucially, despite this scarcity of labeled data, the intrinsic homogeneity in human anatomy across the cohort provides a fundamental basis (or: a promising leverage point) for enhancing model generalization and training efficiency by exploiting inter-instance anatomical complementarity. In this study, we propose a novel semi-supervised approach for medical image segmentation that fully exploits this inter-instance complementarity. The proposed model operates at two levels, integrating a sophisticated copy-paste augmentation module (CPAM) and a trainable region calibration mechanism (TRCM) within the simple mean teacher (MT) framework. Specifically, CPAM is a carefully designed copy-paste strategy that facilitates the exchange of informative regions between samples, thereby enhancing the diversity and robustness of the training data. TRCM leverages the predictions from labeled regions to guide and calibrate the trainable regions in unlabeled data. The calibrated regions typically yield high-quality pseudo-labels, which effectively improve model training. CPAM and TRCM work synergistically, complementing each other to enhance model performance. Experiments on diverse medical image datasets-including LA, ACDC, BraTS2019, and Pancreas-NIH-covering both MRI and CT modalities demonstrate the robust efficacy of our proposed model. In settings with limited annotated data, the model consistently outperforms current state-of-the-art methods across multiple evaluation metrics. The code is available at https://github.com/shuaiaihang/shuaiAIMedcalLab. Shuai Wu 0001, Ruyi Liu 0001, Hang Wei 0005, Linrunjia Liu, Jie Wen 0001, Qiguang Miao |
IEEE Trans. Neural Networks Learn. Syst. | 6 |
| 2026 | Hierarchical Optimization of UAV Deployment and Resource Allocation for ISAC-Enabled Low-Altitude Wireless Networks
Zewei Jing, Qinghai Yang, Ruijin Sun, Qiguang Miao, Jiangzhou Wang, Yuan Wu 0001 |
IEEE Trans. Wirel. Commun. | 5 |
| 2025 | MUCD: Unsupervised Point Cloud Change Detection via Masked Consistencyabstract3D Change Detection (3DCD) has gradually become another research hotspot after image change detection. Recent works focus on using artificial labels for supervised or weakly-supervised training of siamese networks to segment changed points. However, labeling every points of multi-temporal point clouds is very expensive and time-consuming. In addition, these works lack effective self-supervised signals, and existing self-supervised signals often fail to capture sufficiently rich change information. To solve this problem, we assume that the powerful representation of 3D objects should model the consistency information of unchanged regions and distinguish different objects. Based on this assumption, we propose a new unsupervised framework called MUCD to learn change information of multi-temporal point clouds through bidirectional optimization of change segmentor and feature extractor. The training of network is divided into two stages. We first design a foreknowledge point contrastive loss based on the characteristics of the 3DCD task to initialize the feature extractor, and then propose a masked consistency loss to further learn the shared geometric information of unchanged regions in the multi-temporal point clouds, utilizing it as a free and powerful supervised signal to train a change segmentor. In the inference stage, only the segmentor is used to take multi-temporal point clouds as input and produce change segmentation result. Extensive experiments are conducted on SLPCCD and Urb3DCD, two real-world datasets of streets and urban buildings, to verify that our proposed unsupervised method is highly competitive and even outperforms supervised methods in scenes where semantic information changes occur, exhibiting better performance in generalization ability and robustness. Yue Wu 0004, Yongzhe Yuan, Maoguo Gong, Hao Li 0009, Mingyang Zhang 0002, Wenping Ma 0001, Qiguang Miao |
AAAI | 8 |
| 2025 | UrbanWaste: In-the-Bin Dataset for Waste Disposal Inspection with Multi-Granularity Hierarchical LabelsabstractOur world faces the challenge of efficiently and responsibly managing the ever-growing volume of urban waste. Many countries and regions have implemented categorized trash bins and require residents to sort their waste according to specified criteria. Proper waste classification by residents significantly reduces the workload in the waste disposal process. However, due to the lack of effective supervision during classification, the quality of waste sorting is often compromised. This misclassification can lead to higher pollution risks, lower recycling rates, and increased waste management costs and difficulties. To address this issue, we propose using images captured from within trash bins to supervise garbage delivery. We introduce UrbanWaste, an image dataset specifically designed for in-the-bin waste detection and segmentation. The dataset includes 25,254 RGB images and 140,008 annotated items, featuring dense annotations and multi-granularity labels across 193 distinct waste categories. We evaluated state-of-the-art segmentation models to understand their generalization and performance on UrbanWaste. Based on this dataset, we developed a comprehensive workflow for waste classification inspection, which has been deployed in real-world districts to assess the system's effectiveness. We hope UrbanWaste will inspire new directions in AI research for environmental sustainability. Zhuoqi Ma, Zejun You, Xiyue Gao, Qiguang Miao |
AAAI | 6 |
| 2025 | Where Precision Meets Efficiency: Transformation Diffusion Model for Point Cloud RegistrationabstractWe propose a transformation diffusion model for point cloud registration to balance precision and efficiency. Our method formulates point cloud registration as a denoising diffusion process from noisy transformation to object transformation, which is represented by quaternion and translation. Specifically, in training stage, object transformation diffuses from ground-truth transformation to random distribution, and the model learns to reverse this noising process. In sampling stage, the model refines randomly generated transformation to the optimal transformation in a progressive way. We derive the variational bound in closed form for training and provide instantiation of the model. Our diffusion model maps transformation into latent space, and splits the transformation into two components (rotation and translation) based on the fact that they belong to different solution spaces. In addition, our work provides the following crucial findings: (i) Point cloud registration, one of the representative discriminative tasks, can be solved by a generative way and mapped into latent space to obtain new unified probabilistic formulation. (ii) Our model, Transformation Diffusion Model (TDM) can be a plug-and-play agent for point cloud registration, making our method applicable to different deep registration networks. Experimental results on synthetic and real-world datasets demonstrate that, in correspondence-free and correspondence-based scenarios, TDM can both achieve exceeding 60% performance improvements and higher efficiency simultaneously. Yongzhe Yuan, Yue Wu 0004, Xiaolong Fan, Maoguo Gong, Qiguang Miao, Wenping Ma 0001 |
AAAI | 5 |
| 2025 | PointSR: Self-Regularized Point Supervision for Drone-View Object DetectionabstractPoint-Supervised Object Detection (PSOD) in a discriminative style has recently gained significant attention for its impressive detection performance and cost-effectiveness. However, accurately predicting high-quality pseudo-box labels for drone-view images, which often feature densely packed small objects, remains a challenge. This difficulty arises primarily from the limitation of rigid sampling strategies, which hinder the pseudo-box optimization process. To address this, we propose PointSR, an effective and robust point-supervised object detection framework with self-regularized sampling that integrates temporal and informative constraints throughout the pseudo-box generation process. Specifically, the framework comprises three key components: Temporal-Ensembling Encoder (TE Encoder), Coarse Pseudo-box Prediction, and Pseudo-box Refinement. The TE Encoder builds an anchor prototype library by aggregating temporal information for dynamic anchor adjustment. In Coarse Pseudo-box Prediction, anchors are refined using the prototype library, and a set of informative samples is collected for subsequent refinement. During Pseudo-box Refinement, these informative negative samples are used to suppress low-confidence candidate positive samples, thereby improving the quality of the pseudo-boxes. Experimental results on benchmark datasets demonstrate that PointSR significantly outperforms state-of-the-art methods, achieving up to 2.6% ∼ 7.2% higher AP50using only point supervision. Additionally, it exhibits strong robustness to perturbation in human-labeled points. Weizhuo Li, Wenjing Jia, Zehao Zhang, Xiangzeng Liu, Qiguang Miao |
CVPR | 7 |
| 2025 | Enhanced Contrastive Learning with Multi-view Longitudinal Data for Chest X-ray Report GenerationabstractAutomated radiology report generation offers an effective solution to alleviate radiologists’ workload. However, most existing methods focus primarily on single or fixed-view images to model current disease conditions, which limits diagnostic accuracy and overlooks disease progression. Although some approaches utilize longitudinal data to track disease progression, they still rely on single images to analyze current visits. To address these issues, we propose enhanced contrastive learning with Multi-view Longitudinal data to facilitate chest X-ray Report Generation, named MLRG. Specifically, we introduce a multi-view longitudinal contrastive learning method that integrates spatial information from current multi-view images and temporal information from longitudinal data. This method also utilizes the inherent spatiotemporal information of radiology reports to supervise the pre-training of visual and textual representations. Subsequently, we present a tokenized absence encoding technique to flexibly handle missing patient-specific prior knowledge, allowing the model to produce more accurate radiology reports based on available prior knowledge. Extensive experiments on MIMIC-CXR, MIMIC-ABN, and Two-view CXR datasets demonstrate that our MLRG outperforms recent state-of-the-art methods, achieving a 2.3% BLEU-4 improvement on MIMIC-CXR, a 5.5% F1 score improvement on MIMIC-ABN, and a 2.7% F1 RadGraph improvement on Two-view CXR. Kang Liu 0025, Zhuoqi Ma, Xiaolu Kang, Yunan Li 0001, Kun Xie 0011, Zhicheng Jiao, Qiguang Miao |
CVPR | 7 |
| 2025 | Disentangled Pose and Appearance Guidance for Multi-Pose GenerationabstractHuman pose generation is a complex task due to the non-rigid and highly variable nature of human body structures and appearances. However, existing methods often overlook the fundamental differences between spatial transformations of poses and texture generation for appearance, which makes them prone to overfitting. To address this issue, we propose a multi-pose generation framework driven by disentangled pose and appearance guidance. Our approach includes a Global-aware Pose Generation module that iteratively generates pose embeddings, enabling effective control over non-rigid body deformations. Additionally, we introduce the Global-aware Transformer Decoder, which leverages similarity queries and attention mechanisms to achieve spatial transformations and enhance pose consistency through a Global-aware block. In the appearance generation phase, we condition a diffusion model on pose embeddings produced in the initial stage and introduce an Appearance Adapter that extracts high-level contextual semantic information from multi-scale features, enabling further refinement of pose appearance textures and providing appearance guidance. Extensive experiments on the UBC Fashion and TikTok datasets demonstrate that our framework achieves state-of-the-art results in both quality and fidelity, establishing it as a powerful approach for complex pose generation tasks. Tengfei Xiao, Yue Wu 0004, Can Qin, Maoguo Gong, Qiguang Miao, Wenping Ma 0001 |
CVPR | 6 |
| 2025 | Asymmetric Dual-Encoder and Width-Dependent Feature Interaction for Road Extraction in Remote Sensing ImagesabstractRoad extraction from remote sensing images provides basic data for urban planning and intelligent transportation, aiding efficient management and decision-making. Traditional methods struggle to maintain road continuity and integrity due to varying road widths, complex shapes, uneven feature distribution, and occlusions. Despite the powerful feature extraction capabilities of CNNs in this field, their single-branch structure has limited effectiveness in handling roads of varying widths and often results in boundary blurring and disconnections in complex scenes. To address these challenges, we propose an Asymmetric Dual-encoder Feature Interaction Network (ADFINet). This model integrates the local feature extraction capabilities of CNNs with the global context modeling of Transformers. It uses dual encoders to separately extract features for narrow and wide roads and integrates them via a Dual-branch Feature Interaction Fusion Module (DFIFM). Additionally, a Feature Refinement and Redundancy Removal Module (FR3M) enhances feature expression and detail retention. Experiments demonstrate that ADFINet achieves superior performance on the DeepGlobe dataset, significantly improving road extraction accuracy and robustness. Ruyi Liu 0001, Junhong Wu, Qiguang Miao, Panshi Guo |
CW | 3 |
| 2025 | Learning to Suppress Backgrounds and Bidirectionally Fuse Modalities for RGB-D Gesture RecognitionabstractRGB-D video-based gesture recognition is a fundamental task in computer vision, yet it remains challenging due to the small size of hand regions and the presence of redundant background noise. Existing methods often fail to effectively suppress irrelevant background features and inadequately exploit the complementary nature of RGB and depth modalities, leading to suboptimal semantic alignment in feature fusion. To address these issues, we propose a novel end-to-end RGB-D gesture recognition framework that incorporates Spatiotemporal Background Suppression (STBS) and a Bidirectional Modality Fusion Adapter (BMFA). STBS leverages Vision Transformers to construct region-wise tokens and adaptively merges them based on responses scores, suppressing irrelevant background content while preserving fine-grained gesture-related spatiotemporal features. Meanwhile, BMFA enables deep, bidirectional interaction between RGB and depth features across encoder layers, enhancing cross-modal semantic consistency. Extensive experiments on three public RGB-D gesture datasets validate the effectiveness of our method. Experiments conducted on three public RGB-D gesture datasets validate the effectiveness of our approach and demonstrate significant improvements in recognition performance. Our code is available at https://github.com/caicai211/SB-BFM Yunan Li 0001, Yulang Xu, Yilin Zhang 0007, Zixiang Lu, Qiguang Miao |
ECAI | 6 |
| 2025 | KAN-Face: Efficient Resource Usage and Precision Lip-Sync in Talking Head GenerationabstractDespite significant progress in NeRF-based talking head generation, problems like poor lip synchronization and inefficient resource usage remain. To solve these, we propose KANFace, a lightweight framework. In preprocessing, we introduce a Lip-Sync Enhancement Module that uses Wav2Lip to extract high-resolution audio features and map them to an explicit intermediate representation, ensuring precise lip movement alignment with the speaker’s identity. These predicted lip-sync features are combined with fundamental audio-extracted lip features and injected into the rendering module to improve synchronization. For rendering, we introduce FastKAN to map spatial points to color values. As a variant of KAN, FastKAN’s sensitivity to 3D scenes and efficient structure enable precise, fast color prediction. Our framework reduces resource consumption while enhancing lip-sync accuracy and facial reconstruction, making it ideal for talking head generation tasks in resource-limited settings. Project: https://peterfanfan.github.io/KAN-Face/ Guanwen Feng, Zhihao Qian 0001, Yunan Li 0001, Qiguang Miao |
ICASSP | 5 |
| 2025 | Gaussian-Face: Talking Head Generation with Hybrid Density via 3D Gaussian SplattingabstractIn recent years, audio-driven neural radiance field (NeRF)-based talking head generation techniques have achieved impressive results. However, these methods still have some limitations, such as unsynchronized lip movements and visual jitter. Recently, 3D Gaussian splatting has gradually replaced NeRF. Compared to NeRF, 3D Gaussian offers notable advantages, including higher efficiency and better reconstruction quality. Based on this, we propose Gaussian-Face, an audio-driven Gaussian-based facial avatar. With just a few minutes of monocular video and audio, a high-fidelity, driveable facial avatar can be reconstructed within hours. To achieve this, we first use FLAME to obtain the 3D representation of the face and then design a Lip Motion Translator to map audio to 3D lip representations. To model higher-quality facial details, we propose a hybrid density modeling method that balances rendering speed and quality, enabling our approach to render high-fidelity facial avatars at more than 160 FPS. Project page: https://peterfanfan.github.io/Gaussian-Face/ Guanwen Feng, Yilin Zhang 0007, Yunan Li 0001, Qiguang Miao |
ICASSP | 5 |
| 2025 | Sign-Mamba: Advanced Mamba-Based Sign Language GenerationabstractIn the field of sign language generation, Transformer-based models have been widely studied, but their quadratic computational complexity poses challenges. State Space Models (SSMs), like Mamba, offer a promising alternative with efficient long-range interaction modeling and linear complexity. In this study, we propose a two-stage generative framework for sign language generation, called Sign-Mamba, based on the state selection mechanism of SSM. In the first stage, we designed a Mamba-based encoder-decoder architecture, where the encoder captures the latent space representation and a symmetric decoder reconstructs it back into sign language skeletal points. In the second stage, the sign language latent space predicted by the Mamba-based latent predictor is used as a condition to guide the reconstruction network in generating more precise skeletal point sequences. We conducted comprehensive experiments on the PHOENIX and How2Sign datasets, and the results indicate that Sign-Mamba demonstrates competitive performance in sign language generation tasks. Project page: https://peterfanfan.github.io/Sign-Mamba/ Guanwen Feng, Yilin Zhang 0007, Yunan Li 0001, Qiguang Miao |
ICASSP | 5 |
| 2025 | MLNet: Mutual Learning Network to Improve Self-Supervised Representation for Fine-Grained Visual RecognitionabstractHigh-quality annotation of fine-grained visual categorization requires extensive professional knowledge, which is time-consuming and laborious. Therefore, learning fine-grained visual representations from a large number of unlabeled images through self-supervised learning has become a popular alternative solution recently. However, the existing self-supervised learning methods are not effective in fine-grained visual categorization since many features helpful for optimizing self-supervised learning objectives are unsuited to characterize the subtle differences in fine-grained viusal recognition. To deal with this issue, we propose a mutual learning network to enhance the model’s attention towards discriminative semantic features. The key idea is to consider semantic consistency between different augmented views within same image and capture discriminative semantic information. For semantic consistency, our research demonstrates that cross-view attention module between different augmented views can guide our model to capture similar semantic features. Based on this, we further build a GradCAM-guided multi-dimension loss that utilize GradCAM to control our model from different dimensions to discriminative semantic information that are beneficial to fine-grained visual recognition. Experiments on CUB-200-2011, Stanford Cars and Aircrafts datasets demonstrate that the mutual learning network outperforms previous self-supervised learning methods in linear probing and image retrieval. Zixiang Lu, Qiguang Miao |
ICASSP | 4 |
| 2025 | SGAD: Semantic and Geometric-Aware Descriptor for Local Feature MatchingabstractLocal feature matching remains a fundamental challenge in computer vision. Recent Area to Point Matching (A2PM) methods have improved matching accuracy. However, existing research based on this framework relies on inefficient pixel-level comparisons and complex graph matching that limit scalability. In this work, we introduce the Semantic and Geometric-aware Descriptor Network (SGAD), which fundamentally rethinks area-based matching by generating highly discriminative area descriptors that enable direct matching without complex graph optimization. This approach significantly improves both accuracy and efficiency of area matching. We further improve the performance of area matching through a novel supervision strategy that decomposes the area matching task into classification and ranking subtasks. Finally, we introduce the Hierarchical Containment Redundancy Filter (HCRF) to eliminate overlapping areas by analyzing containment graphs. SGAD demonstrates remarkable performance gains, reducing runtime by 60x (0.82s vs. 60.23s) compared to MESA. Extensive evaluations show consistent improvements across multiple point matchers: SGAD+LoFTR reduces runtime compared to DKM, while achieving higher accuracy (0.82s vs. 1.51s, 65.98 vs. 61.11) in outdoor pose estimation, and SGAD+ROMA delivers +7.39% AUC@5° in indoor pose estimation, establishing a new state-of-the-art. Xiangzeng Liu, Guanglu Shi, Qiguang Miao |
ICCV | 5 |
| 2025 | RetouchDiffusion: Unsupervised Personalized Image Retouching via Diffusion ModelsabstractImage retouching aims to enhance the visual quality of images, but existing methods based on style-specific learning often lack the flexibility to cater to individual preferences. To address this limitation, we propose RetouchDiffusion, a personalized image retouching framework that utilizes a user-selected reference image as a guide and leverages the powerful representational capabilities of a pre-trained diffusion model to refine the retouching process. To enable more precise and adaptable adjustments, we introduce the Retouch Network, a dedicated retouching network that preprocesses brightness and tonality, providing controllable auxiliary guidance for the diffusion procedure. Experimental results demonstrate that our method produces high-quality, personalized retouched images more closely aligned with users’ aesthetic preferences across various scenarios. Our code is available at https://github.com/SuperOptimalZ/RetouchDiffusion. Zhuoqi Ma, Zejun You, Yunan Li 0001, Qiguang Miao |
ICME | 5 |
| 2025 | MotionRefineNet: Fine-Grained Pose Sequence Smoothing and RefinementabstractCapturing human motion with existing monocular estimators often results in large errors when dealing with rare poses, occlusions, truncations, and frame blurring, leading to jitter and long-term drift. Although previous methods have introduced post-processing networks for pose refinement, they struggle to balance global smoothing and fine-grained correction. In this work, we propose MotionRefineNet, which leverages the synergy and complementarity between long- and short-term features in the temporal domain and high- and low-frequency features in the frequency domain to address these challenges. The temporal branch is designed as a hierarchical motion structure to learn multi-time scale features, where long-term features learn motion smoothness, and short-term features capture local rapid changes. The frequency branch employs different frequency band learning strategies based on the degrees of freedom (DoF) of body parts. For body parts with low DoF, the focus is on low-frequency features that represent overall motion trends and regular actions. For body parts with high DoF, we design a filter to adaptively extract useful information from all frequency bands, including subtle motion changes in the high-frequency bands. Extensive experiments on multiple datasets and estimators demonstrate that MotionRefineNet outperforms existing methods in refining 2D, 3D, and SMPL poses, achieving superior pose smoothing and deviation correction. Our code is available at: https://github.com/Wheels319/MotionRefineNet. Haolun Li 0001, Weihuang Liu, Jiateng Liu, Zhenhua Tang 0001, Chi-Man Pun, Qiguang Miao, Feng Xu 0005, Hao Gao 0005 |
ACM Multimedia | 6 |
| 2025 | PointTruss: K-Truss for Point Cloud RegistrationabstractPoint cloud registration is a fundamental task in 3D computer vision. Recent advances have shown that graph-based methods are effective for outlier rejection in this context. However, existing clique-based methods impose overly strict constraints and are NP-hard, making it difficult to achieve both robustness and efficiency. While the k-core reduces computational complexity, which only considers node degree and ignores higher-order topological structures such as triangles, limiting its effectiveness in complex scenarios. To overcome these limitations, we introduce the $k$-truss from graph theory into point cloud registration, leveraging triangle support as a constraint for inlier selection. We further propose a consensus voting-based low-scale sampling strategy to efficiently extract the structural skeleton of the point cloud prior to $k$-truss decomposition. Additionally, we design a spatial distribution score that balances coverage and uniformity of inliers, preventing selections that concentrate on sparse local clusters. Extensive experiments on KITTI, 3DMatch, and 3DLoMatch demonstrate that our method consistently outperforms both traditional and learning-based approaches in various indoor and outdoor scenarios, achieving state-of-the-art results. Yue Wu 0004, Yongzhe Yuan, Maoguo Gong, Qiguang Miao, Hao Li 0009, Mingyang Zhang 0002, Wenping Ma 0001 |
NeurIPS | 5 |
| 2025 | FLNet: Filtering and Localizing Network for Fine-Grained Visual Classification
Hang Yao 0001, Weiye Pang, Qiguang Miao |
PRCV (1) | 5 |
| 2025 | Graph reconstruction and attraction method for community detection
Xunlian Wu, Da Teng, Jingqi Hu, Yining Quan, Qiguang Miao, Peng Gang Sun |
Appl. Intell. | 6 |
| 2025 | VMM: Video-Music Mamba for generating background music from videos
Zixiang Lu, Qiguang Miao, Kun Xie 0011 |
Comput. Vis. Image Underst. | 4 |
| 2025 | Multi-scale information sharing and selection network with boundary attention for polyp segmentation
Xiaolu Kang, Zhuoqi Ma, Kang Liu 0025, Yunan Li 0001, Qiguang Miao |
Eng. Appl. Artif. Intell. | 5 |
| 2025 | RECoT: Relation-enhanced Chains-of-Thoughts for knowledge-intensive multi-hop questions answering
Ronghan Li, Haoxiang Jin, Rongcheng Pu, Qiguang Miao |
Neurocomputing | 7 |
| 2025 | Local and Global Spatial-Temporal Transformer for skeleton-based action recognition
Ruyi Liu 0001, Feiyu Gai, Qiguang Miao, Shuai Wu 0001 |
Neurocomputing | 5 |
| 2025 | StrongerCenter: Enhancing TransCenter for robust multi-object tracking
Xiangzeng Liu, Kailai Wang, Bocheng Zhao, Qiguang Miao |
Neurocomputing | 5 |
| 2025 | One-shot handwriting imitation via self-supervised cross spatial transformer networks
Bocheng Zhao, Guanwen Feng, Wenxing Zhang, Yunan Li 0001, Qiguang Miao, Xiangzeng Liu, Ruyi Liu 0001 |
Neurocomputing | 5 |
| 2025 | Motif-based Contrastive Graph Clustering with clustering-oriented prompt
Xunlian Wu, Jingqi Hu, Yining Quan, Qiguang Miao, Peng Gang Sun |
Inf. Process. Manag. | 4 |
| 2025 | Different paths to the same destination: Diversifying LLMs generation for multi-hop open-domain question answering
Ronghan Li, Yu Wang 0313, Zijian Wen, Mingze Cui, Qiguang Miao |
Knowl. Based Syst. | 5 |
| 2025 | A video course enhancement technique utilizing generated talking heads
Zixiang Lu, Bujia Tian, Qiguang Miao, Kun Xie 0011, Ruyi Liu 0001, Yi-Ning Quan |
Neural Comput. Appl. | 4 |
| 2025 | Modeling multi-scale uncertainty with evidence integration for reliable polyp segmentation
Xiaolu Kang, Zhuoqi Ma, Kang Liu 0025, Yunan Li 0001, Qiguang Miao |
Neural Networks | 5 |
| 2025 | Integrated multi-local and global dynamic perception structure for sign language recognition
Siyu Liang 0002, Yunan Li 0001, Huizhou Chen, Qiguang Miao |
Pattern Anal. Appl. | 5 |
| 2025 | Revisiting Siamese-Based 3D Single Object Tracking With a Versatile Transformerabstract3D Single Object Tracking (SOT) plays an important role in real-world visual applications such as autonomous driving and planning. How to realize effective 3D SOT is still a valuable challenge due to its carrier-sparse point clouds and its role-complex influencing factors. Inspired by the remote modeling of popular transformers, we further propose a Versatile Point Tracking Transformer (VPTT) method for 3D SOT, with object guidance from the template point cloud to the search area point cloud under the siamese-based tracking paradigm. Specifically, VPTT employs self- and cross- attention mechanisms and extends four matching operations, resulting in leveraging the contextual information of consecutive frames to improve the tracking results. By constructing a deep network VerFormer consisting of four successive transformer layers, which performs matching operations involving fusional transformation, separative discrimination, intersectional interaction, and unidirectional propagation from shallow to deep. Considering that the tracking task involves multiple processes, VPTT further learns how to forecast intermediate outputs including mask probability, trailing distance, and heading angle at each stage. Such a specialized design allows our VPTT to revisit the end-to-end training paradigm used for 3D tracking while developing a versatile transformer that is a perfect fit for the 3D SOT task. Experiments on three benchmarks, KITTI, nuScenes, and Waymo, show that VPTT achieves state-of-the-art tracking performance on siamese-based tracking running at $\sim$∼62 FPS. Yue Wu 0004, Qiguang Miao, Maoguo Gong, Linghe Kong |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | SG-CLR: Semantic representation-guided contrastive learning for self-supervised skeleton-based action recognition
Ruyi Liu 0001, Wentian Xin, Qiguang Miao, Xiangzeng Liu, Long Li 0005 |
Pattern Recognit. | 5 |
| 2025 | Learning hyperspectral noisy label with global and local hypergraph laplacian energy
Cheng Shi 0002, Linfeng Lu, Minghua Zhao, Xinhong Hei 0001, Chi-Man Pun, Qiguang Miao |
Pattern Recognit. | 6 |
| 2025 | CompNET: Boosting image recognition and writer identification via complementary neural network post-processingabstractIn current classification tasks, an important method to improve accuracy is to pre-train the model using a large-scale domain-specific dataset. However, many tasks such as writer identification (writerID) lack suitable large-scale datasets in practical scenarios. To address this issue, this paper proposes a method that can improve prediction accuracy without relying on significant pre-training but leveraging the diversity of probability distributions predicted by multiple networks, and enhancing the top-1 accuracy through complementary post-processing. Specifically, top-k distributions are sampled from the multiple probability mass functions separately. When the distribution differences of top-k are maximized, the intersection other than the correct category can be narrowed down. Finally, the correct target with suboptimal probability can be rectified by the only intersection. Furthermore, our method has exhibited an intriguing trait during experimentation. Its prediction accuracy enhances concurrently with the incorporation of novel SOTA methods, ultimately surpassing the performance of these new methods. Bocheng Zhao, Xuan Cao, Wenxing Zhang, Xujie Liu, Qiguang Miao, Yunan Li 0001 |
Pattern Recognit. | 5 |
| 2025 | CAETFN: Context Adaptively Enhanced Text-Guided Fusion Network for Multimodal Sentiment AnalysisabstractMultimodal sentiment analysis (MSA) is an active research area in recent years with the exponential development of the internet and social media, which aims to recognize the speaker’s sentiment in the video consisted of text, acoustic and visual cues, and has attracted attention from many applications such as smart education, intelligent medication and social security. The predominant approaches have devoted to developing more complicated fusion strategy to learn efficient multimodal representations. However, information from these modalities usually have different contributions to MSA task. More specifically, the text modality outperforms the non-verbal modalities since its highly condensed semantic information and the maturity of the pre-trained language models. Taking full advantage of the text modality while integrating the non-verbal sentiment-relevant contextual information becomes a substantial challenge. Thus, in this paper, we propose a Context Adaptively Enhanced Text-guided Fusion Network, which is embedded in the pre-trained language model and utilizes the text modality as the guide to reduce the redundancy and exploit the sentiment-relevant information and in turn uses these information to complement itself with the non-verbal sentiment contexts. Moreover, a novelly designed non-verbal feature enhancement module is introduced to capture long-range dependencies in two directions, with the substantial removal of the redundancy and the noise. Extensive experiments on two benchmark datasets CMU-MOSI and CMU-MOSEI demonstrate the competitive performance over the state-of-the-art methods. Ruyi Liu 0001, Qiguang Miao, Di Wang 0011, Xiangzeng Liu |
IEEE Trans. Affect. Comput. | 3 |
| 2025 | Boundary-Aware Sentence-Gloss Alignment With Semantic Similarity Measurement for Continuous Sign Language RecognitionabstractContinuous sign language recognition (CSLR) plays a crucial role in facilitating communication between deaf and hearing individuals. A key aspect of achieving precise CSLR is the alignment of the video segment of each sign with its gloss, namely its corresponding text representation in natural language. However, the coarticulation phenomenon, where contextual dependencies between adjacent signs blur the boundaries of individual signs, poses a significant challenge to this task. In this paper, we propose a novel boundary-aware sentence-gloss alignment network for CSLR to address this challenge. Our network first designs a task-relevant boundary-aware similarity measurement, evaluating sign frames by both appearance and their recognition contribution, mitigating coarticulation-induced transition noise to restore precise boundaries. For enhanced alignment, we propose a hierarchical sentence-gloss alignment: coarse sentence-level alignment reduces cross-modal disparity, while fine-grained gloss-level alignment refines video-to-token mapping. Finally, an adaptive class-divergence loss sharpens gloss decoding by maximizing inter-class discrimination. Our proposed framework provides a simple and effective solution to mitigate the boundary ambiguity caused by coarticulation, optimizing continuous sign language recognition algorithms from a new perspective. Extensive experiments conducted on four public sign language recognition (SLR) datasets demonstrate that our proposed boundary-aware sentence-gloss alignment network learns precise alignments and achieves state-of-the-art performance. Yunan Li 0001, Xi Geng, Zhuoqi Ma, Qiguang Miao, Chi-Man Pun |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Adaptive Occlusion-Aware Network for Occluded Person Re-IdentificationabstractOccluded person re-identification (ReID) is a challenging task due to some of the essential features are interfered by obstacles or other pedestrians. Multi-granularity local feature extraction and recognition can effectively improve the accuracy of ReID under occlusion. However, manual segmentation methods for local features can lead to feature misalignment. Feature alignment based on pose estimation often ignores non-body details (e.g., handbags, backpacks, etc.) while increasing the complexity of the model. To address the above challenges, we propose a novel Adaptive Occlusion-Aware Network (AOANet), which mainly consists of two modules, the Adaptive Position Extractor (APE) and the Occlusion Awareness Module (OAM). In order to adaptively extract distinguishing features of body parts, APE optimizes the representation of multi-granularity features by the guidance of attention mechanism and keypoint features. To further perceive the occluded region, the OAM is developed by adaptively calculating the occlusion weights for body parts. These weights can lead to highlighting the non-occluded parts and suppressing the occluded parts, which in turn improves the accuracy in the occluded situation. Extensive experimental results confirm the advantages of our method on the MSMT17, DukeMTMC-reID, Market-1501, Occluded-Duke and Occluded-ReID datasets. The comparative results demonstrate that our method outperforms comparable methods. Especially on the Occluded-Duke dataset, our method achieved 70.6% mAP and 81.2% Rank-1 accuracy. Xiangzeng Liu, Hao Chen 0187, Qiguang Miao, Ruyi Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2025 | Triple Point MaskingabstractExisting 3D mask learning methods encounter performance bottlenecks under limited data, and our objective is to overcome this limitation. In this paper, we introduce a triple point masking scheme, named TPM, which serves as a scalable plug-and-play framework for MAE pre-training to achieve multi-mask learning for 3D point clouds. Specifically, we augment the baseline methods with two additional mask choices (i.e., medium mask and low mask) as our core insight is that the recovery process of an object can manifest in diverse ways. Previous high-masking schemes focus on capturing the global representation information but lack fine-grained recovery capabilities, so that the generated pre-training weights tend to play a limited role in the fine-tuning process. With the support of the proposed TPM, current methods can exhibit more flexible and accurate completion capabilities, enabling the potential autoencoder in the pre-training stage to consider multiple representations of a single 3D point cloud object. In addition, during the fine-tuning stage, an SVM-guided weight selection module is proposed to fill the encoder parameters for downstream networks with the optimal weight, maximizing linear accuracy and facilitating the acquisition of intricate representations for new objects. Extensive experimental results and theoretical analysis show that five baselines equipped with the proposed TPM achieve comprehensive performance improvements on various downstream tasks. Our code and models are available athttps://github.com/liujia99/TPM. Linghe Kong, Yue Wu 0004, Maoguo Gong, Hao Li 0009, Qiguang Miao, Wenping Ma 0001, Can Qin |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2025 | Equivariance-Based Markov Decision Process for Unsupervised Point Cloud RegistrationabstractUnsupervised point cloud registration is crucial in 3D computer vision. However, most unsupervised methods struggle to construct effective optimization objectives and reliable unsupervised signals to enhance the performance of the model. To address these issues, with the observation of the significant alignment between the registration process and the Markov Decision Process (MDP), we model point cloud registration as MDP, which can provide more reliable unsupervised signals through the reward. We propose a colored noise based cross-entropy method, which introduces colored noise into sampling process, regulating the power spectral density of the action sequence and expanding the search space, improving the registration effect. Particularly, to strengthen constraints on MDP and training in the transformation space, we utilize equivariance theory to construct transformation equivariant constraint as a new optimization objective and derive equivariant constraint solutions for optimization, providing more reliable unsupervised signals. Extensive experiments demonstrate the superior performance of our method on benchmark datasets. Yue Wu 0004, Jiayi Lei, Yongzhe Yuan, Xiaolong Fan, Maoguo Gong, Wenping Ma 0001, Qiguang Miao, Mingyang Zhang 0002 |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2025 | Multitask Multiscale Feature Selection for Point Cloud Registrationabstract3D point cloud registration is a process of solving the geometric transformation between two point clouds. This process is an important issue in computer vision and pattern recognition. The registration methods based on geometric features are highly sensitive to the scale of feature extraction. Changes in scale can introduce inaccuracies in feature descriptions, thereby compromising the reliability of the registration results. To mitigate the impact of feature scale on the outcomes and the high-dimensional issue arising from features of different scales, we propose a method for multi-scale point cloud feature selection. We solve the high-dimensional problem of feature selection by designing a multi-task framework. By designing a mutual information dimensionality reduction method, we decomposed the high-dimensional feature selection task of different descriptors with multi-scale features into multiple related low-dimensional feature selection tasks. Then, by means of the knowledge transfer among these low-dimensional feature selection tasks, we sought the best feature subset to obtain more robust feature information. We evaluate the effectiveness of our method by conducting extensive experiments on various datasets. The experimental results show that the method outperforms other feature descriptors in terms of descriptive power and robustness and improves the effectiveness of point cloud registration. Yue Wu 0004, Chuang Luo, Maoguo Gong, Hangqi Ding, Jinlong Sheng, Qiguang Miao, Hao Li 0009, Wenping Ma 0001 |
IEEE Trans. Evol. Comput. | 6 |
| 2025 | Evolutionary Multitasking Descriptor Optimization for Point Cloud RegistrationabstractPoint cloud registration (PCR) is an important task for other point cloud tasks. Feature-based methods are widely adopted for their speed and efficiency in PCR. The descriptive capability of features extracted by a single geometric descriptor is limited. Descriptive capabilities can be improved by concatenating features extracted from multiple descriptors. However, due to the existence of redundant and irrelevant features, the correct corresponding points are difficult to match, which further affects the registration effect. We propose an evolutionary multitasking point cloud descriptor optimization method. Integrate existing descriptors to optimize descriptors with stronger description ability. Labeling features to calculate the feature importance for the registration and generating multitasks. In optimized processing, approximate evaluation which is calculated by prior correspondence saved in the database replaces the expensive searching correspondences process in the entire point cloud. Finally, a multiscale filter is developed to remove error correspondences by the geometric information from multiple scale descriptor features. Experimental demonstrate that the proposed approach can optimize a feature subset with higher-descriptive capability compared to other methods and show superior PCR performance on 14 point cloud models. This is the first paper on point cloud descriptor optimization, which provides a new idea for PCR research. Yue Wu 0004, Jinlong Sheng, Hangqi Ding, Peiran Gong, Hao Li 0009, Maoguo Gong, Wenping Ma 0001, Qiguang Miao |
IEEE Trans. Evol. Comput. | 8 |
| 2025 | Adaptive Multitype Contrastive Views Generation for Remote Sensing Image Semantic SegmentationabstractSelf-supervised contrastive learning is a powerful pre-training framework for learning the invariant features from the different views of remote sensing images, therefore, the performance of contrastive learning heavily depends on the generation of views. Current view generation is primarily accomplished through different transformations, and the types and parameters of the transformations are require hand-crafted. Hence, the diversity and discriminability of generated views cannot be guaranteed. To address this, we propose a multi-type views optimization method to optimize these transformations. We formulate contrastive learning as a min-max optimization problem, and transformation parameters are optimized by maximizing the contrastive loss. The optimized transformations encourage the negative sample pairs to be close and the positive sample pairs to be far apart. Different from the current adversarial view generation methods, our method can optimize both photometric transformations and geometric transformations. For remote sensing images, the geometric transformation is more critical for view generation, while the existing view optimization methods fail to achieve this. We consider the hue, saturation, brightness, contrast, and geometric rotation transformations in contrastive learning, and evaluate the optimized views on the downstream remote sensing images semantic segmentation task. Extensive experiments are carried on the three remote sensing image segmentation datasets, including ISPRS Potsdam dataset, ISPRS Vaihingen dataset, and LoveDA dataset. Results show that the learned views obtain highly advantages compared to the hand-crafted views and other optimized views. The code associated with this paper has been released and can be accessed at https://github.com/AAAA-CS/AMView. Cheng Shi 0002, Peiwen Han, Minghua Zhao, Qiguang Miao, Chi-Man Pun |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | LightFormer: Vision Transformer for Lightweight Based on Cascade Depthwise Convolution and Mixed Attention
Falin Wang, Jian Ji 0002, Yuan Wang 0081, Zeye Xu, Qiguang Miao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2025 | Adaptive Pitfall: Exploring the Effectiveness of Adaptation in Skeleton-Based Action RecognitionabstractGraph convolution networks (GCNs) have achieved remarkable performance in skeleton-based action recognition by exploiting the adjacency topology of body representation. However, the adaptive strategy adopted by the previous methods to construct the adjacency matrix is not balanced between the performance and the computational cost. We assume this concept ofAdaptive Trap, which can be replaced by multiple autonomous submodules, thereby simultaneously enhancing the dynamic joint representation and effectively reducing network resources. To effectuate the substitution of the adaptive model, we unveil two distinct strategies, both yielding comparable effects. (1) Optimization.Individuality and Commonality GCNs (IC-GCNs)is proposed to specifically optimize the construction method of the associativity adjacency matrix for adaptive processing. The uniqueness and co-occurrence between different joint points and frames in the skeleton topology are effectively captured through methodologies like preferential fusion of physical information, extreme compression of multi-dimensional channels, and simplification of self-attention mechanism. (2) Replacement.Auto-Learning GCNs (AL-GCNs)is proposed to boldly remove popular adaptive modules and cleverly utilize human key points as motion compensation to provide dynamic correlation support. AL-GCNs construct a fully learnable group adjacency matrix in both spatial and temporal dimensions, resulting in an elegant and efficient GCN-based model. In addition, three effective tricks for skeleton-based action recognition (Skip-Block, Bayesian Weight Selection Algorithm, and Simplified Dimensional Attention) are exposed and analyzed in this paper. Finally, we employ the variable channel and grouping method to explore the hardware resource bound of the two proposed models. IC-GCN and AL-GCN exhibit impressive performance across NTU-RGB+D 60, NTU-RGB+D 120, NW-UCLA, and UAV-Human datasets, with an exceptional parameter-cost ratio. Qiguang Miao, Wentian Xin, Ruyi Liu 0001, Cheng Shi 0002, Chi-Man Pun |
IEEE Trans. Multim. | 1 |
| 2025 | Seeking a Hierarchical Prototype for Multimodal Gesture RecognitionabstractGesture recognition has drawn considerable attention from many researchers owing to its wide range of applications. Although significant progress has been made in this field, previous works always focus on how to distinguish between different gesture classes, ignoring the influence of inner-class divergence caused by gesture-irrelevant factors. Meanwhile, for multimodal gesture recognition, feature or score fusion in the final stage is a general choice to combine the information of different modalities. Consequently, the gesture-relevant features in different modalities may be redundant, whereas the complementarity of modalities is not exploited sufficiently. To handle these problems, we propose a hierarchical gesture prototype framework to highlight gesture-relevant features such as poses and motions in this article. This framework consists of a sample-level prototype and a modal-level prototype. The sample-level gesture prototype is established with the structure of a memory bank, which avoids the distraction of gesture-irrelevant factors in each sample, such as the illumination, background, and the performers' appearances. Then the modal-level prototype is obtained via a generative adversarial network (GAN)-based subnetwork, in which the modal-invariant features are extracted and pulled together. Meanwhile, the modal-specific attribute features are used to synthesize the feature of other modalities, and the circulation of modality information helps to leverage their complementarity. Extensive experiments on three widely used gesture datasets demonstrate that our method is effective to highlight gesture-relevant features and can outperform the state-of-the-art methods. Yunan Li 0001, Tianyu Qi, Zhuoqi Ma, Dou Quan, Qiguang Miao |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | One-Nearest Neighborhood Guides Inlier Estimation for Unsupervised Point Cloud RegistrationabstractThe precision of unsupervised point cloud registration methods is typically limited by the lack of reliable inlier estimation and self-supervised signal, especially in partially overlapping scenarios. In this article, we propose an effective inlier estimation method for unsupervised point cloud registration by capturing geometric structure consistency between the source point cloud and its corresponding reference point cloud copy. Specifically, to obtain a high-quality reference point cloud copy, a one-nearest neighborhood (1-NN) point cloud is generated by input point cloud, which facilitates matching map construction and allows for integrating dual neighborhood matching scores of 1-NN point cloud and input point cloud to improve matching confidence. Benefiting from the high-quality reference copy, we argue that the neighborhood graph formed by inlier and its neighborhood should have consistency between source point cloud and its corresponding reference copy. Based on this observation, we construct transformation-invariant geometric structure representations and capture geometric structure consistency to score the inlier confidence for estimated correspondences between source point cloud and its reference copy. This strategy can simultaneously provide the reliable self-supervised signals for model optimization. Finally, we further calculate transformation estimation by the weighted SVD algorithm with the estimated correspondences and the corresponding inlier confidence. We train the proposed model in an unsupervised manner, and extensive experiments on synthetic and real-world datasets illustrate the effectiveness of the proposed method. Yongzhe Yuan, Yue Wu 0004, Maoguo Gong, Qiguang Miao, A. K. Qin 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | M3SOT: Multi-Frame, Multi-Field, Multi-Space 3D Single Object Trackingabstract3D Single Object Tracking (SOT) stands a forefront task of computer vision, proving essential for applications like autonomous driving. Sparse and occluded data in scene point clouds introduce variations in the appearance of tracked objects, adding complexity to the task. In this research, we unveil M3SOT, a novel 3D SOT framework, which synergizes multiple input frames (template sets), multiple receptive fields (continuous contexts), and multiple solution spaces (distinct tasks) in ONE model. Remarkably, M3SOT pioneers in modeling temporality, contexts, and tasks directly from point clouds, revisiting a perspective on the key factors influencing SOT. To this end, we design a transformer-based network centered on point cloud targets in the search area, aggregating diverse contextual representations and propagating target cues by employing historical frames. As M3SOT spans varied processing perspectives, we've streamlined the network—trimming its depth and optimizing its structure—to ensure a lightweight and efficient deployment for SOT applications. We posit that, backed by practical construction, M3SOT sidesteps the need for complex frameworks and auxiliary components to deliver sterling results. Extensive experiments on benchmarks such as KITTI, nuScenes, and Waymo Open Dataset demonstrate that M3SOT achieves state-of-the-art performance at 38 FPS. Our code and models are available at https://github.com/ywu0912/TeamCode.git. Yue Wu 0004, Maoguo Gong, Qiguang Miao, Wenping Ma 0001, Can Qin |
AAAI | 4 |
| 2024 | Evolutionary Multitasking with Compatibility Graph for Point Cloud Registrationabstract3D point cloud registration is a fundamental task in computer vision, aimed at estimating a transformation to align a pair of point clouds. For point cloud registration, where the popular methods are used to build a compatibility graph. Different methods of constructing compatibility graphs can result in different information regions to be searched in the graph, due to varying constraints. Evolutionary multitask optimization has gained attention in the field of evolutionary computation, as it enables knowledge transfer among multitasks to enhance the exploration of information. Inspired by this theory, this paper proposes a method to utilize compatibility graphs as tasks through evolutionary multitasking for solving the problem of point cloud registration. We first construct two compatibility graph tasks with different tightness constraints for point cloud registration. Due to the different constraints of the two proposed tasks, the local information of the graph of interest might be biased. This bias can be effectively utilized in the evolutionary multitasking framework to enhance the ability to discover meaningful consensus relationships in the search space of the graph. Then we map the two tasks to a unified search space by designing clique selection strategies, and carrying out knowledge transfer between the two tasks, aiming to emphasize more on the local consensus information in the graphs. Lastly, the efficacy of our proposed method is further validated on multiple registration datasets. Hangqi Ding, Yue Wu 0004, Hao Li 0009, Maoguo Gong, Wenping Ma 0001, Qiguang Miao |
CEC | 7 |
| 2024 | EFormer: Enhanced Transformer Towards Semantic-Contour Features of Foreground for Portraits MattingabstractThe portrait matting task aims to extract an alpha matte with complete semantics and finely detailed contours. In comparison to CNN-based approaches, transformers with self-attention module have a better capacity to capture long-range dependencies and low-frequency semantic information of a portrait. However, recent research shows that the self-attention mechanism struggles with modeling high-frequency contour information and capturing fine contour details, which can lead to bias while predicting the portrait's contours. To deal with this issue, we propose EFormer to enhance the model's attention towards both the low-frequency semantic and high-frequency contour features. For the high-frequency contours, our research demonstrates that cross-attention module between different resolutions can guide our model to allocate attention appropriately to these contour regions. Supported by this, we can successfully extract the high-frequency detail information around the portrait's contours, which were previously ignored by self-attention. Based on the cross-attention module, we further build a semantic and contour detector (SCD) to accurately capture both the low-frequency semantic and high-frequency contour features. And we design a contour-edge extraction branch and semantic extraction branch to extract refined high-frequency contour features and complete low-frequency semantic information, respectively. Finally, we fuse the two kinds of features and leverage the segmentation head to generate a predicted portrait matte. Experiments on VideoMatte240K (JPEG SD Format) and Adobe Image Matting (AIM) datasets demonstrate that EFormer outperforms previous portrait matte methods. Zitao Wang, Qiguang Miao |
CVPR | 2 |
| 2024 | Inlier Confidence Calibration for Point Cloud RegistrationabstractInliers estimation constitutes a pivotal step in partially overlapping point cloud registration. Existing methods broadly obey coordinate-based scheme, where inlier con-fidence is scored through simply capturing coordinate differences in the context. However, this scheme results in massive inlier misinterpretation readily, consequently affecting the registration performance. In this paper, we explore to extend a new definition called inlier confidence calibration (ICC) to alleviate the above issues. Firstly, we provide finely initial correspondences for ICC in order to generate high quality reference point cloud copy corresponding to the source point cloud. In particular, we develop a soft assignment matrix optimization theorem that offers faster speed and greater precision compared to Sinkhorn. Benefiting from the high quality reference copy, we argue the neighborhood patch formed by inlier and its neighborhood should have consistency between source point cloud and its reference copy. Based on this insight, we construct transformation-invariant geometric constraints and capture geometric structure consistency to calibrate inlier confidence for estimated correspondences between source point cloud and its reference copy. Finally, transformation is further calculated by the weighted SVD algorithm with the calibrated inlier confidence. Our model is trained in an unsupervised manner, and extensive experiments on synthetic and real-world datasets illustrate the effectiveness of the proposed method. Yongzhe Yuan, Yue Wu 0004, Xiaolong Fan, Maoguo Gong, Qiguang Miao, Wenping Ma 0001 |
CVPR | 5 |
| 2024 | Watching it in Dark: A Target-Aware Representation Learning Framework for High-Level Vision Tasks in Low Illumination
Yunan Li 0001, Shoude Li, Dou Quan, Chaoneng Li, Qiguang Miao |
ECCV (75) | 7 |
| 2024 | Evolutionary Multitasking with Two-level Knowledge Transfer for Multi-view Point Cloud RegistrationabstractPoint cloud registration is a hot research topic in the field of computer vision. In recent years, the registration method based on evolutionary computation has attracted more and more attention because of its robustness to initial pose and flexibility of objective function design. However, most of the current evolutionary computation-based point cloud registration methods do not take into account the multi-view problem, that is, to capture the close relationship between point clouds from different perspectives. We fully realize that if these relations are used correctly, the registration performance can be improved. Therefore, this paper proposes an evolutionary multitasking multi-view point cloud registration method, which solves the problem of multi-view error accumulation. To ensure the unity of global and local, a two-level knowledge transfer strategy is proposed, which divides the multi-view cloud registration task into two levels. This strategy unifies the search space of two registration tasks, solves the negative transfer phenomenon, and avoids the problem of falling into the local optimum. Finally, the effectiveness of the method is verified by sufficient experiments. This method has strong robustness to noise and outliers, and can be effectively implemented in various registration scenarios. Hangqi Ding, Haoran Xu 0005, Yue Wu 0004, Hao Li 0009, Maoguo Gong, Wenping Ma 0001, Qiguang Miao, Jiao Shi, Yu Lei 0002 |
GECCO | 7 |
| 2024 | Fusion of Independent and Interactive Features for Human-Object Interaction DetectionabstractHuman-Object Interaction (HOI) detection, which aims to identify humans and objects with interactive behaviors in images and predict the behaviors between them, is of great significance for semantic understanding. The existing works primarily focus on exploring the fine-grained semantic features of humans and objects, as well as the spatial relationships between them. However, these methods do not leverage the contextual information within the interaction area, which could potentially be valuable for predicting interaction behavior. To investigate the impact of contextual information on behavior prediction, we propose a novel approach to extract both independent and interactive features and fuse them. Specifically, our method is capable of extracting interaction features from the interaction region. These features are then merged with fine-grained independent features of humans and objects. Finally, the fused features are utilized to predict interaction behavior. In addition, the feature fusion module does not add extra storage and computation costs to our method. Experiments demonstrate the effectiveness of our method, achieving state-of-the-art performance on two benchmark HOI datasets, namely HCO-DET and V-COCO. Zehai Wu, Lijie Sheng, Songnian Zhang, Qiguang Miao |
ICIP | 4 |
| 2024 | PointMC: Multi-instance Point Cloud Registration based on Maximal CliquesabstractMulti-instance point cloud registration is the problem of estimating multiple rigid transformations between two point clouds. Existing solutions rely on global spatial consistency of ambiguity and the time-consuming clustering of highdimensional correspondence features, making it difficult to handle registration scenarios where multiple instances overlap. To address these problems, we propose a maximal clique based multiinstance point cloud registration framework called PointMC. The key idea is to search for maximal cliques on the correspondence compatibility graph to estimate multiple transformations, and cluster these transformations into clusters corresponding to different instances to efficiently and accurately estimate all poses. PointMC leverages a correspondence embedding module that relies on local spatial consistency to effectively eliminate outliers, and the extracted discriminative features empower the network to circumvent missed pose detection in scenarios involving multiple overlapping instances. We conduct comprehensive experiments on both synthetic and real-world datasets, and the results show that the proposed PointMC yields remarkable performance improvements. Yue Wu 0004, Xidao Hu, Yongzhe Yuan, Xiaolong Fan, Maoguo Gong, Hao Li 0009, Mingyang Zhang 0002, Qiguang Miao, Wenping Ma 0001 |
ICML | 8 |
| 2024 | Structural Entities Extraction and Patient Indications Incorporation for Chest X-Ray Report Generation
Kang Liu 0025, Zhuoqi Ma, Xiaolu Kang, Zhusi Zhong, Zhicheng Jiao, Grayson Baird, Harrison X. Bai, Qiguang Miao |
MICCAI (3) | 8 |
| 2024 | MGR-Dark: A Large Multimodal Video Dataset and RGB-IR Benchmark for Gesture Recognition in Darkness
Yunan Li 0001, Siyu Liang 0002, Huizhou Chen, Qiguang Miao |
ACM Multimedia | 5 |
| 2024 | CISampler: Correlated Information Guided Frame Sampling for Gesture Recognition in Video
Yunan Li 0001, Huizhou Chen, Siyu Liang 0002, Qiguang Miao |
MMAsia | 5 |
| 2024 | Language-Skeleton Pre-training to Collaborate with Self-Supervised Human Action Recognition
Ruyi Liu 0001, Wentian Xin, Qiguang Miao, Yuzhi Hu, Jiahao Qi |
PRCV (7) | 4 |
| 2024 | Flow-Audio-Synth: A Video-to-Audio Model which Captures Dynamic Features
Yupeng Zheng, Zixiang Lu, Qiguang Miao, Xiangzeng Liu |
PRCV (10) | 4 |
| 2024 | Deep Dual Graph attention Auto-Encoder for community detection
Xunlian Wu, Wanying Lu, Yi-Ning Quan, Qiguang Miao, Peng Gang Sun |
Expert Syst. Appl. | 4 |
| 2024 | Explore Bayesian analysis in Cognitive-aware Key-Value Memory Networks for knowledge tracing in online learning
Juli Zhang, Ruoheng Xia, Qiguang Miao, Quan Wang 0006 |
Expert Syst. Appl. | 3 |
| 2024 | DiffTAD: Denoising diffusion probabilistic models for vehicle trajectory anomaly detection
Chaoneng Li, Guanwen Feng, Yunan Li 0001, Ruyi Liu 0001, Qiguang Miao, Liang Chang 0003 |
Knowl. Based Syst. | 5 |
| 2024 | Towards Unified Robustness Against Both Backdoor and Adversarial AttacksabstractDeep Neural Networks (DNNs) are known to be vulnerable to both backdoor and adversarial attacks. In the literature, these two types of attacks are commonly treated as distinct robustness problems and solved separately, since they belong to training-time and inference-time attacks respectively. However, this paper revealed that there is an intriguing connection between them: (1) planting a backdoor into a model will significantly affect the model's adversarial examples and (2) for an infected model, its adversarial examples have similar features as the triggered images. Based on these observations, a novel Progressive Unified Defense (PUD) algorithm is proposed to defend against backdoor and adversarial attacks simultaneously. Specifically, our PUD has a progressive model purification scheme to jointly erase backdoors and enhance the model's adversarial robustness. At the early stage, the adversarial examples of infected models are utilized to erase backdoors. With the backdoor gradually erased, our model purification can naturally turn into a stage to boost the model's robustness against adversarial attacks. Besides, our PUD algorithm can effectively identify poisoned images, which allows the initial extra dataset not to be completely clean. Extensive experimental results show that, our discovered connection between backdoor and adversarial attacks is ubiquitous, no matter what type of backdoor attack. The proposed PUD outperforms the state-of-the-art backdoor defense, including the model repairing-based and data filtering-based methods. Besides, it also has the ability to compete with the most advanced adversarial defense methods. The code is available here. Zhenxing Niu, Yuyao Sun, Qiguang Miao, Rong Jin 0001, Gang Hua 0001 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2024 | Attack-invariant attention feature for adversarial defense in hyperspectral image classification
Cheng Shi 0002, Minghua Zhao, Chi-Man Pun, Qiguang Miao |
Pattern Recognit. | 5 |
| 2024 | Exploration of Class Center for Fine-Grained Visual ClassificationabstractDifferent from large-scale classification tasks, fine-grained visual classification is a challenging task due to two critical problems: 1) evident intra-class variances and subtle inter-class differences, and 2) overfitting owing to fewer training samples in datasets. Most existing methods extract key features to reduce intra-class variances, but pay no attention to subtle inter-class differences in fine-grained visual classification. To address this issue, we propose a loss function named exploration of class center, which consists of a multiple class-center constraint and a class-center label generation. This loss function fully utilizes the information of the class center from the perspective of features and labels. From the feature perspective, the multiple class-center constraint pulls samples closer to the target class center, and pushes samples away from the most similar nontarget class center. Thus, the constraint reduces intra-class variances and enlarges inter-class differences. From the label perspective, the class-center label generation utilizes class-center distributions to generate soft labels to alleviate overfitting. Our method can be easily integrated with existing fine-grained visual classification approaches as a loss function, to further boost excellent performance with only slight training costs. Extensive experiments are conducted to demonstrate consistent improvements achieved by our method on four widely-used fine-grained visual classification datasets. In particular, our method achieves state-of-the-art performance on the FGVC-Aircraft and CUB-200-2011 datasets. Hang Yao 0001, Qiguang Miao, Chaoneng Li, Guanwen Feng, Ruyi Liu 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2024 | Learning Discriminative Features via Multi-Hierarchical Mutual Information for Unsupervised Point Cloud RegistrationabstractExtracting discriminative representations is the key step for correspondence-free point cloud registration. The extracted representations require to be discriminative to transformation, which demands representations to reduce the influence of redundant information irrelevant to transformation. However, recently proposed methods ignore this crucial property, resulting in limited ability to represent point cloud. In addition, researching correspondence-free point cloud registration has stagnated in recent years. In this paper, we try to relieve features redundancy issue for correspondence-free point cloud registration from a new perspective. Specifically, our method comprises two stages: feature extraction stage and rigid body transformation stage. In feature extraction stage, we aim to maximize multi-hierarchical mutual information between different hierarchical features, which can provide discriminative and less redundancy representations to regress transformation parameters for next stage. In rigid body transformation stage, we utilize dual quaternion to estimate transformation parameters, which combines rotation and translation simultaneously within a unified framework and obtains a compact representations for rigid transformation. The proposed model is trained in an unsupervised manner on the ModelNet40 dataset. The experimental results illustrate that our method achieves higher accuracy and robustness compared with existing correspondence-free methods. Yongzhe Yuan, Yue Wu 0004, Mingyu Yue, Maoguo Gong, Xiaolong Fan, Wenping Ma 0001, Qiguang Miao |
IEEE Trans. Circuits Syst. Video Technol. | 7 |
| 2024 | Evolutionary Multiform Optimization With Two-Stage Bidirectional Knowledge Transfer Strategy for Point Cloud RegistrationabstractPoint cloud registration is an important task in computer vision, where the goal is to estimate a transformation to align a pair of point clouds. Most of the existing registration methods face the problems of poor robustness and getting stuck in local optima. Evolutionary multitasking is an effective paradigm to enhance global search capability and improve convergence characteristics through knowledge transfer across multiple related tasks. Inspired by evolutionary multitasking, this article proposes a multiform optimization approach through evolutionary multitasking for solving the point cloud registration problems. We first construct two related registration tasks with different functional landscapes to form a multiform optimization problem. Compared with methods that only focus on a single registration attribute, the two proposed tasks focus on robustness and precision of registration, respectively. Then, a new two-stage bidirectional knowledge transfer strategy is presented, which can implement efficient knowledge transfer among two related tasks. Finally, both simulations and real experiments show the power of our method. The proposed method is robust to noise, outliers, and partial overlaps and is effective in multiple real registration scenarios, such as object registration, scene reconstruction, and simultaneous localization and mapping. Yue Wu 0004, Hangqi Ding, Maoguo Gong, A. K. Qin 0001, Wenping Ma 0001, Qiguang Miao, Kay Chen Tan |
IEEE Trans. Evol. Comput. | 6 |
| 2024 | Masked Topology Convolutional Network for Classification and Segmentation of Remote Sensing ImagesabstractConvolutional Neural Networks have made significant progress in remote sensing image processing. Convolutional networks mostly model the local information of samples based on pixel features, while ignoring the topological relationship among different categories of ground objects. But the non-local topology can better represent the underlying data structure of the image. In order to make up for the insufficiency of Convolutional Neural Networks in extracting non-local topological information, we proposes a new Masked Topology Convolutional Networks (MTC_Net). First, considering the regionality of different categories of ground objects, we use a clustering algorithm to simply cluster the pixels to extract the most obvious area mask, which can strengthen the boundary information among different ground objects. Second, we take the cluster center point as the representative point of the region, and use the graph structure to represent the topological relationship, in addition use graph convolution to further extract the topological structure relationship among different categories of ground objects. Finally, we use convolution modules and multi-level feature attention modules to capture important local and global contextual semantic information and reduce category confusion. We conduct a great deal of experiments on four (two classification and two segmentation) public datasets, and the proposed MTC_Net achieve excellent classification and segmentation performance. Falin Wang, Jian Ji 0002, Yuan Wang 0081, Qiguang Miao |
IEEE Trans. Geosci. Remote. Sens. | 5 |
| 2024 | Detection-Driven Exposure-Correction Network for Nighttime Drone-View Object DetectionabstractDrone-view object detection (DroneDet) models typically suffer a significant performance drop when applied to nighttime scenes. Existing solutions attempt to employ an exposure-adjustment module to reveal objects hidden in dark regions before detection. However, most exposure-adjustment models are only optimized for human perception, where the exposure-adjusted images may not necessarily enhance recognition. To tackle this issue, we propose a novel Detection-driven Exposure-correction network for nighttime DroneDet, called DEDet. The DEDet conducts adaptive, nonlinear adjustment of pixel values in a spatially fine-grained manner to generate DroneDet-friendly images. Specifically, we develop a fine-grained parameter predictor (FPP) to estimate pixelwise parameter maps of the image filters. These filters, along with the estimated parameters, are used to adjust pixel values of the low-light image based on nonuniform illuminations in drone-captured images. In order to learn the nonlinear transformation from the original nighttime images to their DroneDet-friendly counterparts, we propose a progressive filtering module that applies recursive filters to iteratively refine the exposed image. Furthermore, to evaluate the performance of the proposed DEDet, we have built a dataset NightDrone to address the scarcity of the datasets specifically tailored for this purpose. Extensive experiments conducted on four nighttime datasets show that DEDet achieves a superior accuracy compared with the state-of-the-art (SOTA) methods. Furthermore, ablation studies and visualizations demonstrate the validity and interpretability of our approach. Our NightDrone dataset can be downloaded fromhttps://github.com/yuexiemail/NightDrone-Dataset. Wenjing Jia, Qiguang Miao, Junmei Feng, Jinchang Ren |
IEEE Trans. Geosci. Remote. Sens. | 3 |
| 2024 | Pose-Guided Attention Learning for Cloth-Changing Person Re-IdentificationabstractThe change in appearance is a great challenge for cloth-changing person re-identification. Existing methods tackle this challenge by learning the shape features of the human body, however, these features are easily affected by the human pose or camera perspective. Thus, the ability to learn and extract the invariant features of person in varying conditions is crucial to overcome the above challenge. To address the issue of invariant feature extraction for cloth-changing person re-identification, a Pose-Guided Attention Learning (PGAL) framework is proposed in this paper. First, we introduce the human pose estimation network to remove the background effects and align the fine-grained key points features of human body. Then, to fully exploit the available appearance information, we develop a Feature Enhancement Module (FEM) that improves the feature representation of non-key point regions of human body through the Multi-Head Self Attention. Finally, in order to adaptively learn the invariant features of the person, we construct an Attention Learning Module (ALM) to achieve automatic selection of multi-granularity features by utilizing three different loss functions. Comparing with current popular methods on four cloth-changing person Re-ID datasets, the experimental results show the superiority of our method. Xiangzeng Liu, Yi-Ning Quan, Qiguang Miao |
IEEE Trans. Multim. | 6 |
| 2024 | Inter-Modal Masked Autoencoder for Self-Supervised Learning on Point CloudsabstractMasked autoencoder (MAE) is a recently widely used self-supervised learning method that has achieved great success in NLP and computer vision. However, the potential advantages of masked pre-training for point cloud understanding have not been fully explored. There is preliminary work on MAE-based point clouds using the Transformer architecture to explore low-level geometric representations in 3D space, which is insufficient for fine-grained decoding completion and downstream tasks. Inspired by multimodality, we propose Inter-MAE, a inter-modal MAE method for self-supervised learning on point clouds. Specifically, we first use Point-MAE as a baseline to partition point clouds into random low percentage of visible and high percentage of masked point patches. Then, a standard Transformer-based autoencoder is built by asymmetric design and shifting mask operations, and latent features are learned from the visible point patches aiming to recover the masked point patches. In addition, we generate image features based on ViT after point cloud rendering to form inter-modal contrastive learning with the decoded features of the completed point patches. Extensive experiments show that the proposed Inter-MAE generates pre-trained models that are effective and exhibit superior results in various downstream tasks. For example, an accuracy of 85.4% is achieved on ScanObjectNN and 86.3% on ShapeNetPart, outperforming other state-of-the-art self-supervised learning methods. Notably, our work establishes for the first time the feasibility of applying image modality to masked point clouds. Yue Wu 0004, Maoguo Gong, Zhixiao Liu, Qiguang Miao, Wenping Ma 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | A Semi-Supervised Underexposed Image Enhancement Network With Supervised Context Attention and Multi-Exposure FusionabstractRecently, image enhancement approaches yield impressive progress. However, most methods are still based supervised-learning, which requires plenty of paired data. Meanwhile, owing to the complex illumination condition in a real-world scenario, those methods trained on synthetic images cannot restore details in extremely dark or bright areas and lead to exposure errors. The traditional losses that deem all pixels the same in training also produce blurry edges in the result. To handle these problems, in this article, we present an effective semi-supervised framework for severely underexposed image enhancement. Our network consists of a supervised and an unsupervised branch, which shares weights and can make full use of paired data and plenty of unpaired data. Meanwhile, a multi-exposure fusion module is designed to adaptively fuse the corrected images to address the low contrast and color bias issues occurring in some extreme situations. Moreover, we propose a supervised context attention module to better use the edge information as supervision to recover fine image details. Extensive experiments have proved that the proposed method outperforms state-of-the-art approaches in enhancing exposure images. Xiaolong Fu, Yunan Li 0001, Kaibin Miao, Xiangzeng Liu, Bocheng Zhao, Qiguang Miao |
IEEE Trans. Multim. | 7 |
| 2024 | Self-Supervised Intra-Modal and Cross-Modal Contrastive Learning for Point Cloud UnderstandingabstractLearning effective representations from unlabeled data is a challenging task for point cloud understanding. As the human visual system can map concepts learned from 2D images to the 3D world, and inspired by recent multimodal research, we introduce data from point cloud modality and image modality for joint learning. Based on the properties of point clouds and images, we propose CrossNet, a comprehensive intra- and cross-modal contrastive learning method that learns 3D point cloud representations. The proposed method achieves 3D-3D and 3D-2D correspondences of objectives by maximizing the consistency of point clouds and their augmented versions, and with the corresponding rendered images in invariant space. We further distinguish the rendered images into RGB and grayscale images to extract color and geometric features, respectively. These training objectives combine feature correspondences between modalities to combine rich learning signals from point clouds and images. Our CrossNet is simple: we add a feature extraction module and a projection head module to the point cloud and image branches, respectively, to train the backbone network in a self-supervised manner. After the network is pretrained, only the point cloud feature extraction module is required for fine-tuning and directly predicting results for downstream tasks. Our experiments on multiple benchmarks demonstrate improved point cloud classification and segmentation results, and the learned representations can be generalized across domains. Yue Wu 0004, Maoguo Gong, Peiran Gong, Xiaolong Fan, A. K. Qin 0001, Qiguang Miao, Wenping Ma 0001 |
IEEE Trans. Multim. | 7 |
| 2024 | MPCT: Multiscale Point Cloud Transformer With a Residual NetworkabstractThe self-attention (SA) network revisits the essence of data and has achieved remarkable results in text processing and image analysis. SA is conceptualized as a set operator that is insensitive to the order and number of data, making it suitable for point sets embedded in 3D space. However, working with point clouds still poses challenges. To tackle the issue of exponential growth in complexity and singularity induced by the original SA network without position encoding, we modify the attention mechanism by incorporating position encoding to make it linear, thus reducing its computational cost and memory usage and making it more feasible for point clouds. This article presents a new framework called multiscale point cloud transformer (MPCT), which improves upon prior methods in cross-domain applications. The utilization of multiple embeddings enables the complete capture of the remote and local contextual connections within point clouds, as determined by our proposed attention mechanism. Additionally, we use a residual network to facilitate the fusion of multiscale features, allowing MPCT to better comprehend the representations of point clouds at each stage of attention. Experiments conducted on several datasets demonstrate that MPCT outperforms the existing methods, such as achieving accuracies of 94.2% and 84.9% in classification tasks implemented on ModelNet40 and ScanObjectNN, respectively. Yue Wu 0004, Maoguo Gong, Zhixiao Liu, Qiguang Miao, Wenping Ma 0001 |
IEEE Trans. Multim. | 5 |
| 2024 | Deep Reinforcement Learning: A SurveyabstractDeep reinforcement learning (DRL) integrates the feature representation ability of deep learning with the decision-making ability of reinforcement learning so that it can achieve powerful end-to-end learning control capabilities. In the past decade, DRL has made substantial advances in many tasks that require perceiving high-dimensional input and making optimal or near-optimal decisions. However, there are still many challenging problems in the theory and applications of DRL, especially in learning control tasks with limited samples, sparse rewards, and multiple agents. Researchers have proposed various solutions and new theories to solve these problems and promote the development of DRL. In addition, deep learning has stimulated the further development of many subfields of reinforcement learning, such as hierarchical reinforcement learning (HRL), multiagent reinforcement learning, and imitation learning. This article gives a comprehensive overview of the fundamental theories, key algorithms, and primary research domains of DRL. In addition to value-based and policy-based DRL algorithms, the advances in maximum entropy-based DRL are summarized. The future research topics of DRL are also analyzed and discussed. Xu Wang 0043, Xingxing Liang, Dawei Zhao 0003, Jincai Huang 0001, Xin Xu 0001, Bin Dai 0001, Qiguang Miao |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2024 | RORNet: Partial-to-Partial Registration Network With Reliable Overlapping RepresentationsabstractThree-dimensional point cloud registration is an important field in computer vision. Recently, due to the increasingly complex scenes and incomplete observations, many partial-overlap registration methods based on overlap estimation have been proposed. These methods heavily rely on the extracted overlapping regions with their performances greatly degraded when the overlapping region extraction underperforms. To solve this problem, we propose a partial-to-partial registration network (RORNet) to find reliable overlapping representations from the partially overlapping point clouds and use these representations for registration. The idea is to select a small number of key points called reliable overlapping representations from the estimated overlapping points, reducing the side effect of overlap estimation errors on registration. Although it may filter out some inliers, the inclusion of outliers has a much bigger influence than the omission of inliers on the registration task. The RORNet is composed of overlapping points' estimation module and representations' generation module. Different from the previous methods of direct registration after extraction of overlapping areas, RORNet adds the step of extracting reliable representations before registration, where the proposed similarity matrix downsampling method is used to filter out the points with low similarity and retain reliable representations, and thus reduce the side effects of overlap estimation errors on the registration. Besides, compared with previous similarity-based and score-based overlap estimation methods, we use the dual-branch structure to combine the benefits of both, which is less sensitive to noise. We perform overlap estimation experiments and registration experiments on the ModelNet40 dataset, outdoor large scene dataset KITTI, and natural data Stanford Bunny dataset. The experimental results demonstrate that our method is superior to other partial registration methods. Our code is available at https://github.com/superYuezhang/RORNet. Yue Wu 0004, Yue Zhang 0040, Wenping Ma 0001, Maoguo Gong, Xiaolong Fan, Mingyang Zhang 0002, A. K. Qin 0001, Qiguang Miao |
IEEE Trans. Neural Networks Learn. Syst. | 8 |
| 2024 | EGST: Enhanced Geometric Structure Transformer for Point Cloud RegistrationabstractWe explore the effect of geometric structure descriptors on extracting reliable correspondences and obtaining accurate registration for point cloud registration. The point cloud registration task involves the estimation of rigid transformation motion in unorganized point cloud, hence it is crucial to capture the contextual features of the geometric structure in point cloud. Recent coordinates-only methods ignore numerous geometric information in the point cloud which weaken ability to express the global context. We propose Enhanced Geometric Structure Transformer to learn enhanced contextual features of the geometric structure in point cloud and model the structure consistency between point clouds for extracting reliable correspondences, which encodes three explicit enhanced geometric structures and provides significant cues for point cloud registration. More importantly, we report empirical results that Enhanced Geometric Structure Transformer can learn meaningful geometric structure features using none of the following: (i) explicit positional embeddings, (ii) additional feature exchange module such as cross-attention, which can simplify network structure compared with plain Transformer. Extensive experiments on the synthetic dataset and real-world datasets illustrate that our method can achieve competitive results. Yongzhe Yuan, Yue Wu 0004, Xiaolong Fan, Maoguo Gong, Wenping Ma 0001, Qiguang Miao |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2023 | Progressive Backdoor Erasing via connecting Backdoor and Adversarial AttacksabstractDeep neural networks (DNNs) are known to be vulnera-ble to both backdoor attacks as well as adversarial attacks. In the literature, these two types of attacks are commonly treated as distinct problems and solved separately, since they belong to training-time and inference-time attacks respectively. However, in this paper we find an intriguing connection between them: for a model planted with backdoors, we observe that its adversarial examples have similar behaviors as its triggered images, i.e., both activate the same subset of DNN neurons. It indicates that planting a back-door into a model will significantly affect the model's adversarial examples. Based on these observations, a novel Progressive Backdoor Erasing (PBE) algorithm is proposed to progressively purify the infected model by leveraging un-targeted adversarial attacks. Different from previous back-door defense methods, one significant advantage of our approach is that it can erase backdoor even when the clean extra dataset is unavailable. We empirically show that, against 5 state-of-the-art backdoor attacks, our PBE can effectively erase the backdoor without obvious performance degradation on clean samples and outperforms existing de-fense methods. Bingxu Mu, Zhenxing Niu, Le Wang 0003, Xue Wang 0010, Qiguang Miao, Rong Jin 0001, Gang Hua 0001 |
CVPR | 5 |
| 2023 | Learning Robust Representations with Information Bottleneck and Memory Network for RGB-D-based Gesture RecognitionabstractAlthough previous RGB-D-based gesture recognition methods have shown promising performance, researchers often overlook the interference of task-irrelevant cues like illumination and background. These unnecessary factors are learned together with the predictive ones by the network and hinder accurate recognition. In this paper, we propose a convenient and analytical framework to learn a robust feature representation that is impervious to gesture-irrelevant factors. Based on the Information Bottleneck theory, two rules of Sufficiency and Compactness are derived to develop a new information-theoretic loss function, which cultivates a more sufficient and compact representation from the feature encoding and mitigates the impact of gesture-irrelevant information. To highlight the predictive information, we further integrate a memory network. Using our proposed content-based and contextual memory addressing scheme, we weaken the nuisances while preserving the task-relevant information, providing guidance for refining the feature representation. Experiments conducted on three public datasets demonstrate that our approach leads to a better feature representation and achieves better performance than state-of-the-art methods. The code of our method is available at: https://github.com/Carpumpkin/InBoMem. Yunan Li 0001, Huizhou Chen, Guanwen Feng, Qiguang Miao |
ICCV | 4 |
| 2023 | Is Really Correlation Information Represented Well in Self-Attention for Skeleton-based Action Recognition?abstractTransformer has shown significant advantages by various vision tasks. However, the lack of representation of correlation information about data properties makes it difficult to match the excellent results consistent with GCNs in skeleton-based action recognition. In this paper, we propose a Topology and Frames-guided Spatial-Temporal ConvFormer Network (TF-STCFormer), which is well suited for dynamically extracting topological and inter-frame uniqueness & co-occurrence information. Three essential components make up the proposed framework: (1) Grouped Physical-guided Spatial Transformer for focusing on learning essential spatial features and physical topology. (2) Global and Focal Temporal Transformer for promoting the relationship of different joints in consecutive frames and improving the representation of discriminative key-frames. (3) Grouped Dilation Temporal Convolution for connecting the intermediate output obtained by the previous transformers in the feature channels of different dilation. Experiments on four standard datasets (NTU RGB+D, NTU RGB+D 120, NW-UCLA, and UAV-Human) demonstrate that our approach prominently outperforms state-of-the-art methods on all benchmarks. Wentian Xin, Hongkai Lin, Ruyi Liu 0001, Qiguang Miao |
ICME | 5 |
| 2023 | Exploring Dual Representations in Large-Scale Point Clouds: A Simple Weakly Supervised Semantic Segmentation FrameworkabstractExisting work shows that 3D point clouds produce only about a 4% drop in semantic segmentation even at 1% random point annotation, which inspires us to further explore how to achieve better results at lower cost. As scene point clouds provide position and color information and often used in tandem as the only input, with little work going into segmentation by fusing information from dual spaces. To optimize point cloud representations, we propose a novel framework for the dual representation query network (DRQNet). The proposed framework partitions the input point cloud into position and color spaces, using the separately extracted geometric structure and semantic context to create an internal supervisory mechanism that bridges the dual spaces and fuses the information. Adopting sparsely annotated points as the query set, DRQNet provide guidance and perceptual information for multi-stage point clouds through random sampling. More, to differentiate and enhance the features generated by local neighbourhoods within multiple perceptual fields, we design a representation selection module to identify the contributions made by the position and color of each query point, and weight them adaptively according to reliability. The proposed DRQNet is robust to point cloud analysis and eliminates the effects of irregularities and disorder. Our method achieves significant performance gains on three mainstream benchmarks. Yue Wu 0004, Maoguo Gong, Qiguang Miao, Wenping Ma 0001 |
ACM Multimedia | 4 |
| 2023 | Skeleton MixFormer: Multivariate Topology Representation for Skeleton-based Action RecognitionabstractVision Transformer, which performs well in various vision tasks, encounters a bottleneck in skeleton-based action recognition and falls short of advanced GCN-based methods. The root cause is that the current skeleton transformer depends on the self-attention mechanism of the complete channel of the global joint, ignoring the highly discriminative differential correlation within the channel, so it is challenging to learn the expression of the multivariate topology dynamically. To tackle this, we present Skeleton MixFormer, an innovative spatio-temporal architecture to effectively represent the physical correlations and temporal interactivity of the compact skeleton data. Two essential components make up the proposed framework: 1) Spatial MixFormer. The channel-grouping and mix-attention are utilized to calculate the dynamic multivariate topological relationships. Compared with the full-channel self-attention method, Spatial MixFormer better highlights the channel groups' discriminative differences and the joint adjacency's interpretable learning. 2) Temporal MixFormer, which consists of Multiscale Convolution, Temporal Transformer and Sequential Holding Module. The multivariate temporal models ensure the richness of global difference expression and realize the discrimination of crucial intervals in the sequence, thereby enabling more effective learning of long and short-term dependencies in actions. Our Skeleton MixFormer demonstrates state-of-the-art (SOTA) performance across seven different settings on four standard datasets, namely NTU-60, NTU-120, NW-UCLA, and UAV-Human. Related code will be available on https://github.com/ElricXin/Skeleton-MixFormer. Wentian Xin, Qiguang Miao, Ruyi Liu 0001, Chi-Man Pun, Cheng Shi 0002 |
ACM Multimedia | 2 |
| 2023 | Auto-Learning-GCN: An Ingenious Framework for Skeleton-Based Action Recognition
Wentian Xin, Ruyi Liu 0001, Qiguang Miao, Cheng Shi 0002, Chi-Man Pun |
PRCV (1) | 4 |
| 2023 | Personalized recommendation with hybrid feedback by refining implicit data
Junmei Feng, Kunwei Wang, Qiguang Miao, Zhaoqiang Xia |
Expert Syst. Appl. | 3 |
| 2023 | Adaptive bistable stochastic resonance based blind watermark extraction in discrete cosine transform domainabstractAbstract Blind watermark extraction in discrete cosine transform (DCT) domain has a wide application prospect as well as a challenging subject. The imperceptibility of watermark signal makes watermark extraction a weak signal reception issue in essence. For DCT coefficients of host image generally disobey Gaussian distribution, at which the performance of linear correlated reception is no longer optimal. Aiming at this, a novel blind watermark extraction scheme combining the uncorrelated reception with adaptive bistable stochastic resonance (ABSR) technique is proposed. First, by block DCT transformation for host image, an additive watermark embedding algorithm is introduced, in which the watermarked image can be converted to one dimensional time domain weak signal (binary watermark image) reception under additive Laplacian noise (selected DCT coefficients). On this basis, through the key technology research on quantitatively cooperative resonance relationship under Laplacian noise, the ABSR system can be implemented by bistable system parameters self‐adaptive adjustment, in which the ABSR system output signal will be enhanced rather than be weakened by random noise. Finally, the ABSR‐based watermark extraction scheme is investigated, and both the visual effect, bit error ratio performance and robustness of proposed scheme are testified to be superior to that of traditional uncorrelated extraction. Jin Liu 0031, Zan Li 0001, Qiguang Miao, Peihan Qi |
IET Image Process. | 3 |
| 2023 | Deep mutual learning for brain tumor segmentation with the fusion network
Qiguang Miao, Daikai Ma, Ruyi Liu 0001 |
Neurocomputing | 2 |
| 2023 | Focus on hierarchical features: Soft-weighted hierarchical features network
Hongkai Lin, Wentian Xin, Shun Chang, Qianxue Yang, Qiguang Miao, Ruyi Liu 0001, Liang Chang 0003 |
Neurocomputing | 5 |
| 2023 | Transformer for Skeleton-based action recognition: A review of recent advances
Wentian Xin, Ruyi Liu 0001, Qiguang Miao |
Neurocomputing | 6 |
| 2023 | Refined probability distribution module for fine-grained visual categorization
Qiguang Miao, Hongsheng Li 0001, Ruyi Liu 0001, Yi-Ning Quan, Jianfeng Song |
Neurocomputing | 2 |
| 2023 | SACF-Net: Skip-Attention Based Correspondence Filtering Network for Point Cloud RegistrationabstractRigid registration is a transformation estimation problem between two point clouds. The two point clouds captured may partially overlap owing to different viewpoints and acquisition times. Some previous correspondence matching based methods utilize an encoder-decoder network to carry out partial-to-partial registration task and adopt a skip-connection structure to convey information between the encoder and decoder. However, equally revisiting them with skip-connection may introduce the information redundancy, and limit the feature learning ability of the entire network. To address these problems, we propose a skip-attention based correspondence filtering network (SACF-Net) for point cloud registration. A novel feature interaction mechanism is designed to utilize both low-level geometric information and high-level context-aware information to enhance the original pointwise matching map. Additionally, a skip-attention based correspondence filtering method is proposed to selectively revisits features in the encoder at different resolutions, allowing the decoder to extract high-quality correspondences within overlapping regions. We conduct comprehensive experiments on indoor and outdoor scene datasets, and the results show that the proposed SACF-Net yields unprecedented performance improvements. Yue Wu 0004, Xidao Hu, Yue Zhang 0040, Maoguo Gong, Wenping Ma 0001, Qiguang Miao |
IEEE Trans. Circuits Syst. Video Technol. | 6 |
| 2023 | INENet: Inliers Estimation Network With Similarity Learning for Partial Overlapping RegistrationabstractPoint cloud registration is a key problem in the application of computer vision to robotics, autopilot and other fields. However, because the object is partially covered up or the resolution of 3D scanners is different, point clouds collected by the same sense may be inconsistent and even incomplete. Inspired by the recently proposed learning-based approaches, we propose Inliers Estimation Network (INENet) which includes a self-designed threshold prediction network and a probability estimation network with adaptive similarity mutual attention to help to find the overlapping area of the point clouds. In order to solve the above problems, we divide the partially overlapping point cloud registration task into two sub-tasks: overlapping areas detection and registration. The threshold prediction network can automatically calculate the threshold according to the input point clouds, and then the probability estimation network estimates the overlapping points by using threshold. The advantages of the proposed approach include: (1) threshold prediction network avoids bias and the complexity of manually adjusting the threshold. (2) Probability estimation network with similarity matrix can deeply fuse the information between a pair of point clouds, which is helpful to improve the accuracy. (3) INENet can be easily integrated into other overlapping region sensitive algorithms and without adjusting parameters. We conduct experiments on the ModelNet40, S3DIS and 3DMatch data sets. Specifically, the rotation error of the registration algorithm integrated with INENet is improved by at least 25% compared with direct partial overlap registration, our method improves the$F_{1} $score by 5% and has better anti-noise ability compared with the existing overlap detection methods, showing the effectiveness of the proposed method. Yue Wu 0004, Yue Zhang 0040, Xiaolong Fan, Maoguo Gong, Qiguang Miao, Wenping Ma 0001 |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2023 | Universal Object-Level Adversarial Attack in Hyperspectral Image ClassificationabstractThe vulnerability of deep neural networks has garnered significant attention. Various advanced adversarial attack methods have been proposed. However, these methods exhibit higher attack performance on three-band natural images while struggling to handle high-dimensional attacks in terms of attack transferability and robustness. Hyperspectral images, unlike natural images, possess high-dimensional and redundant spectral information. On one hand, different classification models focus on distinct discriminative spectral bands, leading to poor transferability. On the other hand, most existing attack methods are implemented at the pixel-level, making them less resilient to image processing-based defenses. In this paper, we address the improvement of transferability and robustness in high-dimensional attacks and introduce a universal object-level adversarial attack method in hyperspectral image classification. We found that perturbations with higher similarity in a local region can decrease the sensitivity of adversarial attacks to various discriminative spectral patterns and enhance resistance to image processing-based defenses. Consequently, we construct spatial and spectral oversegmented templates by utilizing the local smooth properties of hyperspectral images, aiming to promote similarity among perturbations within a local region. Extensive experiments conducted on two real hyperspectral image datasets validate that our method enhances the attack transferability and robustness of several existing attack methods. By incorporating the object-level adversarial attack with baseline fast gradient sign method (FGSM), momentum iterative FGSM (MI-FGSM), and variance tuning MI-FGSM (VMI-FGSM), the average transferability success rate of the proposed method has increased by 7.38% on the PaviaU dataset and 9.30% on the HoustonU 2018 dataset than the baselines, respectively. Meanwhile, the proposed method outperforms the baselines by an average of 6.19% on the PaviaU dataset and 10.05% on the HoustonU 2018 dataset in attacking image processing-based defense models. The code is available at https://github.com/AAAA-CS/SS_FGSM_HyperspectralAdversarialAttack. Cheng Shi 0002, Mengxin Zhang, Zhiyong Lv, Qiguang Miao, Chi-Man Pun |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2023 | Rearranging 'indivisible' Blocks for Community DetectionabstractUnattributed social networks are more complicated, and it tends not to determine the best division by over-optimizing a theoretical measure for unsupervised algorithms. Nowadays, communities strongly overlap due to the fact that people strongly interact, which makes community detection even more challenging. The paper develops a new algorithm by rearranging ‘indivisible’ blocks (RaidB). In RaidB, we first initialize ‘indivisible’ blocks by disjoint k-clique blocks in a network, and then these blocks are rearranged by moving nodes from one block to another based on maximizing modularity to uncover non-overlapping communities. For identifying overlapping communities, the above blocks are further rearranged, i.e., each block is subdivided and expanded to determine sub-blocks by introducing a dynamic linear threshold (DLT) model for influence interpenetration, and we finally determine a division from these sub-blocks with the minimum size that can cover the network. We compare RaidB with the existing state of the art methods for non-overlapping and overlapping community detection. The results show that RaidB tends to achieve better performance especially on sparse networks with unobvious communities and networks with strongly overlapping communities. Peng Gang Sun, Xunlian Wu, Yi-Ning Quan, Qiguang Miao |
IEEE Trans. Knowl. Data Eng. | 4 |
| 2023 | Image Hazing and Dehazing: From the Viewpoint of Two-Way Image Translation With a Weakly Supervised FrameworkabstractImage dehazing is an important task since it is the prerequisite for many downstream high-level computer vision tasks. Previous dehazing methods depend on either the hand-designed priors/assumptions or supervised learning with plenty of data, which are not easy to implement in practice. Meanwhile, synthesizing hazy images is also significant in many scenes like multi-weather image generation. In this paper, we change the viewpoint of this task to image translation and develop a weakly supervised framework to achieve it. Instead of simply considering the hazy image as the source domain and the haze-free image as the target domain for translation, we design a feature representation scheme that generates a domain indicator, and embed it into the decoder to achieve both hazing and dehazing within one network. This design significantly reduces the complexity of network and can be more easily extended to multi-domain translation tasks than the previous methods, which need one pair of generator-discriminator for each direction of the translation. Meanwhile, aiming at solving the haze-relevant task, we design a haze attention module, which takes the local entropy map as the input. Unlike the previous weakly supervised dehazing methods, our approach only requires unpaired hazy and haze-free images rather than any intermediate supervising data like the transmission map or atmospheric light defined in the atmospheric scattering model. Experimental results on synthetic datasets show our method can achieve competitive results when compared with the state-of-the-art methods and yield more appealing dehazing and hazing results on real-world images. Yunan Li 0001, Huizhou Chen, Qiguang Miao, Siyu Liang 0002, Zhuoqi Ma, Bocheng Zhao |
IEEE Trans. Multim. | 3 |
| 2023 | A Calibrated Force-Based Model for Mixed Traffic SimulationabstractVirtual traffic benefits a variety of applications, including video games, traffic engineering, autonomous driving, and virtual reality. To date, traffic visualization via different simulation models can reconstruct detailed traffic flows. However, each specific behavior of vehicles is always described by establishing an independent control model. Moreover, mutual interactions between vehicles and other road users are rarely modeled in existing simulators. An all-in-one simulator that considers the complex behaviors of all potential road users in a realistic urban environment is urgently needed. In this work, we propose a novel, extensible, and microscopic method to build heterogeneous traffic simulation using the force-based concept. This force-based approach can accurately replicate the sophisticated behaviors of various road users and their interactions in a simple and unified manner. We calibrate the model parameters using real-world traffic trajectory data. The effectiveness of this approach is demonstrated through many simulation experiments, as well as comparisons to real-world traffic data and popular microscopic simulators for traffic animation. Qianwen Chao, Chaoneng Li, Qiguang Miao, Xiaogang Jin 0001 |
IEEE Trans. Vis. Comput. Graph. | 6 |
| 2022 | Fidelity Evaluation of Virtual Traffic Based on Anomalous Trajectory DetectionabstractMeasuring the fidelity of synthesized virtual traffic has become an important and fundamental concern for evaluating the performance of different traffic simulation techniques and applications of autonomous vehicle testing. In this work, we propose a novel method to evaluate the fidelity of any trajectory data from the perspective of anomalous trajectory detection. First, given the trajectory data to be evaluated as input, the method learns spatio-temporal traffic features and reconstructs the input trajectory through a Long Short-Term Memory (LSTM)-based autoencoder architecture. Then, the anomalous trajectories are detected by comparing the reconstructed trajectories and the input ones using the reconstruction error as the benchmark. Our method can detect eight different kinds of anomalous trajectory in terms of changes in velocity and moving direction. In order to evaluate the fidelity of the input trajectory, we design a perceptual evaluation on virtual traffic fidelity and derive a mapping from the reconstruction error to the evaluation score. We demonstrated the effectiveness and robustness of our metric through many experiments on real-world and synthetic trajectory data containing different types of motion anomalies. Chaoneng Li, Qianwen Chao, Guanwen Feng, Qiongyan Wang, Yunan Li 0001, Qiguang Miao |
IROS | 7 |
| 2022 | Guided filter random walk and improved spiking cortical model based image fusion method in NSST domain
Qiguang Miao, Yang Lei 0001, Cong Ren |
Neurocomputing | 2 |
| 2022 | Multimodal medical image fusion using gradient domain guided filter random walk and side window filtering in framelet domain
Qiguang Miao, Ruyi Liu 0001, Yang Lei 0001 |
Inf. Sci. | 2 |
| 2022 | TelecomNet: Tag-Based Weakly-Supervised Modally Cooperative Hashing Network for Image RetrievalabstractWe are concerned with using user-tagged images to learn proper hashing functions for image retrieval. The benefits are two-fold: (1) we could obtain abundant training data for deep hashing models; (2) tagging data possesses richer semantic information which could help better characterize similarity relationships between images. However, tagging data suffers from noises, vagueness and incompleteness. Different from previous unsupervised or supervised hashing learning, we propose a novel weakly-supervised deep hashing framework which consists of two stages: weakly-supervised pre-training and supervised fine-tuning. The second stage is as usual. In the first stage, we propose two formulations Tag-basEd weakLy-supErvised Modally COoperative hashing Network (TelecomNet) and Generalized TelecomNet (GTelecomNet). Rather than performing supervision on tags, TelecomNet first learns an observed semantic embedding vector for each image from attached tags and then uses it to guide hashing learning. GTelecomNet introduces a novel semantic network to exploit more precise semantic information. By carefully designing the optimization problem, they can well leverage tagging information and image content for hashing learning. The framework is general and does not depend on specific deep hashing methods. Empirical results on real world datasets show that they significantly increase the performance of state-of-the-art deep hashing methods. Wei Zhao 0019, Ziyu Guan, Xunlian Wu, Wanqing Zhao, Qiguang Miao, Xiaofei He 0001, Quan Wang 0006 |
IEEE Trans. Pattern Anal. Mach. Intell. | 6 |
| 2022 | ChaLearn Looking at People: IsoGD and ConGD Large-Scale RGB-D Gesture RecognitionabstractThe ChaLearn large-scale gesture recognition challenge has run twice in two workshops in conjunction with the International Conference on Pattern Recognition (ICPR) 2016 and International Conference on Computer Vision (ICCV) 2017, attracting more than 200 teams around the world. This challenge has two tracks, focusing on isolated and continuous gesture recognition, respectively. It describes the creation of both benchmark datasets and analyzes the advances in large-scale gesture recognition based on these two datasets. In this article, we discuss the challenges of collecting large-scale ground-truth annotations of gesture recognition and provide a detailed analysis of the current methods for large-scale isolated and continuous gesture recognition. In addition to the recognition rate and mean Jaccard index (MJI) as evaluation metrics used in previous challenges, we introduce the corrected segmentation rate (CSR) metric to evaluate the performance of temporal segmentation for continuous gesture recognition. Furthermore, we propose a bidirectional long short-term memory (Bi-LSTM) method, determining video division points based on skeleton points. Experiments show that the proposed Bi-LSTM outperforms state-of-the-art methods with an absolute improvement of 8.1% (from 0.8917 to 0.9639) of CSR. Jun Wan 0001, Chi Lin 0002, Longyin Wen, Yunan Li 0001, Qiguang Miao, Sergio Escalera, Gholamreza Anbarjafari, Isabelle Guyon, Guodong Guo, Stan Z. Li |
IEEE Trans. Cybern. | 5 |
| 2022 | Multifeature Collaborative Adversarial Attack in Multimodal Remote Sensing Image ClassificationabstractDeep neural networks have strong feature learning ability, but their vulnerability cannot be ignored. Current research shows that deep learning models are threatened by adversarial examples in remote sensing (RS) classification tasks, and their robustness drops sharply in the face of adversarial attacks. Therefore, many adversarial attack methods have been studied to predict the risks faced by a network. However, the existing adversarial attack methods mainly focus on single-modal image classification networks, and the rapid growth of RS data makes multimodal RS image classification a research hotspot. Generating multimodal adversarial examples needs to consider a high attack success rate, subtle perturbation, and collaborative attack ability between different modalities. In this article, we investigate the vulnerability of multimodal RS classification networks and propose a multifeature collaborative adversarial network (MFCANet) for generating multimodal adversarial examples. Two modality-specific generators are designed to generate the multimodal collaborative perturbations with strong attack ability, and two modality-specific discriminators make the generated multimodal adversarial examples closer to the real instances. In addition, a modality-specific generative loss and a modality-specific discriminative loss are proposed, and an alternating optimization strategy is designed for training the proposed MFCANet. Extensive experiments are carried out on the International Society for Photogrammetry and Remote Sensing (ISPRS) Vaihingen 2D dataset and ISPRS Potsdam 2D dataset. The results show that the attack performance of the proposed method is stronger than that of the fast gradient sign method (FGSM), project gradient descent (PGD), and Carlini and Wagner (C&W) attack methods. Cheng Shi 0002, Yenan Dang, Minghua Zhao, Zhiyong Lv, Qiguang Miao, Chi-Man Pun |
IEEE Trans. Geosci. Remote. Sens. | 6 |
| 2022 | Gradient-Based Pulsed Excitation and Relaxation Encoding in Magnetic Particle ImagingabstractMagnetic particle imaging (MPI) is a radiation-free vessel- and target-imaging modality that can sensitively detect nanoparticles. A static magnetic gradient field, referred to as a selection field, is required in MPI to provide a field-free region (FFR) for spatial encoding. The image resolution of MPI is closely related to the size of the FFR, which is determined by the selection field gradient amplitude. Because of the limitations of existing gradient coil hardware, the image resolution of MPI cannot satisfy the clinical requirements of human in vivo imaging. Pulsed excitation has been confirmed to improve the image resolution of MPI by breaking down the 'relaxation wall.' This work proposes the use of a pulsed waveform magnetic gradient from magnetic resonance imaging to further improve the image resolution of MPI. Through alignment of the gradient direction along the field-free line (FFL), each location on the FFL is able to have a unique excitation field strength that generates a specific relaxation-induced decay signal. Through excitation of nanoparticles on the FFL with many gradient profiles, a high-resolution, one-dimensional (1D) image can be reconstructed on the FFL. For larger magnetic nanoparticles, simulation results revealed that a pulsed excitation field with a greater flat portion generates a 1D bar pattern phantom image with a higher correlation and spatial resolution. With parallel FFL and gradient coil movements, high-resolution, two-dimensional (2D) Shepp-Logan phantom and brain vessel maps were reconstructed through repetition of the spatially resolved measurement of magnetic nanoparticles on the FFL. Guang Jia, Ze Wang 0013, Xiaofeng Liang, Yu Zhang 0147, Qiguang Miao, Kai Hu 0008, Tanping Li, Ying Wang 0142, Li Xi, Xin Feng 0010, Hui Hui, Jie Tian 0001 |
IEEE Trans. Medical Imaging | 7 |
| 2022 | Commonality Autoencoder: Learning Common Features for Change Detection From Heterogeneous ImagesabstractChange detection based on heterogeneous images, such as optical images and synthetic aperture radar images, is a challenging problem because of their huge appearance differences. To combat this problem, we propose an unsupervised change detection method that contains only a convolutional autoencoder (CAE) for feature extraction and the commonality autoencoder for commonalities exploration. The CAE can eliminate a large part of redundancies in two heterogeneous images and obtain more consistent feature representations. The proposed commonality autoencoder has the ability to discover common features of ground objects between two heterogeneous images by transforming one heterogeneous image representation into another. The unchanged regions with the same ground objects share much more common features than the changed regions. Therefore, the number of common features can indicate changed regions and unchanged regions, and then a difference map can be calculated. At last, the change detection result is generated by applying a segmentation algorithm to the difference map. In our method, the network parameters of the commonality autoencoder are learned by the relevance of unchanged regions instead of the labels. Our experimental results on five real data sets demonstrate the promising performance of the proposed framework compared with several existing approaches. Yue Wu 0004, Yongzhe Yuan, A. K. Qin 0001, Qiguang Miao, Maoguo Gong |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2022 | Analysis and Variants of Broad Learning SystemabstractThe broad learning system (BLS) is designed based on the technology of compressed sensing and pseudo-inverse theory, and consists of feature nodes and enhancement nodes, has been proposed recently. Compared with the popular deep learning structures, such as deep neural networks, BLS has the ability of rapid incremental learning and can remodel the system without the usual tedious retraining process. However, given that BLS is still in its infancy, it still needs analysis, improvements, and verification. In this article, we first analyze the principle of fast incremental learning ability of BLS in depth. Second, in order to provide an in-depth analysis of the BLS structure, according to the novel structure design concept of deep neural networks, we present four brand-new BLS variant networks and their incremental realizations. Third, based on our analysis of the effect of feature nodes and enhancement nodes, a new BLS structure with a semantic feature extraction layer has been proposed, which is called SFEBLS. The experimental results show that SFEBLS and its variants can increase the accuracy rate on the NORB dataset 6.18%, Fashion-MNIST dataset by 3.15%, ORL data by 5.00%, street view house number dataset by 12.88%, and CIFAR-10 dataset by 18.42%, respectively, and the four brand-new BLS variant networks also obviously outperform the original BLS. Liang Zhang 0010, Guoqing Lu, Peiyi Shen, Mohammed Bennamoun, Syed Afaq Ali Shah, Qiguang Miao, Guangming Zhu 0001, Ping Li 0030, Xiaoyuan Lu |
IEEE Trans. Syst. Man Cybern. Syst. | 7 |
| 2021 | CA-PMG: Channel attention and progressive multi-granularity training network for fine-grained visual classificationabstractAbstract Fine‐grained visual classification is challenging due to the inherently subtle intra‐class object variations. To solve this issue, a novel framework named channel attention and progressive multi‐granularity training network, is proposed. It first exploits meaningful feature maps through the channel attention module and captures multi‐granularity features by the progressive multi‐granularity training module. For each feature map, the channel attention module is proposed to explore channel‐wise correlation. This allows the model to re‐weight the channels of the feature map according to the impact of their semantic information on performance. Furthermore, the progressive multi‐granularity training module is introduced to fuse features cross multi‐granularity. And the fused features pay more attention to the subtle differences between images. The model can be trained efficiently in an end‐to‐end manner without bounding box or part annotations. Finally, comprehensive experiments are conducted to show that the method achieves state‐of‐the‐art performances on the CUB‐200‐2011, Stanford Cars, and FGVC‐Aircraft datasets. Ablation studies demonstrate the effectiveness of each part in our module. Qiguang Miao, Hang Yao 0001, Xiangzeng Liu, Ruyi Liu 0001, Maoguo Gong |
IET Image Process. | 2 |
| 2021 | Attention adjacency matrix based graph convolutional networks for skeleton-based action recognition
Qiguang Miao, Ruyi Liu 0001, Wentian Xin, Sheng Zhong 0006, Xuesong Gao |
Neurocomputing | 2 |
| 2021 | Community-based k-shell decomposition for identifying influential spreaders
Peng Gang Sun, Qiguang Miao, Steffen Staab |
Pattern Recognit. | 2 |
| 2021 | A tri-objective preference-based uniform weight design method using Delaunay triangulation
Dazhuang Liu, Yutao Qi, Yi-Ning Quan, Xiaodong Li 0001, Qiguang Miao |
Soft Comput. | 6 |
| 2021 | Encoder-Decoder With Cascaded CRFs for Semantic SegmentationabstractWhen dealing with semantic segmentation, how to locate the object boundary information more accurately is a key problem to distinguish different objects better. The existing methods lose some image information more or less in the process of feature extraction, which also includes the boundary and context information. At present, some semantic segmentation methods use CRFs (conditional random fields) to obtain boundary information, but they usually only deal with the final output of the model. In this article, inspired by the skip connection of FCN (Fully convolution network) and the good boundary refinement ability of CRFs, a cascaded CRFs is designed and introduced into the decoder of semantic segmentation model to learn boundary information from multi-layers and enhance the ability of the model in object boundary location. Furthermore, in order to supplement the semantic information of images, the output of the cascaded CRFs is fused with the output of the last decoder, so that the model can enhance the ability of locating the object boundary and get more accurate semantic segmentation results. Finally, a number of experiments on different datasets illustrate the feasibility and efficiency of our method, showing that our method enhances the model’s ability to locate target boundary information. Jian Ji 0002, Sitong Li, Qiguang Miao |
IEEE Trans. Circuits Syst. Video Technol. | 5 |
| 2021 | Pareto Self-Paced Learning Based on Differential EvolutionabstractInspired by the learning rules of humans/animals, self-paced learning (SPL) ranks the samples from easy to complex and assigns real-valued weights to the samples. The current SPL regimes adopt an increasing pace parameter to select the training samples. Obviously, it is difficult to tune the pace parameter during iterations. Furthermore, the current SPL regimes cannot go back to the previous stage even when the current model is worse than the previous one. In this article, a novel Pareto SPL (PSPL) approach is proposed to address the above issues. In PSPL, the SPL problem is considered as a multiobjective optimization problem. Then, a Pareto-based differential evolution algorithm is utilized to optimize the self-paced function and the weighted loss function simultaneously. In PSPL, a new representation method is designed to assign nonpositive weights to the unselected samples. Then, a modified mutation operator inspired by the SPL idea is proposed to generate new population based on ranking the individuals and assigning weights. No pace parameters are introduced in the proposed technique and PSPL can obtain entire solution spectrum to automatically adjust the solution path. The effectiveness of the PSPL has been demonstrated through rigorous studies with matrix factorization and multiclass classification problems. Hao Li 0009, Maoguo Gong, Qiguang Miao |
IEEE Trans. Cybern. | 4 |
| 2021 | Recommendation by Users' Multimodal Preferences for Smart City ApplicationsabstractAs an essential role in smart city applications, personalized recommender systems help users to find their potentially interested items from their historically generated data. Recently, researchers have started to utilize the massive user-generated multimodal contents to improve recommendation performance. However, previous methods have at least one of the following drawbacks: 1) employing shallow models, which cannot well capture high-level conceptual information; 2) failing to capture personalized user visual preference. In this article, we present a deep users’ multimodal preferences-based recommendation (UMPR) method to capture the textual and visual matching of users and items for recommendation. We extract textual matching from historical reviews. We construct users’ visual preference embeddings to model users’ visual preference and match them with items’ visual embeddings to obtain the visual matching. We apply UMPR on two applications related to smart city: restaurant recommendation and product recommendation. Experiments show that UMPR outperforms competitive baseline methods. Ziyu Guan, Wei Zhao 0019, Quanzhou Wu, Meng Yan 0013, Long Chen 0007, Qiguang Miao |
IEEE Trans. Ind. Informatics | 7 |
| 2021 | Review of dynamic gesture recognitionabstractIn recent years, gesture recognition has been widely used in the fields of intelligent driving, virtual reality, and human-computer interaction. With the development of artificial intelligence, deep learning has achieved remarkable success in computer vision. To help researchers better understanding the development status of gesture recognition in video, this article provides a detailed survey of the latest developments in gesture recognition technology for videos based on deep learning. The reviewed methods are broadly categorized into three groups based on the type of neural networks used for recognition: twostream convolutional neural networks, 3D convolutional neural networks, and Long-short Term Memory (LSTM) networks. In this review, we discuss the advantages and limitations of existing technologies, focusing on the feature extraction method of the spatiotemporal structure information in a video sequence, and consider future research directions. Yunan Li 0001, Xiaolong Fu, Kaibin Miao, Qiguang Miao |
Virtual Real. Intell. Hardw. | 5 |
| 2020 | ProFPred: a two-step protein function prediction model based on sequence and evolutionary informationabstractIn post-genomic era, the understanding of protein function has been seriously behind the development of sequencing technology. Experimental verification for protein function is difficult, time consuming and expensive. Meanwhile, the new proteins vary in function. It is difficult for traditional methods to fully and accurately understand its functions. In this work, we present a novel method ProFPred to predict protein function based on sequence and evolutionary information to deal with small samples data corresponding to new diseases or discoveries. Experimental results demonstrate that our method can achieve better or comparable performances compared with current state-of-the-art methods. Ruiquan Ge, Guanwen Feng, Qiguang Miao |
BIBM | 4 |
| 2020 | Image patch prior learning based on random neighbourhood resampling for image denoisingabstractImage patch priors become a popular tool for image denoising. The Gaussian mixture model (GMM) is remarkably effective in modelling natural image patches. However, GMM prior learning using the expectation maximisation (EM) algorithm is sensitive to the initialisation, often leading to low convergence rate of parameter estimation. In this study, a novel sampling method called random neighbourhood resampling (RNR) is proposed to improve the accuracy and efficiency of parameter estimation. An enhanced GMM (EGMM) learning algorithm is further developed by incorporating RNR into the EM algorithm to initialise and update the GMM prior. The learned EGMM prior is applied in the expected patch log‐likelihood (EPLL) framework for image denoising. The effectiveness and performance of the proposed RNR and EGMM algorithm are demonstrated via extensive experimental results comparing with the state‐of‐the‐art image denoising methods, the experimental results show the higher PSNR result of the denoised images using the proposed method. Meanwhile, the authors verified that the proposed method can efficiently reduce the time of image denoising compared with the basic EPLL method. Jian Ji 0002, Jiajie Wei, Mengqi Bai, Qiguang Miao |
IET Image Process. | 6 |
| 2020 | CR-Net: A Deep Classification-Regression Network for Multimodal Apparent Personality Analysis
Yunan Li 0001, Jun Wan 0001, Qiguang Miao, Sergio Escalera, Huijuan Fang, Huizhou Chen, Xiangda Qi, Guodong Guo |
Int. J. Comput. Vis. | 3 |
| 2020 | An adaptive penalty-based boundary intersection method for many-objective optimization problem
Yutao Qi, Dazhuang Liu, Xiaodong Li 0001, Jiaojiao Lei, Xiaoying Xu, Qiguang Miao |
Inf. Sci. | 6 |
| 2020 | A Semi-Supervised High-Level Feature Selection Framework for Road Centerline ExtractionabstractAccurate road centerline extraction is very important for many vital applications. In the road extraction, the acquisition of labeled data is time-consuming; thus, there is only a small amount of labeled samples in reality. To solve the problem of limited labeled samples, a semi-supervised road centerline extraction is proposed, which incorporates high-level feature selection, Markov random field (MRF), and ridge transversal method. The proposed road extraction approach consists of three steps: multiple features extraction, semi-supervised road area extraction, and road centerlines extraction. To get more abstract and discriminative high-level features, we apply multiple-feature adaptive sparse representation in mid-level features in different views generated by different prototype sets. To obtain an accurate road area result, we combine the feature learning framework with MRF. Then, we integrate Gabor filters and nonmaxima suppression with the ridge transversal method to extract centerlines. It is verified the proposed method achieves comparable performance with the state-of-the-art methods in terms of visual and quantitative aspects. Ruyi Liu 0001, Qiguang Miao, Yi Zhang 0033, Maoguo Gong, Pengfei Xu 0003 |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2020 | Dictionary-based Fidelity Measure for Virtual TrafficabstractAiming at objectively measuring the realism of virtual traffic flows and evaluating the effectiveness of different traffic simulation techniques, this paper introduces a general, dictionary-based learning method to evaluate the fidelity of any traffic trajectory data. First, a traffic pattern dictionary that characterizes common patterns of real-world traffic behavior is built offline from pre-collected ground truth traffic data. The corresponding learning error is set as the benchmark of the dictionary-based traffic representation. With the aid of the constructed dictionary, the realism of input simulated traffic flow data can be evaluated by comparing its dictionary-based reconstruction error with the dictionary error benchmark. This evaluation metric can be robustly applied to any simulated traffic flow data; in other words, it is independent of how the traffic data are generated. We demonstrated the effectiveness and robustness of this metric through many experiments on real-world traffic data and various simulated traffic data, comparisons with the state-of-the-art entropy-based similarity metric for aggregate crowd motions, and perceptual evaluation studies. Qianwen Chao, Zhigang Deng 0001, Yangxi Xiao, Dunbang He, Qiguang Miao, Xiaogang Jin 0001 |
IEEE Trans. Vis. Comput. Graph. | 5 |
| 2019 | LAP-Net: Level-Aware Progressive Network for Image DehazingabstractIn this paper, we propose a level-aware progressive network (LAP-Net) for single image dehazing. Unlike previous multi-stage algorithms that generally learn in a coarse-to-fine fashion, each stage of LAP-Net learns different levels of haze with different supervision. Then the network can progressively learn the gradually aggravating haze. With this design, each stage can focus on a region with specific haze level and restore clear details. To effectively fuse the results of varying haze levels at different stages, we develop an adaptive integration strategy to yield the final dehazed image. This strategy is achieved by a hierarchical integration scheme, which is in cooperation with the memory network and the domain knowledge of dehazing to highlight the best-restored regions of each stage. Extensive experiments on both real-world images and two dehazing benchmarks validate the effectiveness of our proposed method. Yunan Li 0001, Qiguang Miao, Wanli Ouyang, Zhenxin Ma, Huijuan Fang, Chao Dong 0005, Yi-Ning Quan |
ICCV | 2 |
| 2019 | Image De-noising by an Effective SURE-Based Weighted Bilateral Filtering
Jian Ji 0002, Sitong Li, Guo-Fei Hou, Fen Ren, Qiguang Miao |
PRCV (2) | 5 |
| 2019 | Architectural Style Classification Based on DNN Model
Qiguang Miao, Ruyi Liu 0001, Jianfeng Song |
PRCV (1) | 2 |
| 2019 | Multiscale road centerlines extraction from high-resolution aerial imagery
Ruyi Liu 0001, Qiguang Miao, Jianfeng Song, Yi-Ning Quan, Yunan Li 0001, Pengfei Xu 0003 |
Neurocomputing | 2 |
| 2019 | Chinese Conference on Computer Vision 2017
Qiguang Miao, Cheng Deng 0002, Mingliang Xu 0001 |
Neurocomputing | 1 |
| 2019 | A spatiotemporal attention-based ResC3D model for large-scale gesture recognition
Yunan Li 0001, Qiguang Miao, Xiangda Qi, Zhenxin Ma, Wanli Ouyang |
Mach. Vis. Appl. | 2 |
| 2019 | Large-scale gesture recognition with a fusion of RGB-D data based on optical flow and the C3D model
Yunan Li 0001, Qiguang Miao, Kuan Tian, Xin Xu 0001, Zhenxin Ma, Jianfeng Song |
Pattern Recognit. Lett. | 2 |
| 2019 | Decomposition-Based Evolutionary Multiobjective Optimization to Self-Paced LearningabstractSelf-paced learning (SPL) is a recently proposed paradigm to imitate the learning process of humans/animals. SPL involves easier samples into training at first and then gradually takes more complex ones into consideration. Current SPL regimes incorporate a self-paced (SP) regularizer into the learning objective with a gradually increasing pace parameter. Therefore, it is difficult to obtain the solution path of the SPL regime and determine where to optimally stop this increasing process. In this paper, a multiobjective SPL method is proposed to optimize the loss function and the SP regularizer simultaneously. A decomposition-based multiobjective particle swarm optimization algorithm is used to simultaneously optimize the two objectives for obtaining the solutions. In the proposed method, a polynomial soft weighting regularizer is proposed to penalize the loss. Theoretical studies are conducted to show that the previous regularizers are roughly particular cases of the proposed polynomial soft weighting regularizer family. Then an implicit decomposition method is proposed to search the solutions with respect to the sample number involved into training. A set of solutions can be obtained by the proposed method and naturally constitute the solution path of the SPL regime. Then a satisfactory solution can be naturally obtained from these solutions by utilizing some effective tools in evolutionary multiobjective optimization. Experiments on matrix factorization and classification problems demonstrate the effectiveness of the proposed technique. Maoguo Gong, Hao Li 0009, Deyu Meng, Qiguang Miao, Jia Liu 0020 |
IEEE Trans. Evol. Comput. | 4 |
| 2018 | Line separation from topographic maps using regional color and spatial informationabstractThe lines in topographic maps are difficult to be separated from each other because of their confusing colors. To solve this problem, we propose a novel line separation method using their regional color and spatial information. Firstly, we divide the lines into lots of circular regions with a certain diameter, and consider these regions as the basic processing units. Then based on a new concept of regional color confusion, we classify all the divided circular regions into two kinds of regions by whether the color is pure or mixed. Further, for pure color regions, a fuzzy clustering algorithm with Gaussian kernel can be used to cluster them into different lines based on their color information. Meanwhile, we determine the memberships of the mixed color regions according to their spatial relations with the clustered pure color regions. The concept of regional color confusion is proposed to reduce the influences of the confusing colors to line separation, and the spatial relations are utilized to solve the problems of the membership determination of the mixed color regions. The experimental results demonstrate that our method can achieve higher accuracy compare with other two state-of-the-art methods, which provides a novel idea for line element segmentation from scanned topographic maps. Pengfei Xu 0003, Qiguang Miao, Tiange Liu, Xiaojiang Chen, Dingyi Fang |
IJCAI | 2 |
| 2018 | Self-Paced Densely Connected Convolutional Neural Network for Visual Tracking
Jianfeng Song, Yutao Qi, Chongxiao Wang, Qiguang Miao |
PRCV (4) | 5 |
| 2018 | Boosting Sparsity-Induced Autoencoder: A Novel Sparse Feature Ensemble Learning for Image Classification
Jian Ji 0002, Qiguang Miao |
PRCV (3) | 4 |
| 2018 | A multi-scale fusion scheme based on haze-relevant features for single image dehazing
Yunan Li 0001, Qiguang Miao, Ruyi Liu 0001, Jianfeng Song, Yi-Ning Quan, Yuhui Huang |
Neurocomputing | 2 |
| 2018 | Interactive active contour with kernel descriptor
Hao Li 0009, Maoguo Gong, Qiguang Miao, Bin Wang 0027 |
Inf. Sci. | 3 |
| 2018 | Identifying influential genes in protein-protein interaction networks
Peng Gang Sun, Yi-Ning Quan, Qiguang Miao, Juan Chi |
Inf. Sci. | 3 |
| 2018 | PSOSAC: Particle Swarm Optimization Sample Consensus Algorithm for Remote Sensing Image RegistrationabstractImage registration is an important preprocessing step for many remote sensing image processing applications, and its result will affect the performance of the follow-up procedures. Establishing reliable matches is a key issue in point matching-based image registration. Due to the significant intensity mapping difference between remote sensing images, it may be difficult to find enough correct matches from the tentative matches. In this letter, particle swarm optimization (PSO) sample consensus algorithm is proposed for remote sensing image registration. Different from random sample consensus (RANSAC) algorithm, the proposed method directly samples the modal transformation parameter rather than randomly selecting tentative matches. Thus, the proposed method is less sensitive to the correct rate than RANSAC, and it has the ability to handle lower correct rate and more matches. Meanwhile, PSO is utilized to optimize parameter as its efficiency. The proposed method is tested on several multisensor remote sensing image pairs. The experimental results indicate that the proposed method yields a better registration performance in terms of both the number of correct matches and aligning accuracy. Yue Wu 0004, Qiguang Miao, Wenping Ma 0001, Maoguo Gong, Shanfeng Wang |
IEEE Geosci. Remote. Sens. Lett. | 2 |
| 2018 | Face detection of golden monkeys via regional color quantization and incremental self-paced curriculum learning
Pengfei Xu 0003, Songtao Guo, Qiguang Miao, Baoguo Li, Xiaojiang Chen, Dingyi Fang |
Multim. Tools Appl. | 3 |
| 2018 | Large-Scale Gesture Recognition With a Fusion of RGB-D Data Based on Saliency Theory and C3D ModelabstractGesture recognition has raised wide attention in computer vision owing to its many applications. However, the task of video-based large-scale gesture recognition yet faces many challenges, since many gesture-irrelevant factors like the background may disturb the recognition accuracy. To better recognize gestures with large-scale videos, we propose a method based on RGB-D data in this paper, where the “RGB-D” means RGB and depth data captured simultaneously by specific devices like Kinect. To learn gesture details better, we first use an adaptive frame unification strategy to unify the frame number of inputs, and then the RGB and depth data are sent to the C3D model to extract spatiotemporal features, respectively. In order to alleviate the interference of gesture-irrelevant factors, the saliency theory is also employed to generate auxiliary data. Next the features of these data are combined to boost the performance, which can also avoid unreasonable synthetic data, since the dimension of C3D features is uniform. Finally the performances of several classifiers are tested and the best one of SVM classifier is selected to output the ultimate accuracy. Our approach achieves 52.04% and 59.43% accuracy on the validation and testing subset of the Chalearn LAP IsoGD, respectively, both of which outperform our results in the chalearn LAP Large-scale Gesture Recognition Challenge as reported in ICPR 2016. Yunan Li 0001, Qiguang Miao, Kuan Tian, Xin Xu 0001, Jianfeng Song |
IEEE Trans. Circuits Syst. Video Technol. | 2 |
| 2018 | Structure Learning for Deep Neural Networks Based on Multiobjective OptimizationabstractThis paper focuses on the connecting structure of deep neural networks and proposes a layerwise structure learning method based on multiobjective optimization. A model with better generalization can be obtained by reducing the connecting parameters in deep networks. The aim is to find the optimal structure with high representation ability and better generalization for each layer. Then, the visible data are modeled with respect to structure based on the products of experts. In order to mitigate the difficulty of estimating the denominator in PoE, the denominator is simplified and taken as another objective, i.e., the connecting sparsity. Moreover, for the consideration of the contradictory nature between the representation ability and the network connecting sparsity, the multiobjective model is established. An improved multiobjective evolutionary algorithm is used to solve this model. Two tricks are designed to decrease the computational cost according to the properties of input data. The experiments on single-layer level, hierarchical level, and application level demonstrate the effectiveness of the proposed algorithm, and the learned structures can improve the performance of deep neural networks. Jia Liu 0020, Maoguo Gong, Qiguang Miao, Xiaogang Wang 0001, Hao Li 0009 |
IEEE Trans. Neural Networks Learn. Syst. | 3 |
| 2017 | Neuron Learning Machine for Representation LearningabstractThis paper presents a novel neuron learning machine (NLM) which can extract hierarchical features from data. We focus on the single-layer neural network architecture and propose to model the network based on the Hebbian learning rule. Hebbian learning rule describes how synaptic weight changes with the activations of presynaptic and postsynaptic neurons. We model the learning rule as the objective function by considering the simplicity of the network and stability of solutions. We make a hypothesis and introduce a correlation based constraint according to the hypothesis. We find that this biologically inspired model has the ability of learning useful features from the perspectives of retaining abstract information. NLM can also be stacked to learn hierarchical features and reformulated into convolutional version to extract features from 2-dimensional data. Jia Liu 0020, Maoguo Gong, Qiguang Miao |
AAAI | 3 |
| 2017 | Modeling Hebb Learning Rule for Unsupervised LearningabstractThis paper presents to model the Hebb learning rule and proposes a neuron learning machine (NLM). Hebb learning rule describes the plasticity of the connection between presynaptic and postsynaptic neurons and it is unsupervised itself. It formulates the updating gradient of the connecting weight in artificial neural networks. In this paper, we construct an objective function via modeling the Hebb rule. We make a hypothesis to simplify the model and introduce a correlation based constraint according to the hypothesis and stability of solutions. By analysis from the perspectives of maintaining abstract information and increasing the energy based probability of observed data, we find that this biologically inspired model has the capability of learning useful features. NLM can also be stacked to learn hierarchical features and reformulated into convolutional version to extract features from 2-dimensional data. Experiments on single-layer and deep networks demonstrate the effectiveness of NLM in unsupervised feature learning. Jia Liu 0020, Maoguo Gong, Qiguang Miao |
IJCAI | 3 |
| 2017 | Filtering LiDAR data based on adjacent triangle of triangulated irregular network
Yi-Ning Quan, Jianfeng Song, Qiguang Miao |
Multim. Tools Appl. | 4 |
| 2017 | Color texture segmentation based on active contour model with multichannel nonlocal and Tikhonov regularization
Guodong Wang 0001, Jingge Lu, Zhenkuan Pan 0001, Qiguang Miao |
Multim. Tools Appl. | 4 |
| 2017 | Superpixel-Based Difference Representation Learning for Change Detection in Multispectral Remote Sensing ImagesabstractWith the rapid technological development of various satellite sensors, high-resolution remotely sensed imagery has been an important source of data for change detection in land cover transition. However, it is still a challenging problem to effectively exploit the available spectral information to highlight changes. In this paper, we present a novel change detection framework for high-resolution remote sensing images, which incorporates superpixel-based change feature extraction and hierarchical difference representation learning by neural networks. First, highly homogenous and compact image superpixels are generated using superpixel segmentation, which makes these image blocks adhere well to image boundaries. Second, the change features are extracted to represent the difference information using spectrum, texture, and spatial features between the corresponding superpixels. Third, motivated by the fact that deep neural network has the ability to learn from data sets that have few labeled data, we use it to learn the semantic difference between the changed and unchanged pixels. The labeled data can be selected from the bitemporal multispectral images via a preclassification map generated in advance. And then, a neural network is built to learn the difference and classify the uncertain samples into changed or unchanged ones. Finally, a robust and high-contrast change detection result can be obtained from the network. The experimental results on the real data sets demonstrate its effectiveness, feasibility, and superiority of the proposed technique. Maoguo Gong, Tao Zhan 0005, Puzhao Zhang, Qiguang Miao |
IEEE Trans. Geosci. Remote. Sens. | 4 |
| 2017 | The Recognition of the Point Symbols in the Scanned Topographic MapsabstractIt is difficult to separate the point symbols from the scanned topographic maps accurately, which brings challenges for the recognition of the point symbols. In this paper, based on the framework of generalized Hough transform (GHT), we propose a new algorithm, which is named shear line segment GHT (SLS-GHT), to recognize the point symbols directly in the scanned topographic maps. SLS-GHT combines the line segment GHT (LS-GHT) and the shear transformation. On the one hand, LS-GHT is proposed to represent the features of the point symbols more completely. Its R-table has double level indices, the first one is the color information of the point symbols, and the other is the slope of the line segment connected a pair of the skeleton points. On the other hand, the shear transformation is introduced to increase the directional features of the point symbols; it can make up for the directional limitation of LS-GHT indirectly. In this way, the point symbols are detected in a series of the sheared maps by LS-GHT, and the final optimal coordinates of the setpoints are gotten from a series of the recognition results. SLS-GHT detects the point symbols directly in the scanned topographic maps, totally different from the traditional pattern of extraction before recognition. Moreover, several experiments demonstrate that the proposed method allows improved recognition in complex scenes than the existing methods. Qiguang Miao, Pengfei Xu 0003, Xuelong Li 0001, Jianfeng Song, Weisheng Li 0001 |
IEEE Trans. Image Process. | 1 |
| 2016 | Multi-Objective Self-Paced LearningabstractCurrent self-paced learning (SPL) regimes adopt the greedy strategy to obtain the solution with a gradually increasing pace parameter while where to optimally terminate this increasing process is difficult to determine.Besides, most SPL implementations are very sensitive to initialization and short of a theoretical result to clarify where SPL converges to with pace parameter increasing.In this paper, we propose a novel multi-objective self-paced learning (MOSPL) method to address these issues.Specifically, we decompose the objective functions as two terms, including the loss and the self-paced regularizer, respectively, and treat the problem as the compromise between these two objectives.This naturally reformulates the SPL problem as a standard multi-objective issue.A multi-objective evolutionary algorithm is used to optimize the two objectives simultaneously to facilitate the rational selection of a proper pace parameter.The proposed technique is capable of ameliorating a set of solutions with respect to a range of pace parameters through finely compromising these solutions inbetween, and making them perform robustly even under bad initialization.A good solution can then be naturally achieved from these solutions by making use of some off-the-shelf tools in multi-objective optimization.Experimental results on matrix factorization and action recognition demonstrate the superiority of the proposed method against the existing issues in current SPL research. Hao Li 0009, Maoguo Gong, Deyu Meng, Qiguang Miao |
AAAI | 4 |
| 2016 | Large-scale gesture recognition with a fusion of RGB-D data based on the C3D modelabstractThe gesture recognition has raised attention in computer vision owing to its many applications. However, video-based large-scale gesture recognition still faces many challenges, since many factors like background may disturb the accuracy. To achieve gesture recognition with large-scale videos, we propose a method based on RGB-D data. To learn gesture details better, the inputs are expanded into 32-frame videos first, and then the RGB and depth videos are sent to the C3D model to extract spatiotemporal features respectively. Next these features are combined to boost the performance, which can also avoid unreasonable synthetic data due to the uniform dimension of C3D features. Our approach achieves 49.2% accuracy on the validation subset of the Chalearn LAP IsoGD Database just with a linear SVM classifier. It also outperforms the baseline and other methods in the challenge and wins the first place at 56.9% on testing set. Yunan Li 0001, Qiguang Miao, Kuan Tian, Xin Xu 0001, Jianfeng Song |
ICPR | 2 |
| 2016 | Special issue on Chinese Conference on Computer Vision 2015
Xinbo Gao 0001, Deyu Meng, Liang Lin 0004, Qiguang Miao |
Neurocomputing | 4 |
| 2016 | Single image haze removal based on haze physical characteristics and adaptive sky region detection
Yunan Li 0001, Qiguang Miao, Jianfeng Song, Yi-Ning Quan, Weisheng Li 0001 |
Neurocomputing | 2 |
| 2016 | Modular ensembles for one-class classification based on density analysis
Qiguang Miao, Jianfeng Song, Yi-Ning Quan |
Neurocomputing | 2 |
| 2016 | Road centerlines extraction from high resolution images based on an improved directional segmentation and road probability
Ruyi Liu 0001, Jianfeng Song, Qiguang Miao, Pengfei Xu 0003 |
Neurocomputing | 3 |
| 2016 | Dynamic character grouping based on four consistency constraints in topographic maps
Pengfei Xu 0003, Qiguang Miao, Ruyi Liu 0001, Xiaojiang Chen, Xunli Fan |
Neurocomputing | 2 |
| 2016 | Graphic-based character grouping in topographic maps
Pengfei Xu 0003, Qiguang Miao, Tiange Liu, Xiaojiang Chen, Weike Nie |
Neurocomputing | 2 |
| 2016 | Self-adaptive multi-objective evolutionary algorithm based on decomposition for large-scale problems: A case study on reservoir flood control operation
Yutao Qi, Liang Bao, Xiaoliang Ma 0001, Qiguang Miao, Xiaodong Li 0001 |
Inf. Sci. | 4 |
| 2016 | Improved road centerlines extraction in high-resolution remote sensing images using shear transform, directional morphological filtering and enhanced broken lines connection
Ruyi Liu 0001, Qiguang Miao, Bormin Huang, Jianfeng Song, Johan Debayle |
J. Vis. Commun. Image Represent. | 2 |
| 2016 | SCTMS: Superpixel based color topographic map segmentation method
Tiange Liu, Qiguang Miao, Kuan Tian, Jianfeng Song, Yutao Qi |
J. Vis. Commun. Image Represent. | 2 |
| 2016 | Artistic information extraction from Chinese calligraphy works via Shear-Guided filter
Pengfei Xu 0003, Xia Zheng, Xiaojun Chang, Qiguang Miao, Zhanyong Tang, Xiaojiang Chen, Dingyi Fang |
J. Vis. Commun. Image Represent. | 4 |
| 2016 | Color topographical map segmentation Algorithm based on linear element features
Tiange Liu, Qiguang Miao, Pengfei Xu 0003, Jianfeng Song, Yi-Ning Quan |
Multim. Tools Appl. | 2 |
| 2016 | A new artistic information extraction method with multi channels and guided filters for calligraphy works
Xia Zheng, Qiguang Miao, Zhenghao Shi, Yachun Fan, Wuyang Shui |
Multim. Tools Appl. | 2 |
| 2016 | Fast structural ensemble for One-Class Classification
Qiguang Miao, Jianfeng Song, Yi-Ning Quan |
Pattern Recognit. Lett. | 2 |
| 2016 | Guided Superpixel Method for Topographic Map ProcessingabstractSuperpixels have been widely used in lots of computer vision and image processing tasks but rarely used in topographic map processing due to the complex distribution of geographic elements in this kind of images. We propose a novel superpixel-generating method based on guided watershed transform (GWT). Before GWT, the cues of geographic element distribution and boundaries between different elements need to be obtained. A linear feature extraction method based on a compound opposite Gaussian filter and a shear transform is presented to acquire the distribution information. Meanwhile, a boundary detection method, which based on the color-opponent mechanisms of the visual system, is employed to get the boundary information. Then, both linear features and boundaries are input to the final partition procedure to obtain superpixels. The experiments show that our method has the best performance in shape control, size control, and boundary adherence, among all the comparison methods, which are classic and state of the art. Furthermore, we verify the low complexity and low cost of memory in our method through experiments, which makes it possible to deal with large-scale topographic maps. Qiguang Miao, Tiange Liu, Jianfeng Song, Maoguo Gong |
IEEE Trans. Geosci. Remote. Sens. | 1 |
| 2016 | Change Detection in Synthetic Aperture Radar Images Based on Deep Neural NetworksabstractThis paper presents a novel change detection approach for synthetic aperture radar images based on deep learning. The approach accomplishes the detection of the changed and unchanged areas by designing a deep neural network. The main guideline is to produce a change detection map directly from two images with the trained deep neural network. The method can omit the process of generating a difference image (DI) that shows difference degrees between multitemporal synthetic aperture radar images. Thus, it can avoid the effect of the DI on the change detection results. The learning algorithm for deep architectures includes unsupervised feature learning and supervised fine-tuning to complete classification. The unsupervised feature learning aims at learning the representation of the relationships between the two images. In addition, the supervised fine-tuning aims at learning the concepts of the changed and unchanged pixels. Experiments on real data sets and theoretical analysis indicate the advantages, feasibility, and potential of the proposed method. Moreover, based on the results achieved by various traditional algorithms, respectively, deep learning can further improve the detection performance. Maoguo Gong, Jiaojiao Zhao, Jia Liu 0020, Qiguang Miao, Licheng Jiao |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2016 | RBoost: Label Noise-Robust Boosting Algorithm Based on a Nonconvex Loss Function and the Numerically Stable Base LearnersabstractAdaBoost has attracted much attention in the machine learning community because of its excellent performance in combining weak classifiers into strong classifiers. However, AdaBoost tends to overfit to the noisy data in many applications. Accordingly, improving the antinoise ability of AdaBoost plays an important role in many applications. The sensitiveness to the noisy data of AdaBoost stems from the exponential loss function, which puts unrestricted penalties to the misclassified samples with very large margins. In this paper, we propose two boosting algorithms, referred to as RBoost1 and RBoost2, which are more robust to the noisy data compared with AdaBoost. RBoost1 and RBoost2 optimize a nonconvex loss function of the classification margin. Because the penalties to the misclassified samples are restricted to an amount less than one, RBoost1 and RBoost2 do not overfocus on the samples that are always misclassified by the previous base learners. Besides the loss function, at each boosting iteration, RBoost1 and RBoost2 use numerically stable ways to compute the base learners. These two improvements contribute to the robustness of the proposed algorithms to the noisy training and testing samples. Experimental results on the synthetic Gaussian data set, the UCI data sets, and a real malware behavior data set illustrate that the proposed RBoost1 and RBoost2 algorithms perform better when the training data sets contain noisy data. Qiguang Miao, Ying Cao 0003, Ge Xia, Maoguo Gong, Jianfeng Song |
IEEE Trans. Neural Networks Learn. Syst. | 1 |
| 2015 | A novel fast image segmentation algorithm for large topographic maps
Qiguang Miao, Pengfei Xu 0003, Tiange Liu, Jianfeng Song, Xiaojiang Chen |
Neurocomputing | 1 |
| 2015 | An Ensemble Cost-Sensitive One-Class Learning Framework for Malware DetectionabstractMachine learning is among the most popular methods in designing unknown and variant malware detection algorithms. However, most of the existing methods take a single type of features to build binary classifiers. In practice, these methods have limited ability in depicting malware characteristics and the binary classification suffers from inadequate sampling of benign samples and extremely imbalanced training samples when detecting malware. In this paper, we present a malware detection Framework based on ENsemble One-Class Learning, namely FENOC. It uses hybrid features at different semantic layers to ensure a comprehensive insight of the program to be analyzed. We construct the malware detector by a novel learning algorithm called Cost-sensitive Twin One-class Classifier (CosTOC), which uses a pair of one-class classifiers to describe malware and benign programs respectively. CosTOC is more flexible and robust in comparison to conventional binary classifiers when training samples are extremely imbalanced or the benign programs are inadequately sampled. Finally, random subspace method and clustering-based ensemble method are developed to enhance the generalization ability of CosTOC. Experimental results show that FENOC gives a comparative detection rate and a lower false positive rate than many other binary classification algorithms, especially when the detector are trained with imbalanced data, or evaluated in terms of false positive rate. Jianfeng Song, Qiguang Miao, Ying Cao 0003, Yi-Ning Quan |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2015 | BoostFS: A Boosting-Based Irrelevant Feature Selection AlgorithmabstractIn a learning process, features play a fundamental role. In this paper, we propose a Boosting-based feature selection algorithm called BoostFS. It extends AdaBoost which is designed for classification problems to feature selection. BoostFS maintains a distribution over training samples which is initialized from the uniform distribution. In each iteration, a decision stump is trained under the sample distribution and then the sample distribution is adjusted so that it is orthogonal to the classification results of all the generated stumps. Because a decision stump can also be regarded as one selected feature, BoostFS is capable to select a subset of features that are irrelevant to each other as much as possible. Experimental results on synthetic datasets, five UCI datasets and a real malware detection dataset all show that the features selected by BoostFS help to improve learning algorithms in classification problems, especially when the original feature set contains redundant features. Qiguang Miao, Ying Cao 0003, Jianfeng Song, Yi-Ning Quan |
Int. J. Pattern Recognit. Artif. Intell. | 1 |
| 2015 | Predicting individual retweet behavior by user similarity: A multi-task learning approach
Xing Tang 0007, Qiguang Miao, Yi-Ning Quan, Jie Tang 0001 |
Knowl. Based Syst. | 2 |
| 2015 | Pan-sharpening via regional division and NSST
Cheng Shi 0002, Fang Liu 0001, Qiguang Miao |
Multim. Tools Appl. | 3 |
| 2015 | Network Structural Balance Based on Evolutionary Multiobjective Optimization: A Two-Step ApproachabstractResearch on network structural balance has been of great concern to scholars from diverse fields. In this paper, a two-step approach is proposed for the first time to address the network structural balance problem. The proposed approach involves evolutionary multiobjective optimization, followed by model selection. In the first step, an improved version of the multiobjective discrete particle swarm optimization framework developed in our previous work is suggested. The suggested framework is then employed to implement network multiresolution clustering. In the second step, a problem-specific model selection strategy is devised to select the best Pareto solution (PS) from the Pareto front produced by the first step. The best PS is then decoded into the corresponding network community structure. Based on the discovered community structure, imbalanced edges are determined. Afterward, imbalanced edges are flipped so as to make the network structurally balanced. Extensive experiments on synthetic and real-world signed networks demonstrate the effectiveness of the proposed approach. Maoguo Gong, Shasha Ruan, Qiguang Miao, Haifeng Du |
IEEE Trans. Evol. Comput. | 4 |
| 2014 | A denoising algorithm via wiener filtering in the shearlet domain
Pengfei Xu 0003, Qiguang Miao, Xing Tang 0007 |
Multim. Tools Appl. | 2 |
| 2013 | A novel algorithm of remote sensing image fusion based on Shearlets and PCNN
Cheng Shi 0002, Qiguang Miao, Pengfei Xu 0003 |
Neurocomputing | 2 |
| 2013 | Linear Feature Separation From Topographic Maps Using Energy Density and the Shear TransformabstractLinear features are difficult to be separated from complicated background in color scanned topographic maps, especially when the color of linear features approximate to that of background in some particular images. This paper presents a method, which is based on energy density and the shear transform, for the separation of lines from background. First, the shear transform, which could add the directional characteristics of the lines, is introduced to overcome the disadvantage that linear information loss would happen if the separation method is used in an image, which is in only one direction. Then templates in the horizontal and vertical directions are built to separate lines from background on account of the fact that the energy concentration of the lines usually reaches a higher level than that of the background in the negtive image. Furthermore, the remaining grid background can be wiped off by grid templates matching. The isolated patches, which include only one pixel or less than ten pixels, are removed according to the connected region area measurement. Finally, using the union operation, the linear features obtained in different sheared images could supplement each other, thus the lines of the final result are more complete. The basic property of this method is introducing the energy density instead of color information commonly used in traditional methods. The experiment results indicate that the proposed method could distinguish the linear features from the background more effectively, and obtain good results for its ability in changing the directions of the lines with the shear transform. Qiguang Miao, Pengfei Xu 0003, Tiange Liu, Weisheng Li 0001 |
IEEE Trans. Image Process. | 1 |
| 2012 | An edge detection algorithm based on the multi-direction shear transform
Pengfei Xu 0003, Qiguang Miao, Cheng Shi 0002, Weisheng Li 0001 |
J. Vis. Commun. Image Represent. | 2 |
| 2007 | A novel image fusion algorithm using FRIT AND PCAabstractThe proposed new fusion algorithm is based on the finite ridgelet transform(FRIT) and PCA. FRIT could capture two and higher dimensional singularity is analyzed. FRIT is used to decompose the image into low and high frequency components. The PCA method is used to fuse the low frequency coefficients. And for the high frequency coefficients, the maximum-method and the region consistency check are adopted. Experiments show that the proposed algorithm outperforms the wavelet transform method and the Laplacian pyramid methods in preserving the edge and texture information. Qiguang Miao, Baoshu Wang |
FUSION | 1 |
| 2004 | Image fusion based on non-negative matrix factorizationabstractNonnegative Matrix Factorization technique (NMF) has been shown to have various applications to image processing, because of its power of local or part-based representation of objects and/or images. In this paper, we present an image fusion method based on NMF, not by the part-based representation feature of NMF, but by its wholly representation of the images needed to be fused: the images are fused by NMF with the parameter r of the NMF to be set to 1. Our experimental results show that the proposed method is efficient and effective for image fusion compared with many other image fusion methods. Le Wei, Qiguang Miao, Yue Joseph Wang |
ICIP | 3 |