EDBT 2026 Demo / reviewers in the wild / expert
Shaoyi Du
dblp:40/4846
· DBLP profile ↗
184ranked-venue papers
16as first author
116since 2021 · last 2026
0000-0002-7092-0596ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 99 · 5 first-author · 68 since 2021Graphics, computer vision, multimedia, augmented reality and games · 71 · 8 first-author · 44 since 2021Applied, interdisciplinary, general and emerging computing · 27 · 2 first-author · 16 since 2021Human-computer interaction and ubiquitous computing · 10 · 1 first-author · 3 since 2021Systems, architecture and hardware · 6 · 1 first-author · 3 since 2021Databases, data management, data science and information retrieval · 3 · 2 since 2021Computer networks · 2 · 2 since 2021Security and privacy · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | DAPE: Harmonizing Content-Position Encoding for Versatile Dense Visual PredictionabstractDense visual prediction tasks, including object detection and segmentation, inherently require precise and discriminative positional information to delineate object boundaries and pixel regions. Recent DETR-based frameworks advance dense prediction tasks through iterative attention applied to content queries, with sampled proposals as position references. However, this paradigm suffers from the misaligned sampling distribution and insufficient interaction between the content and position features, thereby limiting the encoding effectiveness. To overcome these limitations, we investigate the encoding paradigm for content-position harmonization and propose an effective predictor for dense visual tasks, termed DAPE (DETR with hArmonized content-Position Encoding). DAPE introduces explicit position encoding to facilitate content enhancement while maintaining low memory overhead. To achieves this process, DAPE comprises a Shifted Query Sampler (SQS) that enforces strict alignment between the distributions of content and position queries, and a 2D Low-Rank Position Encoder (LRPE) that progressively modulates attention maps based on the aligned representations. DAPE provides a unified solution for various dense prediction tasks. Extensive experiments on object detection, instance segmentation, and few-shot detection benchmarks demonstrate that DAPE achieves state-of-the-art performance while reducing memory consumption. Xiuquan Hou, Meiqin Liu 0001, Senlin Zhang, Shaoyi Du |
AAAI | 4 |
| 2026 | Cog-RAG: Cognitive-Inspired Dual-Hypergraph with Theme Alignment Retrieval-Augmented GenerationabstractRetrieval-Augmented Generation (RAG) enhances the response quality and domain-specific performance of large language models (LLMs) by incorporating external knowledge to combat hallucinations. In recent research, graph structures have been integrated into RAG to enhance the capture of semantic relations between entities. However, it primarily focuses on low-order pairwise entity relations, limiting the high-order associations among multiple entities. Hypergraph-enhanced approaches address this limitation by modeling multi-entity interactions via hyperedges, but they are typically constrained to inter-chunk entity-level representations, overlooking the global thematic organization and alignment across chunks. Drawing inspiration from the top-down cognitive process of human reasoning, we propose a theme-aligned dual-hypergraph RAG framework (Cog-RAG) that uses a theme hypergraph to capture inter-chunk thematic structure and an entity hypergraph to model high-order semantic relations. Furthermore, we design a cognitive-inspired two-stage retrieval strategy that first activates query-relevant thematic content from the theme hypergraph, and then guides fine-grained recall and diffusion in the entity hypergraph, achieving semantic alignment and consistent generation from global themes to local details. Our extensive experiments demonstrate that Cog-RAG significantly outperforms existing state-of-the-art baseline approaches. Yifan Feng 0001, Ruoxue Li, Rundong Xue, Xingliang Hou, Yue Gao 0002, Shaoyi Du |
AAAI | 8 |
| 2026 | Role Hypergraph Contrastive Learning for Multivariate Time-Series AnalysisabstractMultivariate Time-Series (MTS) analysis is crucial across various domains. Considering the spatial and temporal consistency of MTS, existing methods leverage graph structures with temporal augmentation and contrastive learning to achieve robust learning of spatial dependencies and temporal patterns. Given the inherent high-order correlations in MTS, hypergraphs present a promising approach. However, two key challenges limit their further development: 1) Feature-based perspectives capture limited spatial information, while structural perspectives encode richer spatial consistency and evolution dependency; 2) Various semantic patterns (e.g., synergy, inhibition) entangle in sensor correlations, leading to semantic ambiguity. The underlying reason is that conventional hypergraph structures cannot distinguish specific semantic roles within or across hyperedges. Thus, we propose Role Hypergraph Contrastive Learning for MTS analysis. Specifically, we introduce the concept of role to generalize hypergraphs to Role Hypergraphs, enabling precise modeling of sensor correlations by assigning each vertex-hyperedge pair with a semantic role. Building on this structure, we design a role hypergraph contrastive learning paradigm to comprehensively capture the spatial and temporal dependencies: From a structural perspective, role hypergraph structural contrasting captures spatial short-term consistency and long-term evolution; from a feature perspective, alignment of complementary role information ensures sensor-level temporal consistency. Experiments on classification and forecasting tasks demonstrate the effectiveness and interpretability of our method. Rundong Xue, Zhitao Zeng, Xiangmin Han, Shaoyi Du, Yue Gao 0002 |
AAAI | 6 |
| 2026 | MMRAG-RFT: Two-stage Reinforcement Fine-tuning for Explainable Multi-modal Retrieval-augmented GenerationabstractMulti-modal Retrieval-Augmented Generation (MMRAG) enables highly credible generation by integrating external multi-modal knowledge, thus demonstrating impressive performance in complex multi-modal scenarios. However, existing MMRAG methods fail to clarify the reasoning logic behind retrieval and response generation, which limits the explainability of the results. To address this gap, we propose to introduce reinforcement learning into multi-modal retrieval-augmented generation, enhancing the reasoning capabilities of multi-modal large language models through a two-stage reinforcement fine-tuning framework to achieve explainable multi-modal retrieval-augmented generation. Specifically, in the first stage, rule-based reinforcement fine-tuning is employed to perform coarse-grained point-wise ranking of multi-modal documents, effectively filtering out those that are significantly irrelevant. In the second stage, reasoning-based reinforcement fine-tuning is utilized to jointly optimize fine-grained list-wise ranking and answer generation, guiding multi-modal large language models to output explainable reasoning logic in the MMRAG process. Our method achieves state-of-the-art results on WebQA and MultimodalQA, two benchmark datasets for multi-modal retrieval-augmented generation, and its effectiveness is validated through comprehensive ablation experiments. Shengwei Zhao, Jingwen Yao, Sitong Wei, Linhai Xu, Yuying Liu 0007, Dong Zhang 0009, Shaoyi Du |
AAAI | 8 |
| 2026 | PASTA: Peak-Aware Axis-Factorized Spatiotemporal Alignment for Traffic Forecasting
Jingxi Feng, Shaoyi Du |
IV | 5 |
| 2026 | ClustView: Point clustering and depth view fusion for point cloud analysis
Xiaoyang Xiao, Yuanbo Chen, Runzhao Yao, Jue Jiang, Xinhu Zheng, Shaoyi Du, Long Guo |
Expert Syst. Appl. | 7 |
| 2026 | HOBN: A general multi-view high-order brain network learning framework for brain disease diagnosis
Rundong Xue, Shaoyi Du, Xiangmin Han, Dong Zhang 0009, Junchang Li |
Expert Syst. Appl. | 2 |
| 2026 | SoftHGNN: Soft Hypergraph Neural Networks for General Visual Recognition
Mengqi Lei, Siqi Li 0001, Xinhu Zheng, Shaoyi Du, Yue Gao 0002 |
Int. J. Comput. Vis. | 6 |
| 2026 | AsyCMST: Asymmetric cross-modal spatio-temporal learning for multimodal ultrasound nodule recognition
Hongcheng Han, Dong Zhang 0009, Qinbo Guo, Jue Jiang, Shaoyi Du |
Medical Image Anal. | 9 |
| 2026 | SCADA: Sparse cross attention for domain adaptive semantic segmentation
Qizhe Fan, Xiaoqin Shen, Yuanbo Chen, Shihui Ying, Jue Jiang, Shaoyi Du |
Neural Networks | 7 |
| 2026 | Knowledge-Embedded Hypergraph Neural NetworksabstractHypergraph Neural Networks (HGNNs) enhance graph-based modeling by representing complex relationships, with applications in brain network analysis, recommendation systems, and computer vision. However, conventional HGNNs often struggle with effective knowledge extraction and discriminative feature representation, leading to performance limitations. This paper presents Knowledge-Embedded Hypergraph Neural Networks (Knowledge HGNN), a framework that addresses these challenges with two complementary encoders and a multi-dimensional fusion strategy. The High-Order Incidence Encoder (HOI-Encoder) explicitly embeds structural knowledge by capturing permutation-invariant high-order incidence patterns that are typically overlooked by standard HGNNs. In contrast, the Task-Driven Rule Encoder (TDR-Encoder) focuses on feature-level knowledge, extracting task-related rules from vertex attributes through gradient boosted decision tree pre-training and encoding both rule content and positional importance. A Multi-Dimensional Knowledge Fusion module then integrates structural and rule-based embeddings, bridging semantic and dimensional gaps to form enriched vertex representations. The framework includes two implementations: Rule-Driven HGNN, which emphasizes rule-based knowledge, and Dual-Driven HGNN, which jointly leverages structural and rule-based knowledge for comprehensive feature extraction. Extensive experiments on ten datasets, together with ablation studies, demonstrate that Knowledge HGNN significantly improves performance, achieving a 7.3% gain on the Cora dataset and an average improvement of 2.5% across all datasets. These results highlight the effectiveness of explicitly differentiating and fusing structural and rule-based knowledge, setting a new standard for hypergraph applications in complex, data-driven scenarios. Yifan Feng 0001, Shaoyi Du, Shihui Ying, Zongze Wu 0001, Yue Gao 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | HGNN Shield: Defending Hypergraph Neural Networks Against High-Order Structure AttackabstractHypergraph Neural Networks (HGNNs) are crucial in modeling complex high-order correlations in diverse domains, utilizing hyperedges that connect multiple vertices. However, their susceptibility to structural attacks and irrational connections can disrupt message propagation and degrade performance. To address these issues, we introduce the HGNN Shield, a defense framework incorporating two key modules: Hyperedge-Dependent Estimation (HDE) and High-Order Shield (HOS). The HDE module prioritizes vertex dependencies within hyperedges and adapts traditional connectivity measures to hypergraphs, facilitating precise structural modifications. This adaptation allows for a nuanced assessment of vertex relationships within hyperedges, contributing theoretically by extending classical graph-based connection dependency measures to hypergraphs. Following HDE, the HOS module, positioned before convolutional layers, consists of three submodules: Hyperpath Cut, Hyperpath Link, and Hyperpath Refine. These components collectively detect, disconnect, and refine adversarial connections, ensuring robust message propagation. The theoretical contribution of the HOS module lies in maintaining hyperpath integrity and learning trajectory under adversarial conditions, providing a certifiable defense mechanism against high-order structural attacks. Experiments on six hypergraph datasets indicate that HGNN Shield significantly enhances robustness and maintains data integrity against targeted attacks, outperforming existing methods (an average performance improvement of 9.33% over other methods). Our framework not only improves HGNN reliability but also advances security in hypergraph-based applications. Yifan Feng 0001, Shaoyi Du, Shihui Ying, Jun-Hai Yong, Yue Gao 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | Hypergraph Foundation ModelabstractHypergraph neural networks (HGNNs) effectively model complex high-order relationships in domains like protein interactions and social networks by connecting multiple vertices through hyperedges, enhancing modeling capabilities, and reducing information loss. Developing foundation models for hypergraphs is challenging due to their distinct data, which includes both vertex features and intricate structural information. We present Hyper-FM, a Hypergraph Foundation Model for multi-domain knowledge extraction, featuring Hierarchical High-Order Neighbor Guided Vertex Knowledge Embedding for vertex feature representation and Hierarchical Multi-Hypergraph Guided Structural Knowledge Extraction for structural information. Additionally, we curate 11 text-attributed hypergraph datasets to advance research between HGNNs and LLMs. Experiments on these datasets show that Hyper-FM outperforms baseline methods by approximately 13.4%, validating our approach. Furthermore, we propose the first scaling law for hypergraph foundation models, demonstrating that increasing domain diversity significantly enhances performance, unlike merely augmenting vertex and hyperedge counts. This underscores the critical role of domain diversity in scaling hypergraph models. Yue Gao 0002, Yifan Feng 0001, Shiquan Liu, Xiangmin Han, Shaoyi Du, Zongze Wu 0001, Han Hu 0003 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2026 | Graph Quality Matters on Revealing the Semantics Behind the Data in Physical World
Jielong Yan, Shihui Ying, Shaoyi Du, Yue Gao 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2026 | Reinterpreting Hypergraph Kernels: Insights Through Homomorphism AnalysisabstractDesigning expressive hypergraph kernels that can effectively capture high-order structural information is a fundamental challenge in hypergraph learning. In this paper, we propose a novel comparison framework based on hypergraph homomorphisms to evaluate and compare the expressive ability of existing hypergraph kernels. We revisit classical kernels such as Hypergraph Weisfeiler-Lehman (HG WL) and Hypergraph Rooted kernels, providing theoretical conditions under which they fail to distinguish non-isomorphic hypergraphs. Motivated by these insights, we introduce the Hypergraph Subtree-Cycle Kernel, which augments subtree-based features with cycle-based structural patterns to enhance expressiveness. We propose two variants: HG SCKernelv1 and HG SCKernelv2. Extensive experiments on five graph and ten hypergraph classification benchmarks demonstrate the superior performance of our methods, confirming the effectiveness of integrating homomorphism-guided design into hypergraph kernels. Shaoyi Du, Yifan Feng 0001, Shihui Ying, Yue Gao 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2026 | Exploring dynamic interpretable brain networks via hierarchical graph transformer
Rundong Xue, Shaoyi Du, Xiangmin Han, Jingxi Feng, Zeyu Zhang 0006, Wei Zeng 0003, Yue Gao 0002 |
Pattern Recognit. | 3 |
| 2026 | Uncertainty-guided and reliable collaborative perception for open heterogeneous systems
Yihan Tian, ShuaiChen Zhu, Dong Zhang 0009, Yuying Liu 0007, Shaoyi Du |
Pattern Recognit. Lett. | 7 |
| 2026 | UnfoldDet: Advancing Surface Defect Detection With Dual Feature Separation and Relation ReasoningabstractSurface Defect Detection (SDD) aims to accurately localize defects based on predefined category labels in industrial manufacturing. Different from generic object detection, the industrial environment introduces significant challenges due to interference and unrelated background textures, leading to increased confusion between defect and non-defect features. In this work, we identify and analyze the structural characteristics and relations inherent in defect features. This analysis enables effectively distinguishing defects from non-defect areas, thereby enhancing the discriminative power for surface defect detection. Based on this insight, we propose a novel surface defect detection framework, named UnfoldDet. This framework focuses on separating defect and non-defect features and reasoning about the relations among defects. Specifically, we formulate the feature separation as an optimization problem with structural constraints. By expressing its iterations as network stages, we introduce an unfolding fusion module (UFM) to progressively separate and fuse multi-scale features. At the instance level, we propose a hierarchical relation encoder (HRE) to capture the inherent relations among defect instances. Through reasoning on positional and categorical relations, only highly related defect features are enhanced, while unrelated non-defect features are suppressed. Through extensive quantitative and qualitative experiments, as well as ablation studies on real-world datasets including ESD, CSD, and NEU-DET, we demonstrate the effectiveness of the proposed UnfoldDet in terms of both performance and computational efficiency. The code is available at https://github.com/xiuqhou/UnfoldDet. Xiuquan Hou, Meiqin Liu 0001, Shaoyi Du |
IEEE Trans. Circuits Syst. Video Technol. | 3 |
| 2026 | Fine-Grained Domain Alignment for Face Anti-Spoofing With Asymmetric Pseudo-Labels
Jing Yang 0014, Xusheng Cui, Yuehai Chen, Shaoyi Du, Badong Chen, Yuewen Liu |
IEEE Trans. Inf. Forensics Secur. | 4 |
| 2026 | Multi-Scale Temporal Analysis With a Dual-Branch Attention Network for Interpretable Gait-Based Classification of Neurodegenerative DiseasesabstractThe accurate diagnosis of neurodegenerative diseases (NDDs), such as Amyotrophic Lateral Sclerosis (ALS), Huntington's Disease (HD), and Parkinson's Disease (PD), remains a clinical challenge due to the complexity and subtlety of gait abnormalities. This paper proposes the Dual-Branch Attention-Enhanced Residual Network (DAERN), a novel deep learning architecture that integrates Dilated Causal Convolutions (DCCBlock) for local gait pattern extraction and Multi-Head Self-Attention (MHSA) for long-range dependency modeling. A Cross-Attention Fusion module enhances feature integration, while SHapley Additive exPlanations (SHAP) and Integrated Gradients (IG) improve interpretability, providing clinically relevant insights into gait-based NDD classification. Uniform Manifold Approximation and Projection (UMAP) visualizations reveal well-separated clusters corresponding to distinct NDDs categories, demonstrating the model's ability to capture discriminative features. Comprehensive ablation studies validate the contributions of model components and preprocessing strategies, highlighting the significance of each in achieving state-of-the-art classification performance. Experimental evaluations on the Gait in Neurodegenerative Disease (GaitNDD) dataset demonstrate that DAERN achieves an accuracy of 99.64%, an F1-score of 99.65%, and an AUC of 0.9997, significantly outperforming conventional deep learning and machine learning baselines. These findings suggest that DAERN could be a valuable and interpretable tool for clinical gait assessment, aiding in early-stage monitoring and automated screening of NDDs, with potential applications in real-time wearable sensor-based gait analysis. Wei Zeng 0003, Zhangbo Peng, Yang Chen 0045, Shaoyi Du |
IEEE J. Biomed. Health Informatics | 4 |
| 2026 | Keypoint-Guided Medical Video Segmentation Model With Spatiotemporal Feature FusionabstractAtrial fibrillation, characterized by high prevalence and poor prognosis, presents a significant global health burden. Accurate segmentation and measurement of left ventricular and left atrial appendage morphology and function are essential for reliable risk assessment. However, these tasks are hindered by ambiguous boundaries, complex cardiac motion, and sparse annotations. To address these challenges, we propose a Keypoint-Guided Medical Video Segmentation Model with Spatiotemporal Feature Fusion (KG-STS). First, we propose a shape-constrained point encoder that explicitly encodes boundary points to improve the representation of ambiguous boundaries. Next, we introduce a motion-aware alignment module that models cardiac motion by forming coherent motion information across frames. Building on these two modules, we develop a keypoint-guided spatiotemporal feature fusion module that integrates spatial boundary representations with temporal motion cues to enhance decoding features under sparse annotations, enabling temporally consistent segmentation and supporting morphological measurement. We evaluate the segmentation and measurement performance of our method on a self-constructed multi-view transesophageal echocardiography dataset and two publicly available transthoracic echocardiography datasets. The results demonstrate that KG-STS achieves superior temporal consistency in segmentation and higher accuracy in morphological measurements compared to competing methods. Shaoyi Du, Huanhuan Huo, Jue Jiang, Dong Zhang 0009, Hongcheng Han, Shengdi Hou |
IEEE Trans. Medical Imaging | 2 |
| 2026 | 3D Semantic Gaussian via Geometric-Semantic Hypergraph ComputationabstractSemantic labels are inherently tied to geometry and luminance reconstruction, as entities with similar shapes and appearances often share categories. Traditional methods use synthesis-analysis, NeRF, or 3D Gaussian representations to encode semantics and geometry separately. However, 2D methods lack view consistency, NeRF extensions are slow, and faster 3D Gaussian methods risk spatial and channel inconsistencies between semantic and RGB. Moreover, these methods require costly manual dense semantic labels. To alleviate resource demands and achieve effective semantic reconstruction with sparse inputs while enhancing RGB rendering quality, we build upon 3D Gaussian by integrating semantic features from pre-trained models-requiring no additional ground truth input-into Gaussian features, and construct a hypergraph neural network to capture higher-order correlations across RGB and semantic information as well as between different frames. Hypergraphs use hyperedges to link multiple vertices, capturing complex relationships essential for cross-modal tasks. This higher-order structure addresses the limitations of NeRF and Gaussian methods, which lack the capacity for such advanced associations. This framework enables precise novel view synthesis and 2D semantic reconstruction without manual annotations, achieving state-of-the-art results for RGB and semantic tasks on room-scale scenes in the ScanNet and Replica datasets, while supporting real-time rendering speeds of 34 FPS. Dejian Guo, Siqi Li 0001, Shaoyi Du, Xiangmin Han, Yue Gao 0002 |
IEEE Trans. Multim. | 5 |
| 2026 | GLU-Net: Global-Local Fusion Network for Event-Based Monocular Depth Estimation via Uncertainty OptimizationabstractEvent-based monocular depth estimation is crucial for applications such as autonomous driving, obstacle avoidance, and navigation under high-speed scenarios. Events exhibit a unique and irregular modality. To adapt them to neural networks, some studies convert event streams into event voxels or other frame-like representations. However, these approaches tend to lose the temporal characteristics of events. In this study, we propose a network that aggregates global voxel and per-channel temporal local features of event voxels across the temporal dimension, explicitly extracting events’ temporal information. Furthermore, as noise in events can interfere with the training process and is more difficult to predict than that in images, we utilize the uncertainty estimation module to mitigate the impact of uncertain factors and enhance the robustness of the model. Additionally, we employ multi-level depth features for supervisory training, which improves prediction performance compared to methods relying solely on ground-truth depth supervision. Experiments on open source datasets demonstrate the effectiveness of the proposed method. Our code can be found at https://github.com/WuShangjie/GLUNET . Shangjie Wu, Jihua Zhu, Zhikuan Zhou, Siqi Li 0001, Shaoyi Du, Yue Gao 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 5 |
| 2025 | Cross-Template-Based Hypergraph TransformerabstractSingle-template-based brain functional network analysis methods can provide limited functional connectivity information, which constrains the performance of brain disease diagnosis. Previous works have explored multi-template functional network analysis but failed to integrate the high-order correlation information within templates and the complementary information between templates into a unified relationship strength between nodes, and we extract the high-order correlation information within each template through hypergraph convolution. Secondly, for the analysis of functional connectivity between templates, we propose a cross-template Transformer to capture long-range dependencies between templates. A cross-template mask is applied to focus the model’s attention on important connections between templates, thereby enhancing model robustness. Finally, we progressively fuse the high-order information captured within templates with the global information across templates for downstream classification tasks. The proposed method has been validated on the public ABIDE dataset, and it outperforms existing methods in the ASD diagnosis task. Jingxi Feng, Xiangmin Han, Heming Xu, Jue Jiang, Shaoyi Du, Yue Gao 0002 |
ICASSP | 6 |
| 2025 | Beyond Graphs: Can Large Language Models Comprehend Hypergraphs?abstractExisting benchmarks like NLGraph and GraphQA evaluate LLMs on graphs by focusing mainly on pairwise relationships, overlooking the high-order correlations found in real-world data. Hypergraphs, which can model complex beyond-pairwise relationships, offer a more robust framework but are still underexplored in the context of LLMs. To address this gap, we introduce LLM4Hypergraph, the first comprehensive benchmark comprising 21,500 problems across eight low-order, five high-order, and two isomorphism tasks, utilizing both synthetic and real-world hypergraphs from citation networks and protein structures. We evaluate six prominent LLMs, including GPT-4o, demonstrating our benchmark’s effectiveness in identifying model strengths and weaknesses. Our specialized prompt- ing framework incorporates seven hypergraph languages and introduces two novel techniques, Hyper-BAG and Hyper-COT, which enhance high-order reasoning and achieve an average 4% (up to 9%) performance improvement on structure classification tasks. This work establishes a foundational testbed for integrating hypergraph computational capabilities into LLMs, advancing their comprehension. Yifan Feng 0001, Chengwu Yang, Xingliang Hou, Shaoyi Du, Shihui Ying, Zongze Wu 0001, Yue Gao 0002 |
ICLR | 4 |
| 2025 | ERetinex: Event Camera Meets Retinex Theory for Low-Light Image EnhancementabstractLow-light image enhancement aims to restore the under-exposure image captured in dark scenarios. Under such scenarios, traditional frame-based cameras may fail to capture the structure and color information due to the exposure time limitation. Event cameras are bio-inspired vision sensors that respond to pixel-wise brightness changes asynchronously. Event cameras' high dynamic range is pivotal for visual perception in extreme low-light scenarios, surpassing traditional cameras and enabling applications in challenging dark environments. In this paper, inspired by the success of the retinex theory for traditional frame-based low-light image restoration, we introduce the first methods that combine the retinex theory with event cameras and propose a novel retinex-based lowlight image restoration framework named ERetinex. Among our contributions, the first is developing a new approach that leverages the high temporal resolution data from event cameras with traditional image information to estimate scene illumination accurately. This method outperforms traditional image-only techniques, especially in low-light environments, by providing more precise lighting information. Additionally, we propose an effective fusion strategy that combines the high dynamic range data from event cameras with the color information of traditional images to enhance image quality. Through this fusion, we can generate clearer and more detailrich images, maintaining the integrity of visual information even under extreme lighting conditions. The experimental results indicate that our proposed method outperforms state-of-theart (SOTA) methods, achieving a gain of 1.0613 dB in PSNR while reducing FLOPS by 84.28 %. The code is available at https://github.com/lodew920/ERetinex. Xuejian Guo, Yuehang Wang, Siqi Li 0001, Yu Jiang 0006, Shaoyi Du, Yue Gao 0002 |
ICRA | 6 |
| 2025 | Cross-Modal Brain Graph Transformer via Function-Structure Connectivity Network for Brain Disease Diagnosis
Jingxi Feng, Heming Xu, Junhao Cai, Yujie Chang, Dong Zhang 0009, Shaoyi Du |
MICCAI (12) | 6 |
| 2025 | Adaptive Embedding for Long-Range High-Order Dependencies via Time-Varying Transformer on fMRI
Rundong Xue, Xiangmin Han, Zeyu Zhang 0006, Shaoyi Du, Yue Gao 0002 |
MICCAI (12) | 5 |
| 2025 | DHGFormer: Dynamic Hierarchical Graph Transformer for Disorder Brain Disease Diagnosis
Rundong Xue, Zeyu Zhang 0006, Xiangmin Han, Yue Gao 0002, Shaoyi Du |
MICCAI (12) | 7 |
| 2025 | Point-MaDi: Masked Autoencoding with Diffusion for Point Cloud Pre-trainingabstractSelf-supervised pre-training is essential for 3D point cloud representation learning, as annotating their irregular, topology-free structures is costly and labor-intensive. Masked autoencoders (MAEs) offer a promising framework but rely on explicit positional embeddings, such as patch center coordinates, which leak geometric information and limit data-driven structural learning. In this work, we propose Point-MaDi, a novel Point cloud Masked autoencoding Diffusion framework for pre-training that integrates a dual-diffusion pretext task into an MAE architecture to address this issue. Specifically, we introduce a center diffusion mechanism in the encoder, noising and predicting the coordinates of both visible and masked patch centers without ground-truth positional embeddings. These predicted centers are processed using a transformer with self-attention and cross-attention to capture intra- and inter-patch relationships. In the decoder, we design a conditional patch diffusion process, guided by the encoder's latent features and predicted centers to reconstruct masked patches directly from noise. This dual-diffusion design drives comprehensive global semantic and local geometric representations during pre-training, eliminating external geometric priors. Extensive experiments on ScanObjectNN, ModelNet40, ShapeNetPart, S3DIS, and ScanNet demonstrate that Point-MaDi achieves superior performance across downstream tasks, surpassing Point-MAE by 5.51\% on OBJ-BG, 5.17\% on OBJ-ONLY, and 4.34\% on PB-T50-RS for 3D object classification on the ScanObjectNN dataset. Xiaoyang Xiao, Runzhao Yao, Shaoyi Du |
NeurIPS | 4 |
| 2025 | DSGC-Net: A Dual-Stream Graph Convolutional Network for Crowd Counting via Feature Correlation Mining
Jinqiao Wei, Xionghui Zhao, Yidi Li 0001, Shaoyi Du, Bin Ren 0005, Nicu Sebe |
PRCV (17) | 5 |
| 2025 | DSK-YOLO: Feature-level Super Resolution Boosted Industrial Defect DetectionabstractDespite the significant advancements made in industrial defect detection, accurately and timely identifying complex and small-sized defects remains a challenge. Most current lightweight defect detectors are unable to fully extract both global and local contextual information due to their simplified network architectures. To address the above issues, this paper introduces a novel real-time detector DSK-YOLO, which efficiently enhances global and local contextual information with a lower number of parameters. Specifically, DSK-YOLO comprises two key components: DSKblock and DSKSR. The DSKblock employs dilated separable kernels to expand the effective receptive fields (ERFs) without deep layer stacking, thereby identifying complex defects. For small-sized defects detection, we develop a feature-level super resolution (SR) auxiliary branch to enhance local contextual information in the training phase. Moreover, the train-only SR branch brings no extra computational overhead for inference, making it an impressive choice for real-time tasks. Experimental results demonstrate that, on the industrial datasets NEU-DET and ESD, DSK-YOLO achieves mAP of 45.7% and 64.8%, which are 1.4% and 1.0% higher than those of the baseline model YOLOv8n. Our proposed DSK-YOLO offers a favorable tradeoff between precision and parameters compared to state-of-the-art models. Meichen Mu, Meiqin Liu 0001, Senlin Zhang, Shaoyi Du |
SMC | 4 |
| 2025 | Event-enhanced synthetic aperture imaging
Siqi Li 0001, Shaoyi Du, Jun-Hai Yong, Yue Gao 0002 |
Sci. China Inf. Sci. | 2 |
| 2025 | Dual-head detector with point-driven transformer and semantic-spatial gating for liquid crystal display defects
Chaofan Zhou, Meiqin Liu 0001, Senlin Zhang, Shanling Dong, Ronghao Zheng, Shaoyi Du |
Eng. Appl. Artif. Intell. | 6 |
| 2025 | Weak-edge sample extension for enhancing unsupervised feature learning
Yuehai Chen, Shanying Chen, Jing Yang 0014, Badong Chen, Shaoyi Du, Yuewen Liu |
Neurocomputing | 6 |
| 2025 | Hierarchical containment control with cluster consensus for multiagent systems under directional three-layer topology
Jingshu Sang, Dazhong Ma, Shaoyi Du |
Inf. Sci. | 4 |
| 2025 | Cross-sensor contrastive learning-based pre-training for machinery fault diagnosis under sample-limited conditions
Yue Ma 0008, Ruoxue Li, Zhixi Feng, Shuyuan Yang 0001, Shaoyi Du, Yue Gao 0002 |
Knowl. Based Syst. | 6 |
| 2025 | Hyper-YOLO: When Visual Object Detection Meets Hypergraph ComputationabstractWe introduce Hyper-YOLO, a new object detection method that integrates hypergraph computations to capture the complex high-order correlations among visual features. Traditional YOLO models, while powerful, have limitations in their neck designs that restrict the integration of cross-level features and the exploitation of high-order feature interrelationships. To address these challenges, we propose the Hypergraph Computation Empowered Semantic Collecting and Scattering (HGC-SCS) framework, which transposes visual feature maps into a semantic space and constructs a hypergraph for high-order message propagation. This enables the model to acquire both semantic and structural information, advancing beyond conventional feature-focused learning. Hyper-YOLO incorporates the proposed Mixed Aggregation Network (MANet) in its backbone for enhanced feature extraction and introduces the Hypergraph-Based Cross-Level and Cross-Position Representation Network (HyperC2Net) in its neck. HyperC2Net operates across five scales and breaks free from traditional grid structures, allowing for sophisticated high-order interactions across levels and positions. This synergy of components positions Hyper-YOLO as a state-of-the-art architecture in various scale models, as evidenced by its superior performance on the COCO dataset. Specifically, Hyper-YOLO-N significantly outperforms the advanced YOLOv8-N and YOLOv9-T with 12% and 9% improvements. Yifan Feng 0001, Jiangang Huang, Shaoyi Du, Shihui Ying, Jun-Hai Yong, Guiguang Ding, Rongrong Ji, Yue Gao 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 3 |
| 2025 | Self-Supervised Hypergraph Training Framework via Structure-Aware LearningabstractHypergraphs, with their ability to model complex, beyond pair-wise correlations, presents a significant advancement over traditional graphs for capturing intricate relational data across diverse domains. However, the integration of hypergraphs into self-supervised learning (SSL) frameworks has been hindered by the intricate nature of high-order structural variations. This paper introduces the Self-Supervised Hypergraph Training Framework via Structure-Aware Learning (SS-HT), designed to enhance the perception and measurement of these variations within hypergraphs. The SS-HT framework employs a "Masking and Re-Masking" strategy to bolster feature reconstruction in Hypergraph Neural Networks (HGNNs), addressing the limitations of traditional SSL methods. It also introduces a metric strategy for local high-order correlation changes, streamlining the computational efficiency of structural distance calculations. Extensive experiments on 11 datasets demonstrate SS-HT's superior performance over existing SSL methods for both low-order and high-order data. Notably, the framework significantly reduces data labeling dependency, achieving a 32% improvement over HGNN in the downstream task fine-tuning phase under the 1% labeled data setting in the Cora-CC dataset. Ablation studies further validate SS-HT's scalability and its capacity to augment the performance of various HGNN methods, underscoring its robustness and applicability in real-world scenarios. Yifan Feng 0001, Shiquan Liu, Shihui Ying, Shaoyi Du, Zongze Wu 0001, Yue Gao 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Kernelized Hypergraph Neural NetworksabstractHypergraph Neural Networks (HGNNs) have attracted much attention for high-order structural data learning. Existing methods mainly focus on simple mean-based aggregation or manually combining multiple aggregations to capture multiple information on hypergraphs. However, those methods inherently lack continuous non-linear modeling ability and are sensitive to varied distributions. Although some kernel-based aggregations on GNNs and CNNs can capture non-linear patterns to some degree, those methods are restricted in the low-order correlation and may cause unstable computation in training. In this work, we introduce Kernelized Hypergraph Neural Networks (KHGNN) and its variant, Half-Kernelized Hypergraph Neural Networks (H-KHGNN), which synergize mean-based and max-based aggregation functions to enhance representation learning on hypergraphs. KHGNN's kernelized aggregation strategy adaptively captures both semantic and structural information via learnable parameters, offering a mathematically grounded blend of kernelized aggregation approaches for comprehensive feature extraction. H-KHGNN addresses the challenge of overfitting in less intricate hypergraphs by employing non-linear aggregation selectively in the vertex-to-hyperedge message-passing process, thus reducing model complexity. Our theoretical contributions reveal a bounded gradient for kernelized aggregation, ensuring stability during training and inference. Empirical results demonstrate that KHGNN and H-KHGNN outperform state-of-the-art models across 10 graph/hypergraph datasets, with ablation studies demonstrating the effectiveness and computational stability of our method. Yifan Feng 0001, Shihui Ying, Shaoyi Du, Yue Gao 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | Inter-Intra Hypergraph Computation for Survival Prediction on Whole Slide ImagesabstractSurvival prediction on histopathology whole slide images (WSIs) involves the analysis of multi-level complex correlations, such as inter-correlations among patients and intra-correlations within gigapixel histopathology images. However, the current graph-based methods for WSI analysis mainly focus on the exploration of pairwise correlations, resulting in the loss of high-order correlations. Hypergraph-based methods can handle such high-order correlations, while existing hypergraph-based methods fail to integrate multi-level high-order correlations into a unified framework, which limits the representation capability of WSIs. In this work, we propose an inter-intra hypergraph computation (I$^{2}$2HGC) framework to address this issue. The I$^{2}$2HGC framework implements multi-level hypergraph computation for survival prediction on WSIs, namely intra-hypergraph computation and inter-hypergraph computation. Specifically, the intra-hypergraph computation considers each patch sampled from the histopathology WSI as a vertex of the intra-hypergraph and models the high-order correlations among all patches of an individual WSI in both topology and semantic feature spaces using a hypergraph structure. Then, the intra-hypergraph module generates the intra-embedding and intra-risk for each patient. Subsequently, the inter-hypergraph computation employs these intra-embeddings as features for each patient to form the population-level high-order correlations using data- and knowledge-driven hypergraph modeling strategies. Finally, the intra-risks and the inter-risks are fused for the final survival prediction of each patient. Extensive experimental results on four widely used TCGA carcinoma datasets are presented. We demonstrate that the hypergraph structure captures significantly richer correlations than the graph structure, encompassing all pairwise correlations as well as higher-order interactions through hyperedges. For WSIs with a vast number of pixels and complex correlations, hypergraph-based methods effectively capture topological and semantic information while mitigating the exponential growth of pairwise edges, offering practical advantages for large-scale medical image analysis. Xiangmin Han, Huijian Zhou, Shaoyi Du, Yue Gao 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2025 | HSC-T: B-Ultrasound-to-Elastography Translation via Hierarchical Structural Consistency Learning for Thyroid Cancer DiagnosisabstractElastography ultrasound imaging is increasingly important in the diagnosis of thyroid cancer and other diseases, but its reliance on specialized equipment and techniques limits widespread adoption. This paper proposes a novel multimodal ultrasound diagnostic pipeline that expands the application of elastography ultrasound by translating B-ultrasound (BUS) images into elastography images (EUS). Additionally, to address the limitations of existing image-to-image translation methods, which struggle to effectively model inter-sample variations and accurately capture regional-scale structural consistency, we propose a BUS-to-EUS translation method based on hierarchical structural consistency. By incorporating domain-level, sample-level, patch-level, and pixel-level constraints, our approach guides the model in learning a more precise mapping from BUS to EUS, thereby enhancing diagnostic accuracy. Experimental results demonstrate that the proposed method significantly improves the accuracy of BUS-to-EUS translation on the MTUSI dataset and that the generated elastography images enhance nodule diagnostic accuracy compared to solely using BUS images on the STUSI and the BUSI datasets. This advancement highlights the potential for broader application of elastography in clinical practice. Hongcheng Han, Qinbo Guo, Jue Jiang, Shaoyi Du |
IEEE J. Biomed. Health Informatics | 5 |
| 2025 | CSCC: Cross-Scene Crowd Counting via Learning to Diversify for Domain GeneralizationabstractIt is challenging for crowd counting models to generalize to new scenes due to domain shifts in training and test data. Although domain adaptation approaches have made notable progress in bridging the domain gap, they require target domain data. In this paper, we propose a novel framework for cross-scene crowd counting, which unifies domain generalization and adaptation. For domain generalization, we train a model only using single-domain data and the model can be generalized to any scene with satisfying performance. Regarding domain adaptation, we use both source and target domain data to further improve the performance. We first design a generation network that diversifies the generated samples to cover the unseen target domains as much as possible by minimizing mutual information. This approach simulates training data in various domains, thereby enhancing the model's generalization ability. Then we develop a pixel-wise supervised contrastive loss function that pulls the human heads in the source images and generated images closer to each other and pushes them further away from the background. This loss helps extract a domain-invariant feature representation, thus improving the model's generalization ability. Moreover, if information about the target domain is available, our generalization method can be easily applied as an adaptation method by replacing the mutual information minimization loss with the mutual information maximization loss. This can further improve cross-scene crowd counting performance. The experimental results demonstrate the strong generalizability of our method across different datasets. Yuehai Chen, Qingzhong Wang, Jing Yang 0014, Badong Chen, Haoyi Xiong, Shaoyi Du |
IEEE Trans. Multim. | 6 |
| 2025 | RDD: Learning Reinforced 3D Detectors and Descriptors Based on Policy GradientabstractKeypoint detection and descriptor matching are two vital steps in the 3D feature extraction framework, but they are difficult to learn in an end-to-end fashion due to their inherent discreteness. To tackle the non-differentiable operations, we formulate feature extraction as a decision-making problem: the network is treated as a policy pool that can make probabilistic estimations for keypoint selection and feature matching, supervised by maximizing a reward expectation of actions. In this way, we propose a novel end-to-end training paradigm of 3D feature extraction based on the stochastic policy gradient method, named Reinforced Detectors and Descriptors (RDD). Firstly, we propose a local-to-global probabilistic keypoint selection module that formulates the sampling probabilities of keypoints in a local-and-global mechanism to yield sparse and accurate keypoints. Secondly, we regard feature matching as an optimal transport problem and an efficient Sinkhorn method is leveraged to solve the optimal matching probabilities. In particular, we carefully design a reward function and derive gradients of probabilistic actions, thus overcoming the discreteness and providing reinforced supervision signals. Since our reward function is calculated from sampled keypoints rather than from randomly sampled points as in existing methods, the gap between training and inference is bridged. Experimental results demonstrate that our approach exceeds the quality of state-of-the-art methods and shows strong generalization ability. Remarkably, our approach can achieve significantly higher Registration Recall than other advanced methods when aligning scenes with a small number of keypoints, due to our highly accurate and repeatable detector. Wenting Cui, Shaoyi Du, Runzhao Yao, Canhui Tang, Aixue Ye |
IEEE Trans. Multim. | 2 |
| 2025 | Jointly Understand Your Command and Intention: Reciprocal Co-Evolution Between Scene-Aware 3D Human Motion Synthesis and AnalysisabstractAs two intimate reciprocal tasks, scene-aware human motion synthesis and analysis require a joint understanding between multiple modalities, including 3D body motions, 3D scenes, and textual descriptions. In this paper, we integrate these two paired processes into a Co-Evolving Synthesis-Analysis (CESA) pipeline and mutually benefit their learning. Specifically, scene aware text-to-human synthesis generates diverse indoor motion samples from the same textual description to enrich human scene interaction intra-class diversity, thus significantly benefiting training a robust human motion analysis system. Reciprocally, human motion analysis would enforce semantic scrutiny on each synthesized motion sample to ensure its semantic consistency with the given textual description, thus improving realistic motion synthesis. Considering that real-world indoor human motions are goal-oriented and path-guided, we propose a cascaded generation strategy that factorizes text-driven scene-specific human motion generation into three stages: goal inferring, path planning, and pose synthesizing. Coupling CESA with this powerful cascaded motion synthesis model, we jointly improve realistic human motion synthesis and robust human motion analysis in 3D scenes. Xuehao Gao, Yang Yang 0066, Shaoyi Du, Guo-Jun Qi, Junwei Han 0001 |
IEEE Trans. Multim. | 3 |
| 2025 | CCPoint: Contrasting Corrupted Point Clouds for Self-Supervised Representation LearningabstractSelf-supervised Learning (SSL), including mainstream contrastive learning, has achieved significant success in learning visual representations without the need for data annotations in 3D vision. While most contrastive learning methods focus on instance-level information through random affine transformations, they pay limited attention to the intrinsic structures within point clouds. In this work, we propose a novel SSL paradigm for point cloud representation learning, called CCPoint, which incorporates a novel form of data corruption as a negative augmentation strategy. Specifically, we degrade the input point cloud with various corruptions and conduct contrastive learning among the augmented, raw, and corrupted points to learn robust and discriminative representations. To preserve the semantic structure of the point cloud even under heavy degradation, an auxiliary reconstruction decoder is introduced into the corruption branch to provide an additional supervision signal. We explore four families of corruptions—affine, noise, masking, and combined transformations. Different from previous methods that rely on multi-modal data or complex network architectures, CCPoint achieves state-of-the-art performance on three widely used datasets (ModelNet40, ScanObjectNN, and ShapeNetPart) with a lightweight and efficient structure, reaching top linear accuracies of 92.4% and 86.2% on ModelNet40 and ScanObjectNN, respectively. Xiaoyang Xiao, Shaoyi Du, Meiqin Liu 0001, Xinhu Zheng |
IEEE Trans. Multim. | 2 |
| 2025 | Hypergraph Foundation Model for Brain Disease DiagnosisabstractThe goal of the hypergraph foundation model (HGFM) is to learn an encoder based on the hypergraph computational paradigm through self-supervised pretraining on high-order correlation structures, enabling the encoder to rapidly adapt to various downstream tasks in scenarios, where no labeled data or only a small amount of labeled data are available. The initial exploratory work has been applied to brain disease diagnosis tasks. However, existing methods primarily rely on graph-based approaches to learn low-order correlation patterns between brain regions in brain networks, neglecting the modeling and learning of complex correlations between different brain diseases and patients. This article proposes an HGFM for brain disease diagnosis, which conducts multidimensional pretraining tasks to explore latent cross-dimensional high-order correlation patterns on various brain disease datasets. HGFM is a high-order correlation-driven foundation model for brain disease diagnosis and effectively improves prediction performance. Specifically, HGFM first performs brain functional network link prediction tasks on individual brain networks and group interaction network link prediction tasks on group brain networks, constructing an HGFM for brain disease diagnosis. In downstream tasks, it achieves predictions for different brain disease diagnosis tasks through few-shot learning fine-tuning methods. The proposed method is evaluated on functional magnetic resonance imaging (fMRI) data from 4409 patients across four brain diseases. Results show that it outperforms existing state-of-the-art methods in all brain disease diagnosis tasks, demonstrating its potential value in clinical applications. Xiangmin Han, Rundong Xue, Jingxi Feng, Yifan Feng 0001, Shaoyi Du, Jun Shi 0004, Yue Gao 0002 |
IEEE Trans. Neural Networks Learn. Syst. | 5 |
| 2025 | Arbitrary Large-Scale Scene Reconstruction without Annotated Block PartitionsabstractLarge-scale scene reconstruction is a challenging problem. As different parts of the scene could be visible from different collected image frames, previous works manually use distance or geography to decompose the scene into parts and reconstruct each part of the scene separately. However, such manual decomposition is a laborious and time-consuming task when applied to large-scale scene reconstruction in real-world applications. To address this, we propose VisibleNeRF automatically reconstructs large-scale scenes by decomposing scenes into parts based on the part visibility. More specifically, we propose a visibility judgment strategy to decompose the scenes into visible and invisible parts. Then we reconstruct the visible part with the corresponding collected images and continue to decompose the rest of the invisible parts with the proposed visibility judgment strategy. New NeRF modules are re-established for the decomposed invisible parts until the entire scene is reconstructed. To the best of our knowledge, we are the first to propose an online reconstruction of large-scale scenes without manual decomposition. Experimental results on three datasets show that our method successfully reconstructs large-scale scenes in a fully automatic manner. Besides, in the widely used Mission Bay dataset, our model outperforms other state-of-the-art methods by a large margin. Lin Bie, Siqi Li 0001, Dejian Guo, Shaoyi Du, Yue Gao 0002 |
ACM Trans. Multim. Comput. Commun. Appl. | 6 |
| 2024 | PHFormer: Multi-Fragment Assembly Using Proxy-Level Hybrid TransformerabstractFragment assembly involves restoring broken objects to their original geometries, and has many applications, such as archaeological restoration. Existing learning based frameworks have shown potential for solving part assembly problems with semantic decomposition, but cannot handle such geometrical decomposition problems. In this work, we propose a novel assembly framework, proxy level hybrid Transformer, with the core idea of using a hybrid graph to model and reason complex structural relationships between patches of fragments, dubbed as proxies. To this end, we propose a hybrid attention module, composed of intra and inter attention layers, enabling capturing of crucial contextual information within fragments and relative structural knowledge across fragments. Furthermore, we propose an adjacency aware hierarchical pose estimator, exploiting a decompose and integrate strategy. It progressively predicts adjacent probability and relative poses between fragments, and then implicitly infers their absolute poses by dynamic information integration. Extensive experimental results demonstrate that our method effectively reduces assembly errors while maintaining fast inference speed. The code is available at https://github.com/521piglet/PHFormer. Wenting Cui, Runzhao Yao, Shaoyi Du |
AAAI | 3 |
| 2024 | 3D Feature Tracking via Event CameraabstractThis paper presents the first 3D feature tracking method with the corresponding dataset. Our proposed method takes event streams from stereo event cameras as input to pre-dict 3D trajectories of the target features with high-speed motion. To achieve this, our method leverages a joint framework to predict the 2D feature motion offsets and the 3D feature spatial position simultaneously. A motion compensation module is leveraged to overcome the feature deformation. A patch matching module based on bi-polarity hypergraph modeling is proposed to robustly es-timate the feature spatial position. Meanwhile, we collect the first 3D feature tracking dataset with high-speed moving objects and ground truth 3D feature trajectories at 250 FPS, named E-3DTrack, which can be used as the first high-speed 3D feature tracking benchmark. Our code and dataset could be found at: https://github.com/lisiqi19971013/E-3DTrack. Siqi Li 0001, Zhikuan Zhou, Zhou Xue, Shaoyi Du, Yue Gao 0002 |
CVPR | 5 |
| 2024 | ColorPCR: Color Point Cloud Registration with Multi-Stage Geometric-Color FusionabstractPoint cloud registration is still a challenging and open problem. For example, when the overlap between two point clouds is extremely low, geo-only features may be not suf-ficient. Therefore, it is important to further explore how to utilize color data in this task. Under such circumstances, we propose ColorPCR for color point cloud registration with multi-stage geometric-color fusion. We design a Hier-archical Color Enhanced Feature Extraction module to ex-tract multi-level geometric-color features, and a GeoColor Superpoint Matching Module to encode transformation-invariant geo-color global context for robust patch corre-spondences. In this way, both geometric and color data can be used, thus leading to robust performance even under extremely challenging scenarios, such as low overlap between two point clouds. To evaluate the performance of our method, we colorize 3DMatch/3DLoMatch datasets as Color3DMatch/Color3DLoMatch and evaluations on these datasets demonstrate the effectiveness of our proposed method. Our method achieves state-of-the-art registration recall of 97.5%/88.9% on them. Juncheng Mu, Lin Bie, Shaoyi Du, Yue Gao 0002 |
CVPR | 3 |
| 2024 | PARE-Net: Position-Aware Rotation-Equivariant Networks for Robust Point Cloud Registration
Runzhao Yao, Shaoyi Du, Wenting Cui, Canhui Tang, Chengwu Yang |
ECCV (74) | 2 |
| 2024 | Semantic Flow: Learning Semantic Fields of Dynamic Scenes from Monocular VideosabstractIn this work, we pioneer Semantic Flow, a neural semantic representation of dynamic scenes from monocular videos. In contrast to previous NeRF methods that reconstruct dynamic scenes from the colors and volume densities of individual points, Semantic Flow learns semantics from continuous flows that contain rich 3D motion information. As there is 2D-to-3D ambiguity problem in the viewing direction when extracting 3D flow features from 2D video frames, we consider the volume densities as opacity priors that describe the contributions of flow features to the semantics on the frames. More specifically, we first learn a flow network to predict flows in the dynamic scene, and propose a flow feature aggregation module to extract flow features from video frames. Then, we propose a flow attention module to extract motion information from flow features, which is followed by a semantic network to output semantic logits of flows. We integrate the logits with
volume densities in the viewing direction to supervise the flow features with semantic labels on video frames. Experimental results show that our model is able to learn from multiple dynamic scenes and supports a series of new tasks such as instance-level scene editing, semantic completions, dynamic scene tracking and semantic adaption on novel scenes. Fengrui Tian, Yueqi Duan, Angtian Wang, Jianfei Guo, Shaoyi Du |
ICLR | 5 |
| 2024 | Full-Dimensional Optimizable Network: A Channel, Frame and Joint-Specific Network Modeling for Skeleton-Based Action RecognitionabstractRecent human action recognition systems widely adopt graph convolution networks to extract spatial-temporal movement patterns. In graph convolution layers, inter-joint and inter-frame dependencies dominate spatial and temporal feature aggregation and thus are pivotal to representation learning. To enrich learned motion patterns, a powerful feature extractor should introduce its information propagation flexibility into three dimensions: (1) inferring different inter-joint correlations at different frames; (2) inferring different inter-frame correlations at different joints; (3) inferring different inter-joint and interframe correlations at different channels. In this paper, we take a closer look at effective feature aggregation in a skeleton sequence and propose a novel full-dimensional optimizable network with Channel, Frame and Joint-specific Network (CFJ-s Net) modeling for improving action recognition. By promoting dynamic information flows within different channels, frames, and joints, CFJ-s Net significantly extracts richer body posture features and trajectory features from a skeleton sequence. As verified on three large-scale datasets, NTU RGB+D, NTU RGB+D 120, and Northwestern-UCLA, CFJ-s Net achieves substantial improvements over state-of-the-art methods. Yang Yang 0066, Xuehao Gao, Shaoyi Du |
IJCNN | 4 |
| 2024 | ModaLink: Unifying Modalities for Efficient Image-to-PointCloud Place RecognitionabstractPlace recognition is an important task for robots and autonomous cars to localize themselves and close loops in pre-built maps. While single-modal sensor-based methods have shown satisfactory performance, cross-modal place recognition that retrieving images from a point-cloud database remains a challenging problem. Current cross-modal methods transform images into 3D points using depth estimation for modality conversion, which are usually computationally intensive and need expensive labeled data for depth supervision. In this work, we introduce a fast and lightweight framework to encode images and point clouds into place-distinctive descriptors. We propose an effective Field of View (FoV) transformation module to convert point clouds into an analogous modality as images. This module eliminates the necessity for depth estimation and helps subsequent modules achieve real-time performance. We further design a non-negative factorization-based encoder to extract mutually consistent semantic features between point clouds and images. This encoder yields more distinctive global descriptors for retrieval. Experimental results on the KITTI dataset show that our proposed methods achieve state-of-the-art performance while running in real time. Additional evaluation on the HAOMO dataset covering a 17 km trajectory further shows the practical generalization capabilities. We have released the implementation of our methods as open source at: https://github.com/haomo-ai/ModaLink.git. Weidong Xie, Lun Luo, Nanfei Ye, Shaoyi Du, Minhang Wang, Jintao Xu 0001, Rui Ai 0001, Weihao Gu, Xieyuanli Chen |
IROS | 5 |
| 2024 | Inter-intra High-Order Brain Network for ASD Diagnosis via Functional MRIs
Xiangmin Han, Rundong Xue, Shaoyi Du, Yue Gao 0002 |
MICCAI (2) | 3 |
| 2024 | ccRCC Metastasis Prediction via Exploring High-Order Correlations on Multiple WSIs
Huijian Zhou, Xiangmin Han, Shaoyi Du, Yue Gao 0002 |
MICCAI (5) | 4 |
| 2024 | Multi-grained Correspondence Learning of Audio-language Models for Few-shot Audio RecognitionabstractLarge-scale pre-trained audio-language models excel in general multi-modal representation, facilitating their adaptation to downstream audio recognition tasks in a data-efficient manner. However, existing few-shot audio recognition methods based on audio-language models primarily focus on learning coarse-grained correlations, which are not sufficient to capture the intricate matching patterns between the multi-level information of audio and the diverse characteristics of category concepts. To address this gap, we propose multi-grained correspondence learning for bootstrapping audio-language models to improve audio recognition with few training samples. This approach leverages generative models to enrich multi-modal representation learning, mining the multi-level information of audio alongside the diverse characteristics of category concepts. Multi-grained matching patterns are then established through multi-grained key-value cache and multi-grained cross-modal contrast, enhancing the alignment between audio and category concepts. Additionally, we incorporate optimal transport to tackle temporal misalignment and semantic intersection issues in fine-grained correspondence learning, enabling flexible fine-grained matching. Our method achieves state-of-the-art results on multiple benchmark datasets for few-shot audio recognition, with comprehensive ablation experiments validating its effectiveness. Shengwei Zhao, Linhai Xu, Yuying Liu 0007, Shaoyi Du |
ACM Multimedia | 4 |
| 2024 | SCC-CAM: Weakly Supervised Segmentation on Brain Tumor MRI with Similarity Constraint and Causality
Panpan Jiao, Xuejian Guo, Shaoyi Du |
PRCV (2) | 7 |
| 2024 | Robust colored point cloud alignment based on L*a*b* guided and Cauchy kernelabstractAbstract Precision agriculture benefits from point set registration, which can monitor plant health and growth in real time, promote the precise application of fertilizers and pesticides, and provide technical support for achieving sustainable development of agriculture. In this work, we propose a robust point set registration method for precision agriculture based on L*a*b* color guidance, bidirectional search and Cauchy distribution. First, the L*a*b* color guidance is applied to establish accurate correspondences between agricultural RGB‐D data. Second, the bidirectional nearest neighbor search strategy between point sets improves the reliability of establishing correspondences and broadens the convergence domain of the algorithm. Third, Cauchy distribution is utilized as an energy function for noise suppression, which further improves the robustness of the algorithm in dealing with complex vegetation scenes. Finally, results of ablation and simulation experiments indicate that the proposed registration algorithm can achieve more accurate and robust alignment results than other classic and state‐of‐the‐art point cloud registration algorithms to achieve monitoring and comparison of plant growth. Teng Wan, Shaoyi Du, Qiang Zhang 0049, Chunyao Huang, Wei Zeng 0003 |
Comput. Intell. | 2 |
| 2024 | RCI-Seg: Robust click-based interactive segmentation framework with deep reinforcement learning for biomedical images
Yueming He, Yang Li 0111, Shaoyi Du |
Neurocomputing | 5 |
| 2024 | Exploring conditional pixel-independent generation in GAN inversion for image processing
Chunyao Huang, Xiaomei Sun, Shaoyi Du, Wei Zeng 0003 |
Multim. Tools Appl. | 4 |
| 2024 | Hypergraph-Based Multi-Modal Representation for Open-Set 3D Object RetrievalabstractThe traditional 3D object retrieval (3DOR) task is under the close-set setting, which assumes the categories of objects in the retrieval stage are all seen in the training stage. Existing methods under this setting may tend to only lazily discriminate their categories, while not learning a generalized 3D object embedding. Under such circumstances, it is still a challenging and open problem in real-world applications due to the existence of various unseen categories. In this paper, we first introduce the open-set 3DOR task to expand the applications of the traditional 3DOR task. Then, we propose the Hypergraph-Based Multi-Modal Representation (HGM$^{2}$R) framework to learn 3D object embeddings from multi-modal representations under the open-set setting. The proposed framework is composed of two modules, i.e., the Multi-Modal 3D Object Embedding (MM3DOE) module and the Structure-Aware and Invariant Knowledge Learning (SAIKL) module. By utilizing the collaborative information of modalities derived from the same 3D object, the MM3DOE module is able to overcome the distinction across different modality representations and generate unified 3D object embeddings. Then, the SAIKL module utilizes the constructed hypergraph structure to model the high-order correlation among 3D objects from both seen and unseen categories. The SAIKL module also includes a memory bank that stores typical representations of 3D objects. By aligning with those memory anchors in the memory bank, the aligned embeddings can integrate the invariant knowledge to exhibit a powerful generalized capacity toward unseen categories. We formally prove that hypergraph modeling has better representative capability on data correlation than graph modeling. We generate four multi-modal datasets for the open-set 3DOR task, i.e., OS-ESB-core, OS-NTU-core, OS-MN40-core, and OS-ABO-core, in which each 3D object contains three modality representations: multi-view, point clouds, and voxel. Experiments on these four datasets show that the proposed method can significantly outperform existing methods. In particular, the proposed method outperforms the state-of-the-art by 12.12%/12.88% in terms of mAP on the OS-MN40-core/OS-ABO-core dataset, respectively. Results and visualizations demonstrate that the proposed method can effectively extract the generalized 3D object embeddings on the open-set 3DOR task and achieve satisfactory performance. Yifan Feng 0001, Shuyi Ji, Yu-Shen Liu, Shaoyi Du, Qionghai Dai, Yue Gao 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2024 | Hypergraph-Based Multi-View Action Recognition Using Event CamerasabstractAction recognition from video data forms a cornerstone with wide-ranging applications. Single-view action recognition faces limitations due to its reliance on a single viewpoint. In contrast, multi-view approaches capture complementary information from various viewpoints for improved accuracy. Recently, event cameras have emerged as innovative bio-inspired sensors, leading to advancements in event-based action recognition. However, existing works predominantly focus on single-view scenarios, leaving a gap in multi-view event data exploitation, particularly in challenges like information deficit and semantic misalignment. To bridge this gap, we introduceHyperMV, multi-view event-based action recognition framework. HyperMV converts discrete event data into frame-like representations and extracts view-related features using a shared convolutional network. By treating segments as vertices and constructing hyperedges using rule-based and KNN-based strategies, a multi-view hypergraph neural network that captures relationships across viewpoint and temporal features is established. The vertex attention hypergraph propagation is also introduced for enhanced feature fusion. To prompt research in this area, we present the largest multi-view event-based action dataset$\mathbf{THU}^{\mathbf{MV-EACT}}\mathbf{-50}$, comprising 50 actions from 6 viewpoints, which surpasses existing datasets by over tenfold. Experimental results show that HyperMV significantly outperforms baselines in both cross-subject and cross-view scenarios, and also exceeds the state-of-the-arts in frame-based multi-view action recognition. Yue Gao 0002, Jiaxuan Lu, Siqi Li 0001, Shaoyi Du |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Generative Variational-Contrastive Learning for Self-Supervised Point Cloud RepresentationabstractSelf-supervised representation learning for 3D point clouds has attracted increasing attention. However, existing methods in the field of 3D computer vision generally use fixed embeddings to represent the latent features, and impose hard constraints on the embeddings to make the latent feature values of the positive samples converge to consistency, which limits the ability of feature extractors to generalize over different data domains. To address this issue, we propose a Generative Variational-Contrastive Learning (GVC) model, where Gaussian distribution is used to construct a continuous, smoothed representation of the latent features. A distribution constraint and cross-supervision are constructed to improve the transfer ability of the feature extractor over synthetic and real-world data. Specifically, we design a variational contrastive module to constrain the feature distribution instead of feature values corresponding to each sample in the latent space. Moreover, a generative cross-supervision module is introduced to preserve the invariance features and promote the consistency of feature distribution among positive samples. Experimental results demonstrate that GVC achieves SOTA on different downstream tasks. In particular, with only pre-training on the synthetic dataset, GVC achieves a lead of 8.4% and 14.2% when transferring to the real-world dataset in the linear classification and few-shot classification. Bo-Hua Wang, Aixue Ye, Shaoyi Du, Yue Gao 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2024 | Self-adaptive subspace representation from a geometric intuition
Lipeng Cai, Jun Shi 0004, Shaoyi Du, Yue Gao 0002, Shihui Ying |
Pattern Recognit. | 3 |
| 2024 | IA-LSTM: Interaction-Aware LSTM for Pedestrian Trajectory PredictionabstractPredicting the trajectory of pedestrians in crowd scenarios is indispensable in self-driving or autonomous mobile robot field because estimating the future locations of pedestrians around is beneficial for policy decision to avoid collision. It is a challenging issue because humans have different walking motions, and the interactions between humans and objects in the current environment, especially between humans themselves, are complex. Previous researchers focused on how to model human-human interactions but neglected the relative importance of interactions. To address this issue, a novel mechanism based on correntropy is introduced. The proposed mechanism not only can measure the relative importance of human-human interactions but also can build personal space for each pedestrian. An interaction module, including this data-driven mechanism, is further proposed. In the proposed module, the data-driven mechanism can effectively extract the feature representations of dynamic human-human interactions in the scene and calculate the corresponding weights to represent the importance of different interactions. To share such social messages among pedestrians, an interaction-aware architecture based on long short-term memory network for trajectory prediction is designed. Experiments are conducted on two public datasets. Experimental results demonstrate that our model can achieve better performance than several latest methods with good performance. Jing Yang 0014, Yuehai Chen, Shaoyi Du, Badong Chen, José C. Príncipe |
IEEE Trans. Cybern. | 3 |
| 2024 | Learning Discriminative Features for Crowd CountingabstractCrowd counting models in highly congested areas confront two main challenges: weak localization ability and difficulty in differentiating between foreground and background, leading to inaccurate estimations. The reason is that objects in highly congested areas are normally small and high-level features extracted by convolutional neural networks are less discriminative to represent small objects. To address these problems, we propose a learning discriminative features framework for crowd counting, which is composed of a masked feature prediction module (MPM) and a supervised pixel-level contrastive learning module (CLM). The MPM randomly masks feature vectors in the feature map and then reconstructs them, allowing the model to learn about what is present in the masked regions and improving the model's ability to localize objects in high-density regions. The CLM pulls targets close to each other and pushes them far away from background in the feature space, enabling the model to discriminate foreground objects from background. Additionally, the proposed modules can be beneficial in various computer vision tasks, such as crowd counting and object detection, where dense scenes or cluttered environments pose challenges to accurate localization. The proposed two modules are plug-and-play, incorporating the proposed modules into existing models can potentially boost their performance in these scenarios. Yuehai Chen, Qingzhong Wang, Jing Yang 0014, Badong Chen, Haoyi Xiong, Shaoyi Du |
IEEE Trans. Image Process. | 6 |
| 2024 | Multi-Condition Latent Diffusion Network for Scene-Aware Neural Human Motion PredictionabstractInferring 3D human motion is fundamental in many applications, including understanding human activity and analyzing one's intention. While many fruitful efforts have been made to human motion prediction, most approaches focus on pose-driven prediction and inferring human motion in isolation from the contextual environment, thus leaving the body location movement in the scene behind. However, real-world human movements are goal-directed and highly influenced by the spatial layout of their surrounding scenes. In this paper, instead of planning future human motion in a "dark" room, we propose a Multi-Condition Latent Diffusion network (MCLD) that reformulates the human motion prediction task as a multi-condition joint inference problem based on the given historical 3D body motion and the current 3D scene contexts. Specifically, instead of directly modeling joint distribution over the raw motion sequences, MCLD performs a conditional diffusion process within the latent embedding space, characterizing the cross-modal mapping from the past body movement and current scene context condition embeddings to the future human motion embedding. Extensive experiments on large-scale human motion prediction datasets demonstrate that our MCLD achieves significant improvements over the state-of-the-art methods on both realistic and diverse predictions. Xuehao Gao, Yang Yang 0066, Yang Wu 0001, Shaoyi Du, Guo-Jun Qi |
IEEE Trans. Image Process. | 4 |
| 2024 | OTCLDA: Optimal Transport and Contrastive Learning for Domain Adaptive Semantic SegmentationabstractUnsupervised domain adaptive (UDA) semantic segmentation aims to assign a predetermined semantic label to every single pixel of the unannotated target data by exploiting a model that is trained on the labeled source data. Numerous current methods only display concern for grouping similar features together but ignore dispersing those features across various classes, so that some feature representations can not be well-separated. Therefore, we propose to employ contrastive learning (CL) method to increase the similarity of pixel features, propelling similar features closer and dispelling different ones far away. Furthermore, due to the domain shift, the UDA model frequently has poor generalization on the target domain. Accordingly, we design an optimal transport (OT) module to enhance UDA by comparing and aligning sample distributions to minimize transport loss between them. By taking advantage of this, the domain shift can be efficaciously mitigated by bringing the target probability distribution closer to that of the source. Specially, due to its simplicity, our OT module can be integrated into various UDA methods. In light of the aforementioned viewpoints, we put forth an ingenious approach, named OTCLDA, which successfully combines OT and CL while enhancing the performance of the UDA model. Multitudinous experiments demonstrate the importance of our method involving OT and CL. It significantly gains mIoU of 75.1% on benchmark GTA$\rightarrow$Cityscapes, and 66.9% on SYNTHIA$\rightarrow$Cityscapes respectively, displaying a competitive performance compared with previous works. The source code of OTCLDA is publicly available at https://github.com/YYDSDD/OTCLDA. Qizhe Fan, Xiaoqin Shen, Shihui Ying, Shaoyi Du |
IEEE Trans. Intell. Transp. Syst. | 4 |
| 2024 | Adaptive Distraction Recognition via Soft Prototype Learning and Probabilistic Label AlignmentabstractDistracted driving poses a serious threat to traffic safety and remains a widespread problem, highlighting the crucial need for effective recognition of distracted drivers. However, developing models that can generalize across diverse and changing real-world driving conditions is profoundly challenging. Variations in factors like lighting, weather, vehicle type, and drivers complicate generalization. Addressing this critical limitation is key to building recognition models with robust performance for practical deployment. This study presents a novel two-stage unsupervised domain adaptation framework to tackle the important challenge of recognizing distracted drivers across differing environments. The framework first constructs softly assigned class prototypes capturing underlying data structure by aggregating features locally and reweighting sample-prototype relationships globally, which increases the accuracy of class representations. The framework then aligns the probabilities between test samples and prototypes across source and target domains using soft distributional alignment, reducing domain gaps without explicit labeling of the target data. A growth control function balances prototype alignment with classification and adversarial losses. Experiments on distracted driver and object recognition datasets demonstrate this two-stage approach outperforms previous methods, especially under changing driving environments, which is an important problem distracted driving detection research must overcome to effectively enhance road safety. Yuying Liu 0007, Shaoyi Du, Hongcheng Han, Wei Zeng 0003 |
IEEE Trans. Intell. Transp. Syst. | 2 |
| 2024 | Semi-Supervised Domain Adaptation via Joint Transductive and Inductive Subspace LearningabstractMost existing shallow semi-supervised domain adaptation (SSDA) algorithms are based mainly on the framework adopting the maximum mean discrepancy (MMD) criterion, which is unstable and easily becomes stuck in a poor local minimum. Moreover, existing SSDA methods typically assume that the influence of the source domain is equivalent to that of the target domain, which is unreasonable and severely limits their performance. To address such drawbacks, we propose a novel SSDA framework derived from simple least squares regression (LSR) in a joint transductive and inductive learning paradigm, named transferable LSR (TLSR). Specifically, TLSR first learns domain-shared features using transfer component analysis (TCA) in a transductive paradigm. Then, TLSR augments the TCA features into the raw sample feature, formulating them into a block-diagonal matrix and training them in an inductive learning paradigm. This joint transductive and inductive learning paradigm helps alleviate the negative impacts of the MMD criterion of TCA but preserves the useful learned domain-shared knowledge. Moreover, the proposed block-diagonal input structure helps to separate the learned projections into independent domain-specific parts. Owing to the block-diagonal input structure, the influence of each domain can be reweighted, leading to significant improvements in performance. The experimental results demonstrate that the proposed TLSR outperforms the other shallow state-of-the-art competitors in 68 out of 90 cross-domain tasks. The source code of TLSR is available at:https://github.com/Evelhz/TLSR. Kaibing Zhang, Guofa Wang, Shaoyi Du |
IEEE Trans. Multim. | 5 |
| 2024 | Learning Heterogeneous Spatial-Temporal Context for Skeleton-Based Action RecognitionabstractGraph convolution networks (GCNs) have been widely used and achieved fruitful progress in the skeleton-based action recognition task. In GCNs, node interaction modeling dominates the context aggregation and, therefore, is crucial for a graph-based convolution kernel to extract representative features. In this article, we introduce a closer look at a powerful graph convolution formulation to capture rich movement patterns from these skeleton-based graphs. Specifically, we propose a novel heterogeneous graph convolution (HetGCN) that can be considered as the middle ground between the extremes of (2 + 1)-D and 3-D graph convolution. The core observation of HetGCN is that multiple information flows are jointly intertwined in a 3-D convolution kernel, including spatial, temporal, and spatial-temporal cues. Since spatial and temporal information flows characterize different cues for action recognition, HetGCN first dynamically analyzes pairwise interactions between each node and its cross-space-time neighbors and then encourages heterogeneous context aggregation among them. Considering the HetGCN as a generic convolution formulation, we further develop it into two specific instantiations (i.e., intra-scale and inter-scale HetGCN) that significantly facilitate cross-space-time and cross-scale learning on skeleton graphs. By integrating these modules, we propose a strong human action recognition system that outperforms state-of-the-art methods with the accuracy of 93.1% on NTU-60 cross-subject (X-Sub) benchmark, 88.9% on NTU-120 X-Sub benchmark, and 38.4% on kinetics skeleton. Xuehao Gao, Yang Yang 0066, Yang Wu 0001, Shaoyi Du |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | DAGCN: Dynamic and Adaptive Graph Convolutional Network for Salient Object DetectionabstractDeep-learning-based salient object detection (SOD) has achieved significant success in recent years. The SOD focuses on the context modeling of the scene information, and how to effectively model the context relationship in the scene is the key. However, it is difficult to build an effective context structure and model it. In this article, we propose a novel SOD method called dynamic and adaptive graph convolutional network (DAGCN) that is composed of two parts, adaptive neighborhood-wise graph convolutional network (AnwGCN) and spatially restricted K-nearest neighbors (SRKNN). The AnwGCN is novel adaptive neighborhood-wise graph convolution, which is used to model and analyze the saliency context. The SRKNN constructs the topological relationship of the saliency context by measuring the non-Euclidean spatial distance within a limited range. The proposed method constructs the context relationship as a topological graph by measuring the distance of the features in the non-Euclidean space, and conducts comparative modeling of context information through AnwGCN. The model has the ability to learn the metrics from features and can adapt to the hidden space distribution of the data. The description of the feature relationship is more accurate. Through the convolutional kernel adapted to the neighborhood, the model obtains the structure learning ability. Therefore, the graph convolution process can adapt to different graph data. Experimental results demonstrate that our solution achieves satisfactory performance on six widely used datasets and can also effectively detect camouflaged objects. Our code will be available at: https://github.com/CSIM-LUT/DAGCN.git. Ce Li 0001, Fenghua Liu, Shaoyi Du, Yang Wu 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2024 | GUESS: GradUally Enriching SyntheSis for Text-Driven Human Motion GenerationabstractIn this article, we propose a novel cascaded diffusion-based generative framework for text-driven human motion synthesis, which exploits a strategy named GradUally Enriching SyntheSis (GUESS as its abbreviation). The strategy sets up generation objectives by grouping body joints of detailed skeletons in close semantic proximity together and then replacing each of such joint group with a single body-part node. Such an operation recursively abstracts a human pose to coarser and coarser skeletons at multiple granularity levels. With gradually increasing the abstraction level, human motion becomes more and more concise and stable, significantly benefiting the cross-modal motion synthesis task. The whole text-driven human motion synthesis problem is then divided into multiple abstraction levels and solved with a multi-stage generation framework with a cascaded latent diffusion model: an initial generator first generates the coarsest human motion guess from a given text description; then, a series of successive generators gradually enrich the motion details based on the textual description and the previous synthesized results. Notably, we further integrate GUESS with the proposed dynamic multi-condition fusion mechanism to dynamically balance the cooperative effects of the given textual condition and synthesized coarse motion prompt in different generation stages. Extensive experiments on large-scale datasets verify that GUESS outperforms existing state-of-the-art methods by large margins in terms of accuracy, realisticness, and diversity. Xuehao Gao, Yang Yang 0066, Zhenyu Xie, Shaoyi Du, Zhongqian Sun, Yang Wu 0001 |
IEEE Trans. Vis. Comput. Graph. | 4 |
| 2023 | Decompose More and Aggregate Better: Two Closer Looks at Frequency Representation Learning for Human Motion PredictionabstractEncouraged by the effectiveness of encoding temporal dynamics within the frequency domain, recent human motion prediction systems prefer to first convert the motion representation from the original pose space into the frequency space. In this paper, we introduce two closer looks at effective frequency representation learning for robust motion prediction and summarize them as: decompose more and aggregate better. Motivated by these two insights, we develop two powerful units that factorize the frequency representation learning task with a novel decomposition-aggregation two-stage strategy: (1) frequency decomposition unit unweaves multi-view frequency representations from an input body motion by embedding its frequency features into multiple spaces; (2) feature aggregation unit deploys a series of intra-space and inter-space feature aggregation layers to collect comprehensive frequency representations from these spaces for robust human motion prediction. As evaluated on large-scale datasets, we develop a strong baseline model for the human motion prediction task that outperforms state-of-the-art methods by large margins: 8%∼12% on Human3.6M, 3%∼7% on CMU MoCap, and 7%∼10% on 3DPW. Xuehao Gao, Shaoyi Du, Yang Wu 0001, Yang Yang 0066 |
CVPR | 2 |
| 2023 | MonoNeRF: Learning a Generalizable Dynamic Radiance Field from Monocular VideosabstractIn this paper, we target at the problem of learning a generalizable dynamic radiance field from monocular videos. Different from most existing NeRF methods that are based on multiple views, monocular videos only contain one view at each timestamp, thereby suffering from ambiguity along the view direction in estimating point features and scene flows. Previous studies such as DynNeRF disambiguate point features by positional encoding, which is not transferable and severely limits the generalization ability. As a result, these methods have to train one independent model for each scene and suffer from heavy computational costs when applying to increasing monocular videos in real-world applications. To address this, We propose MonoNeRF to simultaneously learn point features and scene flows with point trajectory and feature correspondence constraints across frames. More specifically, we learn an implicit velocity field to estimate point trajectory from temporal features with Neural ODE, which is followed by a flow-based feature aggregation module to obtain spatial features along the point trajectory. We jointly optimize temporal and spatial features in an end-to-end manner. Experiments show that our MonoNeRF is able to learn from multiple scenes and support new applications such as scene editing, unseen frame synthesis, and fast novel scene adaptation. Codes are available at https://github.com/tianfr/MonoNeRF. Fengrui Tian, Shaoyi Du, Yueqi Duan |
ICCV | 2 |
| 2023 | HybridPoint: Point Cloud Registration Based on Hybrid Point Sampling and MatchingabstractPatch-to-point matching has become a robust way of point cloud registration. However, previous patch-matching methods employ superpoints with poor localization precision as nodes, which may lead to ambiguous patch partitions. In this paper, we propose a HybridPoint-based network to find more robust and accurate correspondences. Firstly, we propose to use salient points with prominent local features as nodes to increase patch repeatability, and introduce some uniformly distributed points to complete the point cloud, thus constituting hybrid points. Hybrid points not only have better localization precision but also give a complete picture of the whole point cloud. Furthermore, based on the characteristic of hybrid points, we propose a dual-classes patch matching module, which leverages the matching results of salient points and filters the matching noise of non-salient points. Experiments show that our model achieves state-of-the-art performance on 3DMatch, 3DLoMatch, and KITTI odometry, especially with 93.0% Registration Recall on the 3DMatch dataset. Our code and models are available at https://github.com/liyih/HybridPoint. Canhui Tang, Runzhao Yao, Aixue Ye, Shaoyi Du |
ICME | 6 |
| 2023 | Color-Difference Correntropy Guided Convolution Network for Point Cloud Semantic SegmentationabstractWith the development of data acquisition technology, RGB color information is widely collected to strengthen the 3D point cloud. To some extent, RGB color information contains the prior relation about the spatial position of objects. However, the existing point cloud segmentation networks based on deep learning do not pay much attention to it. To remedy the lack in this area, we propose a color-difference correntropy guided convolution network, which introduces correntropy to optimize the measurement of color-difference. Meanwhile, we select points in the local neighborhood via the color-difference guided module, and construct an ordered sequence of points with correlation information, which not only facilitates the feature extraction by directly applying the convolution but also fully studies the correlation between the color information and spatial position of the point cloud. Moreover, we fuse the sequence features extracted by convolution with the geometric features acquired by MLP to get new features with more abundant semantic information, thus improving the segmentation performance. On both indoor and outdoor datasets, the experimental results demonstrate the effectiveness and superiority of the proposed method by the comparison experiments and ablation experiments. Zhou Jiang 0001, Jing Yang 0014, Chunyu Xuan, Dong Zhang 0009, Shaoyi Du |
IJCNN | 5 |
| 2023 | HD2Reg: Hierarchical Descriptors and Detectors for Point Cloud RegistrationabstractFeature Descriptors and Detectors are two main components of feature-based point cloud registration. However, little attention has been drawn to the explicit representation of local and global semantics in the learning of descriptors and detectors. In this paper, we present a framework that explicitly extracts dual-level descriptors and detectors and performs coarse-to-fine matching with them. First, to explicitly learn local and global semantics, we propose a hierarchical contrastive learning strategy, training the robust matching ability of high-level descriptors, and refining the local feature space using low-level descriptors. Furthermore, we propose to learn dual-level saliency maps that extract two groups of keypoints in two different senses. To overcome the weak supervision of binary matchability labels, we propose a ranking strategy to label the significance ranking of keypoints, and thus provide more fine-grained supervision signals. Finally, we propose a global-to-local matching scheme to obtain robust and accurate correspondences by leveraging the complementary dual-level features. Quantitative experiments on 3DMatch and KITTI odometry datasets show that our method achieves robust and accurate point cloud registration and outperforms recent keypoint-based methods. [code release] Canhui Tang, Shaoyi Du, Guofa Wang |
IV | 3 |
| 2023 | Interpretable Driver Fatigue Estimation Based on Hierarchical Symptom Representations
Jiaqin Lin, Shaoyi Du, Yuying Liu 0007, Nanning Zheng 0001 |
MMM (2) | 2 |
| 2023 | CMFG: Cross-Model Fine-Grained Feature Interaction for Text-Video Retrieval
Shengwei Zhao, Yuying Liu 0007, Shaoyi Du, Linhai Xu |
MMM (2) | 3 |
| 2023 | Multi-grained Representation Learning for Cross-modal RetrievalabstractThe purpose of audio-text retrieval is to learn a cross-modal similarity function between audio and text, enabling a given audio/text to find similar text/audio from a candidate set. Recent audio-text retrieval models aggregate multi-modal features into a single-grained representation. However, single-grained representation is difficult to solve the situation that an audio is described by multiple texts of different granularity levels, because the association pattern between audio and text is complex. Therefore, we propose an adaptive aggregation strategy to automatically find the optimal pool function to aggregate the features into a comprehensive representation, so as to learn valuable multi-grained representation. And multi-grained comparative learning is carried out in order to focus on the complex correlation between audio and text in different granularity. Meanwhile, text-guided token interaction is used to reduce the impact of redundant audio clips. We evaluated our proposed method on two audio-text retrieval benchmark datasets of Audiocaps and Clotho, achieving the state-of-the-art results in text-to-audio and audio-to-text retrieval. Our findings emphasize the importance of learning multi-modal multi-grained representation. Shengwei Zhao, Linhai Xu, Yuying Liu 0007, Shaoyi Du |
SIGIR | 4 |
| 2023 | Video Self-Supervised Cross-Pathway Training Based on Slow and Fast PathwaysabstractIn the field of video self-supervised learning, contrastive instance learning methods suffer from a lack of semantic information, resulting in inadequate generalization in downstream tasks. Although optical flow can provide some semantic information, it requires significant computational cost prior to training. To address this, we propose a Video self-supervised Cross-pathway training model based on Slow and Fast pathways (VCSF). This model separately extracts temporal and spatial features from pure RGB video frames, and uses the complementary representations of the two pathways to conduct cross-pathway training. Additionally, we propose a motion perception module in the low-frame-rate space to enhance the network's ability to perceive rapidly changing human motion. We conducted extensive experiments in downstream missions of UCF101 and HMDB51, and obtained state-of-the-art results in models using the UCF101 data set for self-supervised pre-training, including motion recognition and nearest neighbor retrieval. Jing Yang 0014, Zhou Jiang 0001, Yuehai Chen, Shaoyi Du |
SMC | 5 |
| 2023 | PGF-BIQA: Blind image quality assessment via probability multi-grained cascade forest
Hao Liu 0060, Ce Li 0001, Shangang Jin, Weizhe Gao, Fenghua Liu, Shaoyi Du, Shihui Ying |
Comput. Vis. Image Underst. | 6 |
| 2023 | Curriculum classification network based on margin balancing multi-loss and ensemble learning
Shaoyi Du, Yuying Liu 0007, Xijing Wang, Yuting Chi, Nanning Zheng 0001, Yucheng Guo |
Future Gener. Comput. Syst. | 1 |
| 2023 | Detach and unite: A simple meta-transfer for few-shot learning
Yaoyue Zheng, Xuetao Zhang 0001, Wei Zeng 0003, Shaoyi Du |
Knowl. Based Syst. | 5 |
| 2023 | Glimpse and focus: Global and local-scale graph convolution network for skeleton-based action recognition
Xuehao Gao, Shaoyi Du, Yang Yang 0066 |
Neural Networks | 2 |
| 2023 | Action Recognition and Benchmark Using Event CamerasabstractRecent years have witnessed remarkable achievements in video-based action recognition. Apart from traditional frame-based cameras, event cameras are bio-inspired vision sensors that only record pixel-wise brightness changes rather than the brightness value. However, little effort has been made in event-based action recognition, and large-scale public datasets are also nearly unavailable. In this paper, we propose an event-based action recognition framework calledEV-ACT. The Learnable Multi-Fused Representation (LMFR) is first proposed to integrate multiple event information in a learnable manner. The LMFR with dual temporal granularity is fed into the event-based slow-fast network for the fusion of appearance and motion features. A spatial-temporal attention mechanism is introduced to further enhance the learning capability of action recognition. To prompt research in this direction, we have collected the largest event-based action recognition benchmark namedTHUE-ACT-50and the accompanyingTHUE-ACT-50-CHLdataset under challenging environments, including a total of over 12,830 recordings from 50 action categories, which is over 4 times the size of the previous largest dataset. Experimental results show that our proposed framework could achieve improvements of over 14.5%, 7.6%, 11.2%, and 7.4% compared to previous works on four benchmarks. We have also deployed our proposed EV-ACT framework on a mobile platform to validate its practicality and efficiency. Yue Gao 0002, Jiaxuan Lu, Siqi Li 0001, Nan Ma 0012, Shaoyi Du, Qionghai Dai |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2023 | STORM: Structure-Based Overlap Matching for Partial Point Cloud RegistrationabstractPartial point cloud registration aims to transform partial scans into a common coordinate system. It is an important preprocessing step to generate complete 3D shapes. Although previous registration methods have made great progress in recent decades, traditional registration methods, such as Iterative Closest Point (ICP) and its variants, all these methods highly depend on the sufficient overlaps between two point clouds, because they cannot distinguish outlier correspondences. Note that the overlap between point clouds could always be small, which limits the application of these methods. To tackle this problem, we present a StrucTure-based OveRlap Matching (STORM) method for partial point cloud registration. In our method, an overlap prediction module with differentiable sampling is designed to detect points in overlap utilizing structure information, and facilitates exact partial correspondence generation, which is based on discriminative pointwise feature similarity. The pointwise features which contain effective structural information are extracted by graph-based methods. Experimental results and comparison with state-of-the-art methods demonstrate that STORM can achieve better performance. Moreover, most registration methods perform worse when the overlap ratio decreases, while STORM can still achieve satisfactory performance when the overlap ratio is small. Chenggang Yan 0001, Yutong Feng, Shaoyi Du, Qionghai Dai, Yue Gao 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 4 |
| 2023 | Hunter: Exploring High-Order Consistency for Point Cloud Registration With Severe OutliersabstractAfter decades of investigation, point cloud registration is still a challenging task in practice, especially when the correspondences are contaminated by a large number of outliers. It may result in a rapidly decreasing probability of generating a hypothesis close to the true transformation, leading to the failure of point cloud registration. To tackle this problem, we propose a transformation estimation method, named Hunter, for robust point cloud registration with severe outliers. The core of Hunter is to design a global-to-local exploration scheme to robustly find the correct correspondences. The global exploration aims to exploit guided sampling to generate promising initial alignments. To this end, a hypergraph-based consistency reasoning module is introduced to learn the high-order consistency among correct correspondences, which is able to yield a more distinct inlier cluster that facilitates the generation of all-inlier hypotheses. Moreover, we propose a preference-based local exploration module that exploits the preference information of top- k promising hypotheses to find a better transformation. This module can efficiently obtain multiple reliable transformation hypotheses by using a multi-initialization searching strategy. Finally, we present a distance-angle based hypothesis selection criterion to choose the most reliable transformation, which can avoid selecting symmetrically aligned false transformations. Experimental results on simulated, indoor, and outdoor datasets, demonstrate that Hunter can achieve significant superiority over the state-of-the-art methods, including both learning-based and traditional methods (as shown in Fig. 1). Moreover, experimental results also indicate that Hunter can achieve more stable performance compared with all other methods with severe outliers. Runzhao Yao, Shaoyi Du, Wenting Cui, Aixue Ye, Hongbo Zhang 0004, Yue Gao 0002 |
IEEE Trans. Pattern Anal. Mach. Intell. | 2 |
| 2023 | Coarse-to-fine feature representation based on deformable partition attention for melanoma identification
Dong Zhang 0009, Jing Yang 0014, Shaoyi Du, Hongcheng Han, Yuyan Ge, Longfei Zhu, Ce Li 0001, Meifeng Xu, Nanning Zheng 0001 |
Pattern Recognit. | 3 |
| 2023 | Counting Varying Density Crowds Through Density Guided Adaptive Selection CNN and Transformer EstimationabstractIn real-world crowd counting applications, the crowd densities in an image vary greatly. When facing density variation, humans tend to locate and count the targets in low-density regions, and reason the number in high-density regions. We observe that CNN focus on the local information correlation using a fixed-size convolution kernel and the Transformer could effectively extract the semantic crowd information by using the global self-attention mechanism. Thus, CNN could locate and estimate crowds accurately in low-density regions, while it is hard to properly perceive the densities in high-density regions. On the contrary, Transformer has a high reliability in high-density regions, but fails to locate the targets in sparse regions. Neither CNN nor Transformer can well deal with this kind of density variation. To address this problem, we propose a CNN and Transformer Adaptive Selection Network (CTASNet) which can adaptively select the appropriate counting branch for different density regions. Firstly, CTASNet generates the prediction results of CNN and Transformer. Then, considering that CNN/Transformer is appropriate for low/high-density regions, a density guided adaptive selection module is designed to automatically combine the predictions of CNN and Transformer. Moreover, to reduce the influences of annotation noise, we introduce a Correntropy based optimal transport loss. Extensive experiments on four challenging crowd counting datasets have validated the proposed method. Yuehai Chen, Jing Yang 0014, Badong Chen, Shaoyi Du |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2023 | Tolerating Annotation Displacement in Dense Object Counting via Point Annotation Probability MapabstractCounting objects in crowded scenes remains a challenge to computer vision. The current deep learning based approach often formulate it as a Gaussian density regression problem. Such a brute-force regression, though effective, may not consider the annotation displacement properly which arises from the human annotation process and may lead to different distributions. We conjecture that it would be beneficial to consider the annotation displacement in the dense object counting task. To obtain strong robustness against annotation displacement, generalized Gaussian distribution (GGD) function with a tunable bandwidth and shape parameter is exploited to form the learning target point annotation probability map, PAPM. Specifically, we first present a hand-designed PAPM method (HD-PAPM), in which we design a function based on GGD to tolerate the annotation displacement. For end-to-end training, the hand-designed PAPM may not be optimal for the particular network and dataset. An adaptively learned PAPM method (AL-PAPM) is proposed. To improve the robustness to annotation displacement, we design an effective transport cost function based on GGD. The proposed PAPM is capable of integration with other methods. We also combine PAPM with P2PNet through modifying the matching cost matrix, forming P2P-PAPM. This could also improve the robustness to annotation displacement of P2PNet. Extensive experiments show the superiority of our proposed methods. Yuehai Chen, Jing Yang 0014, Badong Chen, Shaoyi Du, Gang Hua 0001 |
IEEE Trans. Image Process. | 4 |
| 2023 | An Uncertainty-Aware and Sex-Prior Guided Biological Age Estimation From Orthopantomogram ImagesabstractBone age, as a measure of biological age (BA), plays an important role in a variety of fields, including forensics, orthodontics, sports, and immigration. Despite its significance, accurate estimation of BA remains a challenge due to the uncertainty error between BA and chronological age (CA) caused by individual diversity and the difficult integration of multiple factors, such as sex, and identified or measured anatomical structures, into the estimation process. To address problems, we propose an uncertainty-aware and sex-prior guided biological age estimation from orthopantomogram images (OPGs), named UASP-BAE, which models uncertainty errors while setting sex dimorphism as tractive features to enhance age-related specific features, aiming to improve the accuracy of BA estimation. Furthermore, considering the global relevance of the anatomic structure, such as the mandible, teeth, maxillary sinus, etc., a cross-attention module based on CNN and self-attention is proposed to mine the local texture and global semantic features of OPGs. Moreover, we design a novel age composition loss by cross-entropy, probability bias, and regression functions, aiming at evaluating BA's uncertainty errors and results to obtain an accurate and robust model. On 10703 OPGs from 5.00 to 25.00 years of age, our model had a best MAE value of 0.8005 years and higher than the comparison popular algorithms, which also demonstrates the method's potential for improved accuracy in BA estimation. Dong Zhang 0009, Jing Yang 0014, Shaoyi Du, Wenqing Bu, Yu-Cheng Guo |
IEEE J. Biomed. Health Informatics | 3 |
| 2023 | Efficient Spatio-Temporal Contrastive Learning for Skeleton-Based 3-D Action RecognitionabstractIn this paper, we propose a simple yet effective self-supervised method called spatio-temporal contrastive learning (ST-CL) for 3D skeleton-based action recognition. ST-CL acquires action-specific features by regarding the spatio-temporal continuity of motion tendency as the supervisory signal. To yield effective representations, ST-CL first designs some novel contrastive proxy tasks by providing different spatio-temporal observation scenes for the same 3D action and pulling them together in the embedding space. Second, three key components are devised in the action encoding to efficiently extract representations in contrastive tasks: (1) Information Representation introduces the awareness of joint type when analyzing motion dynamics. (2) Non-local GCN learns a data-driven graph topology structure and promotes a spatial message passing among long-range joints in each frame. (3) Multi-Scale TCN makes larger receptive fields for capturing richer longe-range temporal dynamics amomg adjacent frames. In ST-CL, these effective proxy tasks yield useful representations and efficient action encoding further enhances the representation capacity. As validated on four large-scale datasets, ST-CL is a strong baseline with high performance and efficiency for the contrastive learning study of the skeleton data. Compared to previous self-supervised methods, the proposed ST-CL achieves significant improvement consistently with a smaller model size and better training efficiency. Xuehao Gao, Yang Yang 0066, Maosen Li, Jin-Gang Yu, Shaoyi Du |
IEEE Trans. Multim. | 6 |
| 2022 | TCVM: Temporal Contrasting Video Montage Framework for Self-supervised Video Representation Learning
Fengrui Tian, Xie Yu, Shaoyi Du, Meina Song |
ACCV (2) | 4 |
| 2022 | C-CAM: Causal CAM for Weakly Supervised Semantic Segmentation on Medical ImageabstractRecently, many excellent weakly supervised semantic segmentation (WSSS) works are proposed based on class activation mapping (CAM). However, there are few works that consider the characteristics of medical images. In this paper, we find that there are mainly two challenges of medical images in WSSS: i) the boundary of object foreground and background is not clear; ii) the co-occurrence phenomenon is very severe in training stage. We thus propose a Causal CAM (C-CAM) method to overcome the above challenges. Our method is motivated by two cause-effect chains including category-causality chain and anatomy-causality chain. The category-causality chain represents the image content (cause) affects the category (effect). The anatomy-causality chain represents the anatomical structure (cause) affects the organ segmentation (effect). Extensive experiments were conducted on three public medical image data sets. Our C-CAM generates the best pseudo masks with the DSC of 77.26%, 80.34% and 78.15% on ProMRI, ACDC and CHAOS compared with other CAM-like methods. The pseudo masks of C-CAM are further used to improve the segmentation performance for organ segmentation tasks. Our C-CAM achieves DSC of 83.83% on ProMRI and DSC of 87.54% on ACDC, which outperforms state-of-the-art WSSS methods. Our code is available at https://github.com/Tian-lab/C-CAM. Jihua Zhu, Ce Li 0001, Shaoyi Du |
CVPR | 5 |
| 2022 | Adaptive weighted robust iterative closest point
Yu Guo 0006, Luting Zhao, Xuetao Zhang 0001, Shaoyi Du, Fei Wang 0008 |
Neurocomputing | 5 |
| 2022 | Action recognition based on RGB and skeleton data sets: A survey
Rujing Yue, Shaoyi Du |
Neurocomputing | 3 |
| 2022 | A robust registration algorithm based on salient object detection
Runzhao Yao, Shaoyi Du, Teng Wan, Wenting Cui |
Multim. Tools Appl. | 2 |
| 2022 | Region-aware network: Model human's Top-Down visual perception mechanism for crowd counting
Yuehai Chen, Jing Yang 0014, Dong Zhang 0009, Badong Chen, Shaoyi Du |
Neural Networks | 6 |
| 2022 | Hypergraph Learning: Methods and PracticesabstractHypergraph learning is a technique for conducting learning on a hypergraph structure. In recent years, hypergraph learning has attracted increasing attention due to its flexibility and capability in modeling complex data correlation. In this paper, we first systematically review existing literature regarding hypergraph generation, including distance-based, representation-based, attribute-based, and network-based approaches. Then, we introduce the existing learning methods on a hypergraph, including transductive hypergraph learning, inductive hypergraph learning, hypergraph structure updating, and multi-modal hypergraph learning. After that, we present a tensor-based dynamic hypergraph representation and learning framework that can effectively describe high-order correlation in a hypergraph. To study the effectiveness and efficiency of hypergraph generation and learning methods, we conduct comprehensive evaluations on several typical applications, including object and action recognition, Microblog sentiment prediction, and clustering. In addition, we contribute a hypergraph learning development toolkit called THU-HyperG. Yue Gao 0002, Zizhao Zhang 0003, Haojie Lin, Xibin Zhao, Shaoyi Du, Changqing Zou |
IEEE Trans. Pattern Anal. Mach. Intell. | 5 |
| 2022 | Conditional Uncorrelation and Efficient Subset Selection in Sparse RegressionabstractGiven$m~d$-dimensional responsors and$n~d$-dimensional predictors, sparse regression finds at most$k$predictors for each responsor for linear approximation,$1\leq k \leq d-1$. The key problem in sparse regression is subset selection, which usually suffers from high computational cost. In recent years, many improved approximate methods of subset selection have been published. However, less attention has been paid to the nonapproximate method of subset selection, which is very necessary for many questions in data analysis. Here, we consider sparse regression from the view of correlation and propose the formula of conditional uncorrelation. Then, an efficient nonapproximate method of subset selection is proposed in which we do not need to calculate any coefficients in the regression equation for candidate predictors. By the proposed method, the computational complexity is reduced from$O([{1}/{6}]{k^{3}}\!+(m+1)k^{2}\!+\!mkd)$to$O([{1}/{6}]{k^{3}}\!+[{1}/{2}](m+1)k^{2})$for each candidate subset in sparse regression. Because the dimension$d$is generally the number of observations or experiments and large enough, the proposed method can greatly improve the efficiency of nonapproximate subset selection. We also apply the proposed method in real scenarios of dental age assessment and sparse coding to validate the efficiency of the proposed method. Jianji Wang 0001, Qi Liu 0010, Shaoyi Du, Yu-Cheng Guo, Nanning Zheng 0001, Fei-Yue Wang 0001 |
IEEE Trans. Cybern. | 4 |
| 2022 | Rotation-Invariant Point Cloud Representation for 3-D Model RecognitionabstractThree-dimensional (3-D) data have many applications in the field of computer vision and a point cloud is one of the most popular modalities. Therefore, how to establish a good representation for a point cloud is a core issue in computer vision, especially for 3-D object recognition tasks. Existing approaches mainly focus on the invariance of representation under the group of permutations. However, for point cloud data, it should also be rotation invariant. To address such invariance, in this article, we introduce a relation of equivalence under the action of rotation group, through which the representation of point cloud is located in a homogeneous space. That is, two point clouds are regarded as equivalent when they are only different from a rotation. Our network is flexibly incorporated into existing frameworks for point clouds, which guarantees the proposed approach to be rotation invariant. Besides, a sufficient analysis on how to parameterize the group SO(3) into a convolutional network, which captures a relation with all rotations in 3-D Euclidean space [Formula: see text]. We select the optimal rotation as the best representation of point cloud and propose a solution for minimizing the problem on the rotation group SO(3) by using its geometric structure. To validate the rotation invariance, we combine it with two existing deep models and evaluate them on ModelNet40 dataset and its subset ModelNet10. Experimental results indicate that the proposed strategy improves the performance of those existing deep models when the data involve arbitrary rotations. Yan Wang 0076, Shihui Ying, Shaoyi Du, Yue Gao 0002 |
IEEE Trans. Cybern. | 4 |
| 2022 | RGB-D Point Cloud Registration Based on Salient Object DetectionabstractWe propose a robust algorithm for aligning rigid, noisy, and partially overlapping red green blue-depth (RGB-D) point clouds. To address the problems of data degradation and uneven distribution, we offer three strategies to increase the robustness of the iterative closest point (ICP) algorithm. First, we introduce a salient object detection (SOD) method to extract a set of points with significant structural variation in the foreground, which can avoid the unbalanced proportion of foreground and background point sets leading to the local registration. Second, registration algorithms that rely only on structural information for alignment cannot establish the correct correspondences when faced with the point set with no significant change in structure. Therefore, a bidirectional color distance (BCD) is designed to build precise correspondence with bidirectional search and color guidance. Third, the maximum correntropy criterion (MCC) and trimmed strategy are introduced into our algorithm to handle with noise and outliers. We experimentally validate that our algorithm is more robust than previous algorithms on simulated and real-world scene data in most scenarios and achieve a satisfying 3-D reconstruction of indoor scenes. Teng Wan, Shaoyi Du, Wenting Cui, Runzhao Yao, Yuyan Ge, Ce Li 0001, Yue Gao 0002, Nanning Zheng 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 2 |
| 2021 | TAG-Reg: Iterative Accurate Global Registration AlgorithmabstractIn this paper, we propose an accurate global registration (TAG-Reg) algorithm for poor initialization and partially overlapping point clouds registration problem. Firstly, methods based on geometric structure information of points can get the accurate results, which is vulnerable to poor initialization. Meanwhile, existing features based global methods can solve poor initialization problem at a certain extent, but it cannot obtain accurate results. So, we combine the geometric structure information with feature as hybrid feature to solve poor initialization problem completely and obtain accurate results. Secondly, we introduce dynamic trimmed strategy combining with hybrid feature to deal with partially overlapping problem. Then, to improve the accuracy of our method, we utilize the probabilistic method to suppress noise. At last, we establish the TAG-Reg model and propose an iterative algorithm to solve this problem. Experimental results show that our TAG-Reg achieves state-of-the-art performance compared to existing non-deep learning and recent deep learning methods. Our source code will open at https://github.com/BiaoBiaoLi/TAG-Reg. Qixing Xie, Shaoyi Du, Wenting Cui, Runzhao Yao, Yue Gao 0002, Nanning Zheng 0001 |
ICME | 3 |
| 2021 | DWG-Reg: Deep Weight Global RegistrationabstractIn this paper, we propose a deep weight global registration (DWG-Reg) algorithm for poor initialization and partially overlapping point clouds registration problem. Our DWG-Reg is based on three modules: a bidirectional nearest search strategy for correspondence, a convolutional network for correspondence confidence prediction which consists of Hybird Distance Generator, optimal annealing Parameter Prediction network and a robust kernel function, a weighted optimizer algorithm for closed-form pose estimation. Experimental results show that our DWG-Reg achieves state-of-the-art performance compared to existing non-deep learning and recent deep learning methods. Our source code will open at https://github.com/BiaoBiaoLi/DWG-Reg. Qixing Xie, Shaoyi Du, Wenting Cui, Runzhao Yao, Yang Yang 0066, Jing Yang 0014, Lin Wang 0026 |
IJCNN | 3 |
| 2021 | Precise Point Set Registration Based on Feature FusionabstractAbstract This paper proposed a novel precise point set registration method based on feature fusion for three-dimensional data. Firstly, for the prominent foreground with dense and continuous cluster structure, we propose an automatic extraction method combining the principal component analysis projection and density-based clustering method. Secondly, for point sets containing noises, we introduce correntropy measurement into registration to weaken their influence. Thirdly, for the precise registration of uneven distribution of points in the same point set, we propose a feature fusion based algorithm which is distribution specific, using point-to-point measurement for densely distributed foreground and point-to-plane measurement for sparsely distributed background, in case that only one measurement method is used for the whole point set the registration gets trapped into local extremum. Finally, we give the optimization algorithm of the proposed method. We conduct experiments on real orthodontics scenes to verify the effectiveness of our proposed feature extraction method and registration algorithm, and experimental results demonstrate that both the proposed solutions are proper for their respective tasks than other existing methods. Yuying Liu 0007, Shaoyi Du, Wenting Cui, Xijing Wang, Qingnan Mou, Jiamin Zhao, Yucheng Guo |
Comput. J. | 2 |
| 2021 | A generic FPGA-based hardware architecture for recursive least mean p-power extreme learning machine
Jing Yang 0014, Hai-Jun Rong, Shaoyi Du |
Neurocomputing | 4 |
| 2021 | Interactive prostate MR image segmentation based on ConvLSTMs and GGNN
Yaoyue Zheng, Hongcheng Fan, Zhongyu Li 0002, Ce Li 0001, Shaoyi Du |
Neurocomputing | 8 |
| 2021 | Robust registration algorithm based on rational quadratic kernel for point sets with outliers and noise
Runzhao Yao, Shaoyi Du, Teng Wan, Wenting Cui, Yang Yang 0066, Yang Jing, Ce Li 0001 |
Multim. Tools Appl. | 2 |
| 2021 | Robust High-Order Manifold Constrained Low Rank Representation for Subspace ClusteringabstractDue to the effectiveness in learning the subspace structures, low-rank representation (LRR) and its variations have been widely applied in various fields, such as computer vision and pattern recognition. However, in real applications, it is a challenge to handle the complex noises. To address this problem, we propose a novel robust LRR method based on kernel risk-sensitive loss (KRSL) with high-order manifold constraint, called RHLRR, in which the KRSL is introduced to deal with the noises and the multiple hypergraph regularization term is used as a high order manifold constraint to effectively capture the locality, similarity and the intrinsic geometric information in data. Besides, an iterative algorithm based on the half-quadratic (HQ) and the accelerated block coordinate update (BCU) is developed. The experimental results demonstrate that the proposed method can outperform other state-of-the-art LRR variants. Lei Xing 0003, Badong Chen, Jianji Wang 0001, Shaoyi Du, Jiuwen Cao |
IEEE Trans. Circuits Syst. Video Technol. | 4 |
| 2021 | Correntropy-Based Multiview Subspace ClusteringabstractMultiview subspace clustering, which aims to cluster the given data points with information from multiple sources or features into their underlying subspaces, has a wide range of applications in the communities of data mining and pattern recognition. Compared with the single-view subspace clustering, it is challenging to efficiently learn the structure of the representation matrix from each view and make use of the extra information embedded in multiple views. To address the two problems, a novel correntropy-based multiview subspace clustering (CMVSC) method is proposed in this article. The objective function of our model mainly includes two parts. The first part utilizes the Frobenius norm to efficiently estimate the dense connections between the points lying in the same subspace instead of following the standard compressive sensing approach. In the second part, the correntropy-induced metric (CIM) is introduced to characterize the noise in each view and utilize the information embedded in different views from an information-theoretic perspective. Furthermore, an efficient iterative algorithm based on the half-quadratic technique (HQ) and the alternating direction method of multipliers (ADMM) is developed to optimize the proposed joint learning problem, and extensive experimental results on six real-world multiview benchmarks demonstrate that the proposed methods can outperform several state-of-the-art multiview subspace clustering methods. Lei Xing 0003, Badong Chen, Shaoyi Du, Yuantao Gu, Nanning Zheng 0001 |
IEEE Trans. Cybern. | 3 |
| 2021 | Point Set Registration With Similarity and Affine Transformations Based on Bidirectional KMPE LossabstractRobust point set registration is a challenging problem, especially in the cases of noise, outliers, and partial overlapping. Previous methods generally formulate their objective functions based on the mean-square error (MSE) loss and, hence, are only able to register point sets under predefined constraints (e.g., with Gaussian noise). This article proposes a novel objective function based on a bidirectional kernel mean p -power error (KMPE) loss, to jointly deal with the above nonideal situations. KMPE is a nonsecond-order similarity measure in kernel space and shows a strong robustness against various noise and outliers. Moreover, a bidirectional measure is applied to judge the registration, which can avoid the ill-posed problem when a lot of points converges to the same point. In particular, we develop two effective optimization methods to deal with the point set registrations with the similarity and the affine transformations, respectively. The experimental results demonstrate the effectiveness of our methods. Yang Yang 0066, Shaoyi Du, Muyi Wang, Badong Chen, Yue Gao 0002 |
IEEE Trans. Cybern. | 3 |
| 2021 | Effects of Outliers on the Maximum Correntropy Estimation: A Robustness AnalysisabstractRecently, maximum correntropy criterion (MCC) has been widely and successfully used in robust signal processing and machine learning, in which the correntropy is maximized instead of minimizing the popular mean square error (MSE) to improve the robustness with respect to outliers or impulsive noises. A lot of efforts have been devoted to derive different adaptive algorithms under MCC, but to date, little insight has been gained as to how the MCC solution will be influenced by outliers. In this paper, we investigate this problem and our focus is mainly on the parameter estimation of a simple linear errors-in-variables (EIVs) model with scalar variables. Under some conditions, we derive an upper bound on the absolute value of the estimation error and show that the MCC solution can get very close to the true value of the unknown parameter even with arbitrarily large outliers in both the input and output variables. Illustrative examples are provided to verify and clarify the theory. Badong Chen, Lei Xing 0003, Haiquan Zhao 0001, Shaoyi Du, José C. Príncipe |
IEEE Trans. Syst. Man Cybern. Syst. | 4 |
| 2020 | CF-LSTM: Cascaded Feature-Based Long Short-Term Networks for Predicting Pedestrian TrajectoryabstractPedestrian trajectory prediction is an important but difficult task in self-driving or autonomous mobile robot field because there are complex unpredictable human-human interactions in crowded scenarios. There have been a large number of studies that attempt to understand humans' social behavior. However, most of these studies extract location features from previous one time step while neglecting the vital velocity features. In order to address this issue, we propose a novel feature-cascaded framework for long short-term network (CF-LSTM) without extra artificial settings or social rules. In this framework, feature information from previous two time steps are firstly extracted and then integrated as a cascaded feature to LSTM, which is able to capture the previous location information and dynamic velocity information, simultaneously. In addition, this scene-agnostic cascaded feature is the external manifestation of complex human-human interactions, which can also effectively capture dynamic interaction information in different scenes without any other pedestrians' information. Experiments on public benchmark datasets indicate that our model achieves better performance than the state-of-the-art methods and this feature-cascaded framework has the ability to implicitly learn human-human interactions. Yi Xu 0005, Jing Yang 0014, Shaoyi Du |
AAAI | 3 |
| 2020 | CoBigICP: Robust and Precise Point Set Registration using Correntropy Metrics and Bidirectional CorrespondenceabstractIn this paper, we propose a novel probabilistic variant of iterative closest point (ICP) dubbed as CoBigICP. The method leverages both local geometrical information and global noise characteristics. Locally, the 3D structure of both target and source clouds are incorporated into the objective function through bidirectional correspondence. Globally, error metric of correntropy is introduced as noise model to resist outliers. Importantly, the close resemblance between normal-distributions transform (NDT) and correntropy is revealed. To ease the minimization step, an on-manifold parameterization of the special Euclidean group is proposed. Extensive experiments validate that CoBigICP outperforms several well-known and state-of-the-art methods. Pengyu Yin, Di Wang 0028, Shaoyi Du, Shihui Ying, Yue Gao 0002, Nanning Zheng 0001 |
IROS | 3 |
| 2020 | 3-D Oral Shape Retrieval Using Registration Algorithm
Wenting Cui, Shaoyi Du, Teng Wan, Yuying Liu 0007, Yang Yang 0066, Qingnan Mou, Mengqi Han, Yu-Cheng Guo |
MMM (2) | 2 |
| 2020 | Robust RGB-D Data Registration Based on Correntropy and Bi-directional Distance
Teng Wan, Shaoyi Du, Wenting Cui, Qixing Xie, Yuying Liu 0007 |
MMM (2) | 2 |
| 2020 | Pamls Alignment Based On Two-Stage Convolutional Network with a Large in-Plane RotationabstractPalms alignment is an important work for palmprint recognition in uncontrolled environment. Many methods have made progress to achieve alignment. But most of them ignore the palm's angles, which could not satisfy the alignment initialization when the hand has a large in-plane rotation. In this paper, we propose a palms alignment with affine transformation method based on a two-stage convolutional neural network (CNN). The basic idea is to rotate the target palm into the same angle category to avoid the following affine registration has a big matching error at the beginning. At the stage I, the given target palm is classified into two angle categories. At the stage II the upside down palm is firstly rotated 180 degrees, and then inputted into the subsequent feature extraction network, feature matching layer and regression network to achieve the affine alignment. Experimental results have proved the effectiveness of our method. Yang Yang 0066, Guobin Zhang, Wenting Cui, Shaoyi Du |
SMC | 6 |
| 2020 | Robust Point Set Registration Based on Semantic InformationabstractPoint cloud registration a challenging task in situations with poor initial value and scenarios with limited geometric structure. In these cases, the correct correspondence between two point clouds is unknown and difficult to establish. To cope with this problem, the semantic of partial points is introduced in this paper. Firstly, the semantic information is used to find more reasonable correspondence, i.e. semantic point pairs. Secondly, we formulate a novel objective function to integrate the matching error of semantic point pairs as guidance of registration. Thirdly, a hyperparameter is applied to balance the confidence of semantic point pairs. At last, a novel algorithm under the ICP framework is presented to optimize the rigid transformation iteratively. The evaluation of KITTI data set reveals the robustness and accuracy of our method in the complex scenes mentioned above. Qinlong Wang, Yang Yang 0066, Teng Wan, Shaoyi Du |
SMC | 4 |
| 2020 | Individual retrieval based on oral cavity point cloud data and correntropy-based registration algorithmabstractIn this study, the authors present a novel individual retrieval method based on oral cavity point cloud data and correntropy‐based registration algorithm. Since the three‐dimensional oral cavity data contains a large amount of noise and outliers, it may lead to a decrease in registration accuracy, which affects the accuracy of retrieval rate. Therefore, the authors introduce the correntropy into the rigid registration algorithm to solve this problem. Then, they filter the matched point cloud data and then use the mean squared error to judge the individual differences of the model data. Finally, the accurate retrieval of the oral cavity data is realised. Experimental results demonstrate the proposed retrieval three‐dimensional model algorithm can be successfully searched under different model data, which can help forensics use the characteristics of biological individuals to accurately search and identify, and improve recognition efficiency. Wenting Cui, Shaoyi Du, Yuying Liu 0007, Teng Wan, Mengqi Han, Qingnan Mou, Jing Yang 0014, Yu-Cheng Guo |
IET Image Process. | 3 |
| 2020 | Enhanced image no-reference quality assessment based on colour space distributionabstractIn this study, the authors investigate the problem of enhanced image no‐reference (NR) quality assessment. For resolving the problem of the enhanced images, it is difficult to obtain reference images, this study proposes an NR image quality assessment (IQA) model based on colour space distribution. Given an enhanced image, our method first uses a gist to select a clear target image in which the scene, colour and quality are similar to the hypothetical reference images. And then, the colour transfer is used between the input images and target images to construct the reference image. Next, the appropriate IQA method is used to assess enhanced image quality. The absolute colour difference and feature similarity (FSIM) are used to measure the colour and grey‐scale image quality, respectively. Extensive experiments demonstrate that the proposed method is good at evaluating enhanced image quality for X‐ray, dust, underwater and low‐light images. The experimental results are consistent with human subjective evaluation and achieve good assessment effects. Hao Liu 0060, Ce Li 0001, Dong Zhang 0009, Yannan Zhou, Shaoyi Du |
IET Image Process. | 5 |
| 2020 | Adaptive weighted motion averaging with low-rank sparse for robust multi-view registration
Zhongyu Li 0002, Jihua Zhu, Ce Li 0001, Shaoyi Du |
Neurocomputing | 6 |
| 2020 | A review of object detection based on deep learning
Youzi Xiao, Jiachen Yu, Yinshu Zhang, Shaoyi Du, Xuguang Lan |
Multim. Tools Appl. | 6 |
| 2020 | Robust and precise isotropic scaling registration algorithm using bi-directional distance and correntropy
Wenting Cui, Shaoyi Du, Teng Wan, Runzhao Yao, Yuying Liu 0007, Mengqi Han, Qingnan Mou, Yu-Cheng Guo, Nanning Zheng 0001 |
Pattern Recognit. Lett. | 2 |
| 2020 | Robust rigid registration algorithm based on pointwise correspondence and correntropy
Shaoyi Du, Guanglin Xu, Sirui Zhang, Xuetao Zhang 0001, Yue Gao 0002, Badong Chen |
Pattern Recognit. Lett. | 1 |
| 2020 | Predicting COVID-19 in China Using Hybrid AI ModelabstractThe coronavirus disease 2019 (COVID-19) breaking out in late December 2019 is gradually being controlled in China, but it is still spreading rapidly in many other countries and regions worldwide. It is urgent to conduct prediction research on the development and spread of the epidemic. In this article, a hybrid artificial-intelligence (AI) model is proposed for COVID-19 prediction. First, as traditional epidemic models treat all individuals with coronavirus as having the same infection rate, an improved susceptible-infected (ISI) model is proposed to estimate the variety of the infection rates for analyzing the transmission laws and development trend. Second, considering the effects of prevention and control measures and the increase of the public's prevention awareness, the natural language processing (NLP) module and the long short-term memory (LSTM) network are embedded into the ISI model to build the hybrid AI model for COVID-19 prediction. The experimental results on the epidemic data of several typical provinces and cities in China show that individuals with coronavirus have a higher infection rate within the third to eighth days after they were infected, which is more in line with the actual transmission laws of the epidemic. Moreover, compared with the traditional epidemic models, the proposed hybrid AI model can significantly reduce the errors of the prediction results and obtain the mean absolute percentage errors (MAPEs) with 0.52%, 0.38%, 0.05%, and 0.86% for the next six days in Wuhan, Beijing, Shanghai, and countrywide, respectively. Nanning Zheng 0001, Shaoyi Du, Jianji Wang 0001, Wenting Cui, Zijian Kang, Tao Yang 0032, Bin Lou, Yuting Chi, Hong Long, Mei Ma, Dong Zhang 0009, Jingmin Xin |
IEEE Trans. Cybern. | 2 |
| 2020 | Hierarchical U-Shape Attention Network for Salient Object DetectionabstractSalient object detection aims at locating the most conspicuous objects in natural images, which usually acts as a very important pre-processing procedure in many computer vision tasks. In this paper, we propose a simple yet effective Hierarchical U-shape Attention Network (HUAN) to learn a robust mapping function for salient object detection. Firstly, a novel attention mechanism is formulated to improve the well-known U-shape network [1], in which the memory consumption can be extensively reduced and the mask quality can be significantly improved by the resulting U-shape Attention Network (UAN). Secondly, a novel hierarchical structure is constructed to well bridge the low-level and high-level feature representations between different UANs, in which both the intra-network and inter-network connections are considered to explore the salient patterns from a local to global view. Thirdly, a novel Mask Fusion Network (MFN) is designed to fuse the intermediate prediction results, so as to generate a salient mask which is in higher-quality than any of those inputs. Our HUAN can be trained together with any backbone network in an end-to-end manner, and high-quality masks can be finally learned to represent the salient objects. Extensive experimental results on several benchmark datasets show that our method significantly outperforms most of the state-of-the-art approaches. Sanping Zhou, Jinjun Wang, Jimuyang Zhang, Le Wang 0003, Shaoyi Du, Nanning Zheng 0001 |
IEEE Trans. Image Process. | 6 |
| 2019 | A Multi-model Ensemble Method Using CNN and Maximum Correntropy Criterion for Basal Cell Carcinoma and Seborrheic Keratoses ClassificationabstractBasal cell carcinoma is very similar to the clinical traits of seborrheic keratosis, which is still a difficult problem in medical image analysis. To accurately classify it, this paper proposes a multi-model ensemble method based on the maximum correntropy criterion (MCC) and convolutional neural network (CNN). First of all, it is well known that the CNN single models like ResNet, Xception, DensNet, etc. have a good effect on the classification, but the accuracy is still limited, so the multi-model ensemble method is presented to improve the accuracy. Secondly, the traditional multi-model ensemble methods, such as voting and linear regression, can improve the accuracy of the model, but it means that the weight computation of each model does not consider the noise, and could not obtain good results. Therefore, we propose the MCC for the model ensemble, which overcomes the noise in the data and effectively improves the classification accuracy. Finally, our proposed multi-model ensemble algorithm based on the MCC achieved an accuracy of 97.07% in the basal cell carcinoma and seborrheic keratosis classification experiments, surpassing the CNN single model and traditional multi-model ensemble method. Leida Guo, Shaoyi Du, Yuting Chi, Wenting Cui, Panpan Song, Jihua Zhu, Songmei Geng, Meifeng Xu |
IJCNN | 2 |
| 2019 | Precise iterative closest point algorithm with corner point constraint for isotropic scaling registration
Shaoyi Du, Wenting Cui, Liyang Wu, Sirui Zhang, Xuetao Zhang 0001, Guanglin Xu, Meifeng Xu |
Multim. Syst. | 1 |
| 2019 | RGB-D point cloud registration via infrared and color camera
Teng Wan, Shaoyi Du, Yiting Xu, Guanglin Xu, Badong Chen, Yue Gao 0002 |
Multim. Tools Appl. | 2 |
| 2019 | Correntropy based scale ICP algorithm for robust point set registration
Zongze Wu 0001, Hongchen Chen, Shaoyi Du, Minyue Fu 0001, Nanning Zheng 0001 |
Pattern Recognit. | 3 |
| 2019 | Generative adversarial dehaze mapping nets
Ce Li 0001, Zhaoxiang Zhang 0001, Shaoyi Du |
Pattern Recognit. Lett. | 4 |
| 2019 | Granger Causality Analysis Based on Quantized Minimum Error Entropy CriterionabstractLinear regression model (LRM) based on mean square error (MSE) criterion is widely used in Granger causality analysis (GCA), which is the most commonly used method to detect the causality between a pair of time series. However, when signals are seriously contaminated by non-Gaussian noises, the LRM coefficients will be inaccurately identified. This may cause the GCA to detect a wrong causal relationship. Minimum error entropy (MEE) criterion can be used to replace the MSE criterion to deal with the non-Gaussian noises. But its calculation requires a double summation operation, which brings computational bottlenecks to GCA especially when sizes of the signals are large. To address the aforementioned problems, in this letter, we propose a new method called GCA based on the quantized MEE (QMEE) criterion (GCA-QMEE), in which the QMEE criterion is applied to identify the LRM coefficients and the quantized error entropy is used to calculate the causality indexes. Compared with the traditional GCA, the proposed GCA-QMEE not only makes the results more discriminative, but also more robust. Its computational complexity is also not high because of the quantization operation. Illustrative examples on synthetic and EEG datasets are provided to verify the desirable performance and the availability of the GCA-QMEE. Badong Chen, Rongjin Ma, Si-yu Yu, Shaoyi Du, Harry Qin |
IEEE Signal Process. Lett. | 4 |
| 2018 | Accurate Mix-Norm-Based Scan MatchingabstractHighly accurate mapping and localization is of prime importance for mobile robotics, and its core lies in efficient scan matching. Previous research are focusing on designing a robust objective function and the residual error distribution is often ignored or simply assumed as unitary or mixture of simple distributions. In this paper, a mixture of exponential power (MoEP) distributions is proposed to approximate the residual error distribution. The objective function induced by MoEP-based residual error modelling ensembles a mix-norm-based scan matching (MiNoM), which enhances the matching accuracy and convergence characteristic. Both the parameters of transformation (rotation and translation) and residual error distribution are estimated efficiently via an EM-like algorithm. The optimization of MiNoM is iteratively achieved via two phases: An on-line parameter learning (OPL) phase to learn residual error distribution for better representation according to the likelihood field model (LFM), and an iteratively reweighted least squares (IRLS) phase to attain transformation for accuracy and efficiency. Extensive experimental results validate that the proposed MiNoM out-performs several state-of-the-art scan matching algorithms in both convergence characteristic and matching accuracy. Di Wang 0028, Jianru Xue, Zhongxing Tao, Dixiao Cui, Shaoyi Du, Nanning Zheng 0001 |
IROS | 6 |
| 2018 | Accurate Localization in Underground Garages via Cylinder Feature based Map MatchingabstractAutonomous driving in underground garages usually utilizes a 2D/3D occupancy map for localization. However, the real scene is changing, and may not be consistent with the map. Vehicles and other objects not contained in the map are considered as obstacles, which increase the difficulty of localization and affect the accuracy of result. In this paper, we propose a cylinder rotational projection statistics (Cy-RoPS) feature descriptor, which is a local surface feature descriptor to improve the accuracy of localization. The local surface feature motivated by RoPS feature is invariant to rotation of point set enclosed in a cylinder. We also propose to employ the local surface feature for localization in a real underground garage. The experimental results show that the proposed method is robust to dynamic obstacles in the underground garage, and has a higher accuracy in localization, compared with the state-of-the-art methods. Zhongxing Tao, Jianru Xue, Di Wang 0028, Dixiao Cui, Shaoyi Du |
Intelligent Vehicles Symposium | 6 |
| 2018 | Precise Point Set Registration Using Point-to-Plane Distance and Correntropy for LiDAR Based LocalizationabstractIn this paper, we propose a robust point set registration algorithm which combines correntropy and point-to-plane distance, which can register rigid point sets with noises and outliers. Firstly, as correntropy performs well in handling data with non-Gaussian noises, we introduce it to model rigid point set registration problem based on point-to-plane distance; Secondly, we propose an iterative algorithm to solve this problem, which repeats to compute correspondence and transformation parameters respectively in closed form solutions. Simulated experimental results demonstrate the high precision and robustness of the proposed algorithm. In addition, LiDAR based localization experiments on automated vehicle performs satisfactory for localization accuracy and time consumption. Guanglin Xu, Shaoyi Du, Dixiao Cui, Sirui Zhang, Badong Chen, Xuetao Zhang 0001, Jianru Xue, Yue Gao 0002 |
Intelligent Vehicles Symposium | 2 |
| 2018 | Precise Point Set Registration with Color Assisted and Correntropy for 3D ReconstructionabstractIterative closest point (ICP) algorithm, as its accuracy and efficiency, is widely used in rigid registration. However, ICP algorithm is easily failed when point sets lack of structure variety, such as semicircles. To solve this problem, a precise point set registration method for RGB-D data is proposed. Firstly, the color information provides a new information for registration, and the correntropy is introduced to deal with the noises and outliers. With color assisted and correntropy, a more robust objective function is built. Secondly, a variant ICP algorithm is used to deal with optimization problem via multiple iterations. Finally, as shown in the experimental results and scene reconstruction, our method obtains more precise results than other ICP algorithms. Teng Wan, Shaoyi Du, Yiting Xu, Guanglin Xu, Yang Yang 0066, Yue Gao 0002, Badong Chen |
SMC | 2 |
| 2018 | Building Correspondence Based on Matching Triangles for Partial RegistrationabstractAs an important problem in point set registration, partial registration has been solved by some variants of Iterative Closest Point (ICP) algorithm under good initial values. However, the initial parameters remained to be solved for partial registration. This paper presents a parameter initialization algorithm based on matching triangles for partial registration. Experimental results demonstrate that the proposed initialization method can find an appropriate initial transformation for next accurate registration, even the initial rotation angle between two sets is large. Based on the initialization of two point sets, the partial registration can be accomplished by auto trimmed ICP (ATICP) algorithm. Yiting Xu, Shaoyi Du, Teng Wan, Yang Yang 0066, Badong Chen, Yue Gao 0002 |
SMC | 2 |
| 2018 | Nonlinear Semi-Supervised Metric Learning Via Multiple Kernels and Local TopologyabstractChanging the metric on the data may change the data distribution, hence a good distance metric can promote the performance of learning algorithm. In this paper, we address the semi-supervised distance metric learning (ML) problem to obtain the best nonlinear metric for the data. First, we describe the nonlinear metric by the multiple kernel representation. By this approach, we project the data into a high dimensional space, where the data can be well represented by linear ML. Then, we reformulate the linear ML by a minimization problem on the positive definite matrix group. Finally, we develop a two-step algorithm for solving this model and design an intrinsic steepest descent algorithm to learn the positive definite metric matrix. Experimental results validate that our proposed method is effective and outperforms several state-of-the-art ML methods. Yanqin Bai, Yaxin Peng, Shaoyi Du, Shihui Ying |
Int. J. Neural Syst. | 4 |
| 2018 | Beyond Pairwise Matching: Person Reidentification via High-Order Relevance LearningabstractPerson reidentification has attracted extensive research efforts in recent years. It is challenging due to the varied visual appearance from illumination, view angle, background, and possible occlusions, leading to the difficulties when measuring the relevance, i.e., similarities, between probe and gallery images. Existing methods mainly focus on pairwise distance metric learning for person reidentification. In practice, pairwise image matching may limit the data for comparison (just the probe and one gallery subject) and yet lead to suboptimal results. The correlation among gallery data can be also helpful for the person reidentification task. In this paper, we propose to investigate the high-order correlation among the probe and gallery data, not the pairwise matching, to jointly learn the relevance of gallery data to the probe. Recalling recent progresses on feature representation in person reidentification, it is difficult to select the best feature and each type of feature can benefit person description from different aspects. Under such circumstances, we propose a multihypergraph joint learning algorithm to learn the relevance in corporation with multiple features of the imaging data. More specifically, one hypergraph is constructed using one type of feature and multiple hypergraphs can be generated accordingly. Then, the learning process is conducted on the multihypergraph structure, and the identity of a probe is determined by its relevance to each gallery data. The merit of the proposed scheme is twofold. First, different from pairwise image matching, the proposed method jointly explores the relationships among different images. Second, multimodal data, i.e., different features, can be formulated in the multihypergraph structure, which can convey more information in the learning process and can be easily extended. We note that the proposed method is a general framework to incorporate with any combination of features, and thus is flexible in practice. Experimental results and comparisons with the state-of-the-art methods on three public benchmarking data sets demonstrate the superiority of the proposed method. Xibin Zhao, Nan Wang 0015, Yubo Zhang 0006, Shaoyi Du, Yue Gao 0002, Jia-Guang Sun 0001 |
IEEE Trans. Neural Networks Learn. Syst. | 4 |
| 2017 | An Iterative Feature-Pair Updating Framework for Rigid Template Matching with OutliersabstractTo deal with the rigid template matching problem in real-world scenarios, we propose a novel iterative feature-pair updating framework which is also robust to high levels of outliers, such as background changing, complex nonrigid deformation and partial occlusion. Given a pair of template image and target image, we first extract a set of corresponding feature-pairs as candidates. Then, we propose a robust objective function under the iterative framework for discriminatively updating these candidates, where the space distance, appearance distance, and the overlapping percentage of feature pairs are integrated simultaneously. Finally, a hierarchical matching strategy is provided with the parameter discussion. Experimental results compared with the-state-of-art methods on public data sets demonstrate the effectiveness of the proposed method. Yang Yang 0066, Qian Kou, Shaoyi Du, Yuehu Liu, Bangyu Wu |
ISM | 3 |
| 2017 | Precise isotropic scaling iterative closest point algorithm based on corner points for shape registrationabstractThe traditional iterative closest point (ICP) algorithm could register two points sets well, but it is easily affected by local dissimilar. To deal with this problem, this paper proposes an isotropic scaling ICP algorithm with corner point constraint. First, an objective function is proposed under the guidance of the corner points, as the corner points can preserve the similar of the whole shapes. Secondly, a new ICP algorithm is used to complete the isotropic scaling registration. At each step of this new algorithm, the correspondence is built based on the closest point searching, and then a closed-form solution of the transformation is computed. The experimental results demonstrate that our algorithm can prevent the influence of the local dissimilar and improve the registration precision compared with the traditional ICP algorithm. Shaoyi Du, Wenting Cui, Xuetao Zhang 0001, Liyang Wu |
SMC | 1 |
| 2017 | Robust affine registration based on corner point guided ICP algorithmabstractThe traditional affine iterative closest point (ICP) algorithm is fast and accurate for affine registration between two point sets, but it is easy to fall into local minimum. This paper proposes a robust Affine ICP algorithm based on corner points. First, an objective function is established under the guidance of corner points, where the corner points as the shape control point guides the affine registration of 2D point sets. Then, at each step of the algorithm, the affine transformation obtained by the last iterative step is used to establish the correspondence of these two point sets. Next, the new affine transformation is solved by using the objective function under the guidance of the corner points. Experimental results demonstrate that the robustness and convergence of our algorithm are greatly improved compared with the traditional affine registration algorithm. Liyang Wu, Duyan Bi, Shaoyi Du, Wenting Cui |
SMC | 5 |
| 2017 | Robust 2D point set matching with Kernel mean P-power error lossabstractIn this paper, we propose a novel point set matching algorithm to improve the matching precision in the presence of non-Gaussian noises and outliers. In our method, a non-second order similarity measure known as Kernel Mean p-Power Error (KMPE) loss is employed as the matching cost function. We introduce a local optimal solution for computing the rigid transform by repeating the correspondence estimation and parameter updating processes. This new algorithm assigns a non-linear distance evaluation in kernel space according to the current estimation of the correspondence to yield a more accurate matching result between two point sets in practice. Experimental results demonstrate that our algorithm is more robust and accurate than the traditional ICP and the state-of-the-art algorithms. Yang Yang 0066, Weile Chen, Badong Chen, Shaoyi Du |
SMC | 4 |
| 2017 | Improvement of affine iterative closest point algorithm for partial registrationabstractIn this study, partial registration problem with outliers and missing data in the affine case is discussed. To solve this problem, a novel objective function is proposed based on bidirectional distance and trimmed strategy, and then a new affine trimmed iterative closest point algorithm is given. First, when bidirectional distance measurement is applied, the ill‐posed partial registration problem in the affine case is prevented. Second, the overlapping percentage is solved by using trimmed strategy which uses as many correct overlapping points as possible. The authors’ method computes the affine transformation, correspondence and overlapping percentage automatically at each iterative step. In this way, it handles partially overlapping registration with outliers and missing data in the affine case well. Experimental results demonstrate that their method is more robust and precise than the state‐of‐the‐art algorithms. It also has good convergence and similar running time with traditional algorithms. Zhongmin Cai, Shaoyi Du |
IET Comput. Vis. | 3 |
| 2017 | A vision-centered multi-sensor fusing approach to self-localization and obstacle perception for robotic carsabstractMost state-of-the-art robotic cars’ perception systems are quite different from the way a human driver understands traffic environments. First, humans assimilate information from the traffic scene mainly through visual perception, while the machine perception of traffic environments needs to fuse information from several different kinds of sensors to meet safety-critical requirements. Second, a robotic car requires nearly 100% correct perception results for its autonomous driving, while an experienced human driver works well with dynamic traffic environments, in which machine perception could easily produce noisy perception results. In this paper, we propose a vision-centered multi-sensor fusing framework for a traffic environment perception approach to autonomous driving, which fuses camera, LIDAR, and GIS information consistently via both geometrical and semantic constraints for efficient self-localization and obstacle perception. We also discuss robust machine vision algorithms that have been successfully integrated with the framework and address multiple levels of machine vision techniques, from collecting training data, efficiently processing sensor data, and extracting low-level features, to higher-level object and environment mapping. The proposed framework has been tested extensively in actual urban scenes with our self-developed robotic cars for eight years. The empirical results validate its robustness and efficiency. Jianru Xue, Di Wang 0028, Shaoyi Du, Dixiao Cui, Nanning Zheng 0001 |
Frontiers Inf. Technol. Electron. Eng. | 3 |
| 2017 | Precise glasses detection algorithm for face with in-plane rotation
Shaoyi Du, Yuehu Liu, Xuetao Zhang 0001, Jianru Xue |
Multim. Syst. | 1 |
| 2017 | Robust non-rigid point set registration via building tree dynamically
Shaoyi Du, Bo Bi, Guanglin Xu, Jihua Zhu, Xuetao Zhang 0001 |
Multim. Tools Appl. | 1 |
| 2016 | A Modified Non-rigid ICP Algorithm for Registration of Chromosome Images
Qian Kou, Yang Yang 0066, Shaoyi Du, Dongge Cai |
ICIC (2) | 3 |
| 2016 | Robust image registration with rotation, scale and translation using Best-Buddies PairsabstractSince the complex outliers caused by the background and non-rigid deformation, image registration remains a challenging task. This paper proposes a new method for image registration, which is robust to the transformations with rotation, scale and translation. Our method is under the general Iterative Closet Point (ICP) framework which contains two main parts: finding the corresponding points and updating transformation parameters. In our method, to be robust to outliers, Best-Buddies Pairs (BBPs) are used as similarity measure between two images to obtain the corresponding points. While, in order to be invariant to scale transformation, Scaling ICP is employed to update transformation parameters. Besides the image registration, we also apply the method for object tracking in video sequences. Experimental results demonstrate the effectiveness and efficiency of the proposed method. Yang Yang 0066, Shaoyi Du, Qian Kou |
IJCNN | 3 |
| 2016 | Robust affine iterative closest point algorithm based on correntropy for 2D point set registrationabstractThe traditional affine iterative closest point (ICP) algorithm is fast and accuracy for affine registration of point sets, but it performs worse when the point sets with large outliers. This paper introduces a novel algorithm based on correntropy for affine registration of point sets with outliers. First, a novel objective function is proposed by introducing the maximum correntropy criterion (MCC) because of the outlier-rejection property of correntropy. Then, a new affine ICP algorithm is proposed to solve this energy function. This method uses a simple iterative algorithm and computes the affine transformation quickly at each iterative step. Similar to the ICP algorithm, this new algorithm converges monotonically to a local maximum for any given initial parameters. Experimental results demonstrate that our algorithm has the high speed and accuracy for affine registration with outliers compared with the traditional ICP algorithm and the state-of-the-art algorithms. Zongze Wu 0001, Hongchen Chen, Shaoyi Du |
IJCNN | 3 |
| 2016 | Precise 2D point set registration using iterative closest algorithm and correntropyabstractThe iterative closest point (ICP) algorithm is fast and accurate for rigid point set registration, but it works badly when there are many outliers and noises in the point sets. This paper instead proposes a novel method based on the ICP algorithm to deal with this problem. Firstly, correntropy is introduced into the rigid registration problem and then a new energy function based on maximum correntropy criterion is proposed. After that, a new ICP algorithm based on correntropy is proposed, which performs well in dealing with rigid registration with noises and outliers. This new algorithm converges moronically from any given parameters, which is similar to the ICP algorithm. Experimental results demonstrate its accuracy and efficiency compared with the traditional ICP algorithm. Guanglin Xu, Shaoyi Du, Jianru Xue |
IJCNN | 2 |
| 2016 | Robust isotropic scaling ICP algorithm with bidirectional distance and bounded rotation angle
Shaoyi Du, Chunjia Zhang, Zongze Wu 0001, Jianru Xue |
Neurocomputing | 1 |
| 2016 | Robust iterative closest point algorithm with bounded rotation angle for 2D registration
Chunjia Zhang, Shaoyi Du, Jianru Xue, Yuehu Liu |
Neurocomputing | 2 |
| 2016 | New iterative closest point algorithm for isotropic scaling registration of point sets with noise
Shaoyi Du, Bo Bi, Jihua Zhu, Jianru Xue |
J. Vis. Commun. Image Represent. | 1 |
| 2016 | Robust 3D Point Set Registration Using Iterative Closest Point Algorithm with Bounded Rotation Angle
Chunjia Zhang, Shaoyi Du, Jianru Xue |
Signal Process. | 2 |
| 2015 | Accurate non-rigid registration based on heuristic tree for registering point sets with large deformation
Shaoyi Du, Chunjia Zhang, Meifeng Xu, Jianru Xue |
Neurocomputing | 1 |
| 2015 | Probability iterative closest point algorithm for m-D point set registration with noise
Shaoyi Du, Chunjia Zhang, Jihua Zhu |
Neurocomputing | 1 |
| 2015 | Building dynamic population graph for accurate correspondence detection
Shaoyi Du, Yanrong Guo, Gerard Sanroma, Dong Ni 0001, Guorong Wu 0001, Dinggang Shen |
Medical Image Anal. | 1 |
| 2014 | Real-time global localization of intelligent road vehicles in lane-level via lane marking detection and shape registrationabstractIn this paper, we propose an accurate and real-time positioning method for intelligent road vehicles in urban environments. The proposed method uses a robust lane marking detection algorithm, as well as an efficient shape registration algorithm between the detected lane markings and a GPS based road shape prior, to improve the robustness and accuracy of global localization of a road vehicle. We exploit both the state-of-the-art technologies of visual localization based on lane marking detection and the wide availability of Global Positioning System (GPS) based localization. We show that by formulating the positioning problem in a relative sense, we can estimate the vehicle localization in real-time and bound its absolute error in centimeter-level by a cross validation scheme. The validation scheme integrates the vision based lane marking detection with the shape registration, and improves the performance of the overall localization system. The GPS localization can be refined by using lane marking detection when the GPS suffers from frequent satellite signal masking or blockage, while lane marking detection is validated and completing by the GPS based road shape prior when it does not work well in adverse weather conditions or with poor lane signature. We extensively evaluate the proposed method with a single forward-looking camera mounted on an autonomous vehicle which travels at 60km/h through several urban street scenes. Dixiao Cui, Jianru Xue, Shaoyi Du, Nanning Zheng 0001 |
IROS | 3 |
| 2014 | Robust registration of partially overlapping point sets via genetic algorithm with growth operatorabstractRecently, genetic algorithm (GA) has been introduced as an effective method to solve the registration problem. It maintains a population of candidate solutions for the problem and evolves by iteratively applying a set of stochastic operators. Accordingly, a key question is how to reduce the population size. In this study, the authors present two techniques for reducing the population size in the GA for registration of partially overlapping point sets. Based on the trimmed iterative closest point algorithm, they introduce a growth operator into the GA. The growth operator, which is also inspired by the biological evolution, can improve the GA efficiency for registration. Furthermore, they present a technique called centre alignment to confirm the value range of all the registration parameters, which can reduce the search space and allow the well‐designed GA to directly solve the registration problem. Experimental results carried out with the m ‐dimensional point sets illustrate its advantages over previous approaches. Jihua Zhu, Deyu Meng, Zhongyu Li 0002, Shaoyi Du, Zejian Yuan |
IET Image Process. | 4 |
| 2011 | Fast and robust isotropic scaling iterative closest point algorithmabstractThe iterative closest point (ICP) algorithm is an accurate approach for the registration between two point sets on the same scale. However, it can not handle the case with different scales. This paper proposes a fast and robust ICP algorithm for isotropic scaling point sets registration (FRISICP). In order to accurately and directly estimate the scale factor without any constraints, we introduce a bidirection distance measurement method into the least square (LS) problem. Then to keep computational efficiency when the number of points in the set increasing, we further introduce a sparse-to-dense hierarchical model in ICP algorithm to speed up the isotropic scaling point set matching process. Experimental results demonstrate that the proposed FRISICP method outperforms other algorithms on both 2D and 3D point sets. Ce Li 0001, Jianru Xue, Nanning Zheng 0001, Shaoyi Du, Jihua Zhu |
ICIP | 4 |
| 2011 | Expression transfer for facial sketch animation
Yang Yang 0066, Nanning Zheng 0001, Yuehu Liu, Shaoyi Du, Yuanqi Su, Yoshifumi Nishio |
Signal Process. | 4 |
| 2010 | Eye synthesis using the eye curve model
Nanning Zheng 0001, Shaoyi Du, Yuehu Liu |
Image Vis. Comput. | 4 |
| 2010 | Scaling iterative closest point algorithm for registration of m-D point sets
Shaoyi Du, Nanning Zheng 0001, Shihui Ying, Jianru Xue |
J. Vis. Commun. Image Represent. | 1 |
| 2010 | Affine iterative closest point algorithm for point set registration
Shaoyi Du, Nanning Zheng 0001, Shihui Ying |
Pattern Recognit. Lett. | 1 |
| 2009 | Example-based performance driven facial shape animationabstractA novel performance driven facial shape animation method is presented for mapping the expressions from the source face to the target face automatically. Unlike the prior expression cloning approaches, the proposed method aims to animate a new target face with the help of real facial expression samples. The basic idea is to learn the shape deformation from samples for target face to generate corresponding expressions. The process consists of two main stages. First of all, source motion vectors are transferred by statistic face model to generate a reasonable expression on the target face. And then, local deformation constraints are proposed to refine the animation results. In the second part, the local deformation characters for each target organ are learned from the samples, which preserve the personality as well as the expression styles. Experimental results on different facial animation demonstrate the feasibility and effectiveness of the proposed method. Yang Yang 0066, Nanning Zheng 0001, Yuehu Liu, Shaoyi Du, Yoshifumi Nishio |
ICME | 4 |
| 2009 | Facial Expression Synthesis Based on Facial Component ModelabstractStatistical model based facial expression synthesis methods are robust and can be easily used in real environment. But facial expressions of humans are varied. How to represent and synthesize expressions that are is not included in the training set is an unresolved problem in statistical model based researches. In this paper, we propose a two-step method. At first, we propose a statistical appearance model, the facial component model, to represent faces. The model divides the face into seven components, and constructs one global shape model and seven local texture models separately. The motivation to use global shape + local texture strategy is the combination of different components that can generate more types of expression than training sets and the global shape guarantees a "legal" result. Then a neighbor reconstruction framework is proposed to synthesize expressions. The framework estimates the target expression vector by a linear combination of neighbor subject's expression vectors. This paper primarily contributes three things: first, the proposed method can synthesize a wider range of expressions than with the training set. Second, experiments demonstrate that FCM is better than standard AAM in face representation. Third, neighbor reconstruction framework is very flexible. It can be used in multisamples with multitargets and single-sample with single-target applications. Nanning Zheng 0001, Shaoyi Du |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2009 | Lie Group Framework of Iterative Closest Point Algorithm for n-d Data RegistrationabstractThe iterative closet point (ICP) method is a dominant method for data registration that has attracted extensive attention. In this paper, a unified mathematical model of ICP based on Lie group representation is established. Under the framework, the registration problem is formulated into an optimization problem over a certain Lie group. In order to simplify the model and to reduce the dimension of parameter space, the translation part of geometric transformation is eliminated by calibrating the centers of two data sets under registration. As a result, a fast algorithm by solving an iterative linear system is designed for the optimization problem on Lie groups. Moreover, PCA and ICA methods are jointly applied to estimate the initial registration to achieve the global minimum. Finally, several illustrations and comparison experiments are presented to test the performance of the proposed algorithm. Shihui Ying, Shaoyi Du, Hong Qiao |
Int. J. Pattern Recognit. Artif. Intell. | 3 |
| 2009 | A Scale Stretch Method Based on ICP for 3D Data RegistrationabstractIn this paper, we are concerned with the registration of two 3D data sets with large-scale stretches and noises. First, by incorporating a scale factor into the standard iterative closest point (ICP) algorithm, we formulate the registration into a constraint optimization problem over a 7D nonlinear space. Then, we apply the singular value decomposition (SVD) approach to iteratively solving such optimization problem. Finally, we establish a new ICP algorithm, named Scale-ICP algorithm, for registration of the data sets with isotropic stretches. In order to achieve global convergence for the proposed algorithm, we propose a way to select the initial registrations. To demonstrate the performance and efficiency of the proposed algorithm, we give several comparative experiments between Scale-ICP algorithm and the standard ICP algorithm. Shihui Ying, Shaoyi Du, Hong Qiao |
IEEE Trans Autom. Sci. Eng. | 3 |
| 2008 | Illumination transition image: Parameter-based illumination estimation and re-renderingabstractVarying illumination condition is a challenging problem for face recognition and synthesis. The illumination re-rendering technique allows aligning the illumination effects of facial images or relighting them as expected. In this paper, we propose an improved illumination re-rendering method based on more accurate mapping of facial images in the parametric illumination space. This will make the parameter-based illumination alignment more reliable. A clustering-based criterion is designed to evaluate its parameter estimation precision. To guide the image re-rendering between any a parameter pair, an intermediate image called the illumination transition image (ITI) is defined to represent both the illumination variation information and the person-specific facial shape features. The extensive experimental results verify the proposed method outperforms the quotient image approach on both parameter precision and rendering quality. Nanning Zheng 0001, Gaofeng Meng, Shaoyi Du |
ICPR | 5 |
| 2008 | Analysis of Solution for Supervised Graph EmbeddingabstractRecently, Graph Embedding Framework has been proposed for feature extraction. However, it is still an open issue on how to compute robust discriminant transformation for this purpose. In this paper, we show that supervised graph embedding algorithms share a general criterion. Based on the analysis of this criterion, we propose a general solution, called General Solution for Supervised Graph Embedding (GSSGE), for extracting the robust discriminant transformation of Supervised Graph Embedding. Then, we analyze the superiority of our algorithm over traditional algorithms. Extensive experiments on both artificial and real-world data are performed to demonstrate the effectiveness and robustness of our proposed GSSGE. Qubo You, Nanning Zheng 0001, Shaoyi Du, Yang Wu 0001 |
Int. J. Pattern Recognit. Artif. Intell. | 4 |
| 2008 | Solution for supervised graph embedding: A case study
Qubo You, Nanning Zheng 0001, Shaoyi Du |
Signal Process. | 4 |
| 2008 | Affine Registration of Point Sets Using ICP and ICAabstractThis letter proposes a novel algorithm for affine registration of point sets in the way of incorporating an affine transformation into the iterative closest point (ICP) algorithm. At each iterative step of this algorithm, a closed-form solution of the affine transformation is derived. Similar to the ICP algorithm, this new algorithm converges monotonically to a local minimum from any given initial parameters. To get the best affine registration result, good initial parameters are required which are successfully estimated by using independent component analysis (ICA). Experimental results demonstrate the robustness and high accuracy of this algorithm. Shaoyi Du, Nanning Zheng 0001, Gaofeng Meng, Zejian Yuan |
IEEE Signal Process. Lett. | 1 |
| 2008 | Shading Extraction and Correction for Scanned Book ImagesabstractWhen one scans document pages from a bound book, shading artifacts are commonly occurred in the book spine area. In this letter, we propose a general-purpose method for image shading correction based on an assumption that the reflectance function of the page surface is piecewise constant and the illumination function is smooth. The proposed method is able to completely correct more general types of shading artifacts which are nonuniformly distributed along the book spine. Comparison experiments on a synthetic and a variety of real scanned book images demonstrate the feasibility and effectiveness of the proposed method. Gaofeng Meng, Nanning Zheng 0001, Shaoyi Du, Yonghong Song, Yuanlin Zhang 0001 |
IEEE Signal Process. Lett. | 3 |
| 2007 | General Solution for Supervised Graph Embedding
Qubo You, Nanning Zheng 0001, Shaoyi Du, Yang Wu 0001 |
ECML | 3 |
| 2007 | AN Extension of the ICP Algorithm Considering Scale FactorabstractThe ICP algorithm is accurate and fast for registration between two point sets in a same scale, but it doesn't handle the case with different scales. This paper instead introduces a novel approach named the scaling iterative closest point (SICP) algorithm which integrates a scale matrix with boundaries into the original ICP algorithm for scaling registration. This method uses a simple iterative algorithm with the SVD algorithm and the properties of parabola incorporated to compute the translation, rotation and scale transformations at each iterative step, and its convergence is rapid with only a few iterations. The SICP algorithm is independent of shape representation and feature extraction; thereby it is general for scaling registration. Experimental results demonstrate its robustness and fast speed compared with the standard ICP algorithm. Shaoyi Du, Nanning Zheng 0001, Shihui Ying, Qubo You, Yang Wu 0001 |
ICIP (5) | 1 |
| 2007 | Object Recognition by Learning Informative, Biologically Inspired Visual FeaturesabstractThis paper presents a novel, effective way to improve the object recognition performance of a biologically-motivated model by learning informative visual features. The original model has an obvious bottleneck when learning features. Therefore, we propose a circumspect algorithm to solve this problem. First, a novel information factor was designed to find the most informative feature for each image, and then complementary features were selected based on additional information. Finally, an intra-class clustering strategy was used to select the most typical features for each category. By integrating two other improvements, our algorithm performs better than any other system so far based on the same model. Yang Wu 0001, Nanning Zheng 0001, Qubo You, Shaoyi Du |
ICIP (1) | 4 |
| 2007 | ICP with Bounded Scale for Registration of M-D Point SetsabstractThe iterative closest point (ICP) algorithm is an accurate and fast approach for registration between two point sets in a same scale, but it doesn't handle the case with different scales. This paper instead introduces a novel approach named the iterative closest point with bounded scale (ICPBS) algorithm which integrates a scale with boundaries into the traditional ICP algorithm. This proposed technique uses the singular value decomposition algorithm and the properties of parabola to compute the similar transformation at each iterative step, and yields more satisfying robust results than the traditional ICP method in registration between two m-D point sets with different scales. Experimental results demonstrate the presented method is robust and fast for practical use. Shaoyi Du, Nanning Zheng 0001, Shihui Ying, Jishang Wei |
ICME | 1 |
| 2007 | Eye Synthesis Using the Eye Curve ModelabstractEyes are a critical part of facial expressions. Because of the appearance diversity of eyes due to motion, it is difficult to synthesize eye with a particular facial expression. Traditional methods have failed to adequately catch motion-related appearance changes. In order to generate a photorealistic expression eye, we propose a two-step method. Firstly, we propose an eye curve model to represent the eye. The model uses one circle and four skewed elliptical arcs to represent the shape of eyes, and divides the entire eye region into 6 sub-regions that correspond to different anatomical components of the eye. Then we propose a structure-based similarity (SBS) framework to synthesize the expression eye using the eye curve model. This paper primarily contributes three things: first, the proposed eye curve model can represent the diversity of eyes, which is better than some traditional models. Second, when there are many samples, our method can synthesize expression eyes with personal style, which is more reasonable when synthesizing common expressions such as joy. Third, when there is only one sample, our method can clone this expression, which is more useful when synthesizing very special expressions such as a grimace. Experimental results show that all synthesized eyes are realistic and expressive. Nanning Zheng 0001, Qubo You, Shaoyi Du |
ICTAI (2) | 5 |
| 2007 | Neighborhood discriminant projection for face recognition
Qubo You, Nanning Zheng 0001, Shaoyi Du, Yang Wu 0001 |
Pattern Recognit. Lett. | 3 |