Yue Gao 0002

dblp:33/3099-2 · DBLP profile ↗
← Back
285ranked-venue papers
35as first author
140since 2021 · last 2026
0000-0002-4971-590XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Graphics, computer vision, multimedia, augmented reality and games · 157 · 22 first-author · 52 since 2021Artificial intelligence and machine learning · 141 · 10 first-author · 91 since 2021Applied, interdisciplinary, general and emerging computing · 36 · 2 first-author · 21 since 2021Databases, data management, data science and information retrieval · 15 · 4 first-author · 7 since 2021Computer networks · 9 · 3 since 2021Systems, architecture and hardware · 5 · 2 since 2021Human-computer interaction and ubiquitous computing · 2
YearPublicationVenuePosition
2026 Cog-RAG: Cognitive-Inspired Dual-Hypergraph with Theme Alignment Retrieval-Augmented Generation
abstract
Retrieval-Augmented Generation (RAG) enhances the response quality and domain-specific performance of large language models (LLMs) by incorporating external knowledge to combat hallucinations. In recent research, graph structures have been integrated into RAG to enhance the capture of semantic relations between entities. However, it primarily focuses on low-order pairwise entity relations, limiting the high-order associations among multiple entities. Hypergraph-enhanced approaches address this limitation by modeling multi-entity interactions via hyperedges, but they are typically constrained to inter-chunk entity-level representations, overlooking the global thematic organization and alignment across chunks. Drawing inspiration from the top-down cognitive process of human reasoning, we propose a theme-aligned dual-hypergraph RAG framework (Cog-RAG) that uses a theme hypergraph to capture inter-chunk thematic structure and an entity hypergraph to model high-order semantic relations. Furthermore, we design a cognitive-inspired two-stage retrieval strategy that first activates query-relevant thematic content from the theme hypergraph, and then guides fine-grained recall and diffusion in the entity hypergraph, achieving semantic alignment and consistent generation from global themes to local details. Our extensive experiments demonstrate that Cog-RAG significantly outperforms existing state-of-the-art baseline approaches.
Yifan Feng 0001, Ruoxue Li, Rundong Xue, Xingliang Hou, Yue Gao 0002, Shaoyi Du
AAAI7
2026 Role Hypergraph Contrastive Learning for Multivariate Time-Series Analysis
abstract
Multivariate Time-Series (MTS) analysis is crucial across various domains. Considering the spatial and temporal consistency of MTS, existing methods leverage graph structures with temporal augmentation and contrastive learning to achieve robust learning of spatial dependencies and temporal patterns. Given the inherent high-order correlations in MTS, hypergraphs present a promising approach. However, two key challenges limit their further development: 1) Feature-based perspectives capture limited spatial information, while structural perspectives encode richer spatial consistency and evolution dependency; 2) Various semantic patterns (e.g., synergy, inhibition) entangle in sensor correlations, leading to semantic ambiguity. The underlying reason is that conventional hypergraph structures cannot distinguish specific semantic roles within or across hyperedges. Thus, we propose Role Hypergraph Contrastive Learning for MTS analysis. Specifically, we introduce the concept of role to generalize hypergraphs to Role Hypergraphs, enabling precise modeling of sensor correlations by assigning each vertex-hyperedge pair with a semantic role. Building on this structure, we design a role hypergraph contrastive learning paradigm to comprehensively capture the spatial and temporal dependencies: From a structural perspective, role hypergraph structural contrasting captures spatial short-term consistency and long-term evolution; from a feature perspective, alignment of complementary role information ensures sensor-level temporal consistency. Experiments on classification and forecasting tasks demonstrate the effectiveness and interpretability of our method.
Rundong Xue, Zhitao Zeng, Xiangmin Han, Shaoyi Du, Yue Gao 0002
AAAI7
2026 Multi-Channel Batch-Wise Dynamic Hypergraph Network for Multimodal Sentiment Analysis
Wupeng Xie, Zhutian Yang, Yue Gao 0002
ISIT5
2026 SoftHGNN: Soft Hypergraph Neural Networks for General Visual Recognition
Mengqi Lei, Siqi Li 0001, Xinhu Zheng, Shaoyi Du, Yue Gao 0002
Int. J. Comput. Vis.7
2026 M 3 Surv : Fusing Multi-slide and Multi-omics for Memory-augmented robust Survival prediction
abstract
Multimodal survival prediction is crucial for personalized oncology. However, existing methods typically integrate only Formalin-Fixed Paraffin-Embedded (FFPE) slides with a single omics type, such as genomics, overlooking Fresh Frozen (FF) slides that better preserve molecular information, as well as richer multi-omics data like proteomics and transcriptomics. More critically, the complete absence of certain modalities due to clinical constraints ( e.g. , time or cost) severely limits the applicability of conventional fusion models that rely on inter-modality correlations. To address these gaps, we propose M 3 Surv, a framework designed to integrate multi-pathology slides (both FF and FFPE) with multi-omics profiles. For multi-slide fusion, we design a divide-and-conquer hypergraph learning approach to capture both intra-slide higher-order cellular structures and inter-slide relationships, yielding a unified pathology representation. To enrich the biological context, we integrate multi-omics data and employ interactive cross-attention to fuse the pathological and omics modalities. To tackle the missing modality, we introduce a prototype-based memory bank. During training, this memory bank learns and stores representative pathology-omics feature prototypes. At inference, even if a modality is entirely missing, the model can query the bank with available features and robustly impute information from the most similar prototype. Extensive experiments on five TCGA cancer datasets and an in-house dataset demonstrate that M 3 Surv outperforms state-of-the-art methods, achieving an average 2.2% improvement in C-Index. The framework also shows strong stability across various missing modality scenarios, highlighting its clinical potential in real-world, data-incomplete scenarios.
Mingcheng Qu, Donglin Di, Yue Gao 0002, Yang Song 0001, Lei Fan 0007
Medical Image Anal.4
2026 Knowledge-Embedded Hypergraph Neural Networks
abstract
Hypergraph Neural Networks (HGNNs) enhance graph-based modeling by representing complex relationships, with applications in brain network analysis, recommendation systems, and computer vision. However, conventional HGNNs often struggle with effective knowledge extraction and discriminative feature representation, leading to performance limitations. This paper presents Knowledge-Embedded Hypergraph Neural Networks (Knowledge HGNN), a framework that addresses these challenges with two complementary encoders and a multi-dimensional fusion strategy. The High-Order Incidence Encoder (HOI-Encoder) explicitly embeds structural knowledge by capturing permutation-invariant high-order incidence patterns that are typically overlooked by standard HGNNs. In contrast, the Task-Driven Rule Encoder (TDR-Encoder) focuses on feature-level knowledge, extracting task-related rules from vertex attributes through gradient boosted decision tree pre-training and encoding both rule content and positional importance. A Multi-Dimensional Knowledge Fusion module then integrates structural and rule-based embeddings, bridging semantic and dimensional gaps to form enriched vertex representations. The framework includes two implementations: Rule-Driven HGNN, which emphasizes rule-based knowledge, and Dual-Driven HGNN, which jointly leverages structural and rule-based knowledge for comprehensive feature extraction. Extensive experiments on ten datasets, together with ablation studies, demonstrate that Knowledge HGNN significantly improves performance, achieving a 7.3% gain on the Cora dataset and an average improvement of 2.5% across all datasets. These results highlight the effectiveness of explicitly differentiating and fusing structural and rule-based knowledge, setting a new standard for hypergraph applications in complex, data-driven scenarios.
Yifan Feng 0001, Shaoyi Du, Shihui Ying, Zongze Wu 0001, Yue Gao 0002
IEEE Trans. Pattern Anal. Mach. Intell.6
2026 HGNN Shield: Defending Hypergraph Neural Networks Against High-Order Structure Attack
abstract
Hypergraph Neural Networks (HGNNs) are crucial in modeling complex high-order correlations in diverse domains, utilizing hyperedges that connect multiple vertices. However, their susceptibility to structural attacks and irrational connections can disrupt message propagation and degrade performance. To address these issues, we introduce the HGNN Shield, a defense framework incorporating two key modules: Hyperedge-Dependent Estimation (HDE) and High-Order Shield (HOS). The HDE module prioritizes vertex dependencies within hyperedges and adapts traditional connectivity measures to hypergraphs, facilitating precise structural modifications. This adaptation allows for a nuanced assessment of vertex relationships within hyperedges, contributing theoretically by extending classical graph-based connection dependency measures to hypergraphs. Following HDE, the HOS module, positioned before convolutional layers, consists of three submodules: Hyperpath Cut, Hyperpath Link, and Hyperpath Refine. These components collectively detect, disconnect, and refine adversarial connections, ensuring robust message propagation. The theoretical contribution of the HOS module lies in maintaining hyperpath integrity and learning trajectory under adversarial conditions, providing a certifiable defense mechanism against high-order structural attacks. Experiments on six hypergraph datasets indicate that HGNN Shield significantly enhances robustness and maintains data integrity against targeted attacks, outperforming existing methods (an average performance improvement of 9.33% over other methods). Our framework not only improves HGNN reliability but also advances security in hypergraph-based applications.
Yifan Feng 0001, Shaoyi Du, Shihui Ying, Jun-Hai Yong, Yue Gao 0002
IEEE Trans. Pattern Anal. Mach. Intell.6
2026 Hypergraph Foundation Model
abstract
Hypergraph neural networks (HGNNs) effectively model complex high-order relationships in domains like protein interactions and social networks by connecting multiple vertices through hyperedges, enhancing modeling capabilities, and reducing information loss. Developing foundation models for hypergraphs is challenging due to their distinct data, which includes both vertex features and intricate structural information. We present Hyper-FM, a Hypergraph Foundation Model for multi-domain knowledge extraction, featuring Hierarchical High-Order Neighbor Guided Vertex Knowledge Embedding for vertex feature representation and Hierarchical Multi-Hypergraph Guided Structural Knowledge Extraction for structural information. Additionally, we curate 11 text-attributed hypergraph datasets to advance research between HGNNs and LLMs. Experiments on these datasets show that Hyper-FM outperforms baseline methods by approximately 13.4%, validating our approach. Furthermore, we propose the first scaling law for hypergraph foundation models, demonstrating that increasing domain diversity significantly enhances performance, unlike merely augmenting vertex and hyperedge counts. This underscores the critical role of domain diversity in scaling hypergraph models.
Yue Gao 0002, Yifan Feng 0001, Shiquan Liu, Xiangmin Han, Shaoyi Du, Zongze Wu 0001, Han Hu 0003
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 HGNNv2: Stable Hypergraph Neural Networks
abstract
Hypergraph neural networks (HGNNs) are widely used models for analyzing higher-order relational data. HGNNs suffer from the rapid performance degradation with increasing layers. Hypergraph dynamic system (HDS) is a potential way to deal with this challenge. However, hypergraph dynamic system is confined to a time-continuous isotropic model, lacking positional information in the structural space of the hypergraph. In contrast, anisotropic diffusion can capture structural space differences among vertices, providing a more precise representation of the information propagation process in hypergraph structures than isotropic diffusion. In this paper, we introduce HGNNv2, a stable hypergraph neural network, which is built as a hypergraph dynamic system with partial differential equation (PDE). This model incorporates a position-aware anisotropic diffusion term and an external control term. We further present the vertex-rooted subtree method to determine anisotropic diffusion intensity. HGNNv2 has properties that vertices occupying equivalent positions in the structural space share equivalent structural labels and positional features. Experiments on 6 hypergraph datasets and 3 graph datasets reveal that HGNNv2 outperforms all 12 compared methods. HGNNv2 is capable of achieving stable final representations and task accuracy even under noisy conditions. HGNNv2 achieves stable performance with fewer layers than hypergraph dynamic systems employing isotropic diffusion. We provide feature visualizations to illustrate the evolution of representations.
Yue Gao 0002, Jielong Yan, Yifan Feng 0001, Xiangmin Han, Shihui Ying, Zongze Wu 0001, Han Hu 0003
IEEE Trans. Pattern Anal. Mach. Intell.1
2026 Hypergraph-Based High-Order Correlation Analysis for Large-Scale Long-Tailed Data Classification
abstract
High-order correlations, which capture complex interactions among multiple entities, extend beyond traditional graph representations and support a wider range of applications. However, existing neural network models for high-order correlations encounter scalability issues on large datasets due to the substantial computational complexity involved in processing large-scale structures. In addition, long-tailed distributions, which are common in real-world data, result in underrepresented categories and hinder the model's ability to learn effective high-order interaction patterns for rare instances. To address these issues, we introduce a novel framework known as HyperGraph-based High-order Correlation analysis (HGHC) for large-scale long-tailed data classification. Firstly, to tackle the long-tailed distribution problem, HGHC generates synthetic vertices and computes their attributed high-order correlations using an oversampling module inspired by SMOTE, termed HSMOTE, to enhance the representation of tail categories. Secondly, for efficient computational scaling, we treat the data as having two modalities: the structural modality capturing high-order relationships and the feature modality representing individual attributes. We perform computations on both CPU and GPU separately and then fuse the results to achieve a lightweight vertex transformation and aggregation scheme for high-order correlation data. Additionally, we contribute the first benchmark for large-scale long-tailed datasets involving high-order correlations, known as Amazon-LT, which includes multiple datasets with varying imbalance ratios. Our experimental results demonstrate that HGHC achieves state-of-the-art performance in handling high-order correlation analysis issues for large-scale, long-tailed data.
Xiangmin Han, Yubo Zhang 0006, Shihui Ying, Yue Gao 0002
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 Graph Quality Matters on Revealing the Semantics Behind the Data in Physical World
Jielong Yan, Shihui Ying, Shaoyi Du, Yue Gao 0002
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 Reinterpreting Hypergraph Kernels: Insights Through Homomorphism Analysis
abstract
Designing expressive hypergraph kernels that can effectively capture high-order structural information is a fundamental challenge in hypergraph learning. In this paper, we propose a novel comparison framework based on hypergraph homomorphisms to evaluate and compare the expressive ability of existing hypergraph kernels. We revisit classical kernels such as Hypergraph Weisfeiler-Lehman (HG WL) and Hypergraph Rooted kernels, providing theoretical conditions under which they fail to distinguish non-isomorphic hypergraphs. Motivated by these insights, we introduce the Hypergraph Subtree-Cycle Kernel, which augments subtree-based features with cycle-based structural patterns to enhance expressiveness. We propose two variants: HG SCKernelv1 and HG SCKernelv2. Extensive experiments on five graph and ten hypergraph classification benchmarks demonstrate the superior performance of our methods, confirming the effectiveness of integrating homomorphism-guided design into hypergraph kernels.
Shaoyi Du, Yifan Feng 0001, Shihui Ying, Yue Gao 0002
IEEE Trans. Pattern Anal. Mach. Intell.5
2026 Exploring dynamic interpretable brain networks via hierarchical graph transformer
Rundong Xue, Shaoyi Du, Xiangmin Han, Jingxi Feng, Zeyu Zhang 0006, Wei Zeng 0003, Yue Gao 0002
Pattern Recognit.8
2026 Event-based facial expression recognition via large vision-language models
Siqi Li 0001, Yongji Zhang, Yue Gao 0002
Pattern Recognit.4
2026 H3Former: Hypergraph-Based Semantic-Aware Aggregation via Hyperbolic Hierarchical Contrastive Loss for Fine-Grained Visual Classification
abstract
Fine-Grained Visual Classification (FGVC) remains a challenging task due to subtle inter-class differences and large intra-class variations. Existing approaches typically rely on feature-selection mechanisms or region-proposal strategies to localize discriminative regions for semantic analysis. However, these methods often fail to capture discriminative cues comprehensively while introducing substantial category-agnostic redundancy. To address these limitations, we propose $\text {H}^{3}$ Former, a novel token-to-region framework that leverages high-order semantic relations to aggregate local fine-grained representations with structured region-level modeling. Specifically, we propose the Semantic-Aware Aggregation Module (SAAM), which exploits multi-scale contextual cues to dynamically construct a weighted hypergraph among tokens. By applying hypergraph convolution, SAAM captures high-order semantic dependencies and progressively aggregates token features into compact region-level representations. Furthermore, we introduce the Hyperbolic Hierarchical Contrastive Loss (HHCL), which enforces hierarchical semantic constraints in a non-Euclidean embedding space. The HHCL enhances inter-class separability and intra-class consistency while preserving the intrinsic hierarchical relationships among fine-grained categories. Comprehensive experiments conducted on four standard FGVC benchmarks validate the superiority of our $\text {H}^{3}$ Former framework. Code is available at https://github.com/xiaozhangfangyang/H3Former.
Yongji Zhang, Siqi Li 0001, Kuiyang Huang, Yue Gao 0002, Yu Jiang 0006
IEEE Trans. Image Process.4
2026 3D Semantic Gaussian via Geometric-Semantic Hypergraph Computation
abstract
Semantic labels are inherently tied to geometry and luminance reconstruction, as entities with similar shapes and appearances often share categories. Traditional methods use synthesis-analysis, NeRF, or 3D Gaussian representations to encode semantics and geometry separately. However, 2D methods lack view consistency, NeRF extensions are slow, and faster 3D Gaussian methods risk spatial and channel inconsistencies between semantic and RGB. Moreover, these methods require costly manual dense semantic labels. To alleviate resource demands and achieve effective semantic reconstruction with sparse inputs while enhancing RGB rendering quality, we build upon 3D Gaussian by integrating semantic features from pre-trained models-requiring no additional ground truth input-into Gaussian features, and construct a hypergraph neural network to capture higher-order correlations across RGB and semantic information as well as between different frames. Hypergraphs use hyperedges to link multiple vertices, capturing complex relationships essential for cross-modal tasks. This higher-order structure addresses the limitations of NeRF and Gaussian methods, which lack the capacity for such advanced associations. This framework enables precise novel view synthesis and 2D semantic reconstruction without manual annotations, achieving state-of-the-art results for RGB and semantic tasks on room-scale scenes in the ScanNet and Replica datasets, while supporting real-time rendering speeds of 34 FPS.
Dejian Guo, Siqi Li 0001, Shaoyi Du, Xiangmin Han, Yue Gao 0002
IEEE Trans. Multim.7
2026 SkiTrack: An Aerial Skiing Benchmark for Human-Centric Object Tracking
abstract
Aerial skiing is a challenging human-centric sport characterized by rapid motion, large-scale variations, and frequent occlusions. Its extensive spatial range is typically captured by cameras or drones from multiple perspectives, resulting in frequent and complex viewpoint shifts. These challenges encompass nearly all difficulties inherent in human-centric tracking tasks. In this article, we introduce SkiTrack , the first dataset explicitly designed for tracking in aerial skiing. SkiTrack enhances the performance of existing tracking algorithms across a range of human-centric scenarios by providing precise annotations. We observe distinct characteristics in the tracked components, with the skis being rigid and low in visibility and the athlete’s body highly deformable but more visible. To leverage these differences, we propose a components decoupled loss that applies separate constraints to the tracking of the athlete and skis, thereby improving tracking accuracy in skiing scenes. Our experimental results validate the effectiveness of both the SkiTrack dataset and the proposed decoupled loss function, demonstrating consistent improvements in the performance of established models on human-centric tracking tasks. Data are available at https://github.com/xiaozhangfangyang/FineSkiing .
Yu Jiang 0006, Yongji Zhang, Siqi Li 0001, Yuehang Wang, Yue Gao 0002
ACM Trans. Multim. Comput. Commun. Appl.5
2026 GLU-Net: Global-Local Fusion Network for Event-Based Monocular Depth Estimation via Uncertainty Optimization
abstract
Event-based monocular depth estimation is crucial for applications such as autonomous driving, obstacle avoidance, and navigation under high-speed scenarios. Events exhibit a unique and irregular modality. To adapt them to neural networks, some studies convert event streams into event voxels or other frame-like representations. However, these approaches tend to lose the temporal characteristics of events. In this study, we propose a network that aggregates global voxel and per-channel temporal local features of event voxels across the temporal dimension, explicitly extracting events’ temporal information. Furthermore, as noise in events can interfere with the training process and is more difficult to predict than that in images, we utilize the uncertainty estimation module to mitigate the impact of uncertain factors and enhance the robustness of the model. Additionally, we employ multi-level depth features for supervisory training, which improves prediction performance compared to methods relying solely on ground-truth depth supervision. Experiments on open source datasets demonstrate the effectiveness of the proposed method. Our code can be found at https://github.com/WuShangjie/GLUNET .
Shangjie Wu, Jihua Zhu, Zhikuan Zhou, Siqi Li 0001, Shaoyi Du, Yue Gao 0002
ACM Trans. Multim. Comput. Commun. Appl.6
2026 PFF-Net: Patch Feature Fitting for Point Cloud Normal Estimation
abstract
Estimating the normal of a point requires constructing a local patch to provide center-surrounding context, but determining the appropriate neighborhood size is difficult when dealing with different data or geometries. Existing methods commonly employ various parameter-heavy strategies to extract a full feature description from the input patch. However, they still have difficulties in accurately and efficiently predicting normals for various point clouds. In this work, we present a new idea of feature extraction for robust normal estimation of point clouds. We use the fusion of multi-scale features from different neighborhood sizes to address the issue of selecting reasonable patch sizes for various data or geometries. We seek to model a patch feature fitting (PFF) based on multi-scale features to approximate the optimal geometric description for normal estimation and implement the approximation process via multi-scale feature aggregation and cross-scale feature compensation. The feature aggregation module progressively aggregates the patch features of different scales to the center of the patch and shrinks the patch size by removing points far from the center. It not only enables the network to precisely capture the structure characteristic in a wide range, but also describes highly detailed geometries. The feature compensation module ensures the reusability of features from earlier layers of large scales and reveals associated information in different patch sizes. Our approximation strategy based on aggregating the features of multiple scales enables the model to achieve scale adaptation of varying local patches and deliver the optimal feature description. Extensive experiments demonstrate that our method achieves state-of-the-art performance on both synthetic and real-world datasets with fewer network parameters and running time.
Qing Li 0032, Huifang Feng 0002, Kanle Shi, Yue Gao 0002, Yi Fang 0006, Yu-Shen Liu, Zhizhong Han
IEEE Trans. Vis. Comput. Graph.4
2025 A Tutorial on Hypergraph Neural Networks: An In-Depth and Step-By-Step Guide
Sunwoo Kim 0006, Soo Yong Lee, Yue Gao 0002, Alessia Antelmi, Mirko Polato, Kijung Shin
CIKM3
2025 GraphI2P: Image-to-Point Cloud Registration with Exploring Pattern of Correspondence via Graph Learning
abstract
Although the fusion of images and LiDAR point clouds is crucial to many applications in computer vision, the relative poses of cameras and LiDAR scanners are often unknown. However, due to the modality and domain gap between images and LiDAR point clouds, Image-to-Point Cloud Registration is a significant challenge, especially when the image and point cloud come from non-synchronized frames. To tackle these issues, we introduce the virtual point cloud as a bridge to alleviate the cross-modality gap between images and LiDAR point clouds. In this way, the modality gap is converted to the domain gap of point clouds. Moreover, we introduce a virtual-spherical representation achieving orthogonal decoupling between pixel location and predicted depth. As for the domain gap, we propose a distribution-based adaptive sample module to generate a unified distribution of two types of point clouds. Then, we explore the correct correspondence pattern consistency and prune the false correspondences through a graph-based selection process. Experimental results demonstrate that our method outperforms the state-of-the-art methods by more than 10.77% and 12.53% performance on the KITTI Odometry and nuScenes datasets, respectively. The results demonstrate that our method can effectively solve non-synchronized random-frame registration.
Lin Bie, Shouan Pan, Siqi Li 0001, Yue Gao 0002
CVPR5
2025 Cross-Template-Based Hypergraph Transformer
abstract
Single-template-based brain functional network analysis methods can provide limited functional connectivity information, which constrains the performance of brain disease diagnosis. Previous works have explored multi-template functional network analysis but failed to integrate the high-order correlation information within templates and the complementary information between templates into a unified relationship strength between nodes, and we extract the high-order correlation information within each template through hypergraph convolution. Secondly, for the analysis of functional connectivity between templates, we propose a cross-template Transformer to capture long-range dependencies between templates. A cross-template mask is applied to focus the model’s attention on important connections between templates, thereby enhancing model robustness. Finally, we progressively fuse the high-order information captured within templates with the global information across templates for downstream classification tasks. The proposed method has been validated on the public ABIDE dataset, and it outperforms existing methods in the ASD diagnosis task.
Jingxi Feng, Xiangmin Han, Heming Xu, Jue Jiang, Shaoyi Du, Yue Gao 0002
ICASSP7
2025 Hyper-Depth: Hypergraph-Based Multi-Scale Representation Fusion for Monocular Depth Estimation
Lin Bie, Siqi Li 0001, Yifan Feng 0001, Yue Gao 0002
ICCV4
2025 Diff2I2P: Differentiable Image-to-Point Cloud Registration with Diffusion Prior
Juncheng Mu, Chengwei Ren, Weixiang Zhang, Liang Pan, Yue Gao 0002
ICCV6
2025 Beyond Graphs: Can Large Language Models Comprehend Hypergraphs?
abstract
Existing benchmarks like NLGraph and GraphQA evaluate LLMs on graphs by focusing mainly on pairwise relationships, overlooking the high-order correlations found in real-world data. Hypergraphs, which can model complex beyond-pairwise relationships, offer a more robust framework but are still underexplored in the context of LLMs. To address this gap, we introduce LLM4Hypergraph, the first comprehensive benchmark comprising 21,500 problems across eight low-order, five high-order, and two isomorphism tasks, utilizing both synthetic and real-world hypergraphs from citation networks and protein structures. We evaluate six prominent LLMs, including GPT-4o, demonstrating our benchmark’s effectiveness in identifying model strengths and weaknesses. Our specialized prompt- ing framework incorporates seven hypergraph languages and introduces two novel techniques, Hyper-BAG and Hyper-COT, which enhance high-order reasoning and achieve an average 4% (up to 9%) performance improvement on structure classification tasks. This work establishes a foundational testbed for integrating hypergraph computational capabilities into LLMs, advancing their comprehension.
Yifan Feng 0001, Chengwu Yang, Xingliang Hou, Shaoyi Du, Shihui Ying, Zongze Wu 0001, Yue Gao 0002
ICLR7
2025 ERetinex: Event Camera Meets Retinex Theory for Low-Light Image Enhancement
abstract
Low-light image enhancement aims to restore the under-exposure image captured in dark scenarios. Under such scenarios, traditional frame-based cameras may fail to capture the structure and color information due to the exposure time limitation. Event cameras are bio-inspired vision sensors that respond to pixel-wise brightness changes asynchronously. Event cameras' high dynamic range is pivotal for visual perception in extreme low-light scenarios, surpassing traditional cameras and enabling applications in challenging dark environments. In this paper, inspired by the success of the retinex theory for traditional frame-based low-light image restoration, we introduce the first methods that combine the retinex theory with event cameras and propose a novel retinex-based lowlight image restoration framework named ERetinex. Among our contributions, the first is developing a new approach that leverages the high temporal resolution data from event cameras with traditional image information to estimate scene illumination accurately. This method outperforms traditional image-only techniques, especially in low-light environments, by providing more precise lighting information. Additionally, we propose an effective fusion strategy that combines the high dynamic range data from event cameras with the color information of traditional images to enhance image quality. Through this fusion, we can generate clearer and more detailrich images, maintaining the integrity of visual information even under extreme lighting conditions. The experimental results indicate that our proposed method outperforms state-of-theart (SOTA) methods, achieving a gain of 1.0613 dB in PSNR while reducing FLOPS by 84.28 %. The code is available at https://github.com/lodew920/ERetinex.
Xuejian Guo, Yuehang Wang, Siqi Li 0001, Yu Jiang 0006, Shaoyi Du, Yue Gao 0002
ICRA7
2025 Multimodal Cancer Survival Analysis via Hypergraph Learning with Cross-Modality Rebalance
abstract
Multimodal pathology-genomic analysis has become increasingly prominent in cancer survival prediction. However, existing studies mainly utilize multi-instance learning to aggregate patch-level features, neglecting the information loss of contextual and hierarchical details within pathology images. Furthermore, the disparity in data granularity and dimensionality between pathology and genomics leads to a significant modality imbalance. The high spatial resolution inherent in pathology data renders it a dominant role while overshadowing genomics in multimodal integration. In this paper, we propose a multimodal survival prediction framework that incorporates hypergraph learning to effectively capture both contextual and hierarchical details from pathology images. Moreover, it employs a modality rebalance mechanism and an interactive alignment fusion strategy to dynamically reweight the contributions of the two modalities, thereby mitigating the pathology-genomics imbalance. Quantitative and qualitative experiments are conducted on five TCGA datasets, demonstrating that our model outperforms advanced methods by over 3.4% in C-Index performance. Code: https://github.com/MCPathology/MRePath.
Mingcheng Qu, Donglin Di, Tonghua Su, Yue Gao 0002, Yang Song 0001, Lei Fan 0007
IJCAI5
2025 Hypergraph-Guided Federated Distillation Learning for Efficient and Robust Multi-center fMRI Data Analysis
Yidan Xu, Xichun Sheng, Chenggang Yan 0001, Yaoqi Sun, Xiangmin Han, Yue Gao 0002
MICCAI (11)8
2025 Spatially Gene Expression Prediction Using Dual-Scale Contrastive Learning
Mingcheng Qu, Yuncong Wu, Donglin Di, Yue Gao 0002, Tonghua Su, Yang Song 0001, Lei Fan 0007
MICCAI (15)4
2025 Memory-Augmented Incomplete Multimodal Survival Prediction via Cross-Slide and Gene-Attentive Hypergraph Learning
Mingcheng Qu, Donglin Di, Yue Gao 0002, Tonghua Su, Yang Song 0001, Lei Fan 0007
MICCAI (10)4
2025 Adaptive Embedding for Long-Range High-Order Dependencies via Time-Varying Transformer on fMRI
Rundong Xue, Xiangmin Han, Zeyu Zhang 0006, Shaoyi Du, Yue Gao 0002
MICCAI (12)6
2025 DHGFormer: Dynamic Hierarchical Graph Transformer for Disorder Brain Disease Diagnosis
Rundong Xue, Zeyu Zhang 0006, Xiangmin Han, Yue Gao 0002, Shaoyi Du
MICCAI (12)6
2025 Multimodal Hypergraph Guide Learning for Non-invasive CcRCC Survival Prediction
Jielong Yan, Xiangmin Han, Jieyi Zhao, Yue Gao 0002
MICCAI (12)4
2025 Event-enhanced synthetic aperture imaging
Siqi Li 0001, Shaoyi Du, Jun-Hai Yong, Yue Gao 0002
Sci. China Inf. Sci.4
2025 Channel pruning on frequency response
Lin Bie, Chenggang Yan 0001, Xibin Zhao, Yue Gao 0002
Sci. China Inf. Sci.6
2025 Hyper-3DG: Text-to-3D Gaussian Generation via Hypergraph
Donglin Di, Chaofan Luo, Zhou Xue, Wei Chen 0089, Xun Yang 0001, Yue Gao 0002
Int. J. Comput. Vis.7
2025 RGB-D Visual Perception for Occluded Scenes via Event Camera
Siqi Li 0001, Zongze Wu 0001, Zhou Xue, Yu-Shen Liu, Yue Gao 0002
Int. J. Comput. Vis.6
2025 Image Matting and 3D Reconstruction in One Loop
Xinshuang Liu, Siqi Li 0001, Yue Gao 0002
Int. J. Comput. Vis.3
2025 Correction: Multi-source-free Domain Adaptive Object Detection
Sicheng Zhao, Huizai Yao, Chuang Lin 0003, Yue Gao 0002, Guiguang Ding
Int. J. Comput. Vis.4
2025 Cross-sensor contrastive learning-based pre-training for machinery fault diagnosis under sample-limited conditions
Yue Ma 0008, Ruoxue Li, Zhixi Feng, Shuyuan Yang 0001, Shaoyi Du, Yue Gao 0002
Knowl. Based Syst.7
2025 HGTL: A hypergraph transfer learning framework for survival prediction of ccRCC
Xiangmin Han, Wuchao Li, Yan Zhang 0109, Pinhao Li, Jianguo Zhu 0003, Tijiang Zhang, Rongpin Wang, Yue Gao 0002
Medical Image Anal.8
2025 Cross-Modal 3D Shape Retrieval via Heterogeneous Dynamic Graph Representation
abstract
Cross-modal 3D shape retrieval is a crucial and widely applied task in the field of 3D vision. Its goal is to construct retrieval representations capable of measuring the similarity between instances of different 3D modalities. However, existing methods face challenges due to the performance bottlenecks of single-modal representation extractors and the modality gap across 3D modalities. To tackle these issues, we propose a Heterogeneous Dynamic Graph Representation (HDGR) network, which incorporates context-dependent dynamic relations within a heterogeneous framework. By capturing correlations among diverse 3D objects, HDGR overcomes the limitations of ambiguous representations obtained solely from instances. Within the context of varying mini-batches, dynamic graphs are constructed to capture proximal intra-modal relations, and dynamic bipartite graphs represent implicit cross-modal relations, effectively addressing the two challenges above. Subsequently, message passing and aggregation are performed using Dynamic Graph Convolution (DGConv) and Dynamic Bipartite Graph Convolution (DBConv), enhancing features through heterogeneous dynamic relation learning. Finally, intra-modal, cross-modal, and self-transformed features are redistributed and integrated into a heterogeneous dynamic representation for cross-modal 3D shape retrieval. HDGR establishes a stable, context-enhanced, structure-aware 3D shape representation by capturing heterogeneous inter-object relationships and adapting to varying contextual dynamics. Extensive experiments conducted on the ModelNet10, ModelNet40, and real-world ABO datasets demonstrate the state-of-the-art performance of HDGR in cross-modal and intra-modal retrieval tasks. Moreover, under the supervision of robust loss functions, HDGR achieves remarkable cross-modal retrieval against label noise on the 3D MNIST dataset. The comprehensive experimental results highlight the effectiveness and efficiency of HDGR on cross-modal 3D shape retrieval.
Yue Dai 0003, Yifan Feng 0001, Nan Ma 0012, Xibin Zhao, Yue Gao 0002
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 Hyper-YOLO: When Visual Object Detection Meets Hypergraph Computation
abstract
We introduce Hyper-YOLO, a new object detection method that integrates hypergraph computations to capture the complex high-order correlations among visual features. Traditional YOLO models, while powerful, have limitations in their neck designs that restrict the integration of cross-level features and the exploitation of high-order feature interrelationships. To address these challenges, we propose the Hypergraph Computation Empowered Semantic Collecting and Scattering (HGC-SCS) framework, which transposes visual feature maps into a semantic space and constructs a hypergraph for high-order message propagation. This enables the model to acquire both semantic and structural information, advancing beyond conventional feature-focused learning. Hyper-YOLO incorporates the proposed Mixed Aggregation Network (MANet) in its backbone for enhanced feature extraction and introduces the Hypergraph-Based Cross-Level and Cross-Position Representation Network (HyperC2Net) in its neck. HyperC2Net operates across five scales and breaks free from traditional grid structures, allowing for sophisticated high-order interactions across levels and positions. This synergy of components positions Hyper-YOLO as a state-of-the-art architecture in various scale models, as evidenced by its superior performance on the COCO dataset. Specifically, Hyper-YOLO-N significantly outperforms the advanced YOLOv8-N and YOLOv9-T with 12% and 9% improvements.
Yifan Feng 0001, Jiangang Huang, Shaoyi Du, Shihui Ying, Jun-Hai Yong, Guiguang Ding, Rongrong Ji, Yue Gao 0002
IEEE Trans. Pattern Anal. Mach. Intell.9
2025 Self-Supervised Hypergraph Training Framework via Structure-Aware Learning
abstract
Hypergraphs, with their ability to model complex, beyond pair-wise correlations, presents a significant advancement over traditional graphs for capturing intricate relational data across diverse domains. However, the integration of hypergraphs into self-supervised learning (SSL) frameworks has been hindered by the intricate nature of high-order structural variations. This paper introduces the Self-Supervised Hypergraph Training Framework via Structure-Aware Learning (SS-HT), designed to enhance the perception and measurement of these variations within hypergraphs. The SS-HT framework employs a "Masking and Re-Masking" strategy to bolster feature reconstruction in Hypergraph Neural Networks (HGNNs), addressing the limitations of traditional SSL methods. It also introduces a metric strategy for local high-order correlation changes, streamlining the computational efficiency of structural distance calculations. Extensive experiments on 11 datasets demonstrate SS-HT's superior performance over existing SSL methods for both low-order and high-order data. Notably, the framework significantly reduces data labeling dependency, achieving a 32% improvement over HGNN in the downstream task fine-tuning phase under the 1% labeled data setting in the Cora-CC dataset. Ablation studies further validate SS-HT's scalability and its capacity to augment the performance of various HGNN methods, underscoring its robustness and applicability in real-world scenarios.
Yifan Feng 0001, Shiquan Liu, Shihui Ying, Shaoyi Du, Zongze Wu 0001, Yue Gao 0002
IEEE Trans. Pattern Anal. Mach. Intell.6
2025 Kernelized Hypergraph Neural Networks
abstract
Hypergraph Neural Networks (HGNNs) have attracted much attention for high-order structural data learning. Existing methods mainly focus on simple mean-based aggregation or manually combining multiple aggregations to capture multiple information on hypergraphs. However, those methods inherently lack continuous non-linear modeling ability and are sensitive to varied distributions. Although some kernel-based aggregations on GNNs and CNNs can capture non-linear patterns to some degree, those methods are restricted in the low-order correlation and may cause unstable computation in training. In this work, we introduce Kernelized Hypergraph Neural Networks (KHGNN) and its variant, Half-Kernelized Hypergraph Neural Networks (H-KHGNN), which synergize mean-based and max-based aggregation functions to enhance representation learning on hypergraphs. KHGNN's kernelized aggregation strategy adaptively captures both semantic and structural information via learnable parameters, offering a mathematically grounded blend of kernelized aggregation approaches for comprehensive feature extraction. H-KHGNN addresses the challenge of overfitting in less intricate hypergraphs by employing non-linear aggregation selectively in the vertex-to-hyperedge message-passing process, thus reducing model complexity. Our theoretical contributions reveal a bounded gradient for kernelized aggregation, ensuring stability during training and inference. Empirical results demonstrate that KHGNN and H-KHGNN outperform state-of-the-art models across 10 graph/hypergraph datasets, with ablation studies demonstrating the effectiveness and computational stability of our method.
Yifan Feng 0001, Shihui Ying, Shaoyi Du, Yue Gao 0002
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 Inter-Intra Hypergraph Computation for Survival Prediction on Whole Slide Images
abstract
Survival prediction on histopathology whole slide images (WSIs) involves the analysis of multi-level complex correlations, such as inter-correlations among patients and intra-correlations within gigapixel histopathology images. However, the current graph-based methods for WSI analysis mainly focus on the exploration of pairwise correlations, resulting in the loss of high-order correlations. Hypergraph-based methods can handle such high-order correlations, while existing hypergraph-based methods fail to integrate multi-level high-order correlations into a unified framework, which limits the representation capability of WSIs. In this work, we propose an inter-intra hypergraph computation (I$^{2}$2HGC) framework to address this issue. The I$^{2}$2HGC framework implements multi-level hypergraph computation for survival prediction on WSIs, namely intra-hypergraph computation and inter-hypergraph computation. Specifically, the intra-hypergraph computation considers each patch sampled from the histopathology WSI as a vertex of the intra-hypergraph and models the high-order correlations among all patches of an individual WSI in both topology and semantic feature spaces using a hypergraph structure. Then, the intra-hypergraph module generates the intra-embedding and intra-risk for each patient. Subsequently, the inter-hypergraph computation employs these intra-embeddings as features for each patient to form the population-level high-order correlations using data- and knowledge-driven hypergraph modeling strategies. Finally, the intra-risks and the inter-risks are fused for the final survival prediction of each patient. Extensive experimental results on four widely used TCGA carcinoma datasets are presented. We demonstrate that the hypergraph structure captures significantly richer correlations than the graph structure, encompassing all pairwise correlations as well as higher-order interactions through hyperedges. For WSIs with a vast number of pixels and complex correlations, hypergraph-based methods effectively capture topological and semantic information while mitigating the exponential growth of pairwise edges, offering practical advantages for large-scale medical image analysis.
Xiangmin Han, Huijian Zhou, Shaoyi Du, Yue Gao 0002
IEEE Trans. Pattern Anal. Mach. Intell.5
2025 Filter Pruning by High-Order Spectral Clustering
abstract
Large amount of redundancy is widely present in convolutional neural networks (CNNs). Identifying the redundancy in the network and removing the redundant filters is an effective way to compress the CNN model size with a minimal reduction in performance. However, most of the existing redundancy-based pruning methods only consider the distance information between two filters, which can only model simple correlations between filters. Moreover, we point out that distance-based pruning methods are not applicable for high-dimensional features in CNN models by our experimental observations and analysis. To tackle this issue, we propose a new pruning strategy based on high-order spectral clustering. In this approach, we use hypergraph structure to construct complex correlations among filters, and obtain high-order information among filters by hypergraph structure learning. Finally, based on the high-order information, we can perform better clustering on the filters and remove the redundant filters in each cluster. Experiments on various CNN models and datasets demonstrate that our proposed method outperforms the recent state-of-the-art works. For example, with ResNet50, we achieve a 57.1% FLOPs reduction with no accuracy drop on ImageNet, which is the first to achieve lossless pruning with such a high compression ratio.
Yubo Zhang 0006, Lin Bie, Xibin Zhao, Yue Gao 0002
IEEE Trans. Pattern Anal. Mach. Intell.6
2025 Multi-View Spectral Clustering on the Grassmannian Manifold With Hypergraph Representation
abstract
Graph-based multi-view spectral clustering methods have achieved notable progress recently, yet they often fall short in either oversimplifying pairwise relationships or struggling with inefficient spectral decompositions in high-dimensional Euclidean spaces. In this paper, we introduce a novel approach that begins to generate hypergraphs by leveraging sparse representation learning from data points. Based on the generated hypergraph, we propose an optimization function with orthogonality constraints for multi-view hypergraph spectral clustering, which incorporates spectral clustering for each view and ensures consistency across different views. In Euclidean space, solving the orthogonality-constrained optimization problem may yield local maxima and approximation errors. Innovately, we transform this problem into an unconstrained form on the Grassmannian manifold. Finally, we devise an alternating iterative Riemannian optimization algorithm to solve the problem. To validate the effectiveness of the proposed algorithm, we test it on four real-world multi-view datasets and compare its performance with six state-of-the-art multi-view clustering algorithms. The experimental results demonstrate that our method outperforms the baselines in terms of clustering performance due to its superior low-dimensional and resilient feature representation.
Murong Yang, Shihui Ying, Xin-Jian Xu, Yue Gao 0002
IEEE Trans. Big Data4
2025 A Real-World Animation Super-Resolution Benchmark With Color Degradation and Multi-Scale Multi-Frequency Alignment
abstract
Animation super-resolution (SR) aims to generate high-resolution (HR) animation frames from degraded low-resolution (LR) inputs, constituting an important task in real-world SR. Existing animation SR methods typically follow a photorealistic real-world SR computational paradigm. However, digital animation frames commonly suffer from compression and transmission-related degradation, distinct from degradations in camera-captured real-world images. In this paper, we introduce a novel real-world animation super-resolution benchmark designed explicitly for animation frames, named ADASR, featuring both 2D and modern 3D animation content to facilitate industry applications. Additionally, we propose a Color-Aware Animation Super-Resolution (CAASR) method. CAASR, for the first time, incorporates a color degradation simulation mechanism tailored for animations, addressing color banding, blocking, and color shift. Furthermore, we develop a multi-scale multi-frequency alignment mechanism to robustly extract degradation-invariant features. Extensive experiments conducted on both the existing AVC dataset and our newly constructed ADASR dataset demonstrate that our proposed CAASR achieves state-of-the-art performance in restoring HR frames for both 2D and 3D animations. Code and data are available at https://github.com/huangyang-666/CAASR.
Yu Jiang 0006, Yongji Zhang, Siqi Li 0001, Yuehang Wang, Yutong Yao, Yue Gao 0002
IEEE Trans. Image Process.7
2025 Exploring Local and Global Consistent Correlation on Hypergraph for Rotation Invariant Point Cloud Analysis
abstract
Rotation invariant point cloud analysis is essential for many real-world applications where objects can appear in arbitrary orientations. Traditional local rotation-invariant methods rely on lossy region descriptors, limiting the global comprehension of 3D objects. Conversely, global features derived from pose alignment can capture complementary information. To leverage both local and global consistency for enhanced accuracy, we propose the Global-Local-Consistent Hypergraph Cross-Attention Network (GLC-HCAN). This framework includes the Global Consistent Feature (GCF) representation branch, the Local Consistent Feature (LCF) representation branch, and the Hypergraph Cross-Attention (HyperCA) network to model complex correlations through the global-local-consistent hypergraph representation learning. Specifically, the GCF branch employs a multi-pose grouping and aggregation strategy based on PCA for improved global comprehension. Simultaneously, the LCF branch uses local farthest reference point features to enhance local region descriptions. To capture high-order and complex global-local correlations, we construct hypergraphs that integrate both features, mutually enhancing and fusing the representations. The inductive HyperCA module leverages attention techniques to better utilize these high-order relations for comprehensive understanding. Consequently, GLC-HCAN offers an effective and robust rotation-invariant point cloud analysis network, suitable for object classification and shape retrieval tasks in SO(3). Experimental results on both synthetic and scanned point cloud datasets demonstrate that GLC-HCAN outperforms state-of-the-art methods.
Yue Dai 0003, Shihui Ying, Yue Gao 0002
IEEE Trans. Multim.3
2025 EvCSLR: Event-Guided Continuous Sign Language Recognition and Benchmark
abstract
Classical continuous sign language recognition (CSLR) suffers from some main challenges in real-world scenarios: accurate inter-frame movement trajectories may fail to be captured by traditional RGB cameras due to the motion blur, and valid information may be insufficient under low-illumination scenarios. In this paper, we for the first time leverage an event camera to overcome the above-mentioned challenges. Event cameras are bio-inspired vision sensors that could efficiently record high-speed sign language movements under low-illumination scenarios and capture human information while eliminating redundant background interference. To fully exploit the benefits of the event camera for CSLR, we propose a novel event-guided multi-modal CSLR framework, which could achieve significant performance under complex scenarios. Specifically, a time redundancy correction (TRCorr) module is proposed to rectify redundant information in the temporal sequences, directing the model to focus on distinctive features. A multi-modal cross-attention interaction (MCAI) module is proposed to facilitate information fusion between events and frame domains. Furthermore, we construct the first event-based CSLR dataset, namedEvCSLR, which will be released as the first event-based CSLR benchmark. Experimental results demonstrate that our proposed method achieves state-of-the-art performance on EvCSLR and PHOENIX-2014 T datasets.
Yu Jiang 0006, Yuehang Wang, Siqi Li 0001, Yongji Zhang, Qianren Guo, Qi Chu 0010, Yue Gao 0002
IEEE Trans. Multim.7
2025 Hypergraph-Based Remaining Prototype Alignment for Open-Set Cross-Domain Image Retrieval
abstract
Existing cross-domain image retrieval (CDIR) methods exhibit a strong dependency on prior knowledge of training categories, which leads to problems of class confusion and domain shift when encountering unseen categories in open-set environments. In this paper, we explore the CDIR task towards open-set environments and introduce the Hypergraph-Based Remaining Prototype Alignment (RePro) framework for this task. Specifically, to address the problem of unseen class confusion caused by the category differences, we utilize the Remaining Prototype Embedding (RPE) module to generate the remaining embeddings of images and treat these embeddings as domain noise, rather than directly mapping them to the explicit domain-unified prototypes. To overcome the problem of domain shift, our method leverages the high-order correlations among both domains and categories through the Heterogeneous Structure Alignment (HSA) module, by constructing a heterogeneous hypergraph based on intra-domain and inter-category correlations. Besides, we build two multi-domain datasets for open-set cross-domain image retrieval,i.e., OCD-PACS and OCD-VLCS. Each dataset is divided into seen and unseen categories for training and testing, and each class has four different domains of images. Extensive experiments and ablation studies on these two datasets demonstrate the superiority of our method over current state-of-the-art methods.
Yang Xu 0064, Yifan Feng 0001, Xiaopin Zhong, Yue Gao 0002, Zongze Wu 0001
IEEE Trans. Multim.4
2025 Residual Fuzzy Alignment on Hypergraph for Open-Set 3D Cross-Modal Retrieval
abstract
Existing 3D cross-modal retrieval (3CMR) methods heavily rely on prior knowledge of training categories, which leads to the problem of modality shift and unseen center deviation when encountering unseen categories under the open-set environment. Aiming at the open-set 3CMR, this paper introduces the Hypergraph-Based Residual Fuzzy Alignment (ReFA) framework, which revisits the open-set retrieval task and navigates uncertainty of it through the lens ofFuzzy Theory. Facing the challenges of boundaryless space caused by uncertain unseen categories, we explore the representation and measurement in the fuzzy membership space as an alternative to fixed close-set category space. Specifically, to address the problem of modality shift caused by unseen categories, we utilize the Residual Sampling Generation (RSG) module to generate modality sampling embeddings that are independent of seen categories under the guidance of fuzzy representation, which residually decouples the entangled interactions of seen categories and modalities. To overcome the problem of unseen center deviation, we propose the Center Fuzzy Alignment (CFA) module to leverage the high-order fuzzy correlations for generalized metric, by constructing a fuzzy hypergraph based on the inherent and fuzzy correlations among both modalities and categories. The comprehensive evaluations of comparison and ablation studies on the four benchmarks demonstrate the superiority of our proposed framework compared to state-of-the-art methods.
Yang Xu 0064, Yifan Feng 0001, Xu Zhuang, Zongze Wu 0001, Yue Gao 0002
IEEE Trans. Multim.6
2025 Hypergraph Foundation Model for Brain Disease Diagnosis
abstract
The goal of the hypergraph foundation model (HGFM) is to learn an encoder based on the hypergraph computational paradigm through self-supervised pretraining on high-order correlation structures, enabling the encoder to rapidly adapt to various downstream tasks in scenarios, where no labeled data or only a small amount of labeled data are available. The initial exploratory work has been applied to brain disease diagnosis tasks. However, existing methods primarily rely on graph-based approaches to learn low-order correlation patterns between brain regions in brain networks, neglecting the modeling and learning of complex correlations between different brain diseases and patients. This article proposes an HGFM for brain disease diagnosis, which conducts multidimensional pretraining tasks to explore latent cross-dimensional high-order correlation patterns on various brain disease datasets. HGFM is a high-order correlation-driven foundation model for brain disease diagnosis and effectively improves prediction performance. Specifically, HGFM first performs brain functional network link prediction tasks on individual brain networks and group interaction network link prediction tasks on group brain networks, constructing an HGFM for brain disease diagnosis. In downstream tasks, it achieves predictions for different brain disease diagnosis tasks through few-shot learning fine-tuning methods. The proposed method is evaluated on functional magnetic resonance imaging (fMRI) data from 4409 patients across four brain diseases. Results show that it outperforms existing state-of-the-art methods in all brain disease diagnosis tasks, demonstrating its potential value in clinical applications.
Xiangmin Han, Rundong Xue, Jingxi Feng, Yifan Feng 0001, Shaoyi Du, Jun Shi 0004, Yue Gao 0002
IEEE Trans. Neural Networks Learn. Syst.7
2025 Mode Hypergraph Neural Network
abstract
The hypergraph neural network (HGNN) is an emerging powerful tool for modeling and learning complex, high-order correlations among entities upon hypergraph structures. While existing HGNN-based approaches excel in modeling high-order correlations among data using hyperedges, they often have difficulties in distinguishing diverse semantics (e.g., bioactivities between drug and target in biological networks) of different correlations, making it challenging to learn accurate final representations. The underlying reason is that the specific semantic information of each hyperedge cannot be captured and distinguished during the modeling and learning process. To address this, we propose a mode HGNN ( $\textsf {MHGNN}$ ) framework that extends the vanilla hypergraph structure by endowing hyperedges with mode information for encapsulating their semantics and then performs mode-aware high-order message passing upon mode hypergraph for achieving comprehensive node representations. Extensive evaluations on four real-world datasets under two representative tasks have demonstrated the outstanding performance of $\textsf {MHGNN}$ against the state of the arts.
Shuyi Ji, Yifan Feng 0001, Donglin Di, Shihui Ying, Yue Gao 0002
IEEE Trans. Neural Networks Learn. Syst.5
2025 Hierarchical Set-to-Set Representation for 3-D Cross-Modal Retrieval
abstract
Three-dimensional in-domain retrieval has recently achieved significant success, but 3-D cross-modal retrieval still faces problems and challenges. Existing methods only rely on a simple global feature (GF), which overlooks the local information of complex 3-D objects and the connections between similar local features across complex multimodal instances. To tackle this issue, we propose a hierarchical set-to-set representation (HSR) and a corresponding hierarchical similarity that incorporates global-to-global and local-to-local similarity metrics. Specifically, we employ feature extractors for each modality to learn both GFs and local feature sets. We then project these features into their respective common space and use bilinear pooling to generate compact-set features that maintain the invariant for set-to-set similarity measurement. To facilitate effective hierarchical similarity measurement, we design an operation to combine the GF and the compact-set feature to generate the hierarchical representation for 3-D cross-modal retrieval, which preserves hierarchical similarity measurement. To optimize the framework, we adopt the joint loss functions, including cross-modal center loss (CMCL), mean square loss, and cross-entropy loss, to reduce the cross-modal discrepancy for each instance and minimize the distances between the instances in the same category. Experimental results demonstrate that our method outperforms the state-of-the-art methods on the 3-D cross-modal retrieval task on both ModelNet10 and ModelNet40 datasets.
Yu Jiang 0006, Cong Hua, Yifan Feng 0001, Yue Gao 0002
IEEE Trans. Neural Networks Learn. Syst.4
2025 Arbitrary Large-Scale Scene Reconstruction without Annotated Block Partitions
abstract
Large-scale scene reconstruction is a challenging problem. As different parts of the scene could be visible from different collected image frames, previous works manually use distance or geography to decompose the scene into parts and reconstruct each part of the scene separately. However, such manual decomposition is a laborious and time-consuming task when applied to large-scale scene reconstruction in real-world applications. To address this, we propose VisibleNeRF automatically reconstructs large-scale scenes by decomposing scenes into parts based on the part visibility. More specifically, we propose a visibility judgment strategy to decompose the scenes into visible and invisible parts. Then we reconstruct the visible part with the corresponding collected images and continue to decompose the rest of the invisible parts with the proposed visibility judgment strategy. New NeRF modules are re-established for the decomposed invisible parts until the entire scene is reconstructed. To the best of our knowledge, we are the first to propose an online reconstruction of large-scale scenes without manual decomposition. Experimental results on three datasets show that our method successfully reconstructs large-scale scenes in a fully automatic manner. Besides, in the widely used Mission Bay dataset, our model outperforms other state-of-the-art methods by a large margin.
Lin Bie, Siqi Li 0001, Dejian Guo, Shaoyi Du, Yue Gao 0002
ACM Trans. Multim. Comput. Commun. Appl.7
2024 Multi-Energy Guided Image Translation with Stochastic Differential Equations for Near-Infrared Facial Expression Recognition
abstract
Illumination variation has been a long-term challenge in real-world facial expression recognition (FER). Under uncontrolled or non-visible light conditions, near-infrared (NIR) can provide a simple and alternative solution to obtain high-quality images and supplement the geometric and texture details that are missing in the visible (VIS) domain. Due to the lack of large-scale NIR facial expression datasets, directly extending VIS FER methods to the NIR spectrum may be ineffective. Additionally, previous heterogeneous image synthesis methods are restricted by low controllability without prior task knowledge. To tackle these issues, we present the first approach, called for NIR-FER Stochastic Differential Equations (NFER-SDE), that transforms face expression appearance between heterogeneous modalities to the overfitting problem on small-scale NIR data. NFER-SDE can take the whole VIS source image as input and, together with domain-specific knowledge, guide the preservation of modality-invariant information in the high-frequency content of the image. Extensive experiments and ablation studies show that NFER-SDE significantly improves the performance of NIR FER and achieves state-of-the-art results on the only two available NIR FER datasets, Oulu-CASIA and Large-HFE.
Bingjun Luo, Xibin Zhao, Yue Gao 0002
AAAI6
2024 Hypergraph-Guided Disentangled Spectrum Transformer Networks for Near-Infrared Facial Expression Recognition
abstract
With the strong robusticity on illumination variations, near-infrared (NIR) can be an effective and essential complement to visible (VIS) facial expression recognition in low lighting or complete darkness conditions. However, facial expression recognition (FER) from NIR images presents a more challenging problem than traditional FER due to the limitations imposed by the data scale and the difficulty of extracting discriminative features from incomplete visible lighting contents. In this paper, we give the first attempt at deep NIR facial expression recognition and propose a novel method called near-infrared facial expression transformer (NFER-Former). Specifically, to make full use of the abundant label information in the field of VIS, we introduce a Self-Attention Orthogonal Decomposition mechanism that disentangles the expression information and spectrum information from the input image, so that the expression features can be extracted without the interference of spectrum variation. We also propose a Hypergraph-Guided Feature Embedding method that models some key facial behaviors and learns the structure of the complex correlations between them, thereby alleviating the interference of inter-class similarity. Additionally, we construct a large NIR-VIS Facial Expression dataset that includes 360 subjects to better validate the efficiency of NFER-Former. Extensive experiments and ablation studies show that NFER-Former significantly improves the performance of NIR FER and achieves state-of-the-art results on the only two available NIR FER datasets, Oulu-CASIA and Large-HFE.
Bingjun Luo, Xibin Zhao, Yue Gao 0002
AAAI6
2024 3D Feature Tracking via Event Camera
abstract
This paper presents the first 3D feature tracking method with the corresponding dataset. Our proposed method takes event streams from stereo event cameras as input to pre-dict 3D trajectories of the target features with high-speed motion. To achieve this, our method leverages a joint framework to predict the 2D feature motion offsets and the 3D feature spatial position simultaneously. A motion compensation module is leveraged to overcome the feature deformation. A patch matching module based on bi-polarity hypergraph modeling is proposed to robustly es-timate the feature spatial position. Meanwhile, we collect the first 3D feature tracking dataset with high-speed moving objects and ground truth 3D feature trajectories at 250 FPS, named E-3DTrack, which can be used as the first high-speed 3D feature tracking benchmark. Our code and dataset could be found at: https://github.com/lisiqi19971013/E-3DTrack.
Siqi Li 0001, Zhikuan Zhou, Zhou Xue, Shaoyi Du, Yue Gao 0002
CVPR6
2024 ColorPCR: Color Point Cloud Registration with Multi-Stage Geometric-Color Fusion
abstract
Point cloud registration is still a challenging and open problem. For example, when the overlap between two point clouds is extremely low, geo-only features may be not suf-ficient. Therefore, it is important to further explore how to utilize color data in this task. Under such circumstances, we propose ColorPCR for color point cloud registration with multi-stage geometric-color fusion. We design a Hier-archical Color Enhanced Feature Extraction module to ex-tract multi-level geometric-color features, and a GeoColor Superpoint Matching Module to encode transformation-invariant geo-color global context for robust patch corre-spondences. In this way, both geometric and color data can be used, thus leading to robust performance even under extremely challenging scenarios, such as low overlap between two point clouds. To evaluate the performance of our method, we colorize 3DMatch/3DLoMatch datasets as Color3DMatch/Color3DLoMatch and evaluations on these datasets demonstrate the effectiveness of our proposed method. Our method achieves state-of-the-art registration recall of 97.5%/88.9% on them.
Juncheng Mu, Lin Bie, Shaoyi Du, Yue Gao 0002
CVPR4
2024 LightHGNN: Distilling Hypergraph Neural Networks into MLPs for 100x Faster Inference
abstract
Hypergraph Neural Networks (HGNNs) have recently attracted much attention and exhibited satisfactory performance due to their superiority in high-order correlation modeling. However, it is noticed that the high-order modeling capability of hypergraph also brings increased computation complexity, which hinders its practical industrial deployment. In practice, we find that one key barrier to the efficient deployment of HGNNs is the high-order structural dependencies during inference. In this paper, we propose to bridge the gap between the HGNNs and inference-efficient Multi-Layer Perceptron (MLPs) to eliminate the hypergraph dependency of HGNNs and thus reduce computational complexity as well as improve inference speed. Specifically, we introduce LightHGNN and LightHGNN$^+$ for fast inference with low complexity. LightHGNN directly distills the knowledge from teacher HGNNs to student MLPs via soft labels, and LightHGNN$^+$ further explicitly injects reliable high-order correlations into the student MLPs to achieve topology-aware distillation and resistance to over-smoothing. Experiments on eight hypergraph datasets demonstrate that even without hypergraph dependency, the proposed LightHGNNs can still achieve competitive or even better performance than HGNNs and outperform vanilla MLPs by $16.3$ on average. Extensive experiments on three graph datasets further show the average best performance of our LightHGNNs compared with all other methods. Experiments on synthetic hypergraphs with 5.5w vertices indicate LightHGNNs can run $100\times$ faster than HGNNs, showcasing their ability for latency-sensitive deployments.
Yifan Feng 0001, Yihe Luo, Shihui Ying, Yue Gao 0002
ICLR4
2024 Hypergraph Dynamic System
abstract
Recently, hypergraph neural networks (HGNNs) exhibit the potential to tackle tasks with high-order correlations and have achieved success in many tasks. However, existing evolution on the hypergraph has poor controllability and lacks sufficient theoretical support (like dynamic systems), thus yielding sub-optimal performance. One typical scenario is that only one or two layers of HGNNs can achieve good results and more layers lead to degeneration of performance. Under such circumstances, it is important to increase the controllability of HGNNs. In this paper, we first introduce hypergraph dynamic systems (HDS), which bridge hypergraphs and dynamic systems and characterize the continuous dynamics of representations. We then propose a control-diffusion hypergraph dynamic system by an ordinary differential equation (ODE). We design a multi-layer HDS$^{ode}$ as a neural implementation, which contains control steps and diffusion steps. HDS$^{ode}$ has the properties of controllability and stabilization and is allowed to capture long-range correlations among vertices. Experiments on $9$ datasets demonstrate HDS$^{ode}$ beat all compared methods. HDS$^{ode}$ achieves stable performance with increased layers and solves the poor controllability of HGNNs. We also provide the feature visualization of the evolutionary process to demonstrate the controllability and stabilization of HDS$^{ode}$.
Jielong Yan, Yifan Feng 0001, Shihui Ying, Yue Gao 0002
ICLR4
2024 Position: Topological Deep Learning is the New Frontier for Relational Learning
abstract
Topological deep learning (TDL) is a rapidly evolving field that uses topological features to understand and design deep learning models. This paper posits that TDL is the new frontier for relational learning. TDL may complement graph representation learning and geometric deep learning by incorporating topological concepts, and can thus provide a natural choice for various machine learning settings. To this end, this paper discusses open problems in TDL, ranging from practical benefits to theoretical foundations. For each problem, it outlines potential solutions and future research opportunities. At the same time, this paper serves as an invitation to the scientific community to actively participate in TDL research to unlock the potential of this emerging field.
Theodore Papamarkou, Tolga Birdal, Michael M. Bronstein, Gunnar E. Carlsson, Justin Curry, Yue Gao 0002, Mustafa Hajij, Roland Kwitt, Pietro Liò, Paolo Di Lorenzo, Vasileios Maroulas, Nina Miolane, Farzana Nasrin, Karthikeyan Natesan Ramamurthy, Bastian Rieck, Simone Scardapane, Michael T. Schaub, Petar Velickovic, Bei Wang 0001, Yusu Wang 0001, Guo-Wei Wei 0001, Ghada Zamzmi
ICML6
2024 3D-OAE: Occlusion Auto-Encoders for Self-Supervised Learning on Point Clouds
abstract
The manual annotation for large-scale point clouds is still tedious and unavailable for many harsh real-world tasks. Self-supervised learning, which is used on raw and unlabeled data to pre-train deep neural networks, is a promising approach to address this issue. Existing works usually take the common aid from auto-encoders to establish the self-supervision by the self-reconstruction schema. However, the previous auto-encoders merely focus on the global shapes and do not distinguish the local and global geometric features apart. To address this problem, we present a novel and efficient self-supervised point cloud representation learning framework, named 3D Occlusion Auto-Encoder (3D-OAE), to facilitate the detailed supervision inherited in local regions and global shapes. We propose to randomly occlude some local patches of point clouds and establish the supervision via inpainting the occluded patches using the remaining ones. Specifically, we design an asymmetrical encoder-decoder architecture based on standard Transformer, where the encoder operates only on the visible subset of patches to learn local patterns, and a lightweight decoder is designed to leverage these visible patterns to infer the missing geometries via self-attention. We find that occluding a very high proportion of the input point cloud (e.g. 75%) will still yield a nontrivial self-supervisory performance, which enables us to achieve 3-4 times faster during training but also improve accuracy. Experimental results show that our approach outperforms the state-of-the-art on a diverse range of down-stream discriminative and generative tasks. Code is available at https://github.com/junshengzhou/3D-OAE.
Junsheng Zhou, Xin Wen 0003, Baorui Ma, Yu-Shen Liu, Yue Gao 0002, Yi Fang 0006, Zhizhong Han
ICRA5
2024 Negative Prompt Driven Complementary Parallel Representation for Open-World 3D Object Retrieval
Yang Xu 0064, Yifan Feng 0001, Yue Gao 0002
IJCAI3
2024 A Survey on Hypergraph Neural Networks: An In-Depth and Step-By-Step Guide
abstract
Higher-order interactions (HOIs) are ubiquitous in real-world complex systems and applications. Investigation of deep learning for HOIs, thus, has become a valuable agenda for the data mining and machine learning communities. As networks of HOIs are expressed mathematically as hypergraphs, hypergraph neural networks (HNNs) have emerged as a powerful tool for representation learning on hypergraphs. Given the emerging trend, we present the first survey dedicated to HNNs, with an in-depth and step-by-step guide. Broadly, the present survey overviews HNN architectures, training strategies, and applications. First, we break existing HNNs down into four design components: (i) input features, (ii) input structures, (iii) message-passing schemes, and (iv) training strategies. Second, we examine how HNNs address and learn HOIs with each of their components. Third, we overview the recent applications of HNNs in recommendation, bioinformatics and medical science, time series analysis, and computer vision. Lastly, we conclude with a discussion on limitations and future directions.
Sunwoo Kim 0006, Soo Yong Lee, Yue Gao 0002, Alessia Antelmi, Mirko Polato, Kijung Shin
KDD3
2024 Inter-intra High-Order Brain Network for ASD Diagnosis via Functional MRIs
Xiangmin Han, Rundong Xue, Shaoyi Du, Yue Gao 0002
MICCAI (2)4
2024 PathoTune: Adapting Visual Foundation Model to Pathological Specialists
Jiaxuan Lu, Fang Yan 0002, Xiaofan Zhang 0002, Yue Gao 0002, Shaoting Zhang 0001
MICCAI (4)4
2024 ccRCC Metastasis Prediction via Exploring High-Order Correlations on Multiple WSIs
Huijian Zhou, Xiangmin Han, Shaoyi Du, Yue Gao 0002
MICCAI (5)5
2024 Multi-scale Consistency for Robust 3D Registration via Hierarchical Sinkhorn Tree
abstract
We study the problem of retrieving accurate correspondence through multi-scale consistency (MSC) for robust point cloud registration. Existing works in a coarse-to-fine manner either suffer from severe noisy correspondences caused by unreliable coarse matching or struggle to form outlier-free coarse-level correspondence sets. To tackle this, we present Hierarchical Sinkhorn Tree (HST), a pruned tree structure designed to hierarchically measure the local consistency of each coarse correspondence across multiple feature scales, thereby filtering out the local dissimilar ones. In this way, we convert the modeling of MSC for each correspondence into a BFS traversal with pruning of a K-ary tree rooted at the superpoint, with its K nearest neighbors in the feature pyramid serving as child nodes. To achieve efficient pruning and accurate vicinity characterization, we further propose a novel overlap-aware Sinkhorn Distance, which retains only the most likely overlapping points for local measurement and next level exploration. The modeling process essentially involves traversing a pair of HSTs synchronously and aggregating the consistency measures of corresponding tree nodes. Extensive experiments demonstrate HST consistently outperforms the state-of-the-art methods on both indoor and outdoor benchmarks.
Chengwei Ren, Yifan Feng 0001, Weixiang Zhang, Xiao-Ping Zhang 0002, Yue Gao 0002
NeurIPS5
2024 Semi-Open 3D Object Retrieval via Hierarchical Equilibrium on Hypergraph
abstract
Existing open-set learning methods consider only the single-layer labels of objects and strictly assume no overlap between the training and testing sets, leading to contradictory optimization for superposed categories. In this paper, we introduce a more practical Semi-Open Environment setting for open-set 3D object retrieval with hierarchical labels, in which the training and testing set share a partial label space for coarse categories but are completely disjoint from fine categories. We propose the Hypergraph-Based Hierarchical Equilibrium Representation (HERT) framework for this task. Specifically, we propose the Hierarchical Retrace Embedding (HRE) module to overcome the global disequilibrium of unseen categories by fully leveraging the multi-level category information. Besides, tackling the feature overlap and class confusion problem, we perform the Structured Equilibrium Tuning (SET) module to utilize more equilibrial correlations among objects and generalize to unseen categories, by constructing a superposed hypergraph based on the local coherent and global entangled correlations. Furthermore, we generate four semi-open 3DOR datasets with multi-level labels for benchmarking. Results demonstrate that the proposed method can effectively generate the hierarchical embeddings of 3D objects and generalize them towards semi-open environments.
Yang Xu 0064, Yifan Feng 0001, Jun Zhang 0018, Jun-Hai Yong, Yue Gao 0002
NeurIPS5
2024 Assembly Fuzzy Representation on Hypergraph for Open-Set 3D Object Retrieval
abstract
The lack of object-level labels presents a significant challenge for 3D object retrieval in the open-set environment. However, part-level shapes of objects often share commonalities across categories but remain underexploited in existing retrieval methods. In this paper, we introduce the Hypergraph-Based Assembly Fuzzy Representation (HARF) framework, which navigates the intricacies of open-set 3D object retrieval through a bottom-up lens of Part Assembly. To tackle the challenge of assembly isomorphism and unification, we propose the Hypergraph Isomorphism Convolution (HIConv) for smoothing and adopt the Isomorphic Assembly Embedding (IAE) module to generate assembly embeddings with geometric-semantic consistency. To address the challenge of open-set category generalization, our method employs high-order correlations and fuzzy representation to mitigate distribution skew through the Structure Fuzzy Reconstruction (SFR) module, by constructing a leveraged hypergraph based on local certainty and global uncertainty correlations. We construct three open-set retrieval datasets for 3D objects with part-level annotations: OP-SHNP, OP-INTRA, and OP-COSEG. Extensive experiments and ablation studies on these three benchmarks show our method outperforms current state-of-the-art methods.
Yang Xu 0064, Yifan Feng 0001, Jun Zhang 0018, Jun-Hai Yong, Yue Gao 0002
NeurIPS5
2024 Towards Language-Guided Visual Recognition via Dynamic Convolutions
Gen Luo, Yiyi Zhou, Xiaoshuai Sun, Yongjian Wu 0001, Yue Gao 0002, Rongrong Ji
Int. J. Comput. Vis.5
2024 Multi-source-free Domain Adaptive Object Detection
Sicheng Zhao, Huizai Yao, Chuang Lin 0003, Yue Gao 0002, Guiguang Ding
Int. J. Comput. Vis.4
2024 Hypergraph Isomorphism Computation
abstract
The isomorphism problem is a fundamental problem in network analysis, which involves capturing both low-order and high-order structural information. In terms of extracting low-order structural information, graph isomorphism algorithms analyze the structural equivalence to reduce the solver space dimension, which demonstrates its power in many applications, such as protein design, chemical pathways, and community detection. For the more commonly occurring high-order relationships in real-life scenarios, the problem of hypergraph isomorphism, which effectively captures these high-order structural relationships, cannot be straightforwardly addressed using graph isomorphism methods. Besides, the existing hypergraph kernel methods may suffer from high memory consumption or inaccurate sub-structure identification, thus yielding sub-optimal performance. In this paper, to address the abovementioned problems, we first propose the hypergraph Weisfiler-Lehman test algorithm for the hypergraph isomorphism test problem by generalizing the Weisfiler-Lehman test algorithm from graphs to hypergraphs. Secondly, based on the presented algorithm, we propose a general hypergraph Weisfieler-Lehman kernel framework and implement two instances, which are Hypergraph Weisfeiler-Lehamn Subtree Kernel (Hypergraph WL Subtree Kernel) and Hypergraph Weisfeiler-Lehamn Hyperedge Kernel (Hypergraph WL Hyperedge Kernel). The Hypergraph WL Subtree Kernel counts different types of rooted subtrees and generates the final feature vector for a given hypergraph by comparing the number of different types of rooted subtrees. The Hypergraph WL Hyperedge Kernel is developed to process hypergraphs with more degrees of hyperedges, which counts the vertex labels that are connected by each hyperedge to generate the feature vector. Mathematically, we prove the proposed Hypergraph WL Subtree Kernel can degenerate into the typical Graph Weisfeiler-Lehman Subtree Kernel when dealing with low-order graph structures. In order to fulfill our research objectives, a comprehensive set of experiments was meticulously designed, including seven graph classification datasets and 12 hypergraph classification datasets. Results on graph classification datasets indicate that the Hypergraph WL Subtree Kernel can achieve the same performance compared with the classical Graph Weisfeiler-Lehman Subtree Kernel. Results on hypergraph classification datasets show significant improvements compared to other typical kernel-based methods, which demonstrates the effectiveness of the proposed methods. In our evaluation, we found that our proposed methods outperform the second-best method in terms of runtime, running over 80 times faster when handling complex hypergraph structures. This significant speed advantage highlights the great potential of our methods in real-world applications.
Yifan Feng 0001, Jiashu Han, Shihui Ying, Yue Gao 0002
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 Hypergraph-Based Multi-Modal Representation for Open-Set 3D Object Retrieval
abstract
The traditional 3D object retrieval (3DOR) task is under the close-set setting, which assumes the categories of objects in the retrieval stage are all seen in the training stage. Existing methods under this setting may tend to only lazily discriminate their categories, while not learning a generalized 3D object embedding. Under such circumstances, it is still a challenging and open problem in real-world applications due to the existence of various unseen categories. In this paper, we first introduce the open-set 3DOR task to expand the applications of the traditional 3DOR task. Then, we propose the Hypergraph-Based Multi-Modal Representation (HGM$^{2}$R) framework to learn 3D object embeddings from multi-modal representations under the open-set setting. The proposed framework is composed of two modules, i.e., the Multi-Modal 3D Object Embedding (MM3DOE) module and the Structure-Aware and Invariant Knowledge Learning (SAIKL) module. By utilizing the collaborative information of modalities derived from the same 3D object, the MM3DOE module is able to overcome the distinction across different modality representations and generate unified 3D object embeddings. Then, the SAIKL module utilizes the constructed hypergraph structure to model the high-order correlation among 3D objects from both seen and unseen categories. The SAIKL module also includes a memory bank that stores typical representations of 3D objects. By aligning with those memory anchors in the memory bank, the aligned embeddings can integrate the invariant knowledge to exhibit a powerful generalized capacity toward unseen categories. We formally prove that hypergraph modeling has better representative capability on data correlation than graph modeling. We generate four multi-modal datasets for the open-set 3DOR task, i.e., OS-ESB-core, OS-NTU-core, OS-MN40-core, and OS-ABO-core, in which each 3D object contains three modality representations: multi-view, point clouds, and voxel. Experiments on these four datasets show that the proposed method can significantly outperform existing methods. In particular, the proposed method outperforms the state-of-the-art by 12.12%/12.88% in terms of mAP on the OS-MN40-core/OS-ABO-core dataset, respectively. Results and visualizations demonstrate that the proposed method can effectively extract the generalized 3D object embeddings on the open-set 3DOR task and achieve satisfactory performance.
Yifan Feng 0001, Shuyi Ji, Yu-Shen Liu, Shaoyi Du, Qionghai Dai, Yue Gao 0002
IEEE Trans. Pattern Anal. Mach. Intell.6
2024 Hypergraph-Based Multi-View Action Recognition Using Event Cameras
abstract
Action recognition from video data forms a cornerstone with wide-ranging applications. Single-view action recognition faces limitations due to its reliance on a single viewpoint. In contrast, multi-view approaches capture complementary information from various viewpoints for improved accuracy. Recently, event cameras have emerged as innovative bio-inspired sensors, leading to advancements in event-based action recognition. However, existing works predominantly focus on single-view scenarios, leaving a gap in multi-view event data exploitation, particularly in challenges like information deficit and semantic misalignment. To bridge this gap, we introduceHyperMV, multi-view event-based action recognition framework. HyperMV converts discrete event data into frame-like representations and extracts view-related features using a shared convolutional network. By treating segments as vertices and constructing hyperedges using rule-based and KNN-based strategies, a multi-view hypergraph neural network that captures relationships across viewpoint and temporal features is established. The vertex attention hypergraph propagation is also introduced for enhanced feature fusion. To prompt research in this area, we present the largest multi-view event-based action dataset$\mathbf{THU}^{\mathbf{MV-EACT}}\mathbf{-50}$, comprising 50 actions from 6 viewpoints, which surpasses existing datasets by over tenfold. Experimental results show that HyperMV significantly outperforms baselines in both cross-subject and cross-view scenarios, and also exceeds the state-of-the-arts in frame-based multi-view action recognition.
Yue Gao 0002, Jiaxuan Lu, Siqi Li 0001, Shaoyi Du
IEEE Trans. Pattern Anal. Mach. Intell.1
2024 Learning Signed Hyper Surfaces for Oriented Point Cloud Normal Estimation
abstract
We propose a novel method called SHS-Net for point cloud normal estimation by learning signed hyper surfaces, which can accurately predict normals with global consistent orientation from various point clouds. Almost all existing methods estimate oriented normals through a two-stage pipeline, i.e., unoriented normal estimation and normal orientation, and each step is implemented by a separate algorithm. However, previous methods are sensitive to parameter settings, resulting in poor results from point clouds with noise, density variations and complex geometries. In this work, we introduce signed hyper surfaces (SHS), which are parameterized by multi-layer perceptron (MLP) layers, to learn to estimate oriented normals from point clouds in an end-to-end manner. The signed hyper surfaces are implicitly learned in a high-dimensional feature space where the local and global information is aggregated. Specifically, we introduce a patch encoding module and a shape encoding module to encode a 3D point cloud into a local latent code and a global latent code, respectively. Then, an attention-weighted normal prediction module is proposed as a decoder, which takes the local and global latent codes as input to predict oriented normals. Experimental results show that our algorithm outperforms the state-of-the-art methods in both unoriented and oriented normal estimation.
Qing Li 0032, Huifang Feng 0002, Kanle Shi, Yue Gao 0002, Yi Fang 0006, Yu-Shen Liu, Zhizhong Han
IEEE Trans. Pattern Anal. Mach. Intell.4
2024 Generative Variational-Contrastive Learning for Self-Supervised Point Cloud Representation
abstract
Self-supervised representation learning for 3D point clouds has attracted increasing attention. However, existing methods in the field of 3D computer vision generally use fixed embeddings to represent the latent features, and impose hard constraints on the embeddings to make the latent feature values of the positive samples converge to consistency, which limits the ability of feature extractors to generalize over different data domains. To address this issue, we propose a Generative Variational-Contrastive Learning (GVC) model, where Gaussian distribution is used to construct a continuous, smoothed representation of the latent features. A distribution constraint and cross-supervision are constructed to improve the transfer ability of the feature extractor over synthetic and real-world data. Specifically, we design a variational contrastive module to constrain the feature distribution instead of feature values corresponding to each sample in the latent space. Moreover, a generative cross-supervision module is introduced to preserve the invariance features and promote the consistency of feature distribution among positive samples. Experimental results demonstrate that GVC achieves SOTA on different downstream tasks. In particular, with only pre-training on the synthetic dataset, GVC achieves a lead of 8.4% and 14.2% when transferring to the real-world dataset in the linear classification and few-shot classification.
Bo-Hua Wang, Aixue Ye, Shaoyi Du, Yue Gao 0002
IEEE Trans. Pattern Anal. Mach. Intell.6
2024 Self-adaptive subspace representation from a geometric intuition
Lipeng Cai, Jun Shi 0004, Shaoyi Du, Yue Gao 0002, Shihui Ying
Pattern Recognit.4
2024 Multi-View Time-Series Hypergraph Neural Network for Action Recognition
abstract
Recently, action recognition has attracted considerable attention in the field of computer vision. In dynamic circumstances and complicated backgrounds, there are some problems, such as object occlusion, insufficient light, and weak correlation of human body joints, resulting in skeleton-based human action recognition accuracy being very low. To address this issue, we propose a Multi-View Time-Series Hypergraph Neural Network (MV-TSHGNN) method. The framework is composed of two main parts: the construction of a multi-view time-series hypergraph structure and the learning process of multi-view time-series hypergraph convolutions. Specifically, given the multi-view video sequence frames, we first extract the joint features of actions from different views. Then, limb components and adjacent joints spatial hypergraphs based on the joints of different views at the same time are constructed respectively, temporal hypergraphs are constructed joints of the same view at continuous times, which are established high-order semantic relationships and cooperatively generate complementary action features. After that, we design a multi-view time-series hypergraph neural network to efficiently learn the features of spatial and temporal hypergraphs, and effectively improve the accuracy of skeleton-based action recognition. To evaluate the effectiveness and efficiency of MV-TSHGNN, we conduct experiments on NTU RGB+D, NTU RGB+D 120 and imitating traffic police gestures datasets. The experimental results indicate that our proposed method model achieves the new state-of-the-art performance.
Nan Ma 0012, Zhixuan Wu, Yifan Feng 0001, Yue Gao 0002
IEEE Trans. Image Process.5
2024 Attribute Diversity Aware Community Detection on Attributed Graphs Using Three-View Graph Attention Neural Networks
abstract
Community detection is a fundamental yet important task for characterizing and understanding the structure of attributed graphs. Existing methods mainly focus on the structural tightness and attribute similarity among nodes in a community. However, grouping numerous semantically homogeneous nodes will result in information cocoons and thus reduce the robustness of community structure and the efficiency of node collaboration in real-world applications, such as recommendation systems and collaboration networks. Since nodes with closer connections tend to be more similar, finding communities with dense structures and diverse attributes poses great challenges to mining latent relationships between the graph structure and attribute distribution. To our best knowledge, very little research has been conducted to address this challenge. In this article, we propose a novel three-view graph attention neural networks (TvGANN) model to formally address the attribute diversity aware community detection problem. TvGANN reveals correlations between the graph structure and attributes distribution from the perspective of node organization, attribute co-occurrence, and the node-attribute interaction. It effectively captures structural features and attributes distribution by feeding a structural network and an attribute co-occurrence network into graph attention modules through the encoder–decoder framework. It also learns heterogeneous information by feeding a network into a meta-node attention module. Then, it fuzes the three modules and clusters the embedding representations through a Student's t -distribution approach, which iteratively refines the clustering results. The experiments show that our method not only improves the quality in dense community detection but also performs efficiently for attributed graphs.
Yang Zhang 0042, Ting Yu 0004, Shengqiang Chi, Zhen Wang 0037, Yue Gao 0002, Ji Zhang 0001
ACM Trans. Knowl. Discov. Data5
2024 Penalized Flow Hypergraph Local Clustering
abstract
In recent years, hypergraph analysis have attracted increasing attention due to their ability to model complex data correlation, with hypergraph clustering being one of the most important tasks. However, when the scale of hypergraph is large enough, clustering is difficult based on global consistency. Existing flow-based hypergraph local clustering methods have good theoretical cut improvements and runtime guarantees. However, these methods exhibit poor performance when the initial reference node set is small and are prone to causing the output set to shrink into a small subset, resulting in local minima. To address this issue, we propose the Penalized Flow Hypergraph Local Clustering(PFHLC) and provide new conductance guarantees and runtime analyses for our method. First, we use the random walk method to grow the initial seed set, and introduce the random walk information of nodes as penalized flow into the flow-based framework to optimize the output. Second, we propose a generalized objective function containing random walk information, which takes full advantage of the semi-supervised information of the target cluster to protect important nodes. This feature can avoid the local minima of previous flow-based methods. Importantly, our method is strongly-local and can run efficiently on large-scale hypergraphs. We contribute a real-world dataset and the experiments on real-world large-scale datasets show that PFHLC achieves the state-of-the-art significantly.
Yubo Zhang 0006, Chenggang Yan 0001, Zuxing Xuan, Ting Yu 0004, Ji Zhang 0001, Shihui Ying, Yue Gao 0002
IEEE Trans. Knowl. Data Eng.8
2024 Event-Based Low-Illumination Image Enhancement
abstract
Event cameras are bio-inspired vision sensors with a high dynamic range (140 dB for event camerasvs.60 dB for traditional cameras) and can be used to tackle the image degradation problem under extremely low-illumination scenarios, which is still not well-explored yet. In this article, we propose a joint framework to compose the underexposed frames and event streams captured by the event camera to reconstruct clear images with detailed textures under almost dark conditions. A residual fusion module is proposed to reduce the domain gap between event streams and frames by using the residuals of both modalities. A multi-level reconstruction loss based on the variability of the contrast distribution is proposed to reduce the perceptual errors of the output image. In addition, we construct the first real-world low-illumination image enhancement dataset (mainly under 2 lux illumination scenes), named LIE, containing event streams and frames collected under indoor and outdoor low-light scenarios together with the ground truth clear images. Experimental results on our LIE dataset demonstrate that our proposed method could achieve significant improvements compared with existing methods.
Yu Jiang 0006, Yuehang Wang, Siqi Li 0001, Yongji Zhang, Minghao Zhao 0003, Yue Gao 0002
IEEE Trans. Multim.6
2023 Learning Deep Hierarchical Features with Spatial Regularization for One-Class Facial Expression Recognition
abstract
Existing methods on facial expression recognition (FER) are mainly trained in the setting when multi-class data is available. However, to detect the alien expressions that are absent during training, this type of methods cannot work. To address this problem, we develop a Hierarchical Spatial One Class Facial Expression Recognition Network (HS-OCFER) which can construct the decision boundary of a given expression class (called normal class) by training on only one-class data. Specifically, HS-OCFER consists of three novel components. First, hierarchical bottleneck modules are proposed to enrich the representation power of the model and extract detailed feature hierarchy from different levels. Second, multi-scale spatial regularization with facial geometric information is employed to guide the feature extraction towards emotional facial representations and prevent the model from overfitting extraneous disturbing factors. Third, compact intra-class variation is adopted to separate the normal class from alien classes in the decision space. Extensive evaluations on 4 typical FER datasets from both laboratory and wild scenarios show that our method consistently outperforms state-of-the-art One-Class Classification (OCC) approaches.
Bingjun Luo, Sicheng Zhao, Xibin Zhao, Yue Gao 0002
AAAI7
2023 SHS-Net: Learning Signed Hyper Surfaces for Oriented Normal Estimation of Point Clouds
abstract
We propose a novel method called SHS-Net for oriented normal estimation of point clouds by learning signed hyper surfaces, which can accurately predict normals with global consistent orientation from various point clouds. Almost all existing methods estimate oriented normals through a two-stage pipeline, i.e., unoriented normal estimation and normal orientation, and each step is implemented by a separate algorithm. However, previous methods are sensitive to parameter settings, resulting in poor results from point clouds with noise, density variations and complex geometries. In this work, we introduce signed hyper surfaces (SHS), which are parameterized by multi-layer perceptron (MLP) layers, to learn to estimate oriented normals from point clouds in an end-to-end manner. The signed hyper surfaces are implicitly learned in a high-dimensional feature space where the local and global information is aggregated. Specifically, we introduce a patch encoding module and a shape encoding module to encode a 3D point cloud into a local latent code and a global latent code, respectively. Then, an attention-weighted normal prediction module is proposed as a decoder, which takes the local and global latent codes as input to predict oriented normals. Experimental results show that our SHS-Net outperforms the state-of-the-art methods in both unoriented and oriented normal estimation on the widely used benchmarks. The code, data and pretrained models are available at https://github.com/LeoQLi/SHS-Net.
Qing Li 0032, Huifang Feng 0002, Kanle Shi, Yue Gao 0002, Yi Fang 0006, Yu-Shen Liu, Zhizhong Han
CVPR4
2023 LP-DIF: Learning Local Pattern-Specific Deep Implicit Function for 3D Objects and Scenes
abstract
Deep Implicit Function (DIF) has gained much popularity as an efficient 3D shape representation. To capture geometry details, current mainstream methods divide 3D shapes into local regions and then learn each one with a local latent code via a decoder. Such local methods can capture more local details due to less diversity among local regions than global shapes. Although the diversity of local regions has been decreased compared to global approaches, the diversity in different local regions still poses a challenge in learning an implicit function when treating all regions equally using only a single decoder. What is worse, these local regions often exhibit imbalanced distributions, where certain regions have significantly fewer observations. This leads that fine geometry details could not be preserved well. To solve this problem, we propose a novel Local Pattern-specific Implicit Function, named LP-DIF, to represent a shape with clusters of local regions and multiple decoders, where each decoder only focuses on one cluster of local regions which share a certain pattern. Specifically, we first extract local codes for all regions, and then cluster them into multiple groups in the latent space, where similar regions sharing a common pattern fall into one group. After that, we train multiple decoders for mining local patterns of different groups, which simplifies the learning of fine geometric details by reducing the diversity of local regions seen by each decoder. To further alleviate the data-imbalance problem, we introduce a region re-weighting module to each pattern-specific decoder using a kernel density estimator, which dynamically re-weights the regions during learning. Our LP-DIF can restore more geometry details, and thus improve the quality of 3D reconstruction. Experiments demonstrate that our method can achieve the state-of-the-art performance over previous methods. Code is available at https://github.com/gtyxyz/lpdif.
Meng Wang 0001, Yu-Shen Liu, Yue Gao 0002, Kanle Shi, Yi Fang 0006, Zhizhong Han
CVPR3
2023 Variance-Aware Bi-Attention Expression Transformer for Open-Set Facial Expression Recognition in the Wild
abstract
Despite the great accomplishments of facial expression recognition (FER) models in closed-set scenarios, they still lack open-world robustness when it comes to handling unknown samples. To address the demands of operating in an open environment, open-set FER models should improve their performance in rejecting unknown samples while maintaining their efficiency in recognizing known expressions. With this goal in mind, we propose an open-set FER framework named Variance-Aware Bi-Attention Expression Transformer (VBExT), which enhances conventional closed-set FER models with open-world robustness for unknown samples. Specifically, to make full use of the expression representation capabilities of learned features, we introduce a bi-attention feature augmentation mechanism that learns the important regions and integrates the hierarchical features extracted by the emotional CNN backbone. We also propose a variance-aware distribution modeling method that adapts to the diverse distribution of different expression classes in the open environment, thereby enhancing the detection ability of unknown expressions. Additionally, we have constructed a Fine-Grained Light Facial Expression dataset that includes 30 different light brightnesses to better validate the efficiency of VBExT. Extensive experiments and ablation studies show that VBExT significantly improves the performance of open-set FER and achieves state-of-the-art results on CFEE (lab, basic), RAF-DB (wild, basic+compound), and FGL-FE (multiple light brightnesses, basic).
Bingjun Luo, Jinghang Tan, Xibin Zhao, Yue Gao 0002
ACM Multimedia6
2023 NeuralGF: Unsupervised Point Normal Estimation by Learning Neural Gradient Function
abstract
Normal estimation for 3D point clouds is a fundamental task in 3D geometry processing. The state-of-the-art methods rely on priors of fitting local surfaces learned from normal supervision. However, normal supervision in benchmarks comes from synthetic shapes and is usually not available from real scans, thereby limiting the learned priors of these methods. In addition, normal orientation consistency across shapes remains difficult to achieve without a separate post-processing procedure. To resolve these issues, we propose a novel method for estimating oriented normals directly from point clouds without using ground truth normals as supervision. We achieve this by introducing a new paradigm for learning neural gradient functions, which encourages the neural network to fit the input point clouds and yield unit-norm gradients at the points. Specifically, we introduce loss functions to facilitate query points to iteratively reach the moving targets and aggregate onto the approximated surface, thereby learning a global surface representation of the data. Meanwhile, we incorporate gradients into the surface approximation to measure the minimum signed deviation of queries, resulting in a consistent gradient field associated with the surface. These techniques lead to our deep unsupervised oriented normal estimator that is robust to noise, outliers and density variations. Our excellent results on widely used benchmarks demonstrate that our method can learn more accurate normals for both unoriented and oriented normal estimation tasks than the latest methods. The source code and pre-trained model are publicly available.
Qing Li 0032, Huifang Feng 0002, Kanle Shi, Yue Gao 0002, Yi Fang 0006, Yu-Shen Liu, Zhizhong Han
NeurIPS4
2023 GAME: GAussian Mixture Error-based meta-learning architecture
Jinhe Dong, Jun Shi 0004, Yue Gao 0002, Shihui Ying
Neural Comput. Appl.3
2023 Generating Hypergraph-Based High-Order Representations of Whole-Slide Histopathological Images for Survival Prediction
abstract
Patient survival prediction based on gigapixel whole-slide histopathological images (WSIs) has become increasingly prevalent in recent years. A key challenge of this task is achieving an informative survival-specific global representation from those WSIs with highly complicated data correlation. This article proposes a multi-hypergraph based learning framework, called "HGSurvNet," to tackle this challenge. HGSurvNet achieves an effective high-order global representation of WSIs via multilateral correlation modeling in multiple spaces and a general hypergraph convolution network. It has the ability to alleviate over-fitting issues caused by the lack of training data by using a new convolution structure called hypergraph max-mask convolution. Extensive validation experiments were conducted on three widely-used carcinoma datasets: Lung Squamous Cell Carcinoma (LUSC), Glioblastoma Multiforme (GBM), and National Lung Screening Trial (NLST). Quantitative analysis demonstrated that the proposed method consistently outperforms state-of-the-art methods, coupled with the Bayesian Concordance Readjust loss. We also demonstrate the individual effectiveness of each module of the proposed framework and its application potential for pathology diagnosis and reporting empowered by its interpretability potential.
Donglin Di, Changqing Zou, Yifan Feng 0001, Rongrong Ji, Qionghai Dai, Yue Gao 0002
IEEE Trans. Pattern Anal. Mach. Intell.7
2023 HGNN+: General Hypergraph Neural Networks
abstract
Graph Neural Networks have attracted increasing attention in recent years. However, existing GNN frameworks are deployed based upon simple graphs, which limits their applications in dealing with complex data correlation of multi-modal/multi-type data in practice. A few hypergraph-based methods have recently been proposed to address the problem of multi-modal/multi-type data correlation by directly concatenating the hypergraphs constructed from each single individual modality/type, which is difficult to learn an adaptive weight for each modality/type. In this paper, we extend the original conference version HGNN, and introduce a general high-order multi-modal/multi-type data correlation modeling framework called HGNN$^+$to learn an optimal representation in a single hypergraph based framework. It is achieved by bridging multi-modal/multi-type data and hyperedge with hyperedge groups. Specifically, in our method, hyperedge groups are first constructed to represent latent high-order correlations in each specific modality/type with explicit or implicit graph structures. An adaptive hyperedge group fusion strategy is then used to effectively fuse the correlations from different modalities/types in a unified hypergraph. After that a new hypergraph convolution scheme performed in spatial domain is used to learn a general data representation for various tasks. We have evaluated this framework on several popular datasets and compared it with recent state-of-the-art methods. The comprehensive evaluations indicate that the proposed HGNN$^+$framework can consistently outperform existing methods with a significant margin, especially when modeling implicit data correlations. We also release a toolbox called THU-DeepHypergraph for the proposed framework, which can be used for various of applications, such as data classification, retrieval and recommendation.
Yue Gao 0002, Yifan Feng 0001, Shuyi Ji, Rongrong Ji
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 SuperFast: 200× Video Frame Interpolation via Event Camera
abstract
Traditional frame-based video frame interpolation (VFI) methods rely on the linear motion assumption and brightness invariance assumption, which may lead to fatal errors confronting the scenarios with high-speed motions. To tackle the above challenge, inspired by the advantages of event cameras on asynchronously recording brightness changes at each pixel, we propose a Fast-Slow joint synthesis framework for event-enhanced high-speed video frame interpolation, named SuperFast, in this paper, which can generate high frame rate (5000 FPS, 200× faster) video from the input low frame rate (25 FPS) video and the corresponding event stream. In our framework, the task is divided into two sub-tasks, i.e., video frame interpolation for the contents with and without high-speed motions, which are tackled by two corresponding branches, i.e., the fast synthesis pathway and the slow synthesis pathway. The fast synthesis pathway leverages a spiking neural network to encode the input event stream, and combines boundary frames to generate intermediate results through synthesis and refinement, targeting on contents with high-speed motions. The slow synthesis pathway stacks the two input boundary frames and the event stream to synthesize intermediate results, focusing on relatively slow-motion contents. Finally, a fusion module with a comparison loss is utilized to generate the final video frame interpolation results. We also build a hybrid visual acquisition system containing an event camera and a high frame rate camera, and collect the first 5000 FPS High-Speed Event-enhanced Video frame Interpolation (THU[Formula: see text]) dataset. To evaluate the performance of our proposed framework, we have conducted experiments on our THU[Formula: see text] dataset and the existing HS-ERGB dataset. Experimental results demonstrate that our proposed framework can achieve state-of-the-art 200× video frame interpolation performance under high-speed motion scenarios.
Yue Gao 0002, Siqi Li 0001, Yandong Guo, Qionghai Dai
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 Action Recognition and Benchmark Using Event Cameras
abstract
Recent years have witnessed remarkable achievements in video-based action recognition. Apart from traditional frame-based cameras, event cameras are bio-inspired vision sensors that only record pixel-wise brightness changes rather than the brightness value. However, little effort has been made in event-based action recognition, and large-scale public datasets are also nearly unavailable. In this paper, we propose an event-based action recognition framework calledEV-ACT. The Learnable Multi-Fused Representation (LMFR) is first proposed to integrate multiple event information in a learnable manner. The LMFR with dual temporal granularity is fed into the event-based slow-fast network for the fusion of appearance and motion features. A spatial-temporal attention mechanism is introduced to further enhance the learning capability of action recognition. To prompt research in this direction, we have collected the largest event-based action recognition benchmark namedTHUE-ACT-50and the accompanyingTHUE-ACT-50-CHLdataset under challenging environments, including a total of over 12,830 recordings from 50 action categories, which is over 4 times the size of the previous largest dataset. Experimental results show that our proposed framework could achieve improvements of over 14.5%, 7.6%, 11.2%, and 7.4% compared to previous works on four benchmarks. We have also deployed our proposed EV-ACT framework on a mobile platform to validate its practicality and efficiency.
Yue Gao 0002, Jiaxuan Lu, Siqi Li 0001, Nan Ma 0012, Shaoyi Du, Qionghai Dai
IEEE Trans. Pattern Anal. Mach. Intell.1
2023 High-Order Correlation-Guided Slide-Level Histology Retrieval With Self-Supervised Hashing
abstract
Histopathological Whole Slide Images (WSIs) play a crucial role in cancer diagnosis. It is of significant importance for pathologists to search for images sharing similar content with the query WSI, especially in the case-based diagnosis. While slide-level retrieval could be more intuitive and practical in clinical applications, most methods are designed for patch-level retrieval. A few recently unsupervised slide-level methods only focus on integrating patch features directly, without perceiving slide-level information, and thus severely limits the performance of WSI retrieval. To tackle the issue, we propose a High-Order Correlation-Guided Self-Supervised Hashing-Encoding Retrieval (HSHR) method. Specifically, we train an attention-based hash encoder with slide-level representation in a self-supervised manner, enabling it to generate more representative slide-level hash codes of cluster centers and assign weights for each. These optimized and weighted codes are leveraged to establish a similarity-based hypergraph, in which a hypergraph-guided retrieval module is adopted to explore high-order correlations in the multi-pairwise manifold to conduct WSI retrieval. Extensive experiments on multiple TCGA datasets with over 24,000 WSIs spanning 30 cancer subtypes demonstrate that HSHR achieves state-of-the-art performance compared with other unsupervised histology WSI retrieval methods.
Shengrui Li, Jun Zhang 0018, Ting Yu 0004, Ji Zhang 0001, Yue Gao 0002
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 Structure Evolution on Manifold for Graph Learning
abstract
Graph has been widely used in various applications, while how to optimize the graph is still an open question. In this paper, we propose a framework to optimize the graph structure via structure evolution on graph manifold. We first define the graph manifold and search the best graph structure on this manifold. Concretely, associated with the data features and the prediction results of a given task, we define a graph energy to measure how the graph fits the graph manifold from an initial graph structure. The graph structure then evolves by minimizing the graph energy. In this process, the graph structure can be evolved on the graph manifold corresponding to the update of the prediction results. Alternatively iterating these two processes, both the graph structure and the prediction results can be updated until converge. It achieves the suitable structure for graph learning without searching all hyperparameters. To evaluate the performance of the proposed method, we have conducted experiments on eight datasets and compared with the recent state-of-the-art methods. Experiment results demonstrate that our method outperforms the state-of-the-art methods in both transductive and inductive settings.
Hai Wan, Xinwei Zhang 0012, Yubo Zhang 0006, Xibin Zhao, Shihui Ying, Yue Gao 0002
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 STORM: Structure-Based Overlap Matching for Partial Point Cloud Registration
abstract
Partial point cloud registration aims to transform partial scans into a common coordinate system. It is an important preprocessing step to generate complete 3D shapes. Although previous registration methods have made great progress in recent decades, traditional registration methods, such as Iterative Closest Point (ICP) and its variants, all these methods highly depend on the sufficient overlaps between two point clouds, because they cannot distinguish outlier correspondences. Note that the overlap between point clouds could always be small, which limits the application of these methods. To tackle this problem, we present a StrucTure-based OveRlap Matching (STORM) method for partial point cloud registration. In our method, an overlap prediction module with differentiable sampling is designed to detect points in overlap utilizing structure information, and facilitates exact partial correspondence generation, which is based on discriminative pointwise feature similarity. The pointwise features which contain effective structural information are extracted by graph-based methods. Experimental results and comparison with state-of-the-art methods demonstrate that STORM can achieve better performance. Moreover, most registration methods perform worse when the overlap ratio decreases, while STORM can still achieve satisfactory performance when the overlap ratio is small.
Chenggang Yan 0001, Yutong Feng, Shaoyi Du, Qionghai Dai, Yue Gao 0002
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 Hunter: Exploring High-Order Consistency for Point Cloud Registration With Severe Outliers
abstract
After decades of investigation, point cloud registration is still a challenging task in practice, especially when the correspondences are contaminated by a large number of outliers. It may result in a rapidly decreasing probability of generating a hypothesis close to the true transformation, leading to the failure of point cloud registration. To tackle this problem, we propose a transformation estimation method, named Hunter, for robust point cloud registration with severe outliers. The core of Hunter is to design a global-to-local exploration scheme to robustly find the correct correspondences. The global exploration aims to exploit guided sampling to generate promising initial alignments. To this end, a hypergraph-based consistency reasoning module is introduced to learn the high-order consistency among correct correspondences, which is able to yield a more distinct inlier cluster that facilitates the generation of all-inlier hypotheses. Moreover, we propose a preference-based local exploration module that exploits the preference information of top- k promising hypotheses to find a better transformation. This module can efficiently obtain multiple reliable transformation hypotheses by using a multi-initialization searching strategy. Finally, we present a distance-angle based hypothesis selection criterion to choose the most reliable transformation, which can avoid selecting symmetrically aligned false transformations. Experimental results on simulated, indoor, and outdoor datasets, demonstrate that Hunter can achieve significant superiority over the state-of-the-art methods, including both learning-based and traditional methods (as shown in Fig. 1). Moreover, experimental results also indicate that Hunter can achieve more stable performance compared with all other methods with severe outliers.
Runzhao Yao, Shaoyi Du, Wenting Cui, Aixue Ye, Hongbo Zhang 0004, Yue Gao 0002
IEEE Trans. Pattern Anal. Mach. Intell.8
2023 Graph Learning on Millions of Data in Seconds: Label Propagation Acceleration on Graph Using Data Distribution
abstract
Graph-based semi-supervised learning methods have been used in a wide range of real-world applications, e.g., from social relationship mining to multimedia classification and retrieval. However, existing methods are limited along with high computational complexity or not facilitating incremental learning, which may not be powerful to deal with large-scale data, whose scale may continuously increase, in real world. This paper proposes a new method called Data Distribution Based Graph Learning (DDGL) for semi-supervised learning on large-scale data. This method can achieve a fast and effective label propagation and supports incremental learning. The key motivation is to propagate the labels along smaller-scale data distribution model parameters, rather than directly dealing with the raw data as previous methods, which accelerate the data propagation significantly. It also improves the prediction accuracy since the loss of structure information can be alleviated in this way. To enable incremental learning, we propose an adaptive graph updating strategy which can update the model when there is distribution bias between new data and the already seen data. We have conducted comprehensive experiments on multiple datasets with sample sizes increasing from seven thousand to five million. Experimental results on the classification task on large-scale data demonstrate that our proposed DDGL method improves the classification accuracy by a large margin while consuming much less time compared to state-of-the-art methods.
Yubo Zhang 0006, Shuyi Ji, Changqing Zou, Xibin Zhao, Shihui Ying, Yue Gao 0002
IEEE Trans. Pattern Anal. Mach. Intell.6
2023 Cost-Sensitive Hypergraph Learning With F-Measure Optimization
abstract
The imbalanced issue among data is common in many machine-learning applications, where samples from one or more classes are rare. To address this issue, many imbalanced machine-learning methods have been proposed. Most of these methods rely on cost-sensitive learning. However, we note that it is infeasible to determine the precise cost values even with great domain knowledge for those cost-sensitive machine-learning methods. So in this method, due to the superiority of F-measure on evaluating the performance of imbalanced data classification, we employ F-measure to calculate the cost information and propose a cost-sensitive hypergraph learning method with F-measure optimization to solve the imbalanced issue. In this method, we employ the hypergraph structure to explore the high-order relationships among the imbalanced data. Based on the constructed hypergraph structure, we optimize the cost value with F-measure and further conduct cost-sensitive hypergraph learning with the optimized cost information. The comprehensive experiments validate the effectiveness of the proposed method.
Nan Wang 0015, Ruozhou Liang, Xibin Zhao, Yue Gao 0002
IEEE Trans. Cybern.4
2023 Exploring High-Order Spatio-Temporal Correlations From Skeleton for Person Re-Identification
abstract
Person re- identification (Re-ID) has become a hot research topic due to its widespread applications. Conducting person Re-ID in video sequences is a practical requirement, in which the crucial challenge is how to pursue a robust video representation based on spatial and temporal features. However, most of the previous methods only consider how to integrate part-level features in the spatio-temporal range, while how to model and generate the part-correlations is little exploited. In this paper, we propose a skeleton-based dynamic hypergraph framework, namely Skeletal Temporal Dynamic Hypergraph Neural Network (ST-DHGNN) for person Re-ID, which resorts to modeling the high-order correlations among various body parts based on a time series of skeletal information. Specifically, multi-shape and multi-scale patches are heuristically cropped from feature maps, constituting spatial representations in different frames. A joint-centered hypergraph and a bone-centered hypergraph are constructed in parallel from multiple body parts (i.e., head, trunk, and legs) with spatio-temporal multi-granularity in the entire video sequence, in which the graph vertices representing regional features and hyperedges denoting relationships. Dynamic hypergraph propagation containing the re- planning module and the hyperedge elimination module is proposed to better integrate features among vertices. Feature aggregation and attention mechanisms are also adopted to obtain a better video representation for person Re-ID. Experiments show that the proposed method performs significantly better than the state-of-the-art on three video-based person Re-ID datasets, including iLIDS-VID, PRID-2011, and MARS.
Jiaxuan Lu, Hai Wan, Xibin Zhao, Nan Ma 0012, Yue Gao 0002
IEEE Trans. Image Process.6
2023 Knowledge Conditioned Variational Learning for One-Class Facial Expression Recognition
abstract
The openness of application scenarios and the difficulties of data collection make it impossible to prepare all kinds of expressions for training. Hence, detecting expression absent during the training (called alien expression) is important to enhance the robustness of the recognition system. So in this paper, we propose a facial expression recognition (FER) model, named OneExpressNet, to quantify the probability that a test expression sample belongs to the distribution of training data. The proposed model is based on variational auto-encoder and enjoys several merits. First, different from conventional one class classification protocol, OneExpressNet transfers the useful knowledge from the related domain as a constraint condition of the target distribution. By doing so, OneExpressNet will pay more attention to the descriptive region for FER. Second, features from both source and target tasks will aggregate after constructing a skip connection between the encoder and decoder. Finally, to further separate alien expression from training expression, empirical compact variation loss is jointly optimized, so that training expression will concentrate on the compact manifold of feature space. The experimental results show that our method can achieve state-of-the-art results in one class facial expression recognition on small-scale lab-controlled datasets including CFEE and KDEF, and large-scale in-the-wild datasets including RAF-DB and ExpW.
Bingjun Luo, Xibin Zhao, Yue Gao 0002
IEEE Trans. Image Process.6
2023 Adaptive Hypergraph Auto-Encoder for Relational Data Clustering
abstract
The embedded representation and clustering tasks both play important roles in relational data analysis and mining. Traditional methods mainly employ graph structure to describe relational data, but intuitive pairwise connections among nodes are insufficient to model high-order data in the real-world, such as the relations between proteins and polypeptide chains. Hypergraphs are a generalization of graphs, and hypergraphs can well model high-order data. When modeling relational data in the real world, hypergraphs are often accompanied by node attributes, i.e. attributed hypergraphs. Besides this, how to integrate the structural information and attribute information appropriately is another important task, while has not been investigated systematically. In this paper, we propose Adaptive Hypergraph Auto-Encoder(AHGAE) to learn node embeddings in low-dimensional space. Our method can utilize the high-order relation to generate embedding for clustering. It is composed of two procedures, i.e. the adaptive hypergraph Laplacian smoothing filter and the relational reconstruction auto-encoder. It has the advantage of integrating more complex data relations compared with graph-based methods, which leads to better modeling and clustering performance. The proposed method has been evaluated on hypergraph datasets and benchmark graph datasets. Experimental results and comparison with the state-of-the-art methods have demonstrated the effectiveness of our proposed method.
Youpeng Hu, Xunkai Li, Chenggang Yan 0001, Jian Yin 0003, Yue Gao 0002
IEEE Trans. Knowl. Data Eng.8
2023 MISSU: 3D Medical Image Segmentation via Self-Distilling TransUNet
abstract
U-Nets have achieved tremendous success in medical image segmentation. Nevertheless, it may have limitations in global (long-range) contextual interactions and edge-detail preservation. In contrast, the Transformer module has an excellent ability to capture long-range dependencies by leveraging the self-attention mechanism into the encoder. Although the Transformer module was born to model the long-range dependency on the extracted feature maps, it still suffers high computational and spatial complexities in processing high-resolution 3D feature maps. This motivates us to design an efficient Transformer-based UNet model and study the feasibility of Transformer-based network architectures for medical image segmentation tasks. To this end, we propose to self-distill a Transformer-based UNet for medical image segmentation, which simultaneously learns global semantic information and local spatial-detailed features. Meanwhile, a local multi-scale fusion block is first proposed to refine fine-grained details from the skipped connections in the encoder by the main CNN stem through self-distillation, only computed during training and removed at inference with minimal overhead. Extensive experiments on BraTS 2019 and CHAOS datasets show that our MISSU achieves the best performance over previous state-of-the-art methods. Code and models are available at: https://github.com/wangn123/MISSU.git.
Nan Wang 0027, Shaohui Lin, Xiaoxiao Li 0001, Ke Li 0015, Yunhang Shen, Yue Gao 0002, Lizhuang Ma
IEEE Trans. Medical Imaging6
2023 Image Matting With Deep Gaussian Process
abstract
We observe a common characteristic between the classical propagation-based image matting and the Gaussian process (GP)-based regression. The former produces closer alpha matte values for pixels associated with a higher affinity, while the outputs regressed by the latter are more correlated for more similar inputs. Based on this observation, we reformulate image matting as GP and find that this novel matting-GP formulation results in a set of attractive properties. First, it offers an alternative view on and approach to propagation-based image matting. Second, an application of kernel learning in GP brings in a novel deep matting-GP technique, which is pretty powerful for encapsulating the expressive power of deep architecture on the image relative to its matting. Third, an existing scalable GP technique can be incorporated to further reduce the computational complexity to$\mathcal {O}(n)$from$\mathcal {O}(n^{3})$of many conventional matting propagation techniques. Our deep matting-GP provides an attractive strategy toward addressing the limit of widespread adoption of deep learning techniques to image matting for which a sufficiently large labeled dataset is lacking. A set of experiments on both synthetically composited images and real-world images show the superiority of the deep matting-GP to not only the classical propagation-based matting techniques but also modern deep learning-based approaches.
Yuanjie Zheng, Yunshuai Yang, Tongtong Che, Sujuan Hou, Wenhui Huang 0002, Yue Gao 0002, Ping Tan 0002
IEEE Trans. Neural Networks Learn. Syst.6
2022 3D Room Layout Estimation from a Cubemap of Panorama Image via Deep Manhattan Hough Transform
Chao Wen 0001, Zhou Xue, Yue Gao 0002
ECCV (1)4
2022 Rethinking Supervised Pre-Training for Better Downstream Transferring
Yutong Feng, Jianwen Jiang, Mingqian Tang, Rong Jin 0001, Yue Gao 0002
ICLR5
2022 Grow and Merge: A Unified Framework for Continuous Categories Discovery
abstract
Although a number of studies are devoted to novel category discovery, most of them assume a static setting where both labeled and unlabeled data are given at once for finding new categories. In this work, we focus on the application scenarios where unlabeled data are continuously fed into the category discovery system. We refer to it as the {\bf Continuous Category Discovery} ({\bf CCD}) problem, which is significantly more challenging than the static setting. A common challenge faced by novel category discovery is that different sets of features are needed for classification and category discovery: class discriminative features are preferred for classification, while rich and diverse features are more suitable for new category mining. This challenge becomes more severe for dynamic setting as the system is asked to deliver good performance for known classes over time, and at the same time continuously discover new classes from unlabeled data. To address this challenge, we develop a framework of {\bf Grow and Merge} ({\bf GM}) that works by alternating between a growing phase and a merge phase: in the growing phase, it increases the diversity of features through a continuous self-supervised learning for effective category mining, and in the merging phase, it merges the grown model with a static one to ensure satisfying performance for known classes. Our extensive studies verify that the proposed GM framework is significantly more effective than the state-of-the-art approaches for continuous category discovery.
Xinwei Zhang 0012, Jianwen Jiang, Yutong Feng, Zhi-Fan Wu, Xibin Zhao, Hai Wan, Mingqian Tang, Rong Jin 0001, Yue Gao 0002
NeurIPS9
2022 SHREC'22 track: Open-Set 3D Object Retrieval
Yifan Feng 0001, Yue Gao 0002, Xibin Zhao, Yandong Guo, Nihar Bagewadi, Nhat-Tan Bui, Hieu Dao, Shankar Gangisetty, Ripeng Guan, Xie Han 0001, Cong Hua, Chidambar Hunakunti, Yu Jiang 0006, Shichao Jiao, Yuqi Ke, Liqun Kuang, Anan Liu, Dinh-Huan Nguyen, Hai-Dang Nguyen, Weizhi Nie, Bang-Dang Pham, Karthik Raikar, Qingmei Tang, Minh-Triet Tran, Jialong Wan, Chenggang Yan 0001, Haoxuan You, Difei Zhu
Comput. Graph.2
2022 Deepwalk-aware graph convolutional networks
Taisong Jin, Huaqiang Dai, Liujuan Cao, Baochang Zhang 0001, Feiyue Huang, Yue Gao 0002, Rongrong Ji
Sci. China Inf. Sci.6
2022 Attention Mechanism Based on Improved Spatial-Temporal Convolutional Neural Networks for Traffic Police Gesture Recognition
abstract
Human action recognition has attracted extensive research efforts in recent years, in which traffic police gesture recognition is important for self-driving vehicles. One of the crucial challenges in this task is how to find a representation method based on spatial-temporal features. However, existing methods performed poorly in spatial and temporal information fusion, and how to extract features of traffic police gestures has not been well researched. This paper proposes an attention mechanism based on the improved spatial-temporal convolutional neural network (AMSTCNN) for traffic police gesture recognition. This method focuses on the action part of traffic police and uses the correlation between spatial and temporal features to recognize traffic police gestures, so as to ensure that traffic police gesture information is not lost. Specifically, AMSTCNN integrates spatial and temporal information, uses weight matching to pay more attention to the region where human action occurs, and extracts region proposals of the image. Finally, we use Softmax to classify actions after spatial-temporal feature fusion. AMSTCNN can strongly make use of the spatial-temporal information of videos and select effective features to reduce computation. Experiments on AVA and the Chinese traffic police gesture datasets show that our method is superior to several state-of-the-art methods.
Zhixuan Wu, Nan Ma 0002, Yue Gao 0002, Yongqiang Yao
Int. J. Pattern Recognit. Artif. Intell.3
2022 Search-based cost-sensitive hypergraph learning for anomaly detection
Nan Wang 0015, Yubo Zhang 0006, Xibin Zhao, Yingli Zheng, Boya Zhou, Yue Gao 0002
Inf. Sci.7
2022 Heterogeneous Hypergraph Variational Autoencoder for Link Prediction
abstract
Link prediction aims at inferring missing links or predicting future ones based on the currently observed network. This topic is important for many applications such as social media, bioinformatics and recommendation systems. Most existing methods focus on homogeneous settings and consider only low-order pairwise relations while ignoring either the heterogeneity or high-order complex relations among different types of nodes, which tends to lead to a sub-optimal embedding result. This paper presents a method named Heterogeneous Hypergraph Variational Autoencoder (HeteHG-VAE) for link prediction in heterogeneous information networks (HINs). It first maps a conventional HIN to a heterogeneous hypergraph with a certain kind of semantics to capture both the high-order semantics and complex relations among nodes, while preserving the low-order pairwise topology information of the original HIN. Then, deep latent representations of nodes and hyperedges are learned by a Bayesian deep generative framework from the heterogeneous hypergraph in an unsupervised manner. Moreover, a hyperedge attention module is designed to learn the importance of different types of nodes in each hyperedge. The major merit of HeteHG-VAE lies in its ability of modeling multi-level relations in heterogeneous settings. Extensive experiments on real-world datasets demonstrate the effectiveness and efficiency of the proposed method.
Haoyi Fan, Fengbin Zhang, Yuxuan Wei, Changqing Zou, Yue Gao 0002, Qionghai Dai
IEEE Trans. Pattern Anal. Mach. Intell.6
2022 Hypergraph Learning: Methods and Practices
abstract
Hypergraph learning is a technique for conducting learning on a hypergraph structure. In recent years, hypergraph learning has attracted increasing attention due to its flexibility and capability in modeling complex data correlation. In this paper, we first systematically review existing literature regarding hypergraph generation, including distance-based, representation-based, attribute-based, and network-based approaches. Then, we introduce the existing learning methods on a hypergraph, including transductive hypergraph learning, inductive hypergraph learning, hypergraph structure updating, and multi-modal hypergraph learning. After that, we present a tensor-based dynamic hypergraph representation and learning framework that can effectively describe high-order correlation in a hypergraph. To study the effectiveness and efficiency of hypergraph generation and learning methods, we conduct comprehensive evaluations on several typical applications, including object and action recognition, Microblog sentiment prediction, and clustering. In addition, we contribute a hypergraph learning development toolkit called THU-HyperG.
Yue Gao 0002, Zizhao Zhang 0003, Haojie Lin, Xibin Zhao, Shaoyi Du, Changqing Zou
IEEE Trans. Pattern Anal. Mach. Intell.1
2022 View-Aware Geometry-Structure Joint Learning for Single-View 3D Shape Reconstruction
abstract
Reconstructing a 3D shape from a single-view image using deep learning has become increasingly popular recently. Most existing methods only focus on reconstructing the 3D shape geometry based on image constraints. The lack of explicit modeling of structure relations among shape parts yields low-quality reconstruction results for structure-rich man-made shapes. In addition, conventional 2D-3D joint embedding architecture for image-based 3D shape reconstruction often omits the specific view information from the given image, which may lead to degraded geometry and structure reconstruction. We address these problems by introducing VGSNet, an encoder-decoder architecture for view-aware joint geometry and structure learning. The key idea is to jointly learn a multimodal feature representation of 2D image, 3D shape geometry and structure so that both geometry and structure details can be reconstructed from a single-view image. To this end, we explicitly represent 3D shape structures as part relations and employ image supervision to guide the geometry and structure reconstruction. Trained with pairs of view-aligned images and 3D shapes, the VGSNet implicitly encodes the view-aware shape information in the latent feature space. Qualitative and quantitative comparisons with the state-of-the-art baseline methods as well as ablation studies demonstrate the effectiveness of the VGSNet for structure-aware single-view 3D shape reconstruction.
Xuancheng Zhang, Rui Ma 0011, Changqing Zou, Xibin Zhao, Yue Gao 0002
IEEE Trans. Pattern Anal. Mach. Intell.6
2022 Plenty is Plague: Fine-Grained Learning for Visual Question Answering
abstract
Visual Question Answering (VQA) has attracted extensive research focus recently. Along with the ever-increasing data scale and model complexity, the enormous training cost has become an emerging challenge for VQA. In this article, we show such a massive training cost is indeed plague. In contrast, a fine-grained design of the learning paradigm can be extremely beneficial in terms of both training efficiency and model accuracy. In particular, we argue that there exist two essential and unexplored issues in the existing VQA training paradigm that randomly samples data in each epoch, namely, the "difficulty diversity" and the "label redundancy". Concretely, "difficulty diversity" refers to the varying difficulty levels of different question types, while "label redundancy" refers to the redundant and noisy labels contained in individual question type. To tackle these two issues, in this article we propose a fine-grained VQA learning paradigm with an actor-critic based learning agent, termed FG-A1C. Instead of using all training data from scratch, FG-A1C includes a learning agent that adaptively and intelligently schedules the most difficult question types in each training epoch. Subsequently, two curriculum learning based schemes are further designed to identify the most useful data to be learned within each inidividual question type. We conduct extensive experiments on the VQA2.0 and VQA-CP v2 datasets, which demonstrate the significant benefits of our approach. For instance, on VQA-CP v2, with less than 75 percent of the training data, our learning paradigms can help the model achieves better performance than using the whole dataset. Meanwhile, we also shows the effectivenesss of our method in guiding data labeling. Finally, the proposed paradigm can be seamlessly integrated with any cutting-edge VQA models, without modifying their structures.
Yiyi Zhou, Rongrong Ji, Xiaoshuai Sun, Jinsong Su, Deyu Meng, Yue Gao 0002, Chunhua Shen
IEEE Trans. Pattern Anal. Mach. Intell.6
2022 Kullback-Leibler Divergence Metric Learning
abstract
The Kullback-Leibler divergence (KLD), which is widely used to measure the similarity between two distributions, plays an important role in many applications. In this article, we address the KLD metric-learning task, which aims at learning the best KLD-type metric from the distributions of datasets. Concretely, first, we extend the conventional KLD by introducing a linear mapping and obtain the best KLD to well express the similarity of data distributions by optimizing such a linear mapping. It improves the expressivity of data distribution, which means it makes the distributions in the same class close and those in different classes far away. Then, the KLD metric learning is modeled by a minimization problem on the manifold of all positive-definite matrices. To deal with this optimization task, we develop an intrinsic steepest descent method, which preserves the manifold structure of the metric in the iteration. Finally, we apply the proposed method along with ten popular metric-learning approaches on the tasks of 3-D object classification and document classification. The experimental results illustrate that our proposed method outperforms all other methods.
Shuyi Ji, Zizhao Zhang 0003, Shihui Ying, Xibin Zhao, Yue Gao 0002
IEEE Trans. Cybern.6
2022 Deep Correlated Joint Network for 2-D Image-Based 3-D Model Retrieval
abstract
In this article, we propose a novel deep correlated joint network (DCJN) approach for 2-D image-based 3-D model retrieval. First, the proposed method can jointly learn two distinct deep neural networks, which are trained for individual modalities to learn two deep nonlinear transformations for visual feature extraction from the co-embedding feature space. Second, we propose the global loss function for the DCJN, consisting of a discriminative loss and a correlation loss. The discriminative loss aims to minimize the intraclass distance of the extracted features and maximize the interclass distance of such features to a large margin within each modality, while the correlation loss focuses on mitigating the distribution discrepancy across different modalities. Consequently, the proposed method can realize cross-modality feature extraction guided by the defined global loss function to benefit the similarity measure between 2-D images and 3-D models. For a comparison experiment, we contribute the current largest 2-D image-based 3-D model retrieval dataset. Moreover, the proposed method was further evaluated on three popular benchmarks, including the 3-D Shape Retrieval Contest 2014, 2016, and 2018 benchmarks. The extensive comparison experimental results demonstrate the superiority of this method over the state-of-the-art methods.
Weizhi Nie, Anan Liu, Sicheng Zhao, Yue Gao 0002
IEEE Trans. Cybern.4
2022 Rotation-Invariant Point Cloud Representation for 3-D Model Recognition
abstract
Three-dimensional (3-D) data have many applications in the field of computer vision and a point cloud is one of the most popular modalities. Therefore, how to establish a good representation for a point cloud is a core issue in computer vision, especially for 3-D object recognition tasks. Existing approaches mainly focus on the invariance of representation under the group of permutations. However, for point cloud data, it should also be rotation invariant. To address such invariance, in this article, we introduce a relation of equivalence under the action of rotation group, through which the representation of point cloud is located in a homogeneous space. That is, two point clouds are regarded as equivalent when they are only different from a rotation. Our network is flexibly incorporated into existing frameworks for point clouds, which guarantees the proposed approach to be rotation invariant. Besides, a sufficient analysis on how to parameterize the group SO(3) into a convolutional network, which captures a relation with all rotations in 3-D Euclidean space [Formula: see text]. We select the optimal rotation as the best representation of point cloud and propose a solution for minimizing the problem on the rotation group SO(3) by using its geometric structure. To validate the rotation invariance, we combine it with two existing deep models and evaluate them on ModelNet40 dataset and its subset ModelNet10. Experimental results indicate that the proposed strategy improves the performance of those existing deep models when the data involve arbitrary rotations.
Yan Wang 0076, Shihui Ying, Shaoyi Du, Yue Gao 0002
IEEE Trans. Cybern.5
2022 Big-Hypergraph Factorization Neural Network for Survival Prediction From Whole Slide Image
abstract
Survival prediction for patients based on histopa- thological whole-slide images (WSIs) has attracted increasing attention in recent years. Due to the massive pixel data in a single WSI, fully exploiting cell-level structural information (e.g., stromal/tumor microenvironment) from the gigapixel WSI is challenging. Most of the current studies resolve the problem by sampling limited image patches to construct a graph-based model (e.g., hypergraph). However, the sampling scale is a critical bottleneck since it is a fundamental obstacle of broadening samples for transductive learning. To overcome the limitation of the sampling scale for constructing a big hypergraph model, we propose a factorization neural network that embeds the correlation among large-scale vertices and hyperedges into two low-dimensional latent semantic spaces separately, empowering the dense sampling. Thanks to the compressed low-dimensional correlation embedding, the hypergraph convolutional layers generate the high-order global representation for each WSI. To minimize the effect of the uncertainty data as well as to achieve the metric-driven learning, we also propose a multi-level ranking supervision to enable the network learning by a queue of patients on the global horizon. Extensive experiments are conducted on three public carcinoma datasets (i.e., LUSC, GBM, and NLST), and the quantitative results demonstrate the proposed method outperforms state-of-the-art methods across-the-board.
Donglin Di, Jun Zhang 0018, Fuqiang Lei, Qi Tian 0001, Yue Gao 0002
IEEE Trans. Image Process.5
2022 Hematoma Expansion Context Guided Intracranial Hemorrhage Segmentation and Uncertainty Estimation
abstract
Accurate segmentation of the Intracranial Hemorrhage (ICH) in non-contrast CT images is significant for computer-aided diagnosis. Although existing methods have achieved remarkable 1 1 The code will be available from https://github.com/JohnleeHIT/SLEX-Net. results, none of them incorporated ICH's prior information in their methods. In this work, for the first time, we proposed a novel SLice EXpansion Network (SLEX-Net), which incorporated hematoma expansion in the segmentation architecture by directly modeling the hematoma variation among adjacent slices. Firstly, a new module named Slice Expansion Module (SEM) was built, which can effectively transfer contextual information between two adjacent slices by mapping predictions from one slice to another. Secondly, to perceive contextual information from both upper and lower slices, we designed two information transmission paths: forward and backward slice expansion, and aggregated results from those paths with a novel weighing strategy. By further exploiting intra-slice and inter-slice context with the information paths, the network significantly improved the accuracy and continuity of segmentation results. Moreover, the proposed SLEX-Net enables us to conduct an uncertainty estimation with one-time inference, which is much more efficient than existing methods. We evaluated the proposed SLEX-Net and compared it with some state-of-the-art methods. Experimental results demonstrate that our method makes significant improvements in all metrics on segmentation performance and outperforms other existing uncertainty estimation methods in terms of several metrics.
Xiangyu Li 0004, Gongning Luo, Wei Wang 0169, Kuanquan Wang, Yue Gao 0002, Shuo Li 0001
IEEE J. Biomed. Health Informatics5
2022 Unsupervised Spectral Feature Selection With Dynamic Hyper-Graph Learning
abstract
Unsupervised spectral feature selection (USFS) methods could output interpretable and discriminative results by embedding a Laplacian regularizer in the framework of sparse feature selection to keep the local similarity of the training samples. To do this, USFS methods usually construct the Laplacian matrix using either a general-graph or a hyper-graph on the original data. Usually, a general-graph could measure the relationship between two samples while a hyper-graph could measure the relationship among no less than two samples. Obviously, the general-graph is a special case of the hyper-graph and the hyper-graph may capture more complex structure of samples than the general graph. However, in previous USFS methods, the construction of the Laplacian matrix is separated from the process of feature selection. Moreover, the original data usually contain noise. Each of them makes difficult to output reliable feature selection models. In this paper, we propose a novel feature selection method by dynamically constructing a hyper-graph based Laplacian matrix in the framework of sparse feature selection. Experimental results on real datasets showed that our proposed method outperformed the state-of-the-art methods in terms of both clustering and segmentation tasks.
Xiaofeng Zhu 0001, Shichao Zhang 0001, Yonghua Zhu, Pengfei Zhu 0001, Yue Gao 0002
IEEE Trans. Knowl. Data Eng.5
2022 Learning Representation on Optimized High-Order Manifold for Visual Classification
abstract
Graph convolutional networks (GCNs) and graph neural networks (GNNs) have demonstrated convincing performance on many tasks by learning the intrinsic structure of the data. However, it is still valuable and challenging to consider the complex and complete correlations of objects, i.e., high-order manifold structures, for representation learning. In this paper, we present a novel representation learning method that utilizes the optimized high-order manifold of the data for classification tasks of nonstructural data and graph-structure data. In the method, we fully explore the complicated relationship of samples by highlighting the high-order manifold information in a hypergraph. Specifically, we incorporate high-order manifold information by graph$p$-Laplacian into a hypergraph and propose$p$-Laplacian-based hypergraph neural networks (pLapHGNN) to significantly learn hidden layer representations that encode both the high-order structure of data and the high-order manifold geometrical information. Confronting the difficulties of obtaining optimized high-order manifolds of the data, we propose an effective approximate approach by graph$p$-Laplacian representing the relationship of hyperedges in the hypergraph. Furthermore, we study the weights of hyperedges in a hypergraph with high-order manifold information. Experiments on the ModelNet40 dataset and NTU dataset demonstrate that the proposed method is more effective than the other popular methods for 3D shape recognition. Extensive experiments on other visual classification tasks and citation networks also show the superiority of our proposed method for representation learning.
Xueqi Ma, Weifeng Liu 0001, Qi Tian 0001, Yue Gao 0002
IEEE Trans. Multim.4
2022 RGB-D Point Cloud Registration Based on Salient Object Detection
abstract
We propose a robust algorithm for aligning rigid, noisy, and partially overlapping red green blue-depth (RGB-D) point clouds. To address the problems of data degradation and uneven distribution, we offer three strategies to increase the robustness of the iterative closest point (ICP) algorithm. First, we introduce a salient object detection (SOD) method to extract a set of points with significant structural variation in the foreground, which can avoid the unbalanced proportion of foreground and background point sets leading to the local registration. Second, registration algorithms that rely only on structural information for alignment cannot establish the correct correspondences when faced with the point set with no significant change in structure. Therefore, a bidirectional color distance (BCD) is designed to build precise correspondence with bidirectional search and color guidance. Third, the maximum correntropy criterion (MCC) and trimmed strategy are introduced into our algorithm to handle with noise and outliers. We experimentally validate that our algorithm is more robust than previous algorithms on simulated and real-world scene data in most scenarios and achieve a satisfying 3-D reconstruction of indoor scenes.
Teng Wan, Shaoyi Du, Wenting Cui, Runzhao Yao, Yuyan Ge, Ce Li 0001, Yue Gao 0002, Nanning Zheng 0001
IEEE Trans. Neural Networks Learn. Syst.7
2021 Improving Image Captioning by Leveraging Intra- and Inter-layer Global Representation in Transformer Network
abstract
Transformer-based architectures have shown great success in image captioning, where object regions are encoded and then attended into the vectorial representations to guide the caption decoding. However, such vectorial representations only contain region-level information without considering the global information reflecting the entire image, which fails to expand the capability of complex multi-modal reasoning in image captioning. In this paper, we introduce a Global Enhanced Transformer (termed GET) to enable the extraction of a more comprehensive global representation, and then adaptively guide the decoder to generate high-quality captions. In GET, a Global Enhanced Encoder is designed for the embedding of the global feature, and a Global Adaptive Decoder are designed for the guidance of the caption generation. The former models intra- and inter-layer global representation by taking advantage of the proposed Global Enhanced Attention and a layer-wise fusion module. The latter contains a Global Adaptive Controller that can adaptively fuse the global information into the decoder to guide the caption generation. Extensive experiments on MS COCO dataset demonstrate the superiority of our GET over many state-of-the-arts.
Jiayi Ji, Yunpeng Luo, Xiaoshuai Sun, Fuhai Chen, Gen Luo, Yongjian Wu 0001, Yue Gao 0002, Rongrong Ji
AAAI7
2021 Domain General Face Forgery Detection by Learning to Weight
abstract
In this paper, we propose a domain-general model, termed learning-to-weight (LTW), that guarantees face detection performance across multiple domains, particularly the target domains that are never seen before. However, various face forgery methods cause complex and biased data distributions, making it challenging to detect fake faces in unseen domains. We argue that different faces contribute differently to a detection model trained on multiple domains, making the model likely to fit domain-specific biases. As such, we propose the LTW approach based on the meta-weight learning algorithm, which configures different weights for face images from different domains. The LTW network can balance the model's generalizability across multiple domains. Then, the meta-optimization calibrates the source domain's gradient enabling more discriminative features to be learned. The detection ability of the network is further improved by introducing an intra-class compact loss. Extensive experiments on several commonly used deepfake datasets to demonstrate the effectiveness of our method in detecting synthetic faces. Code and supplemental material are available at https://github.com/skJack/LTW.
Ke Sun 0016, Hong Liu 0009, Qixiang Ye, Yue Gao 0002, Jianzhuang Liu, Ling Shao 0001, Rongrong Ji
AAAI4
2021 View-Guided Point Cloud Completion
abstract
This paper presents a view-guided solution for the task of point cloud completion. Unlike most existing methods directly inferring the missing points using shape priors, we address this task by introducing ViPC (view-guided point cloud completion) that takes the missing crucial global structure information from an extra single-view image. By leveraging a framework that sequentially performs effective cross-modality and cross-level fusions, our method achieves significantly superior results over typical existing solutions on a new large-scale dataset we collect for the view-guided point cloud completion task.
Xuancheng Zhang, Yutong Feng, Siqi Li 0001, Changqing Zou, Hai Wan, Xibin Zhao, Yandong Guo, Yue Gao 0002
CVPR8
2021 Event Stream Super-Resolution via Spatiotemporal Constraint Learning
abstract
Event cameras are bio-inspired sensors that respond to brightness changes asynchronously and output in the form of event streams instead of frame-based images. They own outstanding advantages compared with traditional cameras: higher temporal resolution, higher dynamic range, and lower power consumption. However, the spatial resolution of existing event cameras is insufficient and challenging to be enhanced at the hardware level while maintaining the asynchronous philosophy of circuit design. Therefore, it is imperative to explore the algorithm of event stream super-resolution, which is a non-trivial task due to the sparsity and strong spatio-temporal correlation of the events from an event camera. In this paper, we propose an end-to-end framework based on spiking neural network for event stream super-resolution, which can generate high-resolution (HR) event stream from the input low-resolution (LR) event stream. A spatiotemporal constraint learning mechanism is proposed to learn the spatial and temporal distributions of the event stream simultaneously. We validate our method on four large-scale datasets and the results show that our method achieves state-of-the-art performance. The satisfying results on two downstream applications, i.e. object classification and image reconstruction, further demonstrate the usability of our method. To prove the application potential of our method, we deploy it on a mobile platform. The high-quality HR event stream generated by our real-time system demonstrates the effectiveness and efficiency of our method.
Siqi Li 0001, Yutong Feng, Yu Jiang 0006, Changqing Zou, Yue Gao 0002
ICCV6
2021 ReCU: Reviving the Dead Weights in Binary Neural Networks
Mingbao Lin, Jianzhuang Liu, Jie Chen 0001, Ling Shao 0001, Yue Gao 0002, Yonghong Tian 0001, Rongrong Ji
ICCV6
2021 TAG-Reg: Iterative Accurate Global Registration Algorithm
abstract
In this paper, we propose an accurate global registration (TAG-Reg) algorithm for poor initialization and partially overlapping point clouds registration problem. Firstly, methods based on geometric structure information of points can get the accurate results, which is vulnerable to poor initialization. Meanwhile, existing features based global methods can solve poor initialization problem at a certain extent, but it cannot obtain accurate results. So, we combine the geometric structure information with feature as hybrid feature to solve poor initialization problem completely and obtain accurate results. Secondly, we introduce dynamic trimmed strategy combining with hybrid feature to deal with partially overlapping problem. Then, to improve the accuracy of our method, we utilize the probabilistic method to suppress noise. At last, we establish the TAG-Reg model and propose an iterative algorithm to solve this problem. Experimental results show that our TAG-Reg achieves state-of-the-art performance compared to existing non-deep learning and recent deep learning methods. Our source code will open at https://github.com/BiaoBiaoLi/TAG-Reg.
Qixing Xie, Shaoyi Du, Wenting Cui, Runzhao Yao, Yue Gao 0002, Nanning Zheng 0001
ICME6
2021 Future vehicles: interactive wheeled robots
Deyi Li, Yue Gao 0002, Hong Bao, Xinkai Xu 0001, Yuansheng Liu, Zhixuan Wu
Sci. China Inf. Sci.6
2021 Hypergraph learning for identification of COVID-19 with CT imaging
Donglin Di, Feng Shi 0001, Fuhua Yan, Liming Xia, Zhanhao Mo, Zhongxiang Ding, Bin Song 0002, Shengrui Li, Ying Wei 0009, Ying Shao, Miaofei Han, Yaozong Gao, He Sui, Yue Gao 0002, Dinggang Shen
Medical Image Anal.15
2021 Deep Multi-View Enhancement Hashing for Image Retrieval
abstract
Hashing is an efficient method for nearest neighbor search in large-scale data space by embedding high-dimensional feature descriptors into a similarity preserving Hamming space with a low dimension. However, large-scale high-speed retrieval through binary code has a certain degree of reduction in retrieval accuracy compared to traditional retrieval methods. We have noticed that multi-view methods can well preserve the diverse characteristics of data. Therefore, we try to introduce the multi-view deep neural network into the hash learning field, and design an efficient and innovative retrieval model, which has achieved a significant improvement in retrieval performance. In this paper, we propose a supervised multi-view hash model which can enhance the multi-view information through neural networks. This is a completely new hash learning method that combines multi-view and deep learning methods. The proposed method utilizes an effective view stability evaluation method to actively explore the relationship among views, which will affect the optimization direction of the entire network. We have also designed a variety of multi-data fusion methods in the Hamming space to preserve the advantages of both convolution and multi-view. In order to avoid excessive computing resources on the enhancement procedure during retrieval, we set up a separate structure called memory network which participates in training together. The proposed method is systematically evaluated on the CIFAR-10, NUS-WIDE and MS-COCO datasets, and the results show that our method significantly outperforms the state-of-the-art single-view and multi-view hashing methods.
Chenggang Yan 0001, Biao Gong, Yuxuan Wei, Yue Gao 0002
IEEE Trans. Pattern Anal. Mach. Intell.4
2021 Joint segmentation and detection of COVID-19 via a sequential region generation network
Jipeng Wu, Shengchuan Zhang, Xi Li 0011, Jie Chen 0001, Jiawen Zheng, Yue Gao 0002, Yonghong Tian 0001, Yongsheng Liang 0001, Rongrong Ji
Pattern Recognit.7
2021 Point Set Registration With Similarity and Affine Transformations Based on Bidirectional KMPE Loss
abstract
Robust point set registration is a challenging problem, especially in the cases of noise, outliers, and partial overlapping. Previous methods generally formulate their objective functions based on the mean-square error (MSE) loss and, hence, are only able to register point sets under predefined constraints (e.g., with Gaussian noise). This article proposes a novel objective function based on a bidirectional kernel mean p -power error (KMPE) loss, to jointly deal with the above nonideal situations. KMPE is a nonsecond-order similarity measure in kernel space and shows a strong robustness against various noise and outliers. Moreover, a bidirectional measure is applied to judge the registration, which can avoid the ill-posed problem when a lot of points converges to the same point. In particular, we develop two effective optimization methods to deal with the point set registrations with the similarity and the affine transformations, respectively. The experimental results demonstrate the effectiveness of our methods.
Yang Yang 0066, Shaoyi Du, Muyi Wang, Badong Chen, Yue Gao 0002
IEEE Trans. Cybern.6
2021 Group-Wise Learning for Aurora Image Classification With Multiple Representations
abstract
In conventional aurora image classification methods, it is general to employ only one single feature representation to capture the morphological characteristics of aurora images, which is difficult to describe the complicated morphologies of different aurora categories. Although several studies have proposed to use multiple feature representations, the inherent correlation among these representations are usually neglected. To address this problem, we propose a group-wise learning (GWL) method for the automatic aurora image classification using multiple representations. Specifically, we first extract the multiple feature representations for aurora images, and then construct a graph in each of multiple feature spaces. To model the correlation among different representations, we partition multiple graphs into several groups via a clustering algorithm. We further propose a GWL model to automatically estimate class labels for aurora images and optimal weights for the multiple representations in a data-driven manner. Finally, we develop a label fusion approach to make a final classification decision for new testing samples. The proposed GWL method focuses on the diverse properties of multiple feature representations, by clustering the correlated representations into the same group. We evaluate our method on an aurora image data set that contains 12 682 aurora images from 19 days. The experimental results demonstrate that the proposed GWL method achieves approximately 6% improvement in terms of classification accuracy, compared to the methods using a single feature representation.
Jun Zhang 0018, Mingxia Liu 0001, Ke Lu 0002, Yue Gao 0002
IEEE Trans. Cybern.4
2021 Multi-Scale Representation Learning on Hypergraph for 3D Shape Retrieval and Recognition
abstract
Effective 3D shape retrieval and recognition are challenging but important tasks in computer vision research field, which have attracted much attention in recent decades. Although recent progress has shown significant improvement of deep learning methods on 3D shape retrieval and recognition performance, it is still under investigated of how to jointly learn an optimal representation of 3D shapes considering their relationships. To tackle this issue, we propose a multi-scale representation learning method on hypergraph for 3D shape retrieval and recognition, called multi-scale hypergraph neural network (MHGNN). In this method, the correlation among 3D shapes is formulated in a hypergraph and a hypergraph convolution process is conducted to learn the representations. Here, multiple representations can be obtained through different convolution layers, leading to multi-scale representations of 3D shapes. A fusion module is then introduced to combine these representations for 3D shape retrieval and recognition. The main advantages of our method lie in 1) the high-order correlation among 3D shapes can be investigated in the framework and 2) the joint multi-scale representation can be more robust for comparison. Comparisons with state-of-the-art methods on the public ModelNet40 dataset demonstrate remarkable performance improvement of our proposed method on the 3D shape retrieval task. Meanwhile, experiments on recognition tasks also show better results of our proposed method, which indicate the superiority of our method on learning better representation for retrieval and recognition.
Biao Gong, Fuqiang Lei, Chenggang Yan 0001, Yue Gao 0002
IEEE Trans. Image Process.6
2021 DAN: Deep-Attention Network for 3D Shape Recognition
abstract
Due to the wide applications in a rapidly increasing number of different fields, 3D shape recognition has become a hot topic in the computer vision field. Many approaches have been proposed in recent years. However, there remain huge challenges in two aspects: exploring the effective representation of 3D shapes and reducing the redundant complexity of 3D shapes. In this paper, we propose a novel deep-attention network (DAN) for 3D shape representation based on multiview information. More specifically, we introduce the attention mechanism to construct a deep multiattention network that has advantages in two aspects: 1) information selection, in which DAN utilizes the self-attention mechanism to update the feature vector of each view, effectively reducing the redundant information, and 2) information fusion, in which DAN applies attention mechanism that can save more effective information by considering the correlations among views. Meanwhile, deep network structure can fully consider the correlations to continuously fuse effective information. To validate the effectiveness of our proposed method, we conduct experiments on the public 3D shape datasets: ModelNet40, ModelNet10, and ShapeNetCore55. Experimental results and comparison with state-of-the-art methods demonstrate the superiority of our proposed method. Code is released on https://github.com/RiDang/DANN.
Weizhi Nie, Yue Zhao 0042, Dan Song 0006, Yue Gao 0002
IEEE Trans. Image Process.4
2021 Few-Shot Learning by a Cascaded Framework With Shape-Constrained Pseudo Label Assessment for Whole Heart Segmentation
abstract
Automatic and accurate 3D cardiac image segmentation plays a crucial role in cardiac disease diagnosis and treatment. Even though CNN based techniques have achieved great success in medical image segmentation, the expensive annotation, large memory consumption, and insufficient generalization ability still pose challenges to their application in clinical practice, especially in the case of 3D segmentation from high-resolution and large-dimension volumetric imaging. In this paper, we propose a few-shot learning framework by combining ideas of semi-supervised learning and self-training for whole heart segmentation and achieve promising accuracy with a Dice score of 0.890 and a Hausdorff distance of 18.539 mm with only four labeled data for training. When more labeled data provided, the model can generalize better across institutions. The key to success lies in the selection and evolution of high-quality pseudo labels in cascaded learning. A shape-constrained network is built to assess the quality of pseudo labels, and the self-training stages with alternative global-local perspectives are employed to improve the pseudo labels. We evaluate our method on the CTA dataset of the MM-WHS 2017 Challenge and a larger multi-center dataset. In the experiments, our method outperforms the state-of-the-art methods significantly and has great generalization ability on the unseen data. We also demonstrate, by a study of two 4D (3D+T) CTA data, the potential of our method to be applied in clinical practice.
Wenji Wang, Qing Xia 0002, Zhennan Yan, Zhuowei Li 0002, Yue Gao 0002, Dimitris N. Metaxas, Shaoting Zhang 0001
IEEE Trans. Medical Imaging8
2020 Divide and Conquer: Question-Guided Spatio-Temporal Contextual Attention for Video Question Answering
abstract
Understanding questions and finding clues for answers are the key for video question answering. Compared with image question answering, video question answering (Video QA) requires to find the clues accurately on both spatial and temporal dimension simultaneously, and thus is more challenging. However, the relationship between spatio-temporal information and question still has not been well utilized in most existing methods for Video QA. To tackle this problem, we propose a Question-Guided Spatio-Temporal Contextual Attention Network (QueST) method. In QueST, we divide the semantic features generated from question into two separate parts: the spatial part and the temporal part, respectively guiding the process of constructing the contextual attention on spatial and temporal dimension. Under the guidance of the corresponding contextual attention, visual features can be better exploited on both spatial and temporal dimensions. To evaluate the effectiveness of the proposed method, experiments are conducted on TGIF-QA dataset, MSRVTT-QA dataset and MSVD-QA dataset. Experimental results and comparisons with the state-of-the-art methods have shown that our method can achieve superior performance.
Jianwen Jiang, Haojie Lin, Xibin Zhao, Yue Gao 0002
AAAI5
2020 Attention-Based Multi-Modal Fusion Network for Semantic Scene Completion
abstract
This paper presents an end-to-end 3D convolutional network named attention-based multi-modal fusion network (AMFNet) for the semantic scene completion (SSC) task of inferring the occupancy and semantic labels of a volumetric 3D scene from single-view RGB-D images. Compared with previous methods which use only the semantic features extracted from RGB-D images, the proposed AMFNet learns to perform effective 3D scene completion and semantic segmentation simultaneously via leveraging the experience of inferring 2D semantic segmentation from RGB-D images as well as the reliable depth cues in spatial dimension. It is achieved by employing a multi-modal fusion architecture boosted from 2D semantic segmentation and a 3D semantic completion network empowered by residual attention blocks. We validate our method on both the synthetic SUNCG-RGBD dataset and the real NYUv2 dataset and the results show that our method respectively achieves the gains of 2.5% and 2.6% on the synthetic SUNCG-RGBD dataset and the real NYUv2 dataset against the state-of-the-art method.
Siqi Li 0001, Changqing Zou, Xibin Zhao, Yue Gao 0002
AAAI5
2020 Hypergraph Label Propagation Network
abstract
In recent years, with the explosion of information on the Internet, there has been a large amount of data produced, and analyzing these data is useful and has been widely employed in real world applications. Since data labeling is costly, lots of research has focused on how to efficiently label data through semi-supervised learning. Among the methods, graph and hypergraph based label propagation algorithms have been a widely used method. However, traditional hypergraph learning methods may suffer from their high computational cost. In this paper, we propose a Hypergraph Label Propagation Network (HLPN) which combines hypergraph-based label propagation and deep neural networks in order to optimize the feature embedding for optimal hypergraph learning through an end-to-end architecture. The proposed method is more effective and also efficient for data labeling compared with traditional hypergraph learning methods. We verify the effectiveness of our proposed HLPN method on a real-world microblog dataset gathered from Sina Weibo. Experiments demonstrate that the proposed method can significantly outperform the state-of-the-art methods and alternative approaches.
Yubo Zhang 0006, Nan Wang 0015, Changqing Zou, Hai Wan, Xibin Zhao, Yue Gao 0002
AAAI7
2020 CoBigICP: Robust and Precise Point Set Registration using Correntropy Metrics and Bidirectional Correspondence
abstract
In this paper, we propose a novel probabilistic variant of iterative closest point (ICP) dubbed as CoBigICP. The method leverages both local geometrical information and global noise characteristics. Locally, the 3D structure of both target and source clouds are incorporated into the objective function through bidirectional correspondence. Globally, error metric of correntropy is introduced as noise model to resist outliers. Importantly, the close resemblance between normal-distributions transform (NDT) and correntropy is revealed. To ease the minimization step, an on-manifold parameterization of the special Euclidean group is proposed. Extensive experiments validate that CoBigICP outperforms several well-known and state-of-the-art methods.
Pengyu Yin, Di Wang 0028, Shaoyi Du, Shihui Ying, Yue Gao 0002, Nanning Zheng 0001
IROS5
2020 Dual Channel Hypergraph Collaborative Filtering
abstract
Collaborative filtering (CF) is one of the most popular and important recommendation methodologies in the heart of numerous recommender systems today. Although widely adopted, existing CF-based methods, ranging from matrix factorization to the emerging graph-based methods, suffer inferior performance especially when the data for training are very limited. In this paper, we first pinpoint the root causes of such deficiency and observe two main disadvantages that stem from the inherent designs of existing CF-based methods, i.e., 1) inflexible modeling of users and items and 2) insufficient modeling of high-order correlations among the subjects. Under such circumstances, we propose a dual channel hypergraph collaborative filtering (DHCF) framework to tackle the above issues. First, a dual channel learning strategy, which holistically leverages the divide-and-conquer strategy, is introduced to learn the representation of users and items so that these two types of data can be elegantly interconnected while still maintaining their specific properties. Second, the hypergraph structure is employed for modeling users and items with explicit hybrid high-order correlations. The jump hypergraph convolution (JHConv) method is proposed to support the explicit and efficient embedding propagation of high-order correlations. Comprehensive experiments on two public benchmarks and two new real-world datasets demonstrate that DHCF can achieve significant and consistent improvements against other state-of-the-art methods.
Shuyi Ji, Yifan Feng 0001, Rongrong Ji, Xibin Zhao, Wanwan Tang, Yue Gao 0002
KDD6
2020 Ranking-Based Survival Prediction on Histopathological Whole-Slide Images
Donglin Di, Shengrui Li, Jun Zhang 0018, Yue Gao 0002
MICCAI (5)4
2020 IExpressNet: Facial Expression Recognition with Incremental Classes
abstract
Existing methods on facial expression recognition (FER) are mainly trained in the setting when all expression classes are fixed in advance. However, in real applications, expression classes are becoming increasingly fine-grained and incremental. To deal with sequential expression classes, we can fine-tune or re-train these models, but this often results in poor performance or large computing resources consumption. To address these problems, we develop an Incremental Facial Expression Recognition Network (IExpressNet), which can learn a competitive multi-class classifier at any time with a lower requirement of computing resources. Specifically, IExpressNet consists of two novel components. First, we construct an exemplar set by dynamically selecting representative samples from old expression classes. Then, the exemplar set and new expression classes samples constitute the training set. Second, we design a novel center-expression-distilled loss. As for facial expression in the wild, center-expression-distilled loss enhances the discriminative power of the deeply learned features and prevents catastrophic forgetting. Extensive experiments are conducted on two large-scale FER datasets in the wild, RAF-DB and AffectNet. The results demonstrate the superiority of the proposed method as compared to state-of-the-art incremental learning approaches.
Bingjun Luo, Sicheng Zhao, Shihui Ying, Xibin Zhao, Yue Gao 0002
ACM Multimedia6
2020 Future vehicles: learnable wheeled robots
Deyi Li, Yue Gao 0002
Sci. China Inf. Sci.3
2020 Robust rigid registration algorithm based on pointwise correspondence and correntropy
Shaoyi Du, Guanglin Xu, Sirui Zhang, Xuetao Zhang 0001, Yue Gao 0002, Badong Chen
Pattern Recognit. Lett.5
2020 Discrete Probability Distribution Prediction of Image Emotions with Shared Sparse Learning
abstract
Computationally modelling the affective content of images has been extensively studied recently because of its wide applications in entertainment, advertisement, and education. Significant progress has been made on designing discriminative features to bridge the affective gap. Assuming that viewers can reach a consensus on the emotion of images, most existing works focused on assigning the dominant emotion category or the average dimension values to an image. However, the image emotions perceived by viewers are subjective by nature with the influence of personal and situational factors. In this paper, we propose a novel machine learning approach that characterizes the categorical image emotions as a discrete probability distribution (DPD). To associate emotion with the visual features extracted from images, we present shared sparse learning to learn the combination coefficients, with which the DPD of an unseen image is predicted by linearly combining the DPDs of the training images. Furthermore, we extend our method to the setup where multi-features are available and learn the optimal weights for each feature to reflect the importance of different features. Extensive experiments are carried out on Abstract, Emotion6 and IESN datasets and the results demonstrate the superiority of the proposed method, as compared to the state-of-the-art approaches.
Sicheng Zhao, Guiguang Ding, Yue Gao 0002, Xin Zhao 0020, Youbao Tang, Jungong Han, Hongxun Yao, Qingming Huang
IEEE Trans. Affect. Comput.3
2020 Time-Triggered Switch-Memory-Switch Architecture for Time-Sensitive Networking Switches
abstract
Time-sensitive networking (TSN) is a set of extended standards for the IEEE 802.3 Ethernet under development by the IEEE 802.1 TSN task group. TSN depends on two key components, scheduling and fault tolerance, to provide realtime and reliable transmission. There is a strong motivation to replace the widely used field-buses with TSNs in industrial networking applications. However, industrial network devices are typical application-specific embedded systems with limited memory resources. Time-sensitive (TS) transmission certainly prefers on-chip memory, which is even more scarce for embedded systems. As a result, it is critical for TSNs to develop memory-efficient switching techniques with scalable schedulability and elegant fault-tolerance support. This paper proposes a time-triggered switch-memory-switch (SMS) architecture for memory-efficient TSN switches. First, based on the SMS shared memory, our architecture makes it possible to statically schedule memory allocation with full utilization for TS traffic and the remaining memory for other traffic. Compared with perport memory, the shared memory achieves a ratio of (nn/n!) (≈ (en/√(2πn)), n → ∞), where n is the port number, in the feasible solution space under memory constraints and thus significantly improves scheduling memory ability and flexibility. Moreover, we develop a fault-tolerance scheme for reliable transmission. It facilitates a memory-efficient implementation of the popular multiline redundancy in industrial networks. The scheme is validated by five classes of memory conflicts and a case study on two-line redundancy.
Zonghui Li, Hai Wan, Yangdong Deng, Xibin Zhao, Yue Gao 0002, Ming Gu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2020 Model-Based Adaptation of Mixed-Criticality Multiservice Systems for Extreme Physical Environments
abstract
An increasingly important trend in the design of industry-strength embedded systems is the integration of multiple services with varying criticality levels into a common computing platform. Such systems are characterized as mixed-criticality multiservice systems (MCMSs). An MCMS has to survive in rigorous environments posed by industry-level requirements. Such survival, however, is becoming continuously more challenging due to the growing system complexity and integrating more and more services. While existing works typically target reliability-driven design optimization to improve the system robustness rather than deal with the surviving problem of the system in extreme physical environments, this paper addresses the problem by enabling the service capability transitions of an MCMS to adapt to the environments. This paper proposes a service capability model to capture the importance of functional modules for the criticality of different services. A model-based service-capability transition mechanism is designed to automatically identify the maximum allowed service capability under a given physical environment. A case study of the proposed techniques was performed on an industrial Ethernet switch which is a typical MCMS, to validate the capability of adaptation to high and low temperatures. The experimental results demonstrate the significant potential of our approach to improve system survivability under extreme physical environments.
Zonghui Li, Hai Wan, Yangdong Deng, Xibin Zhao, Yue Gao 0002, Ming Gu 0001
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.5
2020 Online Scheduling for Dynamic VM Migration in Multicast Time-Sensitive Networks
abstract
With the development of hardware virtualization and cloud computing, modern industry has a tendency to upgrade from the traditional industrial networks to virtual machine (VM) based networks. To provide firm latency guarantees for control messages in these networks, the time-sensitive network (TSN) is a promising technology due to its determinacy for real-time applications. However, TSN faces the challenge of providing a rapid response to dynamic transmission requirement changes incurred by VM migrations. In this paper, we proposed an online scheduling approach to deal with dynamic VM migrations in multicast TSN. In this approach, we devise a novel online scheduling framework [minimal distance tree (MDT) construction - heuristic breadth first search] containing an offline scheduling phase and an online rescheduling phase. While the offline phase introduces a MDT to increase reusable scheduling results, the online phase proposes a heuristic scheduling approach to reuse the results of the offline phase as much as possible to accelerate the rescheduling process. Experiments show that our framework can provide a rapid response to dynamic VM migrations compared with the existing approaches where the amount of control data does not exceed 50% of the bandwidth.
Qinghan Yu, Hai Wan, Xibin Zhao, Yue Gao 0002, Ming Gu 0001
IEEE Trans. Ind. Informatics4
2020 Hamming Embedding Sensitivity Guided Fusion Network for 3D Shape Representation
abstract
Three-dimensional multi-modal data are used to represent 3D objects in the real world in different ways. Features separately extracted from multimodality data are often poorly correlated. Recent solutions leveraging the attention mechanism to learn a joint-network for the fusion of multimodality features have weak generalization capability. In this paper, we propose a hamming embedding sensitivity network to address the problem of effectively fusing multimodality features. The proposed network called HamNet is the first end-to-end framework with the capacity to theoretically integrate data from all modalities with a unified architecture for 3D shape representation, which can be used for 3D shape retrieval and recognition. HamNet uses the feature concealment module to achieve effective deep feature fusion. The basic idea of the concealment module is to re-weight the features from each modality at an early stage with the hamming embedding of these modalities. The hamming embedding also provides an effective solution for fast retrieval tasks on a large scale dataset. We have evaluated the proposed method on the large-scale ModelNet40 dataset for the tasks of 3D shape classification, single modality and cross-modality retrieval. Comprehensive experiments and comparisons with state-of-the-art methods demonstrate that the proposed approach can achieve superior performance.
Biao Gong, Chenggang Yan 0001, Changqing Zou, Yue Gao 0002
IEEE Trans. Image Process.5
2020 Multi-Atlas Segmentation of Anatomical Brain Structures Using Hierarchical Hypergraph Learning
abstract
Accurate segmentation of anatomical brain structures is crucial for many neuroimaging applications, e.g., early brain development studies and the study of imaging biomarkers of neurodegenerative diseases. Although multi-atlas segmentation (MAS) has achieved many successes in the medical imaging area, this approach encounters limitations in segmenting anatomical structures associated with poor image contrast. To address this issue, we propose a new MAS method that uses a hypergraph learning framework to model the complex subject-within and subject-to-atlas image voxel relationships and propagate the label on the atlas image to the target subject image. To alleviate the low-image contrast issue, we propose two strategies equipped with our hypergraph learning framework. First, we use a hierarchical strategy that exploits high-level context features for hypergraph construction. Because the context features are computed on the tentatively estimated probability maps, we can ultimately turn the hypergraph learning into a hierarchical model. Second, instead of only propagating the labels from the atlas images to the target subject image, we use a dynamic label propagation strategy that can gradually use increasing reliably identified labels from the subject image to aid in predicting the labels on the difficult-to-label subject image voxels. Compared with the state-of-the-art label fusion methods, our results show that the hierarchical hypergraph learning framework can substantially improve the robustness and accuracy in the segmentation of anatomical brain structures with low image contrast from magnetic resonance (MR) images.
Pei Dong, Yanrong Guo, Yue Gao 0002, Peipeng Liang, Yonghong Shi, Guorong Wu 0001
IEEE Trans. Neural Networks Learn. Syst.3
2019 MeshNet: Mesh Neural Network for 3D Shape Representation
abstract
Mesh is an important and powerful type of data for 3D shapes and widely studied in the field of computer vision and computer graphics. Regarding the task of 3D shape representation, there have been extensive research efforts concentrating on how to represent 3D shapes well using volumetric grid, multi-view and point cloud. However, there is little effort on using mesh data in recent years, due to the complexity and irregularity of mesh data. In this paper, we propose a mesh neural network, named MeshNet, to learn 3D shape representation from mesh data. In this method, face-unit and feature splitting are introduced, and a general architecture with available and effective blocks are proposed. In this way, MeshNet is able to solve the complexity and irregularity problem of mesh and conduct 3D shape representation well. We have applied the proposed MeshNet method in the applications of 3D shape classification and retrieval. Experimental results and comparisons with the state-of-the-art methods demonstrate that the proposed MeshNet can achieve satisfying 3D shape classification and retrieval performance, which indicates the effectiveness of the proposed method on 3D shape representation.
Yutong Feng, Yifan Feng 0001, Haoxuan You, Xibin Zhao, Yue Gao 0002
AAAI5
2019 Hypergraph Neural Networks
abstract
In this paper, we present a hypergraph neural networks (HGNN) framework for data representation learning, which can encode high-order data correlation in a hypergraph structure. Confronting the challenges of learning representation for complex data in real practice, we propose to incorporate such data structure in a hypergraph, which is more flexible on data modeling, especially when dealing with complex data. In this method, a hyperedge convolution operation is designed to handle the data correlation during representation learning. In this way, traditional hypergraph learning procedure can be conducted using hyperedge convolution operations efficiently. HGNN is able to learn the hidden layer representation considering the high-order data structure, which is a general framework considering the complex data correlations. We have conducted experiments on citation network classification and visual object recognition tasks and compared HGNN with graph convolutional networks and other traditional methods. Experimental results demonstrate that the proposed HGNN method outperforms recent state-of-theart methods. We can also reveal from the results that the proposed HGNN is superior when dealing with multi-modal data compared with existing methods.
Yifan Feng 0001, Haoxuan You, Zizhao Zhang 0003, Rongrong Ji, Yue Gao 0002
AAAI5
2019 DeepCCFV: Camera Constraint-Free Multi-View Convolutional Neural Network for 3D Object Retrieval
abstract
3D object retrieval has a compelling demand in the field of computer vision with the rapid development of 3D vision technology and increasing applications of 3D objects. 3D objects can be described in different ways such as voxel, point cloud, and multi-view. Among them, multi-view based approaches proposed in recent years show promising results. Most of them require a fixed predefined camera position setting which provides a complete and uniform sampling of views for objects in the training stage. However, this causes heavy over-fitting problems which make the models failed to generalize well in free camera setting applications, particularly when insufficient views are provided. Experiments show the performance drastically drops when the number of views reduces, hindering these methods from practical applications. In this paper, we investigate the over-fitting issue and remove the constraint of the camera setting. First, two basic feature augmentation strategies Dropout and Dropview are introduced to solve the over-fitting issue, and a more precise and more efficient method named DropMax is proposed after analyzing the drawback of the basic ones. Then, by reducing the over-fitting issue, a camera constraint-free multi-view convolutional neural network named DeepCCFV is constructed. Extensive experiments on both single-modal and cross-modal cases demonstrate the effectiveness of the proposed method in free camera settings comparing with existing state-of-theart 3D object retrieval methods.
Zhengyue Huang, Zhehui Zhao, Hengguang Zhou, Xibin Zhao, Yue Gao 0002
AAAI5
2019 MLVCNN: Multi-Loop-View Convolutional Neural Network for 3D Shape Retrieval
abstract
3D shape retrieval has attracted much attention and become a hot topic in computer vision field recently.With the development of deep learning, 3D shape retrieval has also made great progress and many view-based methods have been introduced in recent years. However, how to represent 3D shapes better is still a challenging problem. At the same time, the intrinsic hierarchical associations among views still have not been well utilized. In order to tackle these problems, in this paper, we propose a multi-loop-view convolutional neural network (MLVCNN) framework for 3D shape retrieval. In this method, multiple groups of views are extracted from different loop directions first. Given these multiple loop views, the proposed MLVCNN framework introduces a hierarchical view-loop-shape architecture, i.e., the view level, the loop level, and the shape level, to conduct 3D shape representation from different scales. In the view-level, a convolutional neural network is first trained to extract view features. Then, the proposed Loop Normalization and LSTM are utilized for each loop of view to generate the loop-level features, which considering the intrinsic associations of the different views in the same loop. Finally, all the loop-level descriptors are combined into a shape-level descriptor for 3D shape representation, which is used for 3D shape retrieval. Our proposed method has been evaluated on the public 3D shape benchmark, i.e., ModelNet40. Experiments and comparisons with the state-of-the-art methods show that the proposed MLVCNN method can achieve significant performance improvement on 3D shape retrieval tasks. Our MLVCNN outperforms the state-of-the-art methods by the mAP of 4.84% in 3D shape retrieval task. We have also evaluated the performance of the proposed method on the 3D shape classification task where MLVCNN also achieves superior performance compared with recent methods.
Jianwen Jiang, Di Bao, Xibin Zhao, Yue Gao 0002
AAAI5
2019 PVRNet: Point-View Relation Neural Network for 3D Shape Recognition
abstract
Three-dimensional (3D) shape recognition has drawn much research attention in the field of computer vision. The advances of deep learning encourage various deep models for 3D feature representation. For point cloud and multi-view data, two popular 3D data modalities, different models are proposed with remarkable performance. However the relation between point cloud and views has been rarely investigated. In this paper, we introduce Point-View Relation Network (PVRNet), an effective network designed to well fuse the view features and the point cloud feature with a proposed relation score module. More specifically, based on the relation score module, the point-single-view fusion feature is first extracted by fusing the point cloud feature and each single view feature with point-singe-view relation, then the pointmulti- view fusion feature is extracted by fusing the point cloud feature and the features of different number of views with point-multi-view relation. Finally, the point-single-view fusion feature and point-multi-view fusion feature are further combined together to achieve a unified representation for a 3D shape. Our proposed PVRNet has been evaluated on ModelNet40 dataset for 3D shape classification and retrieval. Experimental results indicate our model can achieve significant performance improvement compared with the state-of-the-art models.
Haoxuan You, Yifan Feng 0001, Xibin Zhao, Changqing Zou, Rongrong Ji, Yue Gao 0002
AAAI6
2019 Universal Adversarial Perturbation via Prior Driven Uncertainty Approximation
abstract
Deep learning models have shown their vulnerabilities to universal adversarial perturbations (UAP), which are quasi-imperceptible. Compared to the conventional supervised UAPs that suffer from the knowledge of training data, the data-independent unsupervised UAPs are more applicable. Existing unsupervised methods fail to take advantage of the model uncertainty to produce robust perturbations. In this paper, we propose a new unsupervised universal adversarial perturbation method, termed as Prior Driven Uncertainty Approximation (PD-UA), to generate a robust UAP by fully exploiting the model uncertainty at each network layer. Specifically, a Monte Carlo sampling method is deployed to activate more neurons to increase the model uncertainty for a better adversarial perturbation. Thereafter, a textural bias prior to revealing a statistical uncertainty is proposed, which helps to improve the attacking performance. The UAP is crafted by the stochastic gradient descent algorithm with a boosted momentum optimizer, and a Laplacian pyramid frequency model is finally used to maintain the statistical uncertainty. Extensive experiments demonstrate that our method achieves well attacking performances on the ImageNet validation set, and significantly improves the fooling rate compared with the state-of-the-art methods.
Hong Liu 0009, Rongrong Ji, Jie Li 0052, Baochang Zhang 0001, Yue Gao 0002, Yongjian Wu 0001, Feiyue Huang
ICCV5
2019 Universal Perturbation Attack Against Image Retrieval
abstract
Universal adversarial perturbations (UAPs), a.k.a. input-agnostic perturbations, has been proved to exist and be able to fool cutting-edge deep learning models on most of the data samples. Existing UAP methods mainly focus on attacking image classification models. Nevertheless, little attention has been paid to attacking image retrieval systems. In this paper, we make the first attempt in attacking image retrieval systems. Concretely, image retrieval attack is to make the retrieval system return irrelevant images to the query at the top ranking list. It plays an important role to corrupt the neighbourhood relationships among features in image retrieval attack. To this end, we propose a novel method to generate retrieval-against UAP to break the neighbourhood relationships of image features via degrading the corresponding ranking metric. To expand the attack method to scenarios with varying input sizes or untouchable network parameters, a multi-scale random resizing scheme and a ranking distillation strategy are proposed. We evaluate the proposed method on four widely-used image retrieval datasets, and report a significant performance drop in terms of different metrics, such as mAP and mP@10. Finally, we test our attack methods on the real-world visual search engine, i.e., Google Images, which demonstrates the practical potentials of our methods.
Jie Li 0052, Rongrong Ji, Hong Liu 0009, Xiaopeng Hong, Yue Gao 0002, Qi Tian 0001
ICCV5
2019 Emotion Recognition from Physiological Signals using Multi-Hypergraph Neural Networks
abstract
Emotion recognition from physiological signals is an effective way to discern the inner state of users. Existing works are lack in the exploration of latent correlation among multiple physiological signals and relationship among different subjects. To tackle this issue, we propose to recognize emotion from physiological signals using multi-hypergraph neural networks (MHGNN). In this method, the correlation among different subjects is formulated in the multi-hypergraph structure, where each type of physiological signal is used to generate one hypergraph. In each hypergraph, the hyperedges are used to represent the connections among the vertices (subject, stimuli). Thus, the emotion recognition task is modeled as classifying each vertex in the multi-hypergraph. Experimental results and comparisons with the state-of-the-art methods in the DEAP dataset demonstrate the superior performance of our method. The comparative experiments based on available biological knowledge verify that MHGNN can depict the real biological response process in a much more precise way.
Xibin Zhao, Han Hu 0003, Yue Gao 0002
ICME4
2019 Dynamic Hypergraph Neural Networks
abstract
In recent years, graph/hypergraph-based deep learning methods have attracted much attention from researchers. These deep learning methods take graph/hypergraph structure as prior knowledge in the model. However, hidden and important relations are not directly represented in the inherent structure. To tackle this issue, we propose a dynamic hypergraph neural networks framework (DHGNN), which is composed of the stacked layers of two modules: dynamic hypergraph construction (DHG) and hypergrpah convolution (HGC). Considering initially constructed hypergraph is probably not a suitable representation for data, the DHG module dynamically updates hypergraph structure on each layer. Then hypergraph convolution is introduced to encode high-order data relations in a hypergraph structure. The HGC module includes two phases: vertex convolution and hyperedge convolution, which are designed to aggregate feature among vertices and hyperedges, respectively. We have evaluated our method on standard datasets, the Cora citation network and Microblog dataset. Our method outperforms state-of-the-art methods. More experiments are conducted to demonstrate the effectiveness and robustness of our method to diverse data distributions.
Jianwen Jiang, Yuxuan Wei, Yifan Feng 0001, Jingxuan Cao, Yue Gao 0002
IJCAI5
2019 Deep Local-Global Refinement Network for Stent Analysis in IVOCT Images
Yuyu Guo 0002, Lei Bi 0001, Ashnil Kumar, Yue Gao 0002, Ruiyan Zhang, David Dagan Feng, Qian Wang 0001, Jinman Kim
MICCAI (5)4
2019 A Deep Reinforcement Learning Framework for Frame-by-Frame Plaque Tracking on Intravascular Optical Coherence Tomography Image
Gongning Luo, Suyu Dong, Kuanquan Wang, Dong Zhang 0009, Yue Gao 0002, Xin Chen 0025, Henggui Zhang, Shuo Li 0001
MICCAI (1)5
2019 Robust ℓ2-Hypergraph and its applications
Taisong Jin, Zhengtao Yu 0001, Yue Gao 0002, Shengxiang Gao, Xiaoshuai Sun, Cuihua Li
Inf. Sci.3
2019 RGB-D point cloud registration via infrared and color camera
Teng Wan, Shaoyi Du, Yiting Xu, Guanglin Xu, Badong Chen, Yue Gao 0002
Multim. Tools Appl.7
2019 Hyper-Clique Graph Matching and Applications
abstract
This paper proposes a method for hyper-clique graph (HCG) generation, which can be considered an extension of classical graphs and hyper-graphs in which the node is replaced with the clique (a set of neighboring nodes in a specific feature space) and the hyper-edge linking multiple nodes is replaced with the hyper-edge linking multiple cliques. In addition, we propose the HCG matching method by preserving global and local structures. Specifically, we embed the clique relations of arbitrary orders in a high-order similarity tensor in a recursive manner. Then, we formulate the objective function of HCG matching with respect to two latent variables: the latent clique structure information in the original graph and the similarity measure of clique sets from pairwise HCGs. Since the objective function is not jointly convex with respect to both latent variables, we decompose it into two consecutive measurements for optimization: 1) a clique-to-clique similarity measurement by preserving local unary and pairwise correspondences and 2) a graph-to-graph similarity measurement by preserving global clique-to-clique correspondence. We suitably adopt the affinity-preserving reweighted random walks to optimize the objective function. We extensively evaluate the HCG matching performance on multiple applications: 1) we evaluate the robustness of HCG with respect to the deformation noise, the number of outliers, and the edge density on synthetic data and explore the effects of both the clique order and hyper-edge order on performance; 2) we explore HCG matching for feature point matching on multiple image data sets (CMU house sequence, Caltech+MSRC, and Car+Motor); and 3) we explore HCG matching for multi-view object retrieval, which is a much more challenging task since multi-view objects contain significant variations of illumination, viewpoint, and so on, using popular data sets (MV-RED and NTU). A comparison against the state-of-the-art methods demonstrates the superior performance of the proposed method.
Weizhi Nie, Anan Liu, Yue Gao 0002, Yuting Su 0001
IEEE Trans. Circuits Syst. Video Technol.3
2019 Desktop Action Recognition From First-Person Point-of-View
abstract
Desktop action recognition from first-person view (egocentric) video is an important task due to its omnipresence in our daily life, and the ideal first-person viewing perspective for observing hand-object interactions. However, no previous research efforts have been dedicated on the benchmark of the task. In this paper, we first release a dataset of daily desktop actions recorded with a wearable camera and publish it as a benchmark for desktop action recognition. Regular desktop activities of six participants were recorded in egocentric video with a wide-angle head-mounted camera. In particular, we focus on five common desktop actions in which hands are involved. We provide original video data, action annotations at frame-level, and hand masks at pixel-level. We also propose a feature representation for the characterization of different desktop actions based on the spatial and temporal information of hands. In experiments, we illustrate the statistical information about the dataset, and evaluate the action recognition performance of different features as a baseline. The proposed method achieves promising performance for five action classes.
Minjie Cai, Feng Lu 0005, Yue Gao 0002
IEEE Trans. Cybern.3
2019 Correntropy-Induced Robust Low-Rank Hypergraph
abstract
Hypergraph learning has been widely exploited in various image processing applications, due to its advantages in modeling the high-order information. Its efficacy highly depends on building an informative hypergraph structure to accurately and robustly formulate the underlying data correlation. However, the existing hypergraph learning methods are sensitive to non- Gaussian noise, which hurts the corresponding performance. In this paper, we present a noise-resistant hypergraph learning model, which provides superior robustness against various non- Gaussian noises. In particular, our model adopts low-rank representation to construct a hypergraph, which captures the globally linear data structure as well as preserving the grouping effect of highly-correlated data. We further introduce a correntropyinduced local metric to measure the reconstruction errors, which is particularly robust to non-Gaussian noises. Finally, the Frobenious-norm based regularization is proposed to combine with the low-rank regularizer, which enables our model to regularize the singular values of the coefficient matrix. By such, the non-zero coefficients are selected to generate a hyperedge set as well as the hyperedge weights. We have evaluated the proposed hypergraph model in the tasks of image clustering and semi-supervised image classification. Quantitatively, our scheme significantly enhances the performance of the state-of-the-art hypergraph models on several benchmark datasets.
Taisong Jin, Rongrong Ji, Yue Gao 0002, Xiaoshuai Sun, Xibin Zhao, Dacheng Tao
IEEE Trans. Image Process.3
2019 Cross-Modality Microblog Sentiment Prediction via Bi-Layer Multimodal Hypergraph Learning
abstract
Microblog sentiment prediction has attracted extensive research focus with wide application prospects. With the increasing proportion of multimodal tweets consisting of images, texts, and emoticons, new challenges have been raised to the existing sentiment prediction schemes. More crucially, it remains an open problem to model the dependency among multiple modalities, where one or more modalities may be missing. In this paper, we present a novel Bi-layer Multimodal Hypergraph learning (Bi-MHG) toward robust sentiment prediction of multimodal tweets to tackle the above challenges. In particular, we design a two-layer structure for the proposed Bi-MHG model: The first layer, that is, a tweet-level hypergraph, learns the tweet-feature correlation and the tweet relevance to predict the sentiments of unlabeled tweets. The second layer, that is, a feature-level hypergraph learns the relevance among different feature modalities (including the midlevel visual features in Sentibank [1]) by leveraging prior multimodal sentiment dictionaries. These two layers are connected by sharing the relevance of multimodal features in a unified bilayer learning scheme. In such a way, Bi-MHG explicitly models the modality relevance rather than implicitly weighting multimodal features adopted in the existing Multimodal Hypergraph learning [2]. Finally, a nested alternating optimization is further proposed for Bi-MHG parameter learning. We have carried out extensive evaluations on a real-world microblog dataset crawled from Sina Weibo. For the task of multimodal sentiment prediction, superior performance is reported over several state-of-the-art and alternative approaches, which demonstrates the merits of the proposed scheme.
Rongrong Ji, Fuhai Chen, Liujuan Cao, Yue Gao 0002
IEEE Trans. Multim.4
2019 Hypergraph-Induced Convolutional Networks for Visual Classification
abstract
At present, convolutional neural networks (CNNs) have become popular in visual classification tasks because of their superior performance. However, CNN-based methods do not consider the correlation of visual data to be classified. Recently, graph convolutional networks (GCNs) have mitigated this problem by modeling the pairwise relationship in visual data. Real-world tasks of visual classification typically must address numerous complex relationships in the data, which are not fit for the modeling of the graph structure using GCNs. Therefore, it is vital to explore the underlying correlation of visual data. Regarding this issue, we propose a framework called the hypergraph-induced convolutional network to explore the high-order correlation in visual data during deep neural networks. First, a hypergraph structure is constructed to formulate the relationship in visual data. Then, the high-order correlation is optimized by a learning process based on the constructed hypergraph. The classification tasks are performed by considering the high-order correlation in the data. Thus, the convolution of the hypergraph-induced convolutional network is based on the corresponding high-order relationship, and the optimization on the network uses each data and considers the high-order correlation of the data. To evaluate the proposed hypergraph-induced convolutional network framework, we have conducted experiments on three visual data sets: the National Taiwan University 3-D model data set, Princeton Shape Benchmark, and multiview RGB-depth object data set. The experimental results and comparison in all data sets demonstrate the effectiveness of our proposed hypergraph-induced convolutional network compared with the state-of-the-art methods.
Heyuan Shi, Yubo Zhang 0006, Zizhao Zhang 0003, Nan Ma 0012, Xibin Zhao, Yue Gao 0002, Jia-Guang Sun 0001
IEEE Trans. Neural Networks Learn. Syst.6
2019 Personalized Emotion Recognition by Personality-Aware High-Order Learning of Physiological Signals
abstract
Due to the subjective responses of different subjects to physical stimuli, emotion recognition methodologies from physiological signals are increasingly becoming personalized. Existing works mainly focused on modeling the involved physiological corpus of each subject, without considering the psychological factors, such as interest and personality. The latent correlation among different subjects has also been rarely examined. In this article, we propose to investigate the influence of personality on emotional behavior in a hypergraph learning framework. Assuming that each vertex is a compound tuple (subject, stimuli), multi-modal hypergraphs can be constructed based on the personality correlation among different subjects and on the physiological correlation among corresponding stimuli. To reveal the different importance of vertices, hyperedges, and modalities, we learn the weights for each of them. As the hypergraphs connect different subjects on the compound vertices, the emotions of multiple subjects can be simultaneously recognized. In this way, the constructed hypergraphs are vertex-weighted multi-modal multi-task ones. The estimated factors, referred to as emotion relevance, are employed for emotion recognition. We carry out extensive experiments on the ASCERTAIN dataset and the results demonstrate the superiority of the proposed method, as compared to the state-of-the-art emotion recognition approaches.
Sicheng Zhao, Amir Gholami, Guiguang Ding, Yue Gao 0002, Jungong Han, Kurt Keutzer
ACM Trans. Multim. Comput. Commun. Appl.4
2019 An Enhanced Reconfiguration for Deterministic Transmission in Time-Triggered Networks
abstract
The emerging momentum of digital transformation of industry, i.e. Industry 4.0, poses strong demands for integrating industrial control networks, and Ethernet to enable the real-time Internet of Things (RT-IoT). Time-triggered (TT) networks provide a cost-efficient integrated solution while RT-IoT arouses the reconfiguration challenges: the network has to be flexible enough to adapt to changes and yet provides deterministic transmission persistently during network reconfiguration. Software defined network benefits the flexible industrial control by configuring the rules handling frames. However, previous reconfiguration mechanisms are mostly oriented to the context of data centers and wide area networks and thus do not consider the deterministic transmission in TT networks. This paper focuses on the reconfiguration (i.e., updates) for the deterministic transmission. To minimize the overhead during updates, namely the minimum number of loss frames and the minimum duration time of updates, we first establish an update theory based on the dependence relationship derived by the conflicts during updates. In addition then the reconfiguration problem is modeled with the dependence graph built by the relationship. On such a basis, we present a reconfiguration mechanism and its implementation to solve the problem. Finally, we evaluate the proposed reconfiguration mechanism in two real industrial network topologies. The experimental results demonstrate that compared with previous methods, our mechanism significantly reduces the number of loss frames and achieves zero loss in almost all cases.
Zonghui Li, Hai Wan, Zaiyu Pang, Qiubo Chen, Yangdong Deng, Xibin Zhao, Yue Gao 0002, Ming Gu 0001
IEEE/ACM Trans. Netw.7
2018 Energy-Efficient Automatic Train Driving by Learning Driving Patterns
abstract
Railway is regarded as the most sustainable means of modern transportation. With the fast-growing of fleet size and the railway mileage, the energy consumption of trains is becoming a serious concern globally. The nature of railway offers a unique opportunity to optimize the energy efficiency of locomotives by taking advantage of the undulating terrains along a route. The derivation of an energy-optimal train driving solution, however, proves to be a significant challenge due to the high dimension, nonlinearity, complex constraints, and time-varying characteristic of the problem. An optimized solution can only be attained by considering both the complex environmental conditions of a given route and the inherent characteristics of a locomotive. To tackle the problem, this paper employs a high-order correlation learning method for online generation of the energy optimized train driving solutions. Based on the driving data of experienced human drivers, a hypergraph model is used to learn the optimal embedding from the specified features for the decision of a driving operation. First, we design a feature set capturing the driving status. Next all the training data are formulated as a hypergraph and an inductive learning process is conducted to obtain the embedding matrix. The hypergraph model can be used for real-time generation of driving operation. We also proposed a reinforcement updating scheme, which offers the capability of sustainable enhancement on the hypergraph model in industrial applications. The learned model can be used to determine an optimized driving operation in real-time tested on the Hardware-in-Loop platform. Validation experiments proved that the energy consumption of the proposed solution is around 10% lower than that of average human drivers.
Jin Huang 0002, Yue Gao 0002, Xibin Zhao, Yangdong Deng, Ming Gu 0001
AAAI2
2018 EMD Metric Learning
abstract
Earth Mover's Distance (EMD), targeting at measuring the many-to-many distances, has shown its superiority and been widely applied in computer vision tasks, such as object recognition, hyperspectral image classification and gesture recognition. However, there is still little effort concentrated on optimizing the EMD metric towards better matching performance. To tackle this issue, we propose an EMD metric learning algorithm in this paper. In our method, the objective is to learn a discriminative distance metric for EMD ground distance matrix generation which can better measure the similarity between compared subjects. More specifically, given a group of labeled data from different categories, we first select a subset of training data and then optimize the metric for ground distance matrix generation. Here, both the EMD metric and the EMD flow-network are alternatively optimized until a steady EMD value can be achieved. This method is able to generate a discriminative ground distance matrix which can further improve the EMD distance measurement. We then apply our EMD metric learning method on two tasks, i.e., multi-view object classification and document classification. The experimental results have shown better performance of our proposed EMD metric learning method compared with the traditional EMD method and the state-of-the-art methods. It is noted that the proposed EMD metric learning method can be also used in other applications.
Zizhao Zhang 0003, Yubo Zhang 0006, Xibin Zhao, Yue Gao 0002
AAAI4
2018 Hypergraph Learning With Cost Interval Optimization
abstract
In many classification tasks, the misclassification costs of different categories usually vary significantly. Under such circumstances, it is essential to identify the importance of different categories and thus assign different misclassification losses in many applications, such as medical diagnosis, saliency detection and software defect prediction. However, we note that it is infeasible to determine the accurate cost value without great domain knowledge. In most common cases, we may just have the information that which category is more important than the other categories, i.e., the identification of defect-prone softwares is more important than that of defect-free. To tackle these issues, in this paper, we propose a hypergraph learning method with cost interval optimization, which is able to handle cost interval when data is formulated using the high-order relationships. In this way, data correlations are modeled by a hypergraph structure, which has the merit to exploit the underlying relationships behind the data. With a cost-sensitive hypergraph structure, in order to improve the performance of the classifier without precise cost value, we further introduce cost interval optimization to hypergraph learning. In this process, the optimization on cost interval achieves better performance instead of choosing uncertain fixed cost in the learning process. To evaluate the effectiveness of the proposed method, we have conducted experiments on two groups of dataset, i.e., the NASA Metrics Data Program (NASA) dataset and UCI Machine Learning Repository (UCI) dataset. Experimental results and comparisons with state-of-the-art methods have exhibited better performance of our proposed method.
Xibin Zhao, Nan Wang 0015, Heyuan Shi, Hai Wan, Jin Huang 0002, Yue Gao 0002
AAAI6
2018 GVCNN: Group-View Convolutional Neural Networks for 3D Shape Recognition
abstract
3D shape recognition has attracted much attention recently. Its recent advances advocate the usage of deep features and achieve the state-of-the-art performance. However, existing deep features for 3D shape recognition are restricted to a view-to-shape setting, which learns the shape descriptor from the view-level feature directly. Despite the exciting progress on view-based 3D shape description, the intrinsic hierarchical correlation and discriminability among views have not been well exploited, which is important for 3D shape representation. To tackle this issue, in this paper, we propose a group-view convolutional neural network (GVCNN) framework for hierarchical correlation modeling towards discriminative 3D shape description. The proposed GVCNN framework is composed of a hierarchical view-group-shape architecture, i.e., from the view level, the group level and the shape level, which are organized using a grouping strategy. Concretely, we first use an expanded CNN to extract a view level descriptor. Then, a grouping module is introduced to estimate the content discrimination of each view, based on which all views can be splitted into different groups according to their discriminative level. A group level description can be further generated by pooling from view descriptors. Finally, all group level descriptors are combined into the shape level descriptor according to their discriminative weights. Experimental results and comparison with state-of-the-art methods show that our proposed GVCNN method can achieve a significant performance gain on both the 3D shape classification and retrieval tasks.
Yifan Feng 0001, Zizhao Zhang 0003, Xibin Zhao, Rongrong Ji, Yue Gao 0002
CVPR5
2018 Iterative Metric Learning for Imbalance Data Classification
abstract
In many classification applications, the amount of data from different categories usually vary significantly, such as software defect predication and medical diagnosis. Under such circumstances, it is essential to propose a proper method to solve the imbalance issue among the data. However, most of the existing methods mainly focus on improving the performance of classifiers rather than searching for an appropriate way to find an effective data space for classification. In this paper, we propose a method named Iterative Metric Learning (IML) to explore the correlations among imbalance data and construct an effective data space for classification. Given the imbalance training data, it is important to select a subset of training samples for each testing data. Thus, we aim to find a more stable neighborhood for testing data using the iterative metric learning strategy. To evaluate the effectiveness of the proposed method, we have conducted experiments on two groups of dataset, i.e., the NASA Metrics Data Program (NASA) dataset and UCI Machine Learning Repository (UCI) dataset. Experimental results and comparisons with state-of-the-art methods have exhibited better performance of our proposed method.
Nan Wang 0015, Xibin Zhao, Yue Gao 0002
IJCAI4
2018 Robust Face Sketch Synthesis via Generative Adversarial Fusion of Priors and Parametric Sigmoid
abstract
Despite the extensive progress in face sketch synthesis, existing methods are mostly workable under constrained conditions, such as fixed illumination, pose, background and ethnic origin that are hardly to control in real-world scenarios. The key issue lies in the difficulty to use data under fixed conditions to train a model against imaging variations. In this paper, we propose a novel generative adversarial network termed pGAN, which can generate face sketches efficiently using training data under fixed conditions and handle the aforementioned uncontrolled conditions. In pGAN, we embed key photo priors into the process of synthesis and design a parametric sigmoid activation function for compensating illumination variations. Compared to the existing methods, we quantitatively demonstrate that the proposed method can work well on face photos in the wild.
Shengchuan Zhang, Rongrong Ji, Jie Hu 0018, Yue Gao 0002, Chia-Wen Lin
IJCAI4
2018 Dynamic Hypergraph Structure Learning
abstract
In recent years, hypergraph modeling has shown its superiority on correlation formulation among samples and has wide applications in classification, retrieval, and other tasks. In all these works, the performance of hypergraph learning highly depends on the generated hypergraph structure. A good hypergraph structure can represent the data correlation better, and vice versa. Although hypergraph learning has attracted much attention recently, most of existing works still rely on a static hypergraph structure, and little effort concentrates on optimizing the hypergraph structure during the learning process. To tackle this problem, we propose a dynamic hypergraph structure learning method in this paper. In this method, given the originally generated hypergraph structure, the objective of our work is to simultaneously optimize the label projection matrix (the common task in hypergraph learning) and the hypergraph structure itself. More specifically, in this formulation, the label projection matrix is related to the hypergraph structure, and the hypergraph structure is associated with the data correlation from both the label space and the feature space. Here, we alternatively learn the optimal label projection matrix and the hypergraph structure, leading to a dynamic hypergraph structure during the learning process. We have applied the proposed method in the tasks of 3D shape recognition and gesture recognition. Experimental results on 4 public datasets show better performance compared with the state-of-the-art methods. We note that the proposed method can be further applied in other tasks.
Zizhao Zhang 0003, Haojie Lin, Yue Gao 0002
IJCAI3
2018 Personality-Aware Personalized Emotion Recognition from Physiological Signals
abstract
Emotion recognition methodologies from physiological signals are increasingly becoming personalized, due to the subjective responses of different subjects to physical stimuli. Existing works mainly focused on modelling the involved physiological corpus of each subject, without considering the psychological factors. The latent correlation among different subjects has also been rarely examined. We propose to investigate the influence of personality on emotional behavior in a hypergraph learning framework. Assuming that each vertex is a compound tuple (subject, stimuli), multi-modal hypergraphs can be constructed based on the personality correlation among different subjects and on the physiological correlation among corresponding stimuli. To reveal the different importance of vertices, hyperedges, and modalities, we assign each of them with weights. The emotion relevance learned on the vertex-weighted multi-modal multi-task hypergraphs is employed for emotion recognition. We carry out extensive experiments on the ASCERTAIN dataset and the results demonstrate the superiority of the proposed method.
Sicheng Zhao, Guiguang Ding, Jungong Han, Yue Gao 0002
IJCAI4
2018 Precise Point Set Registration Using Point-to-Plane Distance and Correntropy for LiDAR Based Localization
abstract
In this paper, we propose a robust point set registration algorithm which combines correntropy and point-to-plane distance, which can register rigid point sets with noises and outliers. Firstly, as correntropy performs well in handling data with non-Gaussian noises, we introduce it to model rigid point set registration problem based on point-to-plane distance; Secondly, we propose an iterative algorithm to solve this problem, which repeats to compute correspondence and transformation parameters respectively in closed form solutions. Simulated experimental results demonstrate the high precision and robustness of the proposed algorithm. In addition, LiDAR based localization experiments on automated vehicle performs satisfactory for localization accuracy and time consumption.
Guanglin Xu, Shaoyi Du, Dixiao Cui, Sirui Zhang, Badong Chen, Xuetao Zhang 0001, Jianru Xue, Yue Gao 0002
Intelligent Vehicles Symposium8
2018 PVNet: A Joint Convolutional Network of Point Cloud and Multi-View for 3D Shape Recognition
abstract
3D object recognition has attracted wide research attention in the field of multimedia and computer vision. With the recent proliferation of deep learning, various deep models with different representations have achieved the state-of-the-art performance. Among them, point cloud and multi-view based 3D shape representations are promising recently, and their corresponding deep models have shown significant performance on 3D shape recognition. However, there is little effort concentrating point cloud data and multi-view data for 3D shape representation, which is, in our consideration, beneficial and compensated to each other. In this paper, we propose the Point-View Network (PVNet), the first framework integrating both the point cloud and the multi-view data towards joint 3D shape recognition. More specifically, an embedding attention fusion scheme is proposed that could employ high-level features from the multi-view data to model the intrinsic correlation and discriminability of different structure features from the point cloud data. In particular, the discriminative descriptions are quantified and leveraged as the soft attention mask to further refine the structure feature of the 3D shape. We have evaluated the proposed method on the ModelNet40 dataset for 3D shape classification and retrieval tasks. Experimental results and comparisons with state-of-the-art methods demonstrate that our framework can achieve superior performance.
Haoxuan You, Yifan Feng 0001, Rongrong Ji, Yue Gao 0002
ACM Multimedia4
2018 Precise Point Set Registration with Color Assisted and Correntropy for 3D Reconstruction
abstract
Iterative closest point (ICP) algorithm, as its accuracy and efficiency, is widely used in rigid registration. However, ICP algorithm is easily failed when point sets lack of structure variety, such as semicircles. To solve this problem, a precise point set registration method for RGB-D data is proposed. Firstly, the color information provides a new information for registration, and the correntropy is introduced to deal with the noises and outliers. With color assisted and correntropy, a more robust objective function is built. Secondly, a variant ICP algorithm is used to deal with optimization problem via multiple iterations. Finally, as shown in the experimental results and scene reconstruction, our method obtains more precise results than other ICP algorithms.
Teng Wan, Shaoyi Du, Yiting Xu, Guanglin Xu, Yang Yang 0066, Yue Gao 0002, Badong Chen
SMC6
2018 Building Correspondence Based on Matching Triangles for Partial Registration
abstract
As an important problem in point set registration, partial registration has been solved by some variants of Iterative Closest Point (ICP) algorithm under good initial values. However, the initial parameters remained to be solved for partial registration. This paper presents a parameter initialization algorithm based on matching triangles for partial registration. Experimental results demonstrate that the proposed initialization method can find an appropriate initial transformation for next accurate registration, even the initial rotation angle between two sets is large. Based on the initialization of two point sets, the partial registration can be accomplished by auto trimmed ICP (ATICP) algorithm.
Yiting Xu, Shaoyi Du, Teng Wan, Yang Yang 0066, Badong Chen, Yue Gao 0002
SMC6
2018 Predicting Personalized Image Emotion Perceptions in Social Networks
abstract
Images can convey rich semantics and induce various emotions to viewers. Most existing works on affective image analysis focused on predicting the dominant emotions for the majority of viewers. However, such dominant emotion is often insufficient in real-world applications, as the emotions that are induced by an image are highly subjective and different with respect to different viewers. In this paper, we propose to predict the personalized emotion perceptions of images for each individual viewer. Different types of factors that may affect personalized image emotion perceptions, including visual content, social context, temporal evolution, and location influence, are jointly investigated. Rolling multi-task hypergraph learning (RMTHG) is presented to consistently combine these factors and a learning algorithm is designed for automatic optimization. For evaluation, we set up a large scale image emotion dataset from Flickr, named Image-Emotion-Social-Net, on both dimensional and categorical emotion representations with over 1 million images and about 8,000 users. Experiments conducted on this dataset demonstrate that the proposed method can achieve significant performance gains on personalized emotion classification, as compared to several state-of-the-art approaches.
Sicheng Zhao, Hongxun Yao, Yue Gao 0002, Guiguang Ding, Tat-Seng Chua
IEEE Trans. Affect. Comput.3
2018 View-Based 3-D Model Retrieval: A Benchmark
abstract
View-based 3-D model retrieval is one of the most important techniques in numerous applications of computer vision. While many methods have been proposed in recent years, to the best of our knowledge, there is no benchmark to evaluate the state-of-the-art methods. To tackle this problem, we systematically investigate and evaluate the related methods by: 1) proposing a clique graph-based method and 2) reimplementing six representative methods. Moreover, we concurrently evaluate both hand-crafted visual features and deep features on four popular datasets (NTU60, NTU216, PSB, and ETH) and one challenging real-world multiview model dataset (MV-RED) prepared by our group with various evaluation criteria to understand how these algorithms perform. By quantitatively analyzing the performances, we discover the graph matching-based method with deep features, especially the clique graph matching algorithm with convolutional neural networks features, can usually outperform the others. We further discuss the future research directions in this field.
Anan Liu, Weizhi Nie, Yue Gao 0002, Yuting Su 0001
IEEE Trans. Cybern.3
2018 Real-Time Multimedia Social Event Detection in Microblog
abstract
Detecting events from massive social media data in social networks can facilitate browsing, search, and monitoring of real-time events by corporations, governments, and users. The short, conversational, heterogeneous, and real-time characteristics of social media data bring great challenges for event detection. The existing event detection approaches rely mainly on textual information, while the visual content of microblogs and the intrinsic correlation among the heterogeneous data are scarcely explored. To deal with the above challenges, we propose a novel real-time event detection method by generating an intermediate semantic level from social multimedia data, named microblog clique (MC), which is able to explore the high correlations among different microblogs. Specifically, the proposed method comprises three stages. First, the heterogeneous data in microblogs is formulated in a hypergraph structure. Hypergraph cut is conducted to group the highly correlated microblogs with the same topics as the MCs, which can address the information inadequateness and data sparseness issues. Second, a bipartite graph is constructed based on the generated MCs and the transfer cut partition is performed to detect the events. Finally, for new incoming microblogs, incremental hypergraph is constructed based on the latest MCs to generate new MCs, which are classified by bipartite graph partition into existing events or new ones. Extensive experiments are conducted on the events in the Brand-Social-Net dataset and the results demonstrate the superiority of the proposed method, as compared to the state-of-the-art approaches.
Sicheng Zhao, Yue Gao 0002, Guiguang Ding, Tat-Seng Chua
IEEE Trans. Cybern.2
2018 Inductive Multi-Hypergraph Learning and Its Application on View-Based 3D Object Classification
abstract
The wide 3D applications have led to increasing amount of 3D object data, and thus effective 3D object classification technique has become an urgent requirement. One important and challenging task for 3D object classification is how to formulate the 3D data correlation and exploit it. Most of the previous works focus on learning optimal pairwise distance metric for object comparison, which may lose the global correlation among 3D objects. Recently, a transductive hypergraph learning has been investigated for classification, which can jointly explore the correlation among multiple objects, including both the labeled and unlabeled data. Although these methods have shown better performance, they are still limited due to 1) a considerable amount of testing data may not be available in practice and 2) the high computational cost to test new coming data. To handle this problem, considering the multi-modal representations of 3D objects in practice, we propose an inductive multi-hypergraph learning algorithm, which targets on learning an optimal projection for the multi-modal training data. In this method, all the training data are formulated in multi-hypergraph based on the features, and the inductive learning is conducted to learn the projection matrices and the optimal multi-hypergraph combination weights simultaneously. Different from the transductive learning on hypergraph, the high cost training process is off-line, and the testing process is very efficient for the inductive learning on hypergraph. We have conducted experiments on two 3D benchmarks, i.e., the NTU and the ModelNet40 data sets, and compared the proposed algorithm with the state-of-the-art methods and traditional transductive multi-hypergraph learning methods. Experimental results have demonstrated that the proposed method can achieve effective and efficient classification performance. We also note that the proposed method is a general framework and has the potential to be applied in other applications in practice.
Zizhao Zhang 0003, Haojie Lin, Xibin Zhao, Rongrong Ji, Yue Gao 0002
IEEE Trans. Image Process.5
2018 Multi-Hypergraph Learning for Incomplete Multimodality Data
abstract
Multi-modality data convey complementary information that can be used to improve the accuracy of prediction models in disease diagnosis. However, effectively integrating multi-modality data remains a challenging problem, especially when the data are incomplete. For instance, more than half of the subjects in the Alzheimer's disease neuroimaging initiative (ADNI) database have no fluorodeoxyglucose positron emission tomography and cerebrospinal fluid data. Currently, there are two commonly used strategies to handle the problem of incomplete data: 1) discard samples having missing features; and 2) impute those missing values via specific techniques. In the first case, a significant amount of useful information is lost and, in the second case, additional noise and artifacts might be introduced into the data. Also, previous studies generally focus on the pairwise relationships among subjects, without considering their underlying complex (e.g., high-order) relationships. To address these issues, in this paper, we propose a multi-hypergraph learning method for dealing with incomplete multimodality data. Specifically, we first construct multiple hypergraphs to represent the high-order relationships among subjects by dividing them into several groups according to the availability of their data modalities. A hypergraph regularized transductive learning method is then applied to these groups for automatic diagnosis of brain diseases. Extensive evaluation of the proposed method using all subjects in the baseline ADNI database indicates that our method achieves promising results in AD/MCI classification, compared with the state-of-the-art methods.
Mingxia Liu 0001, Yue Gao 0002, Pew-Thian Yap, Dinggang Shen
IEEE J. Biomed. Health Informatics2
2018 A Stacked Sparse Autoencoder-Based Detector for Automatic Identification of Neuromagnetic High Frequency Oscillations in Epilepsy
abstract
High-frequency oscillations (HFOs) are spontaneous magnetoencephalography (MEG) patterns that have been acknowledged as a putative biomarker to identify epileptic foci. Correct detection of HFOs in the MEG signals is crucial for the accurate and timely clinical evaluation. Since the visual examination of HFOs is time-consuming, error-prone, and with poor inter-reviewer reliability, an automatic HFOs detector is highly desirable in clinical practice. However, the existing approaches for HFOs detection may not be applicable for MEG signals with noisy background activity. Therefore, we employ the stacked sparse autoencoder (SSAE) and propose an SSAE-based MEG HFOs (SMO) detector to facilitate the clinical detection of HFOs. To the best of our knowledge, this is the first attempt to conduct HFOs detection in MEG using deep learning methods. After configuration optimization, our proposed SMO detector is outperformed other classic peer models by achieving 89.9% in accuracy, 88.2% in sensitivity, and 91.6% in specificity. Furthermore, we have tested the performance consistency of our model using various validation schemes. The distribution of performance metrics demonstrates that our model can achieve steady performance.
Jiayang Guo, Chunli Yin, Jing Xiang, Rongrong Ji, Yue Gao 0002
IEEE Trans. Medical Imaging8
2018 Predicting Microblog Sentiments via Weakly Supervised Multimodal Deep Learning
abstract
Predicting sentiments of multimodal microblogs composed of text, image, and emoticon have attracted ever-increasing research focus recently. The key challenge lies in the difficulty of collecting a sufficient amount of training labels to train a discriminative model for multimodal prediction. One potential solution is to exploit the labels collected from social media users, which is, however, restricted by the negative effect of label noise. Besides, we have quantitatively found that sentiments in different modalities may be independent, which disables the usage of previous multimodal sentiment analysis schemes in our problem. In this paper, we introduce a weakly supervised multimodal deep learning (WS-MDL) scheme toward robust and scalable sentiment prediction. WS-MDL learns convolutional neural networks iteratively and selectively from “weak” emoticon labels, which are cheaply available and noise containing. In particular, to filter out the label noise and to capture the modality dependency, a probabilistic graphical model is introduced to simultaneously learn discriminative multimodal descriptors and infer the confidence of label noise. Extensive evaluations are conducted in a million-scale, real-world microblog sentiment dataset crawled from Sina Weibo. We have validated the merits of the proposed scheme by quantitatively showing its superior performance over several state-of-the-art and alternative approaches.
Fuhai Chen, Rongrong Ji, Jinsong Su, Donglin Cao, Yue Gao 0002
IEEE Trans. Multim.5
2018 Beyond Pairwise Matching: Person Reidentification via High-Order Relevance Learning
abstract
Person reidentification has attracted extensive research efforts in recent years. It is challenging due to the varied visual appearance from illumination, view angle, background, and possible occlusions, leading to the difficulties when measuring the relevance, i.e., similarities, between probe and gallery images. Existing methods mainly focus on pairwise distance metric learning for person reidentification. In practice, pairwise image matching may limit the data for comparison (just the probe and one gallery subject) and yet lead to suboptimal results. The correlation among gallery data can be also helpful for the person reidentification task. In this paper, we propose to investigate the high-order correlation among the probe and gallery data, not the pairwise matching, to jointly learn the relevance of gallery data to the probe. Recalling recent progresses on feature representation in person reidentification, it is difficult to select the best feature and each type of feature can benefit person description from different aspects. Under such circumstances, we propose a multihypergraph joint learning algorithm to learn the relevance in corporation with multiple features of the imaging data. More specifically, one hypergraph is constructed using one type of feature and multiple hypergraphs can be generated accordingly. Then, the learning process is conducted on the multihypergraph structure, and the identity of a probe is determined by its relevance to each gallery data. The merit of the proposed scheme is twofold. First, different from pairwise image matching, the proposed method jointly explores the relationships among different images. Second, multimodal data, i.e., different features, can be formulated in the multihypergraph structure, which can convey more information in the learning process and can be easily extended. We note that the proposed method is a general framework to incorporate with any combination of features, and thus is flexible in practice. Experimental results and comparisons with the state-of-the-art methods on three public benchmarking data sets demonstrate the superiority of the proposed method.
Xibin Zhao, Nan Wang 0015, Yubo Zhang 0006, Shaoyi Du, Yue Gao 0002, Jia-Guang Sun 0001
IEEE Trans. Neural Networks Learn. Syst.5
2017 Active Learning with Cross-Class Similarity Transfer
abstract
How to save labeling efforts for training supervised classifiers is an important research topic in machine learning community. Active learning (AL) and transfer learning (TL) are two useful tools to achieve this goal, and their combination, i.e., transfer active learning (T-AL) has also attracted considerable research interest. However, existing T-AL approaches consider to transfer knowledge from a source/auxiliary domain which has the same class labels as the target domain, but ignore the relationship among classes. In this paper, we investigate a more practical setting where the classes in source domain are related/similar to but different from the target domain classes. Specifically, we propose a novel cross-class T-AL approach to simultaneously transfer knowledge from source domain and actively annotate the most informative samples in target domain so that we can train satisfactory classifiers with as few labeled samples as possible. In particular, based on the class-class similarity and sample-sample similarity, we adopt a similarity propagation to find the source domain samples that can well capture the characteristics of a target class and then transfer the similar samples as the (pseudo) labeled data for the target class. In turn, the labeled and transferred samples are used to train classifiers and actively select new samples for annotation. Extensive experiments on three datasets demonstrate that the proposed approach outperforms significantly the state-of-the-art related approaches.
Guiguang Ding, Yue Gao 0002, Jungong Han
AAAI3
2017 Zero-Shot Recognition via Direct Classifier Learning with Transferred Samples and Pseudo Labels
abstract
As an interesting and emerging topic, zero-shot recognition (ZSR) makes it possible to train a recognition model by specifying the category's attributes when there are no labeled exemplars available. The fundamental idea for ZSR is to transfer knowledge from the abundant labeled data in different but related source classes via the class attributes. Conventional ZSR approaches adopt a two-step strategy in test stage, where the samples are projected into the attribute space in the first step, and then the recognition is carried out based on considering the relationship between samples and classes in the attribute space. Due to this intermediate transformation, information loss is unavoidable, thus degrading the performance of the overall system. Rather than following this two-step strategy, in this paper, we propose a novel one-step approach that is able to perform ZSR in the original feature space by using directly trained classifiers. To tackle the problem that no labeled samples of target classes are available, we propose to assign pseudo labels to samples based on the reliability and diversity, which in turn will be used to train the classifiers. Moreover, we adopt a robust SVM that accounts for the unreliability of pseudo labels. Extensive experiments on four datasets demonstrate consistent performance gains of our approach over the state-of-the-art two-step ZSR approaches.
Guiguang Ding, Jungong Han, Yue Gao 0002
AAAI4
2017 SitNet: Discrete Similarity Transfer Network for Zero-shot Hashing
abstract
Hashing has been widely utilized for fast image retrieval recently. With semantic information as supervision, hashing approaches perform much better, especially when combined with deep convolution neural network(CNN). However, in practice, new concepts emerge every day, making collecting supervised information for re-training hashing model infeasible. In this paper, we propose a novel zero-shot hashing approach, called Discrete Similarity Transfer Network (SitNet), to preserve the semantic similarity between images from both ``seen'' concepts and new ``unseen'' concepts. Motivated by zero-shot learning, the semantic vectors of concepts are adopted to capture the similarity structures among classes, making the model trained with seen concepts generalize well for unseen ones benefiting from the transferability of the semantic vector space. We adopt a multi-task architecture to exploit the supervised information for seen concepts and the semantic vectors simultaneously. Moreover, a discrete hashing layer is integrated into the network for hashcode generating to avoid the information loss caused by real-value relaxation in training phase, which is a critical problem in existing works. Experiments on three benchmarks validate the superiority of SitNet to the state-of-the-arts.
Guiguang Ding, Jungong Han, Yue Gao 0002
IJCAI4
2017 Synthesizing Samples for Zero-shot Learning
abstract
Zero-shot learning (ZSL) is to construct recognition models for unseen target classes that have no labeled samples for training. It utilizes the class attributes or semantic vectors as side information and transfers supervision information from related source classes with abundant labeled samples. Existing ZSL approaches adopt an intermediary embedding space to measure the similarity between a sample and the attributes of a target class to perform zero-shot classification. However, this way may suffer from the information loss caused by the embedding process and the similarity measure cannot fully make use of the data distribution. In this paper, we propose a novel approach which turns the ZSL problem into a conventional supervised learning problem by synthesizing samples for the unseen classes. Firstly, the probability distribution of an unseen class is estimated by using the knowledge from seen classes and the class attributes. Secondly, the samples are synthesized based on the distribution for the unseen class. Finally, we can train any supervised classifiers based on the synthesized samples. Extensive experiments on benchmarks demonstrate the superiority of the proposed approach to the state-of-the-art ZSL approaches.
Guiguang Ding, Jungong Han, Yue Gao 0002
IJCAI4
2017 Vertex-Weighted Hypergraph Learning for Multi-View Object Classification
abstract
3D object classification with multi-view representation has become very popular, thanks to the progress on computer techniques and graphic hardware, and attracted much research attention in recent years. Regarding this task, there are mainly two challenging issues, i.e., the complex correlation among multiple views and the possible imbalance data issue. In this work, we propose to employ the hypergraph structure to formulate the relationship among 3D objects, taking the advantage of hypergraph on high-order correlation modelling. However, traditional hypergraph learning method may suffer from the imbalance data issue. To this end, we propose a vertex-weighted hypergraph learning algorithm for multi-view 3D object classification, introducing an updated hypergraph structure. In our method, the correlation among different objects is formulated in a hypergraph structure and each object (vertex) is associated with a corresponding weight, weighting the importance of each sample in the learning process. The learning process is conducted on the vertex-weighted hypergraph and the estimated object relevance is employed for object classification. The proposed method has been evaluated on two public benchmarks, i.e., the NTU and the PSB datasets. Experimental results and comparison with the state-of-the-art methods and recent deep learning method demonstrate the effectiveness of our proposed method.
Lifan Su, Yue Gao 0002, Xibin Zhao, Hai Wan, Ming Gu 0001, Jia-Guang Sun 0001
IJCAI2
2017 Approximating Discrete Probability Distribution of Image Emotions by Multi-Modal Features Fusion
abstract
Existing works on image emotion recognition mainly assigned the dominant emotion category or average dimension values to an image based on the assumption that viewers can reach a consensus on the emotion of images. However, the image emotions perceived by viewers are subjective by nature and highly related to the personal and situational factors. On the other hand, image emotions can be conveyed by different features, such as semantics and aesthetics. In this paper, we propose a novel machine learning approach that formulates the categorical image emotions as a discrete probability distribution (DPD). To associate emotions with the extracted visual features, we present a weighted multi-modal shared sparse leaning to learn the combination coefficients, with which the DPD of an unseen image can be predicted by linearly integrating the DPDs of the training images. The representation abilities of different modalities are jointly explored and the optimal weight of each modality is automatically learned. Extensive experiments on three datasets verify the superiority of the proposed method, as compared to the state-of-the-art.
Sicheng Zhao, Guiguang Ding, Yue Gao 0002, Jungong Han
IJCAI3
2017 TUCH: Turning Cross-view Hashing into Single-view Hashing via Generative Adversarial Nets
abstract
Cross-view retrieval, which focuses on searching images as response to text queries or vice versa, has received increasing attention recently. Cross-view hashing is to efficiently solve the cross-view retrieval problem with binary hash codes. Most existing works on cross-view hashing exploit multi-view embedding method to tackle this problem, which inevitably causes the information loss in both image and text domains. Inspired by the Generative Adversarial Nets (GANs), this paper presents a new model that is able to Turn Cross-view Hashing into single-view hashing (TUCH), thus enabling the information of image to be preserved as much as possible. TUCH is a novel deep architecture that integrates a language model network T for text feature extraction, a generator network G to generate fake images from text feature and a hashing network H for learning hashing functions to generate compact binary codes. Our architecture effectively unifies joint generative adversarial learning and cross-view hashing. Extensive empirical evidence shows that our TUCH approach achieves state-of-the-art results, especially on text to image retrieval, based on image-sentences datasets, i.e. standard IAPRTC-12 and large-scale Microsoft COCO.
Xin Zhao 0020, Guiguang Ding, Jungong Han, Yue Gao 0002
IJCAI5
2017 Handling scheduling uncertainties through traffic shaping in Time-Triggered train networks
abstract
While trains traditionally relied on field bus to support real-time control applications, next-generation trains are moving toward Ethernet as an integrated, high-bandwidth communication infrastructure for real-time control and best-effort consumer traffic. Time-Triggered Ethernet (TT-Ethernet) is a promising technology for train networks because of its capability to achieve deterministic latencies for real-time applications based on pre-computed transmission schedules. However, the deterministic scheduling approach of TT-Ethernet faces significant challenges in handling scheduling uncertainties caused by switch failures and legacy end devices in train networks. Due to the physical constraints on trains, train networks deal with switch failures by bypassing failed switches using a short circuiting mechanism. Unfortunately, this mechanism incurs scheduling errors as frames bypassing the failed switch may arrive ahead of the pre-computed schedule, resulting in early, unexpected, and out of order arrivals. Furthermore, as trains evolve from traditional communication technologies to TT-Ethernet, the network must support legacy end devices that may generate frames at times unknown to the TT-Ethernet. We propose a novel traffic shaping approach to deal with scheduling uncertainties in TT-Ethernet. The traffic shaper of a TT-Ethernet switch buffers early frames and then releases them at their pre-scheduled arrive time. Furthermore, we devise an efficient buffer management method for the traffic shaper in face of fault scenarios. Finally, we use the traffic shaper to integrate legacy devices into TT-Ethernet. We have implemented the traffic shaping approach in a 24-port TT-Ethernet switch specifically designed for train networks. Experiments show the traffic shaping strategy can effectively deal with scheduling uncertainties incurred by switch failures and legacy devices.
Qinghan Yu, Xibin Zhao, Hai Wan, Yue Gao 0002, Chenyang Lu 0001, Ming Gu 0001
IWQoS4
2017 Learning Visual Emotion Distributions via Multi-Modal Features Fusion
abstract
Current image emotion recognition works mainly classified the images into one dominant emotion category, or regressed the images with average dimension values by assuming that the emotions perceived among different viewers highly accord with each other. However, due to the influence of various personal and situational factors, such as culture background and social interactions, different viewers may react totally different from the emotional perspective to the same image. In this paper, we propose to formulate the image emotion recognition task as a probability distribution learning problem. Motivated by the fact that image emotions can be conveyed through different visual features, such as aesthetics and semantics, we present a novel framework by fusing multi-modal features to tackle this problem. In detail, weighted multi-modal conditional probability neural network (WMMCPNN) is designed as the learning model to associate the visual features with emotion probabilities. By jointly exploring the complementarity and learning the optimal combination coefficients of different modality features, WMMCPNN could effectively utilize the representation ability of each uni-modal feature. We conduct extensive experiments on three publicly available benchmarks and the results demonstrate that the proposed method significantly outperforms the state-of-the-art approaches for emotion distribution prediction.
Sicheng Zhao, Guiguang Ding, Yue Gao 0002, Jungong Han
ACM Multimedia3
2017 Zero-Shot Learning With Transferred Samples
abstract
By transferring knowledge from the abundant labeled samples of known source classes, zero-shot learning (ZSL) makes it possible to train recognition models for novel target classes that have no labeled samples. Conventional ZSL approaches usually adopt a two-step recognition strategy, in which the test sample is projected into an intermediary space in the first step, and then the recognition is carried out by considering the similarity between the sample and target classes in the intermediary space. Due to this redundant intermediate transformation, information loss is unavoidable, thus degrading the performance of overall system. Rather than adopting this two-step strategy, in this paper, we propose a novel one-step recognition framework that is able to perform recognition in the original feature space by using directly trained classifiers. To address the lack of labeled samples for training supervised classifiers for the target classes, we propose to transfer samples from source classes with pseudo labels assigned, in which the transferred samples are selected based on their transferability and diversity. Moreover, to account for the unreliability of pseudo labels of transferred samples, we modify the standard support vector machine formulation such that the unreliable positive samples can be recognized and suppressed in the training phase. The entire framework is fairly general with the possibility of further extensions to several common ZSL settings. Extensive experiments on four benchmark data sets demonstrate the superiority of the proposed framework, compared with the state-of-the-art approaches, in various settings.
Guiguang Ding, Jungong Han, Yue Gao 0002
IEEE Trans. Image Process.4
2017 Event Classification in Microblogs via Social Tracking
abstract
Social media websites have become important information sharing platforms. The rapid development of social media platforms has led to increasingly large-scale social media data, which has shown remarkable societal and marketing values. There are needs to extract important events in live social media streams. However, microblogs event classification is challenging due to two facts, i.e., the short/conversational nature and the incompatible meanings between the text and the corresponding image in social posts, and the rapidly evolving contents. In this article, we propose to conduct event classification via deep learning and social tracking. First, we introduce a Multi-modal Multi-instance Deep Network (M 2 DN) for microblogs classification, which is able to handle the weakly labeled microblogs data oriented from the incompatible meanings inside microblogs. Besides predicting each microblogs as predefined events, we propose to employ social tracking to extract social-related auxiliary information to enrich the testing samples. We extract a set of candidate-relevant microblogs in a short time window by using social connections, such as related users and geographical locations. All these selected microblogs and the testing data are formulated in a Markov Random Field model. The inference on the Markov Random Field is conducted to update the classification results of the testing microblogs. This method is evaluated on the Brand-Social-Net dataset for classification of 20 events. Experimental results and comparison with the state of the arts show that the proposed method can achieve better performance for the event classification task.
Yue Gao 0002, Hanwang Zhang, Xibin Zhao, Shuicheng Yan
ACM Trans. Intell. Syst. Technol.1
2017 Continuous Probability Distribution Prediction of Image Emotions via Multitask Shared Sparse Regression
abstract
Previous works on image emotion analysis mainly focused on predicting the dominant emotion category or the average dimension values of an image for affective image classification and regression. However, this is often insufficient in various real-world applications, as the emotions that are evoked in viewers by an image are highly subjective and different. In this paper, we propose to predict the continuous probability distribution of image emotions which are represented in dimensional valence-arousal space. We carried out large-scale statistical analysis on the constructed Image-Emotion-Social-Net dataset, on which we observed that the emotion distribution can be well-modeled by a Gaussian mixture model. This model is estimated by an expectation-maximization algorithm with specified initializations. Then, we extract commonly used emotion features at different levels for each image. Finally, we formalize the emotion distribution prediction task as a shared sparse regression (SSR) problem and extend it to multitask settings, named multitask shared sparse regression (MTSSR), to explore the latent information between different prediction tasks. SSR and MTSSR are optimized by iteratively reweighted least squares. Experiments are conducted on the Image-Emotion-Social-Net dataset with comparisons to three alternative baselines. The quantitative results demonstrate the superiority of the proposed method.
Sicheng Zhao, Hongxun Yao, Yue Gao 0002, Rongrong Ji, Guiguang Ding
IEEE Trans. Multim.3
2016 Search-Based Depth Estimation via Coupled Dictionary Learning with Large-Margin Structure Inference
Yan Zhang 0109, Rongrong Ji, Xiaopeng Fan 0001, Yan Wang 0059, Feng Guo 0005, Yue Gao 0002, Debin Zhao
ECCV (5)6
2016 Semi-Supervised Active Learning with Cross-Class Sample Transfer
Guiguang Ding, Yue Gao 0002, Jianmin Wang 0001
IJCAI3
2016 Identifying Relationships in Functional and Structural Connectome Data Using a Hypergraph Learning Method
Brent C. Munsell, Guorong Wu 0001, Yue Gao 0002, Nicholas Desisto, Martin Styner
MICCAI (2)3
2016 Predicting Personalized Emotion Perceptions of Social Images
abstract
Images can convey rich semantics and induce various emotions to viewers. Most existing works on affective image analysis focused on predicting the dominant emotions for the majority of viewers. However, such dominant emotion is often insufficient in real-world applications, as the emotions that are induced by an image are highly subjective and different with respect to different viewers. In this paper, we propose to predict the personalized emotion perceptions of images for each individual viewer. Different types of factors that may affect personalized image emotion perceptions, including visual content, social context, temporal evolution, and location influence, are jointly investigated. Rolling multi-task hypergraph learning is presented to consistently combine these factors and a learning algorithm is designed for automatic optimization. For evaluation, we set up a large scale image emotion dataset from Flickr, named Image-Emotion-Social-Net, on both dimensional and categorical emotion representations with over 1 million images and about 8,000 users. Experiments conducted on this dataset demonstrate that the proposed method can achieve significant performance gains on personalized emotion classification, as compared to several state-of-the-art approaches.
Sicheng Zhao, Hongxun Yao, Yue Gao 0002, Rongrong Ji, Wenlong Xie, Xiaolei Jiang, Tat-Seng Chua
ACM Multimedia3
2016 Special issue: When social media meets physical world
Rongrong Ji, Yue Gao 0002, Qi Tian 0001, Qionghai Dai, Ralf Steinmetz
Multim. Syst.2
2016 Recent advances in social multimedia big data mining and applications
Yue Gao 0002, Bing-Kun Bao, Cees Snoek, Qionghai Dai
Multim. Syst.2
2016 Guest Editorial: Image Analysis and Processing Leveraging Additional Information
Luis Herranz, Jian Cheng 0001, Yue Gao 0002, Shuqiang Jiang
Multim. Tools Appl.3
2016 Spectral Multimodal Hashing and Its Application to Multimedia Retrieval
abstract
In recent years, multimedia retrieval has sparked much research interest in the multimedia, pattern recognition, and data mining communities. Although some attempts have been made along this direction, performing fast multimodal search at very large scale still remains a major challenge in the area. While hashing-based methods have recently achieved promising successes in speeding-up large-scale similarity search, most existing methods are only designed for uni-modal data, making them unsuitable for multimodal multimedia retrieval. In this paper, we propose a new hashing-based method for fast multimodal multimedia retrieval. The method is based on spectral analysis of the correlation matrix of different modalities. We also develop an efficient algorithm that learns some parameters from the data distribution for obtaining the binary codes. We empirically compare our method with some state-of-the-art methods on two real-world multimedia data sets.
Yi Zhen, Yue Gao 0002, Dit-Yan Yeung, Hongyuan Zha, Xuelong Li 0001
IEEE Trans. Cybern.2
2016 Large-Scale Cross-Modality Search via Collective Matrix Factorization Hashing
abstract
By transforming data into binary representation, i.e., Hashing, we can perform high-speed search with low storage cost, and thus, Hashing has collected increasing research interest in the recent years. Recently, how to generate Hashcode for multimodal data (e.g., images with textual tags, documents with photos, and so on) for large-scale cross-modality search (e.g., searching semantically related images in database for a document query) is an important research issue because of the fast growth of multimodal data in the Web. To address this issue, a novel framework for multimodal Hashing is proposed, termed as Collective Matrix Factorization Hashing (CMFH). The key idea of CMFH is to learn unified Hashcodes for different modalities of one multimodal instance in the shared latent semantic space in which different modalities can be effectively connected. Therefore, accurate cross-modality search is supported. Based on the general framework, we extend it in the unsupervised scenario where it tries to preserve the Euclidean structure, and in the supervised scenario where it fully exploits the label information of data. The corresponding theoretical analysis and the optimization algorithms are given. We conducted comprehensive experiments on three benchmark data sets for cross-modality search. The experimental results demonstrate that CMFH can significantly outperform several state-of-the-art cross-modality Hashing methods, which validates the effectiveness of the proposed CMFH.
Guiguang Ding, Jile Zhou, Yue Gao 0002
IEEE Trans. Image Process.4
2016 Multi-View 3D Object Retrieval With Deep Embedding Network
abstract
In multi-view 3D object retrieval, each object is characterized by a group of 2D images captured from different views. Rather than using hand-crafted features, in this paper, we take advantage of the strong discriminative power of convolutional neural network to learn an effective 3D object representation tailored for this retrieval task. Specifically, we propose a deep embedding network jointly supervised by classification loss and triplet loss to map the high-dimensional image space into a low-dimensional feature space, where the Euclidean distance of features directly corresponds to the semantic similarity of images. By effectively reducing the intra-class variations while increasing the inter-class ones of the input images, the network guarantees that similar images are closer than dissimilar ones in the learned feature space. Besides, we investigate the effectiveness of deep features extracted from different layers of the embedding network extensively and find that an efficient 3D object representation should be a tradeoff between global semantic information and discriminative local characteristics. Then, with the set of deep features extracted from different views, we can generate a comprehensive description for each 3D object and formulate the multi-view 3D object retrieval as a set-to-set matching problem. Extensive experiments on SHREC'15 data set demonstrate the superiority of our proposed method over the previous state-of-the-art approaches with over 12% performance improvement.
Haiyun Guo, Jinqiao Wang, Yue Gao 0002, Jianqiang Li 0002, Hanqing Lu
IEEE Trans. Image Process.3
2016 Multi-Modal Clique-Graph Matching for View-Based 3D Model Retrieval
abstract
Multi-view matching is an important but a challenging task in view-based 3D model retrieval. To address this challenge, we propose an original multi-modal clique graph (MCG) matching method in this paper. We systematically present a method for MCG generation that is composed of cliques, which consist of neighbor nodes in multi-modal feature space and hyper-edges that link pairwise cliques. Moreover, we propose an image set-based clique/edgewise similarity measure to address the issue of the set-to-set distance measure, which is the core problem in MCG matching. The proposed MCG provides the following benefits: 1) preserves the local and global attributes of a graph with the designed structure; 2) eliminates redundant and noisy information by strengthening inliers while suppressing outliers; and 3) avoids the difficulty of defining high-order attributes and solving hyper-graph matching. We validate the MCG-based 3D model retrieval using three popular single-modal data sets and one novel multi-modal data set. Extensive experiments show the superiority of the proposed method through comparisons. Moreover, we contribute a novel real-world 3D object data set, the multi-view RGB-D object data set. To the best of our knowledge, it is the largest real-world 3D object data set containing multi-modal and multi-view information.
Anan Liu, Weizhi Nie, Yue Gao 0002, Yuting Su 0001
IEEE Trans. Image Process.3
2016 Detecting Anatomical Landmarks for Fast Alzheimer's Disease Diagnosis
abstract
Structural magnetic resonance imaging (MRI) is a very popular and effective technique used to diagnose Alzheimer's disease (AD). The success of computer-aided diagnosis methods using structural MRI data is largely dependent on the two time-consuming steps: 1) nonlinear registration across subjects, and 2) brain tissue segmentation. To overcome this limitation, we propose a landmark-based feature extraction method that does not require nonlinear registration and tissue segmentation. In the training stage, in order to distinguish AD subjects from healthy controls (HCs), group comparisons, based on local morphological features, are first performed to identify brain regions that have significant group differences. In general, the centers of the identified regions become landmark locations (or AD landmarks for short) capable of differentiating AD subjects from HCs. In the testing stage, using the learned AD landmarks, the corresponding landmarks are detected in a testing image using an efficient technique based on a shape-constrained regression-forest algorithm. To improve detection accuracy, an additional set of salient and consistent landmarks are also identified to guide the AD landmark detection. Based on the identified AD landmarks, morphological features are extracted to train a support vector machine (SVM) classifier that is capable of predicting the AD condition. In the experiments, our method is evaluated on landmark detection and AD classification sequentially. Specifically, the landmark detection error (manually annotated versus automatically detected) of the proposed landmark detector is 2.41 mm , and our landmark-based AD classification accuracy is 83.7%. Lastly, the AD classification performance of our method is comparable to, or even better than, that achieved by existing region-based and voxel-based methods, while the proposed method is approximately 50 times faster.
Jun Zhang 0018, Yue Gao 0002, Yaozong Gao, Brent C. Munsell, Dinggang Shen
IEEE Trans. Medical Imaging2
2016 Filtering of Brand-Related Microblogs Using Social-Smooth Multiview Embedding
abstract
In recent years, we have witnessed the boom of social media platforms, through which people have been generating a lot of social media data. This data touches almost every aspect of life and may have significant societal and marketing values for a variety of corporations and organizations. Thus, the development of effective techniques for gathering and analyzing social media content has attracted much research attention. As social media data tend to be heterogeneous, conversational, and fast evolving in content, a recent work reported a multifaceted approach to gather comprehensive brand-related data by crawling data using evolving keywords, key users, similar image content, and known locations. Although such approach has been found to be effective in gathering representative data, it also brings in a lot of noise. This paper aims to develop an accurate classifier to filter out noise by taking into account the multimedia content and social nature of brand-related data. In particular, we develop a microblog filtering method based on a discriminative social-aware multiview embedding. Besides the conventional content-based features, such as textual, low-level visual features, and high-level visual semantic features, that form the three key views of microblogs, we also incorporate the brand and social relations among the microblogs to learn a discriminative and social-aware embedding. With such a learned embedding, an off-the-shelf classifier, such as SVM, can then be trained and applied to microblog filtering. We verify the efficacy of our method on noise filtering in the brand data gathering task on the Brand-Social-Net dataset. Our approach is able to achieve significantly better filtering performance and improve the quality of brand data gathering.
Yue Gao 0002, Yi Zhen, Tat-Seng Chua
IEEE Trans. Multim.1
2016 Estimating 3D Gaze Directions Using Unlabeled Eye Images via Synthetic Iris Appearance Fitting
abstract
Estimating three-dimensional (3D) human eye gaze by capturing a single eye image without active illumination is challenging. Although the elliptical iris shape provides a useful cue, existing methods face difficulties in ellipse fitting due to unreliable iris contour detection. These methods may fail frequently especially with low resolution eye images. In this paper, we propose a synthetic iris appearance fitting (SIAF) method that is model-driven to compute 3D gaze direction from iris shape. Instead of fitting an ellipse based on exactly detected iris contour, our method first synthesizes a set of physically possible iris appearances and then optimizes inside this synthetic space to find the best solution to explain the captured eye image. In this way, the solution is highly constrained and guaranteed to be physically feasible. In addition, the proposed advanced image analysis techniques also help the SIAF method be robust to the unreliable iris contour detection. Furthermore, with multiple eye images, we propose a SIAF-joint method that can further reduce the gaze error by half, and it also resolves the binary ambiguity which is inevitable in conventional methods based on simple ellipse fitting.
Feng Lu 0005, Yue Gao 0002, Xiaowu Chen 0001
IEEE Trans. Multim.2
2016 Image Categorization by Learning a Propagated Graphlet Path
abstract
Spatial pyramid matching is a standard architecture for categorical image retrieval. However, its performance is largely limited by the prespecified rectangular spatial regions when pooling local descriptors. In this paper, we propose to learn object-shaped and directional receptive fields for image categorization. In particular, different objects in an image are seamlessly constructed by superpixels, while the direction captures human gaze shifting path. By generating a number of superpixels in each image, we construct graphlets to describe different objects. They function as the object-shaped receptive fields for image comparison. Due to the huge number of graphlets in an image, a saliency-guided graphlet selection algorithm is proposed. A manifold embedding algorithm encodes graphlets with the semantics of training image tags. Then, we derive a manifold propagation to calculate the postembedding graphlets by leveraging visual saliency maps. The sequentially propagated graphlets constitute a path that mimics human gaze shifting. Finally, we use the learned graphlet path as receptive fields for local image descriptor pooling. The local descriptors from similar receptive fields of pairwise images more significantly contribute to the final image kernel. Thorough experiments demonstrate the advantage of our approach.
Richang Hong, Yue Gao 0002, Rongrong Ji, Qionghai Dai, Xuelong Li 0001
IEEE Trans. Neural Networks Learn. Syst.3
2015 Multimodal hypergraph learning for microblog sentiment prediction
abstract
Microblog sentiment analysis has attracted extensive research attention in the recent literature. However, most existing works mainly focus on the textual modality, while ignore the contribution of visual information that contributes ever increasing proportion in expressing user emotions. In this paper, we propose to employ a hypergraph structure to formulate textual, visual and emoticon information jointly for sentiment prediction. The constructed hypergraph captures the similarities of tweets on different modalities where each vertex represents a tweet and the hyperedge is formed by the “centroid” vertex and its k-nearest neighbors on each modality. Then, the transductive inference is conducted to learn the relevance score among tweets for sentiment prediction. In this way, both intra- and inter- modality dependencies are taken into consideration in sentiment prediction. Experiments conducted on over 6,000 microblog tweets demonstrate the superiority of our method by 86.77% accuracy and 7% improvement compared to the state-of-the-art methods.
Fuhai Chen, Yue Gao 0002, Donglin Cao, Rongrong Ji
ICME2
2015 Medical Image Retrieval Using Multi-graph Learning for MCI Diagnostic Assistance
Yue Gao 0002, Ehsan Adeli-Mosabbeb, Minjeong Kim 0001, Panteleimon Giannakopoulos, Sven Haller, Dinggang Shen
MICCAI (2)1
2015 MCI Identification by Joint Learning on Multiple MRI Data
Yue Gao 0002, Chong-Yaw Wee, Minjeong Kim 0001, Panteleimon Giannakopoulos, Marie-Louise Montandon, Sven Haller, Dinggang Shen
MICCAI (2)1
2015 Multimedia Social Event Detection in Microblog
Yue Gao 0002, Sicheng Zhao, Yang Yang 0002, Tat-Seng Chua
MMM (1)1
2015 Learning for 3D understanding
Yue Gao 0002, Rongrong Ji, Wei Liu 0005, Qionghai Dai
Neurocomputing1
2015 Signal processing and learning methods for 3D semantic analysis
Yue Gao 0002, Rongrong Ji, Xinbo Gao 0001, Pierre-Marc Jodoin
Signal Process.1
2015 Depth Error Elimination for RGB-D Cameras
abstract
The rapid spreading of RGB-D cameras has led to wide applications of 3D videos in both academia and industry, such as 3D entertainment and 3D visual understanding. Under these circumstances, extensive research efforts have been dedicated to RGB-D camera--oriented topics. In these topics, quality promotion of depth videos with the temporal characteristic is emerging and important. Due to the limited exposure time of RGB-D cameras, object movement can easily lead to motion blurs in intensive images, which can further result in obvious artifacts (holes or fake boundaries) in the corresponding depth frames. With regard to this problem, we propose a depth error elimination method based on time series analysis to remove the artifacts in depth images. In this method, we first locate the regions with erroneous depths in intensive images by using motion blur detection based on a time series analysis model. This is based on the fact that the depth image is calculated by intensive color images that are captured synchronously by RGB-D cameras. Then, the artifacts, such as holes or fake boundaries, are fixed by a depth error elimination method. To evaluate the performance of the proposed method, we conducted experiments on 250 images. Experimental results demonstrate that the proposed method can locate the error regions correctly and eliminate these artifacts effectively. The quality of depth video can be improved significantly by using the proposed method.
Yue Gao 0002, You Yang 0002, Yi Zhen, Qionghai Dai
ACM Trans. Intell. Syst. Technol.1
2015 When Location Meets Social Multimedia: A Survey on Vision-Based Recognition and Mining for Geo-Social Multimedia Analytics
abstract
Coming with the popularity of multimedia sharing platforms such as Facebook and Flickr, recent years have witnessed an explosive growth of geographical tags on social multimedia content. This trend enables a wide variety of emerging applications, for example, mobile location search, landmark recognition, scene reconstruction, and touristic recommendation, which range from purely research prototype to commercial systems. In this article, we give a comprehensive survey on these applications, covering recent advances in recognition and mining of geographical-aware social multimedia. We review related work in the past decade regarding to location recognition, scene summarization, tourism suggestion, 3D building modeling, mobile visual search and city navigation. At the end, we further discuss potential challenges, future topics, as well as open issues related to geo-social multimedia computing, recognition, mining, and analytics.
Rongrong Ji, Yue Gao 0002, Wei Liu 0005, Xing Xie 0001, Qi Tian 0001, Xuelong Li 0001
ACM Trans. Intell. Syst. Technol.2
2015 Semantic-Based Location Recommendation With Multimodal Venue Semantics
abstract
In recent years, we have witnessed a flourishing of location -based social networks. A well-formed representation of location knowledge is desired to cater to the need of location sensing, browsing, navigation and querying. In this paper, we aim to study the semantics of point-of-interest (POI) by exploiting the abundant heterogeneous user generated content (UGC) from different social networks. Our idea is to explore the text descriptions, photos, user check-in patterns, and venue context for location semantic similarity measurement. We argue that the venue semantics play an important role in user check-in behavior. Based on this argument, a unified POI recommendation algorithm is proposed by incorporating venue semantics as a regularizer. In addition to deriving user preference based on user-venue check-in information, we place special emphasis on location semantic similarity. Finally, we conduct a comprehensive performance evaluation of location semantic similarity and location recommendation over a real world dataset collected from Foursquare and Instagram. Experimental results show that the UGC information can well characterize the venue semantics, which help to improve the recommendation performance.
Yi-Liang Zhao, Liqiang Nie, Yue Gao 0002, Weizhi Nie, Zhengjun Zha, Tat-Seng Chua
IEEE Trans. Multim.4
2015 Corrections to "Exploiting Web Images for Semantic Video Indexing Via Robust Sample-Specific Loss"
abstract
In the above paper [ibid, vol. 16, no. 6, p. 1679, Oct. 2014], the sentence "Yang et al. [?] integrated semisupervised learning and transfer learning techniques to exploit manually-labeled images for video tagging" should have appeared as "Yang et al. [34] integrated semisupervised learning and transfer learning techniques to exploit manually-labeled images for video tagging."
Yang Yang 0002, Zhengjun Zha, Yue Gao 0002, Xiaofeng Zhu 0001, Tat-Seng Chua
IEEE Trans. Multim.3
2015 Probabilistic Skimlets Fusion for Summarizing Multiple Consumer Landmark Videos
abstract
It is difficult to develop a computational model that can accurately predict the quality of the video summary. This paper proposes a novel algorithm to summarize one-shot landmark videos. The algorithm can optimally combine multiple unedited consumer video skims into an aesthetically pleasing summary. In particular, to effectively select the representative key frames from multiple videos, an active learning algorithm is derived by taking advantage of the locality of the frames within each video. Toward a smooth video summary, we define skimlet, a video clip with adjustable length, starting frame, and positioned by each skim. Thereby, a probabilistic framework is developed to transfer the visual cues from a collection of aesthetically pleasing photos into the video summary. The length and the starting frame of each skimlet are calculated to maximally smoothen the video summary. At the same time, the unstable frames are removed from each skimlet. Experiments on multiple videos taken from different sceneries demonstrated the aesthetics, the smoothness, and the stability of the generated summary.
Yue Gao 0002, Richang Hong, Yuxing Hu, Rongrong Ji, Qionghai Dai
IEEE Trans. Multim.2
2014 Brand Data Gathering From Live Social Media Streams
abstract
Social media streams, such as Twitter, Facebook, and Sina Weibo, have become essential real-time information resources with a wide range of users and applications. The rapidly increasing amount of live information in social media streams has important societal and marketing values for large corporations and government organizations. There is a strong need for effective techniques for data gathering and content analysis. This problem is particularly challenging due to the short and conversational nature of posts, the huge data volume, and the increasing heterogeneous multimedia content in social media streams. Moreover, as the focus of "conversation" often shifts quickly in social media space, the traditional keywords based approach to gather data with respect to a target brand is grossly inadequate. To address these problems, we propose a multi-faceted brand tracking method that gathers relevant data based on not just evolving keywords, but also social factors (users, relations and locations) as well as visual contents as increasing number of social media posts are in multimedia form. For evaluation, we set up a large scale microblog dataset (Brand-Social-Net) on brand/product information, containing 3 million microblogs with over 1.2 million images for 100 famous brands. Experiments on this dataset have demonstrated that the proposed framework is able to gather a more complete set of relevant brand-related data from live social media streams. We have released this dataset to promote social media research.
Yue Gao 0002, Fanglin Wang, Huan-Bo Luan, Tat-Seng Chua
ICMR1
2014 Image Tagging with Social Assistance
abstract
Image tagging, also known as image annotation and image conception detection, has been extensively studied in the literature. However, most existing approaches can hardly achieve satisfactory performance owing to the deficiency and unreliability of the manually-labeled training data. In this paper, we propose a new image tagging scheme, termed social assisted media tagging (SAMT), which leverages the abundant user-generated images and the associated tags as the "social assistance" to learn the classifiers. We focus on addressing the following major challenges: (a) the noisy tags associated to the web images; and (b) the desirable robustness of the tagging model. We present a joint image tagging framework which simultaneously refines the erroneous tags of the web images as well as learns the reliable image classifiers. In particular, we devise a novel tag refinement module for identifying and eliminating the noisy tags by substantially exploring and preserving the low-rank nature of the tag matrix and the structured sparse property of the tag errors. We develop a robust image tagging module based on the l2,p-norm for learning the reliable image classifiers. The correlation of the two modules is well explored within the joint framework to reinforce each other. Extensive experiments on two real-world social image databases illustrate the superiority of the proposed approach as compared to the existing methods.
Yang Yang 0002, Yue Gao 0002, Hanwang Zhang, Jie Shao 0001, Tat-Seng Chua
ICMR2
2014 Perception-Guided Multimodal Feature Fusion for Photo Aesthetics Assessment
abstract
Photo aesthetic quality evaluation is a challenging task in multimedia and computer vision fields. Conventional approaches suffer from the following three drawbacks: 1) the deemphasized role of semantic content that is many times more important than low-level visual features in photo aesthetics; 2) the difficulty to optimally fuse low-level and high-level visual cues in photo aesthetics evaluation; and 3) the absence of a sequential viewing path in the existing models, as humans perceive visually salient regions sequentially when viewing a photo.
Yue Gao 0002, Chao Zhang 0014, Hanwang Zhang, Qi Tian 0001, Roger Zimmermann
ACM Multimedia2
2014 Exploring Principles-of-Art Features For Image Emotion Recognition
abstract
Emotions can be evoked in humans by images. Most previous works on image emotion analysis mainly used the elements-of-art-based low-level visual features. However, these features are vulnerable and not invariant to the different arrangements of elements. In this paper, we investigate the concept of principles-of-art and its influence on image emotions. Principles-of-art-based emotion features (PAEF) are extracted to classify and score image emotions for understanding the relationship between artistic principles and emotions. PAEF are the unified combination of representation features derived from different principles, including balance, emphasis, harmony, variety, gradation, and movement. Experiments on the International Affective Picture System (IAPS), a set of artistic photography and a set of peer rated abstract paintings, demonstrate the superiority of PAEF for affective image classification and regression (with about 5% improvement on classification accuracy and 0.2 decrease in mean squared error), as compared to the state-of-the-art approaches. We then utilize PAEF to analyze the emotions of master paintings, with promising results.
Sicheng Zhao, Yue Gao 0002, Xiaolei Jiang, Hongxun Yao, Tat-Seng Chua, Xiaoshuai Sun
ACM Multimedia2
2014 Improved and Promising Identificationof Human MicroRNAs by Incorporatinga High-Quality Negative Set
abstract
MicroRNA (miRNA) plays an important role as a regulator in biological processes. Identification of (pre-) miRNAs helps in understanding regulatory processes. Machine learning methods have been designed for pre-miRNA identification. However, most of them cannot provide reliable predictive performances on independent testing data sets. We assumed this is because the training sets, especially the negative training sets, are not sufficiently representative. To generate a representative negative set, we proposed a novel negative sample selection technique, and successfully collected negative samples with improved quality. Two recent classifiers rebuilt with the proposed negative set achieved an improvement of ~6 percent in their predictive performance, which confirmed this assumption. Based on the proposed negative set, we constructed a training set, and developed an online system called miRNApre specifically for human pre-miRNA identification. We showed that miRNApre achieved accuracies on updated human and non-human data sets that were 34.3 and 7.6 percent higher than those achieved by current methods. The results suggest that miRNApre is an effective tool for pre-miRNA identification. Additionally, by integrating miRNApre, we developed a miRNA mining tool, mirnaDetect, which can be applied to find potential miRNAs in genome-scale data. MirnaDetect achieved a comparable mining performance on human chromosome 19 data as other existing methods.
Leyi Wei, Minghong Liao, Yue Gao 0002, Rongrong Ji, Zengyou He, Quan Zou 0001
IEEE ACM Trans. Comput. Biol. Bioinform.3
2014 Symbiotic Tracker Ensemble Toward A Unified Tracking Framework
abstract
Tracking people and objects is a fundamental stage toward many video surveillance systems, for which various trackers have been specifically designed in the past decade. However, it comes to a consensus that there is not any specific tracker that works sufficiently well under all circumstances. Therefore, one potential solution is to deploy multiple trackers, with a tracker output fusion step to boost the overall performance. Subsequently, an intelligent fusion design, yet general and orthogonal to any specific tracker, plays a key role in successful tracking. In this paper, we propose a symbiotic tracker ensemble toward a unified tracking framework, which is based on only the output of each individual tracker, without knowing its specific mechanism. In our approach, all trackers run in parallel, without requiring any details for tracker running, which means that all trackers are treated as black boxes. The proposed symbiotic tracker ensemble framework aims at learning an optimal combination of these tracking results. Our method captures the relation among individual trackers robustly from two aspects. First, the consistency between two successive frames is calculated for each tracker. Then, the pair-wise correlation among different trackers is estimated in the new coming frame by a graph-propagation process. Experimental results on the Caremedia dataset and the Caviar dataset demonstrate the effectiveness of the proposed method, with comparisons to several state-of-the-art methods.
Yue Gao 0002, Rongrong Ji, Alex Hauptmann 0001
IEEE Trans. Circuits Syst. Video Technol.1
2014 Image Annotation by Multiple-Instance Learning With Discriminative Feature Mapping and Selection
abstract
Multiple-instance learning (MIL) has been widely investigated in image annotation for its capability of exploring region-level visual information of images. Recent studies show that, by performing feature mapping, MIL can be cast to a single-instance learning problem and, thus, can be solved by traditional supervised learning methods. However, the approaches for feature mapping usually overlook the discriminative ability and the noises of the generated features. In this paper, we propose an MIL method with discriminative feature mapping and feature selection, aiming at solving this problem. Our method is able to explore both the positive and negative concept correlations. It can also select the effective features from a large and diverse set of low-level features for each concept under MIL settings. Experimental results and comparison with other methods demonstrate the effectiveness of our approach.
Richang Hong, Meng Wang 0001, Yue Gao 0002, Dacheng Tao, Xuelong Li 0001, Xindong Wu 0001
IEEE Trans. Cybern.3
2014 Spectral-Spatial Constraint Hyperspectral Image Classification
abstract
Hyperspectral image classification has attracted extensive research efforts in the recent decade. The main difficulty lies in the few labeled samples versus the high dimensional features. To this end, it is a fundamental step to explore the relationship among different pixels in hyperspectral image classification, toward jointly handing both the lack of label and high dimensionality problems. In the hyperspectral images, the classification task can be benefited from the spatial layout information. In this paper, we propose a hyperspectral image classification method to address both the pixel spectral and spatial constraints, in which the relationship among pixels is formulated in a hypergraph structure. In the constructed hypergraph, each vertex denotes a pixel in the hyperspectral image. And the hyperedges are constructed from both the distance between pixels in the feature space and the spatial locations of pixels. More specifically, a feature-based hyperedge is generated by using distance among pixels, where each pixel is connected with its K nearest neighbors in the feature space. Second, a spatial-based hyperedge is generated to model the layout among pixels by linking where each pixel is linked with its spatial local neighbors. Both the learning on the combinational hypergraph is conducted by jointly investigating the image feature and the spatial layout of pixels to seek their joint optimal partitions. Experiments on four data sets are performed to evaluate the effectiveness and and efficiency of the proposed method. Comparisons to the state-of-the-art methods demonstrate the superiority of the proposed method in the hyperspectral image classification.
Rongrong Ji, Yue Gao 0002, Richang Hong, Qiong Liu 0001, Dacheng Tao, Xuelong Li 0001
IEEE Trans. Geosci. Remote. Sens.2
2014 Hyperspectral Image Classification Through Bilayer Graph-Based Learning
abstract
Hyperspectral image classification with limited number of labeled pixels is a challenging task. In this paper, we propose a bilayer graph-based learning framework to address this problem. For graph-based classification, how to establish the neighboring relationship among the pixels from the high dimensional features is the key toward a successful classification. Our graph learning algorithm contains two layers. The first-layer constructs a simple graph, where each vertex denotes one pixel and the edge weight encodes the similarity between two pixels. Unsupervised learning is then conducted to estimate the grouping relations among different pixels. These relations are subsequently fed into the second layer to form a hypergraph structure, on top of which, semisupervised transductive learning is conducted to obtain the final classification results. Our experiments on three data sets demonstrate the merits of our proposed approach, which compares favorably with state of the art.
Yue Gao 0002, Rongrong Ji, Peng Cui 0001, Qionghai Dai, Gang Hua 0001
IEEE Trans. Image Process.1
2014 Weakly Supervised Visual Dictionary Learning by Harnessing Image Attributes
abstract
Bag-of-features (BoFs) representation has been extensively applied to deal with various computer vision applications. To extract discriminative and descriptive BoF, one important step is to learn a good dictionary to minimize the quantization loss between local features and codewords. While most existing visual dictionary learning approaches are engaged with unsupervised feature quantization, the latest trend has turned to supervised learning by harnessing the semantic labels of images or regions. However, such labels are typically too expensive to acquire, which restricts the scalability of supervised dictionary learning approaches. In this paper, we propose to leverage image attributes to weakly supervise the dictionary learning procedure without requiring any actual labels. As a key contribution, our approach establishes a generative hidden Markov random field (HMRF), which models the quantized codewords as the observed states and the image attributes as the hidden states, respectively. Dictionary learning is then performed by supervised grouping the observed states, where the supervised information is stemmed from the hidden states of the HMRF. In such a way, the proposed dictionary learning approach incorporates the image attributes to learn a semantic-preserving BoF representation without any genuine supervision. Experiments in large-scale image retrieval and classification tasks corroborate that our approach significantly outperforms the state-of-the-art unsupervised dictionary learning approaches.
Yue Gao 0002, Rongrong Ji, Wei Liu 0005, Qionghai Dai, Gang Hua 0001
IEEE Trans. Image Process.1
2014 Learning-Based Bipartite Graph Matching for View-Based 3D Model Retrieval
abstract
Distance measure between two sets of views is one central task in view-based 3D model retrieval. In this paper, we introduce a distance metric learning method for bipartite graph matching-based 3D object retrieval framework. In this method, the relationship among 3D models is formulated by a graph structure with semisupervised learning to estimate the model relevance. More specially, we model two sets of views by using a bipartite graph, on which their optimal matching is estimated. Then, we learn a refined distance metric by using the user’s relevance feedback. The proposed method has been evaluated on four data sets and the experimental results and comparison with the state-of-the-art methods demonstrate the effectiveness of the proposed method.
Ke Lu 0002, Rongrong Ji, Jinhui Tang 0001, Yue Gao 0002
IEEE Trans. Image Process.4
2014 Actively Learning Human Gaze Shifting Paths for Semantics-Aware Photo Cropping
abstract
Photo cropping is a widely used tool in printing industry, photography, and cinematography. Conventional cropping models suffer from the following three challenges. First, the deemphasized role of semantic contents that are many times more important than low-level features in photo aesthetics. Second, the absence of a sequential ordering in the existing models. In contrast, humans look at semantically important regions sequentially when viewing a photo. Third, the difficulty of leveraging inputs from multiple users. Experience from multiple users is particularly critical in cropping as photo assessment is quite a subjective task. To address these challenges, this paper proposes semantics-aware photo cropping, which crops a photo by simulating the process of humans sequentially perceiving semantically important regions of a photo. We first project the local features (graphlets in this paper) onto the semantic space, which is constructed based on the category information of the training photos. An efficient learning algorithm is then derived to sequentially select semantically representative graphlets of a photo, and the selecting process can be interpreted by a path, which simulates humans actively perceiving semantics in a photo. Furthermore, we learn a prior distribution of such active graphlet paths from training photos that are marked as aesthetically pleasing by multiple users. The learned priors enforce the corresponding active graphlet path of a test photo to be maximally similar to those from the training photos. Experimental results show that: 1) the active graphlet path accurately predicts human gaze shifting, and thus is more indicative for photo aesthetics than conventional saliency maps and 2) the cropped photos produced by our approach outperform its competitors in both qualitative and quantitative comparisons.
Yue Gao 0002, Rongrong Ji, Yingjie Xia, Qionghai Dai, Xuelong Li 0001
IEEE Trans. Image Process.2
2014 Fusion of Multichannel Local and Global Structural Cues for Photo Aesthetics Evaluation
abstract
Photo aesthetic quality evaluation is a fundamental yet under addressed task in computer vision and image processing fields. Conventional approaches are frustrated by the following two drawbacks. First, both the local and global spatial arrangements of image regions play an important role in photo aesthetics. However, existing rules, e.g., visual balance, heuristically define which spatial distribution among the salient regions of a photo is aesthetically pleasing. Second, it is difficult to adjust visual cues from multiple channels automatically in photo aesthetics assessment. To solve these problems, we propose a new photo aesthetics evaluation framework, focusing on learning the image descriptors that characterize local and global structural aesthetics from multiple visual channels. In particular, to describe the spatial structure of the image local regions, we construct graphlets small-sized connected graphs by connecting spatially adjacent atomic regions. Since spatially adjacent graphlets distribute closely in their feature space, we project them onto a manifold and subsequently propose an embedding algorithm. The embedding algorithm encodes the photo global spatial layout into graphlets. Simultaneously, the importance of graphlets from multiple visual channels are dynamically adjusted. Finally, these post-embedding graphlets are integrated for photo aesthetics evaluation using a probabilistic model. Experimental results show that: 1) the visualized graphlets explicitly capture the aesthetically arranged atomic regions; 2) the proposed approach generalizes and improves four prominent aesthetic rules; and 3) our approach significantly outperforms state-of-the-art algorithms in photo aesthetics prediction.
Yue Gao 0002, Roger Zimmermann, Qi Tian 0001, Xuelong Li 0001
IEEE Trans. Image Process.2
2014 A Probabilistic Associative Model for Segmenting Weakly Supervised Images
abstract
Weakly-supervised image segmentation is an important yet challenging task in image processing and pattern recognition fields. It is defined as: in the training stage, semantic labels are only at the image-level, without regard to their specific object/scene location within the image. Given a test image, the goal is to predict the semantics of every pixel/superpixel. In this paper, we propose a new weakly-supervised image segmentation model, focusing on learning the semantic associations between superpixel sets (graphlets in this work). In particular, we first extract graphlets from each image, where a graphlet is a small-sized graph measures the potential of multiple spatially neighboring superpixels (i.e., the probability of these superpixels sharing a common semantic label, such as the "sky" or the "sea"). To compare dierent-sized graphlets and to incorporate image-level labels, a manifold embedding algorithm is designed to transform all graphlets into equal-length feature vectors. Finally, we present a hierarchical Bayesian network (BN) to capture the semantic associations between post-embedding graphlets, based on which the semantics of each superpixel is inferred accordingly. Experimental results demonstrate that: 1) our approach performs competitively compared with the state-of-the-art approaches on three public data sets, and 2) considerable performance enhancement is achieved when using our approach on segmentation-based photo cropping and image categorization.
Yi Yang 0001, Yue Gao 0002, Yi Yu 0001, Changbo Wang, Xuelong Li 0001
IEEE Trans. Image Process.3
2014 Exploiting Web Images for Semantic Video Indexing Via Robust Sample-Specific Loss
abstract
Semantic video indexing, also known as video annotation or video concept detection in literatures, has been attracting significant attention in recent years. Due to deficiency of labeled training videos, most of the existing approaches can hardly achieve satisfactory performance. In this paper, we propose a novel semantic video indexing approach, which exploits the abundant user-tagged Web images to help learn robust semantic video indexing classifiers. The following two major challenges are well studied: 1) noisy Web images with imprecise and/or incomplete tags; and 2) domain difference between images and videos. Specifically, we first apply a non-parametric approach to estimate the probabilities of images being correctly tagged as confidence scores. We then develop a robust transfer video indexing (RTVI) model to learn reliable classifiers from a limited number of training videos together with the abundance of user-tagged images. The RTVI model is equipped with a novel sample-specific robust loss function, which employs the confidence score of a Web image as prior knowledge to suppress the influence and control the contribution of this image in the learning process. Meanwhile, the RTVI model discovers an optimal kernel space, in which the mismatch between images and videos is minimized for tackling the domain difference problem. Besides, we devise an iterative algorithm to effectively optimize the proposed RTVI model and a theoretical analysis on the convergence of the proposed algorithm is provided as well. Extensive experiments on various real-world multimedia collections demonstrate the effectiveness of the proposed robust semantic video indexing approach.
Yang Yang 0002, Zhengjun Zha, Yue Gao 0002, Xiaofeng Zhu 0001, Tat-Seng Chua
IEEE Trans. Multim.3
2014 Representative Discovery of Structure Cues for Weakly-Supervised Image Segmentation
abstract
Weakly-supervised image segmentation is a challenging problem with multidisciplinary applications in multimedia content analysis and beyond. It aims to segment an image by leveraging its image-level semantics (i.e., tags). This paper presents a weakly-supervised image segmentation algorithm that learns the distribution of spatially structural superpixel sets from image-level labels. More specifically, we first extract graphlets from a given image, which are small-sized graphs consisting of superpixels and encapsulating their spatial structure. Then, an efficient manifold embedding algorithm is proposed to transfer labels from training images into graphlets. It is further observed that there are numerous redundant graphlets that are not discriminative to semantic categories, which are abandoned by a graphlet selection scheme as they make no contribution to the subsequent segmentation. Thereafter, we use a Gaussian mixture model (GMM) to learn the distribution of the selected post-embedding graphlets (i.e., vectors output from the graphlet embedding). Finally, we propose an image segmentation algorithm, termed representative graphlet cut, which leverages the learned GMM prior to measure the structure homogeneity of a test image. Experimental results show that the proposed approach outperforms state-of-the-art weakly-supervised image segmentation methods, on five popular segmentation data sets. Besides, our approach performs competitively to the fully-supervised segmentation models.
Yue Gao 0002, Yingjie Xia, Ke Lu 0002, Jialie Shen 0001, Rongrong Ji
IEEE Trans. Multim.2
2014 Attribute-Augmented Semantic Hierarchy: Towards a Unified Framework for Content-Based Image Retrieval
abstract
This article presents a novel attribute-augmented semantic hierarchy (A 2 SH) and demonstrates its effectiveness in bridging both the semantic and intention gaps in content-based image retrieval (CBIR). A 2 SH organizes semantic concepts into multiple semantic levels and augments each concept with a set of related attributes. The attributes are used to describe the multiple facets of the concept and act as the intermediate bridge connecting the concept and low-level visual content. An hierarchical semantic similarity function is learned to characterize the semantic similarities among images for retrieval. To better capture user search intent, a hybrid feedback mechanism is developed, which collects hybrid feedback on attributes and images. This feedback is then used to refine the search results based on A 2 SH. We use A 2 SH as a basis to develop a unified content-based image retrieval system. We conduct extensive experiments on a large-scale dataset of over one million Web images. Experimental results show that the proposed A 2 SH can characterize the semantic affinities among images accurately and can shape user search intent quickly, leading to more accurate search results as compared to state-of-the-art CBIR solutions.
Hanwang Zhang, Zhengjun Zha, Yang Yang 0002, Shuicheng Yan, Yue Gao 0002, Tat-Seng Chua
ACM Trans. Multim. Comput. Commun. Appl.5
2013 Stereotime: a wireless 2D and 3D switchable video communication system
abstract
Mobile 3D video communication, especially with 2D and 3D compatible, is a new paradigm for both video communication and 3D video processing. Current techniques face challenges in mobile devices when bundled constraints such as computation resource and compatibility should be considered. In this work, we present a wireless 2D and 3D switchable video communication to handle the previous challenges, and name it as Stereotime. The methods of Zig-Zag fast object segmentation, depth cues detection and merging, and texture-adaptive view generation are used for 3D scene reconstruction. We show the functionalities and compatibilities on 3D mobile devices in WiFi network environment.
You Yang 0002, Qiong Liu 0001, Yue Gao 0002, Binbin Xiong, Li Yu 0003, Huan-Bo Luan, Rongrong Ji, Qi Tian 0001
ACM Multimedia3
2013 Attribute-augmented semantic hierarchy: towards bridging semantic gap and intention gap in image retrieval
abstract
This paper presents a novel Attribute-augmented Semantic Hierarchy (A2 SH) and demonstrates its effectiveness in bridging both the semantic and intention gaps in Content-based Image Retrieval (CBIR). A2 SH organizes the semantic concepts into multiple semantic levels and augments each concept with a set of related attributes, which describe the multiple facets of the concept and act as the intermediate bridge connecting the concept and low-level visual content. A hierarchical semantic similarity function is learnt to characterize the semantic similarities among images for retrieval. To better capture user search intent, a hybrid feedback mechanism is developed, which collects hybrid feedbacks on attributes and images. These feedbacks are then used to refine the search results based on A2 SH. We develop a content-based image retrieval system based on the proposed A2 SH. We conduct extensive experiments on a large-scale data set of over one million Web images. Experimental results show that the proposed A2 SH can characterize the semantic affinities among images accurately and can shape user search intent precisely and quickly, leading to more accurate search results as compared to state-of-the-art CBIR solutions.
Hanwang Zhang, Zhengjun Zha, Yang Yang 0002, Shuicheng Yan, Yue Gao 0002, Tat-Seng Chua
ACM Multimedia5
2013 Geographical Retagging
Liujuan Cao, Yue Gao 0002, Qiong Liu 0001, Rongrong Ji
MMM (2)2
2013 Hyperspectral Image Classification by Using Pixel Spatial Correlation
Yue Gao 0002, Tat-Seng Chua
MMM (1)1
2013 Multi-camera Egocentric Activity Detection for Personal Assistant
Yue Gao 0002, Alex Hauptmann 0001
MMM (2)2
2013 Mining spatiotemporal video patterns towards robust action retrieval
Liujuan Cao, Rongrong Ji, Yue Gao 0002, Wei Liu 0005, Qi Tian 0001
Neurocomputing3
2013 Texture-adaptive hole-filling algorithm in raster-order for three-dimensional video applications
Qiong Liu 0001, You Yang 0002, Yue Gao 0002, Richang Hong
Neurocomputing3
2013 A Bayesian framework for dense depth estimation based on spatial-temporal correlation
Qiong Liu 0001, You Yang 0002, Yue Gao 0002, Rongrong Ji, Li Yu 0003
Neurocomputing3
2013 Multimedia encyclopedia construction by mining web knowledge
Richang Hong, Zhengjun Zha, Yue Gao 0002, Tat-Seng Chua, Xindong Wu 0001
Signal Process.3
2013 Visual-Textual Joint Relevance Learning for Tag-Based Social Image Search
abstract
Due to the popularity of social media websites, extensive research efforts have been dedicated to tag-based social image search. Both visual information and tags have been investigated in the research field. However, most existing methods use tags and visual characteristics either separately or sequentially in order to estimate the relevance of images. In this paper, we propose an approach that simultaneously utilizes both visual and textual information to estimate the relevance of user tagged images. The relevance estimation is determined with a hypergraph learning approach. In this method, a social image hypergraph is constructed, where vertices represent images and hyperedges represent visual or textual terms. Learning is achieved with use of a set of pseudo-positive images, where the weights of hyperedges are updated throughout the learning process. In this way, the impact of different tags and visual words can be automatically modulated. Comparative results of the experiments conducted on a dataset including 370+images are presented, which demonstrate the effectiveness of the proposed approach.
Yue Gao 0002, Meng Wang 0001, Zhengjun Zha, Jialie Shen 0001, Xuelong Li 0001, Xindong Wu 0001
IEEE Trans. Image Process.1
2013 View-Based Discriminative Probabilistic Modeling for 3D Object Retrieval and Recognition
abstract
In view-based 3D object retrieval and recognition, each object is described by multiple views. A central problem is how to estimate the distance between two objects. Most conventional methods integrate the distances of view pairs across two objects as an estimation of their distance. In this paper, we propose a discriminative probabilistic object modeling approach. It builds probabilistic models for each object based on the distribution of its views, and the distance between two objects is defined as the upper bound of the Kullback-Leibler divergence of the corresponding probabilistic models. 3D object retrieval and recognition is accomplished based on the distance measures. We first learn models for each object by the adaptation from a set of global models with a maximum likelihood principle. A further adaption step is then performed to enhance the discriminative ability of the models. We conduct experiments on the ETH 3D object dataset, the National Taiwan University 3D model dataset, and the Princeton Shape Benchmark. We compare our approach with different methods, and experimental results demonstrate the superiority of our approach.
Meng Wang 0001, Yue Gao 0002, Ke Lu 0002, Yong Rui
IEEE Trans. Image Process.2
2013 Beyond Text QA: Multimedia Answer Generation by Harvesting Web Information
abstract
Community question answering (cQA) services have gained popularity over the past years. It not only allows community members to post and answer questions but also enables general users to seek information from a comprehensive set of well-answered questions. However, existing cQA forums usually provide only textual answers, which are not informative enough for many questions. In this paper, we propose a scheme that is able to enrich textual answers in cQA with appropriate media data. Our scheme consists of three components: answer medium selection, query generation for multimedia search, and multimedia data selection and presentation. This approach automatically determines which type of media information should be added for a textual answer. It then automatically collects data from the web to enrich the answer. By processing a large set of QA pairs and adding them to a pool, our approach can enable a novel multimedia question answering (MMQA) approach as users can find multimedia answers by matching their questions with those in the pool. Different from a lot of MMQA research efforts that attempt to directly answer questions with image and video data, our approach is built based on community-contributed textual answers and thus it is able to deal with more complex questions. We have conducted extensive experiments on a multi-source QA dataset. The results demonstrate the effectiveness of our approach.
Liqiang Nie, Meng Wang 0001, Yue Gao 0002, Zhengjun Zha, Tat-Seng Chua
IEEE Trans. Multim.3
2013 When Amazon Meets Google: Product Visualization by Exploring Multiple Web Sources
Meng Wang 0001, Guangda Li, Zheng Lu 0002, Yue Gao 0002, Tat-Seng Chua
ACM Trans. Internet Techn.4
2012 Weakly supervised sparse coding with geometric consistency pooling
abstract
Most recently the Bag-of-Features (BoF) representation has been well advocated for image search and classification, with two decent phases named sparse coding and max pooling to compensate quantization loss as well as inject spatial layouts. But still, much information has been discarded by quantizing local descriptors with two-dimensional layouts into a one-dimensional BoF histogram. In this paper, we revisit this popular “sparse coding + max pooling” paradigm by “looking around” the local descriptor context towards an optimal BoF. First, we introduce a Weakly supervised Sparse Coding (WSC) to exploit the Classemes-based attribute labeling to refine the descriptor coding procedure. It is achieved by learning an attribute-to-word co-occurrence prior to impose a label inconsistency distortion over the ℓ1based coding regularizer, such that the descriptor codes can maximally preserve the image semantic similarity. Second, we propose an adaptive feature pooling scheme over “superpixels” rather than over fixed spatial pyramids, named Geometric Consistency Pooling (GCP). As an effect, local descriptors enjoying good geometric consistency are pooled together to ensure a more precise spatial layouts embedding in BoF. Both of our phases are unsupervised, which differ from the existing works in supervised dictionary learning, sparse coding and feature pooling. Therefore, our approach enables potential applications like scalable visual search. We evaluate in both image classification and search benchmarks and report good improvements over the state-of-the-arts.
Liujuan Cao, Rongrong Ji, Yue Gao 0002, Yi Yang 0001, Qi Tian 0001
CVPR3
2012 Weakly supervised topic grouping of YouTube search results
abstract
Recent years have witnessed an explosive growth of user contributed videos on websites like YouTube and Metacafe, which usually provide a query-by-keyword functionality to facilitate the user browsing. For a given query, the returned videos typically contain multiple topics that are mixed up to duplicate the user browsing. Therefore, their diversification and grouping are highly demanded to improve the user experiences. However, the tagging and content qualities of user contributed videos are uncontrolled against their precise grouping. In this paper, we present a weakly supervised topic grouping paradigm to diversify the returned videos of a given keyword query. Our grouping is based on the bag-of-words visual signature quantized over the spatiotemporal STIP descriptor [1] extracted from each returned video. First, we adopt a min-Hashing based visual similarity in combination of the tagging similarity to group the returned videos. Based on the initial grouping configurations, we mine the co-occurred discriminative sub-signatures, based on which we iteratively refine the first step. Such iteration well handles the noise in visual content and tagging, since neither of which is fully trusted during the grouping. We validate our schemes on over 2,000 video clips crawled from a set of YouTube keyword query results. Comparing to alternative approaches, our scheme has shown superior robustness and precision.
Liujuan Cao, Rongrong Ji, Wei Liu 0005, Yue Gao 0002, Ling-Yu Duan, Chaoguang Men
ICIP4
2012 View-based 3D object retrieval by bipartite graph matching
abstract
Bipartite graph matching has been investigated in multiple view matching for 3D object retrieval. However, existing methods employ one-to-one vertex matching scheme while more than two views may share close semantic meanings in practice. In this work, we propose a bipartite graph matching method to measure the distance between two objects based on multiple views. In the proposed method, representative views are first selected by using view clustering for each object, and the corresponding weights are given based on the cluster results. A bipartite graph is constructed by using the two groups of representative views from two compared objects. To calculate the similarity between two objects, the bipartite graph is first partitioned to several subsets, and the views in the same sub-set are with high possibility to be with similar semantic meanings. The distances between two objects within individual subsets are then assembled through the graph to obtain the final similarity. Experimental results and comparison with the state-of-the-art methods demonstrate the effectiveness of the proposed algorithm.
Yue Wen, Yue Gao 0002, Richang Hong, Huan-Bo Luan, Qiong Liu 0001, Jialie Shen 0001, Rongrong Ji
ACM Multimedia2
2012 Attribute feedback
abstract
This demonstration presents a new interactive Content Based Image Retrieval (CBIR) system, termed Attribute Feedback (AF). Unlike traditional relevance feedback purely founded on low-level features, AF system shapes user's search intents more precisely and quickly by collecting feedbacks on intermediate-level semantic attribute. At each interaction iteration, the AF system first determines the most informative binary attributes for feedbacks and then augments the binary attribute feedbacks by a new type of attributes, "affinity attributes", each of which is learnt offline to describe the distance/similarity between user's envisioned image(s) and a retrieved image with respect to the corresponding affinity attribute. Based on the feedbacks on binary and affinity attributes, the images in corpus are further re-ranked towards better fitting user's search intents. The experimental results on two real-world image datasets have demonstrated the superiority of the AF system over other state-of-the-art relevance feedback based CBIR approaches.
Hanwang Zhang, Zhengjun Zha, Jingwen Bian, Yue Gao 0002, Huan-Bo Luan, Tat-Seng Chua
ACM Multimedia4
2012 Symbiotic Black-Box Tracker
Yue Gao 0002, Alex Hauptmann 0001, Rongrong Ji, Boaz J. Super
MMM2
2012 k-Partite graph reinforcement and its application in multimedia information retrieval
Yue Gao 0002, Meng Wang 0001, Rongrong Ji, Zhengjun Zha, Jialie Shen 0001
Inf. Sci.1
2012 Cross-View Down/Up-Sampling Method for Multiview Depth Video Coding
abstract
In this letter, we propose a cross-view down/up-sampling (CDU) method for the framework of reduced resolution multiview depth video coding, which exploits cross-view information to assist the up-sampling at the decoder. In the down-sampling procedure of CDU, the odd-even interlaced extraction is employed to preserve more confident information of the original depth video with reduced resolution. In the decoder, the cross-view information is exploited for up-sampling the reconstructed depth video. An iterative interpolation process is proposed to eliminate the effect of compression distortion on this up-sampling. Experimental results demonstrate the gains of up to 3.88 dB for the proposed algorithm and better quality of synthesized views.
Qiong Liu 0001, You Yang 0002, Rongrong Ji, Yue Gao 0002, Li Yu 0003
IEEE Signal Process. Lett.4
2012 Camera Constraint-Free View-Based 3-D Object Retrieval
abstract
Recently, extensive research efforts have been dedicated to view-based methods for 3-D object retrieval due to the highly discriminative property of multiviews for 3-D object representation. However, most of state-of-the-art approaches highly depend on their own camera array settings for capturing views of 3-D objects. In order to move toward a general framework for 3-D object retrieval without the limitation of camera array restriction, a camera constraint-free view-based (CCFV) 3-D object retrieval algorithm is proposed in this paper. In this framework, each object is represented by a free set of views, which means that these views can be captured from any direction without camera constraint. For each query object, we first cluster all query views to generate the view clusters, which are then used to build the query models. For a more accurate 3-D object comparison, a positive matching model and a negative matching model are individually trained using positive and negative matched samples, respectively. The CCFV model is generated on the basis of the query Gaussian models by combining the positive matching model and the negative matching model. The CCFV removes the constraint of static camera array settings for view capturing and can be applied to any view-based 3-D object database. We conduct experiments on the National Taiwan University 3-D model database and the ETH 3-D object database. Experimental results show that the proposed scheme can achieve better performance than state-of-the-art methods.
Yue Gao 0002, Jinhui Tang 0001, Richang Hong, Shuicheng Yan, Qionghai Dai, Naiyao Zhang, Tat-Seng Chua
IEEE Trans. Image Process.1
2012 3-D Object Retrieval and Recognition With Hypergraph Analysis
abstract
View-based 3-D object retrieval and recognition has become popular in practice, e.g., in computer aided design. It is difficult to precisely estimate the distance between two objects represented by multiple views. Thus, current view-based 3-D object retrieval and recognition methods may not perform well. In this paper, we propose a hypergraph analysis approach to address this problem by avoiding the estimation of the distance between objects. In particular, we construct multiple hypergraphs for a set of 3-D objects based on their 2-D views. In these hypergraphs, each vertex is an object, and each edge is a cluster of views. Therefore, an edge connects multiple vertices. We define the weight of each edge based on the similarities between any two views within the cluster. Retrieval and recognition are performed based on the hypergraphs. Therefore, our method can explore the higher order relationship among objects and does not use the distance between objects. We conduct experiments on the National Taiwan University 3-D model dataset and the ETH 3-D object collection. Experimental results demonstrate the effectiveness of the proposed method by comparing with the state-of-the-art methods.
Yue Gao 0002, Meng Wang 0001, Dacheng Tao, Rongrong Ji, Qionghai Dai
IEEE Trans. Image Process.1
2011 Tag-based social image search with visual-text joint hypergraph learning
abstract
Tag-based social image search has attracted great interest and how to order the search results based on relevance level is a research problem. Visual content of images and tags have both been investigated. However, existing methods usually employ tags and visual content separately or sequentially to learn the image relevance. This paper proposes a tag-based image search with visual-text joint hypergraph learning. We simultaneously investigate the bag-of-words and bag-of-visual-words representations of images and accomplish the relevance estimation with a hypergraph learning approach. Each textual or visual word generates a hyperedge in the constructed hypergraph. We conduct experiments with a real-world data set and experimental results demonstrate the effectiveness of our approach.
Yue Gao 0002, Meng Wang 0001, Huan-Bo Luan, Jialie Shen 0001, Shuicheng Yan, Dacheng Tao
ACM Multimedia1
2011 3D model retrieval using weighted bipartite graph matching
Yue Gao 0002, Qionghai Dai, Meng Wang 0001, Naiyao Zhang
Signal Process. Image Commun.1
2011 Less is More: Efficient 3-D Object Retrieval With Query View Selection
abstract
The explosively increasing 3-D objects make their efficient retrieval technology highly desired. Extensive research efforts have been dedicated to view-based 3-D object retrieval for its advantage of using 2-D views to represent 3-D objects. In this paradigm, typically the retrieval is accomplished by matching the views of the query object with the objects in database. However, using all the query views may not only introduce difficulty in rapid retrieval but also degrade retrieval accuracy when there is a mismatch between the query views and the object views in the database. In this work, we propose an interactive 3-D object retrieval scheme. Given a set of query views, we first perform clustering to obtain several candidates. We then incrementally select query views for object matching: in each round of relevance feedback, we only add the query view that is judged to be the most informative one based on the labeling information. In addition, we also propose an efficient approach to learn a distance metric for the newly selected query view and the weights for combining all of the selected query views. We conduct experiments on the National Taiwan University 3D Model database, ETH 3D object collection, and Shape Retrieval Content of Non-Rigid 3D Model, and results demonstrated that our approach not only significantly speeds up the retrieval process but also achieves encouraging retrieval performance.
Yue Gao 0002, Meng Wang 0001, Zhengjun Zha, Qi Tian 0001, Qionghai Dai, Naiyao Zhang
IEEE Trans. Multim.1
2011 Mining flickr landmarks by modeling reconstruction sparsity
abstract
In recent years, there have been ever-growing geographical tagged photos on the community Web sites such as Flickr. Discovering touristic landmarks from these photos can help us to make better sense of our visual world. In this article, we report our work on mining landmarks from geotagged Flickr photos for city scene summarization and touristic recommendations. We begin by exploring the geographical and visual statistics of the Web users' photographing manner, based on which we conduct landmark mining in two steps: First, we propose to partition each city into geographical regions based on spectral clustering over the geotags of Flickr photos. Second, in each landmark region, we present a representative photo mining scheme based on sparse representation. Our main idea is to regard the landmark mining problem as a process to find photos whose visual signatures can be reconstructed using other photos of this landmark region with a minimal coding length. This sparse reconstruction scheme offers a general perspective to mine the representative photos. Indeed, by simplifying the data correlation constraints in our scheme, several previous works in representative photo discovery and landmark mining can be derived. Finally, we introduce a Hyperlink-Induced Topic Search model to refine our landmark ranking, which incorporates the community knowledge to simulate the landmark ranking problem as a dynamic page ranking problem. We have deployed our proposed landmark mining framework on a city scene summarization and navigation system, which works on one million geotagged Flickr photos coming from twenty worldwide metropolises. We have also quantitatively compared our scheme with several state-of-the-art works.
Rongrong Ji, Yue Gao 0002, Bineng Zhong 0001, Hongxun Yao, Qi Tian 0001
ACM Trans. Multim. Comput. Commun. Appl.2
2010 W2Go: a travel guidance system by automatic landmark ranking
abstract
In this paper, we present a travel guidance system W2Go (Where to Go), which can automatically recognize and rank the landmarks for travellers. In this system, a novel Automatic Landmark Ranking (ALR) method is proposed by utilizing the tag and geo-tag information of photos in Flickr and user knowledge from Yahoo Travel Guide. ALR selects the popular tourist attractions (landmarks) based on not only the subjective opinion of the travel editors as is currently done on sites like WikiTravel and Yahoo Travel Guide, but also the ranking derived from popularity among tourists. Our approach utilizes geo-tag information to locate the positions of the tag-indicated places, and computes the probability of a tag being a landmark/site name. For potential landmarks, impact factors are calculated from the frequency of tags, user numbers in Flickr, and user knowledge in Yahoo Travel Guide. These tags are then ranked based on the impact factors. Several representative views for popular landmarks are generated from the crawled images with geo-tags to describe and present them in context of information derived from several relevant reference sources. The experimental comparisons to the other systems are conducted on eight famous cities over the world. User-based evaluation demonstrates the effectiveness of the proposed ALR method and the W2Go system.
Yue Gao 0002, Jinhui Tang 0001, Richang Hong, Qionghai Dai, Tat-Seng Chua, Ramesh Jain 0001
ACM Multimedia1
2010 Intelligent query: open another door to 3d object retrieval
abstract
The increasing number of available 3D objects makes their efficient retrieval technology highly desired. Extensive research has been dedicated to view-based 3D object retrieval because of its advantage of 2D views for 3D object content representation. In this paradigm, typically the retrieval is accomplished based a set of different views of the query object, and focuses on the 3D object representation, matching and indexing. In this work, we present another aspect towards 3D object retrieval: intelligent query. Intelligent query includes query selection, query description and combination, and assistive query. We will show how this scheme is ideally suit for the 3D object retrieval problem. We conduct experiments on the National Taiwan University 3D Model database and results demonstrated that our approach can improve retrieval performance. Finally, we give insight into the future of the intelligent query for 3D object retrieval.
Yue Gao 0002, Meng Wang 0001, Jialie Shen 0001, Qionghai Dai, Naiyao Zhang
ACM Multimedia1
2010 Representative views re-ranking for 3D model retrieval with multi-bipartite graph reinforcement model
abstract
In this paper, we propose a multi-bipartite graph reinforcement model for representative views re-ranking in 3D model retrieval. Given the views of one query 3D model, all query views are grouped into clusters to generate representative views and corresponding original weights. In the retrieval procedure, labeled positive retrieval results are employed to refine the query information. Each group of views from positive retrieval results and the group of representative query views are employed to construct a bipartite graph, and a multi-bipartite graph reinforcement algorithm is performed on these bipartite graphs to re-rank all views. Then the weights of all representative query views are updated. Experimental results on two 3D model databases are provided to justify the effectiveness of the proposed method.
Yue Gao 0002, You Yang 0002, Qionghai Dai, Naiyao Zhang
ACM Multimedia1
2010 3D object retrieval with bag-of-region-words
abstract
View-based method becomes an essential approach to 3D object retrieval in recent years. In the view-based 3D object retrieval framework, each object is described by a set of views and representative features are extracted from these views to match the objects in database. In this paper, we propose a novel 3D multi-view representation method, Bag-of-Region-Words (BoRW). It first gridly selects points in each view and extracts local SIFT features. Each local feature is encoded into a visual word with a trained visual vocabulary. Then each view is split into several regions, and each region is represented by a bag-of-visual-words feature vector. All the obtained regions are further grouped into clusters based on the bag-of-visual-words feature, and one feature is selected from each cluster with corresponding weight. In this way, each object is described by a set of BoRW. The Earth Movers Distance is employed to estimate the distance between two BoRW feature vectors. Experimental results show that the proposed method can achieve better retrieval performance than existing methods.
Yue Gao 0002, You Yang 0002, Qionghai Dai, Naiyao Zhang
ACM Multimedia1
2010 View-based 3D model retrieval with probabilistic graph model
Yue Gao 0002, Jinhui Tang 0001, Qionghai Dai, Naiyao Zhang
Neurocomputing1
2010 3D model comparison using spatial structure circular descriptor
Yue Gao 0002, Qionghai Dai, Naiyao Zhang
Pattern Recognit.1
2009 Dynamic video summarization using two-level redundancy detection
Yue Gao 0002, Wei-Bo Wang, Jun-Hai Yong, He-Jin Gu
Multim. Tools Appl.1
2008 Shot-based similarity measure for content-based video summarization
abstract
The rapid development of multimedia applications over the past decade requires efficient methods for video browsing. In this paper, we present an algorithm for video summarization with shot comparison. We analyze video content in the shot level, and we calculate the shot distance using the advanced Hausdorff distance. The advanced Hausdorff distance combines the Hausdorff distance and Boolean model, and it could compare two shots from the global view. When the shot similarity matrix is obtained, we group these video shots into several clusters using the affinity propagation cluster method to remove redundant video content. Performance evaluation on ten video sequences are given to illustrate the proposed algorithm.
Yue Gao 0002, Qionghai Dai
ICIP1
2007 Video summarization by redundancy removing and content ranking
abstract
In order to help the user to grasp the long video content quickly, this paper proposes a novel video summarization approach based on redundancy removal and content ranking. By video parsing and cast indexing, the approach first constructs a story board to let user know about the main scenes and the main actors in the video. Then it generates a "story-constraint summary" by key frame clustering and repetitive segment detection. To shorten the video summary length to a target length, our approach constructs a "time-constraint summary" by important factor based content ranking. Extensive experiments are carried out on TV series, movies, and cartoons. Good results demonstrate the effectiveness of the proposed method.
Tao Wang 0003, Yue Gao 0002, Patricia Peng Wang, Eric Q. Li, Wei Hu 0002, Yimin Zhang 0002, Jun-Hai Yong
ACM Multimedia2