Yongli Hu

dblp:72/4503 · DBLP profile ↗
← Back
169ranked-venue papers
13as first author
128since 2021 · last 2026
0000-0003-0440-438XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 81 · 7 first-author · 56 since 2021Graphics, computer vision, multimedia, augmented reality and games · 57 · 5 first-author · 32 since 2021Applied, interdisciplinary, general and emerging computing · 32 · 1 first-author · 31 since 2021Databases, data management, data science and information retrieval · 17 · 1 first-author · 13 since 2021Systems, architecture and hardware · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MFC: Mixed Federated Clustering based on Cross-modal Feature Decoupling
abstract
Existing federated clustering methods typically assume that either all clients supply the same type of single-modal/multi-modal data, or that different clients provide various modalities describing the same object. However, a prevalent real-world scenario involves clients contributing data from diverse and unrelated modalities. Addressing the challenge of uncovering clustering patterns from such heterogeneous modality data distributed across distinct clients is crucial. In this paper, we propose a novel cross-modal feature decoupling-based mixed federated clustering model. To address client heterogeneity, we introduce a cross-modal feature decoupling module for each client, designed to decouple modality-agnostic and modality-specific features through distinct encoders.Only the modality-agnostic encoder parameters and clustering centers of each client are transmitted to the server. This enables the server to aggregate various modality-agnostic encoders, effectively discovering the global clustering structure while avoiding interference from modality-specific noise. Moreover, we develop a global consistent complementary clustering module to integrate the complementary clustering centers from various clients. The global clustering centers are then dispatched back to clients to guide and calibrate their local clustering models. Experimental results on three public datasets show the superiority of the proposed model compared to classic federated clustering methods.
Xiaxia He, Boyue Wang, Junbin Gao, Yongli Hu
KDD (1)4
2026 LocSAM: Modular SAM enhancement for dense object localization in complex scenes
Yong Zhang 0029, Bo Li 0128, Yongli Hu, Bob Zhang 0001
Comput. Vis. Image Underst.4
2026 Multi-agent role-playing by LLMs and LMMs: An explainable open-world multi-modal crisis tweet classification method
Tong Bie, Yongli Hu, Linjia Hao, Tengfei Liu 0005, Huajie Jiang, Junbin Gao
Expert Syst. Appl.2
2026 Multi-faceted contrastive learning with inter-frame difference for traffic video question answering
Kan Guo, Qi Zuo, Yongli Hu, Lanping Qian, Daxin Tian, Jiapu Wang, Guixian Qu, Tingzheng Jia, Junbin Gao
Knowl. Based Syst.3
2026 Adaptive low-quality negative samples partition and weakening for temporal knowledge graph completion
Boyue Wang, Yongli Hu
Multim. Syst.5
2026 Modality and semantic alignments based neighborhood node retrieval for multimodal knowledge graph completion
Dabao Zhang, Boyue Wang, Yongli Hu
Multim. Syst.4
2026 Frequency-domain multi-scale graph learning with information-theoretic constraint for spatio-temporal prediction
Shun Wang 0004, Yong Zhang 0029, Xuanqi Lin, Guangyu Huo, Xinglin Piao, Yongli Hu
Pattern Recognit.6
2026 M3Former: Memory-Guided Multi-Modal Generation and Adaptive Mixture Reasoning for Incomplete-Modality Crisis Event Detection
abstract
Multi-modal crisis event detection is critical for timely situational awareness across diverse real-world emergencies. In practice, however, multi-modal data—typically composed of images and textual reports—are often incomplete, with many instances providing only a single modality because of data loss, platform constraints, or real-time limitations. This modality-incomplete setting poses three key challenges: (1) how to reliably reconstruct missing modalities to restore cross-modal context, (2) how to extract deep semantic clues across original and completed data, and (3) how to bridge the distributional gap between reconstructed and fully-observed samples during model training. To tackle these issues, we propose M3Former, a unified framework tailored for modality-incomplete multi-modal crisis event detection. It consists of four dedicated modules: Memory-Guided Modality Completion builds a memory bank of paired image–text data to retrieve semantically related samples and keywords, guiding powerful pretrained generators—diffusion models for image synthesis and multi-modal large language models for text generation; Crisis-Aware Heterogeneous Dual-Stream Encoders jointly capture modality-specific cues and establish initial cross-modal semantic alignment between original and completed data; Hierarchical Attention Refinement Network progressively refines representations through hierarchical attention and guided cross-modal interaction to suppress noise and semantic drift; Modality-aware Routing Experts designs a gated mixture-of-experts architecture that dynamically selects both modality-specific and shared experts to mitigate distributional shifts and enhance the quality of fused representations. Extensive experiments on two real-world crisis datasets demonstrate that M3Former significantly outperforms existing baselines under a variety of modality-incomplete scenarios. The code is available: https://github.com/lcygky/M3former.
Chenyang Lu 0014, Boyue Wang, Tengfei Liu 0005, Yongli Hu
IEEE Trans. Circuits Syst. Video Technol.5
2026 Modality-Agnostic Hybrid Federated Learning via Knowledge Distillation and Reinforcement Learning-Based Aggregation
abstract
Due to multidimensional heterogeneity, Multimodal Federated Learning (MMFL) confronts fundamental challenges including modality incongruence, modality agnosticism, and modality incompleteness. Existing methods face a trilemma: leveraging external data with privacy risks, isolating features to restrict cross-modal interaction, or incurring high overhead from complex graph-based coordination, all culminating in suboptimal performance. In this paper, we propose a Modality-Agnostic Hybrid Federated Learning (MA-HyFL) framework that synergistically integrates unimodal and multimodal federated processes in modality-agnostic scenarios. Specifically, a bidirectional cross-modal knowledge distillation is employed to promote comprehensive collaboration at inter-client and intra-client levels, enabling robust knowledge transfer among heterogeneous modalities. A reinforcement learning-based aggregation mechanism is further introduced to orchestrate federated workflows through reward-driven policy optimization, dynamically integrating contributive client selection and adaptive aggregation weighting for closed-loop decision-making. Extensive experiments show that MA-HyFL significantly outperforms the other baseline methods in four realistic real-world applications, each exhibiting varying degrees of statistical heterogeneity and missing rates.
Yongli Hu, Huajie Jiang, Junbin Gao
IEEE Trans. Circuits Syst. Video Technol.2
2026 Multimodal Knowledge Graph Completion by Cross-Modal Interaction With Similarity Enhancing and Difference Embracing
abstract
Multimodal knowledge graph completion (MMKGC) enhances the precision and breadth of application of knowledge graphs by integrating rich data from various modalities, steadily increasing its appeal in the research community. Prior studies mainly focus on the common representation of different modalities while neglecting the different and complementary features. On the contrary, some works tend to model triples of each modality separately while overlooking the similarities between modalities. It is challenging to associate the heterogeneous modalities effectively for MMKGC. In this article, we introduce a novel MMKGC framework by cross-modal interaction with similarity-enhancing and difference-embracing (CISEDE), which leverages both the similarities and differences among multimodal entities by a proposed cross-modal interaction mechanism. In the cross-modal interaction, multihead attention is employed to enhance similarity information from multimodal entities and embrace different information by linking various modal triples. Through relation-guided fusion, the modal triples are decoded and merged for MMKGC. The experimental results on three commonly used datasets, FB15k-237, WN9, and WN18RR, show that the proposed method achieves state-of-the-art performance.
Linjia Hao, Yongli Hu, Tong Bie, Huajie Jiang, Junbin Gao
IEEE Trans. Neural Networks Learn. Syst.2
2025 HC-LLM: Historical-Constrained Large Language Models for Radiology Report Generation
abstract
Radiology report generation (RRG) models typically focus on individual exams, often overlooking the integration of historical visual or textual data, which is crucial for patient follow-ups. Traditional methods usually struggle with long sequence dependencies when incorporating historical information, but large language models (LLMs) excel at in-context learning, making them well-suited for analyzing longitudinal medical data. In light of this, we propose a novel Historical-Constrained Large Language Models (HC-LLM) framework for RRG, empowering LLMs with longitudinal report generation capabilities by constraining the consistency and differences between longitudinal images and their corresponding reports. Specifically, our approach extracts both time-shared and time-specific features from longitudinal chest X-rays and diagnostic reports to capture disease progression. Then, we ensure consistent representation by applying intra-modality similarity constraints and aligning various features across modalities with multimodal contrastive and structural constraints. These combined constraints effectively guide the LLMs in generating diagnostic reports that accurately reflect the progression of the disease, achieving state-of-the-art results on the Longitudinal-MIMIC dataset. Notably, our approach performs well even without historical data during testing and can be easily adapted to other multimodal large models, enhancing its versatility.
Tengfei Liu 0005, Jiapu Wang, Yongli Hu, Mingjie Li 0006, Junfei Yi, Xiaojun Chang, Junbin Gao
AAAI3
2025 Visual and Semantic Prompt Collaboration for Generalized Zero-Shot Learning
abstract
Generalized zero-shot learning aims to recognize both seen and unseen classes with the help of semantic information that is shared among different classes. It inevitably requires consistent visual-semantic alignment. Existing approaches fine-tune the visual backbone by seen-class data to obtain semantic-related visual features, which may cause overfitting on seen classes with a limited number of training images. This paper proposes a novel visual and semantic prompt collaboration framework, which utilizes prompt tuning techniques for efficient feature adaptation. Specifically, we design a visual prompt to integrate the visual information for discriminative feature learning and a semantic prompt to integrate the semantic formation for visual-semantic alignment. To achieve effective prompt information integration, we further design a weak prompt fusion mechanism for the shallow layers and a strong prompt fusion mechanism for the deep layers in the network. Through the collaboration of visual and semantic prompts, we can obtain discriminative semantic-related features for generalized zero-shot image recognition. Extensive experiments demonstrate that our framework consistently achieves favorable performance in both conventional zero-shot learning and generalized zero-shot learning benchmarks compared to other state-of-the-art methods.
Huajie Jiang, Zhengxian Li, Xiaohan Yu 0001, Yongli Hu, Jian Yang 0001, Yuankai Qi
CVPR4
2025 Spatial-Temporal Traffic Prediction Based on Multi-Scale Time Difference
Yongli Hu, Qi Zuo, Kan Guo, Zhongfan Sun, Tingzheng Jia
ICIC (11)1
2025 Multi-View Hallucinatory Feature Generation Network for Isolated Sign Language Recognition
abstract
It is widely acknowledged in the field of sign language recognition that the incorporation of supplementary measurements significantly enhances recognition accuracy. Specifically, the utilization of multi-view data effectively mitigates recognition errors arising from variations in observation angles, gesture occlusion, and other related factors. However, in practical scenarios, the availability of multi-view data is often limited. To address this challenge, we propose a novel approach based on dataset distillation, termed the Multi-view Hallucinatory Feature Generation Network (MHFGN). This model is designed to train generators by inputting sequence features derived from multiple distinct views, enabling the generation of hallucinatory features based on a single view. During the inference phase, only single-view data is employed, and the trained generators collaboratively produce multi-view hallucinatory features, thereby achieving the objective of multi-view recognition. Extensive experimental evaluations across multiple datasets demonstrate that the proposed model exhibits simplicity, efficacy, and robustness, particularly in scenarios subject to view changes.
Yongli Hu
IJCNN2
2025 Large-Small Model Synergy with Multimodal Fine-Grained Heuristics for Knowledge-Based Visual Question Answering
abstract
Multimodal Large Language Models (MLLMs) possess extensive knowledge and strong reasoning capabilities, achieving remarkable performance in knowledge-based visual question answering, significantly surpassing traditional small-scale Vision-Language Models (VLMs). However, the distinct training paradigms of MLLMs and small-scale VLMs result in misaligned feature representation spaces and divergent answer prediction distributions. To bridge this gap, we propose a novel end-to-end large-small model synergy framework, where small VLMs and MLLMs collaborate via synergistic optimization of shared objectives while maintaining their co-evolving complementary specializations. Specifically, multimodal fine-grained heuristics are extracted from well-tuned small VLMs and subsequently projected into the textual space of MLLMs through dedicated visual and textual collaboration modules. This enables cross-modal guidance for both visual and textual inputs. Finally, a dual-objective synergy loss promotes alignment toward shared goals, while a visual discrepancy loss preserves specialization diversity. Extensive experiments demonstrate that our framework achieves state-of-the-art performance on both the OK-VQA and A-OKVQA benchmarks.
Zhongfan Sun, Kan Guo, Yongli Hu, Daxin Tian, Qingqing Gao, Jiapu Wang, Junbin Gao
ACM Multimedia3
2025 DiMA: sequence diversity dynamics analyser for viruses
abstract
Sequence diversity is one of the major challenges in the design of diagnostic, prophylactic, and therapeutic interventions against viruses. DiMA is a novel tool that is big data-ready and designed to facilitate the dissection of sequence diversity dynamics for viruses. DiMA stands out from other diversity analysis tools by offering various unique features. DiMA provides a quantitative overview of sequence (DNA/RNA/protein) diversity by use of Shannon's entropy corrected for size bias, applied via a user-defined k-mer sliding window to an input alignment file, and each k-mer position is dissected to various diversity motifs. The motifs are defined based on the probability of distinct sequences at a given k-mer alignment position, whereby an index is the predominant sequence, while all the others are (total) variants to the index. The total variants are sub-classified into the major (most common) variant, minor variants (occurring more than once and of incidence lower than the major), and the unique (singleton) variants. DiMA allows user-defined, sequence metadata enrichment for analyses of the motifs. The application of DiMA was demonstrated for the alignment data of the relatively conserved Spike protein (2,106,985 sequences) of the severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) and the relatively highly diverse pol gene (2637) of the human immunodeficiency virus-1 (HIV-1). The tool is publicly available as a web server (https://dima.bezmialem.edu.tr), as a Python library (via PyPi) and as a command line client (via GitHub).
Shan Tharanga, Eyyüb Selim Ünlü, Yongli Hu, Muhammad Farhan Sjaugi, Muhammet A Çelik, Hilal Hekimoglu, Olivo Miotto, Muhammed Miran Öncel, Asif M. Khan
Briefings Bioinform.3
2025 MRAGNN: Refining urban spatio-temporal prediction of crime occurrence with multi-type crime correlation learning
Shun Wang 0004, Yong Zhang 0029, Xinglin Piao, Xuanqi Lin, Yongli Hu
Expert Syst. Appl.5
2025 MVDT: Multiview Distillation Transformer for View-Invariant Sign Language Translation
abstract
ABSTRACT Sign language translation based on machine learning plays a crucial role in facilitating communication between deaf and hearing individuals. However, due to the complexity and variability of sign language, coupled with limited observation angles, single‐view sign language translation models often underperform in real‐world applications. Although some studies have attempted to improve translation efficiency by incorporating multiview data, challenges, such as feature alignment, fusion, and the high cost of capturing multiview data, remain significant barriers in many practical scenarios. To address these issues, we propose a multiview distillation transformer model (MVDT) for continuous sign language translation. The MVDT introduces a novel distillation mechanism, where a teacher model is designed to learn common features from multiview data, subsequently guiding a student model to extract view‐invariant features using only single‐view input. To evaluate the proposed method, we construct a multiview sign language dataset comprising five distinct views and conduct extensive experiments comparing the MVDT with state‐of‐the‐art methods. Experimental results demonstrate that the proposed model exhibits superior view‐invariant translation capabilities across different views.
Yongli Hu, Huajie Jiang
IET Comput. Vis.2
2025 Multi-view subspace clustering with incomplete graph information
abstract
Abstract The core of multi‐view clustering is how to exploit the shared and specific information of multi‐view data properly. The data missing and incompleteness bring great challenges to multi‐view clustering. In this paper, we propose an innovative multi‐view subspace clustering method with incomplete graph information, so‐called incomplete multiple graphs clustering. Specifically, we creatively separate one shared and multiple specific graphs from multiple raw graph data, and exploit the mask fusion strategy and block diagonal regulariser to obtain the inherent category information. To handle the incomplete multiple graph data, we utilise multiple indicator matrices to mark the missing elements existed in each raw graph. In addition, the weight of each raw graph is adaptively learnt according to the graph importance. The alternative direction optimization algorithm is employed to solve our proposed methods. Finally, we also analyse the algorithm convergence and the computation complexity in detail. The clustering results on six real‐world datasets show that our method obviously outperforms a serious of classic incomplete multi‐view clustering methods.
Xiaxia He, Boyue Wang, Cuicui Luo, Junbin Gao, Yongli Hu
IET Comput. Vis.5
2025 Latent attribute augmented network for few-shot class-incremental learning
Yongli Hu, Jiasen Zhang, Huajie Jiang
Neurocomputing1
2025 PRNet: A Progressive Refinement Network for referring image segmentation
Jing Liu 0059, Huajie Jiang, Yongli Hu
Neurocomputing3
2025 SHKD: A framework for traffic prediction based on Sub-Hypergraph and Knowledge Distillation
Xiangyu Yao, Xinglin Piao, Qitan Shao, Yongli Hu, Yong Zhang 0029
Knowl. Based Syst.4
2025 Multi-view Isolated sign language recognition based on cross-view and multi-level transformer
Yongli Hu, Huajie Jiang
Multim. Syst.2
2025 Memory Transmission Based Referring Video Object Segmentation
Zijin Liu, Lichun Wang 0002, Yongli Hu
Neural Networks3
2025 Multi-Granularity Feature Interaction and Multi-Region Selection Based Triplet Visual Question Answering
abstract
Accurately locating the question-related regions in one given image is crucial for visual question answering (VQA). The current approaches suffer two limitations: (1) Dividing one image into multiple regions may lose parts of semantic information and original relationships between regions; (2) Choosing only one or all image regions to predict the answer may correspondingly result in the insufficiency or redundancy of information. Therefore, how to effectively mine the relationship between image regions and choose the relevant image regions are vital. In this paper, we propose a novelMulti-granularity feature interaction andMulti-region selection-based triplet VQA model (M2TVQA). To tackle the first limitation, we propose the multi-granularity feature interaction strategy that adaptively supplements the global coarse-granularity features with the regional fine-granularity features. To overcome the second limitation, we design the Top-$K$learning strategy to adaptively select$K$most relevant image regions to the question, even if the selected regions are far away in space. Such a strategy can select as many relevant image regions as possible and reduce introducing noise. Finally, we construct the multi-modality triplet to predict the answer of VQA. Extended experiments on two public outside knowledge datasets OK-VQA and KRVQA verify the effectiveness of the proposed model.
Boyue Wang, Junbin Gao, Yongli Hu
IEEE Trans. Big Data6
2025 Training Large-Scale Graph Neural Networks via Graph Partial Pooling
abstract
Graph Neural Networks (GNNs) are powerful tools for graph representation learning, but they face challenges when applied to large-scale graphs due to substantial computational costs and memory requirements. To address scalability limitations, various methods have been proposed, including samplingbased and decoupling-based methods. However, these methods have their limitations: sampling-based methods inevitably discard some link information during the sampling process, while decoupling-based methods require alterations to the model's structure, reducing their adaptability to various GNNs. This paper proposes a novel graph pooling method, Graph Partial Pooling (GPPool), for scaling GNNs to large-scale graphs. GPPool is a versatile and straightforward technique that enhances training efficiency while simultaneously reducing memory requirements. GPPool constructs small-scale pooled graphs by pooling partial nodes into supernodes. Each pooled graph consists of supernodes and unpooled nodes, preserving valuable local and global information. Training GNNs on these graphs reduces memory demands and enhances their performance. Additionally, this paper provides a theoretical analysis of training GNNs using GPPool-constructed graphs from a graph diffusion perspective. It shows that a GNN can be transformed from a large-scale graph into pooled graphs with minimal approximation error. A series of experiments on datasets of varying scales demonstrates the effectiveness of GPPool.
Qi Zhang 0095, Shaofan Wang 0001, Junbin Gao, Yongli Hu
IEEE Trans. Big Data5
2025 Multi-Modal Entity in One Word: Aligning Multi-Level Semantics for Multi-Modal Knowledge Graph Completion
abstract
Current multi-modal knowledge graph completion often incorporates simple fusion neural networks to achieve multi-modal alignment and knowledge completion tasks, which face three major challenges: 1) Inconsistent semantics between images and texts corresponding to the same entity; 2) Discrepancies in semantic spaces resulting from the use of diverse uni-modal feature extractors;3) Inadequate evaluation of semantic alignment using only energy functions or basic contrastive learning losses. To address these challenges, we propose the Multi-modal Entity in One Word (MEOW) model. This model ensures alignment at various levels, including text-image match alignment, feature alignment and distribution alignment. Specificially, the entity image filtering module utilizes a visual-language model to exclude unrelated images by aligning their captions with corresponding text descriptions. A pre-trained CLIP-based encoder is utilized for encoding dense semantic relationships, while a graph attention network based structure encoder handles sparse semantic relationships, yielding a comprehensive semantic representation and enhancing convergence speed. Additionally, a diffusion model is integrated to enhance denoising capabilities. The proposed MEOW further includes a distribution alignment module equipped with dense alignment constraint, integrity alignment constraint, and fusion fidelity constraint to effectively align multi-modal representations. Experiments on two public multi-modal knowledge graph datasets show that MEOW significantly improves link prediction performance. The code of the proposed model is available athttps://github.com/yuyuyuger/MEOW.
Boyue Wang, Junbin Gao, Yongli Hu
IEEE Trans. Big Data5
2025 Dual Prototype Contrastive Network for Generalized Zero-Shot Learning
abstract
Generalized zero-shot learning (GZSL) requires that models are able to recognize classes they were trained on, and new classes they haven't seen before. Feature-generation approaches are popular due to their effectiveness in mitigating overfitting to the training classes. Existing generative approaches usually adopt simple discriminators for distribution or classification supervision, however, thus limiting their ability to generate visual features that are discriminative of and transferable to novel categories. To overcome this limitation and improve the quality of generated features, we propose a dual prototype contrastive augmented discriminator for the generative adversarial network. Specifically, we design a Dual Prototype Contrastive Network (DPCN), which leverages complementary information between visual space and semantic space through multi-task prototype contrastive learning. Contrastive learning of the visual prototypes enhances the ability of the generated features to distinguish between classes, while the contrastive learning of the semantic prototypes improves their transferability. Furthermore, we introduce margins into the contrastive learning process to ensure both intra-class compactness and inter-class separation. To demonstrate the effectiveness of the proposed approach, we conduct experiments on three widely-used zero-shot learning benchmark datasets, where DPCN achieves state-of-the-art performance for GZSL.
Huajie Jiang, Zhengxian Li, Yongli Hu, Jian Yang 0001, Anton van den Hengel, Ming-Hsuan Yang 0001, Yuankai Qi
IEEE Trans. Circuits Syst. Video Technol.3
2025 Tackling Real-World Complexity: Hierarchical Modeling and Dynamic Prompting for Multimodal Long Document Classification
abstract
With the rapid growth of internet content, multimodal long document data has become increasingly prominent, drawing significant attention from researchers. However, most existing methods primarily focus on scenarios where all modalities are present, often overlooking more challenging and realistic cases involving missing image modality. To address this limitation, we propose a robust multimodal long document classification (MLDC) framework that integrates hierarchical modeling and dynamic prompting to handle complex multimodal long document data. Our approach begins by leveraging hierarchical modeling combined with an Adaptive Correlation Multimodal Transformer (ACMT) to effectively capture relationships between text and images at both section and sentence levels. We also introduce a Dynamic Prompt Generation (DPG) module at both levels to enhance the model’s robustness in handling missing image data. By evaluating sample uncertainty, the DPG module dynamically adjusts both the number of prompts and the prompts themselves, allowing the model to better adapt to the varying needs of different samples. Finally, a Hierarchical Heterogeneous Graph (HHG) is introduced to enhance feature interactions across levels, further improving the coherence and accuracy of the model. Extensive experiments on four multi-modal long document datasets demonstrate that our model shows superior performance compared to existing state-of-the-art MLDC classification methods in various conditions.
Tengfei Liu 0005, Yongli Hu, Mingjie Li 0006, Junfei Yi, Xiaojun Chang, Junbin Gao
IEEE Trans. Circuits Syst. Video Technol.2
2025 CMDNet: A Cross-Modality Spatiotemporal Graph Network for Enhanced Air Pollution Prediction With High-Resolution Satellite Data
abstract
Predicting air pollution plays a vital role in urban management and public health by providing early warnings on PM2.5, SO2, and NO2 concentrations, helping to mitigate the adverse effects of these pollutants. Traditional prediction methods, relying on physical and statistical models, often struggle to capture the complex spatio-temporal dependencies and dynamic characteristics of air pollution data. The application of deep learning methods, especially graph neural networks (GNNs), has shown promise in addressing these limitations. However, existing GNN-based methods ignore the integration of rich semantic information provided by high-resolution satellite data. To address this problem, we propose a Cross-Modality Dynamic Spatio-Temporal Graph Neural Network (CMDNet) for air pollution prediction. The model comprises two branches: a dynamic spatio-temporal graph neural network branch and a remote sensing image dynamic encoding network branch. The dynamic spatiotemporal graph neural network branch captures the spatiotemporal dependencies in air pollution data by constructing a dynamic graph structure. The remote sensing image dynamic encoding network branch extracts semantic information from high-resolution satellite data, which improves the model’s power to perceive air pollution conditions in different regions. Experiments on real-world datasets demonstrate that CMDNet achieves better air pollution prediction results than existing SOTA models, with maximum improvements of 2.4% (MAE), 1.8% (RMSE), 1.3% (CSI), 1.6% (FAR), and 1.4% (POD), providing more accurate prediction results.
Shun Wang 0004, Yong Zhang 0029, Xuanqi Lin, Xinglin Piao, Yongli Hu
IEEE Trans. Geosci. Remote. Sens.5
2025 Toward Nonuniformly Distributed Weather Forecasting: Adaptive Filtered Hypergraph Convolution Network
abstract
Weather forecasting, compared to other multivariate time-series prediction tasks, exhibits notably non-uniform distribution of observation sites. The prevailing approaches primarily leverage graph convolutional neural networks to extract spatial features. However, some studies suggest that the poor performance on uneven graphs is primarily due to the fact that traditional graph neural networks (GNNs) are essentially low-pass filters, discarding information beyond low-frequency information on the graph. From another perspective, since the essence of graph convolution is the smoothing of node features, for uneven graphs, there are noticeable differences in the smoothing rates of node features, leading to the coexistence of overfitting and underfitting phenomena. This issue is further exacerbated in higher-order graph structures, such as hypergraphs, due to the irregular and complex nature of hyperedges. To address this issue, we propose a filtered hypergraph neural network. Building on the calculation of hypergraph node smoothing rates, we balance the low-pass and high-pass filter convolutions’ feature extraction through a dual-stream architecture. On uneven graphs, it can be observed that neglecting high-frequency information and concentrating solely on low-frequency information impede the learning of node representations, thereby significantly affecting the performance of downstream prediction tasks. We conducted multi-dimensional time-series prediction experiments using meteorological data, and the results demonstrate that our model performs with high accuracy in node regression tasks across multiple channels.
Yong Zhang 0029, Guodong Jing, Yongli Hu
IEEE Trans. Geosci. Remote. Sens.5
2025 OST-HGCN: Optimized Spatial-Temporal Hypergraph Convolution Network for Trajectory Prediction
abstract
Pedestrian trajectory prediction is a key component for various applications that involve human and vehicle interactions, such as autonomous driving, traffic management and smart city planning. Existing methods based on graph neural networks have limited ability to capture group interactions and precisely model complex associations among multi-agents. To solve these problems, we propose OST-HGCN, an optimized hypergraph convolutional network. It models multi-agent trajectory interactions from both temporal and spatial perspectives using hypergraph structures, and optimizes the spatio-temporal hypergraph structure to enable fine-grained analysis of multi-agent trajectory motion intentions and high-order interactions. We employ OST-HGCN to a CVAE-based prediction framework, and use the optimized hypergraph structure to predict multi-agent plausible trajectories. We conduct extensive experiments on four real trajectory prediction datasets of NBA, NFL, SDD and ETH-UCY, and verify the effectiveness of the proposed OST-HGCN.
Xuanqi Lin, Yong Zhang 0029, Shun Wang 0004, Yongli Hu
IEEE Trans. Intell. Transp. Syst.4
2025 HGSCO: Heterogeneous Graph Structure Contrast Optimization for Trajectory Prediction
abstract
Predicting and planning the future trajectories of various traffic participants is an important task with multiple applications, including autonomous vehicles, service robots, and intelligent transportation. However, the diversity of heterogeneous agents including pedestrians, bicycles, and vehicles in traffic scenarios presents substantial challenges to this task. Current models do not fully capture the implicit and explicit interaction relationships among these heterogeneous agents and often overlook the significance of extracting implicit correlations from agent features. To address these issues, we introduce a novel model for trajectory prediction: Heterogeneous Graph Structure Contrast Optimization (HGSCO). To accurately capture the interaction relationships among heterogeneous agents, HGSCO constructs semantic graph structures representing implicit relationships and meta-path graph structures representing explicit relationships. Then, HGSCO introduces a cross-view contrastive learning approach, which optimizes the heterogeneous graph structure by maximizing mutual information between the two types of graph structures. The model can provide precise interaction relationships among heterogeneous agents by effectively fusing these two graph representations with a gated fusion method. We utilized video data captured by camera sensors in complex environments with multiple agents to conduct experiments. Our proposed model achieved an 11.5% and 6.7% reduction in Average Displacement Error (ADE) across these datasets, respectively, and a reduction of 15.6% and 8.1% in Final Displacement Error (FDE). The results demonstrate that HGSCO significantly surpasses existing state-of-the-art methods regarding trajectory prediction accuracy.
Xuanqi Lin, Yong Zhang 0029, Shun Wang 0004, Xinglin Piao, Yongli Hu
IEEE Trans. Intell. Transp. Syst.5
2025 SAGoG: Similarity-Aware Graph of Graphs Neural Networks for Multivariate Time Series Classification
abstract
Multivariate Time Series Classification (MTSC) has important research significance and practical value. Deep learning models have achieved considerable success in addressing MTSC problems. However, a key challenge faced by existing classification models is how to effectively consider the correlations between time series instances and across channels simultaneously, as well as how to capture the dynamic of these inter-channel correlations over time. Current methods often fall short in these aspects: on one hand, they fail to fully account for the combined effects of inter-instance and inter-channel correlations; on the other hand, they largely overlook the dynamic nature of how inter-channel correlations change over time. To address these issues, we propose a novel graph neural network model, called Similarity-Aware Graph of Graphs neural networks (SAGoG), for multivariate time series classification. This model can comprehensively consider the dependencies between channel-level and instance-level time series, it dynamically learns dependency features through graph structure evolution and graph pooling layers. We conduct experiments on the UEA dataset to validate the SAGoG model, and the results demonstrate its outstanding performance in multivariate time series classification tasks.
Shun Wang 0004, Yong Zhang 0029, Xuanqi Lin, Yongli Hu, Qingming Huang
IEEE Trans. Knowl. Data Eng.4
2025 Hierarchical Multi-Modal Transformer for Cross-Modal Long Document Classification
abstract
Long Document Classification (LDC) has gained significant attention recently. However, multi-modal data in long documents such as texts and images are not being effectively utilized. Prior studies in this area have attempted to integrate texts and images in document-related tasks, but they have only focused on short text sequences and images of pages. How to classify long documents with hierarchical structure texts and embedding images is a new problem and faces multi-modal representation difficulties. In this paper, we propose a novel approach called Hierarchical Multi-modal Transformer (HMT) for cross-modal long document classification. The HMT conducts multi-modal feature interaction and fusion between images and texts in a hierarchical manner. Our approach uses a multi-modal transformer and a dynamic multi-scale multi-modal transformer to model the complex relationships between image features, and the section and sentence features. Furthermore, we introduce a new interaction strategy called the dynamic mask transfer module to integrate these two transformers by propagating features between them. To validate our approach, we conduct cross-modal LDC experiments on two newly created and two publicly available multi-modal long document datasets, and the results show that the proposed HMT outperforms state-of-the-art single-modality and multi-modality methods.
Tengfei Liu 0005, Yongli Hu, Junbin Gao
IEEE Trans. Multim.2
2025 Globality Meets Locality: An Anchor Graph Collaborative Learning Framework for Fast Multiview Subspace Clustering
abstract
Multiview subspace clustering (MSC) maximizes the utilization of complementary description information provided by multiview data and achieves impressive clustering performance. However, most of them are inefficient or even invalid among large-scale scenarios due to expensive computational complexity. Recently, anchor strategy has been developed to address this, which selects a few representative samples as anchor points for representation learning and anchor graph construction. However, most of them only explore single cross-view correlation, i.e., cross-view consistency from the global aspect or cross-view complementarity from the local aspect, which provides insufficient semantic correlation understanding and exploration for complex multiview data. To effectively address this issue, this study proposes a fast multiview subspace clustering (FMSC) with local-global anchor representation collaborative learning. FMSC integrates the discriminative anchor points learning and anchor graph construction with optimal structure into a joint framework. Furthermore, local (view-specific) and global (view-shared) anchor representations are learned collaboratively under two interaction strategies at different levels, providing beneficial guidance from global learning to local learning. Thus, the proposed FMSC can maximize the exploration of the complementarity-consistency among multiview data and capture a more comprehensive semantic correlation. More importantly, an effective algorithm with linear complexity is designed to solve the corresponding optimization problem of FMSC, making it more practical in large-scale clustering tasks. Extensive experimental results confirm the superiority of the proposed FMSC in both clustering performance and computational efficiency.
Jipeng Guo 0001, Xin Ma 0012, Junbin Gao, Yongli Hu, Youqing Wang
IEEE Trans. Neural Networks Learn. Syst.5
2025 Redundancy is Not What You Need: An Embedding Fusion Graph Auto-Encoder for Self-Supervised Graph Representation Learning
abstract
Attribute graphs are a crucial data structure for graph communities. However, the presence of redundancy and noise in the attribute graph can impair the aggregation effect of integrating two different heterogeneous distributions of attribute and structural features, resulting in inconsistent and distorted data that ultimately compromises the accuracy and reliability of attribute graph learning. For instance, redundant or irrelevant attributes can result in overfitting, while noisy attributes can lead to underfitting. Similarly, redundant or noisy structural features can affect the accuracy of graph representations, making it challenging to distinguish between different nodes or communities. To address these issues, we propose the embedded fusion graph auto-encoder framework for self-supervised learning (SSL), which leverages multitask learning to fuse node features across different tasks to reduce redundancy. The embedding fusion graph auto-encoder (EFGAE) framework comprises two phases: pretraining (PT) and downstream task learning (DTL). During the PT phase, EFGAE uses a graph auto-encoder (GAE) based on adversarial contrastive learning to learn structural and attribute embeddings separately and then fuses these embeddings to obtain a representation of the entire graph. During the DTL phase, we introduce an adaptive graph convolutional network (AGCN), which is applied to graph neural network (GNN) classifiers to enhance recognition for downstream tasks. The experimental results demonstrate that our approach outperforms state-of-the-art (SOTA) techniques in terms of accuracy, generalization ability, and robustness.
Mengran Li 0001, Yong Zhang 0029, Shaofan Wang 0001, Yongli Hu
IEEE Trans. Neural Networks Learn. Syst.4
2025 Counterfactual Dual-Bias VQA: A Multimodality Debias Learning for Robust Visual Question Answering
abstract
Visual question answering (VQA) models often face two language bias challenges. First, they tend to rely solely on the question to predict the answer, often overlooking relevant information in the accompanying images. Second, even when considering the question, they may focus only on the wh-words, neglecting other crucial keywords that could enhance interpretability and the question sensitivity. Existing debiasing methods attempt to address this by training a bias model using question-only inputs to enhance the robustness of the target VQA model. However, this approach may not fully capture the language bias present. In this article, we propose a multimodality counterfactual dual-bias model to mitigate the linguistic bias issue in target VQA models. Our approach involves designing a shared-parameterized dual-bias model that incorporates both visual and question counterfactual samples as inputs. By doing so, we aim to fully model language biases, with visual and question counterfactual samples, respectively, emphasizing important objects and keywords to relevant the answers. To ensure that our dual-bias model behaves similarly to an ordinary model, we freeze the parameters of the target VQA model, meanwhile using the cross-entropy and Kullback-Leibler (KL) divergence as the loss function to train the dual-bias model. Subsequently, to mitigate language bias in the target VQA model, we freeze the parameters of the dual-bias model to generate pseudo-labels and then incorporate a margin loss to re-train the target VQA model. Experimental results on the VQA-CP datasets demonstrate the superior effectiveness of our proposed counterfactual dual-bias model. Additionally, we conduct an analysis of the unsatisfactory performance on the VQA v2 dataset. The origin code of the proposed model is available at https://github.com/Arrow2022jv/MCD.
Boyue Wang, Xiaoqian Ju, Junbin Gao, Yongli Hu
IEEE Trans. Neural Networks Learn. Syst.5
2025 Bridging the Cross-Modality Semantic Gap in Visual Question Answering
abstract
The objective of visual question answering (VQA) is to adequately comprehend a question and identify relevant contents in an image that can provide an answer. Existing approaches in VQA often combine visual and question features directly to create a unified cross-modality representation for answer inference. However, this kind of approach fails to bridge the semantic gap between visual and text modalities, resulting in a lack of alignment in cross-modality semantics and the inability to match key visual content accurately. In this article, we propose a model called the caption bridge-based cross-modality alignment and contrastive learning model (CBAC) to address the issue. The CBAC model aims to reduce the semantic gap between different modalities. It consists of a caption-based cross-modality alignment module and a visual-caption (V-C) contrastive learning module. By utilizing an auxiliary caption that shares the same modality as the question and has closer semantic associations with the visual, we are able to effectively reduce the semantic gap by separately matching the caption with both the question and the visual to generate pre-alignment features for each, which are then used in the subsequent fusion process. We also leverage the fact that V-C pairs exhibit stronger semantic connections compared to question-visual (Q-V) pairs to employ a contrastive learning mechanism on visual and caption pairs to further enhance the semantic alignment capabilities of single-modality encoders. Extensive experiments conducted on three benchmark datasets demonstrate that the proposed model outperforms previous state-of-the-art VQA models. Additionally, ablation experiments confirm the effectiveness of each module in our model. Furthermore, we conduct a qualitative analysis by visualizing the attention matrices to assess the reasoning reliability of the proposed model.
Boyue Wang, Yujian Ma, Junbin Gao, Yongli Hu
IEEE Trans. Neural Networks Learn. Syst.5
2025 Modality Perception Learning-Based Determinative Factor Discovery for Multimodal Fake News Detection
abstract
The dissemination of fake news, often fueled by exaggeration, distortion, or misleading statements, significantly jeopardizes public safety and shapes social opinion. Although existing multimodal fake news detection methods focus on multimodal consistency, they occasionally neglect modal heterogeneity, missing the opportunity to unearth the most related determinative information concealed within fake news articles. To address this limitation and extract more decisive information, this article proposes the modality perception learning-based determinative factor discovery (MoPeD) model. MoPeD optimizes the steps of feature extraction, fusion, and aggregation to adaptively discover determinants within both unimodality features and multimodality fusion features for the task of fake news detection. Specifically, to capture comprehensive information, the dual encoding module integrates a modal-consistent contrastive language-image pre-training (CLIP) pretrained encoder with a modal-specific encoder, catering to both explicit and implicit information. Motivated by the prompt strategy, the output features of the dual encoding module are complemented by learnable memory information. To handle modality heterogeneity during fusion, the multilevel cross-modality fusion module is introduced to deeply comprehend the complex implicit meaning within text and image. Finally, for aggregating unimodal and multimodal features, the modality perception learning module gauges the similarity between modalities to dynamically emphasize decisive modality features based on the cross-modal content heterogeneity scores. The experimental evaluations conducted on three public fake news datasets show that the proposed model is superior to other state-of-the-art fake news detection methods.
Boyue Wang, Guangchao Wu, Junbin Gao, Yongli Hu
IEEE Trans. Neural Networks Learn. Syst.5
2024 VIG: Visual Information-Guided Knowledge-Based Visual Question Answering
abstract
A better knowledge-based visual question answering (KBVQA) model needs to rely on visual features, question features, and related external knowledge to solve an open visual question answering task. Although the existing knowledge-based visual question answering works have achieved some accomplishments, there are still the following challenges: 1) There is a serious lack of visual feature information. Image information is worth a thousand words. Only relying on the converted salient text information is difficult to express the original rich information of the image. 2) The external knowledge acquired is not comprehensive enough, and there is a lack of relevant knowledge directly retrieved by visual feature information. To solve these challenges, we propose a Visual Information-Guided knowledge-based visual question answering (VIG) model. It fully considers the utilization of visual features information. Specifically: 1) We introduce multi-granularity visual information that can comprehensively characterize visual feature information. 2) We consider not only the knowledge retrieved through text information but also the knowledge directly retrieved from visual feature information. Finally, we feed the visual features and retrieved multiple text knowledge into an encoder-decoder module to generate an answer. We perform extensive experiments on the OKVQA dataset and achieve state-of-the-art performance of 60.27% accuracy.
Boyue Wang, Yongli Hu
CSCWD5
2024 VQA-PDF: Purifying Debiased Features for Robust Visual Question Answering Task
Yandong Bi, Huajie Jiang, Jing Liu 0059, Yongli Hu
ICIC (12)5
2024 Referring Image Segmentation Without Text Annotations
Jing Liu 0059, Huajie Jiang, Yandong Bi, Yongli Hu
ICIC (12)4
2024 Generating Graph-Based Rules for Enhancing Logical Reasoning
Huajie Jiang, Yongli Hu
ICIC (12)3
2024 Common-Memory Bridged Cross-Modal Adaptive Graph Embedding for Image-Text Retrieval
abstract
The task of image-text retrieval has gained significant attention in the realm of multimodal artificial intelligence. Nonetheless, existing works encounter challenges in efficiently utilizing inter-modal information and adequately leveraging crucial intra-modal details. In this paper, we propose a novel common-Memory Bridged cross-modal Adaptive Graph Embedding (MBAGE) network for image-text retrieval. Initially, we represent images and text as graphs, wherein nodes symbolize salient regions and words. Subsequently, we incorporate a common-memory bank as an intermediate bridge, facilitating interactions between nodes in the two graphs and enabling efficient utilization of inter-modal information. Additionally, we propose an adaptive graph convolutional network to implement intra-modal interaction, which can adaptively suppress the learning of unimportant nodes. Finally, adaptive pooling is employed to retain essential information, yielding a superior holistic embedding. Experimental results on the Flickr30K and MS-COCO datasets demonstrate that the MBAGE network not only achieves compelling retrieval precision but also exhibits high retrieval efficiency.
Yongli Hu, Jiapu Wang, Junbin Gao
ICME2
2024 Large Language Models-guided Dynamic Adaptation for Temporal Knowledge Graph Reasoning
abstract
Temporal Knowledge Graph Reasoning (TKGR) is the process of utilizing temporal information to capture complex relations within a Temporal Knowledge Graph (TKG) to infer new knowledge. Conventional methods in TKGR typically depend on deep learning algorithms or temporal logical rules. However, deep learning-based TKGRs often lack interpretability, whereas rule-based TKGRs struggle to effectively learn temporal rules that capture temporal patterns. Recently, Large Language Models (LLMs) have demonstrated extensive knowledge and remarkable proficiency in temporal reasoning. Consequently, the employment of LLMs for Temporal Knowledge Graph Reasoning (TKGR) has sparked increasing interest among researchers. Nonetheless, LLMs are known to function as black boxes, making it challenging to comprehend their reasoning process. Additionally, due to the resource-intensive nature of fine-tuning, promptly updating LLMs to integrate evolving knowledge within TKGs for reasoning is impractical. To address these challenges, in this paper, we propose a Large Language Models-guided Dynamic Adaptation (LLM-DA) method for reasoning on TKGs. Specifically, LLM-DA harnesses the capabilities of LLMs to analyze historical data and extract temporal logical rules. These rules unveil temporal patterns and facilitate interpretable reasoning. To account for the evolving nature of TKGs, a dynamic adaptation strategy is proposed to update the LLM-generated rules with the latest events. This ensures that the extracted rules always incorporate the most recent knowledge and better generalize to the predictions on future events. Experimental results show that without the need of fine-tuning, LLM-DA significantly improves the accuracy of reasoning over several common datasets, providing a robust framework for TKGR tasks.
Jiapu Wang, Linhao Luo, Yongli Hu, Alan Wee-Chung Liew, Shirui Pan
NeurIPS5
2024 Context-aware relation enhancement and similarity reasoning for image-text retrieval
abstract
Abstract Image‐text retrieval is a fundamental yet challenging task, which aims to bridge a semantic gap between heterogeneous data to achieve precise measurements of semantic similarity. The technique of fine‐grained alignment between cross‐modal features plays a key role in various successful methods that have been proposed. Nevertheless, existing methods cannot effectively utilise intra‐modal information to enhance feature representation and lack powerful similarity reasoning to get a precise similarity score. Intending to tackle these issues, a context‐aware Relation Enhancement and Similarity Reasoning model, called RESR, is proposed, which conducts both intra‐modal relation enhancement and inter‐modal similarity reasoning while considering the global‐context information. For intra‐modal relation enhancement, a novel context‐aware graph convolutional network is introduced to enhance local feature representations by utilising relation and global‐context information. For inter‐modal similarity reasoning, local and global similarity features are exploited by the bidirectional alignment of image and text, and the similarity reasoning is implemented among multi‐granularity similarity features. Finally, refined local and global similarity features are adaptively fused to get a precise similarity score. The experimental results show that our effective model outperforms some state‐of‐the‐art approaches, achieving average improvements of 2.5% and 6.3% in R@sum on the Flickr30K and MS‐COCO dataset.
Yongli Hu
IET Comput. Vis.2
2024 MBMF: Constructing memory banks of multi-scale features for anomaly detection
abstract
Abstract In industrial manufacturing, how to accurately classify defective products and locate the location of defects has always been a concern. Previous studies mainly measured similarity based on extracting single‐scale features of samples. However, only using the features of a single scale is hard to represent different sizes and types of anomalies. Therefore, the authors propose a set of memory banks of multi‐scale features (MBMF) to enrich feature representation and detect and locate various anomalies. To extract features of different scales, different aggregation functions are designed to produce the feature maps at different granularity. Based on the multi‐scale features of normal samples, the MBMF are constructed. Meanwhile, to better adapt to the feature distribution of the training samples, the authors proposed a new iterative updating method for the memory banks. Testing on the widely used and challenging dataset of MVTec AD, the proposed MBMF achieves competitive image‐level anomaly detection performance (Image‐level Area Under the Receiver Operator Curve (AUROC)) and pixel‐level anomaly segmentation performance (Pixel‐level AUROC). To further evaluate the generalisation of the proposed method, we also implement anomaly detection on the BeanTech AD dataset, a commonly used dataset in the field of anomaly detection, and the Fashion‐MNIST dataset, a widely used dataset in the field of image classification. The experimental results also verify the effectiveness of the proposed method.
Yongli Hu, Huajie Jiang
IET Comput. Vis.3
2024 Cross-modal fusion encoder via graph neural network for referring image segmentation
abstract
Abstract Referring image segmentation identifies the object masks from images with the guidance of input natural language expressions. Nowadays, many remarkable cross‐modal decoder are devoted to this task. But there are mainly two key challenges in these models. One is that these models usually lack to extract fine‐grained boundary information and gradient information of images. The other is that these models usually lack to explore language associations among image pixels. In this work, a Multi‐scale Gradient balanced Central Difference Convolution (MG‐CDC) and a Graph convolutional network‐based Language and Image Fusion (GLIF) for cross‐modal encoder, called Graph‐RefSeg, are designed. Specifically, in the shallow layer of the encoder, the MG‐CDC captures comprehensive fine‐grained image features. It could enhance the perception of target boundaries and provide effective guidance for deeper encoding layers. In each encoder layer, the GLIF is used for cross‐modal fusion. It could explore the correlation of every pixel and its corresponding language vectors by a graph neural network. Since the encoder achieves robust cross‐modal alignment and context mining, a light‐weight decoder could be used for segmentation prediction. Extensive experiments show that the proposed Graph‐RefSeg outperforms the state‐of‐the‐art methods on three public datasets. Code and models will be made publicly available at https://github.com/ZYQ111/Graph_refseg .
Yong Zhang 0029, Xinglin Piao, Yongli Hu
IET Image Process.5
2024 Contrastive optimized graph convolution network for traffic forecasting
Kan Guo, Daxin Tian, Yongli Hu, Zhen (sean) Qian, Jianshan Zhou, Junbin Gao
Neurocomputing3
2024 A plug-and-play image enhancement model for end-to-end object detection in low-light condition
Jiaojiao Yuan, Yongli Hu, Boyue Wang
Multim. Syst.2
2024 IE-GAN: a data-driven crowd simulation method via generative adversarial networks
Xuanqi Lin, Yong Zhang 0029, Yongli Hu
Multim. Tools Appl.4
2024 Multi-modal long document classification based on Hierarchical Prompt and Multi-modal Transformer
Tengfei Liu 0005, Yongli Hu, Junbin Gao, Jiapu Wang
Neural Networks2
2024 Multi-scale hypergraph-based feature alignment network for cell localization
Bo Li 0128, Yong Zhang 0029, Xinglin Piao, Yongli Hu
Pattern Recognit.5
2024 Self-supervised knowledge distillation in counterfactual learning for VQA
Yandong Bi, Huajie Jiang, Hanfu Zhang, Yongli Hu
Pattern Recognit. Lett.4
2024 Hierarchical Multi-Granularity Interaction Graph Convolutional Network for Long Document Classification
abstract
With the growing demand for text analytics, long document classification (LDC) has received extensive attention, and great progress has been made. To reveal the complex structure and extract the intrinsic feature, the current approaches focus on modeling a long sequence with sparse attention or representing word-sentence or word-section relations partially. However, the thorough hierarchical structure from words, sentences to sections of long documents remains relatively unexplored. For this purpose, we propose a novel Hierarchical Multi-granularity Interaction Graph Convolutional Network (HMIGCN) for long document classification, in which three different granularity graphs, i.e., section graph, sentence graph and word graph, are constructed hierarchically. The section graph encapsulates the macrostructure of a long document, while the sentence and word graphs delve into the document's microstructure. Notably, within the sentence graph, we introduce a Global-Local Graph Convolutional (GLGC) block to adaptively capture both global and local dependency structures among sentence nodes. Additionally, to integrate the three graph networks as a whole, two well-designed techniques, namely section-guided pooling block and transfer fusion block, are proposed to train the model jointly by promoting each other. Extensive experiments on five long document datasets show that our model outperforms the existing state-of-the-art LDC models.
Tengfei Liu 0005, Yongli Hu, Junbin Gao
IEEE ACM Trans. Audio Speech Lang. Process.2
2024 Multi-Level Interaction Based Knowledge Graph Completion
abstract
With the continuous emergence of new knowledge, Knowledge Graph (KG) typically suffers from the incompleteness problem, hindering the performance of downstream applications. Thus, Knowledge Graph Completion (KGC) has attracted considerable attention. However, existing KGC methods usually capture the coarse-grained information by directly interacting with the entity and relation, ignoring the important fine-grained information in them. To capture the fine-grained information, in this paper, we divide each entity/relation into several segments and propose a novelMulti-LevelInteraction (MLI) based KGC method, which simultaneously interacts with the entity and relation at the fine-grained level and the coarse-grained level. The fine-grained interaction module applies the Gate Recurrent Unit (GRU) mechanism to guarantee the sequentiality between segments, which facilitates the fine-grained feature interaction and does not obviously sacrifice the model complexity. Moreover, the coarse-grained interaction module designs aHigh-orderFactorizedBilinear (HFB) operation to facilitate the coarse-grained interaction between the entity and relation by applying the tensor factorization based multi-head mechanism, which still effectively reduces its parameter scale. Experimental results show that the proposed method achieves state-of-the-art performances on the link prediction task over five well-established knowledge graph completion benchmarks.
Jiapu Wang, Boyue Wang, Junbin Gao, Yongli Hu
IEEE ACM Trans. Audio Speech Lang. Process.5
2024 Dynamic Hypergraph Structure Learning for Multivariate Time Series Forecasting
abstract
Multivariate time series forecasting plays an important role in many domain applications, such as air pollution forecasting and traffic forecasting. Modeling the complex dependencies among time series is a key challenging task in multivariate time series forecasting. Many previous works have used graph structures to learn inter-series correlations, which have achieved remarkable performance. However, graph networks can only capture spatio-temporal dependencies between pairs of nodes, which cannot handle high-order correlations among time series. We propose a Dynamic Hypergraph Structure Learning model (DHSL) to solve the above problems. We generate dynamic hypergraph structures from time series data using the K-Nearest Neighbors method. Then a dynamic hypergraph structure learning module is used to optimize the hypergraph structure to obtain more accurate high-order correlations among nodes. Finally, the hypergraph structures dynamically learned are used in the spatio-temporal hypergraph neural network. We conduct experiments on six real-world datasets. The prediction performance of our model surpasses existing graph network-based prediction models. The experimental results demonstrate the effectiveness and competitiveness of the DHSL model for multivariate time series forecasting.
Shun Wang 0004, Yong Zhang 0029, Xuanqi Lin, Yongli Hu, Qingming Huang
IEEE Trans. Big Data4
2024 Hardware-Friendly 3-D CNN Acceleration With Balanced Kernel Group Sparsity
abstract
Being capable of extracting more information than 2D Convolutional Neural Networks (CNNs), 3D CNNs have been playing a vital role in video analysis tasks like human action recognition, but their massive operations hinder the real-time execution on edge devices with constrained computation and memory resources. Although various model compression techniques have been applied to accelerate 2D CNNs, there are rare efforts in investigating hardware-friendly pruning of 3D CNNs and acceleration on customizable edge platforms like FPGAs. This work starts from proposing a kernel group row-column (KGRC) weight sparsity pattern, which is fine-grained to achieve high pruning ratios with negligible accuracy loss, and balanced across kernel groups to achieve high computation parallelism on hardware. The reweighted pruning algorithm for this sparsity is then presented and performed on 3D CNNs, followed by quantization under different precisions. Along with model compression, FPGA-based accelerators with four modes are designed in support of the kernel group sparsity in multiple dimensions. The co-design framework of the pruning algorithm and the accelerator is tested on two representative 3D CNNs, namely C3D and R(2+1)D, with the Xilinx ZCU102 FPGA platform for action recognition. The experimental results indicate that the accelerator implementation with the KGRC sparsity and 8-bit quantization achieves a good balance between the speedup and model accuracy, leading to acceleration ratios of 4.12× for C3D and 3.85× for R(2+1)D compared with the 16-bit baseline designs supporting only dense models.
Mengshu Sun, Kaidi Xu, Xue Lin 0001, Yongli Hu
IEEE Trans. Comput. Aided Des. Integr. Circuits Syst.4
2024 Large-Scale Traffic Prediction With Hierarchical Hypergraph Message Passing Networks
abstract
Graph convolutional networks (GCNs) are widely used in social computation such as urban traffic prediction. However, when faced with city-level forecasting challenges, the graph-based deep learning methods struggle to process large-scale multivariate data effectively. To address the challenges of limited scalability, a traffic prediction framework based on a hypergraph message passing network (HMSG) is proposed in this article. The model represents the urban transportation network with hypergraph, where nodes denote transportation hubs and hyperedges represent their relationship at geographical and feature level. Compared with pairwise edges, hyperedges are more scalable and flexible, providing a more descriptive representation of traffic information. The HMSG algorithm updates node and hyperedge features in two steps, facilitating effective and efficient integration of hidden spatial features across layers. The proposed framework is evaluated on large-scale historical datasets and demonstrates its completion of city-scale traffic prediction tasks. The results also show that it matches the accuracy of existing traffic prediction methods on small-scale datasets. This validates the potential of the traffic prediction model based on the HMSG algorithm for intelligent transportation applications.
Yong Zhang 0029, Yongli Hu
IEEE Trans. Comput. Soc. Syst.3
2024 Traffic Origin-Destination Demand Prediction via Multichannel Hypergraph Convolutional Networks
abstract
Accurate prediction of origin-destination (OD) demand is critical for service providers to efficiently allocate limited resources in regions with high travel demands. However, OD distributions pose significant challenges, characterized by high sparsity, complex spatial correlations within regions or chains, and potential repetition due to the recurrence of similar semantic contexts. These challenges impede traditional graph-based approaches, which connect two vertices through an edge, from performing effectively in OD prediction. Thus, we present a novel multichannel hypergraph convolutional neural network (MC-HGCN) to overcome the above challenges. The model innovatively extracts distinctive features from the channels of inflows, outflows, and OD flows, to conquer the high sparsity in OD matrices. High-order spatial proximity within regions and OD chains are then modeled by the three adjacency hypergraphs constructed for the above three channels. In each adjacency hypergraph, multiple neighboring stations are treated as vertices, while multiple OD pairs constitute hyperedges. These structures are learned by hypergraph convolutional networks for latent spatial correlations. On this basis, a semantic hypergraph is created for the OD channel to model OD distributions lacking spatial proximity but sharing semantic correlations. It utilizes hyperedges to represent semantic correlations among OD pairs whose origins and destinations both possess similar point-of-interest (POI) functions, before learned by a hypergraph convolutional network (HGCN). Both spatial and semantic correlations intrinsic to OD flows are accordingly captured and embedded into a gated recurrent unit (GRU) to unveil hidden spatiotemporal dependencies among OD distributions. These embedded correlations are ultimately integrated through a multichannel fusion module to enhance the prediction of OD flows, even for minor ones. Our model is validated through experiments on three public datasets, demonstrating its robust performances across long and short time steps. Findings may contribute theoretical insights for practical applications, such as coordinating traffic scheduling or route planning.
Yong Zhang 0029, Xia Zhao 0003, Yongli Hu
IEEE Trans. Comput. Soc. Syst.4
2024 See and Learn More: Dense Caption-Aware Representation for Visual Question Answering
abstract
With the rapid development of deep learning models, great improvements have been achieved in the Visual Question Answering (VQA) field. However, modern VQA models are easily affected by language priors, which ignore image information and learn the superficial relationship between questions and answers, even in the optimal pre-training model. The main reason is that visual information is not fully extracted and utilized, which results in a domain gap between vision and language modalities to a certain extent. In order to mitigate the circumstances, we propose to extract dense captions (auxiliary semantic information) from images to enhance the visual information for reasoning and utilize them to release the gap between vision and language since the dense captions and the questions are from the same language modality (i.e., phrase or sentence). In this paper, we propose a novel dense caption-aware visual question answering model called DenseCapBert to enhance visual reasoning. Specifically, we generate dense captions for the images and propose a multimodal interaction mechanism to fuse dense captions, images, and questions in a unified framework, which makes the VQA models more robust. The experimental results on GQA, GQA-OOD, VQA v2, and VQA-CP v2 datasets show that dense captions are beneficial to improving the model generalization and our model effectively mitigates the language bias problem.
Yandong Bi, Huajie Jiang, Yongli Hu
IEEE Trans. Circuits Syst. Video Technol.3
2024 Fair Attention Network for Robust Visual Question Answering
abstract
As a prevailing cross-modal reasoning task, Visual Question Answering (VQA) has achieved impressive progress in the last few years, where the language bias is widely studied to learn more robust VQA models. However, the visual bias, which also influences the robustness of VQA models, is seldomly considered, resulting in weak inference ability. Therefore, how to balance the effect of language bias and visual bias has become essential in the current VQA task. In this paper, we devise a new reweighting strategy taking both the language bias and visual bias into account, and propose a Fair Attention Network for Robust Visual Question Answering (named as FAN-VQA). It first constructs a question bias branch and a visual bias branch to estimate the bias information from two modalities, which are utilized to judge the importance of samples. Then, adaptive importance weights are learned from the bias information and assigned to the candidate answers to adjust the training losses, enabling the model to shift more attention to the difficult samples that need less-salient visual clues to infer the correct answer. In order to improve the robustness of the VQA model, we design a progressive strategy to balance the influence of original training loss and adjusted training loss. Extensive experiments on the VQA-CP v2, VQA v2, and VQA-CE datasets demonstrate the effectiveness of the proposed FAN-VQA method.
Yandong Bi, Huajie Jiang, Yongli Hu
IEEE Trans. Circuits Syst. Video Technol.3
2024 CFMMC-Align: Coarse-Fine Multi-Modal Contrastive Alignment Network for Traffic Event Video Question Answering
abstract
Traffic video question answering (TrafficVQA) constitutes a specialized VideoQA task designed to enhance the basic comprehension and intricate reasoning capacities of videos, specifically focusing on traffic events. Recent VideoQA models employ pretrained visual and textual encoder models to bridge the feature space gap between visual and textual data. However, in addressing the unique challenges inherent to the TrafficVQA task, three pivotal issues must be addressed: (i) Dimension Gap: Between the pretrained image (appearance feature) and video (motion feature) models, there exists a conspicuous dimension difference in static and dynamic visual data; (ii) Scene Gap: The common real-world datasets and the traffic event datasets differ in visual scene content; (iii) Modality Gap: A pronounced feature distribution discrepancy emerges between traffic video and text data. To alleviate these challenges, we introduce the coarse-fine multimodal contrastive alignment network (CFMMC-Align). This model leverages sequence-level and token-level multimodal features, grounded in an unsupervised visual multimodal contrastive loss to mitigate dimension and scene gaps and a supervised visual-textual contrastive loss to alleviate modality discrepancies. Finally, the model is validated on the challenging public TrafficVQA dataset SUTD-TrafficQA and outperforms the state-of-the-art method by a substantial margin (50.2%compared to46.0%). The code is available at https://github.com/guokan987/CFMMC-Align.
Kan Guo, Daxin Tian, Yongli Hu, Chunmian Lin, Jianshan Zhou, Xuting Duan, Junbin Gao
IEEE Trans. Circuits Syst. Video Technol.3
2024 Domain-Aware Prototype Network for Generalized Zero-Shot Learning
abstract
Generalized zero-shot learning(GZSL) aims to recognize images from seen and unseen classes with side information, such as manually annotated attribute vectors. Traditional methods focus on mapping images and semantics into a common latent space, thus achieving the visual-semantics alignment. Since the unseen classes are unavailable during training, there is a serious problem of recognition bias, which will tend to recognize unseen classes as seen classes. To solve this problem, we propose a Domain-aware Prototype Network(DPN), which splits the GZSL problem into the seen class recognition and unseen class recognition problem. For the seen classes, we design a domain-aware prototype learning branch with a dual attention feature encoder to capture the essential visual information, which aims to recognize the seen classes and discriminate the novel categories. To further recognize the fine-grained unseen classes, a visual-semantic embedding branch is designed, which aims to align the visual and semantic information for unseen-class recognition. Through the multi-task learning of the prototype learning branch and visual-semantic embedding branch, our model can achieve excellent performance on three popular GZSL datasets.
Yongli Hu, Lincong Feng, Huajie Jiang
IEEE Trans. Circuits Syst. Video Technol.1
2024 Hierarchical Multi-Modal Prompting Transformer for Multi-Modal Long Document Classification
abstract
In the context of long document classification (LDC), effectively utilizing multi-modal information encompassing texts and images within these documents has not received adequate attention. This task showcases several notable characteristics. Firstly, the text possesses an implicit or explicit hierarchical structure consisting of sections, sentences, and words. Secondly, the distribution of images is dispersed, encompassing various types such as highly relevant topic images and loosely related reference images. Lastly, intricate and diverse relationships exist between images and text at different levels. To address these challenges, we propose a novel approach called Hierarchical Multi-modal Prompting Transformer (HMPT). Our proposed method constructs the uni-modal and multi-modal transformers at both the section and sentence levels, facilitating effective interaction between features. Notably, we design an adaptive multi-scale multi-modal transformer tailored to capture the multi-granularity correlations between sentences and images. Additionally, we introduce three different types of shared prompts, i.e., shared section, sentence, and image prompts, as bridges connecting the isolated transformers, enabling seamless information interaction across different levels and modalities. To validate the model performance, we conducted experiments on two newly created and two publicly available multi-modal long document datasets. The obtained results show that our method outperforms state-of-the-art single-modality and multi-modality classification methods.
Tengfei Liu 0005, Yongli Hu, Junbin Gao
IEEE Trans. Circuits Syst. Video Technol.2
2024 Multi-Level Dynamic Graph Convolutional Networks for Weakly Supervised Crowd Counting
abstract
Crowd counting is very important in many fields such as public safety, urban planning, and is essential for the intelligent transportation systems. Due to the complexity and diversity of traffic scenes, point-level annotations for pedestrians would cost much human labor. Weakly supervised crowd counting methods are more suitable for these scenes, considering they only require count-level annotations. However, ignoring the uneven distribution of cross-distance crowd region density and multi-scale pedestrian head, existing weakly supervised methods can not achieve similar counting performance as fully supervised crowd counting methods. To solve these issues, we propose a novel multi-level dynamic graph convolutional networks for weakly supervised crowd counting. Within this network, a multi-level region dynamic graph convolutional module is designed to mine the cross-distance intrinsic relationship between crowd regions. A feature enhancement module is used to enhance crowd semantic information. In addition, we design a coarse grained multi-level feature fusion module to aggregate multi-scale pedestrian information. Experiments are conducted on five well-known benchmark crowd counting datasets, achieving state-of-the-art results compared to existing weakly supervised methods and competitive results compared to fully supervised methods.
Zhuangzhuang Miao, Yong Zhang 0029, Yongli Hu
IEEE Trans. Intell. Transp. Syst.4
2024 Self-Attention Graph Convolution Imputation Network for Spatio-Temporal Traffic Data
abstract
Missing data in time series is a pervasive problem that serves as obstacles for subsequent traffic data analysis. Consequently, extensive research works have been conducted on traffic missing data imputation tasks. The state-of-the-art traffic data imputation models are mostly based on recurrent neural networks. However, these methods belong to autoregressive models which are highly susceptible to error propagation. The attention-based methods are non-autoregressive models that can avoid compounding errors and help achieve better imputation quality. Moreover, the attention-based methods in now widely applied and have achieved remarkable results, whereas their application on traffic data imputation is still limited. Thus, this paper proposes Self-Attention Graph Convolution Imputation Network (SAGCIN) for spatio-temporal traffic data. To ensure the accuracy of data imputation, it is necessary to fully capture the spatio-temporal contextual information of traffic data to impute missing values. To this end, the SAGCIN model incorporates self-attention mechanism with diffusion graph convolution network. The SAGCIN model consists of two spatio-temporal blocks with a spatio-temporal encoder and an imputation decoder. The encoder learns spatio-temporal representations specialized for traffic data imputation tasks. Based on the learned representation, the decoder performs two stages of imputation operator for missing data. A joint-optimization training approach of imputation and reconstruction is introduced for SAGCIN to perform missing value imputation for traffic data. Empirical results demonstrate that SAGCIN model outperforms state-of-the-art methods in imputation tasks on relevant real-world benchmarks.
Xiulan Wei, Yong Zhang 0029, Shaofan Wang 0001, Xia Zhao 0003, Yongli Hu
IEEE Trans. Intell. Transp. Syst.5
2024 Cross-modal Multiple Granularity Interactive Fusion Network for Long Document Classification
abstract
Long Document Classification (LDC) has attracted great attention in Natural Language Processing and achieved considerable progress owing to the large-scale pre-trained language models. In spite of this, as a different problem from the traditional text classification, LDC is far from being settled. Long documents, such as news and articles, generally have more than thousands of words with complex structures. Moreover, compared with flat text, long documents usually contain multi-modal content of images, which provide rich information but not yet being utilized for classification. In this article, we propose a novel cross-modal method for long document classification, in which multiple granularity feature shifting networks are proposed to integrate the multi-scale text and visual features of long documents adaptively. Additionally, a multi-modal collaborative pooling block is proposed to eliminate redundant fine-grained text features and simultaneously reduce the computational complexity. To verify the effectiveness of the proposed model, we conduct experiments on the Food101 dataset and two constructed multi-modal long document datasets. The experimental results show that the proposed cross-modal method outperforms the single-modal text methods and defeats the state-of-the-art related multi-modal baselines.
Tengfei Liu 0005, Yongli Hu, Junbin Gao
ACM Trans. Knowl. Discov. Data2
2024 Incorporating Multi-Level Sampling with Adaptive Aggregation for Inductive Knowledge Graph Completion
abstract
In recent years, Graph Neural Networks (GNNs) have achieved unprecedented success in handling graph-structured data, thereby driving the development of numerous GNN-oriented techniques for inductive knowledge graph completion (KGC). A key limitation of existing methods, however, is their dependence on pre-defined aggregation functions, which lack the adaptability to diverse data, resulting in suboptimal performance on established benchmarks. Another challenge arises from the exponential increase in irrelated entities as the reasoning path lengthens, introducing unwarranted noise and consequently diminishing the model’s generalization capabilities. To surmount these obstacles, we design an innovative framework that synergizes M ulti- L evel S ampling with an A daptive A ggregation mechanism (MLSAA). Distinctively, our model couples GNNs with enhanced set transformers, enabling dynamic selection of the most appropriate aggregation function tailored to specific datasets and tasks. This adaptability significantly boosts both the model’s flexibility and its expressive capacity. Additionally, we unveil a unique sampling strategy designed to selectively filter irrelevant entities, while retaining potentially beneficial targets throughout the reasoning process. We undertake an exhaustive evaluation of our novel inductive KGC method across three pivotal benchmark datasets and the experimental results corroborate the efficacy of MLSAA.
Huajie Jiang, Yongli Hu
ACM Trans. Knowl. Discov. Data3
2024 Mixed-Modality Clustering via Generative Graph Structure Matching
abstract
The goal of mixed-modality clustering, which differs from typical multi-modality/view clustering, is to divide samples derived from various modalities into several clusters. This task has to solve two critical semantic gap problems: i) how to generate the missing modalities without the pairwise-modality data; and ii) how to align the representations of heterogeneous modalities. To tackle the above problems, this paper proposes a novel mixedmodality clustering model, which integrates the missing-modality generation and the heterogeneous modality alignment into a unified framework. During the missing-modality generation process, a bidirectional mapping is established between different modalities, enabling generation of preliminary representations for the missing-modality using information from another modality. Then the intra-modality bipartite graphs are constructed to help generate better missing-modality representations by weighted aggregating existing intra-modality neighbors. In this way, a pairwise-modality representation for each sample can be obtained. In the process of heterogeneous modality alignment, each modality is modelled as a graph to capture the global structure among intra-modality samples and is aligned against the heterogeneous modality representations through the adaptive heterogeneous graph matching module. Experimental results on three public datasets show the effectiveness of the proposed model compared to multiple state-of-the-art multi-modality/view clustering methods.
Xiaxia He, Boyue Wang, Junbin Gao, Qianqian Wang 0001, Yongli Hu
IEEE Trans. Knowl. Data Eng.5
2024 Parallelly Adaptive Graph Convolutional Clustering Model
abstract
Benefiting from exploiting the data topological structure, graph convolutional network (GCN) has made considerable improvements in processing clustering tasks. The performance of GCN significantly relies on the quality of the pretrained graph, while the graph structures are often corrupted by noise or outliers. To overcome this problem, we replace the pre-trained and fixed graph in GCN by the adaptive graph learned from the data. In this article, we propose a novel end-to-end parallelly adaptive graph convolutional clustering (AGCC) model with two pathway networks. In the first pathway, an adaptive graph convolutional (AGC) module alternatively updates the graph structure and the data representation layer by layer. The updated graph can better reflect the data relationship than the fixed graph. In the second pathway, the auto-encoder (AE) module aims to extract the latent data features. To effectively connect the AGC and AE modules, we creatively propose an attention-mechanism-based fusion (AMF) module to weight and fuse the data representations of the two modules, and transfer them to the AGC module. This simultaneously avoids the over-smoothing problem of GCN. Experimental results on six public datasets show that the effectiveness of the proposed AGCC compared with multiple state-of-the-art deep clustering methods. The code is available at https://github.com/HeXiax/AGCC.
Xiaxia He, Boyue Wang, Yongli Hu, Junbin Gao
IEEE Trans. Neural Networks Learn. Syst.3
2024 Global and Local Interactive Perception Network for Referring Image Segmentation
abstract
The effective modal fusion and perception between the language and the image are necessary for inferring the reference instance in the referring image segmentation (RIS) task. In this article, we propose a novel RIS network, the global and local interactive perception network (GLIPN), to enhance the quality of modal fusion between the language and the image from the local and global perspectives. The core of GLIPN is the global and local interactive perception (GLIP) scheme. Specifically, the GLIP scheme contains the local perception module (LPM) and the global perception module (GPM). The LPM is designed to enhance the local modal fusion by the correspondence between word and image local semantics. The GPM is designed to inject the global structured semantics of images into the modal fusion process, which can better guide the word embedding to perceive the whole image's global structure. Combined with the local-global context semantics fusion, extensive experiments on several benchmark datasets demonstrate the advantage of the proposed GLIPN over most state-of-the-art approaches.
Jing Liu 0059, Hongchen Tan, Yongli Hu, Huasheng Wang
IEEE Trans. Neural Networks Learn. Syst.3
2024 QDN: A Quadruplet Distributor Network for Temporal Knowledge Graph Completion
abstract
Temporal knowledge graph completion (TKGC) is an extension of the traditional static knowledge graph completion (SKGC) by introducing the timestamp. The existing TKGC methods generally translate the original quadruplet to the form of the triplet by integrating the timestamp into the entity/relation, and then use SKGC methods to infer the missing item. However, such an integrating operation largely limits the expressive ability of temporal information and ignores the semantic loss problem due to the fact that entities, relations, and timestamps are located in different spaces. In this article, we propose a novel TKGC method called the quadruplet distributor network (QDN), which independently models the embeddings of entities, relations, and timestamps in their specific spaces to fully capture the semantics and builds the QD to facilitate the information aggregation and distribution among them. Furthermore, the interaction among entities, relations, and timestamps is integrated using a novel quadruplet-specific decoder, which stretches the third-order tensor to the fourth-order to satisfy the TKGC criterion. Equally important, we design a novel temporal regularization that imposes a smoothness constraint on temporal embeddings. Experimental results show that the proposed method outperforms the existing state-of-the-art TKGC methods. The source codes of this article are available at https://github.com/QDN for Temporal Knowledge Graph Completion.git.
Jiapu Wang, Boyue Wang, Junbin Gao, Yongli Hu
IEEE Trans. Neural Networks Learn. Syst.5
2023 Center Focusing Network for Real-Time LiDAR Panoptic Segmentation
abstract
LiDAR panoptic segmentation facilitates an autonomous vehicle to comprehensively understand the surrounding objects and scenes and is required to run in real time. The recent proposal-free methods accelerate the algorithm, but their effectiveness and efficiency are still limited owing to the difficulty of modeling non-existent instance centers and the costly center-based clustering modules. To achieve accurate and real-time LiDAR panoptic segmentation, a novel center focusing network (CFNet) is introduced. Specifically, the center focusing feature encoding (CFFE) is proposed to explicitly understand the relationships between the original LiDAR points and virtual instance centers by shifting the LiDAR points and filling in the center points. Moreover, to leverage the redundantly detected centers, a fast center deduplication module (CDM) is proposed to select only one center for each instance. Experiments on the SemanticKITTI and nuScenes panoptic segmentation benchmarks demonstrate that our CFNet outperforms all existing methods by a large margin and is 1.6 times faster than the most efficient method.
Boyue Wang, Yongli Hu
CVPR4
2023 Zero-Shot Text Classification with Semantically Extended Textual Entailment
abstract
Zero-shot text classification (0SHOT-TC) aims to detect classes that the model never seen in the training set, and has attracted much attention in the research community of Natural Language Processing (NLP). The emergence of pre-trained language models has fostered the progress of 0SHOT-TC, which turns the task into a textual entailment problem of binary classification. It learns an entailment relatedness (yes/no) between the given sentence (premise) and each category (hypothesis) separately. However, the hypothesis generation paradigms need to be further studied, since the label itself or the label descriptions have limited ability to fully express the category space. Conversely, humans can easily extend a set of words describing the categories to be classified. In this paper, we propose a novel zero-shot text classification method called Semantically Extended Textual Entailment (SETE), which imitates the human's ability in knowledge extension. In the proposed method, three semantic extension methods are used to enrich the categories through a combination of static knowledge (e.g. expert knowledge, knowledge graph) and dynamic knowledge (e.g. language models), and the textual entailment model is finally used for 0SHOT-TC. The experimental results on the benchmarks show that our approach significantly outperforms the current methods in both generalized and non-generalized 0SHOT-TC.
Tengfei Liu 0005, Yongli Hu, Puman Chen
IJCNN2
2023 Breaking the Barrier Between Pre-training and Fine-tuning: A Hybrid Prompting Model for Knowledge-Based VQA
abstract
Considerable performance gains have been achieved for knowledge-based visual question answering due to the visual-language pre-training models with pre-training-then-fine-tuning paradigm. However, because the targets of the pre-training and fine-tuning stages are different, there is an evident barrier that prevents the cross-modal comprehension ability developed in the pre-training stage from fully endowing the fine-tuning task. To break this barrier, in this paper, we propose a novel hybrid prompting model for knowledge-based VQA, which inherits and incorporates the pre-training and fine-tuning tasks with a shared objective. Specifically, based on static declaration prompt, we construct a consistent goal with the fine-tuning via masked language modeling to inherit capabilities of pre-training task, while selecting the top-t relevant knowledge in a dense retrieval manner. Additionally, a dynamic knowledge prompt is learned from retrieved knowledge, which not only alleviates the length constraint on inputs for visual-language pre-trained models but also assists in providing answer features via fine-tuning. Combining and unifying the aims of the two stages could fully exploit the abilities of pre-training and fine-tuning to predict answer. We evaluate the proposed model on the OKVQA dataset, and the result shows that our model outperforms the state-of-the-art methods based on visual-language pre-training models with a noticeable performance gap and even exceeds the large-scale language model of GPT-3, which proves the benefits of the hybrid prompts and the advantages of unifying pre-training to fine-tuning.
Zhongfan Sun, Yongli Hu, Qingqing Gao, Huajie Jiang, Junbin Gao
ACM Multimedia2
2023 DSGEM: Dual scene graph enhancement module-based visual question answering
abstract
Abstract Visual Question Answering (VQA) aims to appropriately answer a text question by understanding the image content. Attention‐based VQA models mine the implicit relationships between objects according to the feature similarity, which neglects the explicit relationships between objects, for example, the relative position. Most Visual Scene Graph‐based VQA models exploit the relative positions or visual relationships between objects to construct the visual scene graph, while they suffer from the semantic insufficiency of visual edge relations. Besides, the scene graph of text modality is often ignored in these works. In this article, a novel Dual Scene Graph Enhancement Module (DSGEM) is proposed that exploits the relevant external knowledge to simultaneously construct two interpretable scene graph structures of image and text modalities, which makes the reasoning process more logical and precise. Specifically, the authors respectively build the visual and textual scene graphs with the help of commonsense knowledge and syntactic structure, which explicitly endows the specific semantics to each edge relation. Then, two scene graph enhancement modules are proposed to propagate the involved external and structural knowledge to explicitly guide the feature interaction between objects (nodes). Finally, the authors embed such two scene graph enhancement modules to existing VQA models to introduce the explicit relation reasoning ability. Experimental results on both VQA V2 and OK‐VQA datasets show that the proposed DSGEM is effective and compatible to various VQA architectures.
Boyue Wang, Yujian Ma, Yongli Hu
IET Comput. Vis.5
2023 A multi-scale feature representation and interaction network for underwater object detection
abstract
Abstract Compared with natural images, underwater images are usually degraded with blur, scale variation, colour shift and texture distortion, which bring much challenge for computer vision tasks like object detection. In this case, generic object detection methods usually fail to achieve satisfactory performance. The main reason is considered that the current methods lack sufficient discriminativeness of feature representation for the degraded underwater images. A a novel multi‐scale feature representation and interaction network for underwater object detection is proposed, in which two core modules are elaborately designed to enhance the discriminativeness of feature representation for underwater images. The first is the Context Integration Module, which extracts rich context information from high‐level features and is integrated with the feature pyramid network to enhance the feature representation in a multi‐scale way. The second is the Dual‐refined Attention Interaction Module, which further enhances the feature representation by sufficient interactions between different levels of features both in channel and spatial domains based on attention mechanism. The proposed model is evaluated on four public underwater datasets. The experimental results compared with state‐of‐the‐art object detection methods show that the proposed model has leading performance, which verifies that it is effective for underwater object detection. In addition, object detection experiments on a foggy dataset of Real‐world Task‐driven Testing Set (RTTS) and the natural image dataset of pattern analysis statistical modelling and computational learning, visual object classes (PASCAL VOC) are conducted. The results show that the proposed model can be applied on the degraded dataset of RTTS but fails on PASCAL VOC.
Jiaojiao Yuan, Yongli Hu
IET Comput. Vis.2
2023 A subgraph sampling method for training large-scale graph convolutional network
Qi Zhang 0095, Yongli Hu, Shaofan Wang 0001
Inf. Sci.3
2023 Graph structure learning layer and its graph convolution clustering application
Xiaxia He, Boyue Wang, Ruikun Li 0001, Junbin Gao, Yongli Hu, Guangyu Huo
Neural Networks5
2023 Logarithmic Schatten-$p$p Norm Minimization for Tensorial Multi-View Subspace Clustering
abstract
The low-rank tensor could characterize inner structure and explore high-order correlation among multi-view representations, which has been widely used in multi-view clustering. Existing approaches adopt the tensor nuclear norm (TNN) as a convex approximation of non-convex tensor rank function. However, TNN treats the different singular values equally and over-penalizes the main rank components, leading to sub-optimal tensor representation. In this paper, we devise a better surrogate of tensor rank, namely the tensor logarithmic Schatten- p norm ([Formula: see text]N), which fully considers the physical difference between singular values by the non-convex and non-linear penalty function. Further, a tensor logarithmic Schatten- p norm minimization ([Formula: see text]NM)-based multi-view subspace clustering ([Formula: see text]NM-MSC) model is proposed. Specially, the proposed [Formula: see text]NM can not only protect the larger singular values encoded with useful structural information, but also remove the smaller ones encoded with redundant information. Thus, the learned tensor representation with compact low-rank structure will well explore the complementary information and accurately characterize the high-order correlation among multi-views. The alternating direction method of multipliers (ADMM) is used to solve the non-convex multi-block [Formula: see text]NM-MSC model where the challenging [Formula: see text]NM problem is carefully handled. Importantly, the algorithm convergence analysis is mathematically established by showing that the sequence generated by the algorithm is of Cauchy and converges to a Karush-Kuhn-Tucker (KKT) point. Experimental results on nine benchmark databases reveal the superiority of the [Formula: see text]NM-MSC model.
Jipeng Guo 0001, Junbin Gao, Yongli Hu
IEEE Trans. Pattern Anal. Mach. Intell.4
2023 Multi-level attention for referring expression comprehension
Yunru Zhang, Huajie Jiang, Yongli Hu
Pattern Recognit. Lett.4
2023 Self-Supervised Nodes-Hyperedges Embedding for Heterogeneous Information Network Learning
abstract
The exploration of self-supervised information mining of heterogeneous datasets has gained significant traction in recent years. Heterogeneous graph neural networks (HGNNs) have emerged as a highly promising method for handling heterogeneous information networks (HINs) due to their superior performance. These networks leverage aggregation functions to convert pairwise relations-based features from raw heterogeneous graphs into embedding vectors. However, real-world HINs contain valuable higher-order relations that are often overlooked but can provide complementary information. To address this issue, we propose a novel method calledSelf-supervisedNodes-HyperedgesEmbedding (SNHE), which leverages hypergraph structures to incorporate higher-order information into the embedding process of HINs. Our method decomposes the raw graph structure into snapshots based on various meta-paths, which are then transformed into hypergraphs to aggregate high-order information within the data and generate embedding representations. Given the complexity of HINs, we develop a dual self-supervised structure that maximizes mutual information in the enhanced graph data space, guides the overall model update, and reduces redundancy and noise. We evaluate our proposed method on various real-world datasets for node classification and clustering tasks, and compare it against state-of-the-art methods. The experimental results demonstrate the efficacy of our method. Our code is available athttps://github.com/limengran98/SNHE.
Mengran Li 0001, Yong Zhang 0029, Wei Zhang 0320, Yi Chu, Yongli Hu
IEEE Trans. Big Data5
2023 STGAN: Spatio-Temporal Generative Adversarial Network for Traffic Data Imputation
abstract
The traffic data corrupted by noise and missing entries often lead to the poor performance of Intelligent Transportation Systems (ITS), such as the bad congestion prediction and route guidance. How to efficiently impute the traffic data is an urgent problem. As a classic deep learning method, Generative Adversarial Network (GAN) achieves remarkable success in image recovery fields, which opens up a new way for the traffic data imputation. In this paper, we propose a novel spatio-temporal GAN model for the traffic data imputation (STGAN). Firstly, we design the generative loss and center loss, which not only minimizes the reconstructed errors of the imputed entries, but also ensures each imputed entry and its neighbors conform to the local spatio-temporal distribution. Then, the discriminator uses the convolution neural network classifier to judge whether the imputed matrix conforms to the global spatio-temporal distribution. As for the network architecture of the generator, we introduce the skip-connection to keep all well preserved data unchanged, and employ the dilated convolution to capture the spatio-temporal correlation in the traffic data. The experimental results show that our proposed method obviously outperforms other competitive traffic data imputation methods.
Yong Zhang 0029, Boyue Wang, Yongli Hu
IEEE Trans. Big Data5
2023 Self-Attention Graph Convolution Residual Network for Traffic Data Completion
abstract
Complete and accurate traffic data is critical in urban traffic management, planning and operation. In fact, real-world traffic data contains missing values due to multiple factors, such as device outages and communication errors. For traffic data completion task, most of the existing methods are matrix/tensor completion methods, which usually enforce low rank constraint on traffic data matrix/tensor. But they neglect the graph structure of traffic data, resulting in low completion performance. Recently, graph convolutional networks have achieved remarkable results in traffic data forecasting due to their abilities of feature extraction and nonlinear fitting on arbitrarily graph-structured data. However, there are few studies based on graph neural networks for traffic data completion task. In this paper, we propose a traffic data completion model based on graph convolutional network model to impute missing values from the perspective of deep learning. This model utilizes graph convolution to model the local spatial dependency. As for global spatial dependency and temporal dependency, this model incorporates self-attention mechanism, which is applied in the spatial and temporal dimensions respectively. The experimental results on the two real-time datasets demonstrate that the proposed model outperforms the baseline methods significantly under arbitrarily missing scenarios.
Yong Zhang 0029, Xiulan Wei, Yongli Hu
IEEE Trans. Big Data4
2023 Hierarchical Spatio-Temporal Graph Convolutional Networks and Transformer Network for Traffic Flow Forecasting
abstract
Graph convolutional networks (GCN) have been applied in the traffic flow forecasting tasks with the graph capability in describing the irregular topology structures of road networks. However, GCN based traffic flow forecasting methods often fail to simultaneously capture the short-term and long-term temporal relations carried by the traffic flow data, and also suffer the over-smoothing problem. To overcome the problems, we propose a hierarchical traffic flow forecasting network by merging newly designed the long-term temporal Transformer network (LTT) and the spatio-temporal graph convolutional networks (STGC). Specifically, LTT aims to learn the long-term temporal relations among the traffic flow data, while the STGC module aims to capture the short-term temporal relations and spatial relations among the traffic flow data, respectively, via cascading between the one-dimensional convolution and the graph convolution. In addition, an attention fusion mechanism is proposed to combine the long-term with the short-term temporal relations as the input of the graph convolution layer in STGC, in order to mitigate the over-smoothing problem of GCN. Experimental results on three public traffic flow datasets prove the effectiveness and robustness of the proposed method.
Guangyu Huo, Yong Zhang 0029, Boyue Wang, Junbin Gao, Yongli Hu
IEEE Trans. Intell. Transp. Syst.5
2023 Substructure-aware subgraph reasoning for inductive relation prediction
Huajie Jiang, Yongli Hu
J. Supercomput.3
2023 Multi-Concept Representation Learning for Knowledge Graph Completion
abstract
Knowledge Graph Completion (KGC) aims at inferring missing entities or relations by embedding them in a low-dimensional space. However, most existing KGC methods generally fail to handle the complex concepts hidden in triplets, so the learned embeddings of entities or relations may deviate from the true situation. In this article, we propose a novel M ulti- c oncept R epresentation L earning (McRL) method for the KGC task, which mainly consists of a multi-concept representation module, a deep residual attention module, and an interaction embedding module. Specifically, instead of the single-feature representation, the multi-concept representation module projects each entity or relation to multiple vectors to capture the complex conceptual information hidden in them. The deep residual attention module simultaneously explores the inter- and intra-connection between entities and relations to enhance the entity and relation embeddings corresponding to the current contextual situation. Moreover, the interaction embedding module further weakens the noise and ambiguity to obtain the optimal and robust embeddings. We conduct the link prediction experiment to evaluate the proposed method on several standard datasets, and experimental results show that the proposed method outperforms existing state-of-the-art KGC methods.
Jiapu Wang, Boyue Wang, Junbin Gao, Yongli Hu
ACM Trans. Knowl. Discov. Data4
2023 CaEGCN: Cross-Attention Fusion Based Enhanced Graph Convolutional Network for Clustering
abstract
With the powerful learning ability of deep convolutional networks, deep clustering methods can extract the most discriminative information from individual data and produce more satisfactory clustering results. However, existing deep clustering methods usually ignore the relationship between the data. Fortunately, the graph convolutional network can handle such relationships, opening a new research direction for deep clustering. In this paper, we propose a cross-attention based deep clustering framework, named Cross-Attention Fusion based Enhanced Graph Convolutional Network (CaEGCN), which contains four main modules: the cross-attention fusion module which innovatively concatenates the Content Auto-encoder module (CAE) relating to the individual data and Graph Convolutional Auto-encoder module (GAE) relating to the relationship between the data in a layer-by-layer manner, and the self-supervised model that highlights the discriminative information for clustering tasks. While the cross-attention fusion module fuses two kinds of heterogeneous representation, the CAE module supplements the content information for the GAE module, which avoids the over-smoothing problem of GCN. In the GAE module, two novel loss functions are proposed that reconstruct the content and relationship between the data, respectively. Finally, the self-supervised module constrains the distributions of the middle layer representations of CAE and GAE to be consistent. Experimental results on different types of datasets prove the superiority and robustness of the proposed CaEGCN.
Guangyu Huo, Yong Zhang 0029, Junbin Gao, Boyue Wang, Yongli Hu
IEEE Trans. Knowl. Data Eng.5
2023 TDN: Triplet Distributor Network for Knowledge Graph Completion
abstract
Conventional Knowledge Graph Completion (KGC) methods typically map entities and relations to a unified space through the shared mapping matrix, and then interact with entities and relations to infer the missing items in the knowledge graph. Although this shared mapping matrix considers the suitability of all triplets, it neglects the specificity of each triplet. To solve this problem, we dynamically learn one information distributor for each triplet to exchange its specific information. In this paper, we propose a novel Triplet Distributor Network (TDN) for the knowledge graph completion task. Specifically, we adaptively learn one Triplet Distributor (TD) for each triplet to assist the interaction between the entity and relation. Furthermore, on the basis of TD, we creatively design the information exchange layer to dynamically propagate the information of the entity and relation, thus mutually enhancing entity and relation representations. Except for several commonly-used knowledge graph datasets, we still implement the link prediction task on the social-relational and medical datasets to test the proposed method. Experimental results demonstrate that the proposed method performs better than existing state-of-the-art KGC methods. The source codes of this paper are available athttps://github.com/TDNfor Knowledge Graph Completion.git.
Jiapu Wang, Boyue Wang, Junbin Gao, Yongli Hu
IEEE Trans. Knowl. Data Eng.5
2023 Hierarchical Graph Convolutional Networks for Structured Long Document Classification
abstract
Long document classification (LDC) has been a focused interest in natural language processing (NLP) recently with the exponential increase of publications. Based on the pretrained language models, many LDC methods have been proposed and achieved considerable progression. However, most of the existing methods model long documents as sequences of text while omitting the document structure, thus limiting the capability of effectively representing long texts carrying structure information. To mitigate such limitation, we propose a novel hierarchical graph convolutional network (HGCN) for structured LDC in this article, in which a section graph network is proposed to model the macrostructure of a document and a word graph network with a decoupled graph convolutional block is designed to extract the fine-grained features of a document. In addition, an interaction strategy is proposed to integrate these two networks as a whole by propagating features between them. To verify the effectiveness of the proposed model, four structured long document datasets are constructed, and the extensive experiments conducted on these datasets and another unstructured dataset show that the proposed method outperforms the state-of-the-art related classification methods.
Tengfei Liu 0005, Yongli Hu, Boyue Wang, Junbin Gao
IEEE Trans. Neural Networks Learn. Syst.2
2023 Domain Adaptation as Optimal Transport on Grassmann Manifolds
abstract
Domain adaptation in the Euclidean space is a challenging task on which researchers recently have made great progress. However, in practice, there are rich data representations that are not Euclidean. For example, many high-dimensional data in computer vision are in general modeled by a low-dimensional manifold. This prompts the demand of exploring domain adaptation between non-Euclidean manifold spaces. This article is concerned with domain adaption over the classic Grassmann manifolds. An optimal transport-based domain adaptation model on Grassmann manifolds has been proposed. The model implements the adaption between datasets by minimizing the Wasserstein distances between the projected source data and the target data on Grassmann manifolds. Four regularization terms are introduced to keep task-related consistency in the adaptation process. Furthermore, to reduce the computational cost, a simplified model preserving the necessary adaption property and its efficient algorithm is proposed and tested. The experiments on several publicly available datasets prove the proposed model outperforms several relevant baseline domain adaptation methods.
Tianhang Long, Junbin Gao, Yongli Hu
IEEE Trans. Neural Networks Learn. Syst.4
2023 Metro Passenger-Flow Representation via Dynamic Mode Decomposition and Its Application
abstract
Passenger-flow anomaly detection and prediction are essential tasks for intelligent operation of the metro system. Accurate passenger-flow representation is the foundation of them. However, spatiotemporal dependencies, complex dynamic changes, and anomalies of passenger-flow data bring great challenges to data representation. Taking advantage of the time-varying characteristics of data, we propose a novel passenger-flow representation model based on low-rank dynamic mode decomposition (DMD), which also integrates the global low-rank nature and sparsity to explore the spatiotemporal consistency of data and depict abrupt data, respectively. The model can detect anomalies and predict short-term passenger flow conveniently and flexibly. For anomaly detection, we further introduce a strong temporal Toeplitz regularization to characterize the temporal periodic change of data, so as to more accurately detect anomalies. We conduct experiments with smart card transaction data from the Beijing metro system to assess the performance of the model in two use cases. In terms of anomaly detection, the experimental results demonstrate that our method can detect anomalies efficiently, especially for time sequence anomalies. As for short-term prediction, our model is superior to other methods in most cases.
Xiulan Wei, Yong Zhang 0029, Yongli Hu, Shuzhen Tong, Wei Huang 0017, Jinde Cao
IEEE Trans. Neural Networks Learn. Syst.4
2022 SHCN: Self-supervised General Hypergraph Clustering Network
abstract
Clustering is a fundamental and hot issue in the unsupervised learning area. With the rapid development of deep learning and graph neural networks (GNNs) techniques, researchers have proposed a series of effective clustering methods. However, most existing approaches adopt a conventional graph to aggregate the neighborhood information, where only the pairwise relations are considered. Moreover, the redundancy/noise in the raw data samples may result in less accurate sample relations and inferior clustering results. In this paper, we proposed a new GNNs based clustering method, which adopts the hypergraph learning approach to explore the high-order relationship for accurate relation learning. Specifically, we first construct two hypergraph representations based on the topology feature and attribute feature from data samples. Then, a self-supervised structure is integrated to learn a cross-correlation matrix from the original hypergraph to act as a higher-order neighborhood with reduced redundancy and noise. Finally the embedding representation of the clustering space is learned in the graph convolution. The proposed method has been evaluated on six public datasets for clustering tasks. Experimental results show that our proposed method outperforms the state-of-the-art ones.
Mengran Li 0001, Xinglin Piao, Yong Zhang 0029, Yongli Hu
IEEE Big Data4
2022 Hierarchical Multiple Granularity Attention Network for Long Document Classification
abstract
Long document classification has aroused tremendous attention in the field of Nature Language Processing, due to the exponential increasing of publications. Although the common text classification methods can be extended for long document classification, they are confined to the length of text and do not have enough expressiveness to model the structure of the long document. To solve these problems, we proposed a hierarchical multiple granularity attention network for long document classification, in which the word and section level features are extracted and fused to represent the complex structure of the long document. Furthermore, a feature-based section pooling module is adopted to eliminate redundant text information and accelerate the computing. A series of experiments are conducted to evaluate the proposed method. The experimental results verify that our method is effective, efficient and competitive compared with the related state-of-the-art methods.
Yongli Hu, Tengfei Liu 0005, Junbin Gao
IJCNN1
2022 Globality constrained adaptive graph regularized non-negative matrix factorization for data representation
abstract
Abstract Benefiting from the good physical interpretations and low computational complexity, non‐negative matrix factorization (NMF) has attracted wide attentions in data representation learning tasks. Some graph‐based NMF approaches make the learned representation encode the topological structure by the local graph Laplacian regularizer, which improves the discriminant ability of data representation. However, the performance of graph‐based NMF methods depend heavily on the quality of the predefined graph and the complexity of models is high. Here, a globality constrained adaptive graph regularized non‐negative matrix factorization for data representation (GCAG‐NMF) model is proposed, which not only uses the self‐representation characteristics of data to learn an adaptive graph to describe the sample relationship more accurately, but also proposes a graph factorization technique to reduce the complexity of the model and improve the discriminative ability of data representation. Then, an iterative optimizing strategy with low complexity and strict convergence guarantee is developed to optimize the objective function. Experimental results on some databases demonstrate the effectiveness of the proposed model.
Jie Wang 0134, Jipeng Guo 0001, Yongli Hu
IET Image Process.4
2022 Multi-graph convolutional clustering network
abstract
Abstract The relationship between objects can be described from different angles. Although multiple kinds of relationships make the connections between objects complex, they bring in more discriminative information for the clustering tasks. Therefore, how to effectively fuse multiple kinds of relationships becomes a critical problem. In this paper, we propose a novel Multi‐graph Convolutional Clustering Network which deeply explores the feature information of nodes and fuses the multiple kinds of relationships between nodes. Unlike most graph convolutional clustering methods that only exploit the single graph or directly fuse multiple graphs into a unified graph before the graph convolution operation, we firstly build multiple parallelled graph convolution layers for each graph to learn diverse data representations, which fully exploits different statistics information between graphs. Then, a designed multi‐graph attention module fuses above data representations and considers the importance of each graph. Besides, the proposed model completes the transition from single graph to multiple graphs, which reduces the dependence of the quality of the single graph and enhances the robustness to graphs. Experimental results verify that the proposed multi‐graph convolution clustering performs better than the traditional single‐graph convolution clustering.
Boyue Wang, Yifan Wang 0004, Xiaxia He, Yongli Hu
IET Signal Process.4
2022 Video Domain Adaptation based on Optimal Transport in Grassmann Manifolds
Tianhang Long, Junbin Gao, Yongli Hu
Inf. Sci.4
2022 Shareability-Exclusivity Representation on Product Grassmann Manifolds for Multi-camera video clustering
Yongli Hu, Cuicui Luo, Junbin Gao, Boyue Wang
J. Vis. Commun. Image Represent.1
2022 Cross-modal alignment with graph reasoning for image-text retrieval
Yongli Hu, Junbin Gao
Multim. Tools Appl.2
2022 CGNN: Caption-assisted graph neural network for image-text retrieval
Yongli Hu, Hanfu Zhang, Huajie Jiang, Yandong Bi
Pattern Recognit. Lett.1
2022 MFFNet: Single facial depth map refinement using multi-level feature fusion
Fan Zhang 0062, Yongli Hu, Fuqing Duan
Signal Process. Image Commun.3
2022 Probabilistic Linear Discriminant Analysis Based on L1-Norm and Its Bayesian Variational Inference
abstract
Probabilistic linear discriminant analysis (PLDA) is a very effective feature extraction approach and has obtained extensive and successful applications in supervised learning tasks. It employs the squared$L_{2}$-norm to measure the model errors, which assumes a Gaussian noise distribution implicitly. However, the noise in real-life applications may not follow a Gaussian distribution. Particularly, the squared$L_{2}$-norm could extremely exaggerate data outliers. To address this issue, this article proposes a robust PLDA model under the assumption of a Laplacian noise distribution, called L1-PLDA. The learning process employs the approach by expressing the Laplacian density function as a superposition of an infinite number of Gaussian distributions via introducing a new latent variable and then adopts the variational expectation–maximization (EM) algorithm to learn parameters. The most significant advantage of the new model is that the introduced latent variable can be used to detect data outliers. The experiments on several public databases show the superiority of the proposed L1-PLDA model in terms of classification and outlier detection.
Xiangjie Hu, Junbin Gao, Yongli Hu, Fujiao Ju
IEEE Trans. Cybern.4
2022 Multi-Attribute Subspace Clustering via Auto-Weighted Tensor Nuclear Norm Minimization
abstract
Self-expressiveness based subspace clustering methods have received wide attention for unsupervised learning tasks. However, most existing subspace clustering methods consider data features as a whole and then focus only on one single self-representation. These approaches ignore the intrinsic multi-attribute information embedded in the original data feature and result in one-attribute self-representation. This paper proposes a novel multi-attribute subspace clustering (MASC) model that understands data from multiple attributes. MASC simultaneously learns multiple subspace representations corresponding to each specific attribute by exploiting the intrinsic multi-attribute features drawn from original data. In order to better capture the high-order correlation among multi-attribute representations, we represent them as a tensor in low-rank structure and propose the auto-weighted tensor nuclear norm (AWTNN) as a superior low-rank tensor approximation. Especially, the non-convex AWTNN fully considers the difference between singular values through the implicit and adaptive weights splitting during the AWTNN optimization procedure. We further develop an efficient algorithm to optimize the non-convex and multi-block MASC model and establish the convergence guarantees. A more comprehensive subspace representation can be obtained via aggregating these multi-attribute representations, which can be used to construct a clustering-friendly affinity matrix. Extensive experiments on eight real-world databases reveal that the proposed MASC exhibits superior performance over other subspace clustering methods.
Jipeng Guo 0001, Junbin Gao, Yongli Hu
IEEE Trans. Image Process.4
2022 Dynamic Graph Convolution Network for Traffic Forecasting Based on Latent Network of Laplace Matrix Estimation
abstract
Traffic forecasting is a challenging problem in the transportation research field as the complexity and non-stationary changing of the traffic data, thus the key to the issue is how to explore proper spatial and temporal characteristics. Based on this thought, many creative methods have been proposed, in which Graph Convolution Network (GCN) based methods have shown promising performance. However, these methods depend on the graph construction, which mainly uses the prior knowledge of the road network. Recently, some works realized the fact of the road network graph changing and tried to construct dynamic graphs for GCN, but they do not fully exploit the spatial and temporal properties of the traffic data in the graph construction. In this paper, we propose a novel dynamic graph convolution network for traffic forecasting, in which a latent network is introduced to extract spatial-temporal features for constructing the dynamic road network graph matrices adaptively. The proposed method is evaluated on several traffic datasets and the experimental results show that it outperforms the state of the art traffic forecasting methods. The website of the code ishttps://github.com/guokan987/DGCN.git.
Kan Guo, Yongli Hu, Sean Qian, Junbin Gao
IEEE Trans. Intell. Transp. Syst.2
2022 Text-to-Traffic Generative Adversarial Network for Traffic Situation Generation
abstract
Traffic situation generation is of importance in the intelligent transportation field, evaluating and simulating the macroscopic traffic conditions. The government often uses the historical traffic data on the same weekday to analyze the future traffic situations, which works unfavorably due to some traffic-related information deficiency, such as weather, location, traffic accidents, social activities and so on. Therefore, how to accurately generate the traffic situation is a challenging problem. Fortunately, massive traffic-related information spread in social media often indicates the traffic situation variation trend, which provides the sufficient information for the traffic situation generation. In this paper, we propose a novel Text-to-Traffic generative adversarial network framework ($\text{T}^{2}$GAN), which fuses the traffic data and the semantic information collected from social media to generate the traffic situation. To reduce the huge gap between the above two modalities and improve the authenticity of the generated traffic situation, we raise a global-local loss. Additionally, we build a heterogeneous dataset containing the traffic-related text data collected from social media and the corresponding traffic passenger flow data. Experimental results show that the proposed methods are obviously better than many outstanding traffic situation generation methods based on neural networks.
Guangyu Huo, Yong Zhang 0029, Boyue Wang, Yongli Hu
IEEE Trans. Intell. Transp. Syst.4
2022 Dual Dynamic Spatial-Temporal Graph Convolution Network for Traffic Prediction
abstract
Recently, Graph Convolution Network (GCN) and Temporal Convolution Network (TCN) are introduced into traffic prediction and achieve state-of-the-art performance due to their good ability for modeling the spatial and temporal property of traffic data. In spite of having good performance, the current methods generally focus on the traffic measurement of road segments, i.e. the nodes of traffic flow graph, while the edges of the graph, which represent the correlation of traffic data of different road segments and form the affinity matrix for GCN, are usually constructed according to the structure of road network, but the spatial and temporal properties are not well exploited in their theories. In this paper, we propose a Dual Dynamic Spatial-Temporal Graph Convolution Network (DDSTGCN), which not only models the dynamic property of the nodes of the traffic flow graph but also captures the dynamic spatial-temporal feature of the edges by transforming the traffic flow graph into its dual hypergraph. The traffic prediction is enhanced by the collaborative convolutions on the traffic flow graph and its dual hypergraph. The proposed method is evaluated by extensive traffic prediction experiments on six real road datasets and the results show that it outperforms state-of-the-art related methods. Source codes are available athttps://github.com/j1o2h3n/DDSTGCN.
Xiangheng Jiang, Yongli Hu, Fuqing Duan, Kan Guo, Boyue Wang, Junbin Gao
IEEE Trans. Intell. Transp. Syst.3
2022 Multitask Hypergraph Convolutional Networks: A Heterogeneous Traffic Prediction Framework
abstract
Traffic prediction methods on a single-source data have achieved excellent results in recent years, especially the Graph Convolutional Networks (GCN) based models with spatio-temporal dependency. In reality, various modes of urban transportation operate simultaneously. They influence and complement each other in common space-time occasions, constituting the transportation system dynamically. Thus, traffic data from multiple sources is ostensibly heterogeneous, but internally correlated. The typical single data driven models are, however, not universally applicable for heterogeneous traffic data. To address this issue, we propose a Multi-task Hypergraph Convolutional Neural Network (MT-HGCN) for the multi-source traffic prediction problem. The framework consists of a main task and a related task. Both tasks are based on Hypergraph Convolutional Neural Networks (HGCN) and are devoted to two prediction problems. Furthermore, the tasks are bridged by a feature compress unit, which models the correlation and shares the latent feature to improve the performance of the main task. The node-level forecasting has been evaluated on historical datasets of Beijing to verify the effectiveness of the proposed method. Compared with the state-of-the-arts, the superior performance of the proposed method can be obtained.
Yong Zhang 0029, Lixun Wang, Yongli Hu, Xinglin Piao
IEEE Trans. Intell. Transp. Syst.4
2022 Urban Traffic Pattern Analysis and Applications Based on Spatio-Temporal Non-Negative Matrix Factorization
abstract
Analyzing the traffic state of large citywide networks is an inherently difficult task. Various data issues, traffic signals, stops signs and other flow inhibitors of the network-level traffic state make the analysis more difficult than that under the small-scale local traffic state. To address this challenge, we propose a method based on spatio-temporal non-negative matrix factorization (ST-NMF), which is used for road network traffic pattern analysis. The method can be further extended to traffic data reconstruction and traffic prediction. In order to analyze traffic patterns, the proposed spatio-temporal non-negative matrix factorization model represents the network traffic as a linear combination of several basic patterns, which is also interpreted as the dynamics of spatial traffic characteristics over time in low-dimensional space. By the visual display of the spatial and temporal patterns and the assistance of clustering methods, the traffic pattern features are extracted. In the extended applications, data reconstruction relies on the sampling representation of missing data by ST-NMF, and data prediction is based on the prediction of the temporal patterns by ST-NMF. Through our method, we can not only obtain a high-quality data foundation, but also explore typical spatio-temporal patterns and general predictions of the future traffic state. The analysis results have important guiding significance on the management of intelligent transportation systems. Experiments on real-world traffic data are provided to verify the validity of our proposed approach.
Yang Wang 0068, Yong Zhang 0029, Lixun Wang, Yongli Hu
IEEE Trans. Intell. Transp. Syst.4
2022 Rank Consistency Induced Multiview Subspace Clustering via Low-Rank Matrix Factorization
abstract
Multiview subspace clustering has been demonstrated to achieve excellent performance in practice by exploiting multiview complementary information. One of the strategies used in most existing methods is to learn a shared self-expressiveness coefficient matrix for all the view data. Different from such a strategy, this article proposes a rank consistency induced multiview subspace clustering model to pursue a consistent low-rank structure among view-specific self-expressiveness coefficient matrices. To facilitate a practical model, we parameterize the low-rank structure on all self-expressiveness coefficient matrices through the tri-factorization along with orthogonal constraints. This specification ensures that self-expressiveness coefficient matrices of different views have the same rank to effectively promote structural consistency across multiviews. Such a model can learn a consistent subspace structure and fully exploit the complementary information from the view-specific self-expressiveness coefficient matrices, simultaneously. The proposed model is formulated as a nonconvex optimization problem. An efficient optimization algorithm with guaranteed convergence under mild conditions is proposed. Extensive experiments on several benchmark databases demonstrate the advantage of the proposed model over the state-of-the-art multiview clustering approaches.
Jipeng Guo 0001, Junbin Gao, Yongli Hu
IEEE Trans. Neural Networks Learn. Syst.4
2022 A Decoder-Free Variational Deep Embedding for Unsupervised Clustering
abstract
In deep clustering frameworks, autoencoder (AE)- or variational AE-based clustering approaches are the most popular and competitive ones that encourage the model to obtain suitable representations and avoid the tendency for degenerate solutions simultaneously. However, for the clustering task, the decoder for reconstructing the original input is usually useless when the model is finished training. The encoder-decoder architecture limits the depth of the encoder so that the learning capacity is reduced severely. In this article, we propose a decoder-free variational deep embedding for unsupervised clustering (DFVC). It is well known that minimizing reconstruction error amounts to maximizing a lower bound on the mutual information (MI) between the input and its representation. That provides a theoretical guarantee for us to discard the bloated decoder. Inspired by contrastive self-supervised learning, we can directly calculate or estimate the MI of the continuous variables. Specifically, we investigate unsupervised representation learning by simultaneously considering the MI estimation of continuous representations and the MI computation of categorical representations. By introducing the data augmentation technique, we incorporate the original input, the augmented input, and their high-level representations into the MI estimation framework to learn more discriminative representations. Instead of matching to a simple standard normal distribution adversarially, we use end-to-end learning to constrain the latent space to be cluster-friendly by applying the Gaussian mixture distribution as the prior. Extensive experiments on challenging data sets show that our model achieves higher performance over a wide range of state-of-the-art clustering approaches.
Junbin Gao, Yongli Hu
IEEE Trans. Neural Networks Learn. Syst.4
2021 Hierarchical Graph Convolution Network for Traffic Forecasting
abstract
Traffic forecasting is attracting considerable interest due to its widespread application in intelligent transportation systems. Given the complex and dynamic traffic data, many methods focus on how to establish a spatial-temporal model to express the non-stationary traffic patterns. Recently, the latest Graph Convolution Network (GCN) has been introduced to learn spatial features while the time neural networks are used to learn temporal features. These GCN based methods obtain state-of-the-art performance. However, the current GCN based methods ignore the natural hierarchical structure of traffic systems which is composed of the micro layers of road networks and the macro layers of region networks, in which the nodes are obtained through pooling method and could include some hot traffic regions such as downtown and CBD etc., while the current GCN is only applied on the micro graph of road networks. In this paper, we propose a novel Hierarchical Graph Convolution Networks (HGCN) for traffic forecasting by operating on both the micro and macro traffic graphs. The proposed method is evaluated on two complex city traffic speed datasets. Compared to the latest GCN based methods like Graph WaveNet, the proposed HGCN gets higher traffic forecasting precision with lower computational cost.The website of the code is https://github.com/guokan987/HGCN.git.
Kan Guo, Yongli Hu, Sean Qian, Junbin Gao
AAAI2
2021 Feature Interaction Based Graph Convolutional Networks for Image-Text Retrieval
Yongli Hu, Feili Gao, Junbin Gao
ICANN (3)1
2021 Bottom-Up Progressive Semantic Alignment for Image-Text Retrieval
Yongli Hu, Junbin Gao
ICONIP (6)2
2021 Hierarchical Attention Transformer Networks for Long Document Classification
abstract
Profiting from the pre-trained language representation models like BERT, the recently proposed document classification methods have obtained considerable improvement. However, most of these methods usually model the document as a sequence of text and omit the structure information, which appears obviously in long document composed of several sections with assigned relations. For this purpose, we propose a novel Hierarchical Attention Transformer Network (HATN) for long document classification, which extracts the structure of the long document by intra- and inter-section attention transformers, and further strengths the feature interaction by two fusion gates: the Residual Fusion Gate (RFG) and the Feature Fusion Gate (FFG). The proposed method is evaluated on three long document datasets and the experimental results show that our approach outperforms the related state-of-the-art methods. The code will be available at https://github.com/TengfeiLiu966/HATN
Yongli Hu, Puman Chen, Tengfei Liu 0005, Junbin Gao
IJCNN1
2021 Attributed Non-negative Matrix Multi-factorization for Data Representation
Jie Wang 0134, Jipeng Guo 0001, Yongli Hu
PRCV (4)4
2021 Complete/incomplete multi-view subspace clustering via soft block-diagonal-induced regulariser
abstract
Abstract This study proposes a novel multi‐view soft block diagonal representation framework for clustering complete and incomplete multi‐view data. First, given that the multi‐view self‐representation model offers better performance in exploring the intrinsic structure of multi‐view data, it can be nicely adopted to individually construct a graph for each view. Second, since an ideal block diagonal graph is beneficial for clustering, a ‘soft’ block diagonal affinity matrix is constructed by fusing multiple previous graphs. The soft diagonal block regulariser encourages a matrix to approximately have (not exactly) diagonal blocks, where is the number of clusters. This strategy adds robustness to noise and outliers. Third, to handle incomplete multi‐view data, multiple indicator matrices are utilised, which can mark the position of missing elements of each view. Finally, the alternative direction of multipliers algorithm is employed to optimise the proposed model, and the corresponding algorithm complexity and convergence are also analysed. Extensive experimental results on several real‐world datasets achieve the best performance among the state‐of‐the‐art complete and incomplete clustering methods, which proves the effectiveness of the proposed methods.
Yongli Hu, Cuicui Luo, Boyue Wang, Junbin Gao
IET Comput. Vis.1
2021 Kronecker-decomposable robust probabilistic tensor discriminant analysis
Fujiao Ju, Junbin Gao, Yongli Hu
Inf. Sci.4
2021 AKM3C: Adaptive K-Multiple-Means for Multi-View Clustering
abstract
With the popularity of cameras and sensors, massive data are captured from various view angles or modalities, which provide abundant complementary information and also bring great challenges for traditional clustering methods. In this article, we propose a novel Adaptive K-Multiple-Means for multi-view clustering method (AKM3C). Unlike traditional multi-view K-means methods by grouping samples into$C$clusters each with a cluster center in every view, the proposed AKM3C employs$M (M>C)$sub-cluster centers in each view to reveal the sub-cluster structure in the multi-view data thus enhances the clustering performance. Additionally, to distinguish the importance of different views, instead of using empirical weights, AKM3C exploits the multi-view combination weights strategy to assign a weight to each view automatically and thus fuses the complementary information of different views properly to get an optimally shared bipartite graph, on which the Laplacian rank constraint is executed and the final clusters are obtained by directly partitioning. An efficient optimization algorithm proposed with complexity and convergence analysis is used to solve the proposed AKM3C method. The extensive experimental results on eight public datasets show that the proposed AKM3C performs better than state-of-the-art multi-view clustering methods. The code can be downloaded athttps://drive.google.com/file/d/1CQ0royrYxKFJdNLnbBQSbDrohtfH71di/view?usp=sharing.
Yongli Hu, Zuolong Song, Boyue Wang, Junbin Gao
IEEE Trans. Circuits Syst. Video Technol.1
2021 Optimized Graph Convolution Recurrent Neural Network for Traffic Prediction
abstract
Traffic prediction is a core problem in the intelligent transportation system and has broad applications in the transportation management and planning, and the main challenge of this field is how to efficiently explore the spatial and temporal information of traffic data. Recently, various deep learning methods, such as convolution neural network (CNN), have shown promising performance in traffic prediction. However, it samples traffic data in regular grids as the input of CNN, thus it destroys the spatial structure of the road network. In this paper, we introduce a graph network and propose an optimized graph convolution recurrent neural network for traffic prediction, in which the spatial information of the road network is represented as a graph. Additionally, distinguishing with most current methods using a simple and empirical spatial graph, the proposed method learns an optimized graph through a data-driven way in the training phase, which reveals the latent relationship among the road segments from the traffic data. Lastly, the proposed method is evaluated on three real-world case studies, and the experimental results show that the proposed method outperforms state-of-the-art traffic prediction methods.
Kan Guo, Yongli Hu, Sean Qian, Hao Liu 0040, Ke Zhang 0016, Junbin Gao
IEEE Trans. Intell. Transp. Syst.2
2021 Metro Passenger Flow Prediction via Dynamic Hypergraph Convolution Networks
abstract
Metro passenger flow prediction is a strategically necessary demand in an intelligent transportation system to alleviate traffic pressure, coordinate operation schedules, and plan future constructions. Graph-based neural networks have been widely used in traffic flow prediction problems. Graph Convolutional Neural Networks (GCN) captures spatial features according to established connections but ignores the high-order relationships between stations and the travel patterns of passengers. In this paper, we utilize a novel representation to tackle this issue - hypergraph. A dynamic spatio-temporal hypergraph neural network to forecast passenger flow is proposed. In the prediction framework, the primary hypergraph is constructed from metro system topology and then extended with advanced hyperedges discovered from pedestrian travel patterns of multiple time spans. Furthermore, hypergraph convolution and spatio-temporal blocks are proposed to extract spatial and temporal features to achieve node-level prediction. Experiments on historical datasets of Beijing and Hangzhou validate the effectiveness of the proposed method, and superior performance of prediction accuracy is achieved compared with the state-of-the-arts.
Yong Zhang 0029, Yongli Hu, Xinglin Piao
IEEE Trans. Intell. Transp. Syst.4
2021 A Low Rank Dynamic Mode Decomposition Model for Short-Term Traffic Flow Prediction
abstract
Traffic flow data has three main characteristics: large amount of noise and incompleteness, temporal and spatial correlation, and dynamic sequential property. Problems of noise, loss and incompleteness could decrease the prediction performance and make it difficult for transportation system management. Inspired by recent work on low rank representation (LRR) and dynamic mode decomposition (DMD), we propose a Low Rank Dynamic Mode Decomposition (LRDMD) model which solves the aforementioned problems simultaneously. LRDMD predicts traffic flow by using a state transition matrix which characterizes the relationship between temporally neighboring fragments of traffic flow with low rank regularization. We conduct experiments of traffic flow prediction of different time intervals using loop coil detector data of Qingdao, and the results show that LRDMD outperforms state-of-the-art methods.
Yadong Yu, Yong Zhang 0029, Sean Qian, Shaofan Wang 0001, Yongli Hu
IEEE Trans. Intell. Transp. Syst.5
2021 Interactive Visual Exploration of Human Mobility Correlation Based on Smart Card Data
abstract
Public transportation agencies call for an intuitive, interactive, and reusable visualization tool to detect patterns of crime (i.e. pickpockets and gangs) or missing commuters on public transportation systems. Few existing visualization techniques have visually explored mobility correlations of targets and their companions, who are characterized in diverse mobility types, by using discrete travel hints extracted from a massive amount of data. To fill this gap, a visual analytical system is provided to conduct a group-based and individual-based exploration of mobility correlations of passengers of interest, based on an auto integration of multiple queries. How passengers differ from or correlate with each other are further examined based on their spatiotemporal distributions in trajectories and ODs. Real-world case studies, as well as user feedback made by 30 participants, demonstrate the effectiveness of the system in detecting specific targets and their companions featured in diverse mobility types, or in characterizing their spatiotemporal aggregation patterns for a further tracking on public transportation systems.
Xia Zhao 0003, Yong Zhang 0029, Yongli Hu, Shun Wang 0004, Yunhui Li, Sean Qian
IEEE Trans. Intell. Transp. Syst.3
2021 Robust Image Representation via Low Rank Locality Preserving Projection
abstract
Locality preserving projection (LPP) is a dimensionality reduction algorithm preserving the neighhorhood graph structure of data. However, the conventional LPP is sensitive to outliers existing in data. This article proposes a novel low-rank LPP model called LR-LPP. In this new model, original data are decomposed into the clean intrinsic component and noise component. Then the projective matrix is learned based on the clean intrinsic component which is encoded in low-rank features. The noise component is constrained by theℓ1-norm which is more robust to outliers. Finally, LR-LPP model is extended to LR-FLPP in which low-dimensional feature is measured by F-norm. LR-FLPP will reduce aggregated error and weaken the effect of outliers, which will make the proposed LR-FLPP even more robust for outliers. The experimental results on public image databases demonstrate the effectiveness of the proposed LR-LPP and LR-FLPP.
Junbin Gao, Yongli Hu, Boyue Wang
ACM Trans. Knowl. Discov. Data4
2021 Learning Adaptive Neighborhood Graph on Grassmann Manifolds for Video/Image-Set Subspace Clustering
abstract
The objective of self-expression based spectral clustering is to learn an affinity matrix which accurately reflects the similarity among data, and the Laplacian constraint is usually exploited to make the affinity matrix preserve the global structure of raw data. However, there exist two drawbacks: firstly, these methods are mostly designed for vectorial data in Euclidean spaces, which are not suitable for multidimensional data with nonlinear manifold structure, e.g., videos and image-sets. Secondly, the clustering performance heavily relies on the quality of a pre-learned Laplacian matrix in which the global structure may be mis-interpreted without considering manifold structures. In this paper, we firstly provide a unified framework about self-expression learning on Grassmann manifolds, which implements the clustering tasks for multidimensional data under subspace views. Then, to assign optimal neighbors to each data depending on the local distance, we adaptively learn the neighborhood relationship from the obtained self-expression coefficient matrix, referred to Learning Adaptive Neighborhood Graph on Grassmann manifolds (GMAN). In the optimization process, the neighborhood relationship can be adaptively learned and updated with the coefficient matrix. The experimental results on five public datasets show that the proposed method is obviously better than many related clustering methods based on Grassmann manifolds, proving the effectiveness of GMAN in multidimensional data clustering.
Boyue Wang, Yongli Hu, Junbin Gao, Fujiao Ju
IEEE Trans. Multim.2
2021 Robust CAPTCHAs Towards Malicious OCR
abstract
Turing test was originally proposed to examine whether machine's behavior is indistinguishable from a human. The most popular and practical Turing test is CAPTCHA, which is to discriminate algorithm from human by offering recognition-alike questions. The recent development of deep learning has significantly advanced the capability of algorithm in solving CAPTCHA questions, forcing CAPTCHA designers to increase question complexity. Instead of designing questions difficult for both algorithm and human, this study attempts to employ the limitations of algorithm to design robust CAPTCHA questions easily solvable to human. Specifically, our data analysis observes that human and algorithm demonstrates different vulnerability to visual distortions: adversarial perturbation is significantly annoying to algorithm yet friendly to human. We are motivated to employ adversarially perturbed images for robust CAPTCHA design in the context of character-based questions. Four modules of multi-target attack, ensemble adversarial training, image preprocessing differentiable approximation, and expectation are proposed to address the characteristics of character-based CAPTCHA cracking. Qualitative and quantitative experimental results demonstrate the effectiveness of the proposed solution. We hope this study can lead to the discussions around adversarial attack/defense in CAPTCHA design and also inspire the future attempts in employing algorithm limitation for practical usage.
Jiaming Zhang 0006, Jitao Sang 0001, Shangxi Wu, Yongli Hu, Jian Yu 0001
IEEE Trans. Multim.7
2021 Adaptive Fusion of Heterogeneous Manifolds for Subspace Clustering
abstract
Multiview clustering (MVC) has recently received great interest due to its pleasing efficacy in combining the abundant and complementary information to improve clustering performance, which overcomes the drawbacks of view limitation existed in the standard single-view clustering. However, the existing MVC methods are mostly designed for vectorial data from linear spaces and, thus, are not suitable for multiple dimensional data with intrinsic nonlinear manifold structures, e.g., videos or image sets. Some works have introduced manifolds' representation methods of data into MVC and obtained considerable improvements, but how to fuse multiple manifolds efficiently for clustering is still a challenging problem. Particularly, for heterogeneous manifolds, it is an entirely new problem. In this article, we propose to represent the complicated multiviews' data as heterogeneous manifolds and a fusion framework of heterogeneous manifolds for clustering. Different from the empirical weighting methods, an adaptive fusion strategy is designed to weight the importance of different manifolds in a data-driven manner. In addition, the low-rank representation is generalized onto the fused heterogeneous manifolds to explore the low-dimensional subspace structures embedded in data for clustering. We assessed the proposed method on several public data sets, including human action video, facial image, and traffic scenario video. The experimental results show that our method obviously outperforms a number of state-of-the-art clustering methods.
Boyue Wang, Yongli Hu, Junbin Gao, Fujiao Ju
IEEE Trans. Neural Networks Learn. Syst.2
2020 Reweighted Non-convex Non-smooth Rank Minimization Based Spectral Clustering on Grassmann Manifold
Xinglin Piao, Yongli Hu, Junbin Gao, Xin Yang 0011
ACCV (5)2
2020 Kernel Clustering On Symmetric Positive Definite Manifolds Via Double Approximated Low Rank Representation
abstract
As an effective descriptor, Symmetric Positive Definite (SPD) matrix is widely used in several areas such as image clustering. Recently, researchers proposed some effective methods based on low rank theory for SPD data clustering with nonlinear metric. However, single nuclear norm is always adopted to formulate the low rank model in these methods, which would lead to suboptimal solution. In this paper, we proposed a novel double low rank representation method for SPD clustering problem, in which matrix factorization and nonconvex rank constraint are combined to reveal the intrinsic property of the data instead of employing the nuclear norm. Meanwhile, kernel method and Log-Euclidean metric are combined to better explore the intrinsic geometry within SPD data. The proposed method has been evaluated on several public datasets and the experimental results demonstrate that the proposed method outperforms the state-of-the-art ones.
Xinglin Piao, Yongli Hu, Junbin Gao, Xin Yang 0011, Wenwu Zhu 0001, Ge Li 0002
ICME2
2020 Low Rank Representation on Product Grassmann Manifolds for Multi-view Subspace Clustering
abstract
Clustering high dimension multi-view data with complex intrinsic properties and nonlinear manifold structure is a challenging task since these data are always embedded in low dimension manifolds. Inspired by Low Rank Representation (LRR), some researchers extended classic LRR on Grassmann manifold or Product Grassmann manifold to represent data with non-linear metrics. However, most of these methods utilized convex nuclear norm to leverage a low-rank structure, which was over-relaxation of true rank and would lead to the results deviated from the true underlying ones. And, the computational complexity of singular value decomposition of matrix is high for nuclear norm minimization. In this paper, we propose a new low rank model for high-dimension multi-view data clustering on Product Grassmann Manifold with the matrix tri-factorization which is used to control the upper bound of true rank of representation matrix. And, the original problem can be transformed into the nuclear norm minimization with smaller scale matrices. An effective solution and theoretical analysis are also provided. The experimental results show that the proposed method obviously outperforms other state-of-the-art methods on several multi-source human/crowd action video datasets.
Jipeng Guo 0001, Junbin Gao, Yongli Hu
ICPR4
2020 Double Manifolds Regularized Non-negative Matrix Factorization for Data Representation
abstract
Non-negative matrix factorization (NMF) is an important method in learning latent data representation. The local geometrical structure can make the learned representation more effectively and significantly improve the performance of NMF. However, most of existing graph-based learning methods are determined by a predefined similarity graph which may be not optimal for specific tasks. To solve the above problem, we propose the Double Manifolds Regularized NMF (DMR-NMF) model which jointly learns an adaptive affinity matrix with the nonnegative matrix factorization. The learned affinity matrix can guide the NMF to fit the clustering task. Moreover, we develop the iterative updating optimization schemes for DMR-NMF, and provide the strict convergence proof of our optimization strategy. Empirical experiments on four different real-world data sets demonstrate the state-of-the-art performance of DMR-NMF in comparison with the other related algorithms.
Jipeng Guo 0001, Yongli Hu
ICPR4
2020 Variational Deep Embedding Clustering by Augmented Mutual Information Maximization
abstract
Clustering is a crucial but challenging task in pattern analysis and machine learning. Recent many deep clustering methods combining representation learning with cluster techniques emerged. These deep clustering methods mainly focus on the correlation among samples and ignore the relationship between samples and their representations. In this paper, we propose a novel end-to-end clustering framework, namely variational deep embedding clustering by augmented mutual information maximization (VCAMI). From the perspective of VAE, we prove that minimizing reconstruction loss is equivalent to maximizing the mutual information of the input and its latent representation. This provides a theoretical guarantee for us to directly maximize the mutual information instead of minimizing reconstruction loss. Therefore we proposed the augmented mutual information which highlights the uniqueness of the representations while discovering invariant information among similar samples. Extensive experiments on several challenging image datasets show that the VCAMI achieves good performance. we achieve state-of-the-art ACC results for clustering on MNIST (99.5%) and CIFAR-10 (65.4%) to the best of our knowledge.
Yongli Hu
ICPR3
2020 Zero-Shot Text Classification with Semantically Extended Graph Convolutional Network
abstract
As a challenging task of Natural Language Processing(NLP), zero-shot text classification has attracted more and more attention recently. It aims to detect classes that the model has never seen in the training set. For this purpose, a feasible way is to construct connection between the seen and unseen classes by semantic extension and classify the unseen classes by information propagation over the connection. Although many related zero-shot text classification methods have been exploited, how to realize semantic extension properly and propagate information effectively are far from solved. In this paper, we propose a novel zero-shot text classification method called Semantically Extended Graph Convolutional Network (SEGCN). In the proposed method, the semantic category knowledge from ConceptNet is utilized to semantic extension for linking seen classes to unseen classes and constructing a graph of all categories. Then, we build upon Graph Convolutional Network (GCN) for predicting the textual classifier for each category, which transfers the category knowledge by the convolution operators on the constructed graph and is trained in a semi-supervised manner using the samples of the seen classes. The experimental results on Dbpedia and 20newsgroup datasets show that our method outperforms the state of the art zero-shot text classification methods.
Tengfei Liu 0005, Yongli Hu, Junbin Gao
ICPR2
2020 A Spectral Clustering on Grassmann Manifold via Double Low Rank Constraint
abstract
Data clustering is a fundamental topic in machine learning and data mining areas. In recent years, researchers have proposed a series of effective methods based on Low Rank Representation (LRR) which could explore low-dimension subspace structure embedded in original data effectively. The traditional LRR methods usually are designed for vectorial data from linear spaces with Euclidean distance. However, high-dimension data (such as video clip or imageset) are always considered as non-linear manifold data such as Grassmann manifold with non-linear metric. In addition, traditional LRR clustering method always adopt single nuclear norm as low rank constraint which would lead to suboptimal solution and decrease the clustering accuracy. In this paper, we proposed a new low rank method on Grassmann manifold for video or imageset data clustering task. In the proposed method, video or imageset data are formulated as sample data on Grassmann manifold first. And then a double low rank constraint is proposed by combining the nuclear norm and bilinear representation for better construct the representation matrix. The experimental results on several public datasets show that the proposed method outperforms the state-of-the-art clustering methods.
Xinglin Piao, Yongli Hu, Junbin Gao, Xin Yang 0011
ICPR2
2020 Adversarial Privacy-preserving Filter
abstract
While widely adopted in practical applications, face recognition has been critically discussed regarding the malicious use of face images and the potential privacy problems, e.g., deceiving payment system and causing personal sabotage. Online photo sharing services unintentionally act as the main repository for malicious crawler and face recognition applications. This work aims to develop a privacy-preserving solution, called Adversarial Privacy-preserving Filter (APF), to protect the online shared face images from being maliciously used. We propose an end-cloud collaborated adversarial attack solution to satisfy requirements of privacy, utility and non-accessibility. Specifically, the solutions consist of three modules: (1) image-specific gradient generation, to extract image-specific gradient in the user end with a compressed probe model; (2) adversarial gradient transfer, to fine-tune the image-specific gradient in the server cloud; and (3) universal adversarial perturbation enhancement, to append image-independent perturbation to derive the final adversarial noise. Extensive experiments on three datasets validate the effectiveness and efficiency of the proposed solution. A prototype application is also released for further evaluation. We hope the end-cloud collaborated attack framework could shed light on addressing the issue of online multimedia sharing privacy-preserving issues from user side.
Jiaming Zhang 0006, Jitao Sang 0001, Xiaowen Huang 0001, Yongli Hu
ACM Multimedia6
2020 Locality preserving projection based on Euler representation
Tianhang Long, Junbin Gao, Yongli Hu
J. Vis. Commun. Image Represent.4
2020 Robust Adaptive Linear Discriminant Analysis with Bidirectional Reconstruction Constraint
abstract
Linear discriminant analysis (LDA) is a well-known supervised method for dimensionality reduction in which the global structure of data can be preserved. The classical LDA is sensitive to the noises, and the projection direction of LDA cannot preserve the main energy. This article proposes a novel feature extraction model with l 2,1 norm constraint based on LDA, termed as RALDA. This model preserves within-class local structure in the latent subspace according to the label information. To reduce information loss, it learns a projection matrix and an inverse projection matrix simultaneously. By introducing an implicit variable and matrix norm transformation, the alternating direction multiple method with updating variables is designed to solve the RALDA model. Moreover, both computational complexity and weak convergence property of the proposed algorithm are investigated. The experimental results on several public databases have demonstrated the effectiveness of our proposed method.
Jipeng Guo 0001, Junbin Gao, Yongli Hu
ACM Trans. Knowl. Discov. Data4
2019 Double Nuclear Norm Based Low Rank Representation on Grassmann Manifolds for Clustering
abstract
Unsupervised clustering for high-dimension data (such as imageset or video) is a hard issue in data processing and data mining area since these data always lie on a manifold (such as Grassmann manifold). Inspired of Low Rank representation theory, researchers proposed a series of effective clustering methods for high-dimension data with non-linear metric. However, most of these methods adopt the traditional single nuclear norm as the relaxation of the rank function, which would lead to suboptimal solution deviated from the original one. In this paper, we propose a new low rank model for high-dimension data clustering task on Grassmann manifold based on the Double Nuclear norm which is used to better approximate the rank minimization of matrix. Further, to consider the inner geometry or structure of data space, we integrated the adaptive Laplacian regularization to construct the local relationship of data samples. The proposed models have been assessed on several public datasets for imageset clustering. The experimental results show that the proposed models outperform the state-of-the-art clustering ones.
Xinglin Piao, Yongli Hu, Junbin Gao
CVPR2
2019 Locality Preserving Projection via Deep Neural Network
abstract
Dimensionality reduction is an essential problem in data mining and machine learning fields. Locality Preserving Projection (LPP) is a well-known dimensionality reduction method which can preserve the neighborhood graph structure of data, and has achieved promising performance. However linear projection makes it difficult to analyze complex data with nonlinear structure. In order to deal with this issue, this paper proposes a novel nonlinear locality preserving projection method via deep neural network, termed as DNLPP, which replaces the linear projection with an appropriate deep neural network. Benefiting from the nonlinearity of neural networks and its powerful representation capability, the proposed method is more discriminative than the conventional LPP. In order to solve the new model, we propose an iterative optimization algorithm. Extensive experiments on several public datasets illustrate that the proposed method is overall superior to the other state-of-art dimensionality reduction methods.
Tianhang Long, Junbin Gao, Mingyan Yang, Yongli Hu
IJCNN4
2019 Domain-invariant representation learning using an unsupervised domain adversarial adaptation deep neural network
Xibin Jia, Ya Jin, Xing Su 0001, Yongli Hu
Neurocomputing4
2019 Maximally Correlated Principal Component Analysis Based on Deep Parameterization Learning
abstract
Dimensionality reduction is widely used to deal with high-dimensional data. As a famous dimensionality reduction method, principal component analysis (PCA) aiming at finding the low dimension feature of original data has made great successes, and many improved PCA algorithms have been proposed. However, most algorithms based on PCA only consider the linear correlation of data features. In this article, we propose a novel dimensionality reduction model called maximally correlated PCA based on deep parameterization learning (MCPCADP), which takes nonlinear correlation into account in the deep parameterization framework for the purpose of dimensionality reduction. The new model explores nonlinear correlation by maximizing Ky-Fan norm of the covariance matrix of nonlinearly mapped data features. A new BP algorithm for model optimization is derived. In order to assess the proposed method, we conduct experiments on both a synthetic database and several real-world databases. The experimental results demonstrate that the proposed algorithm is comparable to several widely used algorithms.
Haoran Chen 0004, Junbin Gao, Yongli Hu
ACM Trans. Knowl. Discov. Data5
2019 Solving Partial Least Squares Regression via Manifold Optimization Approaches
abstract
Partial least squares regression (PLSR) has been a popular technique to explore the linear relationship between two data sets. However, all existing approaches often optimize a PLSR model in Euclidean space and take a successive strategy to calculate all the factors one by one for keeping the mutually orthogonal PLSR factors. Thus, a suboptimal solution is often generated. To overcome the shortcoming, this paper takes statistically inspired modification of PLSR (SIMPLSR) as a representative of PLSR, proposes a novel approach to transform SIMPLSR into optimization problems on Riemannian manifolds, and develops corresponding optimization algorithms. These algorithms can calculate all the PLSR factors simultaneously to avoid any suboptimal solutions. Moreover, we propose sparse SIMPLSR on Riemannian manifolds, which is simple and intuitive. A number of experiments on classification problems have demonstrated that the proposed models and algorithms can get lower classification error rates compared with other linear regression methods in Euclidean space. We have made the experimental code public at https://github.com/Haoran2014.
Haoran Chen 0004, Junbin Gao, Yongli Hu
IEEE Trans. Neural Networks Learn. Syst.4
2019 Probabilistic Linear Discriminant Analysis With Vectorial Representation for Tensor Data
abstract
Linear discriminant analysis (LDA) has been a widely used supervised feature extraction and dimension reduction method in pattern recognition and data analysis. However, facing high-order tensor data, the traditional LDA-based methods take two strategies. One is vectorizing original data as the first step. The process of vectorization will destroy the structure of high-order data and result in high dimensionality issue. Another is tensor LDA-based algorithms that extract features from each mode of high-order data and the obtained representations are also high-order tensor. This paper proposes a new probabilistic LDA (PLDA) model for tensorial data, namely, tensor PLDA. In this model, each tensorial data are decomposed into three parts: the shared subspace component, the individual subspace component, and the noise part. Furthermore, the first two parts are modeled by a linear combination of latent tensor bases, and the noise component is assumed to follow a multivariate Gaussian distribution. Model learning is conducted through a Bayesian inference process. To further reduce the total number of model parameters, the tensor bases are assumed to have tensor CandeComp/PARAFAC (CP) decomposition. Two types of experiments, data reconstruction and classification, are conducted to evaluate the performance of the proposed model with the convincing result, which is superior or comparable against the existing LDA-based methods.
Fujiao Ju, Junbin Gao, Yongli Hu
IEEE Trans. Neural Networks Learn. Syst.4
2018 Locality Preserving Projection Based on F-norm
abstract
Locality preserving projection (LPP) is a well-known method for dimensionality reduction in which the neighborhood graph structure of data is preserved. Traditional LPP employ squared F-norm for distance measurement. This may exaggerate more distance errors, and result in a model being sensitive to outliers. In order to deal with this issue, we propose two novel F-norm-based models, termed as F-LPP and F-2DLPP, which are developed for vector-based and matrix-based data, respectively. In F-LPP and F-2DLPP, the distance of data projected to a low dimensional space is measured by F-norm. Thus it is anticipated that both methods can reduce the influence of outliers. To solve the F-norm-based models, we propose an iterative optimization algorithm, and give the convergence analysis of algorithm. The experimental results on three public databases have demonstrated the effectiveness of our proposed methods.
Xiangjie Hu, Junbin Gao, Yongli Hu
AAAI4
2018 Cascaded Low Rank and Sparse Representation on Grassmann Manifolds
abstract
Inspired by low rank representation and sparse subspace clustering acquiring success, ones attempt to simultaneously perform low rank and sparse constraints on the affinity matrix to improve the performance. However, it is just a trade-off between these two constraints. In this paper, we propose a novel Cascaded Low Rank and Sparse Representation (CLRSR) method for subspace clustering, which seeks the sparse expression on the former learned low rank latent representation. To make our proposed method suitable to multi-dimension or imageset data, we extend CLRSR onto Grassmann manifolds. An effective solution and its convergence analysis are also provided. The excellent experimental results demonstrate the proposed method is more robust than other state-of-the-art clustering methods on imageset data.
Boyue Wang, Yongli Hu, Junbin Gao
IJCAI2
2018 Fast optimization algorithm on Riemannian manifolds and its application in low-rank learning
Haoran Chen 0004, Junbin Gao, Yongli Hu
Neurocomputing4
2018 Low Rank Representation on SPD matrices with Log-Euclidean metric
Boyue Wang, Yongli Hu, Junbin Gao, Muhammad Ali 0005, David Tien
Pattern Recognit.2
2018 Localized LRR on Grassmann Manifold: An Extrinsic View
abstract
Subspace data representation has recently become a common practice in many computer vision tasks. Low-rank representation (LRR) is one of the most successful models for clustering vectorial data according to their subspace structures. This paper explores the possibility of extending LRR for subspace data on Grassmann manifold. Rather than directly embedding the Grassmann manifold into the symmetric matrix space, an extrinsic view is taken to build the self-representation in the local area of the tangent space at each Grassmannian point, resulting in a localized LRR method on Grassmann manifold. A novel algorithm for solving the proposed model is investigated and implemented. The performance of the new clustering algorithm is assessed through experiments on several real-world data sets including MNIST handwritten digits, ballet video clips, SKIG action clips, and DynTex++ data set and highway traffic video clips. The experimental results show that the new method outperforms a number of state-of-the-art clustering methods.
Boyue Wang, Yongli Hu, Junbin Gao
IEEE Trans. Circuits Syst. Video Technol.2
2018 Partial Sum Minimization of Singular Values Representation on Grassmann Manifolds
abstract
Clustering is one of the fundamental topics in data mining and pattern recognition. As a prospective clustering method, the subspace clustering has made considerable progress in recent researches, e.g., sparse subspace clustering (SSC) and low rank representation (LRR). However, most existing subspace clustering algorithms are designed for vectorial data from linear spaces, thus not suitable for high-dimensional data with intrinsic non-linear manifold structure. For high-dimensional or manifold data, few research pays attention to clustering problems. The purpose of clustering on manifolds tends to cluster manifold-valued data into several groups according to the mainfold-based similarity metric. This article proposes an extended LRR model for manifold-valued Grassmann data that incorporates prior knowledge by minimizing partial sum of singular values instead of the nuclear norm, namely Partial Sum minimization of Singular Values Representation (GPSSVR). The new model not only enforces the global structure of data in low rank, but also retains important information by minimizing only smaller singular values. To further maintain the local structures among Grassmann points, we also integrate the Laplacian penalty with GPSSVR. The proposed model and algorithms are assessed on a public human face dataset, some widely used human action video datasets and a real scenery dataset. The experimental results show that the proposed methods obviously outperform other state-of-the-art methods.
Boyue Wang, Yongli Hu, Junbin Gao
ACM Trans. Knowl. Discov. Data2
2018 Vectorial Dimension Reduction for Tensors Based on Bayesian Inference
Fujiao Ju, Junbin Gao, Yongli Hu
IEEE Trans. Neural Networks Learn. Syst.4
2017 Locality Preserving Projections for Grassmann manifold
abstract
Learning on Grassmann manifold has become popular in many computer vision tasks, with the strong capability to extract discriminative information for imagesets and videos. However, such learning algorithms particularly on high-dimensional Grassmann manifold always involve with significantly high computational cost, which seriously limits the applicability of learning on Grassmann manifold in more wide areas. In this research, we propose an unsupervised dimensionality reduction algorithm on Grassmann manifold based on the Locality Preserving Projections (LPP) criterion. LPP is a commonly used dimensionality reduction algorithm for vector-valued data, aiming to preserve local structure of data in the dimension-reduced space. The strategy is to construct a mapping from higher dimensional Grassmann manifold into the one in a relative low-dimensional with more discriminative capability. The proposed method can be optimized as a basic eigenvalue problem. The performance of our proposed method is assessed on several classification and clustering tasks and the experimental results show show its clear advantages over other Grassmann based algorithms.
Boyue Wang, Yongli Hu, Junbin Gao, Haoran Chen 0004, Muhammad Ali 0005
IJCAI2
2017 Matrix variate RBM model with Gaussian distributions
abstract
Restricted Boltzmann Machine (RBM) is a particular type of random neural network models modeling vector data based on the assumption of Bernoulli distribution. For multidimensional and non-binary data, it is necessary to vectorize and discretize the information in order to apply the conventional RBM. It is well-known that vectorization would destroy internal structure of data, and the binary units will limit the applying performance due to fickle real data. To address these issues, this paper proposes a Matrix variate Gaussian Restricted Boltzmann Machine (MVGRBM) model for matrix data whose entries follow Gaussian distributions. Compared with some other RBM algorithms, MVGRBM can model real value data better and it has good performance in image classification. To prove that adding Gaussian parameters could model input data well, we compared the reconstruction performance of the Gaussian parameters updating and fixed.
Simeng Liu, Yongli Hu, Junbin Gao, Fujiao Ju
IJCNN3
2017 Laplacian LRR on Product Grassmann Manifolds for Human Activity Clustering in Multicamera Video Surveillance
abstract
In multicamera video surveillance, it is challenging to represent videos from different cameras properly and fuse them efficiently for specific applications such as human activity recognition and clustering. In this paper, a novel representation for multicamera video data, namely, the product Grassmann manifold (PGM), is proposed to model video sequences as points on the Grassmann manifold and integrate them as a whole in the product manifold form. In addition, with a new geometry metric on the product manifold, the conventional low rank representation (LRR) model is extended onto PGM and the new LRR model can be used for clustering nonlinear data, such as multicamera video data. To evaluate the proposed method, a number of clustering experiments are conducted on several multicamera video data sets of human activity, including the Dongzhimen Transport Hub Crowd action data set, the ACT 42 Human Action data set, and the SKIG action data set. The experiment results show that the proposed method outperforms many state-of-the-art clustering methods.
Boyue Wang, Yongli Hu, Junbin Gao
IEEE Trans. Circuits Syst. Video Technol.2
2016 Product Grassmann Manifold Representation and Its LRR Models
abstract
It is a challenging problem to cluster multi- and high-dimensional data with complex intrinsic properties and non-linear manifold structure. The recently proposed subspace clustering method, Low Rank Representation (LRR), shows attractive performance on data clustering, but it generally does with data in Euclidean spaces. In this paper, we intend to cluster complex high dimensional data with multiple varying factors. We propose a novel representation, namely Product Grassmann Manifold (PGM), to represent these data. Additionally, we discuss the geometry metric of the manifold and expand the conventional LRR model in Euclidean space onto PGM and thus construct a new LRR model. Several clustering experimental results show that the proposed method obtains superior accuracy compared with the clustering methods on manifolds or conventional Euclidean spaces.
Boyue Wang, Yongli Hu, Junbin Gao
AAAI2
2016 Mixture of Bilateral-Projection Two-Dimensional Probabilistic Principal Component Analysis
abstract
The probabilistic principal component analysis (PPCA) is built upon a global linear mapping, with which it is insufficient to model complex data variation. This paper proposes a mixture of bilateral-projection probabilistic principal component analysis model (mixB2DPPCA) on 2D data. With multi-components in the mixture, this model can be seen as a 'soft' cluster algorithm and has capability of modeling data with complex structures. A Bayesian inference scheme has been proposed based on the variational EM (Expectation-Maximization) approach for learning model parameters. Experiments on some publicly available databases show that the performance of mixB2DPPCA has been largely improved, resulting in more accurate reconstruction errors and recognition rates than the existing PCA-based algorithms.
Fujiao Ju, Junbin Gao, Simeng Liu, Yongli Hu
CVPR5
2016 Matrix Variate Restricted Boltzmann Machine
abstract
Restricted Boltzmann Machine (RBM) is an important generative model modeling vectorial data. While applying an RBM in practice to images, the data have to be vectorized. This results in high-dimensional data and valuable spatial information has got lost in vectorization. In this paper, a Matrix-Variate Restricted Boltzmann Machine (MVRBM) model is proposed by generalizing the classic RBM to explicitly model matrix data. In the new RBM model, both input and hidden variables are in matrix forms which are connected by bilinear transforms. The MVRBM has much less model parameters while retaining comparable performance as the classic RBM. The advantages of the MVRBM have been demonstrated on three real-world applications: handwritten digit denoising, reconstruction and recognition.
Guanglei Qi, Junbin Gao, Yongli Hu
IJCNN4
2016 Nonparametric tensor dictionary learning with beta process priors
Fujiao Ju, Junbin Gao, Yongli Hu
Neurocomputing4
2016 Ordered Subspace Clustering With Block-Diagonal Priors
abstract
Many application scenarios involve sequential data, but most existing clustering methods do not well utilize the order information embedded in sequential data. In this paper, we study the subspace clustering problem for sequential data and propose a new clustering method, namely ordered sparse clustering with block-diagonal prior (BD-OSC). Instead of using the sparse normalizer in existing sparse subspace clustering methods, a quadratic normalizer for the data sparse representation is adopted to model the correlation among the data sparse coefficients. Additionally, a block-diagonal prior for the spectral clustering affinity matrix is integrated with the model to improve clustering accuracy. To solve the proposed BD-OSC model, which is a complex optimization problem with quadratic normalizer and block-diagonal prior constraint, an efficient algorithm is proposed. We test the proposed clustering method on several types of databases, such as synthetic subspace data set, human face database, video scene clips, motion tracks, and dynamic 3-D face expression sequences. The experiments show that the proposed method outperforms state-of-the-art subspace clustering methods.
Fei Wu 0020, Yongli Hu, Junbin Gao
IEEE Trans. Cybern.2
2016 Fisher discrimination-based l2, 1-norm sparse representation for face recognition
Yong Zhang 0029, Yongli Hu, Xinglin Piao, Qianjun Wu
Vis. Comput.5
2015 Image Outlier Detection and Feature Extraction via L1-Norm-Based 2D Probabilistic PCA
abstract
This paper introduces an L1-norm-based probabilistic principal component analysis model on 2D data (L1-2DPPCA) based on the assumption of the Laplacian noise model. The Laplacian or L1 density function can be expressed as a superposition of an infinite number of Gaussian distributions. Under this expression, a Bayesian inference can be established based on the variational expectation maximization approach. All the key parameters in the probabilistic model can be learned by the proposed variational algorithm. It has experimentally been demonstrated that the newly introduced hidden variables in the superposition can serve as an effective indicator for data outliers. Experiments on some publicly available databases show that the performance of L1-2DPPCA has largely been improved after identifying and removing sample outliers, resulting in more accurate image reconstruction than the existing PCA-based methods. The performance of feature extraction of the proposed method generally outperforms other existing algorithms in terms of reconstruction errors and classification accuracy.
Fujiao Ju, Junbin Gao, Yongli Hu
IEEE Trans. Image Process.4
2014 Low Rank Representation on Grassmann Manifolds
Boyue Wang, Yongli Hu, Junbin Gao
ACCV (1)2
2014 Craniofacial reconstruction based on multi-linear subspace analysis
Fuqing Duan, Donghua Huang, Yongli Hu, Zhongke Wu
Multim. Tools Appl.4
2014 Color face recognition based on color image correlation similarity discriminant model
Huajie Jia, Yongli Hu
Multim. Tools Appl.3
2013 A hierarchical dense deformable model for 3D face reconstruction from skull
Yongli Hu, Fuqing Duan, Zhongke Wu, Guohua Geng
Multim. Tools Appl.1
2013 3D face recognition using local binary patterns
Hengliang Tang, Yongli Hu
Signal Process.4
2009 The implementation of e-learning tools to enhance undergraduate bioinformatics teaching and learning: a case study in the National University of Singapore
abstract
BACKGROUND: The rapid advancement of computer and information technology in recent years has resulted in the rise of e-learning technologies to enhance and complement traditional classroom teaching in many fields, including bioinformatics. This paper records the experience of implementing e-learning technology to support problem-based learning (PBL) in the teaching of two undergraduate bioinformatics classes in the National University of Singapore. RESULTS: Survey results further established the efficiency and suitability of e-learning tools to supplement PBL in bioinformatics education. 63.16% of year three bioinformatics students showed a positive response regarding the usefulness of the Learning Activity Management System (LAMS) e-learning tool in guiding the learning and discussion process involved in PBL and in enhancing the learning experience by breaking down PBL activities into a sequential workflow. On the other hand, 89.81% of year two bioinformatics students indicated that their revision process was positively impacted with the use of LAMS for guiding the learning process, while 60.19% agreed that the breakdown of activities into a sequential step-by-step workflow by LAMS enhances the learning experience CONCLUSION: We show that e-learning tools are useful for supplementing PBL in bioinformatics education. The results suggest that it is feasible to develop and adopt e-learning tools to supplement a variety of instructional strategies in the future.
Shen Jean Lim, Asif M. Khan, Mark De Silva, Kuan Siong Lim, Yongli Hu, Chay Hoon Tan, Tin Wee Tan
BMC Bioinform.5
2005 Multi-lighting 3D face morphable model based on mesh resampling
abstract
This work describes a multi-lighting 3D face morphable model based on mesh resampling. Despite having realistic results and modeling automatically in 3D face synthesis, the morphable model depends on the unstable optical flow algorithm for model construction, and the model matching is not fitting for complex illumination. To improve modeling results, the mesh resampling method is used to overcome the key problem of model construction; the pixel-to-pixel alignment of prototypic faces. A multi-lighting model is also proposed to evaluate the illumination of the input facial image. The experimental results show these measurements have good performance.
Yongli Hu
ICASSP (2)1
2003 A New Facial Feature Extraction Method Based on Linear Combination Model
abstract
A new facial feature extraction method is proposed. Based on linear combination model, the method locates feature points in facial images precisely. The model uses the knowledge of prototypic faces to interpret novel faces. To get the knowledge, the prototypes are labeled manually on the feature points. Generally, the construction of the linear combination model depends on pixel-wise alignments of prototypes, and the alignments are computed by an optical flow algorithm or bootstrapping algorithm which is a full-scale optimization and not includes local information such as facial feature points. To combine local facial feature with the linear combination model, a restrained optical flow algorithm is proposed to compute the pixel-wise alignments. With the information of labeled feature points, the model matches the input facial images and extracts the feature points automatically. Implementing the feature extraction method on the MPI face database, the experimental results show that the method has good performance.
Yongli Hu, Dehui Kong
Web Intelligence1