Boyue Wang

dblp:161/1804 · DBLP profile ↗
← Back
54ranked-venue papers
15as first author
44since 2021 · last 2026
0000-0002-2677-8342ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 27 · 11 first-author · 21 since 2021Graphics, computer vision, multimedia, augmented reality and games · 19 · 7 first-author · 12 since 2021Databases, data management, data science and information retrieval · 8 · 1 first-author · 7 since 2021Applied, interdisciplinary, general and emerging computing · 8 · 8 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 MARE: Multimodal Analogical Reasoning for Disease Evolution-Aware Radiology Report Generation
abstract
Radiology report generation from longitudinal medical data is critical for assessing disease progression and automating diagnostic workflows. While recent methods incorporate longitudinal information, they primarily rely on multimodal feature fusion, with limited capacity for explicit disease evolution modeling and temporal reasoning. To address this, we propose MARE, an end-to-end framework that formulates longitudinal radiology report generation as a multimodal analogical reasoning task. Inspired by the Abduction–Mapping–Induction paradigm, MARE models latent relational structures underlying disease evolution by aligning lesion-level visual features across time and mapping them to the textual domain for temporally coherent and clinically meaningful report generation. To mitigate the spatial misalignment caused by patient positioning or imaging variation, we introduce an Adaptive Region Alignment (ARA) module for robust temporal correspondence. Additionally, we design Dual Evolution Consistency (DEC) losses to regularize analogical reasoning by enforcing temporal coherence in both visual and textual evolution paths. Extensive experiments on the Longitudinal-MIMIC dataset demonstrate that MARE significantly outperforms state-of-the-art baselines across both natural language generation and clinical effectiveness metrics, highlighting the value of structured analogical reasoning for disease evolution-aware report generation.
Qingqing Gao, Tengfei Liu 0005, Xiaodan Zhang 0003, Zhongfan Sun, Boyue Wang
AAAI6
2026 MFC: Mixed Federated Clustering based on Cross-modal Feature Decoupling
abstract
Existing federated clustering methods typically assume that either all clients supply the same type of single-modal/multi-modal data, or that different clients provide various modalities describing the same object. However, a prevalent real-world scenario involves clients contributing data from diverse and unrelated modalities. Addressing the challenge of uncovering clustering patterns from such heterogeneous modality data distributed across distinct clients is crucial. In this paper, we propose a novel cross-modal feature decoupling-based mixed federated clustering model. To address client heterogeneity, we introduce a cross-modal feature decoupling module for each client, designed to decouple modality-agnostic and modality-specific features through distinct encoders.Only the modality-agnostic encoder parameters and clustering centers of each client are transmitted to the server. This enables the server to aggregate various modality-agnostic encoders, effectively discovering the global clustering structure while avoiding interference from modality-specific noise. Moreover, we develop a global consistent complementary clustering module to integrate the complementary clustering centers from various clients. The global clustering centers are then dispatched back to clients to guide and calibrate their local clustering models. Experimental results on three public datasets show the superiority of the proposed model compared to classic federated clustering methods.
Xiaxia He, Boyue Wang, Junbin Gao, Yongli Hu
KDD (1)2
2026 Adaptive low-quality negative samples partition and weakening for temporal knowledge graph completion
Boyue Wang, Yongli Hu
Multim. Syst.2
2026 Modality and semantic alignments based neighborhood node retrieval for multimodal knowledge graph completion
Dabao Zhang, Boyue Wang, Yongli Hu
Multim. Syst.2
2026 Quadruplet Augmentation With Attribute and Structure Invariance for Online Continual Learning
abstract
Online Continual Learning (OCL) learns from non-independently and identically distributed streaming data with unknown task boundaries during training and testing. Previous methods suffer from the shortcut feature trap and limited plasticity, leading to two requirements: attribute invariance and structure invariance. The former requires to capture the attributes of objects which maintain invariance during all sessions of OCL, while the latter requires to capture the relation of different attributes during OCL. From the causal invariant representation perspective, we propose Quadruplet Augmentation (QuadAug) by preserving attribute and structure invariance via data and channel augmentation with four types of augmentation strategies. First, we build a fine-grained causal graph of OCL to isolate the session-invariant attributes from confounders. Then, by observing different roles of amplitude and phase components of Fourier domain during knowledge transfer, QuadAug preserves attribute invariance by an Amplitude-Phase augmentation (AP-aug) module via a bidirectional data augmentation strategy, to intervene subtle confounders: the single-session class factor and the class-irrelevant factor. Finally, by decomposing the structure invariance into two necessary conditions: channel independence and channel sufficiency, QuadAug preserves structure invariance by an Independence-Sufficiency augmentation (IS-aug) module, which preserves the channel independence property with an inter-channel discrepancy constraint, and the channel sufficiency property with an adversarial augmentation constraint. QuadAug produces significant improvement on four sequential datasets and three blurry datasets for OCL.
Jialu Wu, Shaofan Wang 0001, Boyue Wang
IEEE Trans. Pattern Anal. Mach. Intell.4
2026 M3Former: Memory-Guided Multi-Modal Generation and Adaptive Mixture Reasoning for Incomplete-Modality Crisis Event Detection
abstract
Multi-modal crisis event detection is critical for timely situational awareness across diverse real-world emergencies. In practice, however, multi-modal data—typically composed of images and textual reports—are often incomplete, with many instances providing only a single modality because of data loss, platform constraints, or real-time limitations. This modality-incomplete setting poses three key challenges: (1) how to reliably reconstruct missing modalities to restore cross-modal context, (2) how to extract deep semantic clues across original and completed data, and (3) how to bridge the distributional gap between reconstructed and fully-observed samples during model training. To tackle these issues, we propose M3Former, a unified framework tailored for modality-incomplete multi-modal crisis event detection. It consists of four dedicated modules: Memory-Guided Modality Completion builds a memory bank of paired image–text data to retrieve semantically related samples and keywords, guiding powerful pretrained generators—diffusion models for image synthesis and multi-modal large language models for text generation; Crisis-Aware Heterogeneous Dual-Stream Encoders jointly capture modality-specific cues and establish initial cross-modal semantic alignment between original and completed data; Hierarchical Attention Refinement Network progressively refines representations through hierarchical attention and guided cross-modal interaction to suppress noise and semantic drift; Modality-aware Routing Experts designs a gated mixture-of-experts architecture that dynamically selects both modality-specific and shared experts to mitigate distributional shifts and enhance the quality of fused representations. Extensive experiments on two real-world crisis datasets demonstrate that M3Former significantly outperforms existing baselines under a variety of modality-incomplete scenarios. The code is available: https://github.com/lcygky/M3former.
Chenyang Lu 0014, Boyue Wang, Tengfei Liu 0005, Yongli Hu
IEEE Trans. Circuits Syst. Video Technol.2
2026 Emergency Events Traffic Flow Forecasting Using Text-Prompt-Guided Multimodal Large Language Models
abstract
Emergency events like traffic accidents and natural disasters frequently cause severe disruptions to urban traffic patterns, posing challenges to conventional forecasting methods based on historical data. Recent progress has explored integrating auxiliary textual information from social media, news reports, and incident details to enhance forecasting models. However, these approaches often fail to effectively align semantic context from textual data with the spatio-temporal patterns in urban traffic flow. To bridge this disparity, we propose a novel framework called TPGM-LLM, which leverages text-prompt-guided multimodal large language models to dynamically assimilate real-time data, including incident reports, traffic sensors, and weather conditions. Specifically, the proposed framework integrates a pre-trained LLM-based text encoder to interpret descriptions of emergency events as semantic prompts. Furthermore, it incorporates a dynamic spatio-temporal hypergraph module that employs FastDTW to capture non-local dependencies among different road segments. Additionally, a multimodal feature extraction LLM is utilized to merge textual guidance with traffic dynamics in a hierarchical manner, enhancing the accuracy of long-term traffic forecasting. Experimental results on the Beijing Text-Traffic dataset (BjTT) demonstrate that the proposed model outperforms existing methods, highlighting its effectiveness in complex and dynamic traffic situations. Our code is available athttps://github.com/luyaxuan/TPGM-LLM
Yaxuan Lu, Guangyu Huo, Xiaohui Cui, Boyue Wang, Yong Zhang 0029, Zhiyong Cui
IEEE Trans. Intell. Transp. Syst.4
2025 Deep Fair Multi-View Clustering with Attention KAN
abstract
Multi-view clustering is effective in unsupervised multi-view data analysis and has received considerable attention. However, most existing methods excessively emphasize certain attributes, resulting in unfair clustering outcomes, i.e., certain sensitive attributes dominate the clustering results. Moreover, existing methods struggle to effectively capture complex nonlinear relationships and interactions across views, limiting their ability to achieve optimal clustering performance. Therefore, in this work, we propose a novel method, Deep Fair Multi-View Clustering with Attention Kolmogorov-Arnold Network (DFMVC-AKAN), to generate fair clustering results while maintaining robust performance. DFMVC-AKAN integrates attention mechanisms into Kolmogorov-Arnold Networks (KAN) to exploit the complex nonlinear inter-view relationships. Specifically, KAN provides a nonlinear feature representation capable of efficiently approximating arbitrary multivariate continuous functions, augmented by a hybrid attention mechanism which enables the model to dynamically focus on the most relevant features. Finally, we refine the clustering assignments with a distribution alignment module to ensure fair outcomes across diverse groups while maintaining discriminative ability. Experimental results on four datasets containing sensitive attributes demonstrate that DFMVC-AKAN significantly improves fairness and clustering performance compared to state-of-the-art methods.
Qianqian Wang 0001, Boyue Wang, Quanxue Gao
CVPR3
2025 Scale-aware Multi-head Attention with Explainability for Image Captioning
Yuanzhen Guo, Xiaodan Zhang 0003, Aozhe Jia, Boyue Wang
PRCV (3)4
2025 Multi-view subspace clustering with incomplete graph information
abstract
Abstract The core of multi‐view clustering is how to exploit the shared and specific information of multi‐view data properly. The data missing and incompleteness bring great challenges to multi‐view clustering. In this paper, we propose an innovative multi‐view subspace clustering method with incomplete graph information, so‐called incomplete multiple graphs clustering. Specifically, we creatively separate one shared and multiple specific graphs from multiple raw graph data, and exploit the mask fusion strategy and block diagonal regulariser to obtain the inherent category information. To handle the incomplete multiple graph data, we utilise multiple indicator matrices to mark the missing elements existed in each raw graph. In addition, the weight of each raw graph is adaptively learnt according to the graph importance. The alternative direction optimization algorithm is employed to solve our proposed methods. Finally, we also analyse the algorithm convergence and the computation complexity in detail. The clustering results on six real‐world datasets show that our method obviously outperforms a serious of classic incomplete multi‐view clustering methods.
Xiaxia He, Boyue Wang, Cuicui Luo, Junbin Gao, Yongli Hu
IET Comput. Vis.2
2025 Multi-Granularity Feature Interaction and Multi-Region Selection Based Triplet Visual Question Answering
abstract
Accurately locating the question-related regions in one given image is crucial for visual question answering (VQA). The current approaches suffer two limitations: (1) Dividing one image into multiple regions may lose parts of semantic information and original relationships between regions; (2) Choosing only one or all image regions to predict the answer may correspondingly result in the insufficiency or redundancy of information. Therefore, how to effectively mine the relationship between image regions and choose the relevant image regions are vital. In this paper, we propose a novelMulti-granularity feature interaction andMulti-region selection-based triplet VQA model (M2TVQA). To tackle the first limitation, we propose the multi-granularity feature interaction strategy that adaptively supplements the global coarse-granularity features with the regional fine-granularity features. To overcome the second limitation, we design the Top-$K$learning strategy to adaptively select$K$most relevant image regions to the question, even if the selected regions are far away in space. Such a strategy can select as many relevant image regions as possible and reduce introducing noise. Finally, we construct the multi-modality triplet to predict the answer of VQA. Extended experiments on two public outside knowledge datasets OK-VQA and KRVQA verify the effectiveness of the proposed model.
Boyue Wang, Junbin Gao, Yongli Hu
IEEE Trans. Big Data2
2025 Multi-Modal Entity in One Word: Aligning Multi-Level Semantics for Multi-Modal Knowledge Graph Completion
abstract
Current multi-modal knowledge graph completion often incorporates simple fusion neural networks to achieve multi-modal alignment and knowledge completion tasks, which face three major challenges: 1) Inconsistent semantics between images and texts corresponding to the same entity; 2) Discrepancies in semantic spaces resulting from the use of diverse uni-modal feature extractors;3) Inadequate evaluation of semantic alignment using only energy functions or basic contrastive learning losses. To address these challenges, we propose the Multi-modal Entity in One Word (MEOW) model. This model ensures alignment at various levels, including text-image match alignment, feature alignment and distribution alignment. Specificially, the entity image filtering module utilizes a visual-language model to exclude unrelated images by aligning their captions with corresponding text descriptions. A pre-trained CLIP-based encoder is utilized for encoding dense semantic relationships, while a graph attention network based structure encoder handles sparse semantic relationships, yielding a comprehensive semantic representation and enhancing convergence speed. Additionally, a diffusion model is integrated to enhance denoising capabilities. The proposed MEOW further includes a distribution alignment module equipped with dense alignment constraint, integrity alignment constraint, and fusion fidelity constraint to effectively align multi-modal representations. Experiments on two public multi-modal knowledge graph datasets show that MEOW significantly improves link prediction performance. The code of the proposed model is available athttps://github.com/yuyuyuger/MEOW.
Boyue Wang, Junbin Gao, Yongli Hu
IEEE Trans. Big Data2
2025 Dual-Attention Transformers for Class-Incremental Learning: A Tale of Two Memories
abstract
Class-incremental learning (Class-IL) aims to continuously learn a model from a sequence of tasks, which suffers from the issue of catastrophic forgetting. Recently, a few transformer based methods are proposed to address this issue by transferring self-attention into task-specific attention. However, these methods utilize shared task-specific attention modules across the whole incremental learning process, and are unable to achieve the balance between consolidation and plasticity, i.e., to remember the knowledge learned from previous tasks and absorb the knowledge from the current task simultaneously. Motivated by the mechanism of LSTM and hippocampus memory, we point out that dual attention on long and short-term memories can handle the consolidation-plasticity dilemma of Class-IL. Typically, we propose Dual-Attention Transformers (DAFormer) to learn external attention and internal attention. The former utilizes sample-dependent keys which exclusively focused on the new tasks, while the latter consolidates the knowledge from previous tasks by using sample-agnostic keys. We present two editions of DAFormer: DAFormer-S and DAFormer-M: the former utilizes shared external keys and maintains a small parameter size, while the latter utilizes multiple external keys and enhances the long-term memory. Furthermore, we propose the$K$-nearest neighbor invariant based distillation scheme, which distills knowledge from previous tasks to current task by maintaining the same neighborhood relationship of each sample over old and new models. Experimental results onCIFAR-100,ImageNet-subsetandImageNet-fulldemonstrate that DAFormer significantly outperforms all the state-of-the-art parameter-static and parameter-growing methods.
Shaofan Wang 0001, Zhiyong Wang 0001, Boyue Wang
IEEE Trans. Multim.5
2025 Counterfactual Dual-Bias VQA: A Multimodality Debias Learning for Robust Visual Question Answering
abstract
Visual question answering (VQA) models often face two language bias challenges. First, they tend to rely solely on the question to predict the answer, often overlooking relevant information in the accompanying images. Second, even when considering the question, they may focus only on the wh-words, neglecting other crucial keywords that could enhance interpretability and the question sensitivity. Existing debiasing methods attempt to address this by training a bias model using question-only inputs to enhance the robustness of the target VQA model. However, this approach may not fully capture the language bias present. In this article, we propose a multimodality counterfactual dual-bias model to mitigate the linguistic bias issue in target VQA models. Our approach involves designing a shared-parameterized dual-bias model that incorporates both visual and question counterfactual samples as inputs. By doing so, we aim to fully model language biases, with visual and question counterfactual samples, respectively, emphasizing important objects and keywords to relevant the answers. To ensure that our dual-bias model behaves similarly to an ordinary model, we freeze the parameters of the target VQA model, meanwhile using the cross-entropy and Kullback-Leibler (KL) divergence as the loss function to train the dual-bias model. Subsequently, to mitigate language bias in the target VQA model, we freeze the parameters of the dual-bias model to generate pseudo-labels and then incorporate a margin loss to re-train the target VQA model. Experimental results on the VQA-CP datasets demonstrate the superior effectiveness of our proposed counterfactual dual-bias model. Additionally, we conduct an analysis of the unsatisfactory performance on the VQA v2 dataset. The origin code of the proposed model is available at https://github.com/Arrow2022jv/MCD.
Boyue Wang, Xiaoqian Ju, Junbin Gao, Yongli Hu
IEEE Trans. Neural Networks Learn. Syst.1
2025 Bridging the Cross-Modality Semantic Gap in Visual Question Answering
abstract
The objective of visual question answering (VQA) is to adequately comprehend a question and identify relevant contents in an image that can provide an answer. Existing approaches in VQA often combine visual and question features directly to create a unified cross-modality representation for answer inference. However, this kind of approach fails to bridge the semantic gap between visual and text modalities, resulting in a lack of alignment in cross-modality semantics and the inability to match key visual content accurately. In this article, we propose a model called the caption bridge-based cross-modality alignment and contrastive learning model (CBAC) to address the issue. The CBAC model aims to reduce the semantic gap between different modalities. It consists of a caption-based cross-modality alignment module and a visual-caption (V-C) contrastive learning module. By utilizing an auxiliary caption that shares the same modality as the question and has closer semantic associations with the visual, we are able to effectively reduce the semantic gap by separately matching the caption with both the question and the visual to generate pre-alignment features for each, which are then used in the subsequent fusion process. We also leverage the fact that V-C pairs exhibit stronger semantic connections compared to question-visual (Q-V) pairs to employ a contrastive learning mechanism on visual and caption pairs to further enhance the semantic alignment capabilities of single-modality encoders. Extensive experiments conducted on three benchmark datasets demonstrate that the proposed model outperforms previous state-of-the-art VQA models. Additionally, ablation experiments confirm the effectiveness of each module in our model. Furthermore, we conduct a qualitative analysis by visualizing the attention matrices to assess the reasoning reliability of the proposed model.
Boyue Wang, Yujian Ma, Junbin Gao, Yongli Hu
IEEE Trans. Neural Networks Learn. Syst.1
2025 Modality Perception Learning-Based Determinative Factor Discovery for Multimodal Fake News Detection
abstract
The dissemination of fake news, often fueled by exaggeration, distortion, or misleading statements, significantly jeopardizes public safety and shapes social opinion. Although existing multimodal fake news detection methods focus on multimodal consistency, they occasionally neglect modal heterogeneity, missing the opportunity to unearth the most related determinative information concealed within fake news articles. To address this limitation and extract more decisive information, this article proposes the modality perception learning-based determinative factor discovery (MoPeD) model. MoPeD optimizes the steps of feature extraction, fusion, and aggregation to adaptively discover determinants within both unimodality features and multimodality fusion features for the task of fake news detection. Specifically, to capture comprehensive information, the dual encoding module integrates a modal-consistent contrastive language-image pre-training (CLIP) pretrained encoder with a modal-specific encoder, catering to both explicit and implicit information. Motivated by the prompt strategy, the output features of the dual encoding module are complemented by learnable memory information. To handle modality heterogeneity during fusion, the multilevel cross-modality fusion module is introduced to deeply comprehend the complex implicit meaning within text and image. Finally, for aggregating unimodal and multimodal features, the modality perception learning module gauges the similarity between modalities to dynamically emphasize decisive modality features based on the cross-modal content heterogeneity scores. The experimental evaluations conducted on three public fake news datasets show that the proposed model is superior to other state-of-the-art fake news detection methods.
Boyue Wang, Guangchao Wu, Junbin Gao, Yongli Hu
IEEE Trans. Neural Networks Learn. Syst.1
2024 VIG: Visual Information-Guided Knowledge-Based Visual Question Answering
abstract
A better knowledge-based visual question answering (KBVQA) model needs to rely on visual features, question features, and related external knowledge to solve an open visual question answering task. Although the existing knowledge-based visual question answering works have achieved some accomplishments, there are still the following challenges: 1) There is a serious lack of visual feature information. Image information is worth a thousand words. Only relying on the converted salient text information is difficult to express the original rich information of the image. 2) The external knowledge acquired is not comprehensive enough, and there is a lack of relevant knowledge directly retrieved by visual feature information. To solve these challenges, we propose a Visual Information-Guided knowledge-based visual question answering (VIG) model. It fully considers the utilization of visual features information. Specifically: 1) We introduce multi-granularity visual information that can comprehensively characterize visual feature information. 2) We consider not only the knowledge retrieved through text information but also the knowledge directly retrieved from visual feature information. Finally, we feed the visual features and retrieved multiple text knowledge into an encoder-decoder module to generate an answer. We perform extensive experiments on the OKVQA dataset and achieve state-of-the-art performance of 60.27% accuracy.
Boyue Wang, Yongli Hu
CSCWD2
2024 IME: Integrating Multi-curvature Shared and Specific Embedding for Temporal Knowledge Graph Completion
abstract
Temporal Knowledge Graphs (TKGs) incorporate a temporal dimension, allowing for a precise capture of the evolution of knowledge and reflecting the dynamic nature of the real world. Typically, TKGs contain complex geometric structures, with various geometric structures interwoven. However, existing Temporal Knowledge Graph Completion (TKGC) methods either model TKGs in a single space or neglect the heterogeneity of different curvature spaces, thus constraining their capacity to capture these intricate geometric structures. In this paper, we propose a novel Integrating Multi-curvature shared and specific Embedding (IME) model for TKGC tasks. Concretely, IME models TKGs into multi-curvature spaces, including hyperspherical, hyperbolic, and Euclidean spaces. Subsequently, IME incorporates two key properties, namely space-shared property and space-specific property. The space-shared property facilitates the learning of commonalities across different curvature spaces and alleviates the spatial gap caused by the heterogeneous nature of multi-curvature spaces, while the space-specific property captures characteristic features. Meanwhile, IME proposes an Adjustable Multi-curvature Pooling (AMP) approach to effectively retain important information. Furthermore, IME innovatively designs similarity, difference, and structure loss functions to attain the stated objective. Experimental results clearly demonstrate the superior performance of IME over existing state-of-the-art TKGC models.
Jiapu Wang, Boyue Wang, Shirui Pan, Junbin Gao, Wen Gao 0001
WWW3
2024 A plug-and-play image enhancement model for end-to-end object detection in low-light condition
Jiaojiao Yuan, Yongli Hu, Boyue Wang
Multim. Syst.4
2024 PDRLRR: A novel low-rank representation with projection distance regularization via manifold optimization for clustering
Haoran Chen 0004, Hongwei Tao, Zuhe Li, Boyue Wang
Pattern Recognit.5
2024 Multi-Level Interaction Based Knowledge Graph Completion
abstract
With the continuous emergence of new knowledge, Knowledge Graph (KG) typically suffers from the incompleteness problem, hindering the performance of downstream applications. Thus, Knowledge Graph Completion (KGC) has attracted considerable attention. However, existing KGC methods usually capture the coarse-grained information by directly interacting with the entity and relation, ignoring the important fine-grained information in them. To capture the fine-grained information, in this paper, we divide each entity/relation into several segments and propose a novelMulti-LevelInteraction (MLI) based KGC method, which simultaneously interacts with the entity and relation at the fine-grained level and the coarse-grained level. The fine-grained interaction module applies the Gate Recurrent Unit (GRU) mechanism to guarantee the sequentiality between segments, which facilitates the fine-grained feature interaction and does not obviously sacrifice the model complexity. Moreover, the coarse-grained interaction module designs aHigh-orderFactorizedBilinear (HFB) operation to facilitate the coarse-grained interaction between the entity and relation by applying the tensor factorization based multi-head mechanism, which still effectively reduces its parameter scale. Experimental results show that the proposed method achieves state-of-the-art performances on the link prediction task over five well-established knowledge graph completion benchmarks.
Jiapu Wang, Boyue Wang, Junbin Gao, Yongli Hu
IEEE ACM Trans. Audio Speech Lang. Process.2
2024 MADE: Multicurvature Adaptive Embedding for Temporal Knowledge Graph Completion
abstract
Temporal knowledge graphs (TKGs) are receiving increased attention due to their time-dependent properties and the evolving nature of knowledge over time. TKGs typically contain complex geometric structures, such as hierarchical, ring, and chain structures, which can often be mixed together. However, embedding TKGs into Euclidean space, as is typically done with TKG completion (TKGC) models, presents a challenge when dealing with high-dimensional nonlinear data and complex geometric structures. To address this issue, we propose a novel TKGC model called multicurvature adaptive embedding (MADE). MADE models TKGs in multicurvature spaces, including flat Euclidean space (zero curvature), hyperbolic space (negative curvature), and hyperspherical space (positive curvature), to handle multiple geometric structures. We assign different weights to different curvature spaces in a data-driven manner to strengthen the ideal curvature spaces for modeling and weaken the inappropriate ones. Additionally, we introduce the quadruplet distributor (QD) to assist the information interaction in each geometric space. Ultimately, we develop an innovative temporal regularization to enhance the smoothness of timestamp embeddings by strengthening the correlation of neighboring timestamps. Experimental results show that MADE outperforms the existing state-of-the-art TKGC models.
Jiapu Wang, Boyue Wang, Junbin Gao, Shirui Pan, Tengfei Liu 0005, Wen Gao 0001
IEEE Trans. Cybern.2
2024 Mixed-Modality Clustering via Generative Graph Structure Matching
abstract
The goal of mixed-modality clustering, which differs from typical multi-modality/view clustering, is to divide samples derived from various modalities into several clusters. This task has to solve two critical semantic gap problems: i) how to generate the missing modalities without the pairwise-modality data; and ii) how to align the representations of heterogeneous modalities. To tackle the above problems, this paper proposes a novel mixedmodality clustering model, which integrates the missing-modality generation and the heterogeneous modality alignment into a unified framework. During the missing-modality generation process, a bidirectional mapping is established between different modalities, enabling generation of preliminary representations for the missing-modality using information from another modality. Then the intra-modality bipartite graphs are constructed to help generate better missing-modality representations by weighted aggregating existing intra-modality neighbors. In this way, a pairwise-modality representation for each sample can be obtained. In the process of heterogeneous modality alignment, each modality is modelled as a graph to capture the global structure among intra-modality samples and is aligned against the heterogeneous modality representations through the adaptive heterogeneous graph matching module. Experimental results on three public datasets show the effectiveness of the proposed model compared to multiple state-of-the-art multi-modality/view clustering methods.
Xiaxia He, Boyue Wang, Junbin Gao, Qianqian Wang 0001, Yongli Hu
IEEE Trans. Knowl. Data Eng.2
2024 Parallelly Adaptive Graph Convolutional Clustering Model
abstract
Benefiting from exploiting the data topological structure, graph convolutional network (GCN) has made considerable improvements in processing clustering tasks. The performance of GCN significantly relies on the quality of the pretrained graph, while the graph structures are often corrupted by noise or outliers. To overcome this problem, we replace the pre-trained and fixed graph in GCN by the adaptive graph learned from the data. In this article, we propose a novel end-to-end parallelly adaptive graph convolutional clustering (AGCC) model with two pathway networks. In the first pathway, an adaptive graph convolutional (AGC) module alternatively updates the graph structure and the data representation layer by layer. The updated graph can better reflect the data relationship than the fixed graph. In the second pathway, the auto-encoder (AE) module aims to extract the latent data features. To effectively connect the AGC and AE modules, we creatively propose an attention-mechanism-based fusion (AMF) module to weight and fuse the data representations of the two modules, and transfer them to the AGC module. This simultaneously avoids the over-smoothing problem of GCN. Experimental results on six public datasets show that the effectiveness of the proposed AGCC compared with multiple state-of-the-art deep clustering methods. The code is available at https://github.com/HeXiax/AGCC.
Xiaxia He, Boyue Wang, Yongli Hu, Junbin Gao
IEEE Trans. Neural Networks Learn. Syst.2
2024 QDN: A Quadruplet Distributor Network for Temporal Knowledge Graph Completion
abstract
Temporal knowledge graph completion (TKGC) is an extension of the traditional static knowledge graph completion (SKGC) by introducing the timestamp. The existing TKGC methods generally translate the original quadruplet to the form of the triplet by integrating the timestamp into the entity/relation, and then use SKGC methods to infer the missing item. However, such an integrating operation largely limits the expressive ability of temporal information and ignores the semantic loss problem due to the fact that entities, relations, and timestamps are located in different spaces. In this article, we propose a novel TKGC method called the quadruplet distributor network (QDN), which independently models the embeddings of entities, relations, and timestamps in their specific spaces to fully capture the semantics and builds the QD to facilitate the information aggregation and distribution among them. Furthermore, the interaction among entities, relations, and timestamps is integrated using a novel quadruplet-specific decoder, which stretches the third-order tensor to the fourth-order to satisfy the TKGC criterion. Equally important, we design a novel temporal regularization that imposes a smoothness constraint on temporal embeddings. Experimental results show that the proposed method outperforms the existing state-of-the-art TKGC methods. The source codes of this article are available at https://github.com/QDN for Temporal Knowledge Graph Completion.git.
Jiapu Wang, Boyue Wang, Junbin Gao, Yongli Hu
IEEE Trans. Neural Networks Learn. Syst.2
2023 Center Focusing Network for Real-Time LiDAR Panoptic Segmentation
abstract
LiDAR panoptic segmentation facilitates an autonomous vehicle to comprehensively understand the surrounding objects and scenes and is required to run in real time. The recent proposal-free methods accelerate the algorithm, but their effectiveness and efficiency are still limited owing to the difficulty of modeling non-existent instance centers and the costly center-based clustering modules. To achieve accurate and real-time LiDAR panoptic segmentation, a novel center focusing network (CFNet) is introduced. Specifically, the center focusing feature encoding (CFFE) is proposed to explicitly understand the relationships between the original LiDAR points and virtual instance centers by shifting the LiDAR points and filling in the center points. Moreover, to leverage the redundantly detected centers, a fast center deduplication module (CDM) is proposed to select only one center for each instance. Experiments on the SemanticKITTI and nuScenes panoptic segmentation benchmarks demonstrate that our CFNet outperforms all existing methods by a large margin and is 1.6 times faster than the most efficient method.
Boyue Wang, Yongli Hu
CVPR3
2023 DSGEM: Dual scene graph enhancement module-based visual question answering
abstract
Abstract Visual Question Answering (VQA) aims to appropriately answer a text question by understanding the image content. Attention‐based VQA models mine the implicit relationships between objects according to the feature similarity, which neglects the explicit relationships between objects, for example, the relative position. Most Visual Scene Graph‐based VQA models exploit the relative positions or visual relationships between objects to construct the visual scene graph, while they suffer from the semantic insufficiency of visual edge relations. Besides, the scene graph of text modality is often ignored in these works. In this article, a novel Dual Scene Graph Enhancement Module (DSGEM) is proposed that exploits the relevant external knowledge to simultaneously construct two interpretable scene graph structures of image and text modalities, which makes the reasoning process more logical and precise. Specifically, the authors respectively build the visual and textual scene graphs with the help of commonsense knowledge and syntactic structure, which explicitly endows the specific semantics to each edge relation. Then, two scene graph enhancement modules are proposed to propagate the involved external and structural knowledge to explicitly guide the feature interaction between objects (nodes). Finally, the authors embed such two scene graph enhancement modules to existing VQA models to introduce the explicit relation reasoning ability. Experimental results on both VQA V2 and OK‐VQA datasets show that the proposed DSGEM is effective and compatible to various VQA architectures.
Boyue Wang, Yujian Ma, Yongli Hu
IET Comput. Vis.1
2023 Graph structure learning layer and its graph convolution clustering application
Xiaxia He, Boyue Wang, Ruikun Li 0001, Junbin Gao, Yongli Hu, Guangyu Huo
Neural Networks2
2023 STGAN: Spatio-Temporal Generative Adversarial Network for Traffic Data Imputation
abstract
The traffic data corrupted by noise and missing entries often lead to the poor performance of Intelligent Transportation Systems (ITS), such as the bad congestion prediction and route guidance. How to efficiently impute the traffic data is an urgent problem. As a classic deep learning method, Generative Adversarial Network (GAN) achieves remarkable success in image recovery fields, which opens up a new way for the traffic data imputation. In this paper, we propose a novel spatio-temporal GAN model for the traffic data imputation (STGAN). Firstly, we design the generative loss and center loss, which not only minimizes the reconstructed errors of the imputed entries, but also ensures each imputed entry and its neighbors conform to the local spatio-temporal distribution. Then, the discriminator uses the convolution neural network classifier to judge whether the imputed matrix conforms to the global spatio-temporal distribution. As for the network architecture of the generator, we introduce the skip-connection to keep all well preserved data unchanged, and employ the dilated convolution to capture the spatio-temporal correlation in the traffic data. The experimental results show that our proposed method obviously outperforms other competitive traffic data imputation methods.
Yong Zhang 0029, Boyue Wang, Yongli Hu
IEEE Trans. Big Data3
2023 Hierarchical Spatio-Temporal Graph Convolutional Networks and Transformer Network for Traffic Flow Forecasting
abstract
Graph convolutional networks (GCN) have been applied in the traffic flow forecasting tasks with the graph capability in describing the irregular topology structures of road networks. However, GCN based traffic flow forecasting methods often fail to simultaneously capture the short-term and long-term temporal relations carried by the traffic flow data, and also suffer the over-smoothing problem. To overcome the problems, we propose a hierarchical traffic flow forecasting network by merging newly designed the long-term temporal Transformer network (LTT) and the spatio-temporal graph convolutional networks (STGC). Specifically, LTT aims to learn the long-term temporal relations among the traffic flow data, while the STGC module aims to capture the short-term temporal relations and spatial relations among the traffic flow data, respectively, via cascading between the one-dimensional convolution and the graph convolution. In addition, an attention fusion mechanism is proposed to combine the long-term with the short-term temporal relations as the input of the graph convolution layer in STGC, in order to mitigate the over-smoothing problem of GCN. Experimental results on three public traffic flow datasets prove the effectiveness and robustness of the proposed method.
Guangyu Huo, Yong Zhang 0029, Boyue Wang, Junbin Gao, Yongli Hu
IEEE Trans. Intell. Transp. Syst.3
2023 Multi-Concept Representation Learning for Knowledge Graph Completion
abstract
Knowledge Graph Completion (KGC) aims at inferring missing entities or relations by embedding them in a low-dimensional space. However, most existing KGC methods generally fail to handle the complex concepts hidden in triplets, so the learned embeddings of entities or relations may deviate from the true situation. In this article, we propose a novel M ulti- c oncept R epresentation L earning (McRL) method for the KGC task, which mainly consists of a multi-concept representation module, a deep residual attention module, and an interaction embedding module. Specifically, instead of the single-feature representation, the multi-concept representation module projects each entity or relation to multiple vectors to capture the complex conceptual information hidden in them. The deep residual attention module simultaneously explores the inter- and intra-connection between entities and relations to enhance the entity and relation embeddings corresponding to the current contextual situation. Moreover, the interaction embedding module further weakens the noise and ambiguity to obtain the optimal and robust embeddings. We conduct the link prediction experiment to evaluate the proposed method on several standard datasets, and experimental results show that the proposed method outperforms existing state-of-the-art KGC methods.
Jiapu Wang, Boyue Wang, Junbin Gao, Yongli Hu
ACM Trans. Knowl. Discov. Data2
2023 CaEGCN: Cross-Attention Fusion Based Enhanced Graph Convolutional Network for Clustering
abstract
With the powerful learning ability of deep convolutional networks, deep clustering methods can extract the most discriminative information from individual data and produce more satisfactory clustering results. However, existing deep clustering methods usually ignore the relationship between the data. Fortunately, the graph convolutional network can handle such relationships, opening a new research direction for deep clustering. In this paper, we propose a cross-attention based deep clustering framework, named Cross-Attention Fusion based Enhanced Graph Convolutional Network (CaEGCN), which contains four main modules: the cross-attention fusion module which innovatively concatenates the Content Auto-encoder module (CAE) relating to the individual data and Graph Convolutional Auto-encoder module (GAE) relating to the relationship between the data in a layer-by-layer manner, and the self-supervised model that highlights the discriminative information for clustering tasks. While the cross-attention fusion module fuses two kinds of heterogeneous representation, the CAE module supplements the content information for the GAE module, which avoids the over-smoothing problem of GCN. In the GAE module, two novel loss functions are proposed that reconstruct the content and relationship between the data, respectively. Finally, the self-supervised module constrains the distributions of the middle layer representations of CAE and GAE to be consistent. Experimental results on different types of datasets prove the superiority and robustness of the proposed CaEGCN.
Guangyu Huo, Yong Zhang 0029, Junbin Gao, Boyue Wang, Yongli Hu
IEEE Trans. Knowl. Data Eng.4
2023 TDN: Triplet Distributor Network for Knowledge Graph Completion
abstract
Conventional Knowledge Graph Completion (KGC) methods typically map entities and relations to a unified space through the shared mapping matrix, and then interact with entities and relations to infer the missing items in the knowledge graph. Although this shared mapping matrix considers the suitability of all triplets, it neglects the specificity of each triplet. To solve this problem, we dynamically learn one information distributor for each triplet to exchange its specific information. In this paper, we propose a novel Triplet Distributor Network (TDN) for the knowledge graph completion task. Specifically, we adaptively learn one Triplet Distributor (TD) for each triplet to assist the interaction between the entity and relation. Furthermore, on the basis of TD, we creatively design the information exchange layer to dynamically propagate the information of the entity and relation, thus mutually enhancing entity and relation representations. Except for several commonly-used knowledge graph datasets, we still implement the link prediction task on the social-relational and medical datasets to test the proposed method. Experimental results demonstrate that the proposed method performs better than existing state-of-the-art KGC methods. The source codes of this paper are available athttps://github.com/TDNfor Knowledge Graph Completion.git.
Jiapu Wang, Boyue Wang, Junbin Gao, Yongli Hu
IEEE Trans. Knowl. Data Eng.2
2023 Hierarchical Graph Convolutional Networks for Structured Long Document Classification
abstract
Long document classification (LDC) has been a focused interest in natural language processing (NLP) recently with the exponential increase of publications. Based on the pretrained language models, many LDC methods have been proposed and achieved considerable progression. However, most of the existing methods model long documents as sequences of text while omitting the document structure, thus limiting the capability of effectively representing long texts carrying structure information. To mitigate such limitation, we propose a novel hierarchical graph convolutional network (HGCN) for structured LDC in this article, in which a section graph network is proposed to model the macrostructure of a document and a word graph network with a decoupled graph convolutional block is designed to extract the fine-grained features of a document. In addition, an interaction strategy is proposed to integrate these two networks as a whole by propagating features between them. To verify the effectiveness of the proposed model, four structured long document datasets are constructed, and the extensive experiments conducted on these datasets and another unstructured dataset show that the proposed method outperforms the state-of-the-art related classification methods.
Tengfei Liu 0005, Yongli Hu, Boyue Wang, Junbin Gao
IEEE Trans. Neural Networks Learn. Syst.3
2022 Multi-graph convolutional clustering network
abstract
Abstract The relationship between objects can be described from different angles. Although multiple kinds of relationships make the connections between objects complex, they bring in more discriminative information for the clustering tasks. Therefore, how to effectively fuse multiple kinds of relationships becomes a critical problem. In this paper, we propose a novel Multi‐graph Convolutional Clustering Network which deeply explores the feature information of nodes and fuses the multiple kinds of relationships between nodes. Unlike most graph convolutional clustering methods that only exploit the single graph or directly fuse multiple graphs into a unified graph before the graph convolution operation, we firstly build multiple parallelled graph convolution layers for each graph to learn diverse data representations, which fully exploits different statistics information between graphs. Then, a designed multi‐graph attention module fuses above data representations and considers the importance of each graph. Besides, the proposed model completes the transition from single graph to multiple graphs, which reduces the dependence of the quality of the single graph and enhances the robustness to graphs. Experimental results verify that the proposed multi‐graph convolution clustering performs better than the traditional single‐graph convolution clustering.
Boyue Wang, Yifan Wang 0004, Xiaxia He, Yongli Hu
IET Signal Process.1
2022 Shareability-Exclusivity Representation on Product Grassmann Manifolds for Multi-camera video clustering
Yongli Hu, Cuicui Luo, Junbin Gao, Boyue Wang
J. Vis. Commun. Image Represent.4
2022 Text-to-Traffic Generative Adversarial Network for Traffic Situation Generation
abstract
Traffic situation generation is of importance in the intelligent transportation field, evaluating and simulating the macroscopic traffic conditions. The government often uses the historical traffic data on the same weekday to analyze the future traffic situations, which works unfavorably due to some traffic-related information deficiency, such as weather, location, traffic accidents, social activities and so on. Therefore, how to accurately generate the traffic situation is a challenging problem. Fortunately, massive traffic-related information spread in social media often indicates the traffic situation variation trend, which provides the sufficient information for the traffic situation generation. In this paper, we propose a novel Text-to-Traffic generative adversarial network framework ($\text{T}^{2}$GAN), which fuses the traffic data and the semantic information collected from social media to generate the traffic situation. To reduce the huge gap between the above two modalities and improve the authenticity of the generated traffic situation, we raise a global-local loss. Additionally, we build a heterogeneous dataset containing the traffic-related text data collected from social media and the corresponding traffic passenger flow data. Experimental results show that the proposed methods are obviously better than many outstanding traffic situation generation methods based on neural networks.
Guangyu Huo, Yong Zhang 0029, Boyue Wang, Yongli Hu
IEEE Trans. Intell. Transp. Syst.3
2022 Dual Dynamic Spatial-Temporal Graph Convolution Network for Traffic Prediction
abstract
Recently, Graph Convolution Network (GCN) and Temporal Convolution Network (TCN) are introduced into traffic prediction and achieve state-of-the-art performance due to their good ability for modeling the spatial and temporal property of traffic data. In spite of having good performance, the current methods generally focus on the traffic measurement of road segments, i.e. the nodes of traffic flow graph, while the edges of the graph, which represent the correlation of traffic data of different road segments and form the affinity matrix for GCN, are usually constructed according to the structure of road network, but the spatial and temporal properties are not well exploited in their theories. In this paper, we propose a Dual Dynamic Spatial-Temporal Graph Convolution Network (DDSTGCN), which not only models the dynamic property of the nodes of the traffic flow graph but also captures the dynamic spatial-temporal feature of the edges by transforming the traffic flow graph into its dual hypergraph. The traffic prediction is enhanced by the collaborative convolutions on the traffic flow graph and its dual hypergraph. The proposed method is evaluated by extensive traffic prediction experiments on six real road datasets and the results show that it outperforms state-of-the-art related methods. Source codes are available athttps://github.com/j1o2h3n/DDSTGCN.
Xiangheng Jiang, Yongli Hu, Fuqing Duan, Kan Guo, Boyue Wang, Junbin Gao
IEEE Trans. Intell. Transp. Syst.6
2021 Complete/incomplete multi-view subspace clustering via soft block-diagonal-induced regulariser
abstract
Abstract This study proposes a novel multi‐view soft block diagonal representation framework for clustering complete and incomplete multi‐view data. First, given that the multi‐view self‐representation model offers better performance in exploring the intrinsic structure of multi‐view data, it can be nicely adopted to individually construct a graph for each view. Second, since an ideal block diagonal graph is beneficial for clustering, a ‘soft’ block diagonal affinity matrix is constructed by fusing multiple previous graphs. The soft diagonal block regulariser encourages a matrix to approximately have (not exactly) diagonal blocks, where is the number of clusters. This strategy adds robustness to noise and outliers. Third, to handle incomplete multi‐view data, multiple indicator matrices are utilised, which can mark the position of missing elements of each view. Finally, the alternative direction of multipliers algorithm is employed to optimise the proposed model, and the corresponding algorithm complexity and convergence are also analysed. Extensive experimental results on several real‐world datasets achieve the best performance among the state‐of‐the‐art complete and incomplete clustering methods, which proves the effectiveness of the proposed methods.
Yongli Hu, Cuicui Luo, Boyue Wang, Junbin Gao
IET Comput. Vis.3
2021 TRFH: towards real-time face detection and head pose estimation
abstract
Abstract Nowadays, face detection and head pose estimation have a lot of application such as face recognition, aiding in gaze estimation and modeling attention. For these two tasks, it is usually to design two different models. However, the head pose estimation model often depends on the region of interest (ROI) detected in advance, which means that a serial face detector is needed. Even the lightest face detector will slow down the whole forward inference time and cannot achieve real-time performance when detecting the head pose of multiple people. We can see that both face detection and head pose estimation need face features, so a shared face feature map can be used between them. In this paper, a multi-task learning model is proposed that can solve both problems simultaneously. We directly detect the location of the center point of the bounding box of face; at this location, we calculate the size of the bounding box of face and the head attitude. We evaluate our model’s performance on the AFLW. The proposed model has great competitiveness with the multi-stage face attribute analysis model, and our model can achieve real-time performance.
Shicun Chen, Yong Zhang 0029, Boyue Wang
Pattern Anal. Appl.4
2021 AKM3C: Adaptive K-Multiple-Means for Multi-View Clustering
abstract
With the popularity of cameras and sensors, massive data are captured from various view angles or modalities, which provide abundant complementary information and also bring great challenges for traditional clustering methods. In this article, we propose a novel Adaptive K-Multiple-Means for multi-view clustering method (AKM3C). Unlike traditional multi-view K-means methods by grouping samples into$C$clusters each with a cluster center in every view, the proposed AKM3C employs$M (M>C)$sub-cluster centers in each view to reveal the sub-cluster structure in the multi-view data thus enhances the clustering performance. Additionally, to distinguish the importance of different views, instead of using empirical weights, AKM3C exploits the multi-view combination weights strategy to assign a weight to each view automatically and thus fuses the complementary information of different views properly to get an optimally shared bipartite graph, on which the Laplacian rank constraint is executed and the final clusters are obtained by directly partitioning. An efficient optimization algorithm proposed with complexity and convergence analysis is used to solve the proposed AKM3C method. The extensive experimental results on eight public datasets show that the proposed AKM3C performs better than state-of-the-art multi-view clustering methods. The code can be downloaded athttps://drive.google.com/file/d/1CQ0royrYxKFJdNLnbBQSbDrohtfH71di/view?usp=sharing.
Yongli Hu, Zuolong Song, Boyue Wang, Junbin Gao
IEEE Trans. Circuits Syst. Video Technol.3
2021 Robust Image Representation via Low Rank Locality Preserving Projection
abstract
Locality preserving projection (LPP) is a dimensionality reduction algorithm preserving the neighhorhood graph structure of data. However, the conventional LPP is sensitive to outliers existing in data. This article proposes a novel low-rank LPP model called LR-LPP. In this new model, original data are decomposed into the clean intrinsic component and noise component. Then the projective matrix is learned based on the clean intrinsic component which is encoded in low-rank features. The noise component is constrained by theℓ1-norm which is more robust to outliers. Finally, LR-LPP model is extended to LR-FLPP in which low-dimensional feature is measured by F-norm. LR-FLPP will reduce aggregated error and weaken the effect of outliers, which will make the proposed LR-FLPP even more robust for outliers. The experimental results on public image databases demonstrate the effectiveness of the proposed LR-LPP and LR-FLPP.
Junbin Gao, Yongli Hu, Boyue Wang
ACM Trans. Knowl. Discov. Data5
2021 Learning Adaptive Neighborhood Graph on Grassmann Manifolds for Video/Image-Set Subspace Clustering
abstract
The objective of self-expression based spectral clustering is to learn an affinity matrix which accurately reflects the similarity among data, and the Laplacian constraint is usually exploited to make the affinity matrix preserve the global structure of raw data. However, there exist two drawbacks: firstly, these methods are mostly designed for vectorial data in Euclidean spaces, which are not suitable for multidimensional data with nonlinear manifold structure, e.g., videos and image-sets. Secondly, the clustering performance heavily relies on the quality of a pre-learned Laplacian matrix in which the global structure may be mis-interpreted without considering manifold structures. In this paper, we firstly provide a unified framework about self-expression learning on Grassmann manifolds, which implements the clustering tasks for multidimensional data under subspace views. Then, to assign optimal neighbors to each data depending on the local distance, we adaptively learn the neighborhood relationship from the obtained self-expression coefficient matrix, referred to Learning Adaptive Neighborhood Graph on Grassmann manifolds (GMAN). In the optimization process, the neighborhood relationship can be adaptively learned and updated with the coefficient matrix. The experimental results on five public datasets show that the proposed method is obviously better than many related clustering methods based on Grassmann manifolds, proving the effectiveness of GMAN in multidimensional data clustering.
Boyue Wang, Yongli Hu, Junbin Gao, Fujiao Ju
IEEE Trans. Multim.1
2021 Adaptive Fusion of Heterogeneous Manifolds for Subspace Clustering
abstract
Multiview clustering (MVC) has recently received great interest due to its pleasing efficacy in combining the abundant and complementary information to improve clustering performance, which overcomes the drawbacks of view limitation existed in the standard single-view clustering. However, the existing MVC methods are mostly designed for vectorial data from linear spaces and, thus, are not suitable for multiple dimensional data with intrinsic nonlinear manifold structures, e.g., videos or image sets. Some works have introduced manifolds' representation methods of data into MVC and obtained considerable improvements, but how to fuse multiple manifolds efficiently for clustering is still a challenging problem. Particularly, for heterogeneous manifolds, it is an entirely new problem. In this article, we propose to represent the complicated multiviews' data as heterogeneous manifolds and a fusion framework of heterogeneous manifolds for clustering. Different from the empirical weighting methods, an adaptive fusion strategy is designed to weight the importance of different manifolds in a data-driven manner. In addition, the low-rank representation is generalized onto the fused heterogeneous manifolds to explore the low-dimensional subspace structures embedded in data for clustering. We assessed the proposed method on several public data sets, including human action video, facial image, and traffic scenario video. The experimental results show that our method obviously outperforms a number of state-of-the-art clustering methods.
Boyue Wang, Yongli Hu, Junbin Gao, Fujiao Ju
IEEE Trans. Neural Networks Learn. Syst.1
2020 Spatio-Temporal Memory Attention for Image Captioning
abstract
Visual attention has been successfully applied in image captioning to selectively incorporate the most relevant areas to the language generation procedure. However, the attention in current image captioning methods is only guided by the hidden state of language model, e.g. LSTM (Long-Short Term Memory), indirectly and implicitly, and thus the attended areas are weakly relevant at different time steps. Besides the spatial relationship of attention areas, the temporal relationship in attention is crucial for image captioning according to the attention transmission mechanism of human vision. In this paper, we propose a new spatio-temporal memory attention (STMA) model to learn the spatio-temporal relationship in attention for image captioning. The STMA introduces the memory mechanism to the attention model through a tailored LSTM, where the new cell is used to memorize and propagate the attention information, and the output gate is used to generate attention weights. The attention in STMA transmits with memory adaptively and dependently, which builds strong temporal connections of attentions and learns the spatio-temporal relationship of attended areas simultaneously. Besides, the proposed STMA is flexible to combine with attention-based image captioning frameworks. Experiments on MS COCO dataset demonstrate the superiority of the proposed STMA model in exploring the spatio-temporal relationship in attention and improving the current attention-based image captioning.
Junzhong Ji, Xiaodan Zhang 0003, Boyue Wang, Xinhang Song
IEEE Trans. Image Process.4
2018 Cascaded Low Rank and Sparse Representation on Grassmann Manifolds
abstract
Inspired by low rank representation and sparse subspace clustering acquiring success, ones attempt to simultaneously perform low rank and sparse constraints on the affinity matrix to improve the performance. However, it is just a trade-off between these two constraints. In this paper, we propose a novel Cascaded Low Rank and Sparse Representation (CLRSR) method for subspace clustering, which seeks the sparse expression on the former learned low rank latent representation. To make our proposed method suitable to multi-dimension or imageset data, we extend CLRSR onto Grassmann manifolds. An effective solution and its convergence analysis are also provided. The excellent experimental results demonstrate the proposed method is more robust than other state-of-the-art clustering methods on imageset data.
Boyue Wang, Yongli Hu, Junbin Gao
IJCAI1
2018 Low Rank Representation on SPD matrices with Log-Euclidean metric
Boyue Wang, Yongli Hu, Junbin Gao, Muhammad Ali 0005, David Tien
Pattern Recognit.1
2018 Localized LRR on Grassmann Manifold: An Extrinsic View
abstract
Subspace data representation has recently become a common practice in many computer vision tasks. Low-rank representation (LRR) is one of the most successful models for clustering vectorial data according to their subspace structures. This paper explores the possibility of extending LRR for subspace data on Grassmann manifold. Rather than directly embedding the Grassmann manifold into the symmetric matrix space, an extrinsic view is taken to build the self-representation in the local area of the tangent space at each Grassmannian point, resulting in a localized LRR method on Grassmann manifold. A novel algorithm for solving the proposed model is investigated and implemented. The performance of the new clustering algorithm is assessed through experiments on several real-world data sets including MNIST handwritten digits, ballet video clips, SKIG action clips, and DynTex++ data set and highway traffic video clips. The experimental results show that the new method outperforms a number of state-of-the-art clustering methods.
Boyue Wang, Yongli Hu, Junbin Gao
IEEE Trans. Circuits Syst. Video Technol.1
2018 Partial Sum Minimization of Singular Values Representation on Grassmann Manifolds
abstract
Clustering is one of the fundamental topics in data mining and pattern recognition. As a prospective clustering method, the subspace clustering has made considerable progress in recent researches, e.g., sparse subspace clustering (SSC) and low rank representation (LRR). However, most existing subspace clustering algorithms are designed for vectorial data from linear spaces, thus not suitable for high-dimensional data with intrinsic non-linear manifold structure. For high-dimensional or manifold data, few research pays attention to clustering problems. The purpose of clustering on manifolds tends to cluster manifold-valued data into several groups according to the mainfold-based similarity metric. This article proposes an extended LRR model for manifold-valued Grassmann data that incorporates prior knowledge by minimizing partial sum of singular values instead of the nuclear norm, namely Partial Sum minimization of Singular Values Representation (GPSSVR). The new model not only enforces the global structure of data in low rank, but also retains important information by minimizing only smaller singular values. To further maintain the local structures among Grassmann points, we also integrate the Laplacian penalty with GPSSVR. The proposed model and algorithms are assessed on a public human face dataset, some widely used human action video datasets and a real scenery dataset. The experimental results show that the proposed methods obviously outperform other state-of-the-art methods.
Boyue Wang, Yongli Hu, Junbin Gao
ACM Trans. Knowl. Discov. Data1
2017 Locality Preserving Projections for Grassmann manifold
abstract
Learning on Grassmann manifold has become popular in many computer vision tasks, with the strong capability to extract discriminative information for imagesets and videos. However, such learning algorithms particularly on high-dimensional Grassmann manifold always involve with significantly high computational cost, which seriously limits the applicability of learning on Grassmann manifold in more wide areas. In this research, we propose an unsupervised dimensionality reduction algorithm on Grassmann manifold based on the Locality Preserving Projections (LPP) criterion. LPP is a commonly used dimensionality reduction algorithm for vector-valued data, aiming to preserve local structure of data in the dimension-reduced space. The strategy is to construct a mapping from higher dimensional Grassmann manifold into the one in a relative low-dimensional with more discriminative capability. The proposed method can be optimized as a basic eigenvalue problem. The performance of our proposed method is assessed on several classification and clustering tasks and the experimental results show show its clear advantages over other Grassmann based algorithms.
Boyue Wang, Yongli Hu, Junbin Gao, Haoran Chen 0004, Muhammad Ali 0005
IJCAI1
2017 Laplacian LRR on Product Grassmann Manifolds for Human Activity Clustering in Multicamera Video Surveillance
abstract
In multicamera video surveillance, it is challenging to represent videos from different cameras properly and fuse them efficiently for specific applications such as human activity recognition and clustering. In this paper, a novel representation for multicamera video data, namely, the product Grassmann manifold (PGM), is proposed to model video sequences as points on the Grassmann manifold and integrate them as a whole in the product manifold form. In addition, with a new geometry metric on the product manifold, the conventional low rank representation (LRR) model is extended onto PGM and the new LRR model can be used for clustering nonlinear data, such as multicamera video data. To evaluate the proposed method, a number of clustering experiments are conducted on several multicamera video data sets of human activity, including the Dongzhimen Transport Hub Crowd action data set, the ACT 42 Human Action data set, and the SKIG action data set. The experiment results show that the proposed method outperforms many state-of-the-art clustering methods.
Boyue Wang, Yongli Hu, Junbin Gao
IEEE Trans. Circuits Syst. Video Technol.1
2016 Product Grassmann Manifold Representation and Its LRR Models
abstract
It is a challenging problem to cluster multi- and high-dimensional data with complex intrinsic properties and non-linear manifold structure. The recently proposed subspace clustering method, Low Rank Representation (LRR), shows attractive performance on data clustering, but it generally does with data in Euclidean spaces. In this paper, we intend to cluster complex high dimensional data with multiple varying factors. We propose a novel representation, namely Product Grassmann Manifold (PGM), to represent these data. Additionally, we discuss the geometry metric of the manifold and expand the conventional LRR model in Euclidean space onto PGM and thus construct a new LRR model. Several clustering experimental results show that the proposed method obtains superior accuracy compared with the clustering methods on manifolds or conventional Euclidean spaces.
Boyue Wang, Yongli Hu, Junbin Gao
AAAI1
2016 Low-rank representation based traffic data completion method
abstract
Intelligent Transportation Systems (ITS) plays a significant role in the traffic management, i.e. traffic jam prediction, route guidance. Due to the hardware failure or data transformation failure, some traffic observation data may be occasionally missed, which seriously affect intelligent transportation information service. So, the completion of traffic observation data has now become an issue that requires to be concerned and solved. By analyzing the traffic history data, we find that traffic data tend to have strong spatio-temporal correlation. Considering this feature, we propose a new low-rank representation based traffic data completion method. To further enhance the local correlation, we introduce an ordered regulation into our proposed method. We also give an efficient solution to our proposed methods. In order to verify the performance of our methods, some traffic data completion experiments are conducted on the Beijing metropolitan road speed dataset and the capital airport highway microwave dataset. Experimental results show that the proposed methods are superior to other state-of-the-art traffic data completion methods.
Yong Zhang 0029, Boyue Wang, Hao Liu 0040, Guanglei Qi
IJCNN3
2014 Low Rank Representation on Grassmann Manifolds
Boyue Wang, Yongli Hu, Junbin Gao
ACCV (1)1