Qiuyu Liang

dblp:372/1823 · DBLP profile ↗
← Back
13ranked-venue papers
8as first author
13since 2021 · last 2026
0009-0004-3261-7256ORCID · verified

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 10 · 6 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 5 · 4 first-author · 5 since 2021
YearPublicationVenuePosition
2026 Tensor Decomposition and Language Description for Open-Vocabulary Object Detection
abstract
Open-vocabulary object detection (OVOD) aims at detecting and recognizing objects beyond a fixed set of classes. Although region-word alignment and knowledge distillation have been explored for training a strong open-vocabulary detector, our analysis reveals three main issues (inaccurate alignment, redundant distillation, and low-quality class embedding) that limit OVOD's performance. In this paper, we explore the well-designed Tensor decomposition and Language descriptions for open-vocabulary object Detection (called TLDet). Proposals with the highest similarity score often correspond to discriminative but incomplete regions (e.g., object heads), resulting in inaccurate region-word alignment. To mitigate this issue, we propose a low-rank proposal filtering module that quantitatively assesses the completeness of each proposal by performing singular value decomposition and computing the sum of its singular values. This allows the model to reduce discriminative proposals and enhance the precision of alignment between visual regions and textual concepts. Furthermore, to mitigate redundant knowledge transfer, we introduce a core tensor distillation approach that decomposes teacher and student features into core tensors via Tucker decomposition and performs distillation through optimized tensor alignment. This ensures that the student acquires the most essential knowledge from the teacher. Finally, to improve the quality of class embedding, a language description enhancement method is proposed by exploring the knowledge of LLM to enrich the representations of categories during inference. Extensive experiments on popular datasets demonstrate the superior performance of our TLDet, achieving 36.1% mAP on COCO and 30.1% mask mAP on LVIS, and outperforming existing methods on novel categories.
Qiuyu Liang
AAAI1
2026 Plug-and-Play global and local collaborative fusion for weakly supervised object detection
abstract
• We propose a plug-and-play global and local collaborative fusion method to improve the performance of weakly supervised object detection. • We design a pixel-level global information awareness module that utilizes singular value decomposition for image reconstruction. • We propose a local detail fusion module to enable the visual encoder to learn detailed information about target objects. • We demonstrate the effectiveness and superiority of our plug-and-play method through extensive experiments. Weakly supervised object detection (WSOD) has drawn much attention due to its closeness to practical applications, and researchers have proposed the multi-instance learning (MIL) approach to handle it as a multi-class classification problem. Although these methods have yielded promising results, extraneous information in the images severely affects the model’s feature learning due to the lack of instance-level annotation. To alleviate this limitation, in this paper, a global and local collaborative fusion method is proposed for WSOD by leveraging the complementary information of the original image and its low-rank approximation. Specifically, we design a pixel-level global information awareness (GIA) module to reconstruct the input image and remove redundant noise, which are then fed into a visual encoder to extract the features from a global perspective. Moreover, to compensate for the lack of detail preservation in GIA, we further propose a local detail fusion (LDF) module that fuses image details by leveraging both reconstructed and input images. Our proposed GIA-LDF modules are architecture-agnostic and can be seamlessly embedded into any MIL-based WSOD pipeline. Extensive experiments validate the effectiveness of our plug-and-play GIA-LDF for WSOD. We achieve 60.2%, 57.4%, and 23.2% mAP on PASCAL VOC 2007, VOC 2012, and COCO, respectively, surpassing baseline methods by +2.0%, +1.2%, and +0.3%, and establishing new state-of-the-art performance across all benchmarks.
Qiuyu Liang, Yongqiang Zhang 0007
Knowl. Based Syst.1
2025 Distance-Adaptive Quaternion Knowledge Graph Embedding with Bidirectional Rotation
abstract
Quaternion contains one real part and three imaginary parts, which provided a more expressive hypercomplex space for learning knowledge graph. Existing quaternion embedding models measure the plausibility of a triplet either through semantic matching or distance scoring functions. However, it appears that semantic matching diminishes the separability of entities, while the distance scoring function weakens the semantics of entities. To address this issue, we propose a novel quaternion knowledge graph embedding model. Our model combines semantic matching with entity’s geometric distance to better measure the plausibility of triplets. Specifically, in the quaternion space, we perform a right rotation on the head entity and a reverse rotation on the tail entity to learn the rich semantic features. Then, we utilize distance adaptive translations to learn the geometric distance between entities. Furthermore, we provide mathematical proofs to demonstrate our model can handle complex logical relationships. Extensive experimental results and analyses show our model significantly outperforms previous models on well-known knowledge graph completion benchmark datasets. Our code is available at https://anonymous.4open.science/r/l2730.
Weihua Wang 0006, Qiuyu Liang, Feilong Bao, Guanglai Gao
COLING2
2025 Unifying Dual-Space Embedding for Entity Alignment via Contrastive Learning
abstract
Entity alignment (EA) aims to match identical entities across different knowledge graphs (KGs). Graph neural network-based entity alignment methods have achieved promising results in Euclidean space. However, KGs often contain complex local and hierarchical structures, which are hard to represent in a single space. In this paper, we propose a novel method named as UniEA, which unifies dual-space embedding to preserve the intrinsic structure of KGs. Specifically, we simultaneously learn graph structure embeddings in both Euclidean and hyperbolic spaces to maximize the consistency between embeddings in the two spaces. Moreover, we employ contrastive learning to mitigate the misalignment issues caused by similar entities, where embeddings of similar neighboring entities become too close. Extensive experiments on benchmark datasets demonstrate that our method achieves state-of-the-art performance in structure-based EA. Our code is available at https://github.com/wonderCS1213/UniEA.
Cunda Wang, Weihua Wang 0006, Qiuyu Liang, Feilong Bao, Guanglai Gao
COLING3
2025 Hyperbolic Multimodal Knowledge Graph Embedding
abstract
Multimodal knowledge graph embedding refers to learning multimodal entities and their relation representations in a low-dimensional space. However, existing multimodal embedding models tend to ignore the inherent structure of knowledge graphs. To address this issue, we propose a novel multimodal knowledge graph embedding model to simultaneously learn semantic relation and hierarchical structure of entities within a hyperbolic space. Specifically, we project all modalities features embedding into a hyperbolic space and unify these embeddings to form a multimodal embedding. Then, we model the knowledge graph triplets by treating the relation as a Lorentzian linear transformation from head entity to tail entity. The plausibility of triplets is measured by Lorentz distance. Extensive experiments on multimodal knowledge graph completion benchmarks validate that our model achieves the state-of-the-art results across most metrics. In terms of training speed, our model is one order of magnitude faster than the best one. The visualization results further reveal our model’s ability to capture hierarchical structures. Our code is available at https://github.com/llqy123/HyME.
Qiuyu Liang, Weihua Wang 0006, Cunda Wang, Feilong Bao, Jie Yu 0008
ICASSP1
2025 OTMEA : Multi-modal Entity Alignment via Optimal Transport
abstract
Multi-modal Entity Alignment (MMEA) aims to identify the same entities exhibited in different knowledge graphs (KGs), where the entities are enriched by structure and visual information. Existing MMEA methods learn multi-modal joint entity embeddings by encompassing both modality interaction and modality alignment. However, these approaches predominantly emphasize modality interaction and fail to adequately address the issue of modality heterogeneity. In this paper, we propose a novel approach OTMEA, which leverages optimal transport to mitigate modality heterogeneity from the perspective of modality distributions. Specifically, we view the modality alignment problem as a Wasserstein minimum distance problem involving multimodal distributions. Furthermore, our experiments indicate that employing entity-level attention weights significantly enhances modality alignment through optimal transport. The effectiveness of our method is validated through extensive experiments conducted on five public datasets. The source code is available at https://github.com/wonderCS1213/OTMEA.
Cunda Wang, Weihua Wang 0006, Qiuyu Liang, Feilong Bao
ICASSP4
2025 SAM based Region-Word Clustering and Inference Score Adjusting for Open-Vocabulary Object Detection
Qiuyu Liang, Yongqiang Zhang 0007
ACM Multimedia1
2025 Local and global structure-aware contrastive framework for entity alignment
Cunda Wang, Weihua Wang 0006, Qiuyu Liang, Guanglai Gao
Neurocomputing3
2024 L\²GC: Lorentzian Linear Graph Convolutional Networks for Node Classification
Qiuyu Liang, Weihua Wang 0006, Feilong Bao, Guanglai Gao
LREC/COLING1
2024 Fully Hyperbolic Rotation for Knowledge Graph Embedding
abstract
Hyperbolic rotation is commonly used to effectively model knowledge graphs and their inherent hierarchies. However, existing hyperbolic rotation models rely on logarithmic and exponential mappings for feature transformation. These models only project data features into hyperbolic space for rotation, limiting their ability to fully exploit the hyperbolic space. To address this problem, we propose a novel fully hyperbolic model designed for knowledge graph embedding. Instead of feature mappings, we define the model directly in hyperbolic space with the Lorentz model. Our model considers each relation in knowledge graphs as a Lorentz rotation from the head entity to the tail entity. We adopt the Lorentzian version distance as the scoring function for measuring the plausibility of triplets. Extensive results on standard knowledge graph completion benchmarks demonstrated that our model achieves competitive results with fewer parameters. In addition, our model get the state-of-the-art performance on datasets of CoDEx-s and CoDEx-m, which are more diverse and challenging than before. Our code is available at https://github.com/llqy123/FHRE.
Qiuyu Liang, Weihua Wang 0006, Feilong Bao, Guanglai Gao
ECAI1
2024 Hierarchy-Aware Quaternion Embedding for Knowledge Graph Completion
abstract
Knowledge graph completion is an essential task in the fields of graph mining and graph machine learning. Most contemporary approaches rely on geometric transformation to achieve knowledge graph completion, as geometry offers a well-defined mathematical foundation. For example, rotation transformations in rigid body transformation are frequently employed within quaternion spaces to model complex relation types in knowledge graphs. However, these models cannot effectively handle the hierarchical structure in the knowledge graph. As a result, the performance of knowledge graph completion suffers. To address this shortcoming of quaternion space, we propose a novel model that integrates hyperbolic space. Specifically, we perform a translation transformation in a hyperbolic space to obtain support vector embeddings that imply relation embedding. We then perform a rotation transformation with the Hamilton product in tangent space, treating the relation embedding as a rotation from the head entity embedding to the tail entity embedding. We verify the validity and generalization ability of our model on standard benchmark datasets including WN18RR, FB15k-237 and YAGO3-10. The experimental results show that our model achieves competitive results on MRR and H@K metrics. Our code is publicly available at https://github.com/llqy123/HAQE-master.
Qiuyu Liang, Weihua Wang 0006, Jie Yu 0008, Feilong Bao
IJCNN1
2024 Effective Knowledge Graph Embedding with Quaternion Convolutional Networks
Qiuyu Liang, Weihua Wang 0006, Jie Yu 0008, Feilong Bao
NLPCC (3)1
2024 GSEA: Global Structure-Aware Graph Neural Networks for Entity Alignment
Cunda Wang, Weihua Wang 0006, Qiuyu Liang, Jie Yu 0008, Guanglai Gao
NLPCC (2)3