Weihua Wang 0006

dblp:50/2852-6 · DBLP profile ↗
← Back
26ranked-venue papers
7as first author
20since 2021 · last 2026
0000-0002-3234-7151ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 18 · 6 first-author · 13 since 2021Graphics, computer vision, multimedia, augmented reality and games · 6 · 5 since 2021Databases, data management, data science and information retrieval · 3 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
YearPublicationVenuePosition
2026 CEDAR: A Chinese Evaluation Dataset for Computational Argumentation
abstract
Tian Lan, Jiang Li, Rong Yan, Feilong Bao, Weihua Wang, Guanglai Gao, Xiangdong Su. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jiang Li 0013, Feilong Bao, Weihua Wang 0006, Guanglai Gao, Xiangdong Su
ACL (1)5
2026 Selective Distillation for Continual Named Entity Recognition with Memory Replay
Weihua Wang 0006, Feilong Bao
SIGIR1
2026 How to teach and forget: Towards cross-modal semantic consistency for entity alignment
Cunda Wang, Chenglong Miao, Po Hu 0001, Weihua Wang 0006, Feilong Bao
Inf. Process. Manag.4
2026 Wisteria: A unified multi-scale feature learning framework for DNA language model
Weihua Wang 0006, Haoji Li, Feilong Bao, Guanglai Gao
Pattern Recognit.1
2025 An End-to-End Framework for Reconstructing Continuous Language from fMRI Data
abstract
Traditional approaches to decoding brain fMRI sequences typically learn spatial and temporal features in separate stages to reconstruct language. However, the high dimensionality and low temporal resolution of fMRI signals hinder the effective extraction of semantic features. In this paper, we propose a novel end-to-end framework that jointly models high-dimensional spatiotemporal dynamics. Specifically, a separable spatiotemporal convolution is introduced to capture spatial and temporal dependencies, while a Hilbert curve-based dimensionality reduction preserves local spatial topology. A pretrained language model is then used to decode continuous text from the resulting representations. To prevent data leakage, both subjects and text stimuli are strictly separated between training and evaluation sets. Experimental results demonstrate that our method effectively captures spatiotemporal correlations, maintains spatial locality, and achieves significantly better performance than baseline models.
Changbin Lv, Weihua Wang 0006
BIBM3
2025 Distance-Adaptive Quaternion Knowledge Graph Embedding with Bidirectional Rotation
abstract
Quaternion contains one real part and three imaginary parts, which provided a more expressive hypercomplex space for learning knowledge graph. Existing quaternion embedding models measure the plausibility of a triplet either through semantic matching or distance scoring functions. However, it appears that semantic matching diminishes the separability of entities, while the distance scoring function weakens the semantics of entities. To address this issue, we propose a novel quaternion knowledge graph embedding model. Our model combines semantic matching with entity’s geometric distance to better measure the plausibility of triplets. Specifically, in the quaternion space, we perform a right rotation on the head entity and a reverse rotation on the tail entity to learn the rich semantic features. Then, we utilize distance adaptive translations to learn the geometric distance between entities. Furthermore, we provide mathematical proofs to demonstrate our model can handle complex logical relationships. Extensive experimental results and analyses show our model significantly outperforms previous models on well-known knowledge graph completion benchmark datasets. Our code is available at https://anonymous.4open.science/r/l2730.
Weihua Wang 0006, Qiuyu Liang, Feilong Bao, Guanglai Gao
COLING1
2025 Unifying Dual-Space Embedding for Entity Alignment via Contrastive Learning
abstract
Entity alignment (EA) aims to match identical entities across different knowledge graphs (KGs). Graph neural network-based entity alignment methods have achieved promising results in Euclidean space. However, KGs often contain complex local and hierarchical structures, which are hard to represent in a single space. In this paper, we propose a novel method named as UniEA, which unifies dual-space embedding to preserve the intrinsic structure of KGs. Specifically, we simultaneously learn graph structure embeddings in both Euclidean and hyperbolic spaces to maximize the consistency between embeddings in the two spaces. Moreover, we employ contrastive learning to mitigate the misalignment issues caused by similar entities, where embeddings of similar neighboring entities become too close. Extensive experiments on benchmark datasets demonstrate that our method achieves state-of-the-art performance in structure-based EA. Our code is available at https://github.com/wonderCS1213/UniEA.
Cunda Wang, Weihua Wang 0006, Qiuyu Liang, Feilong Bao, Guanglai Gao
COLING2
2025 Hyperbolic Multimodal Knowledge Graph Embedding
abstract
Multimodal knowledge graph embedding refers to learning multimodal entities and their relation representations in a low-dimensional space. However, existing multimodal embedding models tend to ignore the inherent structure of knowledge graphs. To address this issue, we propose a novel multimodal knowledge graph embedding model to simultaneously learn semantic relation and hierarchical structure of entities within a hyperbolic space. Specifically, we project all modalities features embedding into a hyperbolic space and unify these embeddings to form a multimodal embedding. Then, we model the knowledge graph triplets by treating the relation as a Lorentzian linear transformation from head entity to tail entity. The plausibility of triplets is measured by Lorentz distance. Extensive experiments on multimodal knowledge graph completion benchmarks validate that our model achieves the state-of-the-art results across most metrics. In terms of training speed, our model is one order of magnitude faster than the best one. The visualization results further reveal our model’s ability to capture hierarchical structures. Our code is available at https://github.com/llqy123/HyME.
Qiuyu Liang, Weihua Wang 0006, Cunda Wang, Feilong Bao, Jie Yu 0008
ICASSP2
2025 Dynamic Structure Hypergraph for Document-level Event Extraction
abstract
Document-level Event Extraction (DEE) aims to identify event information from a given document. The two challenges of this task are the event arguments scattering across differrent sentences and the multiple events within a single document. In this paper, we propose a novel Dynamic Structure Hypergraph model to address the issue of limited global modeling capability in traditional graphs. Firstly, we construct a hypergraph to model the global interactions between different sentences and entities in a document. Then, new hyperedges are generated by constructing a mention-mention correlation matrix based on the updated node representations, which evolves the hypergraph into a dynamic structure. This will help the nodes to aware the contextual semantic information in time. Finally, extensive experiments and analysis demonstrate that our method has made significant improvements in addressing the two aforementioned challenges, which outperforms existing state-of-the-art models on two public datasets. Our code is available at https://github.com/1999rq/DSH.
Qi Ren, Weihua Wang 0006, Jie Yu 0008, Guanglai Gao
ICASSP2
2025 OTMEA : Multi-modal Entity Alignment via Optimal Transport
abstract
Multi-modal Entity Alignment (MMEA) aims to identify the same entities exhibited in different knowledge graphs (KGs), where the entities are enriched by structure and visual information. Existing MMEA methods learn multi-modal joint entity embeddings by encompassing both modality interaction and modality alignment. However, these approaches predominantly emphasize modality interaction and fail to adequately address the issue of modality heterogeneity. In this paper, we propose a novel approach OTMEA, which leverages optimal transport to mitigate modality heterogeneity from the perspective of modality distributions. Specifically, we view the modality alignment problem as a Wasserstein minimum distance problem involving multimodal distributions. Furthermore, our experiments indicate that employing entity-level attention weights significantly enhances modality alignment through optimal transport. The effectiveness of our method is validated through extensive experiments conducted on five public datasets. The source code is available at https://github.com/wonderCS1213/OTMEA.
Cunda Wang, Weihua Wang 0006, Qiuyu Liang, Feilong Bao
ICASSP2
2025 Semantic Data Augmentation for Few-Shot Biomedical Named Entity Recognition
abstract
Biomedical Named Entity Recognition (BioNER) aims to identify and classify entities in biomedical text. This task struggles with data scarcity due to limited annotated data. Although data augmentation is effective, existing methods fail to handle the complex semantic mappings on biomedical terminology, which results in the generated samples lacking semantic diversity. In this paper, we propose a data augmentation method to create semantically diverse and coherent training samples for few-shot BioNER. The method utilizes the Unified Medical Language System (UMLS) for entity replacement and combines with a Masked Language Model (MLM) to generate contextually relevant words. Experimental results show that our method improves BioNER performance in few-shot scenarios. Compared to all baseline models, our method achieves an average F1 score improvement of 7.6% and 11.6% on the NCBI and JNLPBA datasets, respectively. Our source code and data are available at https://github.com/ABC8184/SDA
Weihua Wang 0006
ICASSP2
2025 Local and global structure-aware contrastive framework for entity alignment
Cunda Wang, Weihua Wang 0006, Qiuyu Liang, Guanglai Gao
Neurocomputing2
2024 L\²GC: Lorentzian Linear Graph Convolutional Networks for Node Classification
Qiuyu Liang, Weihua Wang 0006, Feilong Bao, Guanglai Gao
LREC/COLING2
2024 Fully Hyperbolic Rotation for Knowledge Graph Embedding
abstract
Hyperbolic rotation is commonly used to effectively model knowledge graphs and their inherent hierarchies. However, existing hyperbolic rotation models rely on logarithmic and exponential mappings for feature transformation. These models only project data features into hyperbolic space for rotation, limiting their ability to fully exploit the hyperbolic space. To address this problem, we propose a novel fully hyperbolic model designed for knowledge graph embedding. Instead of feature mappings, we define the model directly in hyperbolic space with the Lorentz model. Our model considers each relation in knowledge graphs as a Lorentz rotation from the head entity to the tail entity. We adopt the Lorentzian version distance as the scoring function for measuring the plausibility of triplets. Extensive results on standard knowledge graph completion benchmarks demonstrated that our model achieves competitive results with fewer parameters. In addition, our model get the state-of-the-art performance on datasets of CoDEx-s and CoDEx-m, which are more diverse and challenging than before. Our code is available at https://github.com/llqy123/FHRE.
Qiuyu Liang, Weihua Wang 0006, Feilong Bao, Guanglai Gao
ECAI2
2024 Hierarchy-Aware Quaternion Embedding for Knowledge Graph Completion
abstract
Knowledge graph completion is an essential task in the fields of graph mining and graph machine learning. Most contemporary approaches rely on geometric transformation to achieve knowledge graph completion, as geometry offers a well-defined mathematical foundation. For example, rotation transformations in rigid body transformation are frequently employed within quaternion spaces to model complex relation types in knowledge graphs. However, these models cannot effectively handle the hierarchical structure in the knowledge graph. As a result, the performance of knowledge graph completion suffers. To address this shortcoming of quaternion space, we propose a novel model that integrates hyperbolic space. Specifically, we perform a translation transformation in a hyperbolic space to obtain support vector embeddings that imply relation embedding. We then perform a rotation transformation with the Hamilton product in tangent space, treating the relation embedding as a rotation from the head entity embedding to the tail entity embedding. We verify the validity and generalization ability of our model on standard benchmark datasets including WN18RR, FB15k-237 and YAGO3-10. The experimental results show that our model achieves competitive results on MRR and H@K metrics. Our code is publicly available at https://github.com/llqy123/HAQE-master.
Qiuyu Liang, Weihua Wang 0006, Jie Yu 0008, Feilong Bao
IJCNN2
2024 Effective Knowledge Graph Embedding with Quaternion Convolutional Networks
Qiuyu Liang, Weihua Wang 0006, Jie Yu 0008, Feilong Bao
NLPCC (3)2
2024 GSEA: Global Structure-Aware Graph Neural Networks for Entity Alignment
Cunda Wang, Weihua Wang 0006, Qiuyu Liang, Jie Yu 0008, Guanglai Gao
NLPCC (2)2
2023 TableSF: A Structural Bias Framework for Table-To-Text Generation
Weihua Wang 0006, Feilong Bao, Guanglai Gao
ICANN (9)2
2023 Few-Shot Table-to-Text Generation with Structural Bias Attention
Weihua Wang 0006, Feilong Bao, Guanglai Gao
PRICAI (2)2
2022 End-to-End Large-Scale Image Retrieval Network with Convolution and Vision Transformers
Feilong Bao, Xiangdong Su, Weihua Wang 0006, Guanglai Gao
ICANN (4)4
2019 Neural Morphological Segmentation Model for Mongolian
abstract
Morphological segmentation is useful for processing Mongolian. In this paper, we manually build a morphological segmentation data set for Mongolian. We then present a character-based encoder-decoder model with attention mechanism to perform the morphological segmentation task. We further investigate the influence of analogy features extracted from scratch and improve the performance of our model using multi languages setting. Experimental results show that our encoder-decoder model with attention mechanism provides a strong baseline for Mongolian morphological segmentation. The analogy features provide useful information to the model and improve the performance of the system. The use of multi languages data set shows the capability of our model to acquire knowledge through different languages and delivers the best result.
Weihua Wang 0006, Rashel Fam, Feilong Bao, Yves Lepage, Guanglai Gao
IJCNN1
2019 An Automatic Spelling Correction Method for Classical Mongolian
Feilong Bao, Guanglai Gao, Weihua Wang 0006, Hui Zhang 0031
KSEM (2)4
2019 Learning Morpheme Representation for Mongolian Named Entity Recognition
Weihua Wang 0006, Feilong Bao, Guanglai Gao
Neural Process. Lett.1
2016 Mongolian Named Entity Recognition System with Rich Features
abstract
In this paper, we first build a manually annotated named entity corpus of Mongolian. Then, we propose three morphological processing methods and study comprehensive features, including syllable features, lexical features, context features, morphological features and semantic features in Mongolian named entity recognition. Moreover, we also evaluate the influence of word cluster features on the system and combine all features together eventually. The experimental result shows that segmenting each suffix into an individual token achieves better results than deleting suffixes or using the suffixes as feature. The system based on segmenting suffixes with all proposed features yields benchmark result of F-measure=84.65 on this corpus.
Weihua Wang 0006, Feilong Bao, Guanglai Gao
COLING1
2016 Mongolian Named Entity Recognition with Bidirectional Recurrent Neural Networks
abstract
Traditional approaches to Named Entity Recognition almost heavily rely on feature engineering. In this paper, we introduce a kind of bidirectional recurrent neural network with long short memory (BLSTM) to capture bidirectional and long dependencies in a sentence without any feature set. Our model combines BLSTM network with Conditional Random Field (CRF) layer to jointly decode the best output. Additionally, this model inputs the concatenation of Mongolian morpheme and character representation. Experimental results show that the bidirectional recurrent neural networks significantly outperform traditional CRF model using manual features.
Weihua Wang 0006, Feilong Bao, Guanglai Gao
ICTAI1
2013 Segmentation-based Mongolian LVCSR approach
abstract
Mongolian is an agglutinative language. Each root can be followed by several suffixes to formulate new words. This special word formation characteristic results in probably millions of Mongolian words, which is far beyond the coverage of the pronunciation dictionary of any current Mongolian speech recognition system. Moreover, even if the pronunciation dictionary is large enough to cover all of the Mongolian words, the recognition system still cannot perform well due to the problem of sample sparseness. In this paper, we propose a segmentation-based Mongolian Large Vocabulary Continuous Speech Recognition (LVCSR) approach and rebuild the corresponding acoustic model and language model. Experimental results show that, by converting most of these words into their corresponding In-Vocabulary form, the proposed approach effectively recognizes most of the Mongolian words and greatly improves the sample sparseness problem in the language model.
Feilong Bao, Guanglai Gao, Xueliang Yan, Weihua Wang 0006
ICASSP4