EDBT 2026 Demo / reviewers in the wild / expert
Gang Liu 0021
dblp:37/2109-21
· DBLP profile ↗
14ranked-venue papers
8as first author
12since 2021 · last 2026
0000-0002-7032-8429ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 6 · 3 first-author · 5 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 4 since 2021Computer networks · 3 · 3 first-author · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Next Generation Active Learning: Mixture of LLMs in the LoopabstractWith the rapid advancement and strong generalization capabilities of large language models (LLMs), they have been increasingly incorporated into the active learning pipelines as annotators to reduce annotation costs. However, considering the annotation quality, labels generated by LLMs often fall short of real-world applicability. To address this, we propose a novel active learning framework, Mixture of LLMs in the Loop Active Learning, replacing human annotators with labels generated through a Mixture-of-LLMs-based annotation model, aimed at enhancing LLM-based annotation robustness by aggregating the strengths of multiple LLMs. To further mitigate the impact of the noisy labels, we introduce annotation discrepancy and negative learning to identify the unreliable annotations and enhance learning effectiveness. Extensive experiments demonstrate that our framework achieves performance comparable to human annotation and consistently outperforms single-LLM baselines and other LLM-ensemble-based approaches. Moreover, our framework is built on lightweight LLMs, enabling it to operate fully on local machines in real-world applications. Xiaohao Yang, Jueqing Lu, Guoxiang Guo, Joanne Enticott, Gang Liu 0021, Lan Du 0002 |
AAAI | 6 |
| 2026 | MMG-RAG: Multimodal Graph RAG for Medical Report GenerationabstractLarge language models (LLMs) hold great promise for automating radiology report generation. However, these models often suffer from factual hallucinations, which can lead to inaccurate or misleading diagnoses. Although retrieval-augmented generation (RAG) can mitigate this issue by retrieving supplementary knowledge, text-only RAG in multimodal settings may provide incorrect guidance due to the modality mismatch between images and text. In this paper, we propose MMG-RAG, a multimodal graph-based RAG framework for radiology report generation. MMG-RAG improves the factual accuracy of LLMs by retrieving strongly correlated multimodal knowledge. Specifically, it constructs a multimodal knowledge graph that integrates image, text, and entity embeddings. We further design a dense entity retrieval mechanism that retrieves diagnostically relevant multimodal knowledge based on visual similarity and dense entity connections. Finally, the retrieved multimodal knowledge guides report generation through chain-of-thought reasoning. Experiments on two benchmark datasets, MIMIC-CXR and CheXpert, demonstrate that MMG-RAG significantly improves diagnostic accuracy and enhances the clinical quality of generated reports. Gang Liu 0021, Xiaotian Tang, Jiacheng Gan, Tingyao Liu, Shenjun Zhong |
ICMR | 1 |
| 2025 | Corpus Fusion and Text Summarization Extraction for Multi-Feature Enhanced Entity AlignmentabstractCross-lingual entity alignment endeavors to identify semantically similar entities within a knowledge graph, facilitating knowledge complementarity and enriching cross-lingual knowledge. In the context of knowledge-driven tasks such as cross-lingual question answering and knowledge recommendation, cross-lingual entity alignment can effectively enhancing the performance of these applications built upon cross-lingual knowledge graphs. However, the current methodologies exhibit constraints in efficiently extracting and combining features of multiple entities, rendering them unable to fully harness the wealth of extensive information provided by the knowledge graph. To address this challenge, we propose CFSE, a novel multi-feature enhanced fusion model, which includes deep extraction of complex entity relationship, name, and attribute features. Complex entity relationship features are extracted based on corpus fusion and RotatE model. Additionally, an algorithm based on BERT for multilingual text summarization was introduced to extract entity name and attribute features. Through comprehensive entity feature extraction, CFSE not only further improves the alignment accuracy, but also helps to maximize the depth mining of knowledge graph information. The effectiveness of CFSE in cross-lingual entity alignment applications was demonstrated through experimental results on the DBP15K dataset. Gang Liu 0021, Wenli Yang 0003, Tongli Wang, He Zhihao |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |
| 2024 | Robust Representation Learning for Image Clustering
Pengcheng Jiang, Ye Zhu 0002, Yang Cao 0019, Gang Li 0009, Gang Liu 0021, Bo Yang 0002 |
KSEM (4) | 5 |
| 2024 | DSS-DocRE: A DenseNet-Based Semantic Segmentation Model for Document-Level Relation ExtractionabstractDocument-level relation extraction aims to discover inter-sentence relationships within articles through sophisticated reasoning. This task poses three key challenges: (i) recognizing complex entity combinations (one-to-many, many-to-one, many-to-many); (ii) handling entities appearing multiple times across documents, necessitating context-aware embeddings; (iii) focusing on evidence sentences for intricate relational reasoning. In response, we propose DSS-DocRE, a model based on DenseNet. To address the issue of repeated entities, we employ local context pooling, enhancing entity embeddings with additional context relevant to the entity pair. Furthermore, we construct an evidence classifier sub-task model to utilize evidence sentences for relational reasoning, integrating it into a multi-task learning framework for joint training. We also employ a dynamic class loss, replacing the global threshold with a learnable threshold class set, improving the loss function to balance the weights of common and uncommon classes. Finally, noting the current scarcity of document-level relation extraction corpora, a distilled model is used to reason over distant supervision data, enhancing model generalization. Through detailed experiments, we demonstrate that the DSS-DocRE model achieves better performance on the public datasets DocRED and ReDocRED. Additionally, experimental analysis verifies the relation segmentation network's efficacy in document-level relation extraction and its superior interpretability. Gang Liu 0021, Dongze Zhang, Tongli Wang, Shenjun Zhong |
MSN | 1 |
| 2024 | Cross-Modal self-supervised vision language pre-training with multiple objectives for medical visual question answering
Gang Liu 0021, Shenjun Zhong |
J. Biomed. Informatics | 1 |
| 2024 | Kernel-based iVAT with adaptive cluster extractionabstractAbstract Visual Assessment of cluster Tendency (VAT) is a popular method that visually represents the possible clusters found in a dataset as dark blocks along the diagonal of a reordered dissimilarity image (RDI). Although many variants of the VAT algorithm have been proposed to improve the visualisation quality on different types of datasets, they still suffer from the challenge of extracting clusters with varied densities. In this paper, we focus on overcoming this drawback of VAT algorithms by incorporating kernel methods and also propose a novel adaptive cluster extraction strategy, named CER, to effectively identify the local clusters from the RDI. We examine their effects on an improved VAT method (iVAT) and systematically evaluate the clustering performance on 18 synthetic and real-world datasets. The experimental results reveal that the recently proposed data-dependent dissimilarity measure, namely the Isolation kernel, helps to significantly improve the RDI image for easy cluster identification. Furthermore, the proposed cluster extraction method, CER, outperforms other existing methods on most of the datasets in terms of a series of dissimilarity measures. Baojie Zhang, Ye Zhu 0002, Yang Cao 0019, Sutharshan Rajasegarar, Gang Li 0009, Gang Liu 0021 |
Knowl. Inf. Syst. | 6 |
| 2023 | MTLAN: Multi-Task Learning and Auxiliary Network for Enhanced Sentence Embedding
Gang Liu 0021, Tongli Wang, Wenli Yang 0003, Zhizheng Yan, Kai Zhan |
ICONIP (3) | 1 |
| 2023 | MFG-R: Chinese Text Matching with Multi-Information Fusion Graph Embedding and Residual ConnectionsabstractChinese text matching is an important task in natural language processing research, but the current techniques have problems in text feature extraction, such as insufficient word information extraction and lack of deep information in graph convolution networks. In this paper, we propose a model MFG-R for Chinese text matching with multi-information fusion graph embedding and residual connection. The model fuses the word embedding representation of the text obtained by graph convolution network with character-level information and word weight information to extract text features. At the same time, in order to perform deep interaction matching, we construct a word-level similarity interaction matrix between text pairs, and build a text interaction and feature extraction model based on residual network on this basis. Experiments show that MFG-R has excellent performance on two common Chinese datasets, Ant Financial Question Matching Corpus(AFQMC) and Large-scale Chinese Question Matching Corpus(LCQMC). Gang Liu 0021, Tongli Wang, Yichao Dong, Kai Zhan, Wenli Yang 0003 |
ISCC | 1 |
| 2023 | An Improved Visual Assessment with Data-Dependent Kernel for Stream Clustering
Baojie Zhang, Yang Cao 0019, Ye Zhu 0002, Sutharshan Rajasegarar, Gang Liu 0021, Hong Xian Li, Maia Angelova, Gang Li 0009 |
PAKDD (1) | 5 |
| 2021 | Cross-Language Plagiarism Detection Model Based On Multiple FeaturesabstractAs information sharing becomes more and more convenient, a lot of phenomena of plagiarism shows up. The study of cross-language plagiarism is an important problem that the whole academic circle tries to solve it collectively. In this paper, a multiple-features based cross-language plagiarism detection model is proposed, which includes cross-language plagiarism candidate retrieval based on multiple features and cross-language plagiarism detection based on dynamic text alignment. For cross-language plagiarism candidate retrieval, it is mainly based on the translation features. What's more, for cross-language plagiarism detection, a text-alignment based similarity analysis was used to filter the final results between the identified paragraphs. In this step, our approach doesn't use a machine translation system to convert longer text, but uses a dictionary to obtain the translation of a single word. Moreover, experimental results show that our method outperforms the previous methods and achieved the best results in four datasets. Gang Liu 0021, Yichao Dong, Guangxi Li |
ISCC | 1 |
| 2021 | Large-area damage image restoration algorithm based on generative adversarial network
Gang Liu 0021, Xiaofeng Li 0012 |
Neural Comput. Appl. | 1 |
| 2019 | Leveraging external information in topic modelling
He Zhao 0001, Lan Du 0002, Wray L. Buntine, Gang Liu 0021 |
Knowl. Inf. Syst. | 4 |
| 2017 | MetaLDA: A Topic Model that Efficiently Incorporates Meta InformationabstractBesides the text content, documents and their associated words usually come with rich sets of meta information, such as categories of documents and semantic/syntactic features of words, like those encoded in word embeddings. Incorporating such meta information directly into the generative process of topic models can improve modelling accuracy and topic quality, especially in the case where the word-occurrence information in the training data is insufficient. In this paper, we present a topic model, called MetaLDA, which is able to leverage either document or word meta information, or both of them jointly. With two data argumentation techniques, we can derive an efficient Gibbs sampling algorithm, which benefits from the fully local conjugacy of the model. Moreover, the algorithm is favoured by the sparsity of the meta information. Extensive experiments on several real world datasets demonstrate that our model achieves comparable or improved performance in terms of both perplexity and topic quality, particularly in handling sparse texts. In addition, compared with other models using meta information, our model runs significantly faster. He Zhao 0001, Lan Du 0002, Wray L. Buntine, Gang Liu 0021 |
ICDM | 4 |