Jueqing Lu

dblp:276/5821 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
5since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 6 · 2 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Trustworthy machine learning · 31% Information extraction and text analysis · 23% Learning paradigms · 22%

Topics — the 13 heaviest of 14, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Trustworthy machine learning
robustness
1.922026
Next Generation Active Learning: Mixture of LLMs in the Loop · AAAI 2026
Navigating Conflicting Views: Harnessing Trust for Learning · ICML 2025
Machine learning › Efficient and distributed learning
active learning
1.012026
Next Generation Active Learning: Mixture of LLMs in the Loop · AAAI 2026
Machine learning › Trustworthy machine learning › robustness
learning with noisy labels
1.012026
Next Generation Active Learning: Mixture of LLMs in the Loop · AAAI 2026
Natural language and speech › Information extraction and text analysis › data annotation
LLM-based annotation
1.012026
Next Generation Active Learning: Mixture of LLMs in the Loop · AAAI 2026
Machine learning › Learning paradigms › semi-supervised learning
consistency regularization
0.912025
CGMatch: A Different Perspective of Semi-supervised Learning · CVPR 2025
Natural language and speech › Language models and text generation › large language model
large language model integration
0.912025
Neural Topic Modeling with Large Language Models in the Loop · ACL (1) 2025
Natural language and speech › Information extraction and text analysis › topic model
neural topic model
0.912025
Neural Topic Modeling with Large Language Models in the Loop · ACL (1) 2025
Machine learning › Learning paradigms › semi-supervised learning
pseudo-labeling
0.912025
CGMatch: A Different Perspective of Semi-supervised Learning · CVPR 2025
Machine learning › Learning paradigms
semi-supervised learning
0.912025
CGMatch: A Different Perspective of Semi-supervised Learning · CVPR 2025
Natural language and speech › Information extraction and text analysis
topic model
0.912025
Neural Topic Modeling with Large Language Models in the Loop · ACL (1) 2025
Machine learning › Transfer learning and domain adaptation
few-shot learning
0.412020
Multi-label Few/Zero-shot Learning with Knowledge Aggregated from Multiple Label Graphs · EMNLP (1) 2020
Machine learning › Transfer learning and domain adaptation › few-shot learning
multi-label few-shot learning
0.412020
Multi-label Few/Zero-shot Learning with Knowledge Aggregated from Multiple Label Graphs · EMNLP (1) 2020
Knowledge, reasoning and agents › Knowledge representation and reasoning
knowledge graph
0.112020
Multi-label Few/Zero-shot Learning with Knowledge Aggregated from Multiple Label Graphs · EMNLP (1) 2020

Methods — techniques the papers use, named apart from their topics

negative learning · 1.0mixture of LLMs · 1.0annotation discrepancy · 1.0optimal transport · 0.9large language model · 0.9fine-grained dynamic selection · 0.9count-gap metric · 0.9computational trust-based discounting · 0.9word embeddings · 0.4multi-graph aggregation · 0.4
YearPublicationVenuePosition
2026 Next Generation Active Learning: Mixture of LLMs in the Loop
abstract
With the rapid advancement and strong generalization capabilities of large language models (LLMs), they have been increasingly incorporated into the active learning pipelines as annotators to reduce annotation costs. However, considering the annotation quality, labels generated by LLMs often fall short of real-world applicability. To address this, we propose a novel active learning framework, Mixture of LLMs in the Loop Active Learning, replacing human annotators with labels generated through a Mixture-of-LLMs-based annotation model, aimed at enhancing LLM-based annotation robustness by aggregating the strengths of multiple LLMs. To further mitigate the impact of the noisy labels, we introduce annotation discrepancy and negative learning to identify the unreliable annotations and enhance learning effectiveness. Extensive experiments demonstrate that our framework achieves performance comparable to human annotation and consistently outperforms single-LLM baselines and other LLM-ensemble-based approaches. Moreover, our framework is built on lightweight LLMs, enabling it to operate fully on local machines in real-world applications.
Xiaohao Yang, Jueqing Lu, Guoxiang Guo, Joanne Enticott, Gang Liu 0021, Lan Du 0002
AAAI3
2025 Neural Topic Modeling with Large Language Models in the Loop
abstract
Topic modeling is a fundamental task in natural language processing, allowing the discovery of latent thematic structures in text corpora.While Large Language Models (LLMs) have demonstrated promising capabilities in topic discovery, their direct application to topic modeling suffers from issues such as incomplete topic coverage, misalignment of topics, and inefficiency.To address these limitations, we propose LLM-ITL, a novel LLM-in-theloop framework that integrates LLMs with Neural Topic Models (NTMs).In LLM-ITL, global topics and document representations are learned through the NTM.Meanwhile, an LLM refines these topics using an Optimal Transport (OT)-based alignment objective, where the refinement is dynamically adjusted based on the LLM's confidence in suggesting topical words for each set of input words.With the flexibility of being integrated into many existing NTMs, the proposed approach enhances the interpretability of topics while preserving the efficiency of NTMs in learning topics and document representations.Extensive experiments demonstrate that LLM-ITL helps NTMs significantly improve their topic interpretability while maintaining the quality of document representation.
Xiaohao Yang, He Zhao 0001, Weijie Xu, Jueqing Lu, Dinh Q. Phung, Lan Du 0002
ACL (1)5
2025 CGMatch: A Different Perspective of Semi-supervised Learning
abstract
Semi-supervised learning (SSL) has garnered significant attention due to its ability to leverage limited labeled data and a large amount of unlabeled data to improve model generalization performance. Recent approaches achieve impressive successes by combining ideas from both consistency regularization and pseudo-labeling. However, these methods tend to underperform in the more realistic situations with relatively scarce labeled data. We argue that this issue arises because existing methods rely solely on the model’s confidence, making them challenging to accurately assess the model’s state and identify unlabeled examples contributing to the training phase when supervision information is limited, especially during the early stages of model training. In this paper, we propose a novel SSL model called CGMatch, which, for the first time, incorporates a new metric known as Count-Gap (CG). We demonstrate that CG is effective in discovering unlabeled examples beneficial for model training. Along with confidence, a commonly used metric in SSL, we propose a fine-grained dynamic selection (FDS) strategy. This strategy dynamically divides the unlabeled dataset into three subsets with different characteristics: easy-to-learn set, ambiguous set, and hard-to-learn set. By selective filtering subsets, and applying corresponding regularization with selected subsets, we mitigate the negative impact of incorrect pseudo-labels on model optimization and generalization. Extensive experimental results on several common SSL benchmarks indicate the effectiveness of CGMatch especially when the labeled data are particularly limited. Source code is available at https://github.com/BoCheng-96/CGMatch.
Jueqing Lu, Yuan Tian 0016, Haifeng Zhao 0002, Yi Chang 0001, Lan Du 0002
CVPR2
2025 Navigating Conflicting Views: Harnessing Trust for Learning
abstract
Resolving conflicts is critical for improving the reliability of multi-view classification. While prior work focuses on learning consistent and informative representations across views, it often assumes perfect alignment and equal importance of all views, an assumption rarely met in real-world scenarios, as some views may express distinct information. To address this, we develop a computational trust-based discounting method that enhances the Evidential Multi-view framework by accounting for the instance-wise reliability of each view through a probability-sensitive trust mechanism. We evaluate our method on six real-world datasets using Top-1 Accuracy, Fleiss’ Kappa, and a new metric, Multi-View Agreement with Ground Truth, to assess prediction reliability. We also assess the effectiveness of uncertainty in indicating prediction correctness via AUROC. Additionally, we test the scalability of our method through end-to-end training on a large-scale dataset. The experimental results show that computational trust can effectively resolve conflicts, paving the way for more reliable multi-view classification models in real-world applications. Codes available at: https://github.com/OverfitFlow/Trust4Conflict
Jueqing Lu, Wray L. Buntine, Joanna Dipnall, Belinda Gabbe, Lan Du 0002
ICML1
2025 Multi-Label Bayesian Active Learning with Inter-Label Relationships
abstract
The primary challenge of multi-label active learning, differing it from multi-class active learning, lies in assessing the informativeness of an indefinite number of labels while also accounting for the inherited label correlation. Existing studies either require substantial computational resources to leverage correlations or fail to fully explore label dependencies. Additionally, real-world scenarios often require addressing intrinsic biases stemming from imbalanced data distributions. In this paper, we propose a new multi-label active learning strategy to address both challenges. Our method incorporates progressively updated positive and negative correlation matrices to capture co-occurrence and disjoint relationships within the label space of annotated samples, enabling a holistic assessment of uncertainty rather than treating labels as isolated elements. Furthermore, alongside diversity, our model employs ensemble pseudo labeling and beta scoring rules to address data imbalances. Extensive experiments on four realistic datasets demonstrate that our strategy consistently achieves more reliable and superior performance, compared to several established methods.
Jueqing Lu, Xiaohao Yang, Joanne Enticott, Lan Du 0002
UAI2
2020 Multi-label Few/Zero-shot Learning with Knowledge Aggregated from Multiple Label Graphs
abstract
Few/Zero-shot learning is a big challenge of many classifications tasks, where a classifier is required to recognise instances of classes that have very few or even no training samples. It becomes more difficult in multilabel classification, where each instance is la-belled with more than one class. In this paper, we present a simple multi-graph aggregation model that fuses knowledge from multiple label graphs encoding different semantic label relationships in order to study how the aggregated knowledge can benefit multi-label zero/few-shot document classification. The model utilises three kinds of semantic information, i.e., the pre-trained word embeddings, label description, and pre-defined label relations. Experimental results derived on two large clinical datasets (i.e., MIMIC-II and MIMIC-III) and the EU legislation dataset show that methods equipped with the multi-graph knowledge aggregation achieve significant performance improvement across almost all the measures on few/zero-shot labels.
Jueqing Lu, Lan Du 0002, Ming Liu 0028, Joanna Dipnall
EMNLP (1)1