Ningyuan Liu

dblp:399/8333 · DBLP profile ↗
← Back
3ranked-venue papers
0as first author
3since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
3 papers
Language models and text generation · 28% Reinforcement learning · 25% Multi-agent systems · 22%

Topics — the 11 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › hallucination mitigation
hallucination correction
1.012026
LLM-CAS: Dynamic Neuron Perturbation for Real-Time Hallucination Correction · AAAI 2026
Natural language and speech › Language models and text generation
hallucination mitigation
1.012026
LLM-CAS: Dynamic Neuron Perturbation for Real-Time Hallucination Correction · AAAI 2026
Machine learning › Reinforcement learning
bandit
0.912025
KABB: Knowledge-Aware Bayesian Bandits for Dynamic Expert Coordination in Multi-Agent Systems · ICML 2025
Machine learning › Optimization for machine learning
hyperparameter optimization
0.912025
MAT-Agent: Adaptive Multi-Agent Training Optimization · NeurIPS 2025
Knowledge, reasoning and agents › Multi-agent systems
multi-agent coordination
0.912025
KABB: Knowledge-Aware Bayesian Bandits for Dynamic Expert Coordination in Multi-Agent Systems · ICML 2025
Knowledge, reasoning and agents › Multi-agent systems
multi-agent learning
0.912025
MAT-Agent: Adaptive Multi-Agent Training Optimization · NeurIPS 2025
Machine learning › Reinforcement learning
thompson sampling
0.912025
KABB: Knowledge-Aware Bayesian Bandits for Dynamic Expert Coordination in Multi-Agent Systems · ICML 2025
Machine learning › Deep learning architectures and training
training optimization
0.912025
MAT-Agent: Adaptive Multi-Agent Training Optimization · NeurIPS 2025
Machine learning › Reinforcement learning
hierarchical reinforcement learning
0.312026
LLM-CAS: Dynamic Neuron Perturbation for Real-Time Hallucination Correction · AAAI 2026
Natural language and speech › Language models and text generation
large language model
0.312025
KABB: Knowledge-Aware Bayesian Bandits for Dynamic Expert Coordination in Multi-Agent Systems · ICML 2025
Machine learning › Learning paradigms
multi-label classification
0.312025
MAT-Agent: Adaptive Multi-Agent Training Optimization · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

neuron perturbation · 1.0hierarchical reinforcement learning · 1.0upper confidence bound · 0.9thompson sampling · 0.9multi-armed bandit · 0.9mixed-precision training · 0.9knowledge distance model · 0.9epsilon-greedy · 0.9bayesian bandit · 0.9
YearPublicationVenuePosition
2026 LLM-CAS: Dynamic Neuron Perturbation for Real-Time Hallucination Correction
abstract
Large language models (LLMs) often generate hallucinated content lacking factual or contextual grounding, hindering their reliability in critical applications. Traditional methods like supervised fine-tuning and reinforcement learning from human feedback are data-intensive and computationally expensive, while static parameter editing struggles with context-dependent errors and catastrophic forgetting. To overcome these limitations, we introduce LLM-CAS, a framework that formulates real-time hallucination correction as a hierarchical reinforcement learning (HRL) problem. LLM-CAS trains an agent to learn a sophisticated policy, dynamically selecting optimal, temporary neuron perturbations during inference based on the immediate context. This learned, policy-driven approach provides greater adaptability than prior dynamic methods that rely on heuristic or pre-defined adjustments. As a result, LLM-CAS achieves significant performance gains across various LLMs, improving accuracy by 10.98 percentage points on StoryCloze, 2.71 points on TriviaQA, and 2.06 points on TruthfulQA's MC1 score, thereby outperforming static methods like ITI and CAA, as well as the dynamic SADI framework. This context-aware, efficient approach promises enhanced reliability for LLMs in high-stakes domains, with future potential for multimodal extensions.
Jusheng Zhang, Ningyuan Liu, Yijia Fan, Qinglin Zeng, Kaitong Cai, Jian Wang 0100, Keze Wang
AAAI2
2025 KABB: Knowledge-Aware Bayesian Bandits for Dynamic Expert Coordination in Multi-Agent Systems
abstract
As scaling large language models faces prohibitive costs, multi-agent systems emerge as a promising alternative, though challenged by static knowledge assumptions and coordination inefficiencies. We introduce Knowledge-Aware Bayesian Bandits (KABB), a novel framework that enhances multi-agent system coordination through semantic understanding and dynamic adaptation. The framework features three key innovations: a customized knowledge distance model for deep semantic understanding, a dual-adaptation mechanism for continuous expert optimization, and a knowledge-aware Thompson Sampling strategy for efficient expert selection. Extensive evaluation demonstrates KABB achieves an optimal cost-performance balance, maintaining high performance while keeping computational demands relatively low in multi-agent coordination.
Jusheng Zhang, Zimeng Huang, Yijia Fan, Ningyuan Liu, Zhuojie Yang, Jiawei Yao, Jian Wang 0100, Keze Wang
ICML4
2025 MAT-Agent: Adaptive Multi-Agent Training Optimization
abstract
We propose a novel collaborative multi-agent optimization framework for adaptive training in multi-label image classification, fundamentally advancing beyond static decision rules and isolated automation. Our method deploys a set of distributed, task-specific agents, each responsible for dynamically orchestrating critical training components—including data augmentation, optimization methods, learning rate schedules, and loss functions—according to evolving visual-semantic relationships and training states. Each agent employs an advanced non-stationary multi-armed bandit algorithm, integrating both $\epsilon$-greedy and upper confidence bound strategies, to judiciously balance exploration with exploitation throughout the training lifecycle. A hierarchical composite reward mechanism synergizes overall classification accuracy, rare class recognition, and training stability, fostering both independent optimization and implicit collaborative behavior among agents. The framework further leverages refined techniques such as dual-rate exponential moving average smoothing and structured mixed-precision training to enhance robustness and computational efficiency. Extensive experiments across benchmarks including Pascal VOC, COCO, Yeast, and Mediamill demonstrate that our approach achieves superior mean average precision and rare-class F1 scores compared to state-of-the-art methods, while also exhibiting rapid convergence and remarkable cross-domain generalization. Our results indicate that collaborative multi-agent adaptive optimization offers a scalable and principled solution for self-optimizing deep learning in complex multi-label scenarios.
Jusheng Zhang, Kaitong Cai, Yijia Fan, Ningyuan Liu, Keze Wang
NeurIPS4