Eric Hanchen Jiang

dblp:363/7444 · DBLP profile ↗
← Back
7ranked-venue papers
2as first author
7since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 7 · 2 first-author · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Reinforcement learning · 22% Multi-agent systems · 22% Trustworthy machine learning · 14%

Topics — the 10 heaviest of 11, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Question answering and dialogue systems
divide-and-conquer reasoning
1.012026
Training LLMs for Divide-and-Conquer Reasoning Elevates Test-Time Scalability · ACL (1) 2026
Machine learning › Generative modeling › diffusion model
graph diffusion model
1.012026
Dynamic Generation of Multi LLM Agents Communication Topologies with Graph Diffusion Models · ACL (1) 2026
Natural language and speech › Language models and text generation
inference-time intervention
1.012026
Mitigating Over-Refusal in Aligned Large Language Models via Inference-Time Activation Energy · ACL (1) 2026
Machine learning › Trustworthy machine learning › AI safety
over-refusal mitigation
1.012026
Mitigating Over-Refusal in Aligned Large Language Models via Inference-Time Activation Energy · ACL (1) 2026
Machine learning › Reinforcement learning
reward design
1.012026
ENCORE: Entropy-guided Reward Composition for Multi-head Safety Reward Models · AAAI 2026
Machine learning › Reinforcement learning › reward learning
reward modeling
1.012026
ENCORE: Entropy-guided Reward Composition for Multi-head Safety Reward Models · AAAI 2026
Machine learning › Efficient and distributed learning › federated learning
federated graph learning
0.312026
DAWN: Distributed LLM Multi-Agent Workflow Synthesis · AAAI 2026
Machine learning › Efficient and distributed learning
federated learning
0.312026
DAWN: Distributed LLM Multi-Agent Workflow Synthesis · AAAI 2026
Machine learning › Trustworthy machine learning › AI safety › safety alignment
LLM safety alignment
0.312026
ENCORE: Entropy-guided Reward Composition for Multi-head Safety Reward Models · AAAI 2026
Natural language and speech › Language models and text generation
test-time scaling
0.312026
Training LLMs for Divide-and-Conquer Reasoning Elevates Test-Time Scalability · ACL (1) 2026

Methods — techniques the papers use, named apart from their topics

supervised fine-tuning · 1.0structural gravity · 1.0parametric resonance · 1.0inference-time intervention · 1.0gromov-wasserstein distance · 1.0graph diffusion model · 1.0entropy-guided reward composition · 1.0activation energy · 1.0SVD-based denoising · 1.0
YearPublicationVenuePosition
2026 ENCORE: Entropy-guided Reward Composition for Multi-head Safety Reward Models
Xupeng Chen, Jingxuan Fan, Eric Hanchen Jiang, Mingye Gao
AAAI4
2026 DAWN: Distributed LLM Multi-Agent Workflow Synthesis
abstract
Large language models (LLMs) have recently empowered multi-agent systems (MAS) to achieve remarkable advances in collaborative reasoning and complex task automation. The effectiveness of these systems fundamentally depends on the design of adaptive communication graphs—the underlying workflows that coordinate agent interactions. However, in real-world scenarios, strict privacy constraints often silo data across organizations, and client distributions are highly non-IID, posing major challenges for synthesizing such workflows. In this work, we are the first to systematically study distributed multi-agent workflow synthesis under these privacy and heterogeneity constraints, and we introduce the Difficulty-Based Skew (DBS) benchmark to emulate such challenging environments. Drawing inspiration from federated graph learning (FGL)—which has primarily focused on classification over static graphs—we identify a critical gap: existing FGL methods do not address the generative design of communication topologies. We reveal two fundamental obstacles to generative workflow synthesis in this setting: (i) workflow specialization conflict, where agents optimized for different task distributions generate incompatible communication patterns that resist meaningful aggregation, and (ii) structural communication shift, where locally optimal agent interaction graphs fail to compose into globally coherent multi-agent workflows. To address these challenges, we propose DAWN, a federated framework that integrates two key innovations: Parametric Resonance, which robustly aggregates heterogeneous local updates via layer-wise SVD-based denoising and alignment, and Structural Gravity, which regularizes local workflow generation by penalizing the Fusion Gromov-Wasserstein distance to a set of prototype communication graphs, ensuring global structural coherence without stifling local adaptation. Experiments on the DBS benchmark show that DAWN surpasses baselines in global task success and reduces inter-client graph divergence, laying a solid foundation for privacy-preserving, adaptive MAS workflow design in heterogeneous settings.
Guancheng Wan, Xiaoran Shang, Eric Hanchen Jiang, Guibin Zhang, Jinhe Bi, Yunpu Ma, Zaixi Zhang, Ke Liang 0006, Wenke Huang 0003
AAAI5
2026 Dynamic Generation of Multi LLM Agents Communication Topologies with Graph Diffusion Models
abstract
Eric Hanchen Jiang, Levina Li, Frank Wan, Xiao Liang, Sophia Yin, Yuchen Wu, Xinfeng Li, Yizhou Sun, Wei Wang, Kai-Wei Chang, Ying Nian Wu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Eric Hanchen Jiang, Levina Li, Frank Wan, Sophia Yin, Xinfeng Li, Yizhou Sun, Kai-Wei Chang 0001, Ying Nian Wu
ACL (1)1
2026 Mitigating Over-Refusal in Aligned Large Language Models via Inference-Time Activation Energy
abstract
Eric Hanchen Jiang, Weixuan Ou, Run Liu, Shengyuan Pang, Guancheng Wan, Ranjie Duan, Wei Dong, Kai-Wei Chang, XiaoFeng Wang, Ying Nian Wu, Xinfeng Li. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Eric Hanchen Jiang, Weixuan Ou, Run Liu 0005, Shengyuan Pang, Guancheng Wan, Ranjie Duan, Wei Dong 0007, Kai-Wei Chang 0001, Xiaofeng Wang 0001, Ying Nian Wu, Xinfeng Li
ACL (1)1
2026 Training LLMs for Divide-and-Conquer Reasoning Elevates Test-Time Scalability
abstract
Xiao Liang, Zhong-Zhi Li, Zhenghao Lin, Eric Hanchen Jiang, Hengyuan Zhang, Yelong Shen, Kai-Wei Chang, Ying Nian Wu, Yeyun Gong, Weizhu Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhongzhi Li, Zhenghao Lin, Eric Hanchen Jiang, Yelong Shen, Kai-Wei Chang 0001, Ying Nian Wu, Yeyun Gong, Weizhu Chen
ACL (1)4
2025 Statistical Guarantees for Lifelong Reinforcement Learning using PAC-Bayes Theory
abstract
Lifelong reinforcement learning (RL) has been developed as a paradigm for extending single-task RL to more realistic, dynamic settings. In lifelong RL, the "life" of an RL agent is modeled as a stream of tasks drawn from a task distribution. We propose EPIC (Empirical PAC-Bayes that Improves Continuously), a novel algorithm designed for lifelong RL using PAC-Bayes theory. EPIC learns a shared policy distribution, referred to as the world policy, which enables rapid adaptation to new tasks while retaining valuable knowledge from previous experiences. Our theoretical analysis establishes a relationship between the algorithm’s generalization performance and the number of prior tasks preserved in memory. We also derive the sample complexity of EPIC in terms of RL regret. Extensive experiments on a variety of environments demonstrate that EPIC significantly outperforms existing methods in lifelong RL, offering both theoretical guarantees and practical efficacy through the use of the world policy.
Zhi Zhang 0012, Chris Chow, Yasi Zhang, Yanchao Sun, Eric Hanchen Jiang, Han Liu 0001, Furong Huang, Yuchen Cui, Oscar Hernan Madrid Padilla
AISTATS6
2025 CARES: Comprehensive Evaluation of Safety and Adversarial Robustness in Medical LLMs
abstract
Large language models (LLMs) are increasingly deployed in medical contexts, raising critical concerns about safety, alignment, and susceptibility to adversarial manipulation. While prior benchmarks assess model refusal capabilities for harmful prompts, they often lack clinical specificity, graded harmfulness levels, and coverage of jailbreak-style attacks. We introduce CARES (Clinical Adversarial Robustness and Evaluation of Safety), a benchmark for evaluating LLM safety in healthcare. CARES includes over 18,000 prompts spanning eight medical safety principles, four harm levels, and four prompting styles—direct, indirect, obfuscated, and role-play—to simulate both malicious and benign use cases. We propose a three-way response evaluation protocol (Accept, Caution, Refuse) and a fine-grained Safety Score metric to assess model behavior. Our analysis reveals that many state-of-the-art LLMs remain vulnerable to jailbreaks that subtly rephrase harmful prompts, while also over-refusing safe but atypically phrased queries. Finally, we propose a mitigation strategy using a lightweight classifier to detect jailbreak attempts and steer models toward safer behavior via reminder-based conditioning. CARES provides a rigorous framework for testing and improving medical LLM safety under adversarial and ambiguous conditions.
Mengxue Zhang, Eric Hanchen Jiang, Qingcheng Zeng, Chen-Hsiang Yu
NeurIPS4