Jiexiang Xu

dblp:362/9488 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
2since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 2 · 2 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
2 papers
Language models and text generation · 40% Robot manipulation · 20% Multi-agent systems · 20%

Topics — the 5 heaviest of 6, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Knowledge, reasoning and agents › Multi-agent systems › multi-agent systems engineering › multi-agent evaluation
failure attribution
1.012026
ARCHITECT: Uncertainty-Aware Dynamic Tool Learning via Causal Intervention for Open-World Agents · ACL (1) 2026
Machine learning › Trustworthy machine learning › interpretability › model debugging
failure mode analysis
1.012026
Action Boundary Blindness: When LLM Agents Cannot Tell Where One Action Ends and Another Begins · ACL (1) 2026
Natural language and speech › Language models and text generation
LLM agents
1.012026
Action Boundary Blindness: When LLM Agents Cannot Tell Where One Action Ends and Another Begins · ACL (1) 2026
Robotics › Robot manipulation
tool generation
1.012026
ARCHITECT: Uncertainty-Aware Dynamic Tool Learning via Causal Intervention for Open-World Agents · ACL (1) 2026
Natural language and speech › Language models and text generation › LLM agents
tool learning
1.012026
ARCHITECT: Uncertainty-Aware Dynamic Tool Learning via Causal Intervention for Open-World Agents · ACL (1) 2026

Methods — techniques the papers use, named apart from their topics

uncertainty-aware prediction · 1.0structural causal model · 1.0multi-label attribution · 1.0explicit boundary prompting · 1.0event segmentation theory · 1.0causal intervention · 1.0automatic metrics · 1.0
YearPublicationVenuePosition
2026 ARCHITECT: Uncertainty-Aware Dynamic Tool Learning via Causal Intervention for Open-World Agents
abstract
Dynamic tool generation empowers Large Language Model (LLM) agents to synthesize tools on demand, yet a critical challenge remains: 32.4% of generated tools fail on first invocation.We present Causal Tool Diagnosis (CTD), a principled framework that moves beyond blackbox reliability prediction to interpretable failure attribution.CTD constructs a Structural Causal Model (SCM) capturing how specification quality, code characteristics, and execution environment jointly determine tool outcomes.Uniquely leveraging code's intervenability, we conduct controlled sandbox experiments to estimate causal effects-an advantage unavailable in pure text generation.CTD jointly predicts confidence (Spearman rank correlation coefficient ρ=0.90) and root cause attribution (78% accuracy), with attributions directly guiding targeted repairs (+9.6% success rate over error-type classification).Our ARCHI-TECT framework, integrating CTD throughout the tool lifecycle, achieves state-of-the-art on four benchmarks including StableToolBench (+3.8%),MINT (+4.6%),T-Eval (+3.7%), and SWE-bench Lite (+2.4%), with consistent improvements across all settings.
Zhangyi Wang, Jiexiang Xu, Bingnan Yu, Zongze Li 0001
ACL (1)2
2026 Action Boundary Blindness: When LLM Agents Cannot Tell Where One Action Ends and Another Begins
abstract
Large language model (LLM) agents excel at multi-step tasks yet frequently exhibit Action Boundary Blindness-the inability to correctly determine action granularity, scope, and completeness.Grounded in Event Segmentation Theory from cognitive science, we formalize three violation types: granularity confusion, scope creep, and boundary ambiguity.We propose four automatic metrics-Action Boundary Score (ABS), Granularity Alignment Rate (GAR), Scope Violation Rate (SVR), and Boundary-Aware Success Rate (BASR)requiring no human annotation.Experiments on 1,655 tasks across six benchmarks (τ -bench, WebArena, ALFWorld, TheAgentCompany, OSWorld) with seven LLMs reveal that: (1) the best model achieves only 0.424 ABS; (2) using a multi-label attribution framework validated by inter-annotator agreement (κ = 0.78), boundary blindness is the primary failure mode in 37.2% of failures (25.8% as sole cause; 55.9% total involvement including contributing factors); (3) under-action dominates at 48.4%; (4) BASR is consistently ∼4 points lower than traditional success rate, exposing "lucky successes."Critically, Explicit Boundary Prompting (EBP) improves ABS by 0.08-0.13across all models, demonstrating that boundary blindness is better characterized as an elicitation gap rather than a fundamental capability limitation-LLMs possess latent boundary perception not activated by default.This finding has implications for alignment and instruction tuning.We validate metrics through state-based cross-validation and human audit, estimating ∼22% false positive rate from valid alternative paths, with model rankings remaining stable (Spearman ρ = 1.0).
Zhangyi Wang, Bingnan Yu, Jiexiang Xu
ACL (1)3