Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Justin Chih-Yao Chen

dblp:357/5257 · DBLP profile ↗
← Back
9ranked-venue papers
4as first author
9since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 9 · 4 first-author · 9 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Language models and text generation · 54% Planning, search and constraint satisfaction · 13% Efficient and distributed learning · 12%

Topics — the 14 heaviest of 16, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Question answering and dialogue systems
information seeking
1.012026
PRInTS: Reward Modeling for Long-Horizon Information Seeking · ACL (1) 2026
Natural language and speech › Language models and text generation
model routing
1.012026
Routing with Generated Data: Annotation-Free LLM Skill Estimation and Expert Selection · ACL (1) 2026
Machine learning › Reinforcement learning › reward learning
reward modeling
1.012026
PRInTS: Reward Modeling for Long-Horizon Information Seeking · ACL (1) 2026
Natural language and speech › Language models and text generation › evaluation of language models
skill estimation
1.012026
Routing with Generated Data: Annotation-Free LLM Skill Estimation and Expert Selection · ACL (1) 2026
Natural language and speech › Language models and text generation
large language model reasoning
0.912025
System 1.x: Learning to Balance Fast and Slow Planning with Language Models · ICLR 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning › language-based planning
LLM-based planning
0.912025
System 1.x: Learning to Balance Fast and Slow Planning with Language Models · ICLR 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction › planning › hybrid planning
neuro-symbolic planning
0.912025
System 1.x: Learning to Balance Fast and Slow Planning with Language Models · ICLR 2025
Natural language and speech › Language models and text generation › self-improvement
self-refinement
0.912025
MAgICoRe: Multi-Agent, Iterative, Coarse-to-Fine Refinement for Reasoning · EMNLP 2025
Natural language and speech › Language models and text generation
test-time scaling
0.912025
MAgICoRe: Multi-Agent, Iterative, Coarse-to-Fine Refinement for Reasoning · EMNLP 2025
Natural language and speech › Language models and text generation › large language model reasoning
collaborative reasoning
0.812024
ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs · ACL (1) 2024
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.812024
MAGDi: Structured Distillation of Multi-Agent Interaction Graphs Improves Reasoning in Smaller Language Models · ICML 2024
Natural language and speech › Language models and text generation
LLM agents
0.812024
MAGDi: Structured Distillation of Multi-Agent Interaction Graphs Improves Reasoning in Smaller Language Models · ICML 2024
Knowledge, reasoning and agents › Multi-agent systems › LLM-based multi-agent systems
multi-agent debate
0.812024
ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs · ACL (1) 2024
Machine learning › Efficient and distributed learning › model compression › knowledge distillation › LLM distillation
reasoning distillation
0.812024
MAGDi: Structured Distillation of Multi-Agent Interaction Graphs Improves Reasoning in Smaller Language Models · ICML 2024

Methods — techniques the papers use, named apart from their topics

synthetic data generation · 1.0reward modeling · 1.0LLM annotation · 1.0self-consistency · 0.9reward model · 0.9multi-agent · 0.9language model fine-tuning · 0.9a* search · 0.9DFS · 0.9BFS · 0.9
YearPublicationVenuePosition
2026 PRInTS: Reward Modeling for Long-Horizon Information Seeking
abstract
Jaewoo Lee, Archiki Prasad, Justin Chen, Zaid Khan, Elias Stengel-Eskin, Mohit Bansal. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jaewoo Lee 0001, Archiki Prasad, Justin Chih-Yao Chen, Zaid Khan 0001, Elias Stengel-Eskin, Mohit Bansal
ACL (1)3
2026 Routing with Generated Data: Annotation-Free LLM Skill Estimation and Expert Selection
abstract
Tianyi Niu, Justin Chen, Genta Indra Winata, Shi-Xiong Zhang, Supriyo Chakraborty, Sambit Sahu, Yue Zhang, Elias Stengel-Eskin, Mohit Bansal. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Tianyi Niu, Justin Chih-Yao Chen, Genta Indra Winata, Supriyo Chakraborty, Sambit Sahu, Yue Zhang 0004, Elias Stengel-Eskin, Mohit Bansal
ACL (1)2
2025 MAgICoRe: Multi-Agent, Iterative, Coarse-to-Fine Refinement for Reasoning
abstract
Large language model (LLM) reasoning can be improved by scaling test-time compute with aggregation, i.e., generating multiple samples and aggregating over them.While improving performance, this strategy often reaches a saturation point beyond which additional compute provides no return.Refinement offers an alternative by using model-generated feedback to improve answer quality.However, refinement faces three key challenges: (1) Excessive refinement: Uniformly refining all instances can cause over-correction and reduce overall performance.(2) Inability to localize and address errors: LLMs struggle to identify and correct their own mistakes.(3) Insufficient refinement: Stopping refinement too soon could leave errors unaddressed.To tackle these issues, we propose MAGICORE, a framework for Multi-Agent Iteration for Coarse-to-fine Refinement.MAGICORE mitigates excessive refinement by categorizing problems as easy or hard, solving easy problems with coarsegrained aggregation, and solving the hard ones with fine-grained multi-agent refinement.To better localize errors, we incorporate external step-wise reward model scores, and to ensure sufficient refinement, we iteratively refine the solutions using a multi-agent setup.We evaluate MAGICORE on Llama-3-8B and GPT-3.5 and show its effectiveness across seven reasoning datasets.One iteration of MAGI-CORE beats Self-Consistency by 3.4%, Bestof-k by 3.2%, and Self-Refine by 4.0% even when these baselines use k = 120, and MAGI-CORE uses less than 50% of the compute. 1
Justin Chih-Yao Chen, Archiki Prasad, Swarnadeep Saha, Elias Stengel-Eskin, Mohit Bansal
EMNLP1
2025 System 1.x: Learning to Balance Fast and Slow Planning with Language Models
abstract
Language models can be used to solve long-horizon planning problems in two distinct modes. In a fast 'System-1' mode, models directly generate plans without any explicit search or backtracking, and in a slow 'System-2' mode, they plan step-by-step by explicitly searching over possible actions. System-2 planning, while typically more effective, is also computationally more expensive and often infeasible for long plans or large action spaces. Moreover, isolated System-1 or System-2 planning ignores the user's end goals and constraints (e.g., token budget), failing to provide ways for the user to control the model's behavior. To this end, we propose the System-1.x Planner, a framework for controllable planning with language models that is capable of generating hybrid plans and balancing between the two planning modes based on the difficulty of the problem at hand. System-1.x consists of (i) a controller, (ii) a System-1 Planner, and (iii) a System-2 Planner. Based on a user-specified hybridization factor x governing the degree to which the system uses System-1 vs. System-2, the controller decomposes a planning problem into subgoals, and classifies them as easy or hard to be solved by either System-1 or System-2, respectively. We fine-tune all three components on top of a single base LLM, requiring only search traces as supervision. Experiments with two diverse planning tasks -- Maze Navigation and Blocksworld -- show that our System-1.x Planner outperforms a System-1 Planner, a System-2 Planner trained to approximate A* search, and also a symbolic planner (A* search), given a state exploration budget. We also demonstrate the following key properties of our planner: (1) controllability: by adjusting the hybridization factor x (e.g., System-1.75 vs. System-1.5) we can perform more (or less) search, improving performance, (2) flexibility: by building a neuro-symbolic variant composed of a neural System-1 planner and a symbolic System-2 planner, we can take advantage of existing symbolic methods, and (3) generalizability: by learning from different search algorithms (BFS, DFS, A*), we show that our method is robust to the choice of search algorithm used for training.
Swarnadeep Saha, Archiki Prasad, Justin Chih-Yao Chen, Peter Hase, Elias Stengel-Eskin, Mohit Bansal
ICLR3
2025 Reverse Thinking Makes LLMs Stronger Reasoners
abstract
Justin Chen, Zifeng Wang, Hamid Palangi, Rujun Han, Sayna Ebrahimi, Long Le, Vincent Perot, Swaroop Mishra, Mohit Bansal, Chen-Yu Lee, Tomas Pfister. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Justin Chih-Yao Chen, Zifeng Wang 0002, Hamid Palangi, Rujun Han, Sayna Ebrahimi, Long T. Le, Vincent Perot, Swaroop Mishra, Mohit Bansal, Chen-Yu Lee, Tomas Pfister
NAACL (Long Papers)1
2025 MAMM-Refine: A Recipe for Improving Faithfulness in Generation with Multi-Agent Collaboration
abstract
David Wan, Justin Chen, Elias Stengel-Eskin, Mohit Bansal. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
David Wan, Justin Chih-Yao Chen, Elias Stengel-Eskin, Mohit Bansal
NAACL (Long Papers)2
2024 ReConcile: Round-Table Conference Improves Reasoning via Consensus among Diverse LLMs
abstract
Large Language Models (LLMs) still struggle with natural language reasoning tasks.Motivated by the society of minds (Minsky, 1988), we propose RECONCILE, a multi-model multiagent framework designed as a round table conference among diverse LLM agents.RECON-CILE enhances collaborative reasoning between LLM agents via multiple rounds of discussion, learning to convince other agents to improve their answers, and employing a confidenceweighted voting mechanism that leads to a better consensus.In each round, RECONCILE initiates discussion between agents via a 'discussion prompt' that consists of (a) grouped answers and explanations generated by each agent in the previous round, (b) their confidence scores, and (c) demonstrations of answerrectifying human explanations, used for convincing other agents.Experiments on seven benchmarks demonstrate that RECONCILE significantly improves LLMs' reasoning -both individually and as a team -surpassing prior single-agent and multi-agent baselines by up to 11.4% and even outperforming GPT-4 on three datasets.RECONCILE also flexibly incorporates different combinations of agents, including API-based, open-source, and domainspecific models, leading to an 8% improvement on MATH.Finally, we analyze the individual components of RECONCILE, demonstrating that the diversity originating from different models is critical to its superior performance.1 Self-Refine MAD+Judge Multi-Agent Debate (MAD) ReConcile (Group-Discuss-and-Convince) Yes, with 95% confidence No, with 50% confidence No, with 40% confidence yes no no yes no no yes no no Question (Q): Is an ammonia fighting cleaner good for pet owners?Human Explanation (Exp): Ammonia is a component in pet urine.It has an unpleasant odor.
Justin Chih-Yao Chen, Swarnadeep Saha, Mohit Bansal
ACL (1)1
2024 MAGDi: Structured Distillation of Multi-Agent Interaction Graphs Improves Reasoning in Smaller Language Models
abstract
Multi-agent interactions between Large Language Model (LLM) agents have shown major improvements on diverse reasoning tasks. However, these involve long generations from multiple models across several rounds, making them expensive. Moreover, these multi-agent approaches fail to provide a final, single model for efficient inference. To address this, we introduce MAGDi, a new method for structured distillation of the reasoning interactions between multiple LLMs into smaller LMs. MAGDi teaches smaller models by representing multi-agent interactions as graphs, augmenting a base student model with a graph encoder, and distilling knowledge using three objective functions: next-token prediction, a contrastive loss between correct and incorrect reasoning, and a graph-based objective to model the interaction structure. Experiments on seven widely used commonsense and math reasoning benchmarks show that MAGDi improves the reasoning capabilities of smaller models, outperforming several methods that distill from a single teacher and multiple teachers. Moreover, MAGDi also demonstrates an order of magnitude higher efficiency over its teachers. We conduct extensive analyses to show that MAGDi (1) enhances the generalizability to out-of-domain tasks, (2) scales positively with the size and strength of the base student model, and (3) obtains larger improvements (via our multi-teacher training) when applying self-consistency – an inference technique that relies on model diversity.
Justin Chih-Yao Chen, Swarnadeep Saha, Elias Stengel-Eskin, Mohit Bansal
ICML1
2023 Location-Aware Visual Question Generation with Lightweight Models
abstract
This work introduces a novel task, locationaware visual question generation (LocaVQG), which aims to generate engaging questions from data relevant to a particular geographical location.Specifically, we represent such location-aware information with surrounding images and a GPS coordinate.To tackle this task, we present a dataset generation pipeline that leverages GPT-4 to produce diverse and sophisticated questions.Then, we aim to learn a lightweight model that can address the Lo-caVQG task and fit on an edge device, such as a mobile phone.To this end, we propose a method which can reliably generate engaging questions from location-aware information.Our proposed method outperforms baselines regarding human evaluation (e.g., engagement, grounding, coherence) and automatic evaluation metrics (e.g., BERTScore, ROUGE-2).Moreover, we conduct extensive ablation studies to justify our proposed techniques for generating the dataset and solving the task.
Nicholas Collin Suwono, Justin Chih-Yao Chen, Tun-Min Hung, Ting-Hao 'Kenneth' Huang, I-Bin Liao, Yung-Hui Li, Lun-Wei Ku, Shao-Hua Sun
EMNLP2