Robert Tang

dblp:25/9456 · DBLP profile ↗
← Back
8ranked-venue papers
0as first author
7since 2021 · last 2025
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 8 · 7 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Language models and text generation · 44% Multi-agent systems · 32% Planning, search and constraint satisfaction · 16%
Software engineering, system software, and programming languages
2 papers
Program synthesis and code generation · 70% Debugging and program repair · 30%
Interdisciplinary, comprehensive, and emerging computing
1 paper
Bioinformatics and computational biology · 100%
Databases, data mining, and information retrieval
2 papers
Information retrieval · 100%

Topics — the 15 heaviest of 21, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
large language model
1.122025
KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation · NeurIPS 2025
DyFlow: Dynamic Workflow Framework for Agentic Reasoning · NeurIPS 2025
Knowledge, reasoning and agents › Multi-agent systems
agentic AI
0.912025
WebDancer: Towards Autonomous Information Seeking Agency · NeurIPS 2025
Knowledge, reasoning and agents › Multi-agent systems › agentic AI
agentic reasoning
0.912025
DyFlow: Dynamic Workflow Framework for Agentic Reasoning · NeurIPS 2025
Machine learning › Transfer learning and domain adaptation
foundation model evaluation
0.912025
SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks · NeurIPS 2025
Natural language and speech › Language models and text generation › large language model evaluation › capability evaluation
game-based evaluation
0.912025
KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation · NeurIPS 2025
Knowledge, reasoning and agents › Planning, search and constraint satisfaction
game playing
0.912025
KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation · NeurIPS 2025
Knowledge, reasoning and agents › Multi-agent systems › LLM-based multi-agent systems
multi-agent LLM coordination
0.912025
MultiAgentBench : Evaluating the Collaboration and Competition of LLM agents · ACL (1) 2025
Natural language and speech › Language models and text generation › evaluation of language models
reasoning evaluation
0.912025
KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation · NeurIPS 2025
Natural language and speech › Language models and text generation › LLM agents
web agents
0.912025
WebDancer: Towards Autonomous Information Seeking Agency · NeurIPS 2025
Bioinformatics and computational biology › molecular informatics › cheminformatics
chemical reaction prediction
0.912025
Beyond Chemical QA: Evaluating LLM's Chemical Reasoning with Modular Chemical Operations · NeurIPS 2025
Bioinformatics and computational biology › drug discovery
molecular optimization
0.912025
Beyond Chemical QA: Evaluating LLM's Chemical Reasoning with Modular Chemical Operations · NeurIPS 2025
Program synthesis and code generation
code agent
0.912025
LocAgent: Graph-Guided LLM Agents for Code Localization · ACL (1) 2025
Program synthesis and code generation
code generation with language models
0.912025
LocAgent: Graph-Guided LLM Agents for Code Localization · ACL (1) 2025
Debugging and program repair
code localization
0.912025
LocAgent: Graph-Guided LLM Agents for Code Localization · ACL (1) 2025
Natural language and speech › Language models and text generation › large language model evaluation
LLM-as-a-judge
0.312025
SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

fine-tuning · 2.6expert annotation · 2.6chain-of-thought reasoning · 2.6react · 1.7supervised fine-tuning · 0.9reinforcement learning scenarios · 0.9reinforcement learning · 0.9pairwise model comparison · 0.9milestone-based evaluation · 0.9large language model · 0.9interactive evaluation · 0.9graph neural network · 0.9dynamic workflow generation · 0.9coordination protocol analysis · 0.9context-aware parameterization · 0.9community voting · 0.9
YearPublicationVenuePosition
2025 LocAgent: Graph-Guided LLM Agents for Code Localization
abstract
Zhaoling Chen, Robert Tang, Gangda Deng, Fang Wu, Jialong Wu, Zhiwei Jiang, Viktor Prasanna, Arman Cohan, Xingyao Wang. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Zhaoling Chen, Robert Tang, Gangda Deng, Fang Wu 0002, Jialong Wu 0010, Viktor Prasanna 0001, Arman Cohan, Xingyao Wang 0002
ACL (1)2
2025 MultiAgentBench : Evaluating the Collaboration and Competition of LLM agents
abstract
Large Language Models (LLMs) have shown remarkable capabilities as autonomous agents; yet existing benchmarks either focus on single-agent tasks or are confined to narrow domains, failing to capture the dynamics of multi-agent coordination and competition. In this paper, we introduce MultiAgentBench, a comprehensive benchmark designed to evaluate LLM-based multi-agent systems across diverse, interactive scenarios. Our framework measures not only task completion but also the quality of collaboration and competition using novel, milestone-based key performance indicators. Moreover, we evaluate various coordination protocols (including star, chain, tree, and graph topologies) and innovative strategies such as group discussion and cognitive planning. Notably, gpt-4o-mini reaches the average highest task score, graph structure performs the best among coordination protocols in the research scenario,and cognitive planning improves milestone achievement rates by 3%. Code and datasets are publicavailable at https://github.com/ulab-uiuc/MARBLE.
Kunlun Zhu, Hongyi Du, Zhaochen Hong, Xiaocheng Yang, Shuyi Guo, Zhenhailong Wang, Cheng Qian 0008, Robert Tang, Heng Ji 0001, Jiaxuan You
ACL (1)9
2025 Beyond Chemical QA: Evaluating LLM's Chemical Reasoning with Modular Chemical Operations
abstract
While large language models (LLMs) with Chain-of-Thought (CoT) reasoning excel in mathematics and coding, their potential for systematic reasoning in chemistry, a domain demanding rigorous structural analysis for real-world tasks like drug design and reaction engineering, remains untapped. Current benchmarks focus on simple knowledge retrieval, neglecting step-by-step reasoning required for complex tasks such as molecular optimization and reaction prediction. To address this, we introduce ChemCoTBench, a reasoning framework that bridges molecular structure understanding with arithmetic-inspired operations, including addition, deletion, and substitution, to formalize chemical problem-solving into transparent, step-by-step workflows. By treating molecular transformations as modular "chemical operations", the framework enables slow-thinking reasoning, mirroring the logic of mathematical proofs while grounding solutions in real-world chemical constraints. We evaluate models on two high-impact tasks: Molecular Property Optimization and Chemical Reaction Prediction. These tasks mirror real-world challenges while providing structured evaluability. We further provide ChemCoTDataset, a pioneering 22,000-instance chemical reasoning dataset with expert-annotated chains of thought to facilitate LLM fine-tuning. By providing annotated trainable datasets, a reasoning taxonomy, and baseline evaluations, our work bridges the gap between abstract reasoning methods and practical chemical discovery, establishing a foundation for advancing LLMs as tools for AI-driven scientific innovation.
Hao Li 0073, He Cao, Bin Feng 0001, Daniel Shao, Robert Tang, Zhiyuan Yan 0002, Yonghong Tian 0001, Li Yuan 0007, Yu Li 0003
NeurIPS5
2025 KORGym: A Dynamic Game Platform for LLM Reasoning Evaluation
abstract
Recent advancements in large language models (LLMs) underscore the need for more comprehensive evaluation methods to accurately assess their reasoning capabilities. Existing benchmarks are often domain-specific and thus cannot fully capture an LLM’s general reasoning potential. To address this limitation, we introduce the **Knowledge Orthogonal Reasoning Gymnasium (KORGym)**, a dynamic evaluation platform inspired by KOR-Bench and Gymnasium. KORGym offers over fifty games in either textual or visual formats and supports interactive, multi-turn assessments with reinforcement learning scenarios. Using KORGym, we conduct extensive experiments on 19 LLMs and 8 VLMs, revealing consistent reasoning patterns within model families and demonstrating the superior performance of closed-source models. Further analysis examines the effects of modality, reasoning strategies, reinforcement learning techniques, and response length on model performance. We expect KORGym to become a valuable resource for advancing LLM reasoning research and developing evaluation methodologies suited to complex, interactive environments.
Jiajun Shi, Jian Yang 0037, Xingyuan Bu, Jiangjie Chen, Junting Zhou, Kaijing Ma, Zhoufutu Wen, Bingli Wang, Yancheng He, Hualei Zhu, Wei Zhang 0021, Ruibin Yuan, Yunli Wang, Siyuan Fang, Qianyu He, Robert Tang, Yingshui Tan, Wangchunshu Zhou, Zhaoxiang Zhang 0001, Zhoujun Li 0001, Wenhao Huang 0001, Ge Zhang 0009
NeurIPS23
2025 DyFlow: Dynamic Workflow Framework for Agentic Reasoning
abstract
Agent systems based on large language models (LLMs) have shown great potential in complex reasoning tasks, but building efficient and generalizable workflows remains a major challenge. Most existing approaches rely on manually designed processes, which limits their adaptability across different tasks. While a few methods attempt automated workflow generation, they are often tied to specific datasets or query types and make limited use of intermediate feedback, reducing system robustness and reasoning depth. Moreover, their operations are typically predefined and inflexible. To address these limitations, we propose **DyFlow**, a dynamic workflow generation framework that adaptively constructs and adjusts reasoning procedures based on task requirements and real-time intermediate feedback, thereby enhancing cross-task generalization. DyFlow consists of two core components: a designer and an executor. The designer decomposes complex problems into a sequence of sub-goals defined by high-level objectives and dynamically plans the next steps based on intermediate outputs and feedback. These plans are then carried out by the executor, which executes each operation using dynamic operators with context-aware parameterization, enabling flexible and semantically grounded reasoning. We systematically evaluate DyFlow across diverse domains, including social reasoning, biomedical tasks, mathematical problem solving, and code generation. Results demonstrate that DyFlow significantly outperforms existing baselines, achieving substantial Pass@k improvements and exhibiting robust generalization across diverse domains.
Yanbo Wang 0005, Zixiang Xu, Yue Huang 0001, Xiangqi Wang, Zirui Song, Lang Gao, Chenxi Wang 0001, Robert Tang, Yue Zhao 0016, Arman Cohan, Xiangliang Zhang 0001, Xiuying Chen
NeurIPS8
2025 WebDancer: Towards Autonomous Information Seeking Agency
abstract
Addressing intricate real-world problems necessitates in-depth information seeking and multi-step reasoning. Recent progress in agentic systems, exemplified by Deep Research, underscores the potential for autonomous multi-step research. In this work, we present a cohesive paradigm for building end-to-end agentic information seeking agents from a data-centric and training-stage perspective. Our approach consists of four key stages: (1) browsing data construction, (2) trajectories sampling, (3) supervised fine-tuning for effective cold start, and (4) reinforcement learning for enhanced generalisation. We instantiate this framework in a web agent based on the ReAct format, WebDancer. Empirical evaluations on the challenging GAIA and WebWalkerQA benchmarks demonstrate the strong performance of WebDancer, achieving considerable results and highlighting the efficacy of our training paradigm. Further analysis of agent training provides valuable insights and actionable, systematic pathways for developing more capable agentic models.
Jialong Wu 0007, Baixuan Li, Runnan Fang, Wenbiao Yin, Zhenglin Wang, Zhengwei Tao, Dingchu Zhang, Zekun Xi, Robert Tang, Yong Jiang 0005, Pengjun Xie, Fei Huang 0002, Jingren Zhou 0001
NeurIPS10
2025 SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks
abstract
We present SciArena, an open and collaborative platform for evaluating foundation models on scientific literature-grounded tasks. Unlike traditional benchmarks for scientific literature understanding and synthesis, SciArena engages the research community directly, following the Chatbot Arena evaluation approach of community voting on model comparisons.By leveraging collective intelligence, SciArena offers a community-driven evaluation of model performance on open-ended scientific tasks that demand literature-grounded, long-form responses.The platform currently supports 44 open-source and proprietary foundation models and has collected over 19,000 votes from human researchers across diverse scientific domains. Our analysis of the data collected so far confirms its high quality.We discuss the results and insights based on the model ranking leaderboard.To further promote research in building model-based automated evaluation systems for literature tasks, we release SciArena-Eval, a meta-evaluation benchmark based on our collected preference data. The benchmark measures the accuracy of models in judging answer quality by comparing their pairwise assessments with human votes. Our experiments highlight the benchmark’s challenges and emphasize the need for more reliable automated evaluation methods.
Yilun Zhao 0001, Tiansheng Hu, Sihong Wu, Ronan Le Bras 0001, Yixin Liu 0003, Robert Tang, Joseph Chee Chang, Jesse Dodge, Jonathan Bragg, Chen Zhao 0013, Hannaneh Hajishirzi, Doug Downey, Arman Cohan
NeurIPS7
2010 Rocket roll dynamics and disturbance - Minimal modelling and system identification
abstract
The roll dynamics of a 5kg, 1.3 m high sounding rocket are analyzed in a vertical wind tunnel. Significant turbulence in the tunnel makes the system identification of the effective inertia, damping and asymmetry with respect to roll challenging. A novel method is developed which decouples the disturbance from the rocket frame's intrinsic roll dynamics and allows accurate prediction of roll rate and angle. The parameter identification method is integral-based, and treats wind disturbances as equivalent to a movement in the actuator fins. The method is robust, requires minimal computation, and gave a realistic disturbance distribution reflecting the randomness of the turbulent wind flow. The mean absolute roll rate of the rocket frame observed in experiments was 16.4 degree/s and the model predicted the roll rate with a median error of 0.51 degrees/s with a 90th percentile of 1.25 degrees/s. The roll angle (measured by an encoder), was tracked by the model with a median absolute error of 0.25 degrees and a 90th percentile of 0.50 degrees. These results prove the concept of this minimal modeling approach which will be extended to pitch and yaw dynamics in the future.
Christopher E. Hann, Malcolm Snowdon, Avinash Rao, Robert Tang, Agnetha Korevaar, Greg Skinner, Alex Keall, J. Geoffrey Chase
ICARCV4