Terry Jingchen Zhang

dblp:408/5364 · DBLP profile ↗
← Back
4ranked-venue papers
1as first author
4since 2021 · last 2026
0009-0009-1271-4696ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Vision and language · 43% Language models and text generation · 23% Multi-agent systems · 23%
Software engineering, system software, and programming languages
1 paper
Empirical software engineering · 100%

Topics — the 8 heaviest of 10, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › large language model evaluation
benchmark contamination
1.012026
Test of Time: Rethinking Temporal Signal of Benchmark Contamination · ACL (1) 2026
Natural language and speech › Language models and text generation
chain-of-thought reasoning
1.012026
AtomThink: Multimodal Slow Thinking With Atomic Step Reasoning · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Knowledge, reasoning and agents › Multi-agent systems › multi-agent coordination
distributed coordination
1.012026
SILO-BENCH: A Scalable Environment for Evaluating Distributed Coordination in Multi-Agent LLM Systems · ACL (1) 2026
Computer vision › Vision and language › vision-language model
multimodal large language model
1.012026
AtomThink: Multimodal Slow Thinking With Atomic Step Reasoning · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Computer vision › Vision and language
multimodal reasoning
1.012026
AtomThink: Multimodal Slow Thinking With Atomic Step Reasoning · IEEE Trans. Pattern Anal. Mach. Intell. 2026
Computer vision › Vision and language
multimodal benchmark
0.912025
SeePhys: Does Seeing Help Thinking? - Benchmarking Vision-Based Physics Reasoning · NeurIPS 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning › commonsense reasoning
physical reasoning
0.912025
SeePhys: Does Seeing Help Thinking? - Benchmarking Vision-Based Physics Reasoning · NeurIPS 2025
Computer vision › Vision and language
visual reasoning
0.912025
SeePhys: Does Seeing Help Thinking? - Benchmarking Vision-Based Physics Reasoning · NeurIPS 2025

Methods — techniques the papers use, named apart from their topics

temporal analysis · 2.0large language model · 1.9supervised fine-tuning · 1.0reinforcement learning · 1.0multimodal model · 0.9
YearPublicationVenuePosition
2026 Test of Time: Rethinking Temporal Signal of Benchmark Contamination
abstract
Terry Jingchen Zhang, Gopal Dev, Ning Wang, Max Obreiter, Punya Syon Pandey, Keenan Samway, Wenyuan Jiang, Yinya Huang, Bernhard Schölkopf, Mrinmaya Sachan, Zhijing Jin. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Terry Jingchen Zhang, Gopal Dev, Max Obreiter, Wenyuan Jiang, Punya Syon Pandey, Keenan Samway, Yinya Huang, Bernhard Schölkopf, Mrinmaya Sachan, Zhijing Jin 0001
ACL (1)1
2026 SILO-BENCH: A Scalable Environment for Evaluating Distributed Coordination in Multi-Agent LLM Systems
abstract
Yuzhe Zhang, Feiran Liu, Yi Shan, Xinyi Huang, Xin Yang, Yueqi Zhu, Xuxin Cheng, Cao Liu, Ke Zeng, Terry Jingchen Zhang, Wenyuan Jiang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Feiran Liu, Yi Shan 0001, Yueqi Zhu, Xuxin Cheng, Cao Liu, Terry Jingchen Zhang, Wenyuan Jiang
ACL (1)10
2026 AtomThink: Multimodal Slow Thinking With Atomic Step Reasoning
abstract
In this paper, we address the challenging task of multimodal reasoning by incorporating the notion of "slow thinking" into multimodal large language models (MLLMs). Our core idea is that models can learn to adaptively use different levels of reasoning to tackle questions of varying complexity. We propose a novel paradigm of Self-structured Chain of Thought (SCoT), which consists of minimal semantic atomic steps. Unlike existing methods that rely on structured templates or free-form paradigms, our method not only generates flexible CoT structures for various complex tasks but also mitigates the phenomenon of overthinking for easier tasks. To introduce structured reasoning into visual cognition, we design a novel AtomThink framework with four key modules: (i) a data engine to generate high-quality multimodal reasoning paths; (ii) a supervised fine-tuning (SFT) process with serialized inference data; (iii) a policy-guided multi-turn inference method; and (iv) an atomic capability metric to evaluate the single-step utilization rate. Extensive experiments demonstrate that the proposed AtomThink significantly improves the performance of baseline MLLMs, achieving more than 10% average accuracy gains on MathVista and MathVerse. Compared to state-of-the-art structured CoT approaches, our method not only achieves higher accuracy but also improves data utilization by 5 × and boosts inference efficiency by 85.3%.
Kun Xiang, Zhili Liu, Terry Jingchen Zhang, Yinya Huang, Yunshuang Nie, Kaixin Cai, Yiyang Yin, Runhui Huang, Yihan Zeng, Yu-Jie Yuan, Jianhua Han, Lanqing Hong, Hang Xu 0004, Xiaodan Liang
IEEE Trans. Pattern Anal. Mach. Intell.3
2025 SeePhys: Does Seeing Help Thinking? - Benchmarking Vision-Based Physics Reasoning
abstract
We present SeePhys, a large-scale multimodal benchmark for LLM reasoning grounded in physics questions ranging from middle school to PhD qualifying exams. The benchmark covers 7 fundamental domains spanning the physics discipline, incorporating 21 categories of highly heterogeneous diagrams. In contrast to prior works where visual elements mainly serve auxiliary purposes, our benchmark features a substantial proportion of vision-essential problems (75%) that mandate visual information extraction for correct solutions. Through extensive evaluation, we observe that even the most advanced visual reasoning models (e.g., Gemini-2.5-pro and o4-mini) achieve sub-60% accuracy on our benchmark. These results reveal fundamental challenges in current large language models' visual understanding capabilities, particularly in: (i) establishing rigorous coupling between diagram interpretation and physics reasoning, and (ii) overcoming their persistent reliance on textual cues as cognitive shortcuts.Project Page: github.com/SeePhys/seephys-projectHugging Face: huggingface.co/datasets/SeePhys/SeePhys
Kun Xiang, Terry Jingchen Zhang, Yinya Huang, Zirong Liu, Peixin Qu, Jixi He, Yu-Jie Yuan, Jianhua Han, Hang Xu 0004, Mrinmaya Sachan, Xiaodan Liang
NeurIPS3