Juntian Zhang

dblp:242/2198 · DBLP profile ↗
← Back
6ranked-venue papers
1as first author
6since 2021 · last 2026
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Vision and language · 38% Efficient and distributed learning · 20% Multi-agent systems · 18%
Databases, data mining, and information retrieval
1 paper
Recommender systems · 100%

Topics — the 10 heaviest of 12, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Efficient and distributed learning
inference efficiency
1.012026
Forest Before Trees: Latent Superposition for Efficient Visual Reasoning · ACL (1) 2026
Natural language and speech › Language models and text generation › large language model reasoning › inference-time reasoning
latent reasoning
1.012026
Forest Before Trees: Latent Superposition for Efficient Visual Reasoning · ACL (1) 2026
Computer vision › Vision and language
multimodal benchmark
1.012026
DiningBench: A Hierarchical Multi-view Benchmark for Perception and Reasoning in the Dietary Domain · ACL (1) 2026
Natural language and speech › Question answering and dialogue systems
multimodal question answering
1.012026
DiningBench: A Hierarchical Multi-view Benchmark for Perception and Reasoning in the Dietary Domain · ACL (1) 2026
Machine learning › Efficient and distributed learning
token reduction
1.012026
Forest Before Trees: Latent Superposition for Efficient Visual Reasoning · ACL (1) 2026
Computer vision › Vision and language
visual reasoning
1.012026
Forest Before Trees: Latent Superposition for Efficient Visual Reasoning · ACL (1) 2026
Knowledge, reasoning and agents › Multi-agent systems › LLM-based multi-agent systems
LLM-based agent simulation
0.912025
Beyond Static Testbeds: An Interaction-Centric Agent Simulation Platform for Dynamic Recommender Systems · EMNLP 2025
Computer vision › Vision and language › vision-language model
vision-language model training
0.912025
Weaving Context Across Images: Improving Vision-Language Models through Focus-Centric Visual Chains · ACL (1) 2025
Recommender systems
recommender system evaluation
0.912025
Beyond Static Testbeds: An Interaction-Centric Agent Simulation Platform for Dynamic Recommender Systems · EMNLP 2025
Natural language and speech › Language models and text generation
LLM agents
0.312025
The Stepwise Deception: Simulating the Evolution from True News to Fake News with LLM Agents · EMNLP 2025

Methods — techniques the papers use, named apart from their topics

user profiling · 1.7large language model fine-tuning · 1.7chain-of-thought · 1.7self-refined superposition · 1.0hierarchical multi-view benchmark · 1.0dynamic windowed alignment learning · 1.0truth deviation evaluation · 0.9instruction tuning · 0.9data synthesis · 0.9LLM-based simulation · 0.9
YearPublicationVenuePosition
2026 DiningBench: A Hierarchical Multi-view Benchmark for Perception and Reasoning in the Dietary Domain
abstract
Song Jin, Juntian Zhang, Xun Zhang, Zeying Tian, Fei Jiang, Guojun Yin, Wei Lin, Yong Liu, Rui Yan. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Juntian Zhang, Zeying Tian, Guojun Yin, Rui Yan 0001
ACL (1)2
2026 Forest Before Trees: Latent Superposition for Efficient Visual Reasoning
abstract
While Chain-of-Thought empowers Large Vision-Language Models with multi-step reasoning, explicit textual rationales suffer from an information bandwidth bottleneck, where continuous visual details are discarded during discrete tokenization.Recent latent reasoning methods attempt to address this challenge, but often fall prey to premature semantic collapse due to rigid autoregressive objectives.In this paper, we propose Laser, a novel paradigm that reformulates visual deduction via Dynamic Windowed Alignment Learning(DWAL).Instead of forcing a point-wise prediction, Laser aligns the latent state with a dynamic validity window of future semantics.This mechanism enforces a "Forest-before-Trees" cognitive hierarchy, enabling the model to maintain a probabilistic superposition of global features before narrowing down to local details.Crucially, Laser maintains interpretability via decodable trajectories while stabilizing unconstrained learning via Self-Refined Superposition.Extensive experiments on 6 benchmarks demonstrate that Laser achieves state-of-the-art performance among latent reasoning methods, surpassing the strong baseline Monet by 5.03% on average.Notably, it achieves these gains with extreme efficiency, reducing inference tokens by more than 97%, while demonstrating robust generalization to out-of-distribution domains.We hope this work encourages a paradigm shift from explicit next-token prediction to latent visual reasoning: Laser
Juntian Zhang, Nils Lukas, Yuhan Liu 0030
ACL (1)2
2025 Weaving Context Across Images: Improving Vision-Language Models through Focus-Centric Visual Chains
abstract
Vision-language models (VLMs) achieve remarkable success in single-image tasks.However, real-world scenarios often involve intricate multi-image inputs, leading to a notable performance decline as models struggle to disentangle critical information scattered across complex visual features.In this work, we propose Focus-Centric Visual Chain, a novel paradigm that enhances VLMs' perception, comprehension, and reasoning abilities in multi-image scenarios.To facilitate this paradigm, we propose Focus-Centric Data Synthesis, a scalable bottom-up approach for synthesizing high-quality data with elaborate reasoning paths.Through this approach, We construct VISC-150K, a large-scale dataset with reasoning data in the form of Focus-Centric Visual Chain, specifically designed for multi-image tasks.Experimental results on seven multi-image benchmarks demonstrate that our method achieves average performance gains of 3.16% and 2.24% across two distinct model architectures, without compromising the general vision-language capabilities.Our study represents a significant step toward more robust and capable vision-language systems that can handle complex visual scenarios: VISC. * Corresponding authors.Which of the following images contains the same object as the first image and shares the same attribute weight?
Juntian Zhang, Chuanqi Cheng, Yuhan Liu 0023, Wei Liu 0302, Jian Luan 0001, Rui Yan 0001
ACL (1)1
2025 Beyond Static Testbeds: An Interaction-Centric Agent Simulation Platform for Dynamic Recommender Systems
abstract
Evaluating and iterating upon recommender systems is crucial, yet traditional A/B testing is resource-intensive, and offline methods struggle with dynamic user-platform interactions.While agent-based simulation is promising, existing platforms often lack a mechanism for user actions to dynamically reshape the environment.To bridge this gap, we introduce RecInter, a novel agent-based simulation platform for recommender systems featuring a robust interaction mechanism.In RecInter platform, simulated user actions (e.g., likes, reviews, purchases) dynamically update item attributes in real-time, and introduced Merchant Agents can reply, fostering a more realistic and evolving ecosystem.High-fidelity simulation is ensured through Multidimensional User Profiling module, Advanced Agent Architecture, and LLM fine-tuned on Chain-of-Thought (CoT) enriched interaction data.Our platform achieves significantly improved simulation credibility and successfully replicates emergent phenomena like Brand Loyalty and the Matthew Effect.Experiments demonstrate that this interaction mechanism is pivotal for simulating realistic system evolution, establishing our platform as a credible testbed for recommender systems research: RecInter.
Juntian Zhang, Yuhan Liu 0023, Guojun Yin, Rui Yan 0001
EMNLP2
2025 The Stepwise Deception: Simulating the Evolution from True News to Fake News with LLM Agents
abstract
With the growing spread of misinformation online, understanding how true news evolves into fake news has become crucial for early detection and prevention.However, previous research has often assumed fake news inherently exists rather than exploring its gradual formation.To address this gap, we propose FUSE (Fake news evolUtion Simulation framEwork), a novel Large Language Model (LLM)-based simulation approach explicitly focusing on fake news evolution from real news.Our framework model a social network with four distinct types of LLM agents commonly observed in daily interactions: spreaders who propagate information, commentators who provide interpretations, verifiers who fact-check, and bystanders who observe passively to simulate realistic daily interactions that progressively distort true news.To quantify these gradual distortions, we develop FUSE-EVAL, a comprehensive evaluation framework measuring truth deviation along multiple linguistic and semantic dimensions.Results show that FUSE effectively captures fake news evolution patterns and accurately reproduces known fake news, aligning closely with human evaluations.Experiments demonstrate that FUSE accurately reproduces known fake news evolution scenarios, aligns closely with human judgment, and highlights the importance of timely intervention at early stages.Our framework is extensible, enabling future research on broader scenarios of fake news: FUSE.
Yuhan Liu 0030, Zirui Song, Juntian Zhang, Xiaoqing Zhang 0017, Xiuying Chen, Rui Yan 0001
EMNLP3
2025 CausalPrism: A visual analytics approach for subgroup-based causal heterogeneity exploration
Xingyu Liu 0003, Jiehui Zhou, Xumeng Wang, Kamkwai Wong, Wei Zhang 0219, Juntian Zhang, Minfeng Zhu 0001, Wei Chen 0001
Comput. Graph.6