Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Wenzhen Zheng

dblp:382/4761 · DBLP profile ↗
← Back
5ranked-venue papers
1as first author
5since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 5 · 1 first-author · 5 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
5 papers
Language models and text generation · 28% Information extraction and text analysis · 21% Deep learning architectures and training · 20%
Human-computer interaction and pervasive computing
1 paper
Collaborative and social computing · 100%

Topics — the 13 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Machine learning › Deep learning architectures and training
scaling laws
2.632026
Scaling Laws for Code: A More Data-Hungry Regime · ACL (1) 2026
Predictable Scale (Part II) - Farseer: A Refined Scaling Law in LLMs · NeurIPS 2025
Breaking Language Barriers: Cross-Lingual Continual Pre-Training at Scale · EMNLP 2024
Natural language and speech › Information extraction and text analysis
misinformation detection
1.922026
Beyond Detection: Exploring Evidence-based Multi-Agent Debate for Misinformation Intervention and Persuasion · AAAI 2026
Debate-to-Detect: Reformulating Misinformation Detection as a Real-World Debate with Large Language Models · EMNLP 2025
Knowledge, reasoning and agents › Multi-agent systems › LLM-based multi-agent systems
multi-agent debate
1.922026
Beyond Detection: Exploring Evidence-based Multi-Agent Debate for Misinformation Intervention and Persuasion · AAAI 2026
Debate-to-Detect: Reformulating Misinformation Detection as a Real-World Debate with Large Language Models · EMNLP 2025
Natural language and speech › Language models and text generation
code language models
1.012026
Scaling Laws for Code: A More Data-Hungry Regime · ACL (1) 2026
Machine learning › Representation and self-supervised learning › pre-training
pretraining data
1.012026
Scaling Laws for Code: A More Data-Hungry Regime · ACL (1) 2026
Machine learning › Efficient and distributed learning › efficient training
compute-optimal training
0.912025
Predictable Scale (Part II) - Farseer: A Refined Scaling Law in LLMs · NeurIPS 2025
Natural language and speech › Information extraction and text analysis
fact-checking
0.912025
Debate-to-Detect: Reformulating Misinformation Detection as a Real-World Debate with Large Language Models · EMNLP 2025
Natural language and speech › Language models and text generation
large language model training
0.912025
Predictable Scale (Part II) - Farseer: A Refined Scaling Law in LLMs · NeurIPS 2025
Natural language and speech › Language models and text generation › large language model training
continual pre-training
0.812024
Breaking Language Barriers: Cross-Lingual Continual Pre-Training at Scale · EMNLP 2024
Natural language and speech › Language models and text generation › multilingual language models
multilingual pretrained language model
0.812024
Breaking Language Barriers: Cross-Lingual Continual Pre-Training at Scale · EMNLP 2024
Collaborative and social computing › social influence
persuasion
0.312026
Beyond Detection: Exploring Evidence-based Multi-Agent Debate for Misinformation Intervention and Persuasion · AAAI 2026
Natural language and speech › Language models and text generation
large language model reasoning
0.312025
Debate-to-Detect: Reformulating Misinformation Detection as a Real-World Debate with Large Language Models · EMNLP 2025
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer
0.212024
Breaking Language Barriers: Cross-Lingual Continual Pre-Training at Scale · EMNLP 2024

Methods — techniques the papers use, named apart from their topics

large language model · 2.0evidence retrieval · 2.0scaling law analysis · 1.0scaling law fitting · 0.9multi-agent debate · 0.9loss surface modeling · 0.9LLM · 0.9data replay · 0.8continual pre-training · 0.8
YearPublicationVenuePosition
2026 Beyond Detection: Exploring Evidence-based Multi-Agent Debate for Misinformation Intervention and Persuasion
abstract
Multi-agent debate (MAD) frameworks have emerged as promising approaches for misinformation detection by simulating adversarial reasoning. While prior work has focused on detection accuracy, the importance of helping users understand the reasoning behind factual judgments has been overlooked. The debate transcripts generated during MAD offer a rich but underutilized resource for transparent reasoning. In this study, we introduce ED2D, an evidence-based MAD framework that extends previous approach by incorporating factual evidence retrieval. More importantly, ED2D is designed not only as a detection framework but also as a persuasive multi-agent system aimed at correcting user beliefs and discouraging misinformation sharing. We compare the persuasive effects of ED2D-generated debunking transcripts with those authored by human experts. Results demonstrate that ED2D outperforms existing baselines across three misinformation detection benchmarks. When ED2D generates correct predictions, its debunking transcripts exhibit persuasive effects comparable to those of human experts; However, when ED2D misclassifies, its accompanying explanations may inadvertently reinforce users’ misconceptions, even when presented alongside accurate human explanations. Our findings highlight both the promise and the potential risks of deploying MAD systems for misinformation intervention. We further develop a public community website to help users explore ED2D, fostering transparency, critical thinking, and collaborative fact-checking.
Chen Han 0003, Yijia Ma, Wenzhen Zheng, Xijin Tang 0001
AAAI4
2026 Scaling Laws for Code: A More Data-Hungry Regime
abstract
Xianzhen Luo, Wenzhen Zheng, Qingfu Zhu, Rongyi Zhang, Houyi Li, Siming Huang, YuanTao Fan, Wanxiang Che. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xianzhen Luo, Wenzhen Zheng, Qingfu Zhu, Rongyi Zhang, Houyi Li, Siming Huang, YuanTao Fan, Wanxiang Che
ACL (1)2
2025 Debate-to-Detect: Reformulating Misinformation Detection as a Real-World Debate with Large Language Models
abstract
The proliferation of misinformation in digital platforms reveals the limitations of traditional detection methods, which mostly rely on static classification and fail to capture the intricate process of real-world fact-checking.Despite advancements in Large Language Models (LLMs) that enhance automated reasoning, their application to misinformation detection remains hindered by issues of logical inconsistency and superficial verification.Inspired by the idea that "Truth Becomes Clearer Through Debate", we introduce Debate-to-Detect (D2D), a novel Multi-Agent Debate (MAD) framework that reformulates misinformation detection as a structured adversarial debate.Based on fact-checking workflows, D2D assigns domain-specific profiles to each agent and orchestrates a five-stage debate process, including Opening Statement, Rebuttal, Free Debate, Closing Statement, and Judgment.To transcend traditional binary classification, D2D introduces a multi-dimensional evaluation mechanism that assesses each claim across five distinct dimensions: Factuality, Source Reliability, Reasoning Quality, Clarity, and Ethics.Experiments with GPT-4o on two fakenews datasets demonstrate significant improvements over baseline methods, and the case study highlight D2D's capability to iteratively refine evidence while improving decision transparency, representing a substantial advancement towards robust and interpretable misinformation detection.Our code is available at Debate-to-Detect.
Chen Han 0003, Wenzhen Zheng, Xijin Tang 0001
EMNLP2
2025 Predictable Scale (Part II) - Farseer: A Refined Scaling Law in LLMs
abstract
Training Large Language Models (LLMs) is prohibitively expensive, creating a critical scaling gap where insights from small-scale experiments often fail to transfer to resource-intensive production systems, thereby hindering efficient innovation. To bridge this, we introduce Farseer, a novel and refined scaling law offering enhanced predictive accuracy across scales. By systematically constructing a model loss surface $L(N,D)$, Farseer achieves a significantly better fit to empirical data than prior laws (e.g., \Chinchilla's law). Our methodology yields accurate, robust, and highly generalizable predictions, demonstrating excellent extrapolation capabilities, outperforming Chinchilla's law, whose extrapolation error is 433\% higher. This allows for the reliable evaluation of competing training strategies across all $(N,D)$ settings, enabling conclusions from small-scale ablation studies to be confidently extrapolated to predict large-scale performance. Furthermore, Farseer provides new insights into optimal compute allocation, better reflecting the nuanced demands of modern LLM training. To validate our approach, we trained an extensive suite of approximately 1,000 LLMs across diverse scales and configurations, consuming roughly 3 million NVIDIA H100 GPU hours. To foster further research, we are comprehensively open-sourcing all code, data, results (https://github.com/Farseer-Scaling-Law/Farseer), all training logs (https://wandb.ai/billzid/Farseer?nw=nwuserbillzid), all models used in scaling law fitting (https://huggingface.co/Farseer-Scaling-Law).
Houyi Li, Wenzhen Zheng, Zhenyu Ding, Haoying Wang, Shijie Xuyang, Ning Ding 0006, Shuigeng Zhou, Xiangyu Zhang 0005, Daxin Jiang
NeurIPS2
2024 Breaking Language Barriers: Cross-Lingual Continual Pre-Training at Scale
abstract
In recent years, Large Language Models (LLMs) have made significant strides towards Artificial General Intelligence.However, training these models from scratch requires substantial computational resources and vast amounts of text data.In this paper, we explore an alternative approach to constructing an LLM for a new language by continually pre-training (CPT) from existing pre-trained LLMs, instead of using randomly initialized parameters.Based on parallel experiments on 40 model sizes ranging from 40M to 5B parameters, we find that 1) CPT converges faster and saves significant resources in a scalable manner; 2) CPT adheres to an extended scaling law derived from Hoffmann et al. ( 2022) with a joint data-parameter scaling term; 3) The compute-optimal data-parameter allocation for CPT markedly differs based on our estimated scaling factors; 4) The effectiveness of transfer at scale is influenced by training duration and linguistic properties, while robust to data replaying, a method that effectively mitigates catastrophic forgetting in CPT.We hope our findings provide deeper insights into the transferability of LLMs at scale for the research community.
Wenzhen Zheng, Wenbo Pan 0001, Libo Qin 0001, Li Yue 0010
EMNLP1