Zhaochen Hong

dblp:346/4478 · DBLP profile ↗
← Back
7ranked-venue papers
1as first author
7since 2021 · last 2025
0009-0009-6565-1860ORCID · corroborated

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 1 first-author · 4 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
6 papers
Information extraction and text analysis · 35% Multi-agent systems · 22% Language models and text generation · 15%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%

Topics — the 17 heaviest of 18, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis
relation extraction
1.422024
Reading Broadly to Open Your Mind: Improving Open Relation Extraction With Search Documents Under Self-Supervisions · IEEE Trans. Knowl. Data Eng. 2024
Think Rationally about What You See: Continuous Rationale Extraction for Relation Extraction · SIGIR 2023
Knowledge, reasoning and agents › Multi-agent systems
agent-based simulation
0.912025
ResearchTown: Simulator of Human Research Community · ICML 2025
Knowledge, reasoning and agents › Knowledge representation and reasoning
consistency checking
0.912025
ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities · ACL (1) 2025
Natural language and speech › Language models and text generation
large language model evaluation
0.912025
ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities · ACL (1) 2025
Natural language and speech › Language models and text generation
LLM agents
0.912025
ResearchTown: Simulator of Human Research Community · ICML 2025
Knowledge, reasoning and agents › Multi-agent systems › LLM-based multi-agent systems
multi-agent LLM coordination
0.912025
MultiAgentBench : Evaluating the Collaboration and Competition of LLM agents · ACL (1) 2025
Natural language and speech › Information extraction and text analysis › named entity recognition
cross-domain named entity recognition
0.812024
Three Heads Are Better than One: Improving Cross-Domain NER with Progressive Decomposed Network · AAAI 2024
Machine learning › Transfer learning and domain adaptation › domain adaptation
multi-source domain adaptation
0.812024
Three Heads Are Better than One: Improving Cross-Domain NER with Progressive Decomposed Network · AAAI 2024
Machine learning › Transfer learning and domain adaptation
multi-source transfer learning
0.812024
Three Heads Are Better than One: Improving Cross-Domain NER with Progressive Decomposed Network · AAAI 2024
Natural language and speech › Information extraction and text analysis
named entity recognition
0.812024
Three Heads Are Better than One: Improving Cross-Domain NER with Progressive Decomposed Network · AAAI 2024
Natural language and speech › Information extraction and text analysis › relation extraction
open relation extraction
0.812024
Reading Broadly to Open Your Mind: Improving Open Relation Extraction With Search Documents Under Self-Supervisions · IEEE Trans. Knowl. Data Eng. 2024
Machine learning › Trustworthy machine learning › interpretability › rationalization
rationale extraction
0.712023
Think Rationally about What You See: Continuous Rationale Extraction for Relation Extraction · SIGIR 2023
Information retrieval
evidence retrieval
0.712023
Read it Twice: Towards Faithfully Interpretable Fact Verification by Revisiting Evidence · SIGIR 2023
Information retrieval
fact-checking
0.712023
Read it Twice: Towards Faithfully Interpretable Fact Verification by Revisiting Evidence · SIGIR 2023
Natural language and speech › Machine translation
machine translation evaluation
0.312025
ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities · ACL (1) 2025
Natural language and speech › Information extraction and text analysis
entity typing
0.212024
Three Heads Are Better than One: Improving Cross-Domain NER with Progressive Decomposed Network · AAAI 2024
Natural language and speech › Information extraction and text analysis › named entity recognition
mention detection
0.212024
Three Heads Are Better than One: Improving Cross-Domain NER with Progressive Decomposed Network · AAAI 2024

Methods — techniques the papers use, named apart from their topics

tree-based transformation · 0.9reversible transformations · 0.9node-masking prediction · 0.9milestone-based evaluation · 0.9message passing · 0.9coordination protocol analysis · 0.9TextGNN · 0.9LLM-generated benchmarks · 0.9pre-training and fine-tuning · 0.8knowledge distillation · 0.8evidence retrieval · 0.7claim verification · 0.7
YearPublicationVenuePosition
2025 ConsistencyChecker: Tree-based Evaluation of LLM Generalization Capabilities
abstract
Evaluating consistency in Large Language Models (LLMs) is crucial for ensuring reliability, particularly in complex, multi-step interactions between humans and LLMs.Traditional self-consistency methods often miss subtle semantic changes in natural language and functional shifts in code or equations, which can accumulate over multiple transformations.To address this, we propose ConsistencyChecker, a tree-based evaluation framework designed to measure consistency through sequences of reversible transformations, including machine translation tasks and AI-assisted programming tasks.In our proposed framework, nodes represent distinct text states, while edges correspond to pairs of inverse operations.Dynamic and LLM-generated benchmarks ensure a fair assessment of the model's generalization ability and eliminate benchmark leakage.Consistency is quantified based on similarity across different depths of the transformation tree.Experiments on eight models from various families and sizes show that ConsistencyChecker can distinguish the performance of different models.Notably, our consistency scores, computed entirely without using WMT paired data, correlate strongly (r > 0.7) with WMT 2024 auto-ranking, demonstrating the validity of our benchmark-free approach.Our implementation is available at https://github.com
Zhaochen Hong, Haofei Yu, Jiaxuan You
ACL (1)1
2025 MultiAgentBench : Evaluating the Collaboration and Competition of LLM agents
abstract
Large Language Models (LLMs) have shown remarkable capabilities as autonomous agents; yet existing benchmarks either focus on single-agent tasks or are confined to narrow domains, failing to capture the dynamics of multi-agent coordination and competition. In this paper, we introduce MultiAgentBench, a comprehensive benchmark designed to evaluate LLM-based multi-agent systems across diverse, interactive scenarios. Our framework measures not only task completion but also the quality of collaboration and competition using novel, milestone-based key performance indicators. Moreover, we evaluate various coordination protocols (including star, chain, tree, and graph topologies) and innovative strategies such as group discussion and cognitive planning. Notably, gpt-4o-mini reaches the average highest task score, graph structure performs the best among coordination protocols in the research scenario,and cognitive planning improves milestone achievement rates by 3%. Code and datasets are publicavailable at https://github.com/ulab-uiuc/MARBLE.
Kunlun Zhu, Hongyi Du, Zhaochen Hong, Xiaocheng Yang, Shuyi Guo, Zhenhailong Wang, Cheng Qian 0008, Robert Tang, Heng Ji 0001, Jiaxuan You
ACL (1)3
2025 ResearchTown: Simulator of Human Research Community
abstract
Large Language Models (LLMs) have demonstrated remarkable potential in scientific domains, yet a fundamental question remains unanswered: Can we simulate human research communities with LLMs? Addressing this question can deepen our understanding of the processes behind idea brainstorming and inspire the automatic discovery of novel scientific insights. In this work, we propose ResearchTown, a multi-agent framework for research community simulation. Within this framework, the human research community is simplified as an agent-data graph, where researchers and papers are represented as agent-type and data-type nodes, respectively, and connected based on their collaboration relationships. We also introduce TextGNN, a text-based inference framework that models various research activities (e.g., paper reading, paper writing, and review writing) as special forms of a unified message-passing process on the agent-data graph. To evaluate the quality of the research community simulation, we present ResearchBench, a benchmark that uses a node-masking prediction task for scalable and objective assessment based on similarity. Our experiments reveal three key findings: (1) ResearchTown can provide a realistic simulation of collaborative research activities, including paper writing and review writing; (2) ResearchTown can maintain robust simulation with multiple researchers and diverse papers; (3) ResearchTown can generate interdisciplinary research ideas that potentially inspire pioneering research directions.
Haofei Yu, Zhaochen Hong, Zirui Cheng, Kunlun Zhu, Keyang Xuan, Jinwei Yao, Jiaxuan You
ICML2
2024 Three Heads Are Better than One: Improving Cross-Domain NER with Progressive Decomposed Network
abstract
Cross-domain named entity recognition (NER) tasks encourage NER models to transfer knowledge from data-rich source domains to sparsely labeled target domains. Previous works adopt the paradigms of pre-training on the source domain followed by fine-tuning on the target domain. However, these works ignore that general labeled NER source domain data can be easily retrieved in the real world, and soliciting more source domains could bring more benefits. Unfortunately, previous paradigms cannot efficiently transfer knowledge from multiple source domains. In this work, to transfer multiple source domains' knowledge, we decouple the NER task into the pipeline tasks of mention detection and entity typing, where the mention detection unifies the training object across domains, thus providing the entity typing with higher-quality entity mentions. Additionally, we request multiple general source domain models to suggest the potential named entities for sentences in the target domain explicitly, and transfer their knowledge to the target domain models through the knowledge progressive networks implicitly. Furthermore, we propose two methods to analyze in which source domain knowledge transfer occurs, thus helping us judge which source domain brings the greatest benefit. In our experiment, we develop a Chinese cross-domain NER dataset. Our model improved the F1 score by an average of 12.50% across 8 Chinese and English datasets compared to models without source domain data.
Xuming Hu, Zhaochen Hong, Yong Jiang 0005, Zhichao Lin, Xiaobin Wang, Pengjun Xie, Philip S. Yu
AAAI2
2024 Reading Broadly to Open Your Mind: Improving Open Relation Extraction With Search Documents Under Self-Supervisions
abstract
Open relation extraction is the task of extracting open-domain relation facts from natural language sentences. Existing works either utilize distant-supervised annotations to train a supervised classifier over pre-defined relations, or adopt unsupervised methods with additional dependency on external assumptions. However, these works can only obtain information signals from limited existing knowledge bases or datasets. In this work, we propose a self-supervised framework namedWeb-SelfORE, which exploits self-supervised signals by requiring a large pretrained language model to extensively read real-world relevant documents from the web, and obtain contextualized relational features by mixing contextualized representations of entities from different documents. We perform adaptive clustering on contextualized relational features and bootstrap the self-supervised signals by improving contextualized features in relation classification. We additionally compare the effectiveness of self-supervisions brought by different document sources, and introduce relevance and redundancy evaluation metrics to obtain higher-quality self-supervisions. Experimental results on four public datasets show the effectiveness and robustness ofWeb-SelfOREon open-domain relation extraction task when comparing with competitive baselines.
Xuming Hu, Zhaochen Hong, Aiwei Liu, Shiao Meng, Lijie Wen 0001, Irwin King, Philip S. Yu
IEEE Trans. Knowl. Data Eng.2
2023 Read it Twice: Towards Faithfully Interpretable Fact Verification by Revisiting Evidence
abstract
Real-world fact verification task aims to verify the factuality of a claim by retrieving evidence from the source document. The quality of the retrieved evidence plays an important role in claim verification. Ideally, the retrieved evidence should be faithful (reflecting the model's decision-making process in claim verification) and plausible (convincing to humans), and can improve the accuracy of verification task. Although existing approaches leverage the similarity measure of semantic or surface form between claims and documents to retrieve evidence, they all rely on certain heuristics that prevent them from satisfying all three requirements. In light of this, we propose a fact verification model named ReRead to retrieve evidence and verify claim that: (1) Train the evidence retriever to obtain interpretable evidence (i.e., faithfulness and plausibility criteria); (2) Train the claim verifier to revisit the evidence retrieved by the optimized evidence retriever to improve the accuracy. The proposed system is able to achieve significant improvements upon best-reported models under different settings.
Xuming Hu, Zhaochen Hong, Zhijiang Guo, Lijie Wen 0001, Philip S. Yu
SIGIR2
2023 Think Rationally about What You See: Continuous Rationale Extraction for Relation Extraction
abstract
Relation extraction (RE) aims to extract potential relations according to the context of two entities, thus, deriving rational contexts from sentences plays an important role. Previous works either focus on how to leverage the entity information (e.g., entity types, entity verbalization) to inference relations, but ignore context-focused content, or use counterfactual thinking to remove the model's bias of potential relations in entities, but the relation reasoning process will still be hindered by irrelevant content. Therefore, how to preserve relevant content and remove noisy segments from sentences is a crucial task. In addition, retained content needs to be fluent enough to maintain semantic coherence and interpretability. In this work, we propose a novel rationale extraction framework named RE2, which leverages two continuity and sparsity factors to obtain relevant and coherent rationales from sentences. To solve the problem that the gold rationales are not labeled, RE2 applies an optimizable binary mask to each token in the sentence, and adjust the rationales that need to be selected according to the relation label. Experiments on four datasets show that RE2 surpasses baselines.
Xuming Hu, Zhaochen Hong, Irwin King, Philip S. Yu
SIGIR2