EDBT 2026 Demo / reviewers in the wild / expert
Yilun Zhao 0001
dblp:271/8391-1
· DBLP profile ↗
48ranked-venue papers
15as first author
48since 2021 · last 2026
0000-0002-7470-6124ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 46 · 14 first-author · 46 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 first-author · 3 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | SciMDR: Advancing Scientific Multimodal Document ReasoningabstractZiyu Chen, Yilun Zhao, Chengye Wang, Rilyn R. Han, Manasi Patwardhan, Arman Cohan. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yilun Zhao 0001, Chengye Wang, Rilyn Han, Manasi Patwardhan 0001, Arman Cohan |
ACL (1) | 2 |
| 2026 | Experience Retrieval-Augmentation with Electronic Health Records Enables Accurate Discharge QAabstractJustice Ou, Tinglin Huang, Yilun Zhao, Ziyang Yu, Peiqing Lu, Yifei Shen, Rex Ying. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Justice Ou, Tinglin Huang 0001, Yilun Zhao 0001, Peiqing Lu, Yifei Shen 0006, Rex Ying |
ACL (1) | 3 |
| 2026 | MultiFinBen: Benchmarking Large Language Models for Multilingual and Multimodal Financial ApplicationabstractXueqing Peng, Lingfei Qian, Yan Wang, Ruoyu Xiang, Yueru He, Yang Ren, Mingyang Jiang, Vincent Jim Zhang, Yuqing Guo, Jeff Zhao, Huan He, Yi Han, Yun Feng, Yuechen Jiang, Yupeng Cao, Haohang Li, Yangyang Yu, Xiaoyu Wang, Penglei Gao, Shengyuan Lin, Keyi Wang, Shanshan Yang, Yilun Zhao, Zhiwei Liu, Peng Lu, Jerry Huang, Suyuchen Wang, Triantafillos Papadopoulos, Polydoros Giannouris, Efstathia Soufleri, Nuo Chen, Zhiyang Deng, Heming Fu, Yijia Zhao, Mingquan Lin, Meikang Qiu, Kaleb E Smith, Arman Cohan, Xiao-Yang Liu, Jimin Huang, Guojun Xiong, Alejandro Lopez-Lira, Xi Chen, Junichi Tsujii, Jian-Yun Nie, Sophia Ananiadou, Qianqian Xie. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Xueqing Peng, Lingfei Qian, Yan Wang 0015, Ruoyu Xiang, Yueru He, Mingyang Jiang, Vincent Jim Zhang, Jeff Zhao, Yuechen Jiang, Yupeng Cao, Haohang Li, Yangyang Yu, Penglei Gao, Shengyuan Lin, Yilun Zhao 0001, Zhiwei Liu 0003, Peng Lu 0006, Jerry Huang, Suyuchen Wang, Triantafillos Papadopoulos, Polydoros Giannouris, Efstathia Soufleri, Nuo Chen 0002, Zhiyang Deng, Heming Fu, Yijia Zhao, Mingquan Lin, Meikang Qiu, Kaleb E. Smith, Arman Cohan, Xiao-Yang Liu, Jimin Huang, Guojun Xiong, Alejandro Lopez-Lira, Xi Chen 0003, Jun'ichi Tsujii, Jian-Yun Nie, Sophia Ananiadou, Qianqian Xie |
ACL (1) | 23 |
| 2026 | Rethinking Composed Image Retrieval Evaluation: A Fine-Grained Benchmark from Image EditingabstractTingyu Song, Yanzhao Zhang, Mingxin Li, Zhuoning Guo, Dingkun Long, Pengjun Xie, Siyue Zhang, Yilun Zhao, Shu Wu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Tingyu Song, Yanzhao Zhang, Zhuoning Guo, Dingkun Long, Pengjun Xie, Siyue Zhang, Yilun Zhao 0001 |
ACL (1) | 8 |
| 2026 | TexOCR: Advancing Document OCR Models for Compilable Page-to-LaTeX ReconstructionabstractExisting document OCR largely targets plain text or Markdown, discarding the structural and executable properties that make LaTeX essential for scientific publishing.We study page-level reconstruction of scientific PDFs into compilable LaTeX and introduce TEX-OCR-Bench, a benchmark, and TEXOCR-Train, a large-scale training corpus, for this task.TEXOCR-Bench features a multi-dimensional evaluation suite that jointly assesses transcription fidelity, structural faithfulness, and endto-end compilability.Leveraging TEXOCR-Train, we train a 2B-parameter model, TEX-OCR, using supervised fine-tuning (SFT) and reinforcement learning (RL) with verifiable rewards derived from LaTeX unit tests that directly enforce compilability and referential integrity.Experiments across 21 frontier models on TEXOCR-Bench show that existing systems frequently violate key document invariants, including consistent section structure, correct float placement, and valid label-reference links, which undermines compilation reliability and downstream usability.Our analysis further reveals that RL with verifiable rewards yields consistent improvements over SFT alone, particularly on structural and compilation metrics. Chengye Wang, Zexi Kuang, Yilun Zhao 0001 |
ACL (1) | 4 |
| 2026 | A Survey of Reasoning-Intensive Retrieval: Progress and ChallengesabstractReasoning-Intensive Retrieval (RIR) targets retrieval settings where relevance is mediated by latent inferential links between a query and supporting evidence, rather than semantic similarity.Motivated by the emergent reasoning abilities of Large Language Models (LLMs), recent work integrates these capabilities into the IR field, spanning the entire pipeline from benchmarks to retrievers and rerankers.Despite this progress, the field lacks a systematic framework to organize current efforts and articulate a clear path forward.To provide a clear roadmap for this rapidly growing yet fragmented area, this survey (1) systematizes existing RIR benchmarks by knowledge domains and modalities, providing a detailed analysis of the current landscape; (2) introduces a structured taxonomy that categorizes methods based on where and how reasoning is integrated into the retrieval pipeline, alongside an analysis of their tradeoffs and practical applications; and (3) summarizes challenges and future directions to guide research in this evolving field. Yiyang Wei, Tingyu Song, Siyue Zhang, Yilun Zhao 0001 |
ACL (1) | 4 |
| 2026 | Anchor: Branch-Point Data Generation for GUI AgentsabstractEnd-to-end GUI agents for real desktop environments require large amounts of highquality interaction data, yet collecting human demonstrations is expensive and existing synthetic pipelines often suffer from limited task diversity or noisy, goal-drifting trajectories.We present a trajectory expansion framework ANCHOR that bootstraps scalable desktop supervision from a small set of verified seed demonstrations.Starting from each seed, we identify branch points that correspond to meaningful state changes and propose new, stategrounded task variants conditioned on the current GUI context.An executing agent then follows the proposed instructions to generate new trajectories, while a verifier enforces task completion via state-aware checks and trajectorylevel consistency.To improve supervision quality, we further apply task-conditioned steplevel filtering to remove ungrounded actions and denoise post-branch segments to maintain coherent intent.Experiments on standard desktop benchmarks, OSWorld and Win-dowsAgentArena, show that models fine-tuned on our expanded corpus achieve consistent improvements over zero-shot agents and representative synthesis baselines, and generalize across applications and operating systems. Jinbiao Wei, Yilun Zhao 0001, Kangqi Ni, Arman Cohan |
ACL (1) | 2 |
| 2026 | Can AI Be a Good Peer Reviewer? A Survey of Peer Review Process, Evaluation, and the FutureabstractSihong Wu, Owen Jiang, Yilun Zhao, Tiansheng Hu, Yiling Ma, Kaiyan Zhang, Manasi Patwardhan, Arman Cohan. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Sihong Wu, Owen Jiang, Yilun Zhao 0001, Tiansheng Hu, Yiling Ma, Manasi Patwardhan 0001, Arman Cohan |
ACL (1) | 3 |
| 2026 | MMSciCode: Real-world Evaluation of Multilingual Multi-Discipline Scientific Research CodingabstractWe introduce MMSciCode, a comprehensive expert-level, multilingual multi-discipline benchmark for evaluating foundation models in scientific code generation.It includes 624 expert-annotated research coding problems spanning six core scientific disciplines.Compared to prior benchmarks, MMSciCode features three key advancements.First, it challenges models to integrate domain-specific knowledge with algorithmic reasoning to implement core functions from research papers.Second, each problem is meticulously annotated by domain experts through a rigorous papergrounded process, with strict quality controls implemented to ensure dataset integrity and authenticity.Finally, each problem is equipped with comprehensive unit test suites and containerized environments, enabling reproducible and diagnostic evaluation of both functional correctness and domain validity.We conduct an extensive evaluation of 23 state-of-the-art foundation models and 2 coding agents on MMSciCode.We identify substantial performance gaps between models and human experts, providing actionable insights for advancing expert-level scientific code generation. Zheyuan Yang, Arman Cohan, Yilun Zhao 0001 |
ACL (1) | 4 |
| 2026 | A Survey of Multimodal Mathematical Reasoning: From Perception, Alignment to ReasoningabstractTianyu Yang, Sihong Wu, Yilun Zhao, Zhenwen Liang, Lisen Dai, Chen Zhao, Minhao Cheng, Arman Cohan, Xiangliang Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Sihong Wu, Yilun Zhao 0001, Zhenwen Liang, Lisen Dai, Chen Zhao 0013, Minhao Cheng, Arman Cohan, Xiangliang Zhang 0001 |
ACL (1) | 3 |
| 2026 | Rethinking Reasoning-Intensive Retrieval: Evaluating and Advancing Retrievers in Agentic Search SystemsabstractReasoning-intensive retrieval aims to surface evidence that supports downstream reasoning rather than merely matching topical similarity.This capability is increasingly important for agentic search systems, where retrievers must provide complementary evidence across iterative search and synthesis.However, existing work remains limited on both evaluation and training: benchmarks such as BRIGHT provide narrow gold sets and evaluate retrievers in isolation, while synthetic training corpora often optimize single-passage relevance rather than evidence portfolio construction.We introduce BRIGHT-PRO, an expert-annotated benchmark that expands each query with multi-aspect gold evidence and evaluates retrievers under both static and agentic search protocols.We further construct RTriever-Synth, an aspect-decomposed synthetic corpus that generates complementary positives and positive-conditioned hard negatives, and use it to LoRA fine-tune RTriever-4B from Qwen3-Embedding-4B.Experiments across lexical, general-purpose, and reasoningintensive retrievers show that aspect-aware and agentic evaluation expose behaviors hidden by standard metrics, while RTriever-4B substantially improves over its base model. Yilun Zhao 0001, Jinbiao Wei, Tingyu Song, Siyue Zhang, Chen Zhao 0013, Arman Cohan |
ACL (1) | 1 |
| 2025 | AbGen: Evaluating Large Language Models in Ablation Study Design and Evaluation for Scientific ResearchabstractYilun Zhao, Weiyuan Chen, Zhijian Xu, Manasi Patwardhan, Chengye Wang, Yixin Liu, Lovekesh Vig, Arman Cohan. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Yilun Zhao 0001, Weiyuan Chen, Manasi Patwardhan 0001, Chengye Wang, Yixin Liu 0003, Lovekesh Vig, Arman Cohan |
ACL (1) | 1 |
| 2025 | VF-Eval: Evaluating Multimodal LLMs for Generating Feedback on AIGC VideosabstractRecently, multimodal large language models (MLLMs) have been extensively explored in video question answering. However, most existing assessments focus on natural videos, overlooking synthetic videos (e.g., AI-generated content). Meanwhile, some works in video generation rely on MLLMs to evaluate the quality of generated videos, but the capabilities of MLLMs on AIGC videos remain largely underexplored. To address this, we propose a new benchmark, VQ-Eval, which introduces four tasks—coherence validation, error awareness, error type detection, and reasoning evaluation—to comprehensively evaluate the abilities of MLLMs on AIGC videos. We evaluate 13 frontier MLLMs on VQ-Eval and find that even the best-performing model, GPT-4.1, struggles to achieve consistently good performance across all tasks. This highlights the challenging nature of our benchmark. Additionally, to investigate the practical applications of VQ-Eval in improving video generation, we design a re-prompt pipeline, demonstrating that aligning MLLMs more closely with human feedback can benefit the video generation. Tingyu Song, Tongyan Hu, Guo Gan, Yilun Zhao 0001 |
ACL (1) | 4 |
| 2025 | SciVer: Evaluating Foundation Models for Multimodal Scientific Claim VerificationabstractWe introduce SCIVER, the first benchmark specifically designed to evaluate the ability of foundation models to verify claims within a multimodal scientific context.SCIVER consists of 3,000 expert-annotated examples over 1,113 scientific papers, covering four subsets, each representing a common reasoning type in multimodal scientific claim verification.To enable fine-grained evaluation, each example includes expert-annotated supporting evidence.We assess the performance of 21 state-of-the-art multimodal foundation models, including o4mini, Gemini-2.5-Flash,Llama-3.2-Vision, and Qwen2.5-VL.Our experiment reveals a substantial performance gap between these models and human experts on SCIVER.Through an in-depth analysis of retrieval-augmented generation (RAG), and human-conducted error evaluations, we identify critical limitations in current open-source models, offering key insights to advance models' comprehension and reasoning in multimodal scientific literature tasks. Chengye Wang, Yifei Shen 0006, Zexi Kuang, Arman Cohan, Yilun Zhao 0001 |
ACL (1) | 5 |
| 2025 | Can LLMs Identify Critical Limitations within Scientific Research? A Systematic Evaluation on AI Research PapersabstractPeer review is fundamental to scientific research, but the growing volume of publications has intensified the challenges of this expertise-intensive process. While LLMs show promise in various scientific tasks, their potential to assist with peer review, particularly in identifying paper limitations, remains understudied. We first present a comprehensive taxonomy of limitation types in scientific research, with a focus on AI. Guided by this taxonomy, for studying limitations, we present LimitGen, the first comprehensive benchmark for evaluating LLMs’ capability to support early-stage feedback and complement human peer review. Our benchmark consists of two subsets: LimitGen-Syn, a synthetic dataset carefully created through controlled perturbations of high-quality papers, and LimitGen-Human, a collection of real human-written limitations. To improve the ability of LLM systems to identify limitations, we augment them with literature retrieval, which is essential for grounding identifying limitations in prior scientific findings. Our approach enhances the capabilities of LLM systems to generate limitations in research papers, enabling them to provide more concrete and constructive feedback. Yilun Zhao 0001, Manasi Patwardhan 0001, Lovekesh Vig, Arman Cohan |
ACL (1) | 2 |
| 2025 | MMVU: Measuring Expert-Level Multi-Discipline Video UnderstandingabstractWe introduce $\color{Blue}{\text{MMVU}}$, a comprehensive expert-level, multi-discipline benchmark for evaluating foundation models in video understanding. $\color{Blue}{\text{MMVU}}$ includes 3,000 expert-annotated questions spanning 27 subjects across four core disciplines: Science, Healthcare, Humanities & Social Sciences, and Engineering. Compared to prior benchmarks, $\color{Blue}{\text{MMVU}}$ features three key advancements. First, it challenges models to apply domain-specific knowledge and perform expert-level reasoning to analyze specialized-domain videos, moving beyond the basic visual perception typically assessed in current video benchmarks. Second, each example is annotated by human experts from scratch. We implement strict data quality controls to ensure the high quality of the dataset. Finally, each example is enriched with expert-annotated reasoning rationals and relevant domain knowledge, facilitating in-depth analysis. We conduct an extensive evaluation of 36 frontier multimodal foundation models on $\color{Blue}{\text{MMVU}}$. The latest System-2-capable models, o1 and Gemini 2.0 Flash Thinking, achieve the highest performance among the tested models. However, they still fall short of matching human expertise. Through in-depth error analyses and case studies, we offer actionable insights for future advancements in expert-level, knowledge-intensive video understanding for specialized domains. Yilun Zhao 0001, Haowei Zhang 0002, Lujing Xie, Tongyan Hu, Guo Gan, Yitao Long, Weiyuan Chen, Chuhan Li, Chengye Wang, Ziyao Shangguan, Zhenwen Liang, Yixin Liu 0003, Chen Zhao 0013, Arman Cohan |
CVPR | 1 |
| 2025 | SportReason: Evaluating Retrieval-Augmented Reasoning across Tables and Text for Sports Question AnsweringabstractWe present SPORTREASON, a benchmark for retrieval-augmented reasoning on numerical sports questions.Unlike existing benchmarks limited to one or two evidence units, SPORTREASON requires combining and reasoning across free-text, structured tables, and semi-structured infoboxes.We provide 3,000 human-verified QA pairs by repurposing existing QA and table generation datasets, and by prompting large language models (LLMs).Each pair is grounded in multiple evidence from a multi-modal Wikipedia corpus containing 200K knowledge contexts.We evaluate existing retrievers and rerankers, along with agentic Retrieval-Augmented Generation (RAG) systems.The experimental results show that multi-evidence retrieval remains a challenge.Agentic RAG systems (e.g., Search-o1), despite iterative retrieval and reasoning capabilities, fail to improve performance due to imprecise query generation and distracting retrieval information. Kaiyue Feng, Siyue Zhang, Bingsen Chen, Yilun Zhao 0001, Chen Zhao 0013 |
EMNLP | 4 |
| 2025 | FinTrust: A Comprehensive Benchmark of Trustworthiness Evaluation in Finance DomainabstractRecent LLMs have demonstrated promising ability in solving finance related problems.However, applying LLMs in real-world finance application remains challenging due to its high risk and high stakes property.This paper introduces FINTRUST, a comprehensive benchmark specifically designed for evaluating the trustworthiness of LLMs in finance applications.Our benchmark focuses on a wide range of alignment issues based on practical context and features fine-grained tasks for each dimension of trustworthiness evaluation.We assess eleven LLMs on FINTRUST and find that proprietary models like o4-mini outperforms in most tasks such as safety while open-source models like DeepSeek-V3 have advantage in specific areas like industry-level fairness.For challenging task like fiduciary alignment and disclosure, all LLMs fall short, showing a significant gap in legal awareness.We believe that FINTRUST can be a valuable benchmark for LLMs' trustworthiness evaluation in finance domain. Tiansheng Hu, Tongyan Hu, Liuyang Bai, Yilun Zhao 0001, Arman Cohan, Chen Zhao 0013 |
EMNLP | 4 |
| 2025 | LimRank: Less is More for Reasoning-Intensive Information RerankingabstractExisting approaches typically rely on largescale fine-tuning to adapt LLMs for information reranking tasks, which is computationally expensive.In this work, we demonstrate that modern LLMs can be effectively adapted using only minimal, high-quality supervision.To enable this, we design LIMRANK-SYNTHESIZER, a reusable and open-source pipeline for generating diverse, challenging, and realistic reranking examples.Using this synthetic data, we finetune our reranker model, LIMRANK.We evaluate LIMRANK on two challenging benchmarks, i.e., BRIGHT for reasoning-intensive retrieval and FOLLOWIR for instruction-following retrieval.Our experiments demonstrate that LIMRANK achieves competitive performance, while being trained on less than 5% of the data typically used in prior work.Further ablation studies demonstrate the effectiveness of LIMRANK-SYNTHESIZER and the strong generalization capabilities of LIMRANK across downstream tasks, including scientific literature search and retrieval-augmented generation for knowledge-intensive problem solving. yale-nlp/LimRank Tingyu Song, Yilun Zhao 0001, Siyue Zhang, Chen Zhao 0013, Arman Cohan |
EMNLP | 2 |
| 2025 | Table-R1: Inference-Time Scaling for Table Reasoning TasksabstractIn this work, we present the first study to explore inference-time scaling on table reasoning tasks.We develop and evaluate two post-training strategies to enable inferencetime scaling: distillation from frontier model reasoning traces and reinforcement learning with verifiable rewards (RLVR).For distillation, we introduce a large-scale dataset of reasoning traces generated by DeepSeek-R1, which we use to fine-tune LLMs into the Table-R1-SFT model.For RLVR, we propose taskspecific verifiable reward functions and apply the GRPO algorithm to obtain the Table-R1-Zero model.We evaluate our Table-R1series models across diverse table reasoning tasks, including short-form QA, fact verification, and free-form QA.Notably, the Table-R1-Zero model matches or exceeds the performance of GPT-4.1 and DeepSeek-R1, while using only a 7B-parameter LLM.It also demonstrates strong generalization to out-of-domain datasets.Extensive ablation and qualitative analyses reveal the benefits of instruction tuning, model architecture choices, and cross-task generalization, as well as emergence of essential table reasoning skills during RL training.Model huggingface.co/Table-R1Code github.com/Table-R1 Zheyuan Yang, Lyuhao Chen, Arman Cohan, Yilun Zhao 0001 |
EMNLP | 4 |
| 2025 | Diffusion vs. Autoregressive Language Models: A Text Embedding PerspectiveabstractLarge language model (LLM)-based embedding models, benefiting from large scale pretraining and post-training, have begun to surpass BERT and T5-based models on generalpurpose text embedding tasks such as document retrieval.However, a fundamental limitation of LLM embeddings lies in the unidirectional attention used during autoregressive pre-training, which misaligns with the bidirectional nature of text embedding tasks.To this end, we propose adopting diffusion language models for text embeddings, motivated by their inherent bidirectional architecture and recent success in matching or surpassing LLMs especially on reasoning tasks.We present the first systematic study of the diffusion language embedding model, which outperforms the LLM-based embedding model by 20% on long-document retrieval, 8% on reasoning-intensive retrieval, 2% on instruction-following retrieval, and achieve competitive performance on traditional text embedding benchmarks.Our analysis verifies that bidirectional attention is crucial for encoding global context in long and complex text. Siyue Zhang, Yilun Zhao 0001, Liyuan Geng, Arman Cohan, Anh Tuan Luu, Chen Zhao 0013 |
EMNLP | 2 |
| 2025 | TOMATO: Assessing Visual Temporal Reasoning Capabilities in Multimodal Foundation ModelsabstractExisting benchmarks often highlight the remarkable performance achieved by state-of-the-art Multimodal Foundation Models (MFMs) in leveraging temporal context for video understanding.
However, *how well do the models truly perform visual temporal reasoning?*
Our study of existing benchmarks shows that this capability of MFMs is likely overestimated as many questions can be solved by using a single, few, or out-of-order frames.
To systematically examine current visual temporal reasoning tasks, we propose three principles with corresponding metrics:
(1) *Multi-Frame Gain*,
(2) *Frame Order Sensitivity*,
and (3) *Frame Information Disparity*.
Following these principles, we introduce **TOMATO**, **T**emp**O**ral Reasoning **M**ultimod**A**l Evalua**T**i**O**n, a novel benchmark crafted to rigorously assess MFMs' temporal reasoning capabilities in video understanding.
TOMATO comprises 1,484 carefully curated, *human-annotated* questions spanning *six* tasks (i.e. *action count, direction, rotation, shape & trend, velocity & frequency, and visual cues*), applied to 1,417 videos, including 805 self-recorded and -generated videos, that encompass human-centric, real-world, and simulated scenarios.
Our comprehensive evaluation reveals a human-model performance gap of 57.3% with the best-performing model.
Moreover, our in-depth analysis uncovers more fundamental limitations beyond this gap in current MFMs. While they can accurately recognize events in isolated frames, they fail to interpret these frames as a continuous sequence.
We believe TOMATO will serve as a crucial testbed for evaluating the next-generation MFMs and as a call to the community to develop AI systems capable of comprehending the human world dynamics through the video modality. Ziyao Shangguan, Chuhan Li, Yanan Zheng, Yilun Zhao 0001, Tesca Fitzgerald, Arman Cohan |
ICLR | 5 |
| 2025 | ChemAgent: Self-updating Memories in Large Language Models Improves Chemical ReasoningabstractChemical reasoning usually involves complex, multi-step processes that demand precise calculations, where even minor errors can lead to cascading failures. Furthermore, large language models (LLMs) encounter difficulties handling domain-specific formulas, executing reasoning steps accurately, and integrating code ef- effectively when tackling chemical reasoning tasks. To address these challenges, we present ChemAgent, a novel framework designed to improve the performance of LLMs through a dynamic, self-updating library. This library is developed by decomposing chemical tasks into sub-tasks and compiling these sub-tasks into a structured collection that can be referenced for future queries. Then, when presented with a new problem, ChemAgent retrieves and refines pertinent information from the library, which we call memory, facilitating effective task decomposition and the generation of solutions. Our method designs three types of memory and a library-enhanced reasoning component, enabling LLMs to improve over time through experience. Experimental results on four chemical reasoning datasets from SciBench demonstrate that ChemAgent achieves performance gains of up to 46% (GPT-4), significantly outperforming existing methods. Our findings suggest substantial potential for future applications, including tasks such as drug discovery and materials science. Our code can be found at https://github.com/gersteinlab/ChemAgent. Xiangru Tang, Muyang Ye, Yanjun Shao, Xunjian Yin, Siru Ouyang, Wangchunshu Zhou, Pan Lu, Zhuosheng Zhang 0001, Yilun Zhao 0001, Arman Cohan, Mark Gerstein |
ICLR | 10 |
| 2025 | ReIFE: Re-evaluating Instruction-Following EvaluationabstractYixin Liu, Kejian Shi, Alexander Fabbri, Yilun Zhao, PeiFeng Wang, Chien-Sheng Wu, Shafiq Joty, Arman Cohan. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Yixin Liu 0003, Kejian Shi, Alexander R. Fabbri, Yilun Zhao 0001, Peifeng Wang, Chien-Sheng Wu, Shafiq R. Joty, Arman Cohan |
NAACL (Long Papers) | 4 |
| 2025 | IFIR: A Comprehensive Benchmark for Evaluating Instruction-Following in Expert-Domain Information RetrievalabstractTingyu Song, Guo Gan, Mingsheng Shang, Yilun Zhao. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Tingyu Song, Guo Gan, Mingsheng Shang 0001, Yilun Zhao 0001 |
NAACL (Long Papers) | 4 |
| 2025 | Are Multimodal LLMs Robust Against Adversarial Perturbations? RoMMath: A Systematic Evaluation on Multimodal Math ReasoningabstractYilun Zhao, Guo Gan, Chen Zhao, Arman Cohan. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Yilun Zhao 0001, Guo Gan, Chen Zhao 0013, Arman Cohan |
NAACL (Long Papers) | 1 |
| 2025 | Measuring what Matters: Construct Validity in Large Language Model BenchmarksabstractEvaluating large language models (LLMs) is crucial for both assessing their capabilities and identifying safety or robustness issues prior to deployment. Reliably measuring abstract and complex phenomena such as safety' androbustness' requires strong construct validity, that is, having measures that represent what matters to the phenomenon. With a team of 29 expert reviewers, we conduct a systematic review of 445 LLM benchmarks from leading conferences in natural language processing and machine learning. Across the reviewed articles, we find patterns related to the measured phenomena, tasks, and scoring metrics which undermine the validity of the resulting claims. To address these shortcomings, we provide eight key recommendations and detailed actionable guidance to researchers and practitioners in developing LLM benchmarks. Andrew M. Bean 0001, Ryan Othniel Kearns, Angelika Romanou, Franziska Sofia Hafner, Harry Mayne, Jan Batzner, Negar Foroutan Eghlidi, Chris Schmitz, Karolina Korgul, Hunar Batra, Oishi Deb, Emma Beharry, Cornelius Emde, Thomas Foster, Anna Gausen, María Grandury, Simeng Han, Valentin Hofmann, Lujain Ibrahim, Hazel Kim, Hannah Kirk, Fangru Lin, Gabrielle K. Liu, Lennart Luettgau, Jabez Magomere, Jonathan Rystrøm, Anna Sotnikova, Yilun Zhao 0001, Adel Bibi, Antoine Bosselut, Ronald Clark, Arman Cohan, Jakob N. Foerster, Yarin Gal, Scott A. Hale, Inioluwa Deborah Raji, Christopher Summerfield, Philip Torr 0001, Cozmin Ududec, Luc Rocher, Adam Mahdi |
NeurIPS | 29 |
| 2025 | SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded TasksabstractWe present SciArena, an open and collaborative platform for evaluating foundation models on scientific literature-grounded tasks. Unlike traditional benchmarks for scientific literature understanding and synthesis, SciArena engages the research community directly, following the Chatbot Arena evaluation approach of community voting on model comparisons.By leveraging collective intelligence, SciArena offers a community-driven evaluation of model performance on open-ended scientific tasks that demand literature-grounded, long-form responses.The platform currently supports 44 open-source and proprietary foundation models and has collected over 19,000 votes from human researchers across diverse scientific domains. Our analysis of the data collected so far confirms its high quality.We discuss the results and insights based on the model ranking leaderboard.To further promote research in building model-based automated evaluation systems for literature tasks, we release SciArena-Eval, a meta-evaluation benchmark based on our collected preference data. The benchmark measures the accuracy of models in judging answer quality by comparing their pairwise assessments with human votes. Our experiments highlight the benchmark’s challenges and emphasize the need for more reliable automated evaluation methods. Yilun Zhao 0001, Tiansheng Hu, Sihong Wu, Ronan Le Bras 0001, Yixin Liu 0003, Robert Tang, Joseph Chee Chang, Jesse Dodge, Jonathan Bragg, Chen Zhao 0013, Hannaneh Hajishirzi, Doug Downey, Arman Cohan |
NeurIPS | 1 |
| 2024 | TaPERA: Enhancing Faithfulness and Interpretability in Long-Form Table QA by Content Planning and Execution-based ReasoningabstractLong-form Table Question Answering (LFTQA) requires systems to generate paragraph long and complex answers to questions over tabular data.While Large language models based systems have made significant progress, it often hallucinates, especially when the task involves complex reasoning over tables.To tackle this issue, we propose a new LLM-based framework, TAPERA, for LFTQA tasks.Our framework uses a modular approach that decomposes the whole process into three sub-modules: 1) QA-based Content Planner that iteratively decomposes the input question into sub-questions; 2) Execution-based Table Reasoner that produces executable Python program for each sub-question; and 3) Answer Generator that generates long-form answer grounded on the program output.Human evaluation results on the FETAQA and QTSUMM datasets indicate that our framework significantly improves strong baselines on both accuracy and truthfulness, as our modular framework is better at table reasoning, and the long-form answer is always consistent with the program output.Our modular design further provides transparency as users are able to interact with our framework by manually changing the content plans.https://github.com/yilunzhao/TaPERAWithin the Oil and Gas industry, Sinopec Group earns the highest profit -$6,205 million.However, compared to the most profitable company overall, Apple, the profit earned by Sinopec Group is much lower.In fact, Apple earns $51,306 million more profit than Sinopec Group. Yilun Zhao 0001, Lyuhao Chen, Arman Cohan, Chen Zhao 0013 |
ACL (1) | 1 |
| 2024 | DocMath-Eval: Evaluating Math Reasoning Capabilities of LLMs in Understanding Financial DocumentsabstractYilun Zhao, Yitao Long, Hongjun Liu, Ryo Kamoi, Linyong Nan, Lyuhao Chen, Yixin Liu, Xiangru Tang, Rui Zhang, Arman Cohan. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yilun Zhao 0001, Yitao Long, Hongjun Liu 0001, Ryo Kamoi, Linyong Nan, Lyuhao Chen, Yixin Liu 0003, Xiangru Tang, Rui Zhang 0037, Arman Cohan |
ACL (1) | 1 |
| 2024 | KnowledgeFMath: A Knowledge-Intensive Math Reasoning Dataset in Finance DomainsabstractWe introduce FinanceMATH, a novel benchmark designed to evaluate LLMs' capabilities in solving knowledge-intensive math reasoning problems.Compared to prior works, this study features three core advancements.First, FinanceMATH includes 1,200 problems with a hybrid of textual and tabular content.These problems require college-level knowledge in the finance domain for effective resolution.Second, we provide expert-annotated, detailed solution references in Python program format, ensuring a high-quality benchmark for LLM assessment.We also construct a finance-domain knowledge bank and investigate various knowledge integration strategies.Finally, we evaluate a wide spectrum of 51 LLMs with both Chainof-Thought and Program-of-Thought prompting methods.Our experimental results reveal that the current best-performing system (i.e., GPT-4o) achieves only 60.9% accuracy using CoT prompting, leaving substantial room for improvement.Moreover, while augmenting LLMs with external knowledge can improve model performance (e.g., 47.5% → 54.5% for Gemini-1.5-Pro),their accuracy remains significantly lower than the estimated human expert performance of 92%.We believe that Fi-nanceMATH can advance future research in the area of domain-specific knowledge retrieval and integration, particularly within the context of solving reasoning-intensive tasks. * Equal ContributionQuestion: In 2018, Company A had a passive equity ownership interest of 15% in Company B. By the close of 2018, Company A decided to increase its ownership in Company B to 50%, effective as of 1st January 2019, through a cash purchase.There have been no financial transactions between Company A and Company B. Based on the data in the following table with Yilun Zhao 0001, Hongjun Liu 0001, Yitao Long, Rui Zhang 0037, Chen Zhao 0013, Arman Cohan |
ACL (1) | 1 |
| 2024 | FinDVer: Explainable Claim Verification over Long and Hybrid-content Financial DocumentsabstractYilun Zhao, Yitao Long, Tintin Jiang, Chengye Wang, Weiyuan Chen, Hongjun Liu, Xiangru Tang, Yiming Zhang, Chen Zhao, Arman Cohan. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Yilun Zhao 0001, Yitao Long, Tintin Jiang, Chengye Wang, Weiyuan Chen, Hongjun Liu 0001, Xiangru Tang, Chen Zhao 0013, Arman Cohan |
EMNLP | 1 |
| 2024 | FOLIO: Natural Language Reasoning with First-Order LogicabstractSimeng Han, Hailey Schoelkopf, Yilun Zhao, Zhenting Qi, Martin Riddell, Wenfei Zhou, James Coady, David Peng, Yujie Qiao, Luke Benson, Lucy Sun, Alexander Wardle-Solano, Hannah Szabó, Ekaterina Zubova, Matthew Burtell, Jonathan Fan, Yixin Liu, Brian Wong, Malcolm Sailor, Ansong Ni, Linyong Nan, Jungo Kasai, Tao Yu, Rui Zhang, Alexander Fabbri, Wojciech Maciej Kryscinski, Semih Yavuz, Ye Liu, Xi Victoria Lin, Shafiq Joty, Yingbo Zhou, Caiming Xiong, Rex Ying, Arman Cohan, Dragomir Radev. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Simeng Han, Hailey Schoelkopf, Yilun Zhao 0001, Zhenting Qi, Martin Riddell, Wenfei Zhou, James Coady, David Peng, Yujie Qiao, Luke Benson, Lucy Sun, Alexander Wardle-Solano, Hannah Szabó, Ekaterina Zubova, Matthew Burtell, Jonathan Fan 0001, Yixin Liu 0003, Malcolm Sailor, Ansong Ni, Linyong Nan, Jungo Kasai, Tao Yu 0009, Rui Zhang 0037, Alexander R. Fabbri, Wojciech Kryscinski, Semih Yavuz, Ye Liu 0006, Xi Victoria Lin, Shafiq R. Joty, Yingbo Zhou 0002, Caiming Xiong, Rex Ying, Arman Cohan, Dragomir R. Radev |
EMNLP | 3 |
| 2024 | Revisiting Automated Evaluation for Long-form Table Question AnsweringabstractIn the era of data-driven decision-making, Long-Form Table Question Answering (LFTQA) is essential for integrating structured data with complex reasoning.Despite recent advancements in Large Language Models (LLMs) for LFTQA, evaluating their effectiveness remains a significant challenge.We introduce LFTQA-Eval, a meta-evaluation dataset comprising 2,988 human-annotated examples, to rigorously assess the efficacy of current automated metrics in assessing LLM-based LFTQA systems, with a focus on faithfulness and comprehensiveness.Our findings reveal that existing automatic metrics poorly correlate with human judgments and fail to consistently differentiate between factually accurate responses and those that are coherent but factually incorrect.Additionally, our in-depth examination of the limitations associated with automated evaluation methods provides essential insights for the improvement of LFTQA automated evaluation. Lyuhao Chen, Songcheng Cai, Yilun Zhao 0001 |
EMNLP | 5 |
| 2024 | Investigating Data Contamination in Modern Benchmarks for Large Language ModelsabstractChunyuan Deng, Yilun Zhao, Xiangru Tang, Mark Gerstein, Arman Cohan. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Chunyuan Deng, Yilun Zhao 0001, Xiangru Tang, Mark Gerstein, Arman Cohan |
NAACL-HLT | 2 |
| 2024 | Uncertainty of Thoughts: Uncertainty-Aware Planning Enhances Information Seeking in LLMsabstractIn the face of uncertainty, the ability to *seek information* is of fundamental importance. In many practical applications, such as medical diagnosis and troubleshooting, the information needed to solve the task is not initially given, and has to be actively sought by asking follow-up questions (for example, a doctor asking a patient for more details about their symptoms). In this work, we introduce **Uncertainty of Thoughts (UoT)**, an algorithm to augment large language models with the ability to actively seek information by asking effective questions. UoT combines:
1. An *uncertainty-aware simulation approach* which enables the model to simulate possible future scenarios and how likely they are to occur,
2. *Uncertainty-based rewards* motivated by information gain which incentivizes the model to seek information, and
3. A *reward propagation scheme* to select the optimal question to ask in a way that maximizes the expected reward.
In experiments on medical diagnosis, troubleshooting and the `20 Questions' game, UoT achieves an average performance improvement of 38.1% in the rate of successful task completion across multiple LLMs compared with direct prompting, and also improves efficiency (i.e., the number of questions needed to complete the task). Chumin Liu, Xidong Feng, Yilun Zhao 0001, See-Kiong Ng, Anh Tuan Luu, Junxian He, Pang Wei W. Koh, Bryan Hooi |
NeurIPS | 4 |
| 2024 | L2CEval: Evaluating Language-to-Code Generation Capabilities of Large Language ModelsabstractAbstract Recently, large language models (LLMs), especially those that are pretrained on code, have demonstrated strong capabilities in generating programs from natural language inputs. Despite promising results, there is a notable lack of a comprehensive evaluation of these models’ language-to-code generation capabilities. Existing studies often focus on specific tasks, model architectures, or learning paradigms, leading to a fragmented understanding of the overall landscape. In this work, we present L2CEval, a systematic evaluation of the language-to-code generation capabilities of LLMs on 7 tasks across the domain spectrum of semantic parsing, math reasoning, and Python programming, analyzing the factors that potentially affect their performance, such as model size, pretraining data, instruction tuning, and different prompting methods. In addition, we assess confidence calibration, and conduct human evaluations to identify typical failures across different tasks and models. L2CEval offers a comprehensive understanding of the capabilities and limitations of LLMs in language-to-code generation. We release the evaluation framework1 and all model outputs, hoping to lay the groundwork for further future research. All future evaluations (e.g., LLaMA-3, StarCoder2, etc) will be updated on the project website: https://l2c-eval.github.io/. Ansong Ni, Yilun Zhao 0001, Martin Riddell, Troy Feng, Stephen Yin, Ye Liu 0006, Semih Yavuz, Caiming Xiong, Shafiq R. Joty, Yingbo Zhou 0002, Dragomir R. Radev, Arman Cohan |
Trans. Assoc. Comput. Linguistics | 3 |
| 2023 | Revisiting the Gold Standard: Grounding Summarization Evaluation with Robust Human EvaluationabstractYixin Liu, Alex Fabbri, Pengfei Liu, Yilun Zhao, Linyong Nan, Ruilin Han, Simeng Han, Shafiq Joty, Chien-Sheng Wu, Caiming Xiong, Dragomir Radev. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yixin Liu 0003, Alexander R. Fabbri, Pengfei Liu 0003, Yilun Zhao 0001, Linyong Nan, Ruilin Han, Simeng Han, Shafiq R. Joty, Chien-Sheng Wu, Caiming Xiong, Dragomir R. Radev |
ACL (1) | 4 |
| 2023 | RobuT: A Systematic Study of Table QA Robustness Against Human-Annotated Adversarial PerturbationsabstractYilun Zhao, Chen Zhao, Linyong Nan, Zhenting Qi, Wenlin Zhang, Xiangru Tang, Boyu Mi, Dragomir Radev. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Yilun Zhao 0001, Chen Zhao 0013, Linyong Nan, Zhenting Qi, Xiangru Tang, Boyu Mi, Dragomir R. Radev |
ACL (1) | 1 |
| 2023 | LoFT: Enhancing Faithfulness and Diversity for Table-to-Text Generation via Logic Form ControlabstractLogical Table -to-Text (LT2T) generation is tasked with generating logically faithful sentences from tables.There currently exists two challenges in the field: 1) Faithfulness: how to generate sentences that are factually correct given the table content; 2) Diversity: how to generate multiple sentences that offer different perspectives on the table.This work proposes LOFT, which utilizes logic forms as fact verifiers and content planners to control LT2T generation.Experimental results on the LOGICNLG dataset demonstrate that LOFT is the first model that addresses unfaithfulness and lack of diversity issues simultaneously.Our code is publicly available at https: //github.com/Yale-LILY/LoFT. Yilun Zhao 0001, Zhenting Qi, Linyong Nan, Lorenzo Jaime Yu Flores, Dragomir R. Radev |
EACL | 1 |
| 2023 | QTSumm: Query-Focused Summarization over Tabular DataabstractYilun Zhao, Zhenting Qi, Linyong Nan, Boyu Mi, Yixin Liu, Weijin Zou, Simeng Han, Ruizhe Chen, Xiangru Tang, Yumo Xu, Dragomir Radev, Arman Cohan. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Yilun Zhao 0001, Zhenting Qi, Linyong Nan, Boyu Mi, Yixin Liu 0003, Weijin Zou, Simeng Han, Ruizhe Chen, Xiangru Tang, Yumo Xu, Dragomir R. Radev, Arman Cohan |
EMNLP | 1 |
| 2023 | Towards Interpretable and Efficient Automatic Reference-Based Summarization EvaluationabstractYixin Liu, Alexander Fabbri, Yilun Zhao, Pengfei Liu, Shafiq Joty, Chien-Sheng Wu, Caiming Xiong, Dragomir Radev. Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. 2023. Yixin Liu 0003, Alexander R. Fabbri, Yilun Zhao 0001, Pengfei Liu 0003, Shafiq R. Joty, Chien-Sheng Wu, Caiming Xiong, Dragomir R. Radev |
EMNLP | 3 |
| 2022 | MultiHiertt: Numerical Reasoning over Multi Hierarchical Tabular and Textual DataabstractNumerical reasoning over hybrid data containing both textual and tabular content (e.g., financial reports) has recently attracted much attention in the NLP community.However, existing question answering (QA) benchmarks over hybrid data only include a single flat table in each document and thus lack examples of multistep numerical reasoning across multiple hierarchical tables.To facilitate data analytical progress, we construct a new large-scale benchmark, MULTIHIERTT, with QA pairs over Multi Hierarchical Tabular and Textual data.MULTIHIERTT is built from a wealth of financial reports and has the following unique characteristics: 1) each document contain multiple tables and longer unstructured texts; 2) most of tables contained are hierarchical; 3) the reasoning process required for each question is more complex and challenging than existing benchmarks; and 4) fine-grained annotations of reasoning processes and supporting facts are provided to reveal complex numerical reasoning.We further introduce a novel QA model termed MT2Net, which first applies facts retrieving to extract relevant supporting facts from both tables and text and then uses a reasoning module to perform symbolic reasoning over retrieved facts.We conduct comprehensive experiments on various baselines.The experimental results show that MULTIHIERTT presents a strong challenge for existing baselines whose results lag far behind the performance of human experts.The dataset and code are publicly available at https://github. com/psunlpgroup/MultiHiertt.How much is the sum of stock purchase rights in 2018 lower than those in 2017?How many years were the sales and client service expenses higher than software development expenses?How much of US corporate debt securities is there in total (in 2009) without consider gross unrealized gain and gross unrelized loss?How many financing activities continues to increase every year from 2017 to 2021? Yilun Zhao 0001, Chenying Li, Rui Zhang 0037 |
ACL (1) | 1 |
| 2022 | R2D2: Robust Data-to-Text with Replacement DetectionabstractUnfaithful text generation is a common problem for text generation systems.In the case of Data-to-Text (D2T) systems, the factuality of the generated text is particularly crucial for any real-world applications.We introduce R2D2, a training framework that addresses unfaithful Data-to-Text generation by training a system both as a generator and a faithfulness discriminator with additional replacement detection and unlikelihood learning tasks.To facilitate such training, we propose two methods for sampling unfaithful sentences.We argue that the poor entity retrieval capability of D2T systems is one of the primary sources of unfaithfulness, so in addition to the existing metrics, we further propose named entity based metrics to evaluate the fidelity of D2T generations.Our experimental results show that R2D2 systems could effectively mitigate the unfaithful text generation, and they achieve new state-of-the-art results on FeTaQA, LogicNLG, and ToTTo, all with significant improvements. Linyong Nan, Lorenzo Jaime Yu Flores, Yilun Zhao 0001, Yixin Liu 0003, Luke Benson, Weijin Zou, Dragomir R. Radev |
EMNLP | 3 |
| 2022 | ReasTAP: Injecting Table Reasoning Skills During Pre-training via Synthetic Reasoning ExamplesabstractReasoning over tabular data requires both table structure understanding and a broad set of table reasoning skills.Current models with tablespecific architectures and pre-training methods perform well on understanding table structures, but they still struggle with tasks that require various table reasoning skills.In this work, we develop REASTAP to show that high-level table reasoning skills can be injected into models during pre-training without a complex tablespecific architecture design.We define 7 table reasoning skills, such as numerical operation, temporal comparison, and conjunction.Each reasoning skill is associated with one example generator, which synthesizes questions over semi-structured tables according to the sampled templates.We model the table pre-training task as a sequence generation task and pretrain REASTAP to generate precise answers to the synthetic examples.REASTAP is evaluated on four benchmarks covering three downstream tasks including: 1) WIKISQL-WEAK and WIKITQ for Table Question Answering; 2) TABFACT for Table Fact Verification; and 3) LOGICNLG for Faithful Table-to-Text Generation.Experimental results demonstrate that REASTAP achieves new state-of-the-art performance on all benchmarks and delivers a significant improvement on low-resource setting. Yilun Zhao 0001, Linyong Nan, Zhenting Qi, Rui Zhang 0037, Dragomir R. Radev |
EMNLP | 1 |
| 2022 | FinMath: Injecting a Tree-structured Solver for Question Answering over Financial ReportsabstractAnswering questions over financial reports containing both tabular and textual data (hybrid data) is challenging as it requires models to select information from financial reports and perform complex quantitative analyses. Although current models have demonstrated a solid capability to solve simple questions, they struggle with complex questions that require a multiple-step numerical reasoning process. This paper proposes a new framework named FinMath, which improves the model’s numerical reasoning capacity by injecting a tree-structured neural model to perform multi-step numerical reasoning. Specifically, FinMath extracts supporting evidence from the financial reports given the question in the first phase. In the second phase, a tree-structured neural model is applied to generate a tree expression in a top-down recursive way. Experiments on the TAT-QA dataset show that our proposed approach improves the previous best result by 8.5% absolute for Exact Match (EM) score (50.1% to 58.6%) and 6.1% absolute for numeracy-focused F1 score (58.0% to 64.1%). Chenying Li, Wenbo Ye, Yilun Zhao 0001 |
LREC | 3 |
| 2022 | Apparel-Invariant Feature Learning for Person Re-IdentificationabstractWith the rise of deep learning methods, person Re-Identification (ReID) performance has been improved tremendously in many public datasets. However, most public ReID datasets are collected in a short time window in which persons’ appearance rarely changes. In real-world applications such as in a shopping mall, the same person may change their wearings, and different persons may wear similar apparel. It reveals a critical problem that current ReID models heavily rely on a person’s apparel, resulting in an inconsistent ReID performance. Therefore, it is crucial to learn an apparel-invariant person representation under clothes changing or several persons wearing similar clothes cases. In this work, we tackle this problem from the viewpoint of invariant feature representation learning. The main contributions of this work are as follows. (1) We propose the semi-supervised Apparel-invariant Feature Learning (AIFL) framework to learn an apparel-invariant pedestrian representation using images of the same person wearing different clothes. (2) To obtain images of the same person wearing different clothes, we propose an unsupervised apparel-simulation GAN (AS-GAN) to synthesize cloth-changing images according to the target cloth embedding. It is worth noting that the images used in ReID tasks were cropped from real-world low-quality CCTV videos, making it more challenging to synthesize cloth-changing images. Extensive experiments demonstrate that our proposal can improve the ReID performance of the baseline models. Zhengxu Yu, Yilun Zhao 0001, Zhongming Jin 0001, Jianqiang Huang 0001, Deng Cai 0001, Xian-Sheng Hua 0001 |
IEEE Trans. Multim. | 2 |
| 2021 | MusiCoder: A Universal Music-Acoustic Encoder Based on Transformer
Yilun Zhao 0001 |
MMM (1) | 1 |