Chen Zhao 0013

dblp:81/3-13 · DBLP profile ↗
← Back
25ranked-venue papers
4as first author
21since 2021 · last 2026
0000-0001-7672-146XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 22 · 3 first-author · 20 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 2 since 2021Computer networks · 1Databases, data management, data science and information retrieval · 1 · 1 first-authorApplied, interdisciplinary, general and emerging computing · 1 · 1 first-author
YearPublicationVenuePosition
2026 Beyond Single-shot Writing: Deep Research Agents are Unreliable at Multi-turn Report Revision
abstract
Existing benchmarks for Deep Research Agents (DRAs) treat report generation as a single-shot writing task, which fundamentally diverges from how human researchers iteratively draft and revise reports via self-reflection or peer feedback.Whether DRAs can reliably revise reports with user feedback remains unexplored.We introduce MR DRE, an evaluation suite that establishes multi-turn report revision as a new evaluation axis for DRAs.MR DRE consists of (1) a unified long-form report evaluation protocol spanning comprehensiveness, factuality, and presentation, and (2) a human-verified feedback simulation pipeline for multi-turn revision.Our analysis of five diverse DRAs reveals a critical limitation: while agents can address most user feedback, they regress on 16-27% of previously covered content and citation quality.Over multiple revision turns, even the best-performing agent leaves significant headroom, as they continue to disrupt content outside the feedback's scope and fail to preserve earlier edits.We also show that these issues are not easily resolvable through inference-time fixes such as prompt engineering and a dedicated sub-agent for revision 1 . Comprehensiveness
Bingsen Chen, Ping Nie, Yuyu Zhang, Xi Ye 0003, Chen Zhao 0013
ACL (1)6
2026 A Survey of Multimodal Mathematical Reasoning: From Perception, Alignment to Reasoning
abstract
Tianyu Yang, Sihong Wu, Yilun Zhao, Zhenwen Liang, Lisen Dai, Chen Zhao, Minhao Cheng, Arman Cohan, Xiangliang Zhang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Sihong Wu, Yilun Zhao 0001, Zhenwen Liang, Lisen Dai, Chen Zhao 0013, Minhao Cheng, Arman Cohan, Xiangliang Zhang 0001
ACL (1)6
2026 Rethinking Reasoning-Intensive Retrieval: Evaluating and Advancing Retrievers in Agentic Search Systems
abstract
Reasoning-intensive retrieval aims to surface evidence that supports downstream reasoning rather than merely matching topical similarity.This capability is increasingly important for agentic search systems, where retrievers must provide complementary evidence across iterative search and synthesis.However, existing work remains limited on both evaluation and training: benchmarks such as BRIGHT provide narrow gold sets and evaluate retrievers in isolation, while synthetic training corpora often optimize single-passage relevance rather than evidence portfolio construction.We introduce BRIGHT-PRO, an expert-annotated benchmark that expands each query with multi-aspect gold evidence and evaluates retrievers under both static and agentic search protocols.We further construct RTriever-Synth, an aspect-decomposed synthetic corpus that generates complementary positives and positive-conditioned hard negatives, and use it to LoRA fine-tune RTriever-4B from Qwen3-Embedding-4B.Experiments across lexical, general-purpose, and reasoningintensive retrievers show that aspect-aware and agentic evaluation expose behaviors hidden by standard metrics, while RTriever-4B substantially improves over its base model.
Yilun Zhao 0001, Jinbiao Wei, Tingyu Song, Siyue Zhang, Chen Zhao 0013, Arman Cohan
ACL (1)5
2025 MMVU: Measuring Expert-Level Multi-Discipline Video Understanding
abstract
We introduce $\color{Blue}{\text{MMVU}}$, a comprehensive expert-level, multi-discipline benchmark for evaluating foundation models in video understanding. $\color{Blue}{\text{MMVU}}$ includes 3,000 expert-annotated questions spanning 27 subjects across four core disciplines: Science, Healthcare, Humanities & Social Sciences, and Engineering. Compared to prior benchmarks, $\color{Blue}{\text{MMVU}}$ features three key advancements. First, it challenges models to apply domain-specific knowledge and perform expert-level reasoning to analyze specialized-domain videos, moving beyond the basic visual perception typically assessed in current video benchmarks. Second, each example is annotated by human experts from scratch. We implement strict data quality controls to ensure the high quality of the dataset. Finally, each example is enriched with expert-annotated reasoning rationals and relevant domain knowledge, facilitating in-depth analysis. We conduct an extensive evaluation of 36 frontier multimodal foundation models on $\color{Blue}{\text{MMVU}}$. The latest System-2-capable models, o1 and Gemini 2.0 Flash Thinking, achieve the highest performance among the tested models. However, they still fall short of matching human expertise. Through in-depth error analyses and case studies, we offer actionable insights for future advancements in expert-level, knowledge-intensive video understanding for specialized domains.
Yilun Zhao 0001, Haowei Zhang 0002, Lujing Xie, Tongyan Hu, Guo Gan, Yitao Long, Weiyuan Chen, Chuhan Li, Chengye Wang, Ziyao Shangguan, Zhenwen Liang, Yixin Liu 0003, Chen Zhao 0013, Arman Cohan
CVPR15
2025 SportReason: Evaluating Retrieval-Augmented Reasoning across Tables and Text for Sports Question Answering
abstract
We present SPORTREASON, a benchmark for retrieval-augmented reasoning on numerical sports questions.Unlike existing benchmarks limited to one or two evidence units, SPORTREASON requires combining and reasoning across free-text, structured tables, and semi-structured infoboxes.We provide 3,000 human-verified QA pairs by repurposing existing QA and table generation datasets, and by prompting large language models (LLMs).Each pair is grounded in multiple evidence from a multi-modal Wikipedia corpus containing 200K knowledge contexts.We evaluate existing retrievers and rerankers, along with agentic Retrieval-Augmented Generation (RAG) systems.The experimental results show that multi-evidence retrieval remains a challenge.Agentic RAG systems (e.g., Search-o1), despite iterative retrieval and reasoning capabilities, fail to improve performance due to imprecise query generation and distracting retrieval information.
Kaiyue Feng, Siyue Zhang, Bingsen Chen, Yilun Zhao 0001, Chen Zhao 0013
EMNLP5
2025 FinTrust: A Comprehensive Benchmark of Trustworthiness Evaluation in Finance Domain
abstract
Recent LLMs have demonstrated promising ability in solving finance related problems.However, applying LLMs in real-world finance application remains challenging due to its high risk and high stakes property.This paper introduces FINTRUST, a comprehensive benchmark specifically designed for evaluating the trustworthiness of LLMs in finance applications.Our benchmark focuses on a wide range of alignment issues based on practical context and features fine-grained tasks for each dimension of trustworthiness evaluation.We assess eleven LLMs on FINTRUST and find that proprietary models like o4-mini outperforms in most tasks such as safety while open-source models like DeepSeek-V3 have advantage in specific areas like industry-level fairness.For challenging task like fiduciary alignment and disclosure, all LLMs fall short, showing a significant gap in legal awareness.We believe that FINTRUST can be a valuable benchmark for LLMs' trustworthiness evaluation in finance domain.
Tiansheng Hu, Tongyan Hu, Liuyang Bai, Yilun Zhao 0001, Arman Cohan, Chen Zhao 0013
EMNLP6
2025 LimRank: Less is More for Reasoning-Intensive Information Reranking
abstract
Existing approaches typically rely on largescale fine-tuning to adapt LLMs for information reranking tasks, which is computationally expensive.In this work, we demonstrate that modern LLMs can be effectively adapted using only minimal, high-quality supervision.To enable this, we design LIMRANK-SYNTHESIZER, a reusable and open-source pipeline for generating diverse, challenging, and realistic reranking examples.Using this synthetic data, we finetune our reranker model, LIMRANK.We evaluate LIMRANK on two challenging benchmarks, i.e., BRIGHT for reasoning-intensive retrieval and FOLLOWIR for instruction-following retrieval.Our experiments demonstrate that LIMRANK achieves competitive performance, while being trained on less than 5% of the data typically used in prior work.Further ablation studies demonstrate the effectiveness of LIMRANK-SYNTHESIZER and the strong generalization capabilities of LIMRANK across downstream tasks, including scientific literature search and retrieval-augmented generation for knowledge-intensive problem solving. yale-nlp/LimRank
Tingyu Song, Yilun Zhao 0001, Siyue Zhang, Chen Zhao 0013, Arman Cohan
EMNLP4
2025 Diffusion vs. Autoregressive Language Models: A Text Embedding Perspective
abstract
Large language model (LLM)-based embedding models, benefiting from large scale pretraining and post-training, have begun to surpass BERT and T5-based models on generalpurpose text embedding tasks such as document retrieval.However, a fundamental limitation of LLM embeddings lies in the unidirectional attention used during autoregressive pre-training, which misaligns with the bidirectional nature of text embedding tasks.To this end, we propose adopting diffusion language models for text embeddings, motivated by their inherent bidirectional architecture and recent success in matching or surpassing LLMs especially on reasoning tasks.We present the first systematic study of the diffusion language embedding model, which outperforms the LLM-based embedding model by 20% on long-document retrieval, 8% on reasoning-intensive retrieval, 2% on instruction-following retrieval, and achieve competitive performance on traditional text embedding benchmarks.Our analysis verifies that bidirectional attention is crucial for encoding global context in long and complex text.
Siyue Zhang, Yilun Zhao 0001, Liyuan Geng, Arman Cohan, Anh Tuan Luu, Chen Zhao 0013
EMNLP6
2025 Are Multimodal LLMs Robust Against Adversarial Perturbations? RoMMath: A Systematic Evaluation on Multimodal Math Reasoning
abstract
Yilun Zhao, Guo Gan, Chen Zhao, Arman Cohan. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Yilun Zhao 0001, Guo Gan, Chen Zhao 0013, Arman Cohan
NAACL (Long Papers)3
2025 SciArena: An Open Evaluation Platform for Non-Verifiable Scientific Literature-Grounded Tasks
abstract
We present SciArena, an open and collaborative platform for evaluating foundation models on scientific literature-grounded tasks. Unlike traditional benchmarks for scientific literature understanding and synthesis, SciArena engages the research community directly, following the Chatbot Arena evaluation approach of community voting on model comparisons.By leveraging collective intelligence, SciArena offers a community-driven evaluation of model performance on open-ended scientific tasks that demand literature-grounded, long-form responses.The platform currently supports 44 open-source and proprietary foundation models and has collected over 19,000 votes from human researchers across diverse scientific domains. Our analysis of the data collected so far confirms its high quality.We discuss the results and insights based on the model ranking leaderboard.To further promote research in building model-based automated evaluation systems for literature tasks, we release SciArena-Eval, a meta-evaluation benchmark based on our collected preference data. The benchmark measures the accuracy of models in judging answer quality by comparing their pairwise assessments with human votes. Our experiments highlight the benchmark’s challenges and emphasize the need for more reliable automated evaluation methods.
Yilun Zhao 0001, Tiansheng Hu, Sihong Wu, Ronan Le Bras 0001, Yixin Liu 0003, Robert Tang, Joseph Chee Chang, Jesse Dodge, Jonathan Bragg, Chen Zhao 0013, Hannaneh Hajishirzi, Doug Downey, Arman Cohan
NeurIPS11
2024 TaPERA: Enhancing Faithfulness and Interpretability in Long-Form Table QA by Content Planning and Execution-based Reasoning
abstract
Long-form Table Question Answering (LFTQA) requires systems to generate paragraph long and complex answers to questions over tabular data.While Large language models based systems have made significant progress, it often hallucinates, especially when the task involves complex reasoning over tables.To tackle this issue, we propose a new LLM-based framework, TAPERA, for LFTQA tasks.Our framework uses a modular approach that decomposes the whole process into three sub-modules: 1) QA-based Content Planner that iteratively decomposes the input question into sub-questions; 2) Execution-based Table Reasoner that produces executable Python program for each sub-question; and 3) Answer Generator that generates long-form answer grounded on the program output.Human evaluation results on the FETAQA and QTSUMM datasets indicate that our framework significantly improves strong baselines on both accuracy and truthfulness, as our modular framework is better at table reasoning, and the long-form answer is always consistent with the program output.Our modular design further provides transparency as users are able to interact with our framework by manually changing the content plans.https://github.com/yilunzhao/TaPERAWithin the Oil and Gas industry, Sinopec Group earns the highest profit -$6,205 million.However, compared to the most profitable company overall, Apple, the profit earned by Sinopec Group is much lower.In fact, Apple earns $51,306 million more profit than Sinopec Group.
Yilun Zhao 0001, Lyuhao Chen, Arman Cohan, Chen Zhao 0013
ACL (1)4
2024 KnowledgeFMath: A Knowledge-Intensive Math Reasoning Dataset in Finance Domains
abstract
We introduce FinanceMATH, a novel benchmark designed to evaluate LLMs' capabilities in solving knowledge-intensive math reasoning problems.Compared to prior works, this study features three core advancements.First, FinanceMATH includes 1,200 problems with a hybrid of textual and tabular content.These problems require college-level knowledge in the finance domain for effective resolution.Second, we provide expert-annotated, detailed solution references in Python program format, ensuring a high-quality benchmark for LLM assessment.We also construct a finance-domain knowledge bank and investigate various knowledge integration strategies.Finally, we evaluate a wide spectrum of 51 LLMs with both Chainof-Thought and Program-of-Thought prompting methods.Our experimental results reveal that the current best-performing system (i.e., GPT-4o) achieves only 60.9% accuracy using CoT prompting, leaving substantial room for improvement.Moreover, while augmenting LLMs with external knowledge can improve model performance (e.g., 47.5% → 54.5% for Gemini-1.5-Pro),their accuracy remains significantly lower than the estimated human expert performance of 92%.We believe that Fi-nanceMATH can advance future research in the area of domain-specific knowledge retrieval and integration, particularly within the context of solving reasoning-intensive tasks. * Equal ContributionQuestion: In 2018, Company A had a passive equity ownership interest of 15% in Company B. By the close of 2018, Company A decided to increase its ownership in Company B to 50%, effective as of 1st January 2019, through a cash purchase.There have been no financial transactions between Company A and Company B. Based on the data in the following table with
Yilun Zhao 0001, Hongjun Liu 0001, Yitao Long, Rui Zhang 0037, Chen Zhao 0013, Arman Cohan
ACL (1)5
2024 Parallel Structures in Pre-training Data Yield In-Context Learning
abstract
Pre-trained language models (LMs) are capable of in-context learning (ICL): they can adapt to a task with only a few examples given in the prompt without any parameter update.However, it is unclear where this capability comes from as there is a stark distribution shift between pre-training text and ICL prompts.In this work, we study what patterns of the pretraining data contribute to ICL.We find that LMs' ICL ability depends on parallel structures in the pre-training data-pairs of phrases following similar templates in the same context window.Specifically, we detect parallel structures by checking whether training on one phrase improves prediction of the other, and conduct ablation experiments to study their effect on ICL.We show that removing parallel structures in the pre-training data reduces LMs' ICL accuracy by 51% (vs 2% from random ablation).This drop persists even when excluding common patterns such as n-gram repetitions and long-range dependency, showing the diversity and generality of parallel structures.A closer look at the detected parallel structures indicates that they cover diverse linguistic tasks and span long distances in the data.
Yanda Chen, Chen Zhao 0013, Kathy McKeown, He He 0001
ACL (1)2
2024 FinDVer: Explainable Claim Verification over Long and Hybrid-content Financial Documents
abstract
Yilun Zhao, Yitao Long, Tintin Jiang, Chengye Wang, Weiyuan Chen, Hongjun Liu, Xiangru Tang, Yiming Zhang, Chen Zhao, Arman Cohan. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024.
Yilun Zhao 0001, Yitao Long, Tintin Jiang, Chengye Wang, Weiyuan Chen, Hongjun Liu 0001, Xiangru Tang, Chen Zhao 0013, Arman Cohan
EMNLP9
2024 Do Models Explain Themselves? Counterfactual Simulatability of Natural Language Explanations
abstract
Large language models (LLMs) are trained to imitate humans to explain human decisions. However, do LLMs explain themselves? Can they help humans build mental models of how LLMs process different inputs? To answer these questions, we propose to evaluate $\textbf{counterfactual simulatability}$ of natural language explanations: whether an explanation can enable humans to precisely infer the model’s outputs on diverse counterfactuals of the explained input. For example, if a model answers ”$\textit{yes}$” to the input question ”$\textit{Can eagles fly?}$” with the explanation ”$\textit{all birds can fly}$”, then humans would infer from the explanation that it would also answer ”$\textit{yes}$” to the counterfactual input ”$\textit{Can penguins fly?}$”. If the explanation is precise, then the model’s answer should match humans’ expectations. We implemented two metrics based on counterfactual simulatability: precision and generality. We generated diverse counterfactuals automatically using LLMs. We then used these metrics to evaluate state-of-the-art LLMs (e.g., GPT-4) on two tasks: multi-hop factual reasoning and reward modeling. We found that LLM’s explanations have low precision and that precision does not correlate with plausibility. Therefore, naively optimizing human approvals (e.g., RLHF) may be insufficient.
Yanda Chen, Ruiqi Zhong, Narutatsu Ri, Chen Zhao 0013, He He 0001, Jacob Steinhardt, Kathy McKeown
ICML4
2024 Large Language Models Help Humans Verify Truthfulness - Except When They Are Convincingly Wrong
abstract
Chenglei Si, Navita Goyal, Tongshuang Wu, Chen Zhao, Shi Feng, Hal Daumé Iii, Jordan Boyd-Graber. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024.
Chenglei Si, Navita Goyal, Sherry Tongshuang Wu, Chen Zhao 0013, Shi Feng 0005, Hal Daumé III, Jordan L. Boyd-Graber
NAACL-HLT4
2024 : Visualization of AI-Assisted Task Guidance in AR
abstract
The concept of augmented reality (AR) assistants has captured the human imagination for decades, becoming a staple of modern science fiction. To pursue this goal, it is necessary to develop artificial intelligence (AI)-based methods that simultaneously perceive the 3D environment, reason about physical tasks, and model the performer, all in real-time. Within this framework, a wide variety of sensors are needed to generate data across different modalities, such as audio, video, depth, speech, and time-of-flight. The required sensors are typically part of the AR headset, providing performer sensing and interaction through visual, audio, and haptic feedback. AI assistants not only record the performer as they perform activities, but also require machine learning (ML) models to understand and assist the performer as they interact with the physical world. Therefore, developing such assistants is a challenging task. We propose ARGUS, a visual analytics system to support the development of intelligent AR assistants. Our system was designed as part of a multi-year-long collaboration between visualization researchers and ML and AR experts. This co-design process has led to advances in the visualization of ML in AR. Our system allows for online visualization of object, action, and step detection as well as offline analysis of previously recorded AR sessions. It visualizes not only the multimodal sensor data streams but also the output of the ML models. This allows developers to gain insights into the performer activities as well as the ML models, helping them troubleshoot, improve, and fine-tune the components of the AR assistant.
Sonia Castelo Quispe, João Rulff, Erin McGowan, Bea Steers, Guande Wu, Shaoyu Chen, Irán R. Román, Roque Lopez, Ethan Brewer, Chen Zhao 0013, Kyunghyun Cho, He He 0001, Qi Sun 0003, Huy T. Vo, Juan Pablo Bello, Michael Krone, Cláudio T. Silva
IEEE Trans. Vis. Comput. Graph.10
2023 RobuT: A Systematic Study of Table QA Robustness Against Human-Annotated Adversarial Perturbations
abstract
Yilun Zhao, Chen Zhao, Linyong Nan, Zhenting Qi, Wenlin Zhang, Xiangru Tang, Boyu Mi, Dragomir Radev. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Yilun Zhao 0001, Chen Zhao 0013, Linyong Nan, Zhenting Qi, Xiangru Tang, Boyu Mi, Dragomir R. Radev
ACL (1)2
2021 What's in a Name? Answer Equivalence For Open-Domain Question Answering
abstract
A flaw in QA evaluation is that annotations often only provide one gold answer.Thus, model predictions semantically equivalent to the answer but superficially different are considered incorrect.This work explores mining alias entities from knowledge bases and using them as additional gold answers (i.e., equivalent answers).We incorporate answers for two settings: evaluation with additional answers and model training with equivalent answers.We analyse three QA benchmarks: Natural Questions, TriviaQA and SQuAD.Answer expansion increases the exact match score on all datasets for evaluation, while incorporating it helps model training over real-world datasets.We ensure the additional answers are valid through a human post hoc evaluation. 1
Chenglei Si, Chen Zhao 0013, Jordan L. Boyd-Graber
EMNLP (1)2
2021 Distantly-Supervised Dense Retrieval Enables Open-Domain Question Answering without Evidence Annotation
abstract
Open-domain question answering answers a question based on evidence retrieved from a large corpus.State-of-the-art neural approaches require intermediate evidence annotations for training.However, such intermediate annotations are expensive, and methods that rely on them cannot transfer to the more common setting, where only questionanswer pairs are available.This paper investigates whether models can learn to find evidence from a large corpus, with only distant supervision from answer labels for model training, thereby generating no additional annotation cost.We introduce a novel approach (DISTDR) that iteratively improves over a weak retriever by alternately finding evidence from the up-to-date model and encouraging the model to learn the most likely evidence.Without using any evidence labels, DISTDR is on par with fully-supervised state-of-theart methods on both multi-hop and singlehop QA benchmarks.Our analysis confirms that DISTDR finds more accurate evidence over iterations, which leads to model improvements.The code is available at https:// github.com/henryzhao5852/DistDR.
Chen Zhao 0013, Chenyan Xiong, Jordan L. Boyd-Graber, Hal Daumé III
EMNLP (1)1
2021 Multi-Step Reasoning Over Unstructured Text with Beam Dense Retrieval
abstract
Chen Zhao, Chenyan Xiong, Jordan Boyd-Graber, Hal Daumé III. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Chen Zhao 0013, Chenyan Xiong, Jordan L. Boyd-Graber, Hal Daumé III
NAACL-HLT1
2020 Transformer-XH: Multi-Evidence Reasoning with eXtra Hop Attention
Chen Zhao 0013, Chenyan Xiong, Corby Rosset, Paul N. Bennett, Saurabh Tiwary
ICLR1
2020 Complex Factoid Question Answering with a Free-Text Knowledge Graph
abstract
We introduce delft, a factoid question answering system which combines the nuance and depth of knowledge graph question answering approaches with the broader coverage of free-text. delft builds a free-text knowledge graph from Wikipedia, with entities as nodes and sentences in which entities co-occur as edges. For each question, delft finds the subgraph linking question entity nodes to candidates using text sentences as edges, creating a dense and high coverage semantic graph. A novel graph neural network reasons over the free-text graph—combining evidence on the nodes via information along edge sentences—to select a final answer. Experiments on three question answering datasets show delft can answer entity-rich questions better than machine reading based models, bert-based answer ranking and memory networks. delft’s advantage comes from both the high coverage of its free-text knowledge graph—more than double that of dbpedia relations—and the novel graph neural network which reasons on the rich but noisy free-text evidence.
Chen Zhao 0013, Chenyan Xiong, Jordan L. Boyd-Graber
WWW1
2018 A dataset and baselines for sequential open-domain question answering
abstract
Previous work on question-answering systems mainly focuses on answering individual questions, assuming they are independent and devoid of context.Instead, we investigate sequential question answering, asking multiple related questions.We present QBLink, a new dataset of fully human-authored questions.We extend existing strong question answering frameworks to include previous questions to improve the overall question-answering accuracy in open-domain question answering.The dataset is publicly available at http:// sequential.qanta.org.
Ahmed Elgohary, Chen Zhao 0013, Jordan L. Boyd-Graber
EMNLP2
2016 Atom Decomposition with Adaptive Basis Selection Strategy for Matrix Completion
abstract
Estimating missing entries in matrices has attracted much attention due to its wide range of applications like image inpainting and video denoising, which are usually considered as low-rank matrix completion problems theoretically. It is common to consider nuclear norm as a surrogate of the rank operator since it is the tightest convex lower bound of the rank operator under certain conditions. However, most approaches based on nuclear norm minimization involve a number of singular value decomposition (SVD) operations. Given a matrix X ∈ R m × n , the time complexity of the SVD operation is O ( mn 2 ), which brings prohibitive computational burden on large-scale matrices, limiting the further usage of these methods in real applications. Motivated by this observation, a series of atom-decomposition-based matrix completion methods have been studied. The key to these methods is to reconstruct the target matrix by pursuit methods in a greedy way, which only involves the computation of the top SVD and has great advantages in efficiency compared with the SVD-based matrix completion methods. However, due to gradually serious accumulation errors, atom-decomposition-based methods usually result in unsatisfactory reconstruction accuracy. In this article, we propose a new efficient and scalable atom decomposition algorithm for matrix completion called Adaptive Basis Selection Strategy ( ABSS ). Different from traditional greedy atom decomposition methods, a two-phase strategy is conducted to generate the basis separately via different strategies according to their different nature. At first, we globally prune the basis space to eliminate the unimportant basis as much as possible and locate the probable subspace containing the most informative basis. Then, another group of basis spaces are learned to improve the recovery accuracy based on local information. In this way, our proposed algorithm breaks through the accuracy bottleneck of traditional atom-decomposition-based matrix completion methods; meanwhile, it reserves the innate efficiency advantages over SVD-based matrix completion methods. We empirically evaluate the proposed algorithm ABSS on real visual image data and large-scale recommendation datasets. Results have shown that ABSS has much better reconstruction accuracy with comparable cost to atom-decomposition-based methods. At the same time, it outperforms the state-of-the-art SVD-based matrix completion algorithms by similar or better reconstruction accuracy with enormous advantages on efficiency.
Yao Hu 0002, Chen Zhao 0013, Deng Cai 0001, Xiaofei He 0001, Xuelong Li 0001
ACM Trans. Multim. Comput. Commun. Appl.2