EDBT 2026 Demo / reviewers in the wild / expert
Sercan Ö. Arik
dblp:78/10697 · also Sercan Ömer Arik
· DBLP profile ↗
44ranked-venue papers
8as first author
28since 2021 · last 2025
0000-0001-6333-1729ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 41 · 6 first-author · 28 since 2021Graphics, computer vision, multimedia, augmented reality and games · 9 · 4 first-author · 3 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Astute RAG: Overcoming Imperfect Retrieval Augmentation and Knowledge Conflicts for Large Language ModelsabstractRetrieval-augmented generation (RAG), while effective in integrating external knowledge to enhance large language models (LLMs), can be undermined by imperfect retrieval, which may introduce irrelevant, misleading, or even malicious information.Despite its importance, previous studies have rarely explored the behavior of RAG with errors from imperfect retrieval, and how potential conflicts arise between the LLMs' internal knowledge and external sources.We show that imperfect retrieval augmentation might be inevitable and quite harmful, through controlled analysis under realistic conditions.Knowledge conflicts between LLM-internal and external knowledge from retrieval is a bottleneck to overcome in the post-retrieval stage of RAG.To render LLMs resilient to imperfect retrieval, we propose ASTUTE RAG, a novel RAG approach that adaptively elicits essential information from LLMs' internal knowledge, iteratively consolidates internal and external knowledge with source-awareness, and finalizes the answer according to information reliability.Our experiments with Gemini and Claude demonstrate that ASTUTE RAG significantly outperforms previous robustness-enhanced RAG methods.Notably, ASTUTE RAG is the only approach that matches or exceeds the performance of LLMs without RAG under worst-case scenarios.ASTUTE RAG effectively resolves knowledge conflicts, improving the reliability and trustworthiness of RAG systems.45.4% LLM correct RAG correct LLM incorrect RAG incorrect Both sides are wrong.It is hard to improve, but combining internal and external knowledge may help.Previous work leverages RAG to address LLMs' knowledge gap.Zonia receives from Reuben a letter in the play.Zonia receives from Reuben a kiss in the play Xingchen Wan, Ruoxi Sun 0002, Jiefeng Chen 0001, Sercan Ö. Arik |
ACL (1) | 5 |
| 2025 | Embedding-Converter: A Unified Framework for Cross-Model Embedding TransformationabstractEmbedding models play a crucial role in machine learning.However, the continuous development of new models presents a major challenge: migrating to a potentially superior model often requires the computationally expensive process of re-embedding entire datasets-without any guarantee of performance improvement.This paper presents Embedding-Converter, a novel framework for efficiently transforming embeddings between different models, thus avoiding costly 'reembedding'.The proposed approach achieves 100 times faster and cheaper computations in real-world applications.Experiments show that Embedding-Converter not only streamlines transitions to new models, but can also improve upon the source model's performance, approaching that of the target model.This facilitates efficient evaluation and broader adoption of new embedding models by significantly reducing the overhead of model switching.Furthermore, Embedding-Converter addresses latency limitations by enabling the use of smaller models for online tasks while still benefiting from the performance of larger models offline.By promoting the release of converters alongside new embedding models, Embedding-Converter fosters a more dynamic and accessible ecosystem for embedding model development and deployment. Jinsung Yoon, Sercan Ö. Arik |
ACL (1) | 2 |
| 2025 | Learning to Clarify: Multi-turn Conversations with Action-Based Contrastive Self-TrainingabstractLarge language models (LLMs), optimized through human feedback, have rapidly emerged as a leading paradigm for developing intelligent conversational assistants. However, despite their strong performance across many benchmarks, LLM-based agents might still lack conversational skills such as disambiguation -- when they are faced with ambiguity, they often overhedge or implicitly guess users' true intents rather than asking clarification questions. Under task-specific settings, high-quality conversation samples are often limited, constituting a bottleneck for LLMs' ability to learn optimal dialogue action policies. We propose Action-Based Contrastive Self-Training (ACT), a quasi-online preference optimization algorithm based on Direct Preference Optimization (DPO), that enables data-efficient dialogue policy learning in multi-turn conversation modeling. We demonstrate ACT's efficacy under data-efficient tuning scenarios, even when there is no action label available, using multiple real-world conversational tasks: tabular-grounded question-answering, machine reading comprehension, and AmbigSQL, a novel task for disambiguating information-seeking requests for complex SQL generation towards data analysis agents. Additionally, we propose evaluating LLMs' ability to function as conversational agents by examining whether they can implicitly recognize and reason about ambiguity in conversation. ACT demonstrates substantial conversation modeling improvements over standard tuning approaches like supervised fine-tuning and DPO. Maximillian Chen 0001, Ruoxi Sun 0002, Tomas Pfister, Sercan Ö. Arik |
ICLR | 4 |
| 2025 | Long-Context LLMs Meet RAG: Overcoming Challenges for Long Inputs in RAGabstractRetrieval-augmented generation (RAG) empowers large language models (LLMs) to utilize external knowledge sources. The increasing capacity of LLMs to process longer input sequences opens up avenues for providing more retrieved information, to potentially enhance the quality of generated outputs. From a long-context LLM perspective, it assumes that a larger retrieval set would contain more relevant information (higher recall), that might result in improved performance. However, our empirical findings demonstrate that for many long-context LLMs, the quality of generated output initially improves first, but then subsequently declines as the number of retrieved passages increases. This paper investigates this phenomenon, identifying the detrimental impact of retrieved "hard negatives" as a key contributor. To mitigate this and enhance the robustness of long-context LLM-based RAG, we propose both training-free and training-based approaches. We first showcase the effectiveness of retrieval reordering as a simple yet powerful training-free optimization. Furthermore, we explore training-based methods, specifically RAG-specific implicit LLM fine-tuning and RAG-oriented fine-tuning with intermediate reasoning, demonstrating their capacity for substantial performance gains. Finally, we conduct a systematic analysis of design choices for these training-based methods, including data distribution, retriever selection, and training context length. Bowen Jin, Jinsung Yoon, Jiawei Han 0001, Sercan Ö. Arik |
ICLR | 4 |
| 2025 | CHASE-SQL: Multi-Path Reasoning and Preference Optimized Candidate Selection in Text-to-SQLabstractWe present CHASE-SQL, a novel framework addressing large language model (LLM) performance challenges for Text-to-SQL tasks by leveraging multi-agent modeling and test-time compute for improved candidate generation and selection. CHASE-SQL uses LLMs to generate diverse SQL candidates with: (1) a divide-and-conquer approach to break down complex queries, (2) chain-of-thought reasoning based on query execution plans, and (3) instance-aware synthetic example generation for tailored few-shot demonstrations. A selection agent ranks candidates via pairwise comparisons using a fine-tuned binary selection LLM, offering robust performance. This framework improves SQL query quality and diversity, achieving state-of-the-art execution accuracy of 73.0% on the BIRD Text-to-SQL benchmark test set, topping the leaderboard at the time of submission. Mohammadreza Pourreza, Ruoxi Sun 0002, Yeounoh Chung, Shayan Talaei, Gaurav Tarlok Kakkar, Amin Saberi, Fatma Özcan 0001, Sercan Ö. Arik |
ICLR | 10 |
| 2025 | Mitigating Object Hallucination in MLLMs via Data-augmented Phrase-level AlignmentabstractDespite their significant advancements, Multimodal Large Language Models
(MLLMs) often generate factually inaccurate information, referred to as hallucination.
In this work, we address object hallucinations in MLLMs, where information
is generated about an object not present in the input image. We introduce Data-augmented
Phrase-level Alignment (DPA), a novel loss which can be applied to
instruction-tuned off-the-shelf MLLMs to mitigate hallucinations, while preserving
their general vision-language capabilities. To fine-tune MLLMs with DPA, we first
generate a set of 'hallucinated' and 'correct' response pairs through generative data
augmentation by selectively altering the ground-truth information of the correct
responses at a phrase level. The DPA loss is then used to train MLLMs to reduce
the likelihood of hallucinated phrases compared to the correct ones. Our thorough
evaluation on various benchmarks confirms the effectiveness of DPA in mitigating
hallucination while retaining the out-of-the-box performance of the MLLMs on
general tasks. For instance, MLLMs finetuned with DPA, which we refer to as Hallucination
Attenuated Language and Vision Assistant (HALVA), improve F1 by up
to 13.4% on hallucination visual question-answering and reduce the hallucination
rate by up to 4.2% on image description tasks. Pritam Sarkar, Sayna Ebrahimi, Ali Etemad, Ahmad Beirami, Sercan Ö. Arik, Tomas Pfister |
ICLR | 5 |
| 2025 | Learn-by-interact: A Data-Centric Framework For Self-Adaptive Agents in Realistic EnvironmentsabstractAutonomous agents powered by large language models (LLMs) have the potential to enhance human capabilities, assisting with digital tasks from sending emails to performing data analysis. The abilities of existing LLMs at such tasks are often hindered by the lack of high-quality agent data from the corresponding environments they interact with. We propose LEARN-BY-INTERACT, a data-centric framework to adapt LLM agents to any given environments without human annotations. LEARN-BY-INTERACT synthesizes trajectories of agent-environment interactions based on documentations, and constructs instructions by summarizing or abstracting the interaction histories, a process called backward construction. We assess the quality of our synthetic data by using them in both training-based scenarios and training-free in-context learning (ICL), where we craft innovative retrieval approaches optimized for agents. Extensive experiments on SWE-bench, WebArena, OSWorld, and Spider2-V spanning across realistic coding, web, and desktop environments show the effectiveness of LEARN-BY-INTERACT in various downstream agentic tasks — baseline results are improved up to 11.1% for ICL with Claude-3.5 and 23.1% for training with Codestral-22B. We further demonstrate the critical role of backward construction, which provides up to 10.6% improvement for training. Our ablation studies demonstrate the efficiency provided by our synthesized data in ICL and the superiority of our retrieval pipeline over alternative approaches like conventional retrieval-augmented generation (RAG). We expect that LEARN-BY-INTERACT will serve as a foundation for agent data synthesis as LLMs are increasingly deployed at real-world environments. Hongjin Su, Ruoxi Sun 0002, Jinsung Yoon, Tao Yu 0009, Sercan Ö. Arik |
ICLR | 6 |
| 2025 | BRIGHT: A Realistic and Challenging Benchmark for Reasoning-Intensive RetrievalabstractExisting retrieval benchmarks primarily consist of information-seeking queries (e.g., aggregated questions from search engines) where keyword or semantic-based retrieval is usually sufficient. However, many complex real-world queries require in-depth reasoning to identify relevant documents that go beyond surface form matching. For example, finding documentation for a coding question requires understanding the logic and syntax of the functions involved. To better benchmark retrieval on such challenging queries, we introduce BRIGHT, the first text retrieval benchmark that requires intensive reasoning to retrieve relevant documents. Our dataset consists of 1,398 real-world queries spanning diverse domains such as economics, psychology, mathematics, coding, and more. These queries are drawn from naturally occurring or carefully curated human data. Extensive evaluation reveals that even state-of-the-art retrieval models perform poorly on BRIGHT. The leading model on the MTEB leaderboard (Muennighoff et al., 2023), which achieves a score of 59.0 nDCG@10,1 produces a score of nDCG@10 of 18.0 on BRIGHT. We show that incorporating explicit reasoning about the query improves retrieval performance by up to 12.2 points. Moreover, incorporating retrieved documents from the top-performing retriever boosts question answering performance by over 6.6 points. We believe that BRIGHT paves the way for future research on retrieval systems in more realistic and challenging settings. Hongjin Su, Howard Yen, Mengzhou Xia, Niklas Muennighoff, Han-yu Wang, Haisu Liu, Zachary S. Siegel, Michael Tang, Ruoxi Sun 0002, Jinsung Yoon, Sercan Ö. Arik, Danqi Chen 0001, Tao Yu 0009 |
ICLR | 13 |
| 2025 | From Few to Many: Self-Improving Many-Shot Reasoners Through Iterative Optimization and GenerationabstractRecent advances in long-context large language models (LLMs) have led to the emerging paradigm of many-shot in-context learning (ICL), where it is observed that scaling many more demonstrating examples beyond the conventional few-shot setup in the context can lead to performance benefits. However, despite its promise, it is unclear what aspects dominate the benefits and whether simply scaling to more examples is the most effective way of improving many-shot ICL. In this work, we first provide an analysis on the factors driving many-shot ICL, and we find that 1) many-shot performance can still be attributed to often a few disproportionately influential examples and 2) identifying such influential examples ("optimize") and using them as demonstrations to regenerate new examples ("generate") can lead to further improvements. Inspired by the findings, we propose BRIDGE, an algorithm that alternates between the optimize step with Bayesian optimization to discover the influential sets of examples and the generate step to reuse this set to expand the reasoning paths of the examples back to the many-shot regime automatically. On Gemini, Claude, and Mistral LLMs of different sizes, we show BRIDGE led to significant improvements across a diverse set of tasks including symbolic reasoning, numerical reasoning and code generation. Xingchen Wan, Han Zhou 0010, Ruoxi Sun 0002, Sercan Ö. Arik |
ICLR | 4 |
| 2025 | Retrieval Augmented Time Series ForecastingabstractTime series forecasting uses historical data to predict future trends, leveraging the relationships between past observations and available features. In this paper, we propose RAFT, a retrieval-augmented time series forecasting method to provide sufficient inductive biases and complement the model’s learning capacity. When forecasting the subsequent time frames, we directly retrieve historical data candidates from the training dataset with patterns most similar to the input, and utilize the future values of these candidates alongside the inputs to obtain predictions. This simple approach augments the model’s capacity by externally providing information about past patterns via retrieval modules. Our empirical evaluations on ten benchmark datasets show that RAFT consistently outperforms contemporary baselines with an average win ratio of 86%. Sungwon Han 0001, SeungEon Lee 0001, Meeyoung Cha, Sercan Ö. Arik, Jinsung Yoon |
ICML | 4 |
| 2025 | LLM Alignment as Retriever Optimization: An Information Retrieval PerspectiveabstractLarge Language Models (LLMs) have revolutionized artificial intelligence with capabilities in reasoning, coding, and communication, driving innovation across industries. Their true potential depends on effective alignment to ensure correct, trustworthy and ethical behavior, addressing challenges like misinformation, hallucinations, bias and misuse. While existing Reinforcement Learning (RL)-based alignment methods are notoriously complex, direct optimization approaches offer a simpler alternative. In this work, we introduce a novel direct optimization approach for LLM alignment by drawing on established Information Retrieval (IR) principles. We present a systematic framework that bridges LLM alignment and IR methodologies, mapping LLM generation and reward models to IR’s retriever-reranker paradigm. Building on this foundation, we propose LLM Alignment as Retriever Preference Optimization (LarPO), a new alignment method that enhances overall alignment quality. Extensive experiments validate LarPO’s effectiveness with 38.9 % and 13.7 % averaged improvement on AlpacaEval2 and MixEval-Hard respectively. Our work opens new avenues for advancing LLM alignment by integrating IR foundations, offering a promising direction for future research. Bowen Jin, Jinsung Yoon, Zhen Qin 0001, Ziqi Wang 0003, Wei Xiong 0015, Yu Meng 0001, Jiawei Han 0001, Sercan Ö. Arik |
ICML | 8 |
| 2025 | MLE-STAR: Machine Learning Engineering Agent via Search and Targeted RefinementabstractAgents based on large language models (LLMs) for machine learning engineering (MLE) can automatically implement ML models via code generation. However, existing approaches to build such agents often rely heavily on inherent LLM knowledge and employ coarse exploration strategies that modify the entire code structure at once. This limits their ability to select effective task-specific models and perform deep exploration within specific components, such as experimenting extensively with feature engineering options. To overcome these, we propose MLE-STAR, a novel approach to build MLE agents. MLE-STAR first leverages external knowledge by using a search engine to retrieve effective models from the web, forming an initial solution, then iteratively refines it by exploring various strategies targeting specific ML components. This exploration is guided by ablation studies analyzing the impact of individual code blocks. Furthermore, we introduce a novel ensembling method using an effective strategy suggested by MLE-STAR. Our experimental results show that MLE-STAR achieves medals in 64% of the Kaggle competitions on the MLE-bench, significantly outperforming the best alternative. Jaehyun Nam, Jinsung Yoon, Jiefeng Chen 0001, Jinwoo Shin, Sercan Ö. Arik, Tomas Pfister |
NeurIPS | 5 |
| 2024 | Search-Adaptor: Embedding Customization for Information RetrievalabstractEmbeddings extracted by pre-trained Large Language Models (LLMs) have significant potential to improve information retrieval and search.Beyond the zero-shot setup in which they are being conventionally used, being able to take advantage of the information from the relevant query-corpus paired data can further boost the LLM capabilities.In this paper, we propose a novel method, Search-Adaptor, for customizing LLMs for information retrieval in an efficient and robust way.Search-Adaptor modifies the embeddings generated by pre-trained LLMs, and can be integrated with any LLM, including those only available via prediction APIs.On multiple English, multilingual, and multimodal retrieval datasets, we show consistent and significant performance benefits for Search-Adaptor -e.g., more than 5% improvements for Google Embedding APIs in nDCG@10 averaged over 14 BEIR datasets. Jinsung Yoon, Yanfei Chen, Sercan Ö. Arik, Tomas Pfister |
ACL (1) | 3 |
| 2024 | Matryoshka-Adaptor: Unsupervised and Supervised Tuning for Smaller Embedding DimensionsabstractEmbeddings from Large Language Models (LLMs) have emerged as critical components in various applications, particularly for information retrieval.While high-dimensional embeddings generally demonstrate superior performance as they contain more salient information, their practical application is frequently hindered by elevated computational latency and the associated higher cost.To address these challenges, we propose Matryoshka-Adaptor, a novel tuning framework designed for the customization of LLM embeddings.Matryoshka-Adaptor facilitates substantial dimensionality reduction while maintaining comparable performance levels, thereby achieving a significant enhancement in computational efficiency and costeffectiveness.Our framework directly modifies the embeddings from pre-trained LLMs which is designed to be seamlessly integrated with any LLM architecture, encompassing those accessible exclusively through blackbox APIs.Also, it exhibits efficacy in both unsupervised and supervised learning settings.A rigorous evaluation conducted across a diverse corpus of English, multilingual, and multimodal datasets consistently reveals substantial gains with Matryoshka-Adaptor.Notably, with Google and OpenAI Embedding APIs, Matryoshka-Adaptor achieves a reduction in dimensionality ranging from twoto twelve-fold without compromising performance across multiple BEIR datasets. Jinsung Yoon, Rajarishi Sinha, Sercan Ö. Arik, Tomas Pfister |
EMNLP | 3 |
| 2024 | TEMPO: Prompt-based Generative Pre-trained Transformer for Time Series ForecastingabstractThe past decade has witnessed significant advances in time series modeling with deep learning. While achieving state-of-the-art results, the best-performing architectures vary highly across applications and domains. Meanwhile, for natural language processing, the Generative Pre-trained Transformer (GPT) has demonstrated impressive performance via training one general-purpose model across various textual datasets. It is intriguing to explore whether GPT-type architectures can be effective for time series, capturing the intrinsic dynamic attributes and leading to significant accuracy improvements. In this paper, we propose a novel framework, TEMPO, that can effectively learn time series representations. We focus on utilizing two essential inductive biases of the time series task for pre-trained models: (i) decomposition of the complex interaction between trend, seasonal and residual components; and (ii) introducing the design of prompts to facilitate distribution adaptation in different types of time series. TEMPO expands the capability for dynamically modeling real-world temporal phenomena from data within diverse domains. Our experiments demonstrate the superior performance of TEMPO over state-of-the-art methods on zero shot setting for a number of time series benchmark datasets. This performance gain is observed not only in scenarios involving previously unseen datasets but also in scenarios with multi-modal inputs. This compelling finding highlights TEMPO's potential to constitute a foundational model-building framework. Defu Cao, Furong Jia 0002, Sercan Ö. Arik, Tomas Pfister, Yixiang Zheng, Wen Ye 0001, Yan Liu 0002 |
ICLR | 3 |
| 2024 | Large Language Models Can Automatically Engineer Features for Few-Shot Tabular LearningabstractLarge Language Models (LLMs), with their remarkable ability to tackle challenging and unseen reasoning problems, hold immense potential for tabular learning, that is vital for many real-world applications. In this paper, we propose a novel in-context learning framework, FeatLLM, which employs LLMs as feature engineers to produce an input data set that is optimally suited for tabular predictions. The generated features are used to infer class likelihood with a simple downstream machine learning model, such as linear regression and yields high performance few-shot learning. The proposed FeatLLM framework only uses this simple predictive model with the discovered features at inference time. Compared to existing LLM-based approaches, FeatLLM eliminates the need to send queries to the LLM for each sample at inference time. Moreover, it merely requires API-level access to LLMs, and overcomes prompt size limitations. As demonstrated across numerous tabular datasets from a wide range of domains, FeatLLM generates high-quality rules, significantly (10% on average) outperforming alternatives such as TabLLM and STUNT. Sungwon Han 0001, Jinsung Yoon, Sercan Ö. Arik, Tomas Pfister |
ICML | 3 |
| 2024 | Effective Large Language Model Adaptation for Improved Grounding and Citation GenerationabstractXi Ye, Ruoxi Sun, Sercan Arik, Tomas Pfister. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Xi Ye 0003, Ruoxi Sun 0002, Sercan Ö. Arik, Tomas Pfister |
NAACL-HLT | 3 |
| 2024 | Teach Better or Show Smarter? On Instructions and Exemplars in Automatic Prompt OptimizationabstractLarge language models have demonstrated remarkable capabilities but their performance is heavily reliant on effective prompt engineering. Automatic prompt optimization (APO) methods are designed to automate this and can be broadly categorized into those targeting instructions (instruction optimization, IO) vs. those targeting exemplars (exemplar optimization, EO). Despite their shared objective, these have evolved rather independently, with IO receiving more research attention recently. This paper seeks to bridge this gap by comprehensively comparing the performance of representative IO and EO techniques both isolation and combination on a diverse set of challenging tasks. Our findings reveal that intelligently reusing model-generated input-output pairs obtained from evaluating prompts on the validation set as exemplars, consistently improves performance on top of IO methods but is currently under-investigated. We also find that despite the recent focus on IO, how we select exemplars can outweigh how we optimize instructions, with EO strategies as simple as random search outperforming state-of-the-art IO methods with seed instructions without any optimization. Moreover, we observe a synergy between EO and IO, with optimal combinations surpassing the individual contributions. We conclude that studying exemplar optimization both as a standalone method and its optimal combination with instruction optimization remain a crucial aspect of APO and deserve greater consideration in future research, even in the era of highly capable instruction-following models. Xingchen Wan, Ruoxi Sun 0002, Hootan Nakhost, Sercan Ö. Arik |
NeurIPS | 4 |
| 2024 | Chain of Agents: Large Language Models Collaborating on Long-Context TasksabstractAddressing the challenge of effectively processing long contexts has become a critical issue for Large Language Models (LLMs). Two common strategies have emerged: 1) reducing the input length, such as retrieving relevant chunks by Retrieval-Augmented Generation (RAG), and 2) expanding the context window limit of LLMs. However, both strategies have drawbacks: input reduction has no guarantee of covering the part with needed information, while window extension struggles with focusing on the pertinent information for solving the task. To mitigate these limitations, we propose Chain-of-Agents (CoA), a novel framework that harnesses multi-agent collaboration through natural language to enable information aggregation and context reasoning across various LLMs over long-context tasks. CoA consists of multiple worker agents who sequentially communicate to handle different segmented portions of the text, followed by a manager agent who synthesizes these contributions into a coherent final output. CoA processes the entire input by interleaving reading and reasoning, and it mitigates long context focus issues by assigning each agent a short context. We perform a comprehensive evaluation of CoA on a wide range of long-context tasks in question answering, summarization, and code completion, demonstrating significant improvements by up to 10% over strong baselines of RAG, Full-Context, and multi-agent LLMs. Yusen Zhang 0001, Ruoxi Sun 0002, Yanfei Chen, Tomas Pfister, Rui Zhang 0037, Sercan Ö. Arik |
NeurIPS | 6 |
| 2023 | Neural Spline Search for Quantile Probabilistic ModelingabstractAccurate estimation of output quantiles is crucial in many use cases, where it is desired to model the range of possibility. Modeling target distribution at arbitrary quantile levels and at arbitrary input attribute levels are important to offer a comprehensive picture of the data, and requires the quantile function to be expressive enough. The quantile function describing the target distribution using quantile levels is critical for quantile regression. Although various parametric forms for the distributions (that the quantile function specifies) can be adopted, an everlasting problem is selecting the most appropriate one that can properly approximate the data distributions. In this paper, we propose a non-parametric and data-driven approach, Neural Spline Search (NSS), to represent the observed data distribution without parametric assumptions. NSS is flexible and expressive for modeling data distributions by transforming the inputs with a series of monotonic spline regressions guided by symbolic operators. We demonstrate that NSS outperforms previous methods on synthetic, real-world regression and time-series forecasting tasks. Ruoxi Sun 0002, Chun-Liang Li, Sercan Ö. Arik, Michael Dusenberry, Chen-Yu Lee, Tomas Pfister |
AAAI | 3 |
| 2023 | Universal Self-Adaptive PromptingabstractA hallmark of modern large language models (LLMs) is their impressive general zero-shot and few-shot abilities, often elicited through in-context learning (ICL) via prompting.However, while highly coveted and being the most general, zero-shot performances in LLMs are still typically weaker due to the lack of guidance and the difficulty of applying existing automatic prompt design methods in general tasks when ground-truth labels are unavailable.In this study, we address this by presenting Universal Self-Adaptive Prompting (USP), an automatic prompt design approach specifically tailored for zero-shot learning (while compatible with few-shot).Requiring only a small amount of unlabeled data and an inferenceonly LLM, USP is highly versatile: to achieve universal prompting, USP categorizes a possible NLP task into one of the three possible task types and then uses a corresponding selector to select the most suitable queries and zero-shot model-generated responses as pseudo-demonstrations, thereby generalizing ICL to the zero-shot setup in a fully automated way.We evaluate USP with PaLM and PaLM 2 models and demonstrate performances that are considerably stronger than standard zero-shot baselines and often comparable to or even superior to few-shot baselines across more than 40 natural language understanding, natural language generation, and reasoning tasks. Xingchen Wan, Ruoxi Sun 0002, Hootan Nakhost, Hanjun Dai, Julian Martin Eisenschlos, Sercan Ö. Arik, Tomas Pfister |
EMNLP | 6 |
| 2023 | Koopman Neural Operator Forecaster for Time-series with Temporal Distributional Shifts
Rui Wang 0086, Yihe Dong, Sercan Ö. Arik, Rose Yu |
ICLR | 3 |
| 2023 | Data-Efficient and Interpretable Tabular Anomaly DetectionabstractAnomaly detection (AD) plays an important role in numerous applications. In this paper, we focus on two understudied aspects of AD that are critical for integration into real-world applications. First, most AD methods cannot incorporate labeled data that are often available in practice in small quantities and can be crucial to achieve high accuracy. Second, most AD methods are not interpretable, a bottleneck that prevents stakeholders from understanding the reason behind the anomalies. In this paper, we propose a novel AD framework, DIAD, that adapts a white-box model class, Generalized Additive Models, to detect anomalies using a partial identification objective which naturally handles noisy or heterogeneous features. DIAD can incorporate a small amount of labeled data to further boost AD performances in semi-supervised settings. We demonstrate the superiority of DIAD compared to previous work in both unsupervised and semi-supervised settings on multiple datasets. We also present explainability capabilities of DIAD, on its rationale behind predicting certain samples as anomalies. Chun-Hao Chang, Jinsung Yoon, Sercan Ö. Arik, Madeleine Udell, Tomas Pfister |
KDD | 3 |
| 2022 | Nested Hierarchical Transformer: Towards Accurate, Data-Efficient and Interpretable Visual UnderstandingabstractHierarchical structures are popular in recent vision transformers, however, they require sophisticated designs and massive datasets to work well. In this paper, we explore the idea of nesting basic local transformers on non-overlapping image blocks and aggregating them in a hierarchical way. We find that the block aggregation function plays a critical role in enabling cross-block non-local information communication. This observation leads us to design a simplified architecture that requires minor code changes upon the original vision transformer. The benefits of the proposed judiciously-selected design are threefold: (1) NesT converges faster and requires much less training data to achieve good generalization on both ImageNet and small datasets like CIFAR; (2) when extending our key ideas to image generation, NesT leads to a strong decoder that is 8 times faster than previous transformer-based generators; and (3) we show that decoupling the feature learning and abstraction processes via this nested hierarchy in our design enables constructing a novel method (named GradCAT) for visually interpreting the learned model. Source code is available https://github.com/google-research/nested-transformer. Han Zhang 0010, Long Zhao 0003, Ting Chen 0001, Sercan Ö. Arik, Tomas Pfister |
AAAI | 5 |
| 2022 | Decoupling Local and Global Representations of Time SeriesabstractReal-world time series data are often generated from several sources of variation. Learning representations that capture the factors contributing to this variability enables better understanding of the data via its underlying generative process and can lead to improvements in performance on downstream machine learning tasks. In this paper, we propose a novel generative approach for learning representations for the global and local factors of variation in time series data. The local representation of each sample models non-stationarity over time with a stochastic process prior, and the global representation of the sample encodes the time-independent characteristics. To encourage decoupling between the representations, we introduce a counterfactual regularization that minimizes the mutual information between the two variables. In experiments, we demonstrate successful recovery of the true local and global factors of variability on simulated data, and show that representations learned using our method lead to superior performance on downstream tasks on real-world datasets. We believe that the proposed way of defining representations is beneficial for data modelling and can yield better insights into the complexity of the real-world data. Sana Tonekaboni, Chun-Liang Li, Sercan Ö. Arik, Anna Goldenberg, Tomas Pfister |
AISTATS | 3 |
| 2022 | Self-Supervised Learning with an Information Maximization CriterionabstractSelf-supervised learning allows AI systems to learn effective representations from large amounts of data using tasks that do not require costly labeling. Mode collapse, i.e., the model producing identical representations for all inputs, is a central problem to many self-supervised learning approaches, making self-supervised tasks, such as matching distorted variants of the inputs, ineffective. In this article, we argue that a straightforward application of information maximization among alternative latent representations of the same input naturally solves the collapse problem and achieves competitive empirical results. We propose a self-supervised learning method, CorInfoMax, that uses a second-order statistics-based mutual information measure that reflects the level of correlation among its arguments. Maximizing this correlative information measure between alternative representations of the same input serves two purposes: (1) it avoids the collapse problem by generating feature vectors with non-degenerate covariances; (2) it establishes relevance among alternative representations by increasing the linear dependence among them. An approximation of the proposed information maximization objective simplifies to a Euclidean distance-based objective function regularized by the log-determinant of the feature covariance matrix. The regularization term acts as a natural barrier against feature space degeneracy. Consequently, beyond avoiding complete output collapse to a single point, the proposed approach also prevents dimensional collapse by encouraging the spread of information across the whole feature space. Numerical experiments demonstrate that CorInfoMax achieves better or competitive performance results relative to the state-of-the-art SSL approaches. Serdar Ozsoy, Shadi Hamdan, Sercan Ö. Arik, Deniz Yuret, Alper T. Erdogan |
NeurIPS | 3 |
| 2021 | TabNet: Attentive Interpretable Tabular LearningabstractWe propose a novel high-performance and interpretable canonical deep tabular data learning architecture, TabNet. TabNet uses sequential attention to choose which features to reason from at each decision step, enabling interpretability and more efficient learning as the learning capacity is used for the most salient features. We demonstrate that TabNet outperforms other variants on a wide range of non-performance-saturated tabular datasets and yields interpretable feature attributions plus insights into its global behavior. Finally, we demonstrate self-supervised learning for tabular data, significantly improving performance when unlabeled data is abundant. Sercan Ö. Arik, Tomas Pfister |
AAAI | 1 |
| 2021 | Controlling Neural Networks with Rule RepresentationsabstractWe propose a novel training method that integrates rules into deep learning, in a way the strengths of the rules are controllable at inference. Deep Neural Networks with Controllable Rule Representations (DeepCTRL) incorporates a rule encoder into the model coupled with a rule-based objective, enabling a shared representation for decision making. DeepCTRL is agnostic to data type and model architecture. It can be applied to any kind of rule defined for inputs and outputs. The key aspect of DeepCTRL is that it does not require retraining to adapt the rule strength -- at inference, the user can adjust it based on the desired operation point on accuracy vs. rule verification ratio. In real-world domains where incorporating rules is critical -- such as Physics, Retail and Healthcare -- we show the effectiveness of DeepCTRL in teaching rules for deep learning. DeepCTRL improves the trust and reliability of the trained models by significantly increasing their rule verification ratio, while also providing accuracy gains at downstream tasks. Additionally, DeepCTRL enables novel use cases such as hypothesis testing of the rules on data samples, and unsupervised adaptation based on shared rules between datasets. Sungyong Seo, Sercan Ö. Arik, Jinsung Yoon, Kihyuk Sohn, Tomas Pfister |
NeurIPS | 2 |
| 2020 | Distilling Effective Supervision From Severe Label NoiseabstractCollecting large-scale data with clean labels for supervised training of neural networks is practically challenging. Although noisy labels are usually cheap to acquire, existing methods suffer a lot from label noise. This paper targets at the challenge of robust training at high label noise regimes. The key insight to achieve this goal is to wisely leverage a small trusted set to estimate exemplar weights and pseudo labels for noisy data in order to reuse them for supervised training. We present a holistic framework to train deep neural networks in a way that is highly invulnerable to label noise. Our method sets the new state of the art on various types of label noise and achieves excellent performance on large-scale datasets with real-world label noise. For instance, on CIFAR100 with a 40% uniform noise ratio and only 10 trusted labeled data per class, our method achieves 80.2% classification accuracy, where the error rate is only 1.4% higher than a neural network trained without label noise. Moreover, increasing the noise ratio to 80%, our method still maintains a high accuracy of 75.5%, compared to the previous best accuracy 48.2%. Han Zhang 0010, Sercan Ö. Arik, Honglak Lee, Tomas Pfister |
CVPR | 3 |
| 2020 | Consistency-Based Semi-supervised Active Learning: Towards Minimizing Labeling Cost
Mingfei Gao, Sercan Ö. Arik, Larry Davis 0001, Tomas Pfister |
ECCV (10) | 4 |
| 2020 | Learning to Transfer Learn: Reinforcement Learning-Based Selection for Adaptive Transfer Learning
Linchao Zhu, Sercan Ö. Arik, Yi Yang 0001, Tomas Pfister |
ECCV (27) | 2 |
| 2020 | Distance-Based Learning from Errors for Confidence Calibration
Chen Xing, Sercan Ö. Arik, Tomas Pfister |
ICLR | 2 |
| 2020 | Data Valuation using Reinforcement LearningabstractQuantifying the value of data is a fundamental problem in machine learning and has multiple important use cases: (1) building insights about the dataset and task, (2) domain adaptation, (3) corrupted sample discovery, and (4) robust learning. We propose Data Valuation using Reinforcement Learning (DVRL), to adaptively learn data values jointly with the predictor model. DVRL uses a data value estimator (DVE) to learn how likely each datum is used in training of the predictor model. DVE is trained using a reinforcement signal that reflects performance on the target task. We demonstrate that DVRL yields superior data value estimates compared to alternative methods across numerous datasets and application scenarios. The corrupted sample discovery performance of DVRL is close to optimal in many regimes (i.e. as if the noisy samples were known apriori), and for domain adaptation and robust learning DVRL significantly outperforms state-of-the-art by 14.6% and 10.8%, respectively. Jinsung Yoon, Sercan Ö. Arik, Tomas Pfister |
ICML | 2 |
| 2020 | Interpretable Sequence Learning for Covid-19 ForecastingabstractWe propose a novel approach that integrates machine learning into compartmental disease modeling (e.g., SEIR) to predict the progression of COVID-19. Our model is explainable by design as it explicitly shows how different compartments evolve and it uses interpretable encoders to incorporate covariates and improve performance. Explainability is valuable to ensure that the model's forecasts are credible to epidemiologists and to instill confidence in end-users such as policy makers and healthcare institutions. Our model can be applied at different geographic resolutions, and we demonstrate it for states and counties in the United States. We show that our model provides more accurate forecasts compared to the alternatives, and that it provides qualitatively meaningful explanatory insights. Sercan Ö. Arik, Chun-Liang Li, Jinsung Yoon, Rajarishi Sinha, Arkady Epshteyn, Long T. Le, Vikas Menon, Shashank Singh 0005, Leyou Zhang, Martin Nikoltchev, Yash Sonthalia, Hootan Nakhost, Elli Kanal, Tomas Pfister |
NeurIPS | 1 |
| 2020 | On Completeness-aware Concept-Based Explanations in Deep Neural NetworksabstractHuman explanations of high-level decisions are often expressed in terms of key concepts the decisions are based on. In this paper, we study such concept-based explainability for Deep Neural Networks (DNNs). First, we define the notion of \emph{completeness}, which quantifies how sufficient a particular set of concepts is in explaining a model's prediction behavior based on the assumption that complete concept scores are sufficient statistics of the model prediction. Next, we propose a concept discovery method that aims to infer a complete set of concepts that are additionally encouraged to be interpretable, which addresses the limitations of existing methods on concept explanations. To define an importance score for each discovered concept, we adapt game-theoretic notions to aggregate over sets and propose \emph{ConceptSHAP}. Via proposed metrics and user studies, on a synthetic dataset with apriori-known concept explanations, as well as on real-world image and language datasets, we validate the effectiveness of our method in finding concepts that are both complete in explaining the decisions and interpretable. Chih-Kuan Yeh, Been Kim, Sercan Ö. Arik, Chun-Liang Li, Tomas Pfister, Pradeep Ravikumar |
NeurIPS | 3 |
| 2020 | ProtoAttend: Attention-Based Prototypical LearningabstractWe propose a novel inherently interpretable machine learning method that bases decisions on few relevant examples that we call prototypes. Our method, ProtoAttend, can be integrated into a wide range of neural network architectures including pre-trained models. It utilizes an attention mechanism that relates the encoded representations to samples in order to determine prototypes. Protoattend yields superior results in three high impact problems without sacrificing accuracy of the original model: (1)it enables high-quality interpretability that outputs samples most relevant to the decision-making (i.e. a sample-based interpretability method); (2) it achieves state of the art confidence estimation by quantifying the mismatch across prototype labels; and (3) it obtains state of the art in distribution mismatch detection. All these can be achieved with minimal additional test time and a practically viable training time computational cost. Sercan Ö. Arik, Tomas Pfister |
J. Mach. Learn. Res. | 1 |
| 2019 | Fast Spectrogram Inversion Using Multi-Head Convolutional Neural NetworksabstractWe propose the multi-head convolutional neural network (MCNN) for waveform synthesis from spectrograms. Nonlinear interpolation in MCNN is employed with transposed convolution layers in parallel heads. MCNN enables significantly better utilization of modern multi-core processors than commonly used iterative algorithms like Griffin-Lim, and yields very fast (more than 300 × real time) runtime. For training of MCNN, we use a large-scale speech recognition dataset and losses defined on waveforms that are related to perceptual audio quality. We demonstrate that MCNN constitutes a very promising approach for high-quality speech synthesis, without any iterative algorithms or autoregression in computations. Sercan Ö. Arik, Heewoo Jun, Gregory Frederick Diamos |
IEEE Signal Process. Lett. | 1 |
| 2018 | Deep Voice 3: Scaling Text-to-Speech with Convolutional Sequence Learning
Wei Ping, Kainan Peng, Andrew Gibiansky, Sercan Ö. Arik, Ajay Kannan, Sharan Narang, Jonathan Raiman, John Miller 0001 |
ICLR (Poster) | 4 |
| 2018 | Neural Voice Cloning with a Few SamplesabstractVoice cloning is a highly desired feature for personalized speech interfaces. We introduce a neural voice cloning system that learns to synthesize a person's voice from only a few audio samples. We study two approaches: speaker adaptation and speaker encoding. Speaker adaptation is based on fine-tuning a multi-speaker generative model. Speaker encoding is based on training a separate model to directly infer a new speaker embedding, which will be applied to a multi-speaker generative model. In terms of naturalness of the speech and similarity to the original speaker, both approaches can achieve good performance, even with a few cloning audios. While speaker adaptation can achieve slightly better naturalness and similarity, cloning time and required memory for the speaker encoding approach are significantly less, making it more favorable for low-resource deployment. Sercan Ö. Arik, Jitong Chen, Kainan Peng, Wei Ping, Yanqi Zhou |
NeurIPS | 1 |
| 2017 | Scaling optical networks using full-spectrum spatial switchingabstractAccommodating sustained exponential traffic growth in optical networks requires scaling the spatial dimension using space-division multiplexing. Numerous uncoupled spatial channels may be realized by activating multiple parallel fibers or cores in multicore fibers. A multiplicity of uncoupled spatial channels will render the granularity provided by multiple wavelength channels less essential in enabling optical switching. In this paper, we investigate a possible paradigm shift in optical node architectures, in which nodes based on wavelength-selective switching are replaced by those based on simple spatial switching. We compare spatial switching of full-spectrum superchannels to wavelength switching of uncoupled spatial superchannels, considering the evolution of traffic over time. Our results show that spatial switching may achieve more efficient scaling than wavelength switching in roughly 10 to 17 years, depending on the traffic growth rate assumed. Alaelson C. Jatoba-Neto, Christian Esteve Rothenberg, Darli A. A. Mello, Sercan Ö. Arik, Joseph M. Kahn |
HPSR | 4 |
| 2017 | Deep Voice: Real-time Neural Text-to-SpeechabstractWe present Deep Voice, a production-quality text-to-speech system constructed entirely from deep neural networks. Deep Voice lays the groundwork for truly end-to-end neural speech synthesis. The system comprises five major building blocks: a segmentation model for locating phoneme boundaries, a grapheme-to-phoneme conversion model, a phoneme duration prediction model, a fundamental frequency prediction model, and an audio synthesis model. For the segmentation model, we propose a novel way of performing phoneme boundary detection with deep neural networks using connectionist temporal classification (CTC) loss. For the audio synthesis model, we implement a variant of WaveNet that requires fewer parameters and trains faster than the original. By using a neural network for each component, our system is simpler and more flexible than traditional text-to-speech systems, where each component requires laborious feature engineering and extensive domain expertise. Finally, we show that inference with our system can be performed faster than real time and describe optimized WaveNet inference kernels on both CPU and GPU that achieve up to 400x speedups over existing implementations. Sercan Ö. Arik, Mike Chrzanowski, Adam Coates 0002, Gregory Frederick Diamos, Andrew Gibiansky, Yongguo Kang, John Miller 0001, Andrew Y. Ng, Jonathan Raiman, Shubho Sengupta, Mohammad Shoeybi |
ICML | 1 |
| 2017 | Convolutional Recurrent Neural Networks for Small-Footprint Keyword SpottingabstractKeyword spotting (KWS) constitutes a major component of human-technology interfaces.Maximizing the detection accuracy at a low false alarm (FA) rate, while minimizing the footprint size, latency and complexity are the goals for KWS.Towards achieving them, we study Convolutional Recurrent Neural Networks (CRNNs).Inspired by large-scale state-ofthe-art speech recognition systems, we combine the strengths of convolutional layers and recurrent layers to exploit local structure and long-range context.We analyze the effect of architecture parameters, and propose training strategies to improve performance.With only ~230k parameters, our CRNN model yields acceptably low latency, and achieves 97.71% accuracy at 0.5 FA/hour for 5 dB signal-to-noise ratio. Sercan Ö. Arik, Markus Kliegl, Rewon Child, Joel Hestness, Andrew Gibiansky, Christopher Fougner, Ryan Prenger, Adam Coates 0002 |
INTERSPEECH | 1 |
| 2017 | Deep Voice 2: Multi-Speaker Neural Text-to-SpeechabstractWe introduce a technique for augmenting neural text-to-speech (TTS) with low-dimensional trainable speaker embeddings to generate different voices from a single model. As a starting point, we show improvements over the two state-of-the-art approaches for single-speaker neural TTS: Deep Voice 1 and Tacotron. We introduce Deep Voice 2, which is based on a similar pipeline with Deep Voice 1, but constructed with higher performance building blocks and demonstrates a significant audio quality improvement over Deep Voice 1. We improve Tacotron by introducing a post-processing neural vocoder, and demonstrate a significant audio quality improvement. We then demonstrate our technique for multi-speaker speech synthesis for both Deep Voice 2 and Tacotron on two multi-speaker TTS datasets. We show that a single neural TTS system can learn hundreds of unique voices from less than half an hour of data per speaker, while achieving high audio quality synthesis and preserving the speaker identities almost perfectly. Andrew Gibiansky, Sercan Ö. Arik, Gregory Frederick Diamos, John Miller 0001, Kainan Peng, Wei Ping, Jonathan Raiman, Yanqi Zhou |
NIPS | 2 |
| 2011 | Alignment of uncalibrated images for multi-view classificationabstractEfficient solutions for the classification of multi-view images can be built on graph-based algorithms when little information is known about the scene or cameras. Such methods typically require a pair-wise similarity measure between images, where a common choice is the Euclidean distance. However, the accuracy of the Euclidean distance as a similarity measure is restricted to cases where images are captured from nearby viewpoints. In settings with large transformations and viewpoint changes, alignment of images is necessary prior to distance computation. We propose a method for the registration of uncalibrated images that capture the same 3D scene or object. We model the depth map of the scene as an algebraic surface, which yields a warp model in the form of a rational function between image pairs. The warp model is computed by minimizing the registration error, where the registered image is a weighted combination of two images generated with two different warp functions estimated from feature matches and image intensity functions in order to provide robust registration. We demonstrate the flexibility of our alignment method by experimentation on several wide-baseline image pairs with arbitrary scene geometries and texture levels. Moreover, the results on multi-view image classification suggest that the proposed alignment method can be effectively used in graph-based classification algorithms for the computation of pairwise distances where it achieves significant improvements over distance computation without prior alignment. Sercan Ö. Arik, Elif Vural, Pascal Frossard |
ICIP | 1 |