EDBT 2026 Demo / reviewers in the wild / expert
Julian Martin Eisenschlos
dblp:262/3990 · also Julian Eisenschlos
· DBLP profile ↗
17ranked-venue papers
4as first author
14since 2021 · last 2025
—ORCID · none
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 17 · 4 first-author · 14 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | TANQ: An Open Domain Dataset of Table Answered QuestionsabstractAbstract Language models, potentially augmented with tool usage such as retrieval, are becoming the go-to means of answering questions. Understanding and answering questions in real-world settings often requires retrieving information from different sources, processing and aggregating data to extract insights, and presenting complex findings in form of structured artifacts such as novel tables, charts, or infographics. In this paper, we introduce TANQ,1 the first open-domain question answering dataset where the answers require building tables from information across multiple sources. We release the full source attribution for every cell in the resulting table and benchmark state-of-the-art language models in open, oracle, and closed book setups. Our best-performing baseline, Gemini Flash, reaches an overall F1 score of 60.7, lagging behind human performance by 12.3 points. We analyze baselines’ performance across different dataset attributes such as different skills required for this task, including multi-hop reasoning, math operations, and unit conversions. We further discuss common failures in model-generated answers, suggesting that TANQ is a complex task with many challenges ahead. Mubashara Akhtar, Chenxi Pang, Andreea Marzoca, Yasemin Altun, Julian Martin Eisenschlos |
Trans. Assoc. Comput. Linguistics | 5 |
| 2024 | Faithful Chart Summarization with ChaTS-PiabstractChart-to-summary generation can help explore data, communicate insights, and help the visually impaired people.Multi-modal generative models have been used to produce fluent summaries, but they can suffer from factual and perceptual errors.In this work we present CHATS-CRITIC , a reference-free chart summarization metric for scoring faithfulness.CHATS-CRITIC is composed of an image-to-text model to recover the table from a chart, and a tabular entailment model applied to score the summary sentence by sentence.We find that CHATS-CRITIC evaluates the summary quality according to human ratings better than reference-based metrics, either learned or n-gram based, and can be further used to fix candidate summaries by removing not supported sentences.We then introduce CHATS-PI , a chart-to-summary pipeline that leverages CHATS-CRITIC during inference to fix and rank sampled candidates from any chart-summarization model.We evaluate CHATS-PI and CHATS-CRITIC using human raters, establishing state-of-the-art results on two popular chart-to-summary datasets.1 Syrine Krichene, Francesco Piccinno, Fangyu Liu 0001, Julian Martin Eisenschlos |
ACL (1) | 4 |
| 2024 | Chain-of-Table: Evolving Tables in the Reasoning Chain for Table UnderstandingabstractTable-based reasoning with large language models (LLMs) is a promising direction to tackle many table understanding tasks, such as table-based question answering and fact verification. Compared with generic reasoning, table-based reasoning requires the extraction of underlying semantics from both free-form questions and semi-structured tabular data. Chain-of-Thought and its similar approaches incorporate the reasoning chain in the form of textual context, but it is still an open question how to effectively leverage tabular data in the reasoning chain. We propose the Chain-of-Table framework, where tabular data is explicitly used in the reasoning chain as a proxy for intermediate thoughts. Specifically, we guide LLMs using in-context learning to iteratively generate operations and update the table to represent a tabular reasoning chain. LLMs can therefore dynamically plan the next operation based on the results of the previous ones. This continuous evolution of the table forms a chain, showing the reasoning process for a given tabular problem. The chain carries structured information of the intermediate results, enabling more accurate and reliable predictions. Chain-of-Table achieves new state-of-the-art performance on WikiTQ, FeTaQA, and TabFact benchmarks across multiple LLM choices. Zilong Wang 0002, Chun-Liang Li, Julian Martin Eisenschlos, Vincent Perot, Zifeng Wang 0002, Lesly Miculicich, Yasuhisa Fujii, Jingbo Shang, Chen-Yu Lee, Tomas Pfister |
ICLR | 4 |
| 2024 | TableRAG: Million-Token Table Understanding with Language ModelsabstractRecent advancements in language models (LMs) have notably enhanced their ability to reason with tabular data, primarily through program-aided mechanisms that manipulate and analyze tables.
However, these methods often require the entire table as input, leading to scalability challenges due to the positional bias or context length constraints.
In response to these challenges, we introduce TableRAG, a Retrieval-Augmented Generation (RAG) framework specifically designed for LM-based table understanding.
TableRAG leverages query expansion combined with schema and cell retrieval to pinpoint crucial information before providing it to the LMs.
This enables more efficient data encoding and precise retrieval, significantly reducing prompt lengths and mitigating information loss.
We have developed two new million-token benchmarks from the Arcade and BIRD-SQL datasets to thoroughly evaluate TableRAG's effectiveness at scale.
Our results demonstrate that TableRAG's retrieval design achieves the highest retrieval quality, leading to the new state-of-the-art performance on large-scale table understanding. Si-An Chen, Lesly Miculicich, Julian Martin Eisenschlos, Zifeng Wang 0002, Zilong Wang 0002, Yanfei Chen, Yasuhisa Fujii, Hsuan-Tien Lin, Chen-Yu Lee, Tomas Pfister |
NeurIPS | 3 |
| 2023 | MatCha: Enhancing Visual Language Pretraining with Math Reasoning and Chart DerenderingabstractFangyu Liu, Francesco Piccinno, Syrine Krichene, Chenxi Pang, Kenton Lee, Mandar Joshi, Yasemin Altun, Nigel Collier, Julian Eisenschlos. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Fangyu Liu 0001, Francesco Piccinno, Syrine Krichene, Chenxi Pang, Kenton Lee, Mandar Joshi, Yasemin Altun, Nigel Collier, Julian Martin Eisenschlos |
ACL (1) | 9 |
| 2023 | DiffQG: Generating Questions to Summarize Factual ChangesabstractJeremy R. Cole, Palak Jain, Julian Martin Eisenschlos, Michael J.Q. Zhang, Eunsol Choi, Bhuwan Dhingra. Proceedings of the 17th Conference of the European Chapter of the Association for Computational Linguistics. 2023. Jeremy R. Cole, Palak Jain 0006, Julian Martin Eisenschlos, Michael J. Q. Zhang, Eunsol Choi, Bhuwan Dhingra |
EACL | 3 |
| 2023 | WinoDict: Probing language models for in-context word acquisitionabstractWe introduce a new in-context learning paradigm to measure Large Language Models' (LLMs) ability to learn novel words during inference.In particular, we rewrite Winogradstyle co-reference resolution problems by replacing the key concept word with a synthetic but plausible word that the model must understand to complete the task.Solving this task requires the model to make use of the dictionary definition of the new word given in the prompt.This benchmark addresses word acquisition, one important aspect of the diachronic degradation known to afflict LLMs.As LLMs are frozen in time at the moment they are trained, they are normally unable to reflect the way language changes over time.We show that the accuracy of LLMs compared to the original Winograd tasks decreases radically in our benchmark, thus identifying a limitation of current models and providing a benchmark to measure future improvements in LLMs ability to do in-context learning. Julian Martin Eisenschlos, Jeremy R. Cole, Fangyu Liu 0001, William W. Cohen |
EACL | 1 |
| 2023 | Selectively Answering Ambiguous QuestionsabstractTrustworthy language models should abstain from answering questions when they do not know the answer.However, the answer to a question can be unknown for a variety of reasons.Prior research has focused on the case in which the question is clear and the answer is unambiguous but possibly unknown.But the answer to a question can also be unclear due to uncertainty of the questioner's intent or context.We investigate question answering from this perspective, focusing on answering a subset of questions with a high degree of accuracy, from a set of questions in which many are inherently ambiguous.In this setting, we find that the most reliable approach to decide when to abstain involves quantifying repetition within sampled model outputs, rather than the model's likelihood or self-verification as used in prior work.We find this to be the case across different types of uncertainty and model scales, and with or without instruction tuning.Our results suggest that sampling-based confidence scores help calibrate answers to relatively unambiguous questions, with more dramatic improvements on ambiguous questions. Jeremy R. Cole, Michael J. Q. Zhang, Daniel Gillick, Julian Martin Eisenschlos, Bhuwan Dhingra, Jacob Eisenstein |
EMNLP | 4 |
| 2023 | Universal Self-Adaptive PromptingabstractA hallmark of modern large language models (LLMs) is their impressive general zero-shot and few-shot abilities, often elicited through in-context learning (ICL) via prompting.However, while highly coveted and being the most general, zero-shot performances in LLMs are still typically weaker due to the lack of guidance and the difficulty of applying existing automatic prompt design methods in general tasks when ground-truth labels are unavailable.In this study, we address this by presenting Universal Self-Adaptive Prompting (USP), an automatic prompt design approach specifically tailored for zero-shot learning (while compatible with few-shot).Requiring only a small amount of unlabeled data and an inferenceonly LLM, USP is highly versatile: to achieve universal prompting, USP categorizes a possible NLP task into one of the three possible task types and then uses a corresponding selector to select the most suitable queries and zero-shot model-generated responses as pseudo-demonstrations, thereby generalizing ICL to the zero-shot setup in a fully automated way.We evaluate USP with PaLM and PaLM 2 models and demonstrate performances that are considerably stronger than standard zero-shot baselines and often comparable to or even superior to few-shot baselines across more than 40 natural language understanding, natural language generation, and reasoning tasks. Xingchen Wan, Ruoxi Sun 0002, Hootan Nakhost, Hanjun Dai, Julian Martin Eisenschlos, Sercan Ö. Arik, Tomas Pfister |
EMNLP | 5 |
| 2023 | Pix2Struct: Screenshot Parsing as Pretraining for Visual Language UnderstandingabstractVisually-situated language is ubiquitous---sources range from textbooks with diagrams to web pages with images and tables, to mobile apps with buttons and forms. Perhaps due to this diversity, previous work has typically relied on domain-specific recipes with limited sharing of the underlying data, model architectures, and objectives. We present Pix2Struct, a pretrained image-to-text model for purely visual language understanding, which can be finetuned on tasks containing visually-situated language. Pix2Struct is pretrained by learning to parse masked screenshots of web pages into simplified HTML. The web, with its richness of visual elements cleanly reflected in the HTML structure, provides a large source of pretraining data well suited to the diversity of downstream tasks. Intuitively, this objective subsumes common pretraining signals such as OCR, language modeling, and image captioning. In addition to the novel pretraining strategy, we introduce a variable-resolution input representation and a more flexible integration of language and vision inputs, where language prompts such as questions are rendered directly on top of the input image. For the first time, we show that a single pretrained model can achieve state-of-the-art results in six out of nine tasks across four domains: documents, illustrations, user interfaces, and natural images. Kenton Lee, Mandar Joshi, Iulia Turc, Hexiang Hu, Fangyu Liu 0001, Julian Martin Eisenschlos, Urvashi Khandelwal, Peter Shaw 0004, Ming-Wei Chang, Kristina Toutanova |
ICML | 6 |
| 2022 | Time-Aware Language Models as Temporal Knowledge BasesabstractAbstract Many facts come with an expiration date, from the name of the President to the basketball team Lebron James plays for. However, most language models (LMs) are trained on snapshots of data collected at a specific moment in time. This can limit their utility, especially in the closed-book setting where the pretraining corpus must contain the facts the model should memorize. We introduce a diagnostic dataset aimed at probing LMs for factual knowledge that changes over time and highlight problems with LMs at either end of the spectrum—those trained on specific slices of temporal data, as well as those trained on a wide range of temporal data. To mitigate these problems, we propose a simple technique for jointly modeling text with its timestamp. This improves memorization of seen facts from the training time period, as well as calibration on predictions about unseen facts from future time periods. We also show that models trained with temporal context can be efficiently “refreshed” as new data arrives, without the need for retraining from scratch. Bhuwan Dhingra, Jeremy R. Cole, Julian Martin Eisenschlos, Daniel Gillick, Jacob Eisenstein, William W. Cohen |
Trans. Assoc. Comput. Linguistics | 3 |
| 2021 | MATE: Multi-view Attention for Table Transformer EfficiencyabstractThis work presents a sparse-attention Transformer architecture for modeling documents that contain large tables.Tables are ubiquitous on the web, and are rich in information.However, more than 20% of relational tables on the web have 20 or more rows (Cafarella et al., 2008), and these large tables present a challenge for current Transformer models, which are typically limited to 512 tokens.Here we propose MATE, a novel Transformer architecture designed to model the structure of web tables.MATE uses sparse attention in a way that allows heads to efficiently attend to either rows or columns in a table.This architecture scales linearly with respect to speed and memory, and can handle documents containing more than 8000 tokens with current accelerators.MATE also has a more appropriate inductive bias for tabular data, and sets a new state-of-the-art for three table reasoning datasets.For HY-BRIDQA (Chen et al., 2020b), a dataset that involves large documents containing tables, we improve the best prior result by 19 points. Julian Martin Eisenschlos, Maharshi Gor, Thomas Müller 0009, William W. Cohen |
EMNLP (1) | 1 |
| 2021 | Fool Me Twice: Entailment from Wikipedia GamificationabstractJulian Eisenschlos, Bhuwan Dhingra, Jannis Bulian, Benjamin Börschinger, Jordan Boyd-Graber. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Julian Martin Eisenschlos, Bhuwan Dhingra, Jannis Bulian, Benjamin Börschinger, Jordan L. Boyd-Graber |
NAACL-HLT | 1 |
| 2021 | Open Domain Question Answering over Tables via Dense RetrievalabstractJonathan Herzig, Thomas Müller, Syrine Krichene, Julian Eisenschlos. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Jonathan Herzig, Thomas Müller 0009, Syrine Krichene, Julian Martin Eisenschlos |
NAACL-HLT | 4 |
| 2020 | TaPas: Weakly Supervised Table Parsing via Pre-trainingabstractAnswering natural language questions over tables is usually seen as a semantic parsing task.To alleviate the collection cost of full logical forms, one popular approach focuses on weak supervision consisting of denotations instead of logical forms.However, training semantic parsers from weak supervision poses difficulties, and in addition, the generated logical forms are only used as an intermediate step prior to retrieving the denotation.In this paper, we present TAPAS, an approach to question answering over tables without generating logical forms.TAPAS trains from weak supervision, and predicts the denotation by selecting table cells and optionally applying a corresponding aggregation operator to such selection.TAPAS extends BERT's architecture to encode tables as input, initializes from an effective joint pre-training of text segments and tables crawled from Wikipedia, and is trained end-to-end.We experiment with three different semantic parsing datasets, and find that TAPAS outperforms or rivals semantic parsing models by improving state-of-the-art accuracy on SQA from 55.1 to 67.2 and performing on par with the state-of-the-art on WIKISQL and WIKITQ, but with a simpler model architecture.We additionally find that transfer learning, which is trivial in our setting, from WIK-ISQL to WIKITQ, yields 48.7 accuracy, 4.2 points above the state-of-the-art. Jonathan Herzig, Pawel Krzysztof Nowak, Thomas Müller 0009, Francesco Piccinno, Julian Martin Eisenschlos |
ACL | 5 |
| 2020 | SoftSort: A Continuous Relaxation for the argsort OperatorabstractWhile sorting is an important procedure in computer science, the argsort operator - which takes as input a vector and returns its sorting permutation - has a discrete image and thus zero gradients almost everywhere. This prohibits end-to-end, gradient-based learning of models that rely on the argsort operator. A natural way to overcome this problem is to replace the argsort operator with a continuous relaxation. Recent work has shown a number of ways to do this, but the relaxations proposed so far are computationally complex. In this work we propose a simple continuous relaxation for the argsort operator which has the following qualities: it can be implemented in three lines of code, achieves state-of-the-art performance, is easy to reason about mathematically - substantially simplifying proofs - and is faster than competing approaches. We open source the code to reproduce all of the experiments and results. Sebastian Prillo, Julian Martin Eisenschlos |
ICML | 2 |
| 2019 | MultiFiT: Efficient Multi-lingual Language Model Fine-tuningabstractJulian Eisenschlos, Sebastian Ruder, Piotr Czapla, Marcin Kadras, Sylvain Gugger, Jeremy Howard. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Julian Martin Eisenschlos, Sebastian Ruder, Piotr Czapla, Marcin Kardas, Sylvain Gugger, Jeremy Howard |
EMNLP/IJCNLP (1) | 1 |