EDBT 2026 Demo / reviewers in the wild / expert
Daniel Hershcovich
dblp:145/9324
· DBLP profile ↗
37ranked-venue papers
5as first author
29since 2021 · last 2026
0000-0002-3966-8708ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 33 · 5 first-author · 26 since 2021Databases, data management, data science and information retrieval · 2 · 2 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Systems, architecture and hardware · 1Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Culinary Crossroads: A RAG Framework for Enhancing Diversity in Cross-Cultural Recipe AdaptationabstractTianyi Hu, Andrea Morales-Garzón, Jingyi Zheng, Maria Maistro, Daniel Hershcovich. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Andrea Morales-Garzón, Jingyi Zheng, Maria Maistro, Daniel Hershcovich |
ACL (1) | 5 |
| 2026 | NOVELSUM: Evaluating Long-Form Summary Generation for Historical Scandinavian Novels
Ali Allaith, Alexander Conroy, Kirstine Nielsen Degn, Jens Bjerring-Hansen, Daniel Hershcovich |
LREC | 5 |
| 2026 | ScholarLens: Tracking the growing penetration of Large Language Models in scholarly writing and peer reviewabstractAlthough the widespread use of Large Language Models (LLMs) brings convenience, it also raises concerns about the credibility of academic research and scholarly processes. To better understand the extent and characteristics of LLM use in scholarly writing and peer review, the penetration of LLMs across academic workflows is evaluated from multiple perspectives and dimensions, providing compelling evidence of their growing influence. A framework consisting of two components is proposed: ScholarLens , a curated dataset of human-written and LLM-generated content across scholarly writing and peer review for multi-perspective evaluation, and LLMetrica , a tool for assessing LLM penetration using rule-based metrics and model-based detectors for multi-dimensional evaluation. The effectiveness of LLMetrica is demonstrated through experiments, revealing the increasing role of LLMs in scholarly processes. These findings emphasize the need for transparency, accountability, and ethical practices in the use of LLMs to maintain academic credibility. Li Zhou 0010, Xunlian Dai, Daniel Hershcovich, Lihui Wang 0002 |
Eng. Appl. Artif. Intell. | 4 |
| 2025 | Evaluating Multimodal Language Models as Visual Assistants for Visually Impaired UsersabstractAntonia Karamolegkou, Malvina Nikandrou, Georgios Pantazopoulos, Danae Sanchez Villegas, Phillip Rust, Ruchira Dhar, Daniel Hershcovich, Anders Søgaard. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Antonia Karamolegkou, Malvina Nikandrou, Georgios Pantazopoulos, Danae Sanchez Villegas, Phillip Rust, Ruchira Dhar, Daniel Hershcovich, Anders Søgaard |
ACL (1) | 7 |
| 2025 | Dying or Departing? Euphemism Detection for Death Discourse in Historical TextsabstractEuphemisms are a linguistic device used to soften discussions of sensitive or uncomfortable topics, with death being a prominent example. In this paper, we present a study on the detection of death-related euphemisms in historical literary texts from a corpus containing Danish and Norwegian novels from the late 19th century. We introduce an annotated dataset of euphemistic and literal references to death, including both common and rare euphemisms, ranging from well-established terms to more culturally nuanced expressions. We evaluate the performances of state-of-the-art pre-trained language models fine-tuned for euphemism detection. Our findings show that fixed, literal expressions of death became less frequent over time, while metaphorical euphemisms grew in prevalence. Additionally, euphemistic language was more common in historical novels, whereas contemporary novels tended to refer to death more literally, reflecting the rise of secularism. These results shed light on the shifting discourse on death during a period when the concept of death as final became prominent. Ali Allaith, Alexander Conroy, Jens Bjerring-Hansen, Bolette S. Pedersen, Carsten Levisen, Daniel Hershcovich |
COLING | 6 |
| 2025 | Beyond Demographics: Enhancing Cultural Value Survey Simulation with Multi-Stage Personality-Driven Cognitive ReasoningabstractIntroducing MARK, the Multi-stAge Reasoning frameworK for cultural value survey response simulation, designed to enhance the accuracy, steerability, and interpretability of large language models in this task.The system is inspired by the type dynamics theory in the MBTI psychological framework for personality research.It effectively predicts and utilizes human demographic information for simulation: life-situational stress analysis, group-level personality prediction, and self-weighted cognitive imitation.Experiments on the World Values Survey show that MARK outperforms existing baselines by 10% accuracy and reduces the divergence between model predictions and human preferences.This highlights the potential of our framework to improve zero-shot personalization and help social scientists interpret model predictions.1 Chao Gao 0014, Yong Cao 0001, Daniel Hershcovich, Jinguang Gu |
EMNLP | 7 |
| 2025 | Specializing Large Language Models to Simulate Survey Response Distributions for Global PopulationsabstractYong Cao, Haijiang Liu, Arnav Arora, Isabelle Augenstein, Paul Röttger, Daniel Hershcovich. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Yong Cao 0001, Arnav Arora, Isabelle Augenstein, Paul Röttger, Daniel Hershcovich |
NAACL (Long Papers) | 6 |
| 2025 | Does Mapo Tofu Contain Coffee? Probing LLMs for Food-related Cultural KnowledgeabstractLi Zhou, Taelin Karidi, Wanlong Liu, Nicolas Garneau, Yong Cao, Wenyu Chen, Haizhou Li, Daniel Hershcovich. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Li Zhou 0010, Taelin Karidi, Wanlong Liu, Nicolas Garneau, Yong Cao 0001, Wenyu Chen 0001, Haizhou Li 0001, Daniel Hershcovich |
NAACL (Long Papers) | 8 |
| 2025 | Towards realistic evaluation of cultural value alignment in large language models: Diversity enhancement for survey response simulation
Yong Cao 0001, Chen Qiu 0005, Jinguang Gu, Maofu Liu, Daniel Hershcovich |
Inf. Process. Manag. | 7 |
| 2024 | Development and Evaluation of Pre-trained Language Models for Historical Danish and Norwegian Literary TextsabstractWe develop and evaluate the first pre-trained language models specifically tailored for historical Danish and Norwegian texts. Three models are trained on a corpus of 19th-century Danish and Norwegian literature: two directly on the corpus with no prior pre-training, and one with continued pre-training. To evaluate the models, we utilize an existing sentiment classification dataset, and additionally introduce a new annotated word sense disambiguation dataset focusing on the concept of fate. Our assessment reveals that the model employing continued pre-training outperforms the others in two downstream NLP tasks on historical texts. Specifically, we observe substantial improvement in sentiment classification and word sense disambiguation compared to models trained on contemporary texts. These results highlight the effectiveness of continued pre-training for enhancing performance across various NLP tasks in historical text analysis. Ali Allaith, Alexander Conroy, Jens Bjerring-Hansen, Daniel Hershcovich |
LREC/COLING | 4 |
| 2024 | Geo-Encoder: A Chunk-Argument Bi-Encoder Framework for Chinese Geographic Re-RankingabstractYong Cao, Ruixue Ding, Boli Chen, Xianzhi Li, Min Chen, Daniel Hershcovich, Pengjun Xie, Fei Huang. Proceedings of the 18th Conference of the European Chapter of the Association for Computational Linguistics (Volume 1: Long Papers). 2024. Yong Cao 0001, Ruixue Ding, Boli Chen, Xianzhi Li 0001, Min Chen 0003, Daniel Hershcovich, Pengjun Xie, Fei Huang 0002 |
EACL (1) | 6 |
| 2024 | Noise, Novels, Numbers. A Framework for Detecting and Categorizing Noise in Danish and Norwegian LiteratureabstractAli Al-Laith, Daniel Hershcovich, Jens Bjerring-Hansen, Jakob Ingemann Parby, Alexander Conroy, Timothy R Tangherlini. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Ali Allaith, Daniel Hershcovich, Jens Bjerring-Hansen, Jakob Parby, Alexander Conroy, Timothy Tangherlini |
EMNLP | 2 |
| 2024 | Bridging Cultures in the Kitchen: A Framework and Benchmark for Cross-Cultural Recipe RetrievalabstractThe cross-cultural adaptation of recipes is an important application of identifying and bridging cultural differences in language.The challenge lies in retaining the essence of the original recipe while also aligning with the writing and dietary habits of the target culture.Information Retrieval (IR) offers a way to address the challenge because it retrieves results from the culinary practices of the target culture while maintaining relevance to the original recipe.We introduce a novel task about cross-cultural recipe retrieval and present a unique Chinese-English cross-cultural recipe retrieval benchmark.Our benchmark is manually annotated under limited resource, utilizing various retrieval models to generate a pool of candidate results for manual annotation.The dataset provides retrieval samples that are culturally adapted but textually diverse, presenting greater challenges.We propose CARROT, a plug-and-play culturalaware recipe information retrieval framework that incorporates cultural-aware query rewriting and reranking methods and evaluate it both on our benchmark and intuitive human judgments.The results show that our framework significantly enhances the preservation of the original recipe and its cultural appropriateness for the target culture.We believe these insights will significantly contribute to future research on cultural adaptation. Maria Maistro, Daniel Hershcovich |
EMNLP | 3 |
| 2024 | FoodieQA: A Multimodal Dataset for Fine-Grained Understanding of Chinese Food CultureabstractWenyan Li, Crystina Zhang, Jiaang Li, Qiwei Peng, Raphael Tang, Li Zhou, Weijia Zhang, Guimin Hu, Yifei Yuan, Anders Søgaard, Daniel Hershcovich, Desmond Elliott. Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. 2024. Wenyan Li 0001, Xinyu Zhang 0018, Jiaang Li 0002, Qiwei Peng 0003, Raphael Tang, Li Zhou 0010, Weijia Zhang 0004, Guimin Hu, Yifei Yuan 0002, Anders Søgaard, Daniel Hershcovich, Desmond Elliott |
EMNLP | 11 |
| 2024 | Cultural Adaptation of RecipesabstractAbstract Building upon the considerable advances in Large Language Models (LLMs), we are now equipped to address more sophisticated tasks demanding a nuanced understanding of cross-cultural contexts. A key example is recipe adaptation, which goes beyond simple translation to include a grasp of ingredients, culinary techniques, and dietary preferences specific to a given culture. We introduce a new task involving the translation and cultural adaptation of recipes between Chinese- and English-speaking cuisines. To support this investigation, we present CulturalRecipes, a unique dataset composed of automatically paired recipes written in Mandarin Chinese and English. This dataset is further enriched with a human-written and curated test set. In this intricate task of cross-cultural recipe adaptation, we evaluate the performance of various methods, including GPT-4 and other LLMs, traditional machine translation, and information retrieval techniques. Our comprehensive analysis includes both automatic and human evaluation metrics. While GPT-4 exhibits impressive abilities in adapting Chinese recipes into English, it still lags behind human expertise when translating English recipes into Chinese. This underscores the multifaceted nature of cultural adaptations. We anticipate that these insights will significantly contribute to future research on culturally aware language models and their practical application in culturally diverse contexts. Yong Cao 0001, Yova Kementchedjhieva, Ruixiang Cui, Antonia Karamolegkou, Li Zhou 0010, Megan Dare, Lucia Donatelli, Daniel Hershcovich |
Trans. Assoc. Comput. Linguistics | 8 |
| 2024 | CreoleVal: Multilingual Multitask Benchmarks for CreolesabstractAbstract Creoles represent an under-explored and marginalized group of languages, with few available resources for NLP research. While the genealogical ties between Creoles and a number of highly resourced languages imply a significant potential for transfer learning, this potential is hampered due to this lack of annotated data. In this work we present CreoleVal, a collection of benchmark datasets spanning 8 different NLP tasks, covering up to 28 Creole languages; it is an aggregate of novel development datasets for reading comprehension relation classification, and machine translation for Creoles, in addition to a practical gateway to a handful of preexisting benchmarks. For each benchmark, we conduct baseline experiments in a zero-shot setting in order to further ascertain the capabilities and limitations of transfer learning for Creoles. Ultimately, we see CreoleVal as an opportunity to empower research on Creoles in NLP and computational linguistics, and in general, a step towards more equitable language technology around the globe. Heather C. Lent, Kushal Tatariya, Raj Dabre, Yiyi Chen 0002, Marcell Fekete, Esther Ploeger, Li Zhou 0010, Ruth-Ann Armstrong, Abee Eijansantos, Catriona Malau, Hans Erik Heje, Ernests Lavrinovics, Diptesh Kanojia, Paul Belony, Marcel Bollmann, Loïc Grobol, Miryam de Lhoneux, Daniel Hershcovich, Michel DeGraff, Anders Søgaard, Johannes Bjerva |
Trans. Assoc. Comput. Linguistics | 18 |
| 2023 | What does the Failure to Reason with "Respectively" in Zero/Few-Shot Settings Tell Us about Language Models?abstractHumans can effortlessly understand the coordinate structure of sentences such as “Niels Bohr and Kurt Cobain were born in Copenhagen and Seattle, respectively”. In the context of natural language inference (NLI), we examine how language models (LMs) reason with respective readings (Gawron and Kehler, 2004) from two perspectives: syntactic-semantic and commonsense-world knowledge. We propose a controlled synthetic dataset WikiResNLI and a naturally occurring dataset NatResNLI to encompass various explicit and implicit realizations of “respectively”. We show that fine-tuned NLI models struggle with understanding such readings without explicit supervision. While few-shot learning is easy in the presence of explicit cues, longer training is required when the reading is evoked implicitly, leaving models to rely on common sense inferences. Furthermore, our fine-grained analysis indicates models fail to generalize across different constructions. To conclude, we demonstrate that LMs still lag behind humans in generalizing to the long tail of linguistic constructions. Ruixiang Cui, Seolhwa Lee, Daniel Hershcovich, Anders Søgaard |
ACL (1) | 3 |
| 2023 | What's the Meaning of Superhuman Performance in Today's NLU?abstractSimone Tedeschi, Johan Bos, Thierry Declerck, Jan Hajič, Daniel Hershcovich, Eduard Hovy, Alexander Koller, Simon Krek, Steven Schockaert, Rico Sennrich, Ekaterina Shutova, Roberto Navigli. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023. Simone Tedeschi, Johan Bos, Thierry Declerck, Jan Hajic 0001, Daniel Hershcovich, Eduard H. Hovy, Alexander Koller, Simon Krek, Steven Schockaert, Rico Sennrich, Ekaterina Shutova, Roberto Navigli |
ACL (1) | 5 |
| 2023 | On Evaluating Multilingual Compositional Generalization with Translated DatasetsabstractCompositional generalization allows efficient learning and human-like inductive biases.Since most research investigating compositional generalization in NLP is done on English, important questions remain underexplored.Do the necessary compositional generalization abilities differ across languages?Can models compositionally generalize crosslingually?As a first step to answering these questions, recent work used neural machine translation to translate datasets for evaluating compositional generalization in semantic parsing.However, we show that this entails critical semantic distortion.To address this limitation, we craft a faithful rule-based translation of the MCWQ dataset (Cui et al., 2022) from English to Chinese and Japanese.Even with the resulting robust benchmark, which we call MCWQ-R, we show that the distribution of compositions still suffers due to linguistic divergences, and that multilingual models still struggle with cross-lingual compositional generalization.Our dataset and methodology will be useful resources for the study of cross-lingual compositional generalization in other tasks. 1 Daniel Hershcovich |
ACL (1) | 2 |
| 2023 | A Two-Sided Discussion of Preregistration of NLP ResearchabstractVan Miltenburg et al. (2021) suggest NLP research should adopt preregistration to prevent fishing expeditions and to promote publication of negative results.At face value, this is a very reasonable suggestion, seemingly solving many methodological problems with NLP research.We discuss pros and cons-some old, some new: a) Preregistration is challenged by the practice of retrieving hypotheses after the results are known; b) preregistration may bias NLP toward confirmatory research; c) preregistration must allow for reclassification of research as exploratory; d) preregistration may increase publication bias; e) preregistration may increase flag-planting; f) preregistration may increase p-hacking; and finally, g) preregistration may make us less risk tolerant.We cast our discussion as a dialogue, presenting both sides of the debate. Anders Søgaard, Daniel Hershcovich, Miryam de Lhoneux |
EACL | 2 |
| 2022 | Challenges and Strategies in Cross-Cultural NLPabstractDaniel Hershcovich, Stella Frank, Heather Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Piqueras, Ilias Chalkidis, Ruixiang Cui, Constanza Fierro, Katerina Margatina, Phillip Rust, Anders Søgaard. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Daniel Hershcovich, Stella Frank, Heather C. Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Piqueras, Ilias Chalkidis, Ruixiang Cui, Constanza Fierro, Aikaterini Margatina, Phillip Rust, Anders Søgaard |
ACL (1) | 1 |
| 2022 | Scaling Creative Inspiration with Fine-Grained Functional Aspects of IdeasabstractLarge repositories of products, patents and scientific papers offer an opportunity for building systems that scour millions of ideas and help users discover inspirations. However, idea descriptions are typically in the form of unstructured text, lacking key structure that is required for supporting creative innovation interactions. Prior work has explored idea representations that were either limited in expressivity, required significant manual effort from users, or dependent on curated knowledge bases with poor coverage. We explore a novel representation that automatically breaks up products into fine-grained functional aspects capturing the purposes and mechanisms of ideas, and use it to support important creative innovation interactions: functional search for ideas, and exploration of the design space around a focal problem by viewing related problem perspectives pooled from across many products. In user studies, our approach boosts the quality of creative search and inspirations, substantially outperforming strong baselines by 50-60%. Tom Hope, Ronen Tamari, Daniel Hershcovich, Hyeonsu B. Kang, Joel Chan, Aniket Kittur, Dafna Shahaf |
CHI | 3 |
| 2022 | Towards Climate Awareness in NLP ResearchabstractThe climate impact of AI, and NLP research in particular, has become a serious issue given the enormous amount of energy that is increasingly being used for training and running computational models.Consequently, increasing focus is placed on efficient NLP.However, this important initiative lacks simple guidelines that would allow for systematic climate reporting of NLP research.We argue that this deficiency is one of the reasons why very few publications in NLP report key figures that would allow a more thorough examination of environmental impact, and present a quantitative survey to demonstrate this.As a remedy, we propose a climate performance model card with the primary purpose of being practically usable with only limited information about experiments and the underlying computer hardware.We describe why this step is essential to increase awareness about the environmental impact of NLP research and, thereby, paving the way for more thorough discussions.1 Daniel Hershcovich, Nicolas Webersinke, Mathias Kraus, Julia Anna Bingler, Markus Leippold |
EMNLP | 1 |
| 2022 | Evaluating Deep Taylor Decomposition for Reliability Assessment in the Wild
Stephanie Brandl, Daniel Hershcovich, Anders Søgaard |
ICWSM | 2 |
| 2022 | Generalized Quantifiers as a Source of Error in Multilingual NLU BenchmarksabstractLogical approaches to representing language have developed and evaluated computational models of quantifier words since the 19th century, but today's NLU models still struggle to capture their semantics.We rely on Generalized Quantifier Theory for languageindependent representations of the semantics of quantifier words, to quantify their contribution to the errors of NLU models.We find that quantifiers are pervasive in NLU benchmarks, and their occurrence at test time is associated with performance drops.Multilingual models also exhibit unsatisfying quantifier reasoning abilities, but not necessarily worse for non-English languages.To facilitate directlytargeted probing, we present an adversarial generalized quantifier NLI task (GQNLI) and show that pre-trained language models have a clear lack of robustness in generalized quantifier reasoning. Ruixiang Cui, Daniel Hershcovich, Anders Søgaard |
NAACL-HLT | 2 |
| 2022 | Compositional Generalization in Multilingual Semantic Parsing over WikidataabstractAbstract Semantic parsing (SP) allows humans to leverage vast knowledge resources through natural interaction. However, parsers are mostly designed for and evaluated on English resources, such as CFQ (Keysers et al., 2020), the current standard benchmark based on English data generated from grammar rules and oriented towards Freebase, an outdated knowledge base. We propose a method for creating a multilingual, parallel dataset of question-query pairs, grounded in Wikidata. We introduce such a dataset, which we call Multilingual Compositional Wikidata Questions (MCWQ), and use it to analyze the compositional generalization of semantic parsers in Hebrew, Kannada, Chinese, and English. While within- language generalization is comparable across languages, experiments on zero-shot cross- lingual transfer demonstrate that cross-lingual compositional generalization fails, even with state-of-the-art pretrained multilingual encoders. Furthermore, our methodology, dataset, and results will facilitate future research on SP in more realistic and diverse settings than has been possible with existing resources. Ruixiang Cui, Rahul Aralikatte, Heather C. Lent, Daniel Hershcovich |
Trans. Assoc. Comput. Linguistics | 4 |
| 2021 | Joint Semantic Analysis with Document-Level Cross-Task Coherence RewardsabstractCoreference resolution and semantic role labeling are NLP tasks that capture different aspects of semantics, indicating respectively, which expressions refer to the same entity, and what semantic roles expressions serve in the sentence. However, they are often closely interdependent, and both generally necessitate natural language understanding. Do they form a coherent abstract representation of documents? We present a neural network architecture for joint coreference resolution and semantic role labeling for English, and train graph neural networks to model the 'coherence' of the combined shallow semantic graph. Using the resulting coherence score as a reward for our joint semantic analyzer, we use reinforcement learning to encourage global coherence over the document and between semantic annotations. This leads to improvements on both tasks in multiple datasets from different domains, and across a range of encoders of different expressivity, calling, we believe, for a more holistic approach for semantics in NLP. Rahul Aralikatte, Mostafa Abdou, Heather C. Lent, Daniel Hershcovich, Anders Søgaard |
AAAI | 4 |
| 2021 | Can Language Models Encode Perceptual Structure Without Grounding? A Case Study in ColorabstractPretrained language models have been shown to encode relational information, such as the relations between entities or concepts in knowledge-bases -(Paris, Capital, France).However, simple relations of this type can often be recovered heuristically and the extent to which models implicitly reflect topological structure that is grounded in world, such as perceptual structure, is unknown.To explore this question, we conduct a thorough case study on color.Namely, we employ a dataset of monolexemic color terms and color chips represented in CIELAB, a color space with a perceptually meaningful distance metric.Using two methods of evaluating the structural alignment of colors in this space with textderived color term representations, we find significant correspondence.Analyzing the differences in alignment across the color spectrum, we find that warmer colors are, on average, better aligned to the perceptual color space than cooler ones, suggesting an intriguing connection to findings from recent work on efficient communication in color naming.Further analysis suggests that differences in alignment are, in part, mediated by collocationality and differences in syntactic usage, posing questions as to the relationship between color perception and usage and context. Mostafa Abdou, Artur Kulmizev, Daniel Hershcovich, Stella Frank, Ellie Pavlick, Anders Søgaard |
CoNLL | 3 |
| 2021 | A Multilingual Benchmark for Probing Negation-Awareness with Minimal PairsabstractMareike Hartmann, Miryam de Lhoneux, Daniel Hershcovich, Yova Kementchedjhieva, Lukas Nielsen, Chen Qiu, Anders Søgaard. Proceedings of the 25th Conference on Computational Natural Language Learning. 2021. Mareike Hartmann, Miryam de Lhoneux, Daniel Hershcovich, Yova Kementchedjhieva, Lukas Nielsen, Chen Qiu 0005, Anders Søgaard |
CoNLL | 3 |
| 2020 | Comparison by Conversion: Reverse-Engineering UCCA from Syntax and Lexical SemanticsabstractBuilding robust natural language understanding systems will require a clear characterization of whether and how various linguistic meaning representations complement each other.To perform a systematic comparative analysis, we evaluate the mapping between meaning representations from different frameworks using two complementary methods: (i) a rule-based converter, and (ii) a supervised delexicalized parser that parses to one framework using only information from the other as features.We apply these methods to convert the STREUSLE corpus (with syntactic and lexical semantic annotations) to UCCA (a graph-structured full-sentence meaning representation).Both methods yield surprisingly accurate target representations, close to fully supervised UCCA parser quality-indicating that UCCA annotations are partially redundant with STREUSLE annotations.Despite this substantial convergence between frameworks, we find several important areas of divergence. Daniel Hershcovich, Nathan Schneider 0001, Dotan Dvir, Jakob Prange, Miryam de Lhoneux, Omri Abend |
COLING | 1 |
| 2019 | Argument Invention from First PrinciplesabstractYonatan Bilu, Ariel Gera, Daniel Hershcovich, Benjamin Sznajder, Dan Lahav, Guy Moshkowich, Anael Malet, Assaf Gavron, Noam Slonim. Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. 2019. Yonatan Bilu, Ariel Gera, Daniel Hershcovich, Benjamin Sznajder, Dan Lahav, Guy Moshkowich, Anael Malet, Assaf Gavron, Noam Slonim |
ACL (1) | 3 |
| 2019 | The Language of Legal and Illegal Activity on the DarknetabstractThe non-indexed parts of the Internet (the Darknet) have become a haven for both legal and illegal anonymous activity.Given the magnitude of these networks, scalably monitoring their activity necessarily relies on automated tools, and notably on NLP tools.However, little is known about what characteristics texts communicated through the Darknet have, and how well off-the-shelf NLP tools do on this domain.This paper tackles this gap and performs an in-depth investigation of the characteristics of legal and illegal text in the Darknet, comparing it to a clear net website with similar content as a control condition.Taking drug-related websites as a test case, we find that texts for selling legal and illegal drugs have several linguistic characteristics that distinguish them from one another, as well as from the control condition, among them the distribution of POS tags, and the coverage of their named entities in Wikipedia. 1 Leshem Choshen, Dan Eldad, Daniel Hershcovich, Elior Sulem, Omri Abend |
ACL (1) | 3 |
| 2019 | Rewarding Coreference Resolvers for Being Consistent with World KnowledgeabstractRahul Aralikatte, Heather Lent, Ana Valeria Gonzalez, Daniel Herschcovich, Chen Qiu, Anders Sandholm, Michael Ringaard, Anders Søgaard. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Rahul Aralikatte, Heather C. Lent, Ana Valeria González-Garduño, Daniel Hershcovich, Chen Qiu 0005, Anders Sandholm 0001, Michael Ringaard, Anders Søgaard |
EMNLP/IJCNLP (1) | 4 |
| 2018 | Multitask Parsing Across Semantic RepresentationsabstractThe ability to consolidate information of different types is at the core of intelligence, and has tremendous practical value in allowing learning for one task to benefit from generalizations learned for others.In this paper we tackle the challenging task of improving semantic parsing performance, taking UCCA parsing as a test case, and AMR, SDP and Universal Dependencies (UD) parsing as auxiliary tasks.We experiment on three languages, using a uniform transition-based system and learning architecture for all parsing tasks.Despite notable conceptual, formal and domain differences, we show that multitask learning significantly improves UCCA parsing in both in-domain and out-of-domain settings.Our code is publicly available. Daniel Hershcovich, Omri Abend, Ari Rappoport |
ACL (1) | 1 |
| 2017 | A Transition-Based Directed Acyclic Graph Parser for UCCAabstractWe present the first parser for UCCA, a cross-linguistically applicable framework for semantic representation, which builds on extensive typological work and supports rapid annotation.UCCA poses a challenge for existing parsing techniques, as it exhibits reentrancy (resulting in DAG structures), discontinuous structures and non-terminal nodes corresponding to complex semantic units.To our knowledge, the conjunction of these formal properties is not supported by any existing parser.Our transition-based parser, which uses a novel transition set and features based on bidirectional LSTMs, has value not just for UCCA parsing: its ability to handle more general graph structures can inform the development of parsers for other semantic DAG structures, and in languages that frequently use discontinuous structures. Daniel Hershcovich, Omri Abend, Ari Rappoport |
ACL (1) | 1 |
| 2014 | Context Dependent Claim Detection
Ran Levy 0001, Yonatan Bilu, Daniel Hershcovich, Ehud Aharoni, Noam Slonim |
COLING | 3 |
| 2014 | Verification of Transactional Memory in POWER8abstractTransactional memory is a promising mechanism for synchronizing concurrent programs that eliminates locks at the expense of hardware complexity. Transactional memory is a hard feature to verify. First, transactions comprise several instructions that must be observed as a single global atomic operation. In addition, there are many reasons a transaction can fail. This results in a high level of non-determinism which must be tamed by the verification methodology. This paper describes the innovation that was applied to tools and methodology in pre-silicon simulation, acceleration and post-silicon in order to verify transactional memory in the IBM POWER8 processor core. Allon Adir, Dave Goodman, Daniel Hershcovich, Oz Hershkovitz, Bryan G. Hickerson, Karen Holtz, Wisam Kadry, Anatoly Koyfman, John M. Ludden, Charles Meissner, Amir Nahir, Randall R. Pratt, Mike Schiffli, Brett St. Onge, Brian W. Thompto, Elena Tsanko, Avi Ziv |
DAC | 3 |