VLDB 2026 Research / reviewers in the wild / expert
Shay B. Cohen
dblp:04/5629
· DBLP profile ↗
98ranked-venue papers
24as first author
30since 2021 · last 2026
0000-0003-4753-8353ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 93 · 22 first-author · 27 since 2021Databases, data management, data science and information retrieval · 3 · 3 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 1 first-authorSoftware engineering, systems software and programming languages · 2 · 2 first-authorTheory of computation · 2 · 1 first-author · 1 since 2021Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Thinking in Schemas: Robust Syllogistic Reasoning in LLMsabstractLLMs often mistake what sounds true for what is formally valid.This limitation is especially evident in syllogistic reasoning, where plausible arguments can lead models to endorse conclusions that are logically invalid, a phenomenon known as content effect (CE).We present Boethius, a schema-guided framework for syllogistic reasoning that disentangles semantic plausibility from logical validity.Boethius adopts an auditable, quasi-formal reasoning process with two complementary stages: a Schema Module, which deduces the underlying logical form by analysing the formal structure of the premises, and an Instantiation Module, which instantiates this form over the concrete argument and evaluates validity independently of content-level semantics.Our results show that Boethius consistently outperforms existing approaches, improving syllogistic reasoning accuracy while substantially reducing CE.These gains hold for both large models in a pure in-context learning setting and smaller models trained via schema-guided trajectories using supervised fine-tuning and optimisation-based refinement."Some winged animals are not chickadees.It is certain that all chickadees are birds.Therefore, some winged animals are not birds." SCHEMA MODULE INSTANTIATION MODULE PREMISES IDENTIFICATION:P1: Some winged animals are not chickadees.P2: It is certain that all chickadees are birds. Federico Ranaldi, Leonardo Ranaldi, Fabio Massimo Zanzotto, Shay B. Cohen |
ACL (1) | 4 |
| 2025 | Theorem Prover as a Judge for Synthetic Data GenerationabstractThe demand for synthetic data in mathematical reasoning has increased due to its potential to enhance the mathematical capabilities of large language models (LLMs).However, ensuring the validity of intermediate reasoning steps remains a significant challenge, affecting data quality.While formal verification via theorem provers effectively validates LLM reasoning, the autoformalisation of mathematical proofs remains error-prone.We introduce iterative autoformalisation, an approach that iteratively refines theorem prover formalisation to mitigate errors, thereby increasing the execution rate on the Lean prover from 60% to 87%.Building upon that, we introduce Theorem Prover as a Judge (TP-as-a-Judge), a method that makes use of theorem prover formalisation to rigorously assess LLM intermediate reasoning, effectively integrating autoformalisation with synthetic data generation.Finally, we present Reinforcement Learning from Theorem Prover Feedback (RLTPF), a framework that replaces human annotation with theorem prover feedback in Reinforcement Learning from Human Feedback (RLHF).Across multiple LLMs, applying TP-as-a-Judge and RLTPF improves benchmarks with only 3,508 samples, achieving 5.56% accuracy gain on Mistral-7B for MultiArith, 6.00% on Llama-2-7B for SVAMP, and 3.55% on Llama-3.1-8B for AQUA. 1 Joshua Ong Jun Leang, Giwon Hong, Shay B. Cohen |
ACL (1) | 4 |
| 2025 | People Attribute Purpose to Autonomous Vehicles When Explaining Their Behavior: Insights from Cognitive Science for Explainable AI
Balint Gyevnar, Stephanie Droop, Tadeg Quillien, Shay B. Cohen, Neil Bramley, Christopher G. Lucas, Stefano V. Albrecht |
CHI | 4 |
| 2025 | CoMAT: Chain of Mathematically Annotated Thought Improves Mathematical ReasoningabstractMathematical reasoning remains a significant challenge for large language models (LLMs), despite progress in prompting techniques such as Chain-of-Thought (CoT).We present Chain of Mathematically Annotated Thought (Co-MAT), which enhances reasoning through two stages: Symbolic Conversion (converting natural language queries into symbolic form) and Reasoning Execution (deriving answers from symbolic representations).CoMAT operates entirely with a single LLM and without external solvers.Across four LLMs, CoMAT outperforms traditional CoT on six out of seven benchmarks, achieving gains of 4.48% on MMLU-Redux (MATH) and 4.58% on GaoKao MCQ.In addition to improved performance, CoMAT ensures faithfulness and verifiability, offering a transparent reasoning process for complex mathematical tasks 1 . Joshua Ong Jun Leang, Aryo Pradipta Gema, Shay B. Cohen |
EMNLP | 3 |
| 2025 | Iterative Multilingual Spectral Attribute ErasureabstractMultilingual representations embed words with similar meanings to share a common semantic space across languages, creating opportunities to transfer debiasing effects between languages.However, existing methods for debiasing are unable to exploit this opportunity because they operate on individual languages.We present Iterative Multilingual Spectral Attribute Erasure (IMSAE), which identifies and mitigates joint bias subspaces across multiple languages through iterative SVD-based truncation.Evaluating IMSAE across eight languages and five demographic dimensions, we demonstrate its effectiveness in both standard and zero-shot settings, where target language data is unavailable, but linguistically similar languages can be used for debiasing.Our comprehensive experiments across diverse language models (BERT, Llama, Mistral) show that IMSAE outperforms traditional monolingual and cross-lingual approaches while maintaining model utility. 1 Shun Shao, Yftah Ziser, Zheng Zhao 0005, Yifu Qiu, Shay B. Cohen, Anna Korhonen |
EMNLP | 5 |
| 2025 | DEPfold: RNA Secondary Structure Prediction as Dependency ParsingabstractRNA secondary structure prediction is critical for understanding RNA function
but remains challenging due to complex structural elements like pseudoknots and
limited training data. We introduce DEPfold, a novel deep learning approach that
re-frames RNA secondary structure prediction as a dependency parsing problem.
DEPfold presents three key innovations: (1) a biologically motivated transformation of RNA structures into labeled dependency trees, (2) a biaffine attention
mechanism for joint prediction of base pairings and their types, and (3) an optimal
tree decoding algorithm that enforces valid RNA structural constraints. Unlike traditional energy-based methods, DEPfold learns directly from annotated data and
leverages pretrained language models to predict RNA structure. We evaluate DEPfold on both within-family and cross-family RNA datasets, demonstrating significant performance improvements over existing methods. DEPfold shows strong
performance in cross-family generalization when trained on data augmented by
traditional energy-based models, outperforming existing methods on the bpRNAnew dataset. This demonstrates DEPfold’s ability to effectively learn structural
information beyond what traditional methods capture. Our approach bridges natural language processing (NLP) with RNA biology, providing a computationally
efficient and adaptable tool for advancing RNA structure prediction and analysis Shay B. Cohen |
ICLR | 2 |
| 2025 | PoisonBench: Assessing Language Model Vulnerability to Poisoned Preference DataabstractPreference learning is a central component for aligning current LLMs, but this process can be vulnerable to data poisoning attacks. To address this concern, we introduce PoisonBench, a benchmark for evaluating large language models' susceptibility to data poisoning during preference learning. Data poisoning attacks can manipulate large language model responses to include hidden malicious content or biases, potentially causing the model to generate harmful or unintended outputs while appearing to function normally. We deploy two distinct attack types across eight realistic scenarios, assessing 22 widely-used models. Our findings reveal concerning trends: (1) Scaling up parameter size does not always enhance resilience against poisoning attacks and the influence on model resilience varies among different model suites. (2) There exists a log-linear relationship between the effects of the attack and the data poison ratio; (3) The effect of data poisoning can generalize to extrapolated triggers that are not included in the poisoned data.
These results expose weaknesses in current preference learning techniques, highlighting the urgent need for more robust defenses against malicious models and data manipulation. Tingchen Fu, Mrinank Sharma, Philip Torr 0001, Shay B. Cohen, David Krueger 0001, Fazl Barez |
ICML | 4 |
| 2025 | TSPRank: Bridging Pairwise and Listwise Methods with a Bilinear Travelling Salesman ModelabstractTraditional Learning-To-Rank (LETOR) approaches, including pairwise methods like RankNet and LambdaMART, often fall short by solely focusing on pairwise comparisons, leading to sub-optimal global rankings. Conversely, deep learning based listwise methods, while aiming to optimise entire lists, require complex tuning and yield only marginal improvements over robust pairwise models. To overcome these limitations, we introduce Travelling Salesman Problem Rank (TSPRank), a hybrid pairwise-listwise ranking method. TSPRank reframes the ranking problem as a Travelling Salesman Problem (TSP), a well-known combinatorial optimisation challenge that has been extensively studied for its numerous solution algorithms and applications. This approach enables the modelling of pairwise relationships and leverages combinatorial optimisation to determine the listwise ranking. TSPRank can be directly integrated as an additional component into embeddings generated by existing backbone models to enhance ranking performance. Our extensive experiments across three backbone models on diverse tasks, including stock ranking, information retrieval, and historical events ordering, demonstrate that TSPRank significantly outperforms both pure pairwise and listwise methods. Our qualitative analysis reveals that TSPRank's main advantage over existing methods is its ability to harness global information better while ranking. TSPRank's robustness and superior performance across different domains highlight its potential as a versatile and effective LETOR solution. Weixian Waylon Li, Yftah Ziser, Yifei Xie 0002, Shay B. Cohen, Tiejun Ma |
KDD (1) | 4 |
| 2025 | Pre-training Time Series Models with Stock Data CustomizationabstractStock selection, which aims to predict stock prices and identify the most profitable ones, is a crucial task in finance. While existing methods primarily focus on developing model structures and building graphs for improved selection, pre-training strategies remain underexplored in this domain. Current stock series pre-training follows methods from other areas without adapting to the unique characteristics of financial data, particularly overlooking stock-specific contextual information and the non-stationary nature of stock prices. Consequently, the latent statistical features inherent in stock data are underutilized. In this paper, we propose three novel pre-training tasks tailored to stock data characteristics: stock code classification, stock sector classification, and moving average prediction. We develop the Stock Specialized Pre-trained Transformer (SSPT) based on a two-layer transformer architecture. Extensive experimental results validate the effectiveness of our pre-training methods and provide detailed guidance on their application. Evaluations on five stock datasets, including four markets and two time periods, demonstrate that SSPT consistently outperforms the market and existing methods in terms of both cumulative investment return ratio and Sharpe ratio. Additionally, our experiments on simulated data investigate the underlying mechanisms of our methods, providing insights into understanding price series. Our code is publicly available at: https://github.com/astudentuser/Pre-training-Time-Series-Models-with-Stock-Data-Customization. Tiejun Ma, Shay B. Cohen |
KDD (2) | 3 |
| 2024 | Layer by Layer: Uncovering Where Multi-Task Learning Happens in Instruction-Tuned Large Language ModelsabstractFine-tuning pre-trained large language models (LLMs) on a diverse array of tasks has become a common approach for building models that can solve various natural language processing (NLP) tasks. However, where and to what extent these models retain task-specific knowledge remains largely unexplored. This study investigates the task-specific information encoded in pre-trained LLMs and the effects of instruction tuning on their representations across a diverse set of over 60 NLP tasks. We use a set of matrix analysis tools to examine the differences between the way pre-trained and instruction-tuned LLMs store task-specific information. Our findings reveal that while some tasks are already encoded within the pre-trained LLMs, others greatly benefit from instruction tuning. Additionally, we pinpointed the layers in which the model transitions from high-level general representations to more task-oriented representations. This finding extends our understanding of the governing mechanisms of LLMs and facilitates future research in the fields of parameter-efficient transfer learning and multi-task learning. Our code is available at: https://github.com/zsquaredz/layerbylayer/ Zheng Zhao 0005, Yftah Ziser, Shay B. Cohen |
EMNLP | 3 |
| 2024 | Interpreting Context Look-ups in Transformers: Investigating Attention-MLP InteractionsabstractUnderstanding the inner workings of large language models (LLMs) is crucial for advancing their theoretical foundations and real-world applications.While the attention mechanism and multi-layer perceptrons (MLPs) have been studied independently, their interactions remain largely unexplored.This study investigates how attention heads and next-token neurons interact in LLMs to predict new words.We propose a methodology to identify next-token neurons, find prompts that highly activate them, and determine the upstream attention heads responsible.We then generate and evaluate explanations for the activity of these attention heads in an automated manner.Our findings reveal that some attention heads recognize specific contexts relevant to predicting a token and activate a downstream token-predicting neuron accordingly.This mechanism provides a deeper understanding of how attention heads work with MLP neurons to perform next-token prediction.Our approach offers a foundation for further research into the intricate workings of LLMs and their impact on text generation and understanding. Clement Neo, Shay B. Cohen, Fazl Barez |
EMNLP | 2 |
| 2024 | LeanReasoner: Boosting Complex Logical Reasoning with LeanabstractDongwei Jiang, Marcio Fonseca, Shay Cohen. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Dongwei Jiang, Marcio Fonseca, Shay B. Cohen |
NAACL-HLT | 3 |
| 2024 | Are Large Language Model Temporally Grounded?abstractYifu Qiu, Zheng Zhao, Yftah Ziser, Anna Korhonen, Edoardo Ponti, Shay Cohen. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2024. Yifu Qiu, Zheng Zhao 0005, Yftah Ziser, Anna Korhonen, Edoardo Maria Ponti, Shay B. Cohen |
NAACL-HLT | 6 |
| 2024 | einspace: Searching for Neural Architectures from Fundamental OperationsabstractNeural architecture search (NAS) finds high performing networks for a given task. Yet the results of NAS are fairly prosaic; they did not e.g. create a shift from convolutional structures to transformers. This is not least because the search spaces in NAS often aren’t diverse enough to include such transformations *a priori*. Instead, for NAS to provide greater potential for fundamental design shifts, we need a novel expressive search space design which is built from more fundamental operations. To this end, we introduce `einspace`, a search space based on a parameterised probabilistic context-free grammar. Our space is versatile, supporting architectures of various sizes and complexities, while also containing diverse network operations which allow it to model convolutions, attention components and more. It contains many existing competitive architectures, and provides flexibility for discovering new ones. Using this search space, we perform experiments to find novel architectures as well as improvements on existing ones on the diverse Unseen NAS datasets. We show that competitive architectures can be obtained by searching from scratch, and we consistently find large improvements when initialising the search with strong baselines. We believe that this work is an important advancement towards a transformative NAS paradigm where search space expressivity and strategic search initialisation play key roles. Linus Ericsson, Miguel Espinosa, Chenhongyi Yang, Antreas Antoniou, Amos J. Storkey, Shay B. Cohen, Elliot Crowley |
NeurIPS | 6 |
| 2024 | Spectral Editing of Activations for Large Language Model AlignmentabstractLarge language models (LLMs) often exhibit undesirable behaviours, such as generating untruthful or biased content. Editing their internal representations has been shown to be effective in mitigating such behaviours on top of the existing alignment methods. We propose a novel inference-time editing method, namely spectral editing of activations (SEA), to project the input representations into directions with maximal covariance with the positive demonstrations (e.g., truthful) while minimising covariance with the negative demonstrations (e.g., hallucinated). We also extend our method to non-linear editing using feature functions. We run extensive experiments on benchmarks concerning truthfulness and bias with six open-source LLMs of different sizes and model families. The results demonstrate the superiority of SEA in effectiveness, generalisation to similar tasks, as well as computation and data efficiency. We also show that SEA editing only has a limited negative impact on other model capabilities. Yifu Qiu, Zheng Zhao 0005, Yftah Ziser, Anna Korhonen, Edoardo Maria Ponti, Shay B. Cohen |
NeurIPS | 6 |
| 2024 | CivilSum: A Dataset for Abstractive Summarization of Indian Court DecisionsabstractExtracting relevant information from legal documents is a challenging task due to the technical complexity and volume of their content. These factors also increase the costs of annotating large datasets, which are required to train state-of-the-art summarization systems. To address these challenges, we introduce CivilSum, a collection of 23,350 legal case decisions from the Supreme Court of India and other Indian High Courts paired with human-written summaries. Compared to previous datasets such as IN-Abs, CivilSum not only has more legal decisions but also provides shorter and more abstractive summaries, thus offering a challenging benchmark for legal summarization. Unlike other domains such as news articles, our analysis shows the most important content tends to appear at the end of the documents. We measure the effect of this tail bias on summarization performance using strong architectures for long-document abstractive summarization, and the results highlight the importance of long sequence modeling for the proposed task. CivilSum and related code are publicly available to the research community to advance text summarization in the legal domain. Manuj Malik, Zheng Zhao 0005, Marcio Fonseca, Shrisha Rao 0001, Shay B. Cohen |
SIGIR | 5 |
| 2024 | On the Trade-off between Redundancy and Cohesiveness in Extractive SummarizationabstractExtractive summaries are usually presented as lists of sentences with no expected cohesion between them and with plenty of redundant information if not accounted for. In this paper, we investigate the trade-offs incurred when aiming to control for inter-sentential cohesion and redundancy in extracted summaries, and their impact on their informativeness. As case study, we focus on the summarization of long, highly redundant documents and consider two optimization scenarios, reward-guided and with no supervision. In the reward-guided scenario, we compare systems that control for redundancy and cohesiveness during sentence scoring. In the unsupervised scenario, we introduce two systems that aim to control all three properties --informativeness, redundancy, and cohesiveness-- in a principled way. Both systems implement a psycholinguistic theory that simulates how humans keep track of relevant content units and how cohesiveness and non-redundancy constraints are applied in short-term memory during reading. Extensive automatic and human evaluations reveal that systems optimizing for --among other properties-- cohesiveness are capable of better organizing content in summaries compared to systems that optimize only for redundancy, while maintaining comparable informativeness. We find that the proposed unsupervised systems manage to extract highly cohesive summaries across varying levels of document redundancy, although sacrificing informativeness in the process. Finally, we lay evidence as to how simulated cognitive processes impact the trade-off between the analysed summary properties. Ronald Cardenas, Matthias Shen, Shay B. Cohen |
J. Artif. Intell. Res. | 3 |
| 2023 | BERT Is Not The Count: Learning to Match Mathematical Statements with ProofsabstractWe introduce a task consisting in matching a proof to a given mathematical statement.The task fits well within current research on Mathematical Information Retrieval and, more generally, mathematical article analysis (Mathematical Sciences, 2014).We present a dataset for the task (the MATCH dataset) consisting of over 180k statement-proof pairs extracted from modern mathematical research articles.1 We find this dataset highly representative of our task, as it consists of relatively new findings useful to mathematicians.We propose a bilinear similarity model and two decoding methods to match statements to proofs effectively.While the first decoding method matches a proof to a statement without being aware of other statements or proofs, the second method treats the task as a global matching problem.Through a symbol replacement procedure, we analyze the "insights" that pre-trained language models have in such mathematical article analysis and show that while these models perform well on this task with the best performing mean reciprocal rank of 73.7, they follow a relatively shallow symbolic analysis and matching to achieve that performance.2 * Work mostly done at the University of Edinburgh. 1 Our dataset and code are available at https:// github.com/waylonli/MATcH.2 Like Bert, The Count (or Count von Count; ) is a character from the television show Sesame Street.The Count likes counting, and his main role in the show is to teach this skill to children. Weixian Waylon Li, Yftah Ziser, Maximin Coavoux, Shay B. Cohen |
EACL | 4 |
| 2023 | Gold Doesn't Always Glitter: Spectral Removal of Linear and Nonlinear Guarded Attribute InformationabstractWe describe a simple and effective method (Spectral Attribute removaL; SAL) to remove private or guarded information from neural representations.Our method uses matrix decomposition to project the input representations into directions with reduced covariance with the guarded information rather than maximal covariance as factorization methods normally use.We begin with linear information removal and proceed to generalize our algorithm to the case of nonlinear information removal using kernels.Our experiments demonstrate that our algorithm retains better main task performance after removing the guarded information compared to previous work.In addition, our experiments demonstrate that we need a relatively small amount of guarded attribute data to remove information about these attributes, which lowers the exposure to sensitive data and is more suitable for low-resource scenarios.1 Shun Shao, Yftah Ziser, Shay B. Cohen |
EACL | 3 |
| 2023 | AMR Parsing is Far from Solved: GrAPES, the Granular AMR Parsing Evaluation SuiteabstractMessage from the General Chair I am happy to welcome you to EMNLP-2023 in Singapore!Like EMNLP-2021, EMNLP-2022, and other ACL-related meetings, we decided to host EMNLP-2023 as another hybrid conference having both in-person and virtual presentations and participants.We are not sure how long this style of our meetings will last.However, we have already accustomed to this style of conferences, which has its own advantages, while it causes a heavy burden to those organizing such events.The past one year has been a terrific and thrilling year since the advent of ChatGPT and other Large Language Models.Any people having access to those models has posed a big impact on people's impression about AI and has started to give them a feeling of fear.People now can do not only natural conversation with AI but also conduct various natural language tasks using our own languages.We now know it is difficult to guarantee that Large Language Models produce honest and harmless outputs.We have found that good prompting, demonstrations and complex ones like the Chain of Thought prompting draw out or enhance the emergent abilities of Large Language Models.However, we still don't know precisely why and how such in-context learning works.This year's EMNLP highlights a theme track, "Large Language Models and the Future of NLP."I hope we can see enthusiastic discussions and innovative ideas will be presented in EMNLP-2023.One big trial is that the Program Chairs decided to use OpenReview as the cradle of the main conference papers, for making reviews and author responses publicly available.The motivation and effects of this trial will be explained by the PC Chairs.Another important trial is to rent out the Universal Studio Singapore for our Social Event.I hope everyone will enjoy this event.EMNLP-2023 is the biggest conference ever in the SIGDAT history.Organizing such a big event is very difficult.As the General Chair, the most important and difficult task is to organize all the committees by a group of enthusiastic and talented people.I was very fortunate to be able to collect great committee members.Without such a wonderful group of colleagues, it almost has been impossible to make this great event happen.I would like to send my sincere thanks to all the members of our organization teams.Here, I only list the chairs by names, but I also like to send gratitude from my heart to all the people involved in EMNLP-2023, including keynote speakers, panelists, workshop organizers, tutorial tutors, senior area chairs, area chairs, reviewers, volunteers, sponsors, the Underline team, and all of you attending EMNLP-2023 in-person or virtually.• The program chairs -Houda Bouamor, Juan Pino, and Kalika Bali -who made a number of innovations and handled a huge number of submitted papers.I cannot help but be grateful for their tireless work.• The Local Chair and the Local Team -Haizhou Li the Chair organized and lead a wonderful group of people.While I cannot name every one of them, weekly meetings with the team members including related Chairs made our communication smooth and worked as a good time-keeper.For the remaining committee chairs, I only list them by names, as I cannot give all my gratitude only with short messages. Jonas Groschwitz, Shay B. Cohen, Lucia Donatelli, Meaghan Fowlie |
EMNLP | 2 |
| 2023 | Detecting and Mitigating Hallucinations in Multilingual SummarisationabstractHallucinations pose a significant challenge to the reliability of neural models for abstractive summarisation.While automatically generated summaries may be fluent, they often lack faithfulness to the original document.This issue becomes even more pronounced in lowresource languages, where summarisation requires cross-lingual transfer.With the existing faithful metrics focusing on English, even measuring the extent of this phenomenon in crosslingual settings is hard.To address this, we first develop a novel metric, mFACT, evaluating the faithfulness of non-English summaries, leveraging translation-based transfer from multiple English faithfulness metrics.Through extensive experiments in multiple languages, we demonstrate that mFACT is best suited to detect hallucinations compared to alternative metrics.With mFACT, we assess a broad range of multilingual large language models, and find that they all tend to hallucinate often in languages different from English.We then propose a simple but effective method to reduce hallucinations in cross-lingual transfer, which weighs the loss of each training example by its faithfulness score.This method drastically increases both performance and faithfulness according to both automatic and human evaluation when compared to strong baselines for cross-lingual transfer such as MAD-X. Yifu Qiu, Yftah Ziser, Anna Korhonen, Edoardo Maria Ponti, Shay B. Cohen |
EMNLP | 5 |
| 2023 | Erasure of Unaligned Attributes from Neural RepresentationsabstractAbstract We present the Assignment-Maximization Spectral Attribute removaL (AMSAL) algorithm, which erases information from neural representations when the information to be erased is implicit rather than directly being aligned to each input example. Our algorithm works by alternating between two steps. In one, it finds an assignment of the input representations to the information to be erased, and in the other, it creates projections of both the input representations and the information to be erased into a joint latent space. We test our algorithm on an extensive array of datasets, including a Twitter dataset with multiple guarded attributes, the BiasBios dataset, and the BiasBench benchmark. The latter benchmark includes four datasets with various types of protected attributes. Our results demonstrate that bias can often be removed in our setup. We also discuss the limitations of our approach when there is a strong entanglement between the main task and the information to be erased.1 Shun Shao, Yftah Ziser, Shay B. Cohen |
Trans. Assoc. Comput. Linguistics | 3 |
| 2022 | Factorizing Content and Budget Decisions in Abstractive Summarization of Long DocumentsabstractWe argue that disentangling content selection from the budget used to cover salient content improves the performance and applicability of abstractive summarizers.Our method, FAC-TORSUM 1 , does this disentanglement by factorizing summarization into two steps through an energy function: (1) generation of abstractive summary views covering salient information in subsets of the input document (document views) ; (2) combination of these views into a final summary, following a budget and content guidance.This guidance may come from different sources, including from an advisor model such as BART or BigBird, or in oracle modefrom the reference.This factorization achieves significantly higher ROUGE scores on multiple benchmarks for long document summarization, namely PubMed, arXiv, and GovReport.Notably, our model is effective for domain adaptation.When trained only on PubMed, it achieves a 46.29 ROUGE-1 score on arXiv, outperforming PEGASUS trained in domain by a large margin.Our experimental results indicate that the performance gains are due to more flexible budget adaptation and processing of shorter contexts provided by partial document views. Marcio Fonseca, Yftah Ziser, Shay B. Cohen |
EMNLP | 3 |
| 2022 | Sentence-Incremental Neural Coreference ResolutionabstractWe propose a sentence-incremental neural coreference resolution system which incrementally builds clusters after marking mention boundaries in a shift-reduce method.The system is aimed at bridging two recent approaches at coreference resolution: (1) state-of-the-art non-incremental models that incur quadratic complexity in document length with high computational cost, and (2) memory networkbased models which operate incrementally but do not generalize beyond pronouns.For comparison, we simulate an incremental setting by constraining non-incremental systems to form partial coreference chains before observing new sentences.In this setting, our system outperforms comparable state-of-the-art methods by 2 F1 on OntoNotes and 6.8 F1 on the CODI-CRAC 2021 corpus.In a conventional coreference setup, our system achieves 76.3 F1 on OntoNotes and 45.5 F1 on CODI-CRAC 2021, which is comparable to state-of-the-art baselines.We also analyze variations of our system and show that the degree of incrementality in the encoder has a surprisingly large effect on the resulting performance.1 1 Code is available at: https://github.com/mgrenander/sentence-incremental-coref Matt Grenander, Shay B. Cohen, Mark Steedman |
EMNLP | 2 |
| 2022 | Abstractive Summarization Guided by Latent Hierarchical Document StructureabstractSequential abstractive neural summarizers often do not use the underlying structure in the input article or dependencies between the input sentences.This structure is essential to integrate and consolidate information from different parts of the text.To address this shortcoming, we propose a hierarchy-aware graph neural network (HierGNN) which captures such dependencies through three main steps: 1) learning a hierarchical document structure through a latent structure tree learned by a sparse matrixtree computation; 2) propagating sentence information over this structure using a novel message-passing node propagation mechanism to identify salient information; 3) using graphlevel attention to concentrate the decoder on salient information.Experiments confirm Hi-erGNN improves strong sequence models such as BART, with a 0.55 and 0.75 margin in average ROUGE-1/2/L for CNN/DM and XSum.Further human evaluation demonstrates that summaries produced by our model are more relevant and less redundant than the baselines, into which HierGNN is incorporated.We also find HierGNN synthesizes summaries by fusing multiple source sentences more, rather than compressing a single source sentence, and that it processes long inputs more effectively.1 Yifu Qiu, Shay B. Cohen |
EMNLP | 2 |
| 2021 | A Differentiable Relaxation of Graph Segmentation and Alignment for AMR ParsingabstractMeaning Representations (AMR) are a broad-coverage semantic formalism which represents sentence meaning as a directed acyclic graph.To train most AMR parsers, one needs to segment the graph into subgraphs and align each such subgraph to a word in a sentence; this is normally done at preprocessing, relying on hand-crafted rules.In contrast, we treat both alignment and segmentation as latent variables in our model and induce them as part of end-to-end training.As marginalizing over the structured latent variables is infeasible, we use the variational autoencoding framework.To ensure end-to-end differentiable optimization, we introduce a differentiable relaxation of the segmentation and alignment problems.We observe that inducing segmentation yields substantial gains over using a 'greedy' segmentation heuristic.The performance of our method also approaches that of a model that relies on the segmentation rules of Lyu and Titov (2018), which were hand-crafted to handle individual AMR constructions. Chunchuan Lyu, Shay B. Cohen, Ivan Titov 0001 |
EMNLP (1) | 2 |
| 2021 | A Root of a Problem: Optimizing Single-Root Dependency ParsingabstractWe describe two approaches to single-root dependency parsing that yield significant speed ups in such parsing.One approach has been previously used in dependency parsers in practice, but remains undocumented in the parsing literature, and is considered a heuristic.We show that this approach actually finds the optimal dependency tree.The second approach relies on simple reweighting of the inference graph being input to the dependency parser and has an optimal running time.Here, we again show that this approach is fully correct and identifies the highest-scoring parse tree.Our experiments demonstrate a manyfold speed up compared to a previous graph-based state-of-the-art parser without any loss in accuracy or optimality. 1 Milos Stanojevic, Shay B. Cohen |
EMNLP (1) | 2 |
| 2021 | Text Generation from Discourse Representation StructuresabstractWe propose neural models to generate text from formal meaning representations based on Discourse Representation Structures (DRSs).DRSs are document-level representations which encode rich semantic detail pertaining to rhetorical relations, presupposition, and co-reference within and across sentences.We formalize the task of neural DRS-to-text generation and provide modeling solutions for the problems of condition ordering and variable naming which render generation from DRSs non-trivial.Our generator relies on a novel sibling treeLSTM model which is able to accurately represent DRS structures and is more generally suited to trees with wide branches.We achieve competitive performance (59.48 BLEU) on the GMB benchmark against several strong baselines. Jiangming Liu, Shay B. Cohen, Mirella Lapata |
NAACL-HLT | 2 |
| 2021 | Universal Discourse Representation Structure ParsingabstractAbstract We consider the task of crosslingual semantic parsing in the style of Discourse Representation Theory (DRT) where knowledge from annotated corpora in a resource-rich language is transferred via bitext to guide learning in other languages. We introduce Universal Discourse Representation Theory (UDRT), a variant of DRT that explicitly anchors semantic representations to tokens in the linguistic input. We develop a semantic parsing framework based on the Transformer architecture and utilize it to obtain semantic resources in multiple languages following two learning schemes. The Many-to-One approach translates non-English text to English, and then runs a relatively accurate English parser on the translated text, while the One-to-Many approach translates gold standard English to non-English text and trains multiple parsers (one per language) on the translations. Experimental results on the Parallel Meaning Bank show that our proposal outperforms strong baselines by a wide margin and can be used to construct (silver-standard) meaning banks for 99 languages. Jiangming Liu, Shay B. Cohen, Mirella Lapata, Johan Bos |
Comput. Linguistics | 2 |
| 2021 | Bottom-up unranked tree-to-graph transducers for translation into semantic graphsabstractWe develop a finite-state transducer for translating unranked trees into general graphs. This work is motivated by recent progress in semantic parsing for natural language, where sentences are first mapped into tree-shaped syntactic representations, and then these trees are translated into graph semantic representations. We investigate formal properties of our tree-to-graph transducers and develop a polynomial time algorithm for translating a weighted language of input trees into a packed representation, from which best-score graphs can be efficiently recovered. Johanna Björklund, Shay B. Cohen, Frank Drewes, Giorgio Satta |
Theor. Comput. Sci. | 2 |
| 2020 | Learning Dialog Policies from Weak DemonstrationsabstractDeep reinforcement learning is a promising approach to training a dialog manager, but current methods struggle with the large state and action spaces of multi-domain dialog systems.Building upon Deep Q-learning from Demonstrations (DQfD), an algorithm that scores highly in difficult Atari games, we leverage dialog data to guide the agent to successfully respond to a user's requests.We make progressively fewer assumptions about the data needed, using labeled, reduced-labeled, and even unlabeled data to train expert demonstrators.We introduce Reinforced Fine-tune Learning, an extension to DQfD, enabling us to overcome the domain gap between the datasets and the environment.Experiments in a challenging multi-domain dialog system framework validate our approaches, and get high success rates even when trained on outof-domain data. Gabriel Gordon-Hall, Philip John Gorinski, Shay B. Cohen |
ACL | 3 |
| 2020 | Machine Reading of Historical EventsabstractMachine reading is an ambitious goal in NLP that subsumes a wide range of text understanding capabilities.Within this broad framework, we address the task of machine reading the time of historical events, compile datasets for the task, and develop a model for tackling it.Given a brief textual description of an event, we show that good performance can be achieved by extracting relevant sentences from Wikipedia, and applying a combination of taskspecific and general-purpose feature embeddings for the classification.Furthermore, we establish a link between the historical event ordering task and the event focus time task from the information retrieval literature, showing they also provide a challenging test case for machine reading algorithms. 1 1 Code and data are available at https://github.com/ltorroba/ machine-reading-historical-events.* Equal contribution.Year Event text OTD 2005 107 die in Amagasaki rail crash in Japan.1939 BMI (Broadcast Music Incorporated) formed.1864 General Sherman's armies reach Savannah & 12 day siege begins.WOTD 1887 Buffalo Bill Cody's Wild West Show opens in London.1399 Henry IV is proclaimed King of England.1943 First Flight of the Gloster Meteor, Britain's first combat jet aircraft. Or Honovich, Lucas Torroba Hennigen, Omri Abend, Shay B. Cohen |
ACL | 4 |
| 2020 | Dscorer: A Fast Evaluation Metric for Discourse Representation Structure ParsingabstractDiscourse representation structures (DRSs) are scoped semantic representations for texts of arbitrary length.Evaluation of the accuracy of predicted DRSs plays a key role in developing semantic parsers and improving their performance.DRSs are typically visualized as nested boxes, in a way that is not straightforward to process automatically.COUNTER, an evaluation algorithm for DRSs, transforms them to clauses and measures clause overlap by searching for variable mappings between two DRSs.Unfortunately, COUNTER is computationally costly (with respect to memory and CPU time) and does not scale with longer texts.We introduce DSCORER, an efficient new metric which converts box-style DRSs to graphs and then measures the overlap of n-grams in the graphs.Experiments show that DSCORER computes accuracy scores that correlate with scores from COUNTER at a fraction of the time. Jiangming Liu, Shay B. Cohen, Mirella Lapata |
ACL | 2 |
| 2020 | Multi-Step Inference for Reasoning Over ParagraphsabstractComplex reasoning over text requires understanding and chaining together free-form predicates and logical connectives.Prior work has largely tried to do this either symbolically or with black-box transformers.We present a middle ground between these two extremes: a compositional model reminiscent of neural module networks that can perform chained logical reasoning.This model first finds relevant sentences in the context and then chains them together using neural modules.Our model gives significant performance improvements (up to 29% relative error reduction when combined with a reranker) on ROPES, a recentlyintroduced complex reasoning dataset. Jiangming Liu, Matt Gardner 0001, Shay B. Cohen, Mirella Lapata |
EMNLP (1) | 3 |
| 2020 | Lightweight, Dynamic Graph Convolutional Networks for AMR-to-Text GenerationabstractAMR-to-text generation is used to transduce Abstract Meaning Representation structures (AMR) into text.A key challenge in this task is to efficiently learn effective graph representations.Previously, Graph Convolution Networks (GCNs) were used to encode input AMRs, however, vanilla GCNs are not able to capture non-local information and additionally, they follow a local (first-order) information aggregation scheme.To account for these issues, larger and deeper GCN models are required to capture more complex interactions.In this paper, we introduce a dynamic fusion mechanism, proposing Lightweight Dynamic Graph Convolutional Networks (LDGCNs) that capture richer non-local interactions by synthesizing higher order information from the input graphs.We further develop two novel parameter saving strategies based on the group graph convolutions and weight tied convolutions to reduce memory usage and model complexity.With the help of these strategies, we are able to train a model with fewer parameters while maintaining the model capacity.Experiments demonstrate that LDGCNs outperform stateof-the-art models on two benchmark datasets for AMR-to-text generation with significantly fewer parameters. Yan Zhang 0004, Zhijiang Guo, Zhiyang Teng, Wei Lu 0011, Shay B. Cohen, Zuozhu Liu, Lidong Bing |
EMNLP (1) | 5 |
| 2020 | Compositional languages emerge in a neural iterated learning model
Shangmin Guo, Matthieu Labeau, Shay B. Cohen, Simon Kirby |
ICLR | 4 |
| 2020 | Learning Latent Forests for Medical Relation ExtractionabstractThe goal of medical relation extraction is to detect relations among entities, such as genes, mutations and drugs in medical texts. Dependency tree structures have been proven useful for this task. Existing approaches to such relation extraction leverage off-the-shelf dependency parsers to obtain a syntactic tree or forest for the text. However, for the medical domain, low parsing accuracy may lead to error propagation downstream the relation extraction pipeline. In this work, we propose a novel model which treats the dependency structure as a latent variable and induces it from the unstructured text in an end-to-end fashion. Our model can be understood as composing task-specific dependency forests that capture non-local interactions for better relation extraction. Extensive results on four datasets show that our model is able to significantly outperform state-of-the-art systems without relying on any direct tree supervision or pre-training. Zhijiang Guo, Guoshun Nan, Wei Lu 0011, Shay B. Cohen |
IJCAI | 4 |
| 2019 | Duality of Link Prediction and Entailment Graph InductionabstractLink prediction and entailment graph induction are often treated as different problems.In this paper, we show that these two problems are actually complementary.We train a link prediction model on a knowledge graph of assertions extracted from raw text.We propose an entailment score that exploits the new facts discovered by the link prediction model, and then form entailment graphs between relations.We further use the learned entailments to predict improved link prediction scores.Our results show that the two tasks can benefit from each other.The new entailment score outperforms prior state-of-the-art results on a standard entialment dataset and the new link prediction scores show improvements over the raw link prediction scores. Mohammad Javad Hosseini, Shay B. Cohen, Mark Johnson 0001, Mark Steedman |
ACL (1) | 2 |
| 2019 | Discourse Representation Parsing for Sentences and DocumentsabstractWe introduce a novel semantic parsing task based on Discourse Representation Theory (DRT; Kamp and Reyle 1993).Our model operates over Discourse Representation Tree Structures which we formally define for sentences and documents.We present a general framework for parsing discourse structures of arbitrary length and granularity.We achieve this with a neural model equipped with a supervised hierarchical attention mechanism and a linguistically-motivated copy strategy.Experimental results on sentence-and documentlevel benchmarks show that our model outperforms competitive baselines by a wide margin. Jiangming Liu, Shay B. Cohen, Mirella Lapata |
ACL (1) | 2 |
| 2019 | Wide-Coverage Neural A* Parsing for Minimalist GrammarsabstractMinimalist Grammars (Stabler, 1997) are a computationally oriented, and rigorous formalisation of many aspects of Chomsky's (1995) Minimalist Program.This paper presents the first ever application of this formalism to the task of realistic wide-coverage parsing.The parser uses a linguistically expressive yet highly constrained grammar, together with an adaptation of the A* search algorithm currently used in CCG parsing (Lewis and Steedman, 2014;Lewis et al., 2016), with supertag probabilities provided by a bi-LSTM neural network supertagger trained on MGbank, a corpus of MG derivation trees.We report on some promising initial experimental results for overall dependency recovery as well as on the recovery of certain unbounded long distance dependencies.Finally, although like other MG parsers, ours has a high order polynomial worst case time complexity, we show that in practice its expected time complexity is O(n 3 ).The parser is publicly available. John Torr, Milos Stanojevic, Mark Steedman, Shay B. Cohen |
ACL (1) | 4 |
| 2019 | Experimenting with Power Divergences for Language ModelingabstractMatthieu Labeau, Shay B. Cohen. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Matthieu Labeau, Shay B. Cohen |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Semantic Role Labeling with Iterative Structure RefinementabstractChunchuan Lyu, Shay B. Cohen, Ivan Titov. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Chunchuan Lyu, Shay B. Cohen, Ivan Titov 0001 |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Partners in Crime: Multi-view Sequential Inference for Movie UnderstandingabstractNikos Papasarantopoulos, Lea Frermann, Mirella Lapata, Shay B. Cohen. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Nikos Papasarantopoulos, Lea Frermann, Mirella Lapata, Shay B. Cohen |
EMNLP/IJCNLP (1) | 4 |
| 2019 | What is this Article about? Extreme Summarization with Topic-aware Convolutional Neural NetworksabstractWe introduce "extreme summarization," a new single-document summarization task which aims at creating a short, one-sentence news summary answering the question "What is the article about?". We argue that extreme summarization, by nature, is not amenable to extractive strategies and requires an abstractive modeling approach. In the hope of driving research on this task further: (a) we collect a real-world, large scale dataset by harvesting online articles from the British Broadcasting Corporation (BBC); and (b) propose a novel abstractive model which is conditioned on the article's topics and based entirely on convolutional neural networks. We demonstrate experimentally that this architecture captures long-range dependencies in a document and recognizes pertinent content, outperforming an oracle extractive system and state-of-the-art abstractive approaches when evaluated automatically and by humans on the extreme summarization dataset. Shashi Narayan, Shay B. Cohen, Mirella Lapata |
J. Artif. Intell. Res. | 2 |
| 2019 | Unlexicalized Transition-based Discontinuous Constituency ParsingabstractAbstract Lexicalized parsing models are based on the assumptions that (i) constituents are organized around a lexical head and (ii) bilexical statistics are crucial to solve ambiguities. In this paper, we introduce an unlexicalized transition-based parser for discontinuous constituency structures, based on a structure-label transition system and a bi-LSTM scoring system. We compare it with lexicalized parsing models in order to address the question of lexicalization in the context of discontinuous constituency parsing. Our experiments show that unlexicalized models systematically achieve higher results than lexicalized models, and provide additional empirical evidence that lexicalization is not necessary to achieve strong parsing results. Our best unlexicalized model sets a new state of the art on English and German discontinuous constituency treebanks. We further provide a per-phenomenon analysis of its errors on discontinuous constituents. Maximin Coavoux, Benoît Crabbé, Shay B. Cohen |
Trans. Assoc. Comput. Linguistics | 3 |
| 2018 | Canonical Correlation Inference for Mapping Abstract Scenes to Text
Nikos Papasarantopoulos, Helen Jiang, Shay B. Cohen |
AAAI | 3 |
| 2018 | Stock Movement Prediction from Tweets and Historical PricesabstractStock movement prediction is a challenging problem: the market is highly stochastic, and we make temporally-dependent predictions from chaotic data. We treat these three complexities and present a novel deep generative model jointly exploiting text and price signals for this task. Unlike the case with discriminative or topic modeling, our model introduces recurrent, continuous latent variables for a better treatment of stochasticity, and uses neural variational inference to address the intractable posterior inference. We also provide a hybrid objective with temporal auxiliary to flexibly capture predictive dependencies. We demonstrate the state-of- the-art performance of our proposed model on a new stock movement prediction dataset which we collected.11https://github.com/yumoxu/stocknet-dataset Yumo Xu, Shay B. Cohen |
ACL (1) | 2 |
| 2018 | Discourse Representation Structure ParsingabstractWe introduce an open-domain neural semantic parser which generates formal meaning representations in the style of Discourse Representation Theory (DRT; Kamp and Reyle 1993).We propose a method which transforms Discourse Representation Structures (DRSs) to trees and develop a structure-aware model which decomposes the decoding process into three stages: basic DRS structure prediction, condition prediction (i.e., predicates and relations), and referent prediction (i.e., variables).Experimental results on the Groningen Meaning Bank (GMB) show that our model outperforms competitive baselines by a wide margin. Jiangming Liu, Shay B. Cohen, Mirella Lapata |
ACL (1) | 2 |
| 2018 | Document Modeling with External Attention for Sentence ExtractionabstractShashi Narayan, Ronald Cardenas, Nikos Papasarantopoulos, Shay B. Cohen, Mirella Lapata, Jiangsheng Yu, Yi Chang. Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2018. Shashi Narayan, Ronald Cardenas, Nikos Papasarantopoulos, Shay B. Cohen, Mirella Lapata, Jiangsheng Yu, Yi Chang 0001 |
ACL (1) | 4 |
| 2018 | Local String Transduction as Sequence LabelingabstractWe show that the general problem of string transduction can be reduced to the problem of sequence labeling. While character deletion and insertions are allowed in string transduction, they do not exist in sequence labeling. We show how to overcome this difference. Our approach can be used with any sequence labeling algorithm and it works best for problems in which string transduction imposes a strong notion of locality (no long range dependencies). We experiment with spelling correction for social media, OCR correction, and morphological inflection, and we see that it behaves better than seq2seq models and yields state-of-the-art results in several cases. Joana Ribeiro, Shashi Narayan, Shay B. Cohen, Xavier Carreras |
COLING | 3 |
| 2018 | Privacy-preserving Neural Representations of TextabstractThis article deals with adversarial attacks towards deep learning systems for Natural Language Processing (NLP), in the context of privacy protection.We study a specific type of attack: an attacker eavesdrops on the hidden representations of a neural text classifier and tries to recover information about the input text.Such scenario may arise in situations when the computation of a neural network is shared across multiple devices, e.g.some hidden representation is computed by a user's device and sent to a cloud-based model.We measure the privacy of a hidden representation by the ability of an attacker to predict accurately specific private information from it and characterize the tradeoff between the privacy and the utility of neural representations.Finally, we propose several defense methods based on modified training objectives and show that they improve the privacy of neural representations. Maximin Coavoux, Shashi Narayan, Shay B. Cohen |
EMNLP | 3 |
| 2018 | Multilingual Clustering of Streaming NewsabstractClustering news across languages enables efficient media monitoring by aggregating articles from multilingual sources into coherent stories.Doing so in an online setting allows scalable processing of massive news streams.To this end, we describe a novel method for clustering an incoming stream of multilingual documents into monolingual and crosslingual story clusters.Unlike typical clustering approaches that consider a small and known number of labels, we tackle the problem of discovering an ever growing number of cluster labels in an online fashion, using real news datasets in multiple languages.Our method is simple to implement, computationally efficient and produces state-of-the-art results on datasets in German, English and Spanish. Sebastião Miranda, Arturs Znotins, Shay B. Cohen, Guntis Barzdins |
EMNLP | 3 |
| 2018 | Don't Give Me the Details, Just the Summary! Topic-Aware Convolutional Neural Networks for Extreme SummarizationabstractWe introduce extreme summarization, a new single-document summarization task which does not favor extractive strategies and calls for an abstractive modeling approach.The idea is to create a short, one-sentence news summary answering the question "What is the article about?".We collect a real-world, large scale dataset for this task by harvesting online articles from the British Broadcasting Corporation (BBC).We propose a novel abstractive model which is conditioned on the article's topics and based entirely on convolutional neural networks.We demonstrate experimentally that this architecture captures longrange dependencies in a document and recognizes pertinent content, outperforming an oracle extractive system and state-of-the-art abstractive approaches when evaluated automatically and by humans. 1 Shashi Narayan, Shay B. Cohen, Mirella Lapata |
EMNLP | 2 |
| 2018 | Cross-Lingual Abstract Meaning Representation ParsingabstractMarco Damonte, Shay B. Cohen. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Marco Damonte, Shay B. Cohen |
NAACL-HLT | 2 |
| 2018 | Abstract Meaning Representation for Paraphrase DetectionabstractFuad Issa, Marco Damonte, Shay B. Cohen, Xiaohui Yan, Yi Chang. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Fuad Issa, Marco Damonte, Shay B. Cohen, Yi Chang 0001 |
NAACL-HLT | 3 |
| 2018 | Ranking Sentences for Extractive Summarization with Reinforcement LearningabstractShashi Narayan, Shay B. Cohen, Mirella Lapata. Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers). 2018. Shashi Narayan, Shay B. Cohen, Mirella Lapata |
NAACL-HLT | 2 |
| 2018 | Bayesian Analysis in Natural Language Processing Shay Cohen (University of Edinburgh)Morgan & Claypool (Synthesis Lectures on Human Language Technologies, edited by Graeme Hirst, volume 35), 2016, xxvii+246 pp; paperback, ISBN 9781627058735, $85.00; ebook, ISBN 9781627054218, $68.00; doi: 10.2200/S00719ED1V01Y201605HLT035abstractBayesian techniques are useful tools for modeling a wide range of data and phenomena. Natural language is no exception. The Bayesian approach works as follows:1. Model the data x probabilistically with p(x|θ), where θ are some unknown parameters. For example, this could be a generative story for a sentence x, based on some unknown context-free grammar parameters θ.2. Represent the uncertainty about θ by a prior distribution p(θ).3. Given data, apply Bayes theorem p(θ|x) ∝ p(x|θ)p(θ) to find the posterior distribution for the quantities of interest.This approach enables an elegant and unified way to incorporate prior knowledge and manage uncertainty over parameters. It can also be used to provide capacity control for complex models as an alternative to smoothing. There have been many successful applications of Bayesian techniques in natural language processing (NLP). Some examples include: word segmentation (Goldwater et al. 2009), syntax (Johnson et al. 2007), morphology (Snyder & Barzilay 2008), coreference resolution (Haghighi & Klein 2007), and machine translation (Blunsom et al. 2009).Cohen’s book provides an accessible yet in-depth introduction to Bayesian techniques. It is aimed at a researcher or student who is already familiar with statistical modeling in natural language (i.e., at the level of introductory books such as Manning & Schütze [1999], Jurafsky & Martin [2009]). The stated goal of the book is to “cover the methods and algorithms that are needed to fluently read Bayesian learning papers in NLP and to do research in the area.” I believe Cohen successfully achieves this goal, striking a nice balance between breadth and depth of material.Chapter 1 is a brief review of probability and statistics. It covers prerequisite concepts such as independence, conditional independence, and exchangeability of random variables. The differences between Bayesian and frequentist philosophies are discussed, albeit briefly. In general, the book maintains a pragmatic approach, focusing more on the mathematics and less on the philosophy.Chapter 2 motivates Bayesian techniques in NLP with two distinct examples: latent Dirichlet allocation (LDA) for unsupervised topic modeling and Bayesian linear regression for supervised text analytics. Although Bayesian techniques are most often used for unsupervised problems in NLP (e.g., for addressing problems where the Expectation-Maximization [EM] algorithm fails to find good solutions), the supervised example demonstrates their broader applicability. One aspect that I appreciate about Cohen’s exposition is that he strives to distinguish between what is technically feasible with Bayesian techniques vs. what is frequently used in research. This helps avoid potential misconceptions.Chapter 3 describes the priors p(θ) that are common in NLP: for example, the Dirichlet distribution, the logistic normal distribution, and non-informative priors such as the Jeffreys prior.Chapters 4–6 explain Bayesian inference in detail and form the core of the book. How does one combine the data likelihood and the prior to compute the posterior distribution p(θ|x) ∝ p(x|θ)p(θ)? This can be a computationally difficult problem. Chapter 4 discusses the maximum a posterior (MAP) estimation of p(θ|x). Chapter 5 covers sampling methods, particularly Gibbs sampling, Metropolis-Hastings sampling, and slice sampling. Chapter 6 describes variational inference, the mean-field approximation, and the variational EM algorithm. In each chapter, algorithm variants that are popular in NLP research are discussed in detail. An example is blocked Gibbs sampling using dynamic programming.Chapter 7 focuses on Bayesian nonparameterics. These models allow for an infinite-dimensional parameter space, but the actual number of parameters grows with sample size. This is an active field of research. Cohen does a laudable job of explaining the basic mathematics behind the Dirichlet process, which generalizes the Dirichlet distribution. He describes the stick-breaking and Chinese Restaurant Process viewpoints, and shows how the Dirichlet process can be used as a prior in a nonparametric mixture model and a hierarchical topic model.Chapter 8, the final chapter, demonstrates how Bayesian techniques can be applied to grammar models in NLP. It focuses on parametric and non-parameteric Bayesian models for probabilistic context-free grammars. I wished there was an additional chapter on other applications besides grammar models. However, I think a reader who completes this book would have gained the technical background to do such a survey by themselves.In summary, Cohen’s Bayesian Analysis in Natural Language Processing is a good starting point for a researcher or a student who wishes to learn more about Bayesian techniques. It covers the necessary and sufficient knowledge needed to understand papers in this area, and leaves the remaining details as references. It can be viewed as a concise introduction that complements the many excellent statistics textbooks on Bayesian techniques, which tend to be more detailed.I wish I had this book when I first started to learn Bayesian techniques back in 2009. I recall struggling through several advanced statistics textbooks, before being rescued by Kevin Knight’s classic workbook, “Bayesian Inference with Tears.”1 Cohen’s book might have saved me some Kleenex. Shay B. Cohen, Graeme Hirst, Kevin Duh |
Comput. Linguistics | 1 |
| 2018 | Whodunnit? Crime Drama as a Case for Natural Language UnderstandingabstractIn this paper we argue that crime drama exemplified in television programs such as CSI: Crime Scene Investigation is an ideal testbed for approximating real-world natural language understanding and the complex inferences associated with it. We propose to treat crime drama as a new inference task, capitalizing on the fact that each episode poses the same basic question (i.e., who committed the crime) and naturally provides the answer when the perpetrator is revealed. We develop a new dataset based on CSI episodes, formalize perpetrator identification as a sequence labeling problem, and develop an LSTM-based model which learns from multi-modal data. Experimental results show that an incremental inference strategy is key to making accurate guesses as well as learning from representations fusing textual, visual, and acoustic input. Lea Frermann, Shay B. Cohen, Mirella Lapata |
Trans. Assoc. Comput. Linguistics | 2 |
| 2018 | Learning Typed Entailment Graphs with Global Soft ConstraintsabstractThis paper presents a new method for learning typed entailment graphs from text. We extract predicate-argument structures from multiple-source news corpora, and compute local distributional similarity scores to learn entailments between predicates with typed arguments (e.g., person contracted disease). Previous work has used transitivity constraints to improve local decisions, but these constraints are intractable on large graphs. We instead propose a scalable method that learns globally consistent similarity scores based on new soft constraints that consider both the structures across typed entailment graphs and inside each graph. Learning takes only a few hours to run over 100K predicates and our results show large improvements over local similarity scores on two entailment data sets. We further show improvements over paraphrases and entailments from the Paraphrase Database, and prior state-of-the-art entailment graphs. We show that the entailment graphs improve performance in a downstream task. Mohammad Javad Hosseini, Nathanael Chambers, Siva Reddy, Xavier R. Holt, Shay B. Cohen, Mark Johnson 0001, Mark Steedman |
Trans. Assoc. Comput. Linguistics | 5 |
| 2017 | An Incremental Parser for Abstract Meaning RepresentationabstractMeaning Representation (AMR) is a semantic representation for natural language that embeds annotations related to traditional tasks such as named entity recognition, semantic role labeling, word sense disambiguation and co-reference resolution.We describe a transition-based parser for AMR that parses sentences leftto-right, in linear time.We further propose a test-suite that assesses specific subtasks that are helpful in comparing AMR parsers, and show that our parser is competitive with the state of the art on the LDC2015E86 dataset and that it outperforms state-of-the-art parsers for recovering named entities and handling polarity. Marco Damonte, Shay B. Cohen, Giorgio Satta |
EACL (1) | 2 |
| 2017 | Split and RephraseabstractWe propose a new sentence simplification task (Split-and-Rephrase) where the aim is to split a complex sentence into a meaning preserving sequence of shorter sentences.Like sentence simplification, splitting-and-rephrasing has the potential of benefiting both natural language processing and societal applications.Because shorter sentences are generally better processed by NLP systems, it could be used as a preprocessing step which facilitates and improves the performance of parsers, semantic role labelers and machine translation systems.It should also be of use for people with reading disabilities because it allows the conversion of longer sentences into shorter ones.This paper makes two contributions towards this new task.First, we create and make available a benchmark consisting of 1,066,115 tuples mapping a single complex sentence to a sequence of sentences expressing the same meaning.1 Second, we propose five models (vanilla sequence-to-sequence to semantically-motivated models) to understand the difficulty of the proposed task. Shashi Narayan, Claire Gardent, Shay B. Cohen, Anastasia Shimorina |
EMNLP | 3 |
| 2016 | Optimizing Spectral Learning for ParsingabstractWe describe a search algorithm for optimizing the number of latent states when estimating latent-variable PCFGs with spectral methods.Our results show that contrary to the common belief that the number of latent states for each nonterminal in an L-PCFG can be decided in isolation with spectral methods, parsing results significantly improve if the number of latent states for each nonterminal is globally optimized, while taking into account interactions between the different nonterminals.In addition, we contribute an empirical analysis of spectral algorithms on eight morphologically rich languages: Basque, French, German, Hebrew, Hungarian, Korean, Polish and Swedish.Our results show that our estimation consistently performs better or close to coarse-to-fine expectation-maximization techniques for these languages. Shashi Narayan, Shay B. Cohen |
ACL (1) | 2 |
| 2016 | Low-Rank Approximation of Weighted Tree AutomataabstractWe describe a technique to minimize weighted tree automata (WTA), a powerful formalisms that subsumes probabilistic context-free grammars (PCFGs) and latent-variable PCFGs. Our method relies on a singular value decomposition of the underlying Hankel matrix defined by the WTA. Our main theoretical result is an efficient algorithm for computing the SVD of an infinite Hankel matrix implicitly represented as a WTA. We evaluate our method on real-world data originating in newswire treebank. We show that the model achieves lower perplexity than previous methods for PCFG minimization, and also is much more stable due to the absence of local optima. Guillaume Rabusseau, Borja Balle, Shay B. Cohen |
AISTATS | 3 |
| 2016 | Semi-Supervised Learning of Sequence Models with Method of MomentsabstractWe propose a fast and scalable method for semi-supervised learning of sequence models, based on anchor words and moment matching. Our method can handle hidden Markov models with feature-based log-linear emissions. Unlike other semi-supervised methods, no decoding passes are necessary on the unlabeled data and no graph needs to be constructed— only one pass is necessary to collect moment statistics. The model parameters are estimated by solving a small quadratic program for each feature. Experiments on part-of-speech (POS) tagging for Twitter and for a low-resource language (Malagasy) show that our method can learn from very few annotated sentences. Zita Marinho, André F. T. Martins, Shay B. Cohen, Noah A. Smith |
EMNLP | 3 |
| 2016 | Paraphrase Generation from Latent-Variable PCFGs for Semantic ParsingabstractOne of the limitations of semantic parsing approaches to open-domain question answering is the lexicosyntactic gap between natural language questions and knowledge base entries -there are many ways to ask a question, all with the same answer.In this paper we propose to bridge this gap by generating paraphrases of the input question with the goal that at least one of them will be correctly mapped to a knowledge-base query.We introduce a novel grammar model for paraphrase generation that does not require any sentence-aligned paraphrase corpus.Our key idea is to leverage the flexibility and scalability of latent-variable probabilistic context-free grammars to sample paraphrases.We do an extrinsic evaluation of our paraphrases by plugging them into a semantic parser for Freebase.Our evaluation experiments on the WebQuestions benchmark dataset show that the performance of the semantic parser improves over strong baselines. Shashi Narayan, Siva Reddy, Shay B. Cohen |
INLG | 3 |
| 2016 | Parsing Linear Context-Free Rewriting Systems with Fast Matrix MultiplicationabstractWe describe a recognition algorithm for a subset of binary linear context-free rewriting systems (LCFRS) with running time O(nωd) where M(m) = O(mω) is the running time for m × m matrix multiplication and d is the “contact rank” of the LCFRS—the maximal number of combination and non-combination points that appear in the grammar rules. We also show that this algorithm can be used as a subroutine to obtain a recognition algorithm for general binary LCFRS with running time O(nωd+1). The currently best known ω is smaller than 2.38. Our result provides another proof for the best known result for parsing mildly context-sensitive formalisms such as combinatory categorial grammars, head grammars, linear indexed grammars, and tree-adjoining grammars, which can be parsed in time O(n4.76). It also shows that inversion transduction grammars can be parsed in time O(n5.76). In addition, binary LCFRS subsumes many other formalisms and types of grammars, for some of which we also improve the asymptotic complexity of parsing. Shay B. Cohen, Daniel Gildea |
Comput. Linguistics | 1 |
| 2016 | Encoding Prior Knowledge with Eigenword EmbeddingsabstractCanonical correlation analysis (CCA) is a method for reducing the dimension of data represented using two views. It has been previously used to derive word embeddings, where one view indicates a word, and the other view indicates its context. We describe a way to incorporate prior knowledge into CCA, give a theoretical justification for it, and test it by deriving word embeddings and evaluating them on a myriad of datasets. Dominique Osborne, Shashi Narayan, Shay B. Cohen |
Trans. Assoc. Comput. Linguistics | 3 |
| 2015 | A Coactive Learning View of Online Structured Prediction in Statistical Machine TranslationabstractWe present a theoretical analysis of online parameter tuning in statistical machine translation (SMT) from a coactive learning view.This perspective allows us to give regret and generalization bounds for latent perceptron algorithms that are common in SMT, but fall outside of the standard convex optimization scenario.Coactive learning also introduces the concept of weak feedback, which we apply in a proofof-concept experiment to SMT, showing that learning from feedback that consists of slight improvements over predictions leads to convergence in regret and translation error rate.This suggests that coactive learning might be a viable framework for interactive machine translation.Furthermore, we find that surrogate translations replacing references that are unreachable in the decoder search space can be interpreted as weak feedback and lead to convergence in learning, if they admit an underlying linear model. Artem Sokolov 0001, Stefan Riezler, Shay B. Cohen |
CoNLL | 3 |
| 2015 | Conversation Trees: A Grammar Model for Topic Structure in ForumsabstractOnline forum discussions proceed differently from face-to-face conversations and any single thread on an online forum contains posts on different subtopics.This work aims to characterize the content of a forum thread as a conversation tree of topics.We present models that jointly perform two tasks: segment a thread into subparts, and assign a topic to each part.Our core idea is a definition of topic structure using probabilistic grammars.By leveraging the flexibility of two grammar formalisms, Context-Free Grammars and Linear Context-Free Rewriting Systems, our models create desirable structures for forum threads: our topic segmentation is hierarchical, links non-adjacent segments on the same topic, and jointly labels the topic during segmentation.We show that our models outperform a number of tree generation baselines. Annie Louis, Shay B. Cohen |
EMNLP | 2 |
| 2015 | Diversity in Spectral Learning for Natural Language ParsingabstractWe describe an approach to create a diverse set of predictions with spectral learning of latent-variable PCFGs (L-PCFGs).Our approach works by creating multiple spectral models where noise is added to the underlying features in the training set before the estimation of each model.We describe three ways to decode with multiple models.In addition, we describe a simple variant of the spectral algorithm for L-PCFGs that is fast and leads to compact models.Our experiments for natural language parsing, for English and German, show that we get a significant improvement over baselines comparable to state of the art.For English, we achieve the F 1 score of 90.18, and for German we achieve the F 1 score of 83.38. Shashi Narayan, Shay B. Cohen |
EMNLP | 2 |
| 2015 | Lexical Event Ordering with an Edge-Factored ModelabstractExtensive lexical knowledge is necessary for temporal analysis and planning tasks.We address in this paper a lexical setting that allows for the straightforward incorporation of rich features and structural constraints.We explore a lexical event ordering task, namely determining the likely temporal order of events based solely on the identity of their predicates and arguments.We propose an "edgefactored" model for the task that decomposes over the edges of the event graph.We learn it using the structured perceptron.As lexical tasks require large amounts of text, we do not attempt manual annotation and instead use the textual order of events in a domain where this order is aligned with their temporal order, namely cooking recipes. Omri Abend, Shay B. Cohen, Mark Steedman |
HLT-NAACL | 2 |
| 2014 | Lexical Inference over Multi-Word Predicates: A Distributional ApproachabstractRepresenting predicates in terms of their argument distribution is common practice in NLP.Multi-word predicates (MWPs) in this context are often either disregarded or considered as fixed expressions.The latter treatment is unsatisfactory in two ways: (1) identifying MWPs is notoriously difficult, (2) MWPs show varying degrees of compositionality and could benefit from taking into account the identity of their component parts.We propose a novel approach that integrates the distributional representation of multiple sub-sets of the MWP's words.We assume a latent distribution over sub-sets of the MWP, and estimate it relative to a downstream prediction task.Focusing on the supervised identification of lexical inference relations, we compare against state-of-the-art baselines that consider a single sub-set of an MWP, obtaining substantial improvements.To our knowledge, this is the first work to address lexical relations between MWPs of varying degrees of compositionality within distributional semantics. Omri Abend, Shay B. Cohen, Mark Steedman |
ACL (1) | 2 |
| 2014 | A Provably Correct Learning Algorithm for Latent-Variable PCFGsabstractWe introduce a provably correct learning algorithm for latent-variable PCFGs. The algorithm relies on two steps: first, the use of a matrix-decomposition algorithm ap-plied to a co-occurrence matrix estimated from the parse trees in a training sample; second, the use of EM applied to a convex objective derived from the training sam-ples in combination with the output from the matrix decomposition. Experiments on parsing and a language modeling problem show that the algorithm is efficient and ef-fective in practice. 1 Shay B. Cohen, Michael Collins 0001 |
ACL (1) | 1 |
| 2014 | Spectral Unsupervised Parsing with Additive Tree MetricsabstractWe propose a spectral approach for unsupervised constituent parsing that comes with theoretical guarantees on latent structure recovery.Our approach is grammarless -we directly learn the bracketing structure of a given sentence without using a grammar model.The main algorithm is based on lifting the concept of additive tree metrics for structure learning of latent trees in the phylogenetic and machine learning communities to the case where the tree structure varies across examples.Although finding the "minimal" latent tree is NP-hard in general, for the case of projective trees we find that it can be found using bilexical parsing algorithms.Empirically, our algorithm performs favorably compared to the constituent context model of Klein and Manning (2002) without the need for careful initialization. Ankur P. Parikh, Shay B. Cohen, Eric P. Xing |
ACL (1) | 2 |
| 2014 | Latent-Variable Synchronous CFGs for Hierarchical TranslationabstractData-driven refinement of non-terminal categories has been demonstrated to be a reliable technique for improving mono-lingual parsing with PCFGs. In this pa-per, we extend these techniques to learn latent refinements of single-category syn-chronous grammars, so as to improve translation performance. We compare two estimators for this latent-variable model: one based on EM and the other is a spec-tral algorithm based on the method of mo-ments. We evaluate their performance on a Chinese–English translation task. The re-sults indicate that we can achieve signifi-cant gains over the baseline with both ap-proaches, but in particular the moments-based estimator is both faster and performs better than EM. 1 Avneesh Saluja, Chris Dyer, Shay B. Cohen |
EMNLP | 3 |
| 2014 | Spectral learning of latent-variable PCFGs: algorithms and sample complexity
Shay B. Cohen, Karl Stratos, Michael Collins 0001, Dean P. Foster, Lyle H. Ungar |
J. Mach. Learn. Res. | 1 |
| 2014 | Online Adaptor Grammars with Hybrid InferenceabstractAdaptor grammars are a flexible, powerful formalism for defining nonparametric, unsupervised models of grammar productions. This flexibility comes at the cost of expensive inference. We address the difficulty of inference through an online algorithm which uses a hybrid of Markov chain Monte Carlo and variational inference. We show that this inference strategy improves scalability without sacrificing performance on unsupervised word segmentation and topic modeling tasks. Ke Zhai 0001, Jordan L. Boyd-Graber, Shay B. Cohen |
Trans. Assoc. Comput. Linguistics | 3 |
| 2013 | The effect of non-tightness on Bayesian estimation of PCFGs
Shay B. Cohen, Mark Johnson 0001 |
ACL (1) | 1 |
| 2013 | Spectral Learning of Refinement HMMs
Karl Stratos, Alexander M. Rush, Shay B. Cohen, Michael Collins 0001 |
CoNLL | 3 |
| 2013 | Spectral Learning Algorithms for Natural Language Processing
Shay B. Cohen, Michael Collins 0001, Dean P. Foster, Karl Stratos, Lyle H. Ungar |
HLT-NAACL | 1 |
| 2013 | Approximate PCFG Parsing Using Tensor Decomposition
Shay B. Cohen, Giorgio Satta, Michael Collins 0001 |
HLT-NAACL | 1 |
| 2013 | Experiments with Spectral Learning of Latent-Variable PCFGs
Shay B. Cohen, Karl Stratos, Michael Collins 0001, Dean P. Foster, Lyle H. Ungar |
HLT-NAACL | 1 |
| 2012 | Spectral Learning of Latent-Variable PCFGs
Shay B. Cohen, Karl Stratos, Michael Collins 0001, Dean P. Foster, Lyle H. Ungar |
ACL (1) | 1 |
| 2012 | Tensor Decomposition for Fast Parsing with Latent-Variable PCFGsabstractWe describe an approach to speed-up inference with latent variable PCFGs, which have been shown to be highly effective for natural language parsing. Our approach is based on a tensor formulation recently introduced for spectral estimation of latent-variable PCFGs coupled with a tensor decomposition algorithm well-known in the multilinear algebra literature. We also describe an error bound for this approximation, which bounds the difference between the probabilities calculated by the algorithm and the true probabilities that the approximated model gives. Empirical evaluation on real-world natural language parsing data demonstrates a significant speed-up at minimal cost for parsing performance. Shay B. Cohen, Michael Collins 0001 |
NIPS | 1 |
| 2012 | Empirical Risk Minimization for Probabilistic Grammars: Sample Complexity and Hardness of LearningabstractProbabilistic grammars are generative statistical models that are useful for compositional and sequential structures. They are used ubiquitously in computational linguistics. We present a framework, reminiscent of structural risk minimization, for empirical risk minimization of probabilistic grammars using the log-loss. We derive sample complexity bounds in this framework that apply both to the supervised setting and the unsupervised setting. By making assumptions about the underlying distribution that are appropriate for natural language scenarios, we are able to derive distribution-dependent sample complexity bounds for probabilistic grammars. We also give simple algorithms for carrying out empirical risk minimization using this framework in both the supervised and unsupervised settings. In the unsupervised case, we show that the problem of minimizing empirical risk is NP-hard. We therefore suggest an approximate algorithm, similar to expectation-maximization, to minimize the empirical risk. Shay B. Cohen, Noah A. Smith |
Comput. Linguistics | 1 |
| 2011 | Unsupervised Structure Prediction with Non-Parallel Multilingual Guidance
Shay B. Cohen, Dipanjan Das 0001, Noah A. Smith |
EMNLP | 1 |
| 2011 | Exact Inference for Generative Probabilistic Non-Projective Dependency Parsing
Shay B. Cohen, Carlos Gómez-Rodríguez, Giorgio Satta |
EMNLP | 1 |
| 2011 | Products of weighted logic programsabstractAbstract Weighted logic programming, a generalization of bottom-up logic programming, is a well-suited framework for specifying dynamic programming algorithms. In this setting, proofs correspond to the algorithm's output space, such as a path through a graph or a grammatical derivation, and are given a real-valued score (often interpreted as a probability) that depends on the real weights of the base axioms used in the proof. The desired output is a function over all possible proofs, such as a sum of scores or an optimal score. We describe theproducttransformation, which can merge two weighted logic programs into a new one. The resulting program optimizes a product of proof scores from the original programs, constituting a scoring function known in machine learning as a “product of experts.” Through the addition of intuitive constraining side conditions, we show that several important dynamic programming algorithms can be derived by applyingproductto weighted logic programs corresponding tosimplerweighted logic programs. In addition, we show how the computation of Kullback–Leibler divergence, an information-theoretic measure, can be interpreted usingproduct. Shay B. Cohen, Robert J. Simmons, Noah A. Smith |
Theory Pract. Log. Program. | 1 |
| 2010 | Viterbi Training for PCFGs: Hardness Results and Competitiveness of Uniform Initialization
Shay B. Cohen, Noah A. Smith |
ACL | 1 |
| 2010 | Variational Inference for Adaptor Grammars
Shay B. Cohen, David M. Blei, Noah A. Smith |
HLT-NAACL | 1 |
| 2010 | Empirical Risk Minimization with Approximations of Probabilistic GrammarsabstractProbabilistic grammars are generative statistical models that are useful for compositional and sequential structures. We present a framework, reminiscent of structural risk minimization, for empirical risk minimization of the parameters of a fixed probabilistic grammar using the log-loss. We derive sample complexity bounds in this framework that apply both to the supervised setting and the unsupervised setting. Shay B. Cohen, Noah A. Smith |
NIPS | 1 |
| 2010 | Covariance in Unsupervised Learning of Probabilistic Grammars
Shay B. Cohen, Noah A. Smith |
J. Mach. Learn. Res. | 1 |
| 2009 | Shared Logistic Normal Distributions for Soft Parameter Tying in Unsupervised Grammar Induction
Shay B. Cohen, Noah A. Smith |
HLT-NAACL | 1 |
| 2008 | Dynamic Programming Algorithms as Products of Weighted Logic Programs
Shay B. Cohen, Robert J. Simmons, Noah A. Smith |
ICLP | 1 |
| 2008 | Logistic Normal Priors for Unsupervised Probabilistic Grammar InductionabstractWe explore a new Bayesian model for probabilistic grammars, a family of distributions over discrete structures that includes hidden Markov models and probabilistic context-free grammars. Our model extends the correlated topic model framework to probabilistic grammars, exploiting the logistic normal distribution as a prior over the grammar parameters. We derive a variational EM algorithm for that model, and then experiment with the task of unsupervised grammar induction for natural language dependency parsing. We show that our model achieves superior results over previous models that use different priors. Shay B. Cohen, Kevin Gimpel, Noah A. Smith |
NIPS | 1 |
| 2007 | Joint Morphological and Syntactic Disambiguation
Shay B. Cohen, Noah A. Smith |
EMNLP-CoNLL | 1 |
| 2007 | Feature Selection via Coalitional Game TheoryabstractWe present and study the contribution-selection algorithm (CSA), a novel algorithm for feature selection. The algorithm is based on the multiperturbation shapley analysis (MSA), a framework that relies on game theory to estimate usefulness. The algorithm iteratively estimates the usefulness of features and selects them accordingly, using either forward selection or backward elimination. It can optimize various performance measures over unseen data such as accuracy, balanced error rate, and area under receiver-operator-characteristic curve. Empirical comparison with several other existing feature selection methods shows that the backward elimination variant of CSA leads to the most accurate classification results on an array of data sets. Shay B. Cohen, Gideon Dror, Eytan Ruppin |
Neural Comput. | 1 |
| 2005 | Feature Selection Based on the Shapley Value
Shay B. Cohen, Eytan Ruppin, Gideon Dror |
IJCAI | 1 |