Demonstration venue · read-only. Every page can be browsed; the buttons that would change it are switched off. Create an account to run TaxoReview on your own data.

Matthew R. Gormley

dblp:116/0475 · also Matthew Gormley 0001 · DBLP profile ↗
← Back
25ranked-venue papers
5as first author
8since 2021 · last 2025
0000-0003-4207-5986ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 24 · 5 first-author · 7 since 2021Human-computer interaction and ubiquitous computing · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
14 papers
Information extraction and text analysis · 30% Language models and text generation · 14% Trustworthy machine learning · 12%
Interdisciplinary, comprehensive, and emerging computing
2 papers
Medical and health informatics · 100%
Databases, data mining, and information retrieval
1 paper
Information retrieval · 100%
Software engineering, system software, and programming languages
1 paper
Programming languages and type systems · 100%

Topics — the 28 heaviest of 32, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Information extraction and text analysis › text mining › biomedical text mining
clinical text analysis
0.712023
MDACE: MIMIC Documents Annotated with Code Evidence · ACL (1) 2023
Natural language and speech › Language models and text generation › language modeling › long-context language modeling › context utilization › long-context modeling
long-context language model
0.712023
Unlimiformer: Long-Range Transformers with Unlimited Length Input · NeurIPS 2023
Medical and health informatics › clinical text processing
clinical text annotation
0.712023
MDACE: MIMIC Documents Annotated with Code Evidence · ACL (1) 2023
Information retrieval › similarity search
nearest neighbor search
0.712023
Unlimiformer: Long-Range Transformers with Unlimited Length Input · NeurIPS 2023
Machine learning › Trustworthy machine learning › robustness
out-of-distribution detection
0.612022
AdaFocal: Calibration-aware Adaptive Focal Loss · NeurIPS 2022
Machine learning › Trustworthy machine learning
uncertainty and calibration
0.612022
AdaFocal: Calibration-aware Adaptive Focal Loss · NeurIPS 2022
Natural language and speech › Information extraction and text analysis
semantic role labeling
0.522017
Semantic Proto-Role Labeling · AAAI 2017
Low-Resource Semantic Role Labeling · ACL (1) 2014
Machine learning › Probabilistic and Bayesian machine learning › structured models › graphical models › structure learning
graphical model learning
0.412020
Training for Gibbs Sampling on Conditional Random Fields with Neural Scoring Factors · EMNLP (1) 2020
Natural language and speech › Information extraction and text analysis
named entity recognition
0.412020
Training for Gibbs Sampling on Conditional Random Fields with Neural Scoring Factors · EMNLP (1) 2020
Machine learning › Probabilistic and Bayesian machine learning › structured models › latent variable model › latent generative model
noisy channel model
0.412020
Phonetic and Visual Priors for Decipherment of Informal Romanization · ACL 2020
Natural language and speech › Machine translation
transliteration
0.412020
Phonetic and Visual Priors for Decipherment of Informal Romanization · ACL 2020
Natural language and speech › Machine translation
bilingual lexicon induction
0.412019
Bilingual Lexicon Induction with Semi-supervision in Non-Isometric Embedding Spaces · ACL (1) 2019
Machine learning › Efficient and distributed learning › automated machine learning
neural architecture search
0.412019
Towards modular and programmable architecture search · NeurIPS 2019
Machine learning › Efficient and distributed learning › automated machine learning › neural architecture search
search space design
0.412019
Towards modular and programmable architecture search · NeurIPS 2019
Machine learning › Representation and self-supervised learning › word representation
word embedding alignment
0.412019
Bilingual Lexicon Induction with Semi-supervision in Non-Isometric Embedding Spaces · ACL (1) 2019
Programming languages and type systems
domain-specific languages
0.412019
Towards modular and programmable architecture search · NeurIPS 2019
Natural language and speech › Language models and text generation › decoding › decoding strategy
beam search
0.312018
Learning Beam Search Policies via Imitation Learning · NeurIPS 2018
Machine learning › Reinforcement learning
imitation learning
0.312018
Learning Beam Search Policies via Imitation Learning · NeurIPS 2018
Natural language and speech › Information extraction and text analysis › morphological analysis
morphological tagging
0.312018
Neural Factor Graph Models for Cross-lingual Morphological Tagging · ACL (1) 2018
Machine learning › Learning paradigms
multi-label classification
0.312017
Semantic Proto-Role Labeling · AAAI 2017
Natural language and speech › Information extraction and text analysis
relation extraction
0.212015
Improved Relation Extraction with Feature-Rich Compositional Embedding Models · EMNLP 2015
Natural language and speech › Language models and text generation
text summarization
0.212023
Unlimiformer: Long-Range Transformers with Unlimited Length Input · NeurIPS 2023
Computer vision › Image recognition and object detection
image classification
0.212022
AdaFocal: Calibration-aware Adaptive Focal Loss · NeurIPS 2022
Natural language and speech › Speech recognition and synthesis
phonetic modeling
0.112020
Phonetic and Visual Priors for Decipherment of Informal Romanization · ACL 2020
Natural language and speech › Information extraction and text analysis
sequence labeling
0.112020
Training for Gibbs Sampling on Conditional Random Fields with Neural Scoring Factors · EMNLP (1) 2020
Machine learning › Transfer learning and domain adaptation
cross-lingual transfer
0.112018
Neural Factor Graph Models for Cross-lingual Morphological Tagging · ACL (1) 2018
Natural language and speech › Language models and text generation › grammar induction
unsupervised grammar induction
0.112014
Low-Resource Semantic Role Labeling · ACL (1) 2014
Machine learning › Probabilistic and Bayesian machine learning › structured models
latent variable model
0.012013
Nonconvex Global Optimization for Latent-Variable Models · ACL (1) 2013

Methods — techniques the papers use, named apart from their topics

focal loss · 1.6k-nearest-neighbor index · 1.3cross-attention offloading · 1.3annotation · 1.3squeeze-and-excitation · 1.0convolutional attention network · 1.0adaptive loss weighting · 0.6weighted finite-state transducer · 0.4unsupervised learning · 0.4gibbs sampling · 0.4hyperparameter optimization · 0.4
YearPublicationVenuePosition
2025 Predicting the Past: Estimating Historical Appraisals with OCR and Machine Learning
abstract
Despite well-documented consequences of the U.S. government's 1930s housing policies on racial wealth disparities, scholars have struggled to quantify its precise financial effects due to the inaccessibility of historical property appraisal records. Many counties still store these records in physical formats, making large-scale quantitative analysis difficult. We present an approach scholars can use to digitize historical housing assessment data, applying it to build and release a dataset for one county. Starting from publicly available scanned documents, we manually annotated property cards for over 12,000 properties to train and validate our methods. We use OCR to label data for an additional 50,000 properties, based on our two-stage approach combining classical computer vision techniques with deep learning-based OCR. For cases where OCR cannot be applied, such as when scanned documents are not available, we show how a regression model based on building feature data can estimate the historical values, and test the generalizability of this model to other counties. With these cost-effective tools, scholars, community activists, and policy makers can better analyze and understand the historical impacts of redlining.
Mihir Bhaskar, Jun Tao Luo, Zihan Geng, Asmita Hajra, Junia Howell, Matthew R. Gormley
COMPASS6
2025 In-Context Learning with Long-Context Models: An In-Depth Exploration
abstract
Amanda Bertsch, Maor Ivgi, Emily Xiao, Uri Alon, Jonathan Berant, Matthew R. Gormley, Graham Neubig. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Amanda Bertsch, Maor Ivgi, Emily Xiao, Uri Alon 0002, Jonathan Berant, Matthew R. Gormley, Graham Neubig
NAACL (Long Papers)6
2025 Larger than Life In-Class Demonstrations for Introductory Machine Learning
Henry Chai, Matthew R. Gormley
SIGCSE (1)2
2023 MDACE: MIMIC Documents Annotated with Code Evidence
abstract
Hua Cheng, Rana Jafari, April Russell, Russell Klopfer, Edmond Lu, Benjamin Striner, Matthew Gormley. Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2023.
Rana Jafari, April Russell, Russell Klopfer, Edmond Lu, Benjamin Striner, Matthew R. Gormley
ACL (1)7
2023 Unlimiformer: Long-Range Transformers with Unlimited Length Input
abstract
Since the proposal of transformers, these models have been limited to bounded input lengths, because of their need to attend to every token in the input. In this work, we propose Unlimiformer: a general approach that wraps any existing pretrained encoder-decoder transformer, and offloads the cross-attention computation to a single $k$-nearest-neighbor ($k$NN) index, while the returned $k$NN distances are the attention dot-product scores. This $k$NN index can be kept on either the GPU or CPU memory and queried in sub-linear time; this way, we can index practically unlimited input sequences, while every attention head in every decoder layer retrieves its top-$k$ keys, instead of attending to every key. We evaluate Unlimiformer on several long-document and book-summarization benchmarks, showing that it can process even **500k** token-long inputs from the BookSum dataset, without any input truncation at test time. We demonstrate that Unlimiformer improves pretrained models such as BART and Longformer by extending them to unlimited inputs without additional learned weights and without modifying their code. Our code and models are publicly available at https://github.com/abertsch72/unlimiformer , and support LLaMA-2 as well.
Amanda Bertsch, Uri Alon 0002, Graham Neubig, Matthew R. Gormley
NeurIPS4
2022 AdaFocal: Calibration-aware Adaptive Focal Loss
abstract
Much recent work has been devoted to the problem of ensuring that a neural network's confidence scores match the true probability of being correct, i.e. the calibration problem. Of note, it was found that training with focal loss leads to better calibration than cross-entropy while achieving similar level of accuracy \cite{mukhoti2020}. This success stems from focal loss regularizing the entropy of the model's prediction (controlled by the parameter $\gamma$), thereby reining in the model's overconfidence. Further improvement is expected if $\gamma$ is selected independently for each training sample (Sample-Dependent Focal Loss (FLSD-53) \cite{mukhoti2020}). However, FLSD-53 is based on heuristics and does not generalize well. In this paper, we propose a calibration-aware adaptive focal loss called AdaFocal that utilizes the calibration properties of focal (and inverse-focal) loss and adaptively modifies $\gamma_t$ for different groups of samples based on $\gamma_{t-1}$ from the previous step and the knowledge of model's under/over-confidence on the validation set. We evaluate AdaFocal on various image recognition and one NLP task, covering a wide variety of network architectures, to confirm the improvement in calibration while achieving similar levels of accuracy. Additionally, we show that models trained with AdaFocal achieve a significant boost in out-of-distribution detection.
Thomas Schaaf, Matthew R. Gormley
NeurIPS3
2021 Effective Convolutional Attention Network for Multi-label Clinical Document Classification
abstract
Multi-label document classification (MLDC) problems can be challenging, especially for long documents with a large label set and a long-tail distribution over labels.In this paper, we present an effective convolutional attention network for the MLDC problem with a focus on medical code prediction from clinical documents.Our innovations are three-fold: (1) we utilize a deep convolution-based encoder with the squeeze-and-excitation networks and residual networks to aggregate the information across the document and learn meaningful document representations that cover different ranges of texts; (2) we explore multilayer and sum-pooling attention to extract the most informative features from these multiscale representations; (3) we combine binary cross entropy loss and focal loss to improve performance for rare labels.We focus our evaluation study on MIMIC-III, a widely used dataset in the medical domain.Our models outperform prior work on medical coding and achieve new state-of-the-art results on multiple metrics.We also demonstrate the language independent nature of our approach by applying it to two non-English datasets.Our model outperforms prior best model and a multilingual Transformer model by a substantial margin.
Russell Klopfer, Matthew R. Gormley, Thomas Schaaf
EMNLP (1)4
2021 Limitations of Autoregressive Models and Their Alternatives
abstract
Chu-Cheng Lin, Aaron Jaech, Xin Li, Matthew R. Gormley, Jason Eisner. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021.
Chu-Cheng Lin, Aaron Jaech, Matthew R. Gormley, Jason Eisner
NAACL-HLT4
2020 Phonetic and Visual Priors for Decipherment of Informal Romanization
abstract
Informal romanization is an idiosyncratic process used by humans in informal digital communication to encode non-Latin script languages into Latin character sets found on common keyboards. Character substitution choices differ between users but have been shown to be governed by the same main principles observed across a variety of languages—namely, character pairs are often associated through phonetic or visual similarity. We propose a noisy-channel WFST cascade model for deciphering the original non-Latin script from observed romanized text in an unsupervised fashion. We train our model directly on romanized data from two languages: Egyptian Arabic and Russian. We demonstrate that adding inductive bias through phonetic and visual priors on character mappings substantially improves the model’s performance on both languages, yielding results much closer to the supervised skyline. Finally, we introduce a new dataset of romanized Russian, collected from a Russian social network website and partially annotated for our experiments.
Maria Ryskina, Matthew R. Gormley, Taylor Berg-Kirkpatrick
ACL2
2020 Training for Gibbs Sampling on Conditional Random Fields with Neural Scoring Factors
abstract
Most recent improvements in NLP come from changes to the neural network architectures modeling the text input.Yet, state-of-theart models often rely on simple approaches to model the label space, e.g.bigram Conditional Random Fields (CRFs) in sequence tagging.More expressive graphical models are rarely used due to their prohibitive computational cost.In this work, we present an approach for efficiently training and decoding hybrids of graphical models and neural networks based on Gibbs sampling.Our approach is the natural adaptation of SampleRank (Wick et al., 2011) to neural models, and is widely applicable to tasks beyond sequence tagging.We apply our approach to named entity recognition and present a neural skipchain CRF model, for which exact inference is impractical.The skip-chain model improves over a strong baseline on three languages from CoNLL-02/03.We obtain new state-of-the-art results on Dutch. 1
Sida Gao, Matthew R. Gormley
EMNLP (1)2
2019 Bilingual Lexicon Induction with Semi-supervision in Non-Isometric Embedding Spaces
abstract
Recent work on bilingual lexicon induction (BLI) has frequently depended either on aligned bilingual lexicons or on distribution matching, often with an assumption about the isometry of the two spaces.We propose a technique to quantitatively estimate this assumption of the isometry between two embedding spaces and empirically show that this assumption weakens as the languages in question become increasingly etymologically distant.We then propose Bilingual Lexicon Induction with Semi-Supervision (BLISS) -a semi-supervised approach that relaxes the isometric assumption while leveraging both limited aligned bilingual lexicons and a larger set of unaligned word embeddings, as well as a novel hubness filtering technique.Our proposed method obtains state of the art results on 15 of 18 language pairs on the MUSE dataset, and does particularly well when the embedding spaces don't appear to be isometric.In addition, we also show that adding supervision stabilizes the learning procedure, and is effective even with minimal supervision.⇤
Barun Patra, Joel Ruben Antony Moniz, Sarthak Garg, Matthew R. Gormley, Graham Neubig
ACL (1)4
2019 Towards modular and programmable architecture search
abstract
Neural architecture search methods are able to find high performance deep learning architectures with minimal effort from an expert. However, current systems focus on specific use-cases (e.g. convolutional image classifiers and recurrent language models), making them unsuitable for general use-cases that an expert might wish to write. Hyperparameter optimization systems are general-purpose but lack the constructs needed for easy application to architecture search. In this work, we propose a formal language for encoding search spaces over general computational graphs. The language constructs allow us to write modular, composable, and reusable search space encodings and to reason about search space design. We use our language to encode search spaces from the architecture search literature. The language allows us to decouple the implementations of the search space and the search algorithm, allowing us to expose search spaces to search algorithms through a consistent interface. Our experiments show the ease with which we can experiment with different combinations of search spaces and search algorithms without having to implement each combination from scratch. We release an implementation of our language with this paper.
Renato Negrinho, Matthew R. Gormley, Geoffrey J. Gordon, Darshan Patil, Nghia Le
NeurIPS2
2018 Neural Factor Graph Models for Cross-lingual Morphological Tagging
abstract
Morphological analysis involves predicting the syntactic traits of a word (e.g.{POS: Noun, Case: Acc, Gender: Fem}).Previous work in morphological tagging improves performance for low-resource languages (LRLs) through cross-lingual training with a high-resource language (HRL) from the same family, but is limited by the strict-often false-assumption that tag sets exactly overlap between the HRL and LRL.In this paper we propose a method for cross-lingual morphological tagging that aims to improve information sharing between languages by relaxing this assumption.The proposed model uses factorial conditional random fields with neural network potentials, making it possible to (1) utilize the expressive power of neural network representations to smooth over superficial differences in the surface forms, (2) model pairwise and transitive relationships between tags, and (3) accurately generate tag sets that are unseen or rare in the training data.Experiments on four languages from the Universal Dependencies Treebank (Nivre et al., 2017) demonstrate superior tagging accuracies over existing cross-lingual approaches. 1
Chaitanya Malaviya, Matthew R. Gormley, Graham Neubig
ACL (1)2
2018 Learning Beam Search Policies via Imitation Learning
abstract
Beam search is widely used for approximate decoding in structured prediction problems. Models often use a beam at test time but ignore its existence at train time, and therefore do not explicitly learn how to use the beam. We develop an unifying meta-algorithm for learning beam search policies using imitation learning. In our setting, the beam is part of the model and not just an artifact of approximate decoding. Our meta-algorithm captures existing learning algorithms and suggests new ones. It also lets us show novel no-regret guarantees for learning beam search policies.
Renato Negrinho, Matthew R. Gormley, Geoffrey J. Gordon
NeurIPS2
2017 Semantic Proto-Role Labeling
abstract
The semantic function tags of Bonial, Stowe, and Palmer (2013) and the ordinal, multi-property annotations of Reisinger et al. (2015) draw inspiration from Ddowty's semantic proto-role theory. We approach proto-role labeling as a multi-label classification problem and establish strong results for the task by adapting a successful model of traditional semantic role labeling. We achieve a proto-role micro-averaged F1 of 81.7 using gold syntax and explore joint and conditional models of proto-roles and categorical roles. In comparing the effect of Bonial, Stowe, and Palmer's tags to PropBank ArgN-style role labels, we are surprised that neither annotations greatly improve proto-role prediction; however, we observe that ArgN models benefit much from observed syntax and from observed or modeled proto-roles while our models of the semantic function tags do not.
Adam R. Teichert, Adam Poliak, Benjamin Van Durme, Matthew R. Gormley
AAAI4
2016 Embedding Lexical Features via Low-Rank Tensors
abstract
Mo Yu, Mark Dredze, Raman Arora, Matthew R. Gormley. Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2016.
Mo Yu, Mark Dredze, Raman Arora, Matthew R. Gormley
HLT-NAACL4
2015 Improved Relation Extraction with Feature-Rich Compositional Embedding Models
abstract
Compositional embedding models build a representation (or embedding) for a linguistic structure based on its component word embeddings.We propose a Feature-rich Compositional Embedding Model (FCM) for relation extraction that is expressive, generalizes to new domains, and is easy-to-implement.The key idea is to combine both (unlexicalized) handcrafted features with learned word embeddings.The model is able to directly tackle the difficulties met by traditional compositional embeddings models, such as handling arbitrary types of sentence annotations and utilizing global information for composition.We test the proposed model on two relation extraction tasks, and demonstrate that our model outperforms both previous compositional models and traditional feature rich models on the ACE 2005 relation extraction task, and the SemEval 2010 relation classification task.The combination of our model and a loglinear classifier with hand-crafted features gives state-of-the-art results.We made our implementation available for general use 1 .
Matthew R. Gormley, Mo Yu, Mark Dredze
EMNLP1
2015 A Concrete Chinese NLP Pipeline
abstract
Nanyun Peng, Francis Ferraro, Mo Yu, Nicholas Andrews, Jay DeYoung, Max Thomas, Matthew R. Gormley, Travis Wolfe, Craig Harman, Benjamin Van Durme, Mark Dredze. Proceedings of the 2015 Conference of the North American Chapter of the Association for Computational Linguistics: Demonstrations. 2015.
Nanyun Peng 0001, Francis Ferraro, Mo Yu, Nicholas Andrews, Jay DeYoung, Max Thomas, Matthew R. Gormley, Travis Wolfe, Craig Harman, Benjamin Van Durme, Mark Dredze
HLT-NAACL7
2015 Combining Word Embeddings and Feature Embeddings for Fine-grained Relation Extraction
abstract
Compositional embedding models build a rep-resentation for a linguistic structure based on its component word embeddings. While re-cent work has combined these word embed-dings with hand crafted features for improved performance, it was restricted to a small num-ber of features due to model complexity, thus limiting its applicability. We propose a new model that conjoins features and word em-beddings while maintaing a small number of parameters by learning feature embeddings jointly with the parameters of a compositional model. The result is a method that can scale to more features and more labels, while avoiding overfitting. We demonstrate that our model at-tains state-of-the-art results on ACE and ERE fine-grained relation extraction. 1
Mo Yu, Matthew R. Gormley, Mark Dredze
HLT-NAACL2
2015 Approximation-Aware Dependency Parsing by Belief Propagation
abstract
We show how to train the fast dependency parser of Smith and Eisner (2008) for improved accuracy. This parser can consider higher-order interactions among edges while retaining O( n3) runtime. It outputs the parse with maximum expected recall—but for speed, this expectation is taken under a posterior distribution that is constructed only approximately, using loopy belief propagation through structured factors. We show how to adjust the model parameters to compensate for the errors introduced by this approximation, by following the gradient of the actual loss on training data. We find this gradient by back-propagation. That is, we treat the entire parser (approximations and all) as a differentiable circuit, as others have done for loopy CRFs (Domke, 2010; Stoyanov et al., 2011; Domke, 2011; Stoyanov and Eisner, 2012). The resulting parser obtains higher accuracy with fewer iterations of belief propagation than one trained by conditional log-likelihood.
Matthew R. Gormley, Mark Dredze, Jason Eisner
Trans. Assoc. Comput. Linguistics1
2014 Low-Resource Semantic Role Labeling
abstract
We explore the extent to which highresource manual annotations such as treebanks are necessary for the task of semantic role labeling (SRL).We examine how performance changes without syntactic supervision, comparing both joint and pipelined methods to induce latent syntax.This work highlights a new application of unsupervised grammar induction and demonstrates several approaches to SRL in the absence of supervised syntax.Our best models obtain competitive results in the high-resource setting and state-ofthe-art results in the low resource setting, reaching 72.48% F1 averaged across languages.We release our code for this work along with a larger toolkit for specifying arbitrary graphical structure.1
Matthew R. Gormley, Margaret Mitchell, Benjamin Van Durme, Mark Dredze
ACL (1)1
2013 Nonconvex Global Optimization for Latent-Variable Models
Matthew R. Gormley, Jason Eisner
ACL (1)1
2013 Topic Models and Metadata for Visualizing Text Corpora
Justin Snyder, Rebecca Knowles, Mark Dredze, Matthew R. Gormley, Travis Wolfe
HLT-NAACL4
2012 Shared Components Topic Models
Matthew R. Gormley, Mark Dredze, Benjamin Van Durme, Jason Eisner
HLT-NAACL1
2012 Entity Clustering Across Languages
Spence Green, Nicholas Andrews, Matthew R. Gormley, Mark Dredze, Christopher D. Manning
HLT-NAACL3