David Mueller

dblp:224/2296 · DBLP profile ↗
← Back
6ranked-venue papers
2as first author
1since 2021 · last 2024
—ORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 4 · 2 first-author · 1 since 2021Systems, architecture and hardware · 1Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
4 papers
Information extraction and text analysis · 22% Efficient and distributed learning · 18% Language models and text generation · 13%

Topics — the 12 heaviest of 13, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation
in-context learning
0.812024
Where does In-context Learning Happen in Large Language Models? · NeurIPS 2024
Machine learning › Learning theory › neural network theory › neural network analysis
layer-wise analysis
0.812024
Where does In-context Learning Happen in Large Language Models? · NeurIPS 2024
Computer vision › Video understanding and tracking › activity recognition
task recognition
0.812024
Where does In-context Learning Happen in Large Language Models? · NeurIPS 2024
Natural language and speech › Information extraction and text analysis
named entity recognition
0.622020
Sources of Transfer in Multilingual Named Entity Recognition · ACL 2020
Ensemble Distillation for Structured Prediction: Calibrated, Accurate, Fast - Choose Three · EMNLP (1) 2020
Natural language and speech › Information extraction and text analysis › named entity recognition
cross-lingual named entity recognition
0.412020
Sources of Transfer in Multilingual Named Entity Recognition · ACL 2020
Machine learning › Efficient and distributed learning › model compression › knowledge distillation
ensemble distillation
0.412020
Ensemble Distillation for Structured Prediction: Calibrated, Accurate, Fast - Choose Three · EMNLP (1) 2020
Machine learning › Efficient and distributed learning › model compression
knowledge distillation
0.412020
Ensemble Distillation for Structured Prediction: Calibrated, Accurate, Fast - Choose Three · EMNLP (1) 2020
Machine learning › Trustworthy machine learning › calibration
model calibration
0.412020
Ensemble Distillation for Structured Prediction: Calibrated, Accurate, Fast - Choose Three · EMNLP (1) 2020
Machine learning › Transfer learning and domain adaptation › cross-lingual transfer
multilingual transfer
0.412020
Sources of Transfer in Multilingual Named Entity Recognition · ACL 2020
Machine learning › Representation and self-supervised learning › word representation
contextual representation
0.312018
Effective Use of Context in Noisy Entity Linking · EMNLP 2018
Natural language and speech › Information extraction and text analysis
entity linking
0.312018
Effective Use of Context in Noisy Entity Linking · EMNLP 2018
Machine learning › Efficient and distributed learning
parameter-efficient fine-tuning
0.212024
Where does In-context Learning Happen in Large Language Models? · NeurIPS 2024

Methods — techniques the papers use, named apart from their topics

layer-wise context-masking · 0.8attention analysis · 0.8temperature scaling · 0.4isotonic regression · 0.4fine-tuning · 0.4ensemble distillation · 0.4sparse features · 0.3convolutional neural network · 0.3attention mechanism · 0.3
YearPublicationVenuePosition
2024 Where does In-context Learning Happen in Large Language Models?
abstract
Self-supervised large language models have demonstrated the ability to perform various tasks via in-context learning, but little is known about where the model locates the task with respect to prompt instructions and demonstration examples. In this work, we attempt to characterize the region where large language models transition from recognizing the task to performing the task. Through a series of layer-wise context-masking experiments on GPTNeo2.7B, Bloom3B, Starcoder2-7B, Llama3.1-8B, Llama3.1-8B-Instruct, on Machine Translation and Code generation, we demonstrate evidence of a "task recognition" point where the task is encoded into the input representations and attention to context is no longer necessary. Taking advantage of this redundancy results in 45% computational savings when prompting with 5 examples, and task recognition achieved at layer 14 / 32 using an example with Machine Translation. Our findings also have implications for resource and parameter efficient fine-tuning; we observe a correspondence between strong fine-tuning performance of individual LoRA layers and the task recognition layers.
Suzanna Sia, David Mueller, Kevin Duh
NeurIPS2
2020 Sources of Transfer in Multilingual Named Entity Recognition
abstract
Named-entities are inherently multilingual, and annotations in any given language may be limited.This motivates us to consider polyglot named-entity recognition (NER), where one model is trained using annotated data drawn from more than one language.However, a straightforward implementation of this simple idea does not always work in practice: naive training of NER models using annotated data drawn from multiple languages consistently underperforms models trained on monolingual data alone, despite having access to more training data.The starting point of this paper is a simple solution to this problem, in which polyglot models are fine-tuned on monolingual data to consistently and significantly outperform their monolingual counterparts.To explain this phenomena, we explore the sources of multilingual transfer in polyglot NER models and examine the weight structure of polyglot models compared to their monolingual counterparts.We find that polyglot models efficiently share many parameters across languages and that fine-tuning may utilize a large number of those parameters.
David Mueller, Nicholas Andrews, Mark Dredze
ACL1
2020 Ensemble Distillation for Structured Prediction: Calibrated, Accurate, Fast - Choose Three
abstract
Modern neural networks do not always produce well-calibrated predictions, even when trained with a proper scoring function such as cross-entropy.In classification settings, simple methods such as isotonic regression or temperature scaling may be used in conjunction with a held-out dataset to calibrate model outputs.However, extending these methods to structured prediction is not always straightforward or effective; furthermore, a held-out calibration set may not always be available.In this paper, we study ensemble distillation as a general framework for producing wellcalibrated structured prediction models while avoiding the prohibitive inference-time cost of ensembles.We validate this framework on two tasks: named-entity recognition and machine translation.We find that, across both tasks, ensemble distillation produces models which retain much of, and occasionally improve upon, the performance and calibration benefits of ensembles, while only requiring a single model during test-time.
Steven Reich, David Mueller, Nicholas Andrews
EMNLP (1)2
2020 A2Cloud-RF: A random forest based statistical framework to guide resource selection for high-performance scientific computing on the cloud
abstract
Summary This article proposes a random‐forest based A2Cloud framework to match scientific applications with Cloud providers and their instances for high performance. The framework leverages four engines for this task: PERF engine, Cloud trace engine, A2Cloud‐ext engine, and the random forest classifier (RFC) engine. The PERF engine profiles the application to obtain performance characteristics, including the number of single‐precision (SP) floating‐point operations (FLOPs), double‐precision (DP) FLOPs, x87 operations, memory accesses, and disk accesses. The Cloud trace engine obtains the corresponding performance characteristics of the selected Cloud instances including: SP floating point operations per second (FLOPS), DP FLOPS, x87 operations per second, memory bandwidth, and disk bandwidth. The A2Cloud‐ext engine uses the application and Cloud instance characteristics to generate objective scores that represent the application‐to‐Cloud match. The RFC engine uses these objective scores to generate two types of random forests to assist users with rapid analysis: application‐specific random forests (ARF) and application‐class based random forests. The ARF consider only the input application's characteristics to generate a random forest and provide numerical ratings to the selected Cloud instances. To generate the application‐class based random forests, the RFC engine downloads the application profiles and scores of previously tested applications that perform similar to the input application. Using these data, the RFC engine creates a random forest for instance recommendation. We exhaustively test this framework using eight real‐world applications across 12 instances from different Cloud providers. Our tests show significant statistical agreement between the instance ratings given by the framework and the ratings obtained via actual Cloud executions.
David Samuel, Syeduzzaman Khan, Cody J. Balos, Zachariah Abuelhaj, Anthony D. Dutoi, Chadi Kari, David Mueller, Vivek K. Pallipuram
Concurr. Comput. Pract. Exp.7
2018 A2Cloud: An Analytical Model for Application-to-Cloud Matching to Empower Scientific Computing
abstract
We present an analytical model that matches scientific applications to effective Cloud instances for high application performance. The model constructs two vectors namely, the application vector and the Cloud vector. The application vector consists of application performance components such as the number of single-precision (SP) floating-point operations (FLOPs) and double-precision (DP) FLOPs, main memory accesses, and disk accesses. The Cloud vector comprises corresponding Cloud instance performance components such as the benchmarked SP and DP floating-point operations per second (FLOPS), memory bandwidth, and disk bandwidth. The model performs an inner product of the two vectors to produce an Application-to-Cloud (A2Cloud) score, which quantifies the application-to-Cloud match. We encapsulate the A2Cloud model in a user-friendly A2Cloud framework that inputs a test application and a target Cloud instance, profiles them, and executes the A2Cloud model to generate the A2Cloud score. We demonstrate the model by conducting 162 application executions across nine Cloud instances. Our tests yield an average A2Cloud matching rate of 6 for every 9 application-instance pairs with a mean absolute difference of ±1.08 ranks.
Cody J. Balos, David de la Vega, Zachariah Abuelhaj, Chadi Kari, David Mueller, Vivek K. Pallipuram
IEEE CLOUD5
2018 Effective Use of Context in Noisy Entity Linking
abstract
To disambiguate between closely related concepts, entity linking systems need to effectively distill cues from a mention's textual context.We investigate several techniques for using these cues in the task of noisy entity linking on short texts.Our starting point is a stateof-the-art attention-based model from prior work; while this model's attention typically identifies context that is topically relevant, it fails to identify some of the most indicative context words, especially those exhibiting lexical overlap with the true title.Augmenting the model with convolutional networks over characters still leaves it largely unable to pick up on these cues compared to sparse features that target them directly, indicating that automatically learning how to identify relevant character-level context features is a hard problem.Armed with these sparse features, our final system 1 outperforms past work on the WikilinksNED test set by 2.8% absolute.
David Mueller, Greg Durrett
EMNLP1