Derek Merck

dblp:60/1661 · DBLP profile ↗
← Back
2ranked-venue papers
0as first author
0since 2021 · last 2020
—ORCID · unresolved

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1Human-computer interaction and ubiquitous computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Language models and text generation · 100%

Topics — the 3 heaviest of 5, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › trustworthy language model › large language model reliability
factuality
0.412020
Optimizing the Factual Correctness of a Summary: A Study of Summarizing Radiology Reports · ACL 2020
Natural language and speech › Language models and text generation › text summarization › biomedical summarization
radiology report summarization
0.412020
Optimizing the Factual Correctness of a Summary: A Study of Summarizing Radiology Reports · ACL 2020
Natural language and speech › Language models and text generation
text summarization
0.412020
Optimizing the Factual Correctness of a Summary: A Study of Summarizing Radiology Reports · ACL 2020

Methods — techniques the papers use, named apart from their topics

reinforcement learning · 0.4information extraction · 0.4fact-checking · 0.4
YearPublicationVenuePosition
2020 Optimizing the Factual Correctness of a Summary: A Study of Summarizing Radiology Reports
abstract
Neural abstractive summarization models are able to generate summaries which have high overlap with human references.However, existing models are not optimized for factual correctness, a critical metric in real-world applications.In this work, we develop a general framework where we evaluate the factual correctness of a generated summary by factchecking it automatically against its reference using an information extraction module.We further propose a training strategy which optimizes a neural summarization model with a factual correctness reward via reinforcement learning.We apply the proposed method to the summarization of radiology reports, where factual correctness is a key requirement.On two separate datasets collected from hospitals, we show via both automatic and human evaluation that the proposed approach substantially improves the factual correctness and overall quality of outputs over a competitive neural summarization system, producing radiology summaries that approach the quality of humanauthored ones.Background: radiographic examination of the chest.clinical history: 80 years of age, male ... Findings: frontal radiograph of the chest demonstrates repositioning of the right atrial lead possibly into the ivc.... a right apical pneumothorax can be seen from the image.moderate right and small left pleural effusions continue.no pulmonary edema is observed.heart size is upper limits of normal. Human Summary: pneumothorax is seen. bilateral pleural effusions continue.Summary A (ROUGE-L = 0.77): no pneumothorax is observed.bilateral pleural effusions continue.Summary B (ROUGE-L = 0.44): pneumothorax is observed on radiograph.bilateral pleural effusions continue to be seen.
Yuhao Zhang 0004, Derek Merck, Emily Bao Tsai, Christopher D. Manning, Curt Langlotz
ACL2
2018 Relating Task Demand, Mental Effort and Task Difficulty with Physicians' Performance during Interactions with Electronic Health Records (EHRs)
abstract
Objective was to assess the relationship between task demand, mental effort, task difficulty, and performance during physicians’ interaction with electronic health records (EHRs). Seventeen physicians performed three EHR-based scenarios with varying task demands. Mental effort was measured using eye tracking measures via task evoked pupillary responses (TEPR), blink frequency, and gaze speed; task difficulty (or user behavior) was measured using frequent mouse click patterns and task flow; user performance was quantified using two types of omission errors: (i) omission errors with no evidence of trying to complete the task and (ii) omission error with evidence of trying but unable to complete the task. The results indicated that task demand significantly increased mental effort, but not task difficulty. Task demand, mental effort, and task difficulty all predicted performance. Specifically, there was a significant relationship between (i) task demand, TEPR and omission errors with no evidence of trying to complete the task, and (ii) blink frequency, repeated search clicks and omission error with evidence of trying but unable to complete the task. In concert, results suggest that physicians’ performance during EHR interaction was negatively affected by task demands and increase in mental effort. This highlight the need for implementation of appropriate quality assurance (QA) measures, in addition to EHR usability improvement, to minimize omission errors and improve physician’s performance. Additionally, the lack of relationship between task demand and task difficulty highlights a need for further methodological and empirical studies to advance our understanding from theory to application during physician–EHR interaction.
Prithima Mosaly, Lukasz Mazur, Fei Yu 0013, Hua Guo 0003, Derek Merck, David H. Laidlaw, Carlton Moore, Lawrence B. Marks, Javed Mostafa
Int. J. Hum. Comput. Interact.5