Terrence Neumann

dblp:172/9370 · DBLP profile ↗
← Back
2ranked-venue papers
1as first author
1since 2021 · last 2026
—ORCID · none

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 1 · 1 first-author · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 first-author · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1

Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.

Artificial intelligence
1 paper
Language models and text generation · 77% Trustworthy machine learning · 23%

Topics — the 1 heaviest of 2, each with the papers that count most for it

TopicWeightPapersLastEvidence papers
Natural language and speech › Language models and text generation › LLM agents
LLM-based simulation
1.012026
Should You Use LLMs to Simulate Opinions? Quality Checks for Early-Stage Deliberation · AAAI 2026

Methods — techniques the papers use, named apart from their topics

prompt engineering · 1.0in-context learning · 1.0fine-tuning · 1.0
YearPublicationVenuePosition
2026 Should You Use LLMs to Simulate Opinions? Quality Checks for Early-Stage Deliberation
abstract
The emergent capabilities of large language models (LLMs) have prompted interest in using them as surrogates for human subjects in opinion surveys. However, prior evaluations of LLM-based opinion simulation have relied heavily on costly, domain-specific survey data, and mixed empirical results leave their reliability in question. To enable cost-effective, early-stage evaluation, we introduce a quality control assessment designed to test the viability of LLM-simulated opinions on Likert-scale tasks without requiring large-scale human data for validation. This assessment comprises two key tests: logical consistency and alignment with stakeholder expectations, offering a low-cost, domain-adaptable validation tool. We apply our quality control assessment to an opinion simulation task relevant to AI-assisted content moderation and fact-checking workflows---a socially impactful use case---and evaluate nine LLMs using a baseline prompt engineering method (backstory prompting), as well as fine-tuning and in-context learning variants. None of the models or methods pass the full assessment, revealing several failure modes. We conclude with a discussion of the risk management implications and release TopicMisinfo, a benchmark dataset with paired human and LLM annotations simulated by various models and approaches, to support future research.
Terrence Neumann, Maria De-Arteaga, Sina Fazelpour
AAAI1
2015 A Novel Scoring Based Distributed Protein Docking Application to Improve Enrichment
abstract
Molecular docking is a computational technique which predicts the binding energy and the preferred binding mode of a ligand to a protein target. Virtual screening is a tool which uses docking to investigate large chemical libraries to identify ligands that bind favorably to a protein target. We have developed a novel scoring based distributed protein docking application to improve enrichment in virtual screening. The application addresses the issue of time and cost of screening in contrast to conventional systematic parallel virtual screening methods in two ways. Firstly, it automates the process of creating and launching multiple independent dockings on a high performance computing cluster. Secondly, it uses a Nȧi̇ve Bayes scoring function to calculate binding energy of un-docked ligands to identify and preferentially dock (Autodock predicted) better binders. The application was tested on four proteins using a library of 10,573 ligands. In all the experiments, (i). 200 of the 1,000 best binders are identified after docking only ~14 percent of the chemical library, (ii). 9 or 10 best-binders are identified after docking only ~19 percent of the chemical library, and (iii). no significant enrichment is observed after docking ~70 percent of the chemical library. The results show significant increase in enrichment of potential drug leads in early rounds of virtual screening.
Prachi Pradeep, Craig A. Struble, Terrence Neumann, Daniel S. Sem, Stephen J. Merrill
IEEE ACM Trans. Comput. Biol. Bioinform.3