VLDB 2026 Research / reviewers in the wild / expert
Amy Wang
dblp:55/2777
· DBLP profile ↗
10ranked-venue papers
2as first author
3since 2021 · last 2026
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Systems, architecture and hardware · 6 · 2 first-authorArtificial intelligence and machine learning · 2 · 2 since 2021Software engineering, systems software and programming languages · 1Databases, data management, data science and information retrieval · 1Human-computer interaction and ubiquitous computing · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
2 papers |
Generative modeling · 40% Language models and text generation · 20% Trustworthy machine learning · 20% | |
| Human-computer interaction and pervasive computing
1 paper |
Human-AI interaction · 77% Health and well-being technologies · 23% | |
| Interdisciplinary, comprehensive, and emerging computing
3 papers |
Bioinformatics and computational biology · 96% Medical and health informatics · 4% | |
| Software engineering, system software, and programming languages
1 paper |
Concurrent programming · 67% Runtime systems and virtual machines · 33% | |
| Computer architecture, parallel and distributed computing, and storage systems
2 papers |
Parallel and multicore computing · 64% High-performance computing · 25% Processor architecture and microarchitecture · 11% |
Topics — the 20 heaviest of 21, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Human-AI interaction
conversational agents |
1.0 | 1 | 2026 | Towards Better Health Conversations: The Benefits of Context-seeking · CHI 2026 |
Machine learning › Generative modeling
diffusion model |
0.9 | 1 | 2025 | Fine-Tuning Discrete Diffusion Models via Reward Optimization with Applications to DNA and Protein Design · ICLR 2025 |
Machine learning › Generative modeling › diffusion model
discrete diffusion model |
0.9 | 1 | 2025 | Fine-Tuning Discrete Diffusion Models via Reward Optimization with Applications to DNA and Protein Design · ICLR 2025 |
Machine learning › Trustworthy machine learning
interpretability |
0.9 | 1 | 2025 | Concept Bottleneck Language Models For Protein Design · ICLR 2025 |
Natural language and speech › Language models and text generation › neural language model
protein language model |
0.9 | 1 | 2025 | Concept Bottleneck Language Models For Protein Design · ICLR 2025 |
Machine learning › Reinforcement learning
reward maximization |
0.9 | 1 | 2025 | Fine-Tuning Discrete Diffusion Models via Reward Optimization with Applications to DNA and Protein Design · ICLR 2025 |
Bioinformatics and computational biology
protein design |
0.5 | 2 | 2025 | Fine-Tuning Discrete Diffusion Models via Reward Optimization with Applications to DNA and Protein Design · ICLR 2025 Concept Bottleneck Language Models For Protein Design · ICLR 2025 |
Health and well-being technologies › health informatics
health information seeking |
0.3 | 1 | 2026 | Towards Better Health Conversations: The Benefits of Context-seeking · CHI 2026 |
Bioinformatics and computational biology › synthetic biology
DNA sequence design |
0.3 | 1 | 2025 | Fine-Tuning Discrete Diffusion Models via Reward Optimization with Applications to DNA and Protein Design · ICLR 2025 |
Bioinformatics and computational biology › protein design
inverse protein folding |
0.3 | 1 | 2025 | Fine-Tuning Discrete Diffusion Models via Reward Optimization with Applications to DNA and Protein Design · ICLR 2025 |
Concurrent programming › transactional memory
hardware transactional memory |
0.2 | 1 | 2015 | Software Support and Evaluation of Hardware Transactional Memory on Blue Gene/Q · IEEE Trans. Computers 2015 |
Concurrent programming
transactional memory |
0.2 | 1 | 2015 | Software Support and Evaluation of Hardware Transactional Memory on Blue Gene/Q · IEEE Trans. Computers 2015 |
Runtime systems and virtual machines › parallel runtime systems
transactional memory runtime |
0.2 | 1 | 2015 | Software Support and Evaluation of Hardware Transactional Memory on Blue Gene/Q · IEEE Trans. Computers 2015 |
High-performance computing
performance optimization at scale |
0.1 | 1 | 2012 | What scientific applications can benefit from hardware transactional memory? · SC 2012 |
Parallel and multicore computing
synchronization |
0.1 | 1 | 2012 | What scientific applications can benefit from hardware transactional memory? · SC 2012 |
Parallel and multicore computing
transactional memory |
0.1 | 1 | 2012 | What scientific applications can benefit from hardware transactional memory? · SC 2012 |
Processor architecture and microarchitecture › transactional execution
hardware transactional memory support |
0.1 | 1 | 2015 | Software Support and Evaluation of Hardware Transactional Memory on Blue Gene/Q · IEEE Trans. Computers 2015 |
Medical and health informatics › clinical informatics
clinical data management |
0.0 | 1 | 2004 | Managing Healthcare Data Hippocratically · SIGMOD Conference 2004 |
Parallel and multicore computing › parallel programming models › directive-based programming
OpenMP |
0.0 | 1 | 2012 | What scientific applications can benefit from hardware transactional memory? · SC 2012 |
Parallel and multicore computing
parallel programming models |
0.0 | 1 | 2012 | What scientific applications can benefit from hardware transactional memory? · SC 2012 |
Methods — techniques the papers use, named apart from their topics
reinforcement learning · 1.7masked language modeling · 1.7linear probing · 1.7gumbel-softmax · 1.7KL divergence · 1.7randomized controlled trial · 1.0mixed-methods study · 1.0software transactional memory · 0.4best-effort HTM · 0.4performance characterization · 0.1best practices · 0.1
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Towards Better Health Conversations: The Benefits of Context-seekingabstractNavigating health questions can be daunting in the modern information landscape. Large language models (LLMs) may provide tailored, accessible information, but also risk being inaccurate, biased or misleading. We present insights from 5 mixed-methods studies (total N=261), examining how people interact with LLMs for their own health questions. Qualitative studies revealed the importance of context-seeking in conversational AIs to elicit specific details a person may not volunteer or know to share. Context-seeking by LLMs was valued by participants, even if it meant deferring an answer for several turns. Incorporating these insights, we developed a “Wayfinding AI” to proactively solicit context. In two randomized, blinded studies, participants rated the Wayfinding AI as more helpful, relevant, and tailored to their concerns compared to a baseline AI. These results demonstrate the strong impact of proactive context-seeking on conversational dynamics, and suggest design patterns for conversational AI to help navigate health topics. Rory Sayres, Yuexing Hao, Abbi Ward, Amy Wang, Beverly Freeman, Serena Zhan, Diego Ardila, I-Ching Lee, Anna Iurchenko, Siyi Kou, Kartikeya Badola, Jimmy Hu, Bhawesh Kumar, Keith Y. Johnson, Supriya Vijay, Justin Krogue, Avinatan Hassidim, Yossi Matias, Dale R. Webster, Sunny Virmani, Yun Liu 0013, Quang Duong 0004, Mike Schaekermann |
CHI | 4 |
| 2025 | Concept Bottleneck Language Models For Protein DesignabstractWe introduce Concept Bottleneck Protein Language Models (CB-pLM), a generative masked language model with a layer where each neuron corresponds to an interpretable concept. Our architecture offers three key benefits: i) Control: We can intervene on concept values to precisely control the properties of generated proteins, achieving a 3$\times$ larger change in desired concept values compared to baselines. ii) Interpretability: A linear mapping between concept values and predicted tokens allows transparent analysis of the model's decision-making process. iii) Debugging: This transparency facilitates easy debugging of trained models. Our models achieve pre-training perplexity and downstream task performance comparable to traditional masked protein language models, demonstrating that interpretability does not compromise performance. While adaptable to any language model, we focus on masked protein language models due to their importance in drug discovery and the ability to validate our model's capabilities through real-world experiments and expert knowledge. We scale our CB-pLM from 24 million to 3 billion parameters, making them the largest Concept Bottleneck Models trained and the first capable of generative language modeling. Aya Abdelsalam Ismail, Tuomas P. Oikarinen, Amy Wang, Julius Adebayo, Samuel Stanton, Héctor Corrada Bravo, Kyunghyun Cho, Nathan C. Frey |
ICLR | 3 |
| 2025 | Fine-Tuning Discrete Diffusion Models via Reward Optimization with Applications to DNA and Protein DesignabstractRecent studies have demonstrated the strong empirical performance of diffusion models on discrete sequences (i.e., discrete diffusion models) across domains such as natural language and biological sequence generation. For example, in the protein inverse folding task, where the goal is to generate a protein sequence from a given backbone structure, conditional diffusion models have achieved impressive results in generating "natural" sequences that fold back into the original structure. However, practical design tasks often require not only modeling a conditional distribution but also optimizing specific task objectives. For instance, in the inverse folding task, we may prefer proteins with high stability. To address this, we consider the scenario where we have pre-trained discrete diffusion models that can generate "natural" sequences, as well as reward models that map sequences to task objectives. We then formulate the reward maximization problem within discrete diffusion models, analogous to reinforcement learning (RL), while minimizing the KL divergence against pre-trained diffusion models to preserve naturalness. To solve this RL problem, we propose a novel algorithm that enables direct backpropagation of rewards through entire trajectories generated by diffusion models, by making the originally non-differentiable trajectories differentiable using the Gumbel-Softmax trick. Our theoretical analysis indicates that our approach can generate sequences that are both "natural" (i.e., have a high probability under a pre-trained model) and yield high rewards. While similar tasks have been recently explored in diffusion models for continuous domains, our work addresses unique algorithmic and theoretical challenges specific to discrete diffusion models, which arise from their foundation in continuous-time Markov chains rather than Brownian motion. Finally, we demonstrate the effectiveness of our algorithm in generating DNA and protein sequences that optimize enhancer activity and protein stability, respectively, important tasks for gene therapies and protein-based therapeutics. The code is available at https://github.com/ChenyuWang-Monica/DRAKES. Chenyu Wang 0003, Masatoshi Uehara, Yichun He, Amy Wang, Avantika Lal, Tommi S. Jaakkola, Sergey Levine, Aviv Regev, Hanchen Wang 0002, Tommaso Biancalani |
ICLR | 4 |
| 2015 | Software Support and Evaluation of Hardware Transactional Memory on Blue Gene/QabstractThis paper describes an end-to-end system implementation of a transactional memory (TM) programming model on top of the hardware transactional memory (HTM) of the Blue Gene/Q machine. The TM programming model supports most C/C++ programming constructs using a best-effort HTM and the help of a complete software stack including the compiler, the kernel, and the TM runtime. An extensive evaluation of the STAMP and the RMS-TM benchmark suites on BG/Q is the first of its kind in understanding characteristics of running TM workloads on real hardware TM. The study reveals several interesting insights on the overhead and the scalability of BG/Q HTM with respect to sequential execution, coarse-grain locking, and software TM. Amy Wang, Matthew Gaudet, Peng Wu 0001, Martin Ohmacht, José Nelson Amaral, Christopher Barton, Raúl Silvera, Maged M. Michael |
IEEE Trans. Computers | 1 |
| 2012 | Evaluation of blue Gene/Q hardware support for transactional memoriesabstractThis paper describes an end-to-end system implementation of the transactional memory (TM) programming model on top of the hardware transactional memory (HTM) of the Blue Gene/Q (BG/Q) machine. The TM programming model supports most C/C++ programming constructs on top of a best-effort HTM with the help of a complete software stack including the compiler, the kernel, and the TM runtime. Amy Wang, Matthew Gaudet, Peng Wu 0001, José Nelson Amaral, Martin Ohmacht, Christopher Barton, Raúl Silvera, Maged M. Michael |
PACT | 1 |
| 2012 | What scientific applications can benefit from hardware transactional memory?abstractAchieving efficient and correct synchronization of multiple threads is a difficult and error-prone task at small scale and, as we march towards extreme scale computing, will be even more challenging when the resulting application is supposed to utilize millions of cores efficiently. Transactional Memory (TM) is a promising technique to ease the burden on the programmer, but only recently has become available on commercial hardware in the new Blue Gene/Q system and hence the real benefit for realistic applications has not been studied yet. This paper presents the first performance results of TM embedded into OpenMP on a prototype system of BG/Q and characterizes code properties that will likely lead to benefits when augmented with TM primitives. We first study the influence of thread count, environment variables and memory layout on TM performance and identify code properties that will yield performance gains with TM. Second, we evaluate the combination of OpenMP with multiple synchronization primitives on top of MPI to determine suitable task to thread ratios per node. Finally, we condense our findings into a set of best practices. These are applied to a Monte Carlo Benchmark and a Smoothed Particle Hydrodynamics method. In both cases an optimized TM version, executed with 64 threads on one node, outperforms a simple TM implementation. MCB with optimized TM yields a speedup of 27.45 over baseline. Martin Schindewolf, Barna L. Bihari, John C. Gyllenhaal, Martin Schulz 0001, Amy Wang, Wolfgang Karl |
SC | 5 |
| 2005 | Efficient SIMD Code Generation for Runtime Alignment and Length ConversionabstractWhen generating codes for today's multimedia extensions, one of the major challenges is to deal with memory alignment issues. While hand programming still yields best performing SIMD codes, it is both time consuming and error prone. Compiler technology has greatly improved, including techniques that simdize loops with misaligned accesses by automatically rearranging misaligned memory streams in registers. Current techniques are applicable to runtime alignments, but they aggressively reduce the alignment overhead only when all alignments are known at compile time. This paper presents two major enhancements to the state of the art, improving both performance and coverage. First, we propose a novel technique to simdize loops with runtime alignment nearly as efficiently as those with compile-time misalignment. Runtime alignment is pervasive in real applications because it is either part of the algorithms, or it is an artifact of the compiler's inability to extract accurate alignment information from complex applications. Second, we incorporate length conversion operations, e.g., conversions between data of different sizes, into the alignment handling framework. Length conversions are pervasive in multimedia applications where mixed integer types are often used. Supporting length conversion can greatly improve the coverage of simdizable loops. Experimental results indicate that our runtime alignment technique achieves a 19% to 32% speedup increase over prior art for a benchmark stressing the impact of misaligned data. We also demonstrate speedup factors of up to 8.11 for real benchmarks over sequential execution. Peng Wu 0001, Alexandre E. Eichenberger, Amy Wang |
CGO | 3 |
| 2005 | An integrated simdization framework using virtual vectorsabstractAutomatic simdization for multimedia extensions faces several new challenges that are not present in traditional vectorization. Some of the new issues are due to the more restrictive SIMD architectures designed for multimedia extensions. Among them are alignment constraints, lack of memory gather and scatter support, and the short and fixed-length nature of SIMD vectors. Since these constraints affect some very basic components of a program, a compiler must not only provide solid solutions to individual issues, but also take an integrated approach to address these constraints in combination.In this paper, we propose a simdization framework that addresses several orthogonal aspects of simdization, such as alignment handling, simdization of loops with mixed data lengths, and SIMD parallelism extraction from different program scopes (from basic blocks to inner loops). The novelty of this framework is its ability to facilitate interactions between different techniques based on the simple intermediate representation of virtual vectors. Measurements on a PPC970 with a VMX SIMD unit indicate speedup factors of up to 8.11 for numerical/video/communication kernels and speedup factors of up to 2.16 for benchmarks, when automatic simdization is turned on. Peng Wu 0001, Alexandre E. Eichenberger, Amy Wang |
ICS | 3 |
| 2004 | Managing Healthcare Data HippocraticallyabstractNo abstract available. Rakesh Agrawal 0001, Ameet Kini, Kristen LeFevre, Amy Wang, Yirong Xu, Diana Zhou |
SIGMOD Conference | 4 |
| 2004 | Excitation, Observation, and ELF-MD: Optimization Criteria for High Quality Test SetsabstractIn previous work, we have shown that optimizing the number of site observations leads to more defect detection. However, for increasingly difficult defects, optimizing patterns for balanced random excitation also enhances test effectiveness. We can also reduce the effect of undetected defects by choosing tests that minimize the likelihood of field failures. Jennifer Dworak, David Dorsey, Amy Wang, M. Ray Mercer |
VTS | 3 |