VLDB 2026 Research / reviewers in the wild / expert
Kevin He
dblp:147/5857
· DBLP profile ↗
16ranked-venue papers
5as first author
13since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 8 · 4 first-author · 6 since 2021Theory of computation · 8 · 4 first-author · 6 since 2021Systems, architecture and hardware · 5 · 5 since 2021Applied, interdisciplinary, general and emerging computing · 3 · 1 first-author · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | DEMOTIC: A Differentiable Sampler for Multi-Level Digital CircuitsabstractEfficient sampling of satisfying formulas for circuit satisfiability (CircuitSAT), a well-known NP-complete problem, is essential in modern front-end applications for thorough testing and verification of digital circuits. Generating such samples is a hard computational problem due to the inherent complexity of digital circuits, size of the search space, and resource constraints involved in the process. Addressing these challenges has prompted the development of specialized algorithms that heavily rely on heuristics. However, these heuristic-based approaches frequently encounter scalability issues when tasked with sampling from a larger number of solutions, primarily due to their sequential nature. Different from such heuristic algorithms, we propose a novel differentiable sampler for multi-level digital circuits, called Demotic, that utilizes gradient descent (GD) to solve the CircuitSAT problem and obtain a wide range of valid and distinct solutions. Demotic leverages the circuit structure of the problem instance to learn valid solutions using GD by re-framing the CircuitSAT problem as a supervised multi-output regression task. This differentiable approach allows bit-wise operations to be performed independently on each element of a tensor, enabling parallel execution of learning operations, and accordingly, GPU-accelerated sampling with significant runtime improvements compared to state-of-the-art heuristic samplers. We demonstrate the superior runtime performance of Demotic in the sampling task across various CircuitSAT instances from the ISCAS-85 benchmark suite. Specifically, Demotic outperforms the state-of-the-art sampler by more than two orders of magnitude in most cases. Arash Ardakani, Kevin He, Qijing Huang 0001, Vighnesh M. Iyer, Suhong Moon, John Wawrzynek |
ASP-DAC | 3 |
| 2025 | High-Throughput SAT SamplingabstractIn this work, we present a novel technique for GPU-accelerated Boolean satisfiability (SAT) sampling. Unlike conventional sampling algorithms that directly operate on conjunctive normal form (CNF), our method transforms the logical constraints of SAT problems by factoring their CNF representations into simplified multilevel, multi-output Boolean functions. It then leverages gradient-based optimization to guide the search for a diverse set of valid solutions. Our method operates directly on the circuit structure of refactored SAT instances, reinterpreting the SAT problem as a supervised multi-output regression task. This differentiable technique enables independent bit-wise operations on each tensor element, allowing parallel execution of learning processes. As a result, we achieve GPU-accelerated sampling with significant runtime improvements ranging from 33.6x to 523.6x over state-of-the-art heuristic samplers. We demonstrate the superior performance of our sampling method through an extensive evaluation on 60 instances from a public domain benchmark suite utilized in previous studies. Arash Ardakani, Kevin He, Qijing Huang 0001, John Wawrzynek |
DATE | 3 |
| 2025 | Human Misperception of Generative-AI Alignment: A Laboratory ExperimentabstractWe conduct an incentivized laboratory experiment to study people's perception of generative artificial intelligence (GenAI) alignment in the context of economic decision-making. Using a panel of economic problems spanning the domains of risk, time preference, social preference, and strategic interactions, we ask human subjects to make choices for themselves and to predict the choices made by GenAI on behalf of a human user. These problems confront agents with trade-offs (e.g., higher payoff vs. earlier payoff, efficiency vs. equity, riskier but potentially higher rewards vs. safer but lower rewards) and the optimal choices depend on the agent's preferences. We find that people overestimate the degree to which GenAI choices are aligned with human preferences in general (anthropomorphic projection), and with their personal references in particular (self projection). On average, human subjects' predictions about GenAI's choices in every decision environment are much closer to the average human-subject choice than to the average GenAI choice. At the individual level, human subjects' predictions about GenAI's choices in a given environment are highly correlated with their own choices in the same environment. Kevin He, Ran I. Shorrer, Mengjia Xia |
EC | 1 |
| 2024 | Late Breaking Results: Differential and Massively Parallel Sampling of SAT FormulasabstractDiverse solutions to the Boolean satisfiability (SAT) problem are essential for thorough testing and verification of software and hardware designs, ensuring reliability and applicability to real-world scenarios. We introduce a novel differentiable sampling method, called DiffSampler, which employs gradient descent (GD) to learn diverse solutions to the SAT problem. By formulating SAT as a supervised multi-output regression task and minimizing its loss function using GD, our approach enables performing the learning operations in parallel, leading to GPU-accelerated sampling and comparable run time performance w.r.t. heuristic samplers. We demonstrate that DiffSampler can generate diverse uniform-like solutions similar to conventional samplers. Arash Ardakani, Kevin He, Vighnesh M. Iyer, Suhong Moon, John Wawrzynek |
DAC | 3 |
| 2024 | A Distributed Stealth Address Generation Protocol for Threshold SignaturesabstractStealth addresses protect recipient identity privacy in blockchain systems by allowing a sender to derive a stealth address using the recipient’s public key, with the receiver deriving a corresponding one-time private key. In threshold signatures, where distributed recipient entities only hold private key shares and the full private key never appears, this becomes challenging. To address this, we propose a distributed stealth address generation (DSAG) protocol for threshold signatures. In our approach, distributed sender entities derive the stealth address, and distributed recipient entities derive corresponding one-time private key shares using Lagrange interpolation. The correctness of these shares is verified by ensuring that they satisfy the necessary Lagrange interpolation and discrete logarithm relationships. Our protocol avoids bilinear mappings and relies solely on elliptic curve exponentiations, leading to higher efficiency. Experimental tests demonstrate that our scheme instance takes just 4 milliseconds confirming its practicality. Yudi Zou, Kevin He |
ISPA | 5 |
| 2024 | Learning from Viral ContentabstractIn recent years, viral content on social media platforms has become a major source of news and information for many people. Which stories go viral is jointly determined by the algorithms generating platform news feeds and users' actions on the platforms. We study this process with an equilibrium model of users interacting with shared news stories, focusing on how the design of news feeds affects how users learn. In our model, rational users arrive sequentially, observe an original story (i.e., a private signal) and a sample of predecessors' stories in a news feed, and then decide which stories to share. The observed sample of stories depends on what predecessors share as well as the sampling algorithm generating news feeds. Krishna Dasaratha, Kevin He |
EC | 2 |
| 2023 | Distributed Key Derivation for Multi-Party Management of Blockchain Digital AssetsabstractThe current Single-User Key Derivation (SKD) caters to individual management of blockchain’s tree-structured assets but falls short for threshold signatures aimed at multi-party control of blockchain assets. We introduce a Distributed Key Derivation (DKD) for collaborative management of these assets. Our DKD aligns with SKD and the prevalent Decentralized Key Generation (DKG) in threshold signatures, employing the GG18 DKG [1] protocol to safely refresh child keys without affecting the parent key or other nodes. It maintains the hierarchical structure while preventing privilege escalation attacks. Experimental tests show that our DKD’s key derivation time is about 240 μs, minor compared to GG18’s DKG time. Yong Ding 0005, Kevin He |
ICPADS | 5 |
| 2022 | Screening p-Hackers: Dissemination Noise as BaitabstractWe show that adding noise to data before making data public is effective at screening p-hacked findings: spurious explanations of the outcome variable produced by attempting multiple econometric specifications. Noise creates "baits'' that affect two types of researchers differently. Uninformed p-hackers who engage in data mining with no prior information about the true causal mechanism often fall for baits and report verifiably wrong results when evaluated with the original data. But informed researchers who start with an ex-ante hypothesis about the causal mechanism before seeing any data are minimally affected by noise. We characterize the optimal level of dissemination noise and highlight the relevant trade-offs in a simple theoretical model. Dissemination noise is a tool that statistical agencies (e.g., the US Census Bureau) currently use to protect privacy, and we show this existing practice can be repurposed to improve research credibility. Federico Echenique, Kevin He |
EC | 2 |
| 2022 | Private Private InformationabstractIn a private private information structure, agents' signals contain no information about the signals of their peers. We study how informative such structures can be, and characterize those that are on the Pareto frontier, in the sense that it is impossible to give more information to any agent without violating privacy. In our main application, we show how to optimally disclose information about an unknown state under the constraint of not revealing anything about a correlated variable that contains sensitive information. Kevin He, Fedor Sandomirskiy, Omer Tamuz |
EC | 1 |
| 2021 | Aggregative Efficiency of Bayesian Learning in NetworksabstractIn social-learning settings where individuals receive private signals and observe network neighbors' actions, the network structure often obstructs information aggregation. We consider sequential social learning with rational agents and Gaussian signals and ask how the efficiency of signal aggregation changes with the network. Rational actions in our model are a log-linear function of observations and admit a signal-counting interpretation of accuracy. This leads to a detailed ranking of networks for social learning based on their aggregative efficiency index. Networks where agents observe multiple neighbors but not their common predecessors confound information, and we show confounding can make learning very inefficient. In a class of networks where agents move in generations and observe the previous generation, aggregative efficiency is a simple function of network parameters: increasing in observations and decreasing in confounding. Generations after the first contribute very little additional information due to confounding, even when generations are arbitrarily large. Krishna Dasaratha, Kevin He |
EC | 2 |
| 2021 | Evolutionarily Stable (Mis)specifications: Theory and ApplicationsabstractWe introduce an evolutionary framework to evaluate competing (mis)specifications in strategic situations, focusing on which misspecifications can persist over a correct specification. Agents with heterogeneous specifications coexist in a society and repeatedly match against random opponents to play a stage game. They draw Bayesian inferences about the environment based on personal experience, so their learning depends on the distribution of specifications and matching assortativity in the society. One specification is evolutionarily stable against another if, whenever sufficiently prevalent, its adherents obtain higher expected objective payoffs than their counterparts. The learning channel leads to novel stability phenomena compared to frameworks where the heritable unit of cultural transmission is a single belief instead of a specification (i.e., set of feasible beliefs). We apply the framework to linear-quadratic-normal games where players receive correlated signals but possibly misperceive the information structure. The correct specification is not evolutionarily stable against a correlational error, whose direction depends on matching assortativity. As another application, the framework also endogenizes coarse analogy classes in centipede games. The full paper can be found at https://kevinhe.net/papers/theory_evolution.pdf Kevin He, Jonathan Libgober |
EC | 1 |
| 2021 | Cox-nnet v2.0: improved neural-network-based survival prediction extended to large-scale EMR dataabstractSUMMARY: Cox-nnet is a neural-network-based prognosis prediction method, originally applied to genomics data. Here, we propose the version 2 of Cox-nnet, with significant improvement on efficiency and interpretability, making it suitable to predict prognosis based on large-scale population data, including those electronic medical records (EMR) datasets. We also add permutation-based feature importance scores and the direction of feature coefficients. When applied on a kidney transplantation dataset, Cox-nnet v2.0 reduces the training time of Cox-nnet up to 32-folds (n =10 000) and achieves better prediction accuracy than Cox-PH (P<0.05). It also achieves similarly superior performance on a publicly available SUPPORT data (n=8000). The high efficiency and accuracy make Cox-nnet v2.0 a desirable method for survival prediction in large-scale EMR data. AVAILABILITY AND IMPLEMENTATION: Cox-nnet v2.0 is freely available to the public at https://github.com/lanagarmire/Cox-nnet-v2.0. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Zheng Jing, Kevin He, Lana X. Garmire |
Bioinform. | 3 |
| 2021 | Advancement in predicting interactions between drugs used to treat psoriasis and its comorbidities by integrating molecular and clinical resourcesabstractOBJECTIVE: Drug-drug interactions (DDIs) can result in adverse and potentially life-threatening health consequences; however, it is challenging to predict potential DDIs in advance. We introduce a new computational approach to comprehensively assess the drug pairs which may be involved in specific DDI types by combining information from large-scale gene expression (984 transcriptomic datasets), molecular structure (2159 drugs), and medical claims (150 million patients). MATERIALS AND METHODS: Features were integrated using ensemble machine learning techniques, and we evaluated the DDIs predicted with a large hospital-based medical records dataset. Our pipeline integrates information from >30 different resources, including >10 000 drugs and >1.7 million drug-gene pairs. We applied our technique to predict interactions between 37 611 drug pairs used to treat psoriasis and its comorbidities. RESULTS: Our approach achieves >0.9 area under the receiver operator curve (AUROC) for differentiating 11 861 known DDIs from 25 750 non-DDI drug pairs. Significantly, we demonstrate that the novel DDIs we predict can be confirmed through independent data sources and supported using clinical medical records. CONCLUSIONS: By applying machine learning and taking advantage of molecular, genomic, and health record data, we are able to accurately predict potential new DDIs that can have an impact on public health. Matthew T. Patrick, Redina Bardhi, Kalpana Raja, Kevin He, Lam C. Tsoi |
J. Am. Medical Informatics Assoc. | 4 |
| 2020 | An Experiment on Network Density and Sequential LearningabstractWe conduct a sequential social-learning experiment where subjects take turns guessing a hidden state based on private signals and the guesses of a subset of their predecessors. A network determines the observable predecessors, and we compare subjects' accuracy on sparse and dense networks. Accuracy gains from social learning are twice as large on sparse networks compared to dense networks. Models of naive inference where agents ignore correlation between observations predict this comparative static in network density, while the finding is difficult to reconcile with rational-learning models. Krishna Dasaratha, Kevin He |
EC | 2 |
| 2016 | Component-wise gradient boosting and false discovery control in survival analysis with high-dimensional covariatesabstractMOTIVATION: Technological advances that allow routine identification of high-dimensional risk factors have led to high demand for statistical techniques that enable full utilization of these rich sources of information for genetics studies. Variable selection for censored outcome data as well as control of false discoveries (i.e. inclusion of irrelevant variables) in the presence of high-dimensional predictors present serious challenges. This article develops a computationally feasible method based on boosting and stability selection. Specifically, we modified the component-wise gradient boosting to improve the computational feasibility and introduced random permutation in stability selection for controlling false discoveries. RESULTS: We have proposed a high-dimensional variable selection method by incorporating stability selection to control false discovery. Comparisons between the proposed method and the commonly used univariate and Lasso approaches for variable selection reveal that the proposed method yields fewer false discoveries. The proposed method is applied to study the associations of 2339 common single-nucleotide polymorphisms (SNPs) with overall survival among cutaneous melanoma (CM) patients. The results have confirmed that BRCA2 pathway SNPs are likely to be associated with overall survival, as reported by previous literature. Moreover, we have identified several new Fanconi anemia (FA) pathway SNPs that are likely to modulate survival of CM patients. AVAILABILITY AND IMPLEMENTATION: The related source code and documents are freely available at https://sites.google.com/site/bestumich/issues. CONTACT: [email protected]. Kevin He, Jeffrey E. Lee, Christopher I. Amos, Terry Hyslop, Jiashun Jin, Huazhen Lin, Qinyi Wei, Yi Li 0019 |
Bioinform. | 1 |
| 2014 | Differentially private and incentive compatible recommendation system for the adoption of network goodsabstractWe study the problem of designing a recommendation system for network goods under the constraint of differential privacy. Agents living on a graph face the introduction of a new good and undergo two stages of adoption. The first stage consists of private, random adoptions. In the second stage, remaining non-adopters decide whether to adopt with the help of a recommendation system A. The good has network complimentarity, making it socially desirable for A to reveal the adoption status of neighboring agents. The designer's problem, however, is to find the socially optimal A that preserves privacy. We derive feasibility conditions for this problem and characterize the optimal solution. Kevin He, Xiaosheng Mu |
EC | 1 |