VLDB 2026 Research / reviewers in the wild / expert
Samuel C. Hoffman
dblp:220/5755
· DBLP profile ↗
6ranked-venue papers
0as first author
3since 2021 · last 2026
0009-0000-8730-2789ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 5 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 3 · 3 since 2021Databases, data management, data science and information retrieval · 1
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Artificial intelligence
4 papers |
Generative modeling · 60% Trustworthy machine learning · 40% | |
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 100% | |
| Theoretical computer science
1 paper |
Mathematical optimization · 100% |
Topics — the 12 heaviest of 13, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Machine learning › Generative modeling
molecular generation |
1.4 | 2 | 2026 | GP-MoLFormer-Sim: Test Time Molecular Optimization Through Contextual Similarity Guidance · AAAI 2026 CogMol: Target-Specific and Selective Drug Design for COVID-19 Using Deep Generative Models · NeurIPS 2020 |
Machine learning › Trustworthy machine learning › interpretability › explainable AI
explanation methods |
1.0 | 2 | 2022 | AI Explainability 360: Impact and Design · AAAI 2022 AI Explainability 360: An Extensible Toolkit for Understanding Data and Machine Learning Models · J. Mach. Learn. Res. 2020 |
Machine learning › Trustworthy machine learning
interpretability |
1.0 | 2 | 2022 | AI Explainability 360: Impact and Design · AAAI 2022 AI Explainability 360: An Extensible Toolkit for Understanding Data and Machine Learning Models · J. Mach. Learn. Res. 2020 |
Machine learning › Generative modeling › molecular generation
molecular optimization |
1.0 | 1 | 2026 | GP-MoLFormer-Sim: Test Time Molecular Optimization Through Contextual Similarity Guidance · AAAI 2026 |
Bioinformatics and computational biology
drug discovery |
0.7 | 2 | 2026 | CogMol: Target-Specific and Selective Drug Design for COVID-19 Using Deep Generative Models · NeurIPS 2020 GP-MoLFormer-Sim: Test Time Molecular Optimization Through Contextual Similarity Guidance · AAAI 2026 |
Machine learning › Trustworthy machine learning › interpretability
explanation evaluation |
0.6 | 2 | 2022 | AI Explainability 360: An Extensible Toolkit for Understanding Data and Machine Learning Models · J. Mach. Learn. Res. 2020 AI Explainability 360: Impact and Design · AAAI 2022 |
Machine learning › Generative modeling
variational autoencoder |
0.4 | 1 | 2020 | CogMol: Target-Specific and Selective Drug Design for COVID-19 Using Deep Generative Models · NeurIPS 2020 |
Bioinformatics and computational biology › drug discovery › drug design
de novo drug design |
0.4 | 1 | 2020 | CogMol: Target-Specific and Selective Drug Design for COVID-19 Using Deep Generative Models · NeurIPS 2020 |
Mathematical optimization
black-box optimization |
0.4 | 1 | 2020 | Combinatorial Black-Box Optimization with Expert Advice · KDD 2020 |
Mathematical optimization
combinatorial optimization |
0.4 | 1 | 2020 | Combinatorial Black-Box Optimization with Expert Advice · KDD 2020 |
Mathematical optimization › metaheuristic optimization
simulated annealing |
0.4 | 1 | 2020 | Combinatorial Black-Box Optimization with Expert Advice · KDD 2020 |
Bioinformatics and computational biology › drug discovery › drug design
structure-based drug design |
0.3 | 1 | 2026 | GP-MoLFormer-Sim: Test Time Molecular Optimization Through Contextual Similarity Guidance · AAAI 2026 |
Methods — techniques the papers use, named apart from their topics
genetic algorithm · 2.0contextual similarity guidance · 2.0autoregressive sampling · 2.0toxicity prediction · 0.9retrosynthesis prediction · 0.9docking simulation · 0.9controlled sampling · 0.9binding affinity prediction · 0.9explainability methods · 0.6software toolkit · 0.4simulated annealing · 0.4multilinear polynomials · 0.4exponential weight update · 0.4expert advice · 0.4
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | GP-MoLFormer-Sim: Test Time Molecular Optimization Through Contextual Similarity GuidanceabstractThe ability to design molecules while preserving similarity to a target molecule and/or property is crucial for various applications in drug discovery, chemical design, and biology. We introduce in this paper an efficient training-free method for navigating and sampling from the molecular space with a generative Chemical Language Model (CLM), while using the molecular similarity to the target as a guide. Our method leverages the contextual representations learned from the CLM itself to estimate the molecular similarity, which is then used to adjust the autoregressive sampling strategy of the CLM. At each step of the decoding process, the method tracks the distance of the current generations from the target and updates the logits to encourage the preservation of similarity in generations. We implement the method using a recently proposed ~47M parameter SMILES-based CLM, GP-MoLFormer, and therefore refer to the method as GP-MoLFormer-Sim, which enables a test-time update of the deep generative policy to reflect the contextual similarity to a set of guide molecules. The method is further integrated into a genetic algorithm (GA) and tested on a set of standard molecular optimization benchmarks involving property optimization, molecular rediscovery, and structure-based drug design. Results show that, GP-MoLFormer-Sim, combined with GA (GP-MoLFormer-Sim+GA) outperforms existing training-free baseline methods, when the oracle remains black-box. The findings in this work are a step forward in understanding and guiding the generative mechanisms of CLMs. Jirí Navrátil 0001, Jarret Ross, Youssef Mroueh, Samuel C. Hoffman, Vijil Chenthamarakshan, Brian M. Belgodere |
AAAI | 5 |
| 2022 | AI Explainability 360: Impact and DesignabstractAs artificial intelligence and machine learning algorithms become increasingly prevalent in society, multiple stakeholders are calling for these algorithms to provide explanations. At the same time, these stakeholders, whether they be affected citizens, government regulators, domain experts, or system developers, have different explanation needs. To address these needs, in 2019, we created AI Explainability 360, an open source software toolkit featuring ten diverse and state-of-the-art explainability methods and two evaluation metrics. This paper examines the impact of the toolkit with several case studies, statistics, and community feedback. The different ways in which users have experienced AI Explainability 360 have resulted in multiple types of impact and improvements in multiple metrics, highlighted by the adoption of the toolkit by the independent LF AI & Data Foundation. The paper also describes the flexible design of the toolkit, examples of its use, and the significant educational material and documentation available to its users. Vijay Arya, Rachel K. E. Bellamy, Amit Dhurandhar, Michael Hind, Samuel C. Hoffman, Stephanie Houde, Qingzi Vera Liao, Ronny Luss, Aleksandra Mojsilovic, Sami Mourad, Pablo Pedemonte, Ramya Raghavendra, John T. Richards, Prasanna Sattigeri, Karthikeyan Shanmugam 0001, Moninder Singh, Kush R. Varshney, Dennis Wei |
AAAI | 6 |
| 2022 | Augmenting Molecular Deep Generative Models with Topological Data Analysis RepresentationsabstractDeep generative models have emerged as a powerful tool for learning useful molecular representations and designing novel molecules with desired properties, with applications in drug discovery and material design. However, most existing deep generative models are restricted due to lack of spatial information. Here we propose augmentation of deep generative models with topological data analysis (TDA) representations, known as persistence images, for robust encoding of 3D molecular geometry. We show that the TDA augmentation of a character-based Variational Auto-Encoder (VAE) outperforms state-of-the-art generative neural nets in accurately modeling the structural composition of the QM9 benchmark. Generated molecules are valid, novel, and diverse, while exhibiting distinct electronic property distribution, namely higher sample population with small HOMO-LUMO gap. These results demonstrate that TDA features indeed provide crucial geometric signal for learning abstract structures, which is non-trivial for existing generative models operating on string, graph, or 3D point sets to capture. Yair Schiff, Vijil Chenthamarakshan, Samuel C. Hoffman, Karthikeyan Natesan Ramamurthy |
ICASSP | 3 |
| 2020 | Combinatorial Black-Box Optimization with Expert AdviceabstractWe consider the problem of black-box function optimization over the Boolean hypercube. Despite the vast literature on black-box function optimization over continuous domains, not much attention has been paid to learning models for optimization over combinatorial domains until recently. However, the computational complexity of the recently devised algorithms are prohibitive even for moderate numbers of variables; drawing one sample using the existing algorithms is more expensive than a function evaluation for many black-box functions of interest. To address this problem, we propose a computationally efficient model learning algorithm based on multilinear polynomials and exponential weight updates. In the proposed algorithm, we alternate between simulated annealing with respect to the current polynomial representation and updating the weights using monomial experts' advice. Numerical experiments on various datasets in both unconstrained and sum-constrained Boolean optimization indicate the competitive performance of the proposed algorithm, while improving the computational time up to several orders of magnitude compared to state-of-the-art algorithms in the literature. Hamid Dadkhahi, Karthikeyan Shanmugam 0001, Jesus Rios, Samuel C. Hoffman, Troy D. Loeffler, Subramanian Sankaranarayanan |
KDD | 5 |
| 2020 | CogMol: Target-Specific and Selective Drug Design for COVID-19 Using Deep Generative ModelsabstractThe novel nature of SARS-CoV-2 calls for the development of efficient de novo drug design approaches. In this study, we propose an end-to-end framework, named CogMol (Controlled Generation of Molecules), for designing new drug-like small molecules targeting novel viral proteins with high affinity and off-target selectivity. CogMol combines adaptive pre-training of a molecular SMILES Variational Autoencoder (VAE) and an efficient multi-attribute controlled sampling scheme that uses guidance from attribute predictors trained on latent features. To generate novel and optimal drug-like molecules for unseen viral targets, CogMol leverages a protein-molecule binding affinity predictor that is trained using SMILES VAE embeddings and protein sequence embeddings learned unsupervised from a large corpus. We applied the CogMol framework to three SARS-CoV-2 target proteins: main protease, receptor-binding domain of the spike protein, and non-structural protein 9 replicase. The generated candidates are novel at both the molecular and chemical scaffold levels when compared to the training data. CogMol also includes insilico screening for assessing toxicity of parent molecules and their metabolites with a multi-task toxicity classifier, synthetic feasibility with a chemical retrosynthesis predictor, and target structure binding with docking simulations. Docking reveals favorable binding of generated molecules to the target protein structure, where 87--95\% of high affinity molecules showed docking free energy $<$ -6 kcal/mol. When compared to approved drugs, the majority of designed compounds show low predicted parent molecule and metabolite toxicity and high predicted synthetic feasibility. In summary, CogMol can handle multi-constraint design of synthesizable, low-toxic, drug-like molecules with high target specificity and selectivity, even to novel protein target sequences, and does not need target-dependent fine-tuning of the framework or target structure information. Vijil Chenthamarakshan, Samuel C. Hoffman, Hendrik Strobelt, Inkit Padhi, Kar Wai Lim, Benjamin Hoover, Matteo Manica, Jannis Born, Teodoro Laino, Aleksandra Mojsilovic |
NeurIPS | 3 |
| 2020 | AI Explainability 360: An Extensible Toolkit for Understanding Data and Machine Learning ModelsabstractAs artificial intelligence algorithms make further inroads in high-stakes societal applications, there are increasing calls from multiple stakeholders for these algorithms to explain their outputs. To make matters more challenging, different personas of consumers of explanations have different requirements for explanations. Toward addressing these needs, we introduce AI Explainability 360, an open-source Python toolkit featuring ten diverse and state-of-the-art explainability methods and two evaluation metrics. Equally important, we provide a taxonomy to help entities requiring explanations to navigate the space of interpretation and explanation methods, not only those in the toolkit but also in the broader literature on explainability. For data scientists and other users of the toolkit, we have implemented an extensible software architecture that organizes methods according to their place in the AI modeling pipeline. The toolkit is not only the software, but also guidance material, tutorials, and an interactive web demo to introduce AI explainability to different audiences. Together, our toolkit and taxonomy can help identify gaps where more explainability methods are needed and provide a platform to incorporate them as they are developed. Vijay Arya, Rachel K. E. Bellamy, Amit Dhurandhar, Michael Hind, Samuel C. Hoffman, Stephanie Houde, Qingzi Vera Liao, Ronny Luss, Aleksandra Mojsilovic, Sami Mourad, Pablo Pedemonte, Ramya Raghavendra, John T. Richards, Prasanna Sattigeri, Karthikeyan Shanmugam 0001, Moninder Singh, Kush R. Varshney, Dennis Wei |
J. Mach. Learn. Res. | 6 |