EDBT 2026 Demo / reviewers in the wild / expert
Tom Gibbs
dblp:236/5876
· DBLP profile ↗
4ranked-venue papers
0as first author
4since 2021 · last 2025
0000-0002-9196-5830ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 2 · 2 since 2021Systems, architecture and hardware · 2 · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
Expertise — from the expertise taxonomy: the topics of the expert's papers under the CCF categories. A weight counts papers with recency: 1 for a paper about the topic, 0.3 when the topic is its context, halved every five years.
| Interdisciplinary, comprehensive, and emerging computing
2 papers |
Bioinformatics and computational biology · 90% Computational social science and digital humanities · 10% | |
| Artificial intelligence
2 papers |
Multi-agent systems · 51% Language models and text generation · 49% | |
| Computer architecture, parallel and distributed computing, and storage systems
1 paper |
High-performance computing · 77% GPUs and heterogeneous computing · 23% |
Topics — the 10 heaviest of 10, each with the papers that count most for it
| Topic | Weight | Papers | Last | Evidence papers |
|---|---|---|---|---|
Knowledge, reasoning and agents › Multi-agent systems › agent-based simulation › social simulation
agent-based social simulation |
0.9 | 1 | 2025 | SandboxSocial: A Sandbox for Social Media Using Multimodal AI Agents · IJCAI 2025 |
High-performance computing
scientific computing systems |
0.8 | 1 | 2024 | MProt-DPO: Breaking the ExaFLOPS Barrier for Multimodal Protein Design Workflows with Direct Preference Optimization · SC 2024 |
Natural language and speech › Language models and text generation › neural language model
protein language model |
0.6 | 1 | 2022 | ProtTrans: Toward Understanding the Language of Life Through Self-Supervised Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2022 |
Bioinformatics and computational biology
protein function prediction |
0.6 | 1 | 2022 | ProtTrans: Toward Understanding the Language of Life Through Self-Supervised Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2022 |
Bioinformatics and computational biology
protein structure prediction |
0.6 | 1 | 2022 | ProtTrans: Toward Understanding the Language of Life Through Self-Supervised Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2022 |
Bioinformatics and computational biology › protein function prediction
protein subcellular localization prediction |
0.6 | 1 | 2022 | ProtTrans: Toward Understanding the Language of Life Through Self-Supervised Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2022 |
Bioinformatics and computational biology › protein structure prediction
secondary structure prediction |
0.6 | 1 | 2022 | ProtTrans: Toward Understanding the Language of Life Through Self-Supervised Learning · IEEE Trans. Pattern Anal. Mach. Intell. 2022 |
Natural language and speech › Language models and text generation
LLM agents |
0.3 | 1 | 2025 | SandboxSocial: A Sandbox for Social Media Using Multimodal AI Agents · IJCAI 2025 |
Computational social science and digital humanities
social media analysis |
0.3 | 1 | 2025 | SandboxSocial: A Sandbox for Social Media Using Multimodal AI Agents · IJCAI 2025 |
GPUs and heterogeneous computing
GPU and heterogeneous computing |
0.2 | 1 | 2024 | MProt-DPO: Breaking the ExaFLOPS Barrier for Multimodal Protein Design Workflows with Direct Preference Optimization · SC 2024 |
Methods — techniques the papers use, named apart from their topics
multimodal AI agents · 1.7large language model · 1.7t5 · 1.1self-supervised learning · 1.1albert · 1.1XLNet · 1.1Transformer-XL · 1.1ELECTRA · 1.1BERT · 1.1multimodal generative models · 0.8mixed precision · 0.8direct preference optimization · 0.8
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | SandboxSocial: A Sandbox for Social Media Using Multimodal AI AgentsabstractThe online information ecosystem enables influence campaigns of unprecedented scale and impact. We urgently need empirically grounded approaches to counter the growing threat of malicious campaigns, now amplified by generative AI. But, developing defenses in real-world settings is impractical. Social system simulations with agents modelled using Large Language Models (LLMs) are a promising alternative approach and a growing area of research. However, existing simulators lack features needed to capture the complex information-sharing dynamics of platform-based social networks. To bridge this gap, we present SandboxSocial, a new simulator that includes several key innovations, mainly: (1) a virtual social media platform (modelled as Mastodon and mirrored in an actual Mastodon server) that enables a realistic setting in which agents interact; (2) an adapter that uses real-world user data to create more grounded agents and social media content; and (3) multi-modal capabilities that enable our agents to interact using both text and images---just as humans do on social media. We make the simulator more useful to researchers by providing measurement and analysis tools that track simulation dynamics and compute evaluation metrics to compare experimental results. Maximilian Puelma Touzel, Sneheel Sarangi, Gayatri Krishnakumar, Busra Tugce Gurbuz, Austin Welch, Zachary Yang, Andreea Musulan, Ethan Kosak-Hine, Tom Gibbs, Camille Thibault, Reihaneh Rabbany, Jean-François Godbout, Kellin Pelrine |
IJCAI | 10 |
| 2024 | MProt-DPO: Breaking the ExaFLOPS Barrier for Multimodal Protein Design Workflows with Direct Preference OptimizationabstractWe present a scalable, end-to-end workflow for protein design. By augmenting protein sequences with natural language descriptions of their biochemical properties, we train generative models that can be preferentially aligned with protein fitness landscapes. Through complex experimental-and simulation-based observations, we integrate these measures as preferred parameters for generating new protein variants and demonstrate our workflow on five diverse supercomputers. We achieve >1 ExaFLOPS sustained performance in mixed precision on each supercomputer and a maximum sustained performance of 4.11 Ex-aFLOPS and peak performance of 5.57 ExaFLOPS. We establish the scientific performance of our model on two tasks: (1) across a predetermined benchmark dataset of deep mutational scanning experiments to optimize the fitness-determining mutations in the yeast protein HIS7, and (2) in optimizing the design of the enzyme malate dehydrogenase to achieve lower activation barriers (and therefore increased catalytic rates) using simulation data. Our implementation thus sets high watermarks for multimodal protein design workflows. Gautham Dharuman, Kyle Hippe, Alex Brace, Sam Foreman, Väinö Hatanpää, Varuni Sastry 0001, Huihuo Zheng, Logan T. Ward, Servesh Muralidharan, Archit Vasan, Bharat Kale, Carla M. Mann, Yun-Hsuan Cheng, Yuliana Zamora, Shengchao Liu, Chaowei Xiao, Murali Emani, Tom Gibbs, Mahidhar Tatineni, Deepak Canchi, Jerome Mitchell, Koichi Yamada, María Jesús Garzarán, Michael E. Papka, Ian T. Foster, Rick L. Stevens, Anima Anandkumar, Venkatram Vishwanath, Arvind Ramanathan |
SC | 19 |
| 2022 | ProtTrans: Toward Understanding the Language of Life Through Self-Supervised LearningabstractComputational biology and bioinformatics provide vast data gold-mines from protein sequences, ideal for Language Models (LMs) taken from Natural Language Processing (NLP). These LMs reach for new prediction frontiers at low inference costs. Here, we trained two auto-regressive models (Transformer-XL, XLNet) and four auto-encoder models (BERT, Albert, Electra, T5) on data from UniRef and BFD containing up to 393 billion amino acids. The protein LMs (pLMs) were trained on the Summit supercomputer using 5616 GPUs and TPU Pod up-to 1024 cores. Dimensionality reduction revealed that the raw pLM-embeddings from unlabeled data captured some biophysical features of protein sequences. We validated the advantage of using the embeddings as exclusive input for several subsequent tasks: (1) a per-residue (per-token) prediction of protein secondary structure (3-state accuracy Q3=81%-87%); (2) per-protein (pooling) predictions of protein sub-cellular location (ten-state accuracy: Q10=81%) and membrane versus water-soluble (2-state accuracy Q2=91%). For secondary structure, the most informative embeddings (ProtT5) for the first time outperformed the state-of-the-art without multiple sequence alignments (MSAs) or evolutionary information thereby bypassing expensive database searches. Taken together, the results implied that pLMs learned some of the grammar of the language of life. All our models are available through https://github.com/agemagician/ProtTrans. Ahmed Elnaggar, Michael Heinzinger, Christian Dallago, Ghalia Rehawi, Yu Wang 0008, Llion Jones, Tom Gibbs, Tamas Feher, Christoph Angerer, Martin Steinegger, Debsindhu Bhowmik, Burkhard Rost |
IEEE Trans. Pattern Anal. Mach. Intell. | 7 |
| 2021 | IMPECCABLE: Integrated Modeling PipelinE for COVID Cure by Assessing Better LEadsabstractThe drug discovery process currently employed in the pharmaceutical industry typically requires about 10 years and $2–3 billion to deliver one new drug. This is both too expensive and too slow, especially in emergencies like the COVID-19 pandemic. In silico methodologies need to be improved both to select better lead compounds, so as to improve the efficiency of later stages in the drug discovery protocol, and to identify those lead compounds more quickly. No known methodological approach can deliver this combination of higher quality and speed. Here, we describe an Integrated Modeling PipEline for COVID Cure by Assessing Better LEads (IMPECCABLE) that employs multiple methodological innovations to overcome this fundamental limitation. We also describe the computational framework that we have developed to support these innovations at scale, and characterize the performance of this framework in terms of throughput, peak performance, and scientific results. We show that individual workflow components deliver 100 × to 1000 × improvement over traditional methods, and that the integration of methods, supported by scalable infrastructure, speeds up drug discovery by orders of magnitudes. IMPECCABLE has screened ∼ 1011 ligands and has been used to discover a promising drug candidate. These capabilities have been used by the US DOE National Virtual Biotechnology Laboratory and the EU Centre of Excellence in Computational Biomedicine. Aymen Alsaadi, Dario Alfè, Yadu N. Babuji, Agastya Bhati, Ben Blaiszik, Alex Brace, Thomas S. Brettin, Kyle Chard, Ryan Chard, Austin Clyde, Peter V. Coveney, Ian T. Foster, Tom Gibbs, Shantenu Jha, Kristopher Keipert, Dieter Kranzlmüller, Thorsten Kurth, Hyungro Lee, Zhuozhao Li, Gerald Mathias, André Merzky, Alexander Partin, Arvind Ramanathan, Ashka Shah, Abraham C. Stern, Rick L. Stevens, Mikhail Titov, Anda Trifan, Aristeidis Tsaris, Matteo Turilli, Huub J. J. Van Dam, Shunzhou Wan, David Wifling, Junqi Yin |
ICPP | 13 |