VLDB 2026 Research / reviewers in the wild / expert
Jiarui Lu
dblp:255/8650
· DBLP profile ↗
14ranked-venue papers
5as first author
14since 2021 · last 2025
—ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 10 · 4 first-author · 10 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 first-author · 2 since 2021Applied, interdisciplinary, general and emerging computing · 2 · 2 since 2021Software engineering, systems software and programming languages · 1 · 1 first-author · 1 since 2021Databases, data management, data science and information retrieval · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2025 | Structure Language Models for Protein Conformation GenerationabstractProteins adopt multiple structural conformations to perform their diverse biological functions, and understanding these conformations is crucial for advancing drug discovery. Traditional physics-based simulation methods often struggle with sampling equilibrium conformations and are computationally expensive. Recently, deep generative models have shown promise in generating protein conformations as a more efficient alternative. However, these methods predominantly rely on the diffusion process within a 3D geometric space, which typically centers around the vicinity of metastable states and is often inefficient in terms of runtime. In this paper, we introduce Structure Language Modeling (SLM) as a novel framework for efficient protein conformation generation. Specifically, the protein structures are first encoded into a compact latent space using a discrete variational auto-encoder, followed by conditional language modeling that effectively captures sequence-specific conformation distributions. This enables a more efficient and interpretable exploration of diverse ensemble modes compared to existing methods. Based on this general framework, we instantiate SLM with various popular LM architectures as well as proposing the ESMDiff, a novel BERT-like structure language model fine-tuned from ESM3 with masked diffusion. We verify our approach in various scenarios, including the equilibrium dynamics of BPTI, conformational change pairs, and intrinsically disordered proteins. SLM provides a highly efficient solution, offering a 20-100x speedup than existing methods in generating diverse conformations, shedding light on promising avenues for future research. Jiarui Lu, Xiaoyin Chen, Stephen Zhewen Lu, Chence Shi, Yoshua Bengio, Jian Tang 0005 |
ICLR | 1 |
| 2025 | Aligning Protein Conformation Ensemble Generation with Physical FeedbackabstractProtein dynamics play a crucial role in protein biological functions and properties, and their traditional study typically relies on time-consuming molecular dynamics (MD) simulations conducted in silico. Recent advances in generative modeling, particularly denoising diffusion models, have enabled efficient accurate protein structure prediction and conformation sampling by learning distributions over crystallographic structures. However, effectively integrating physical supervision into these data-driven approaches remains challenging, as standard energy-based objectives often lead to intractable optimization. In this paper, we introduce Energy-based Alignment (EBA), a method that aligns generative models with feedback from physical models, efficiently calibrating them to appropriately balance conformational states based on their energy differences. Experimental results on the MD ensemble benchmark demonstrate that EBA achieves state-of-the-art performance in generating high-quality protein ensembles. By improving the physical plausibility of generated structures, our approach enhances model predictions and holds promise for applications in structural biology and drug discovery. Jiarui Lu, Xiaoyin Chen, Stephen Zhewen Lu, Aurélie C. Lozano, Vijil Chenthamarakshan, Jian Tang 0005 |
ICML | 1 |
| 2025 | EilMoB: Emotion-aware Incongruity Learning and Modality Bridging Network for Multi-modal Sarcasm Detection
Yongxiu Xu, Xinkui Lin, Jiarui Lu |
ICMR | 4 |
| 2025 | Measuring Scientific Capabilities of Language Models with a Systems Biology Dry LababstractDesigning experiments and result interpretations are core scientific competencies, particularly in biology, where researchers perturb complex systems to uncover the underlying systems. Recent efforts to evaluate the scientific capabilities of large language models (LLMs) fail to test these competencies because wet-lab experimentation is prohibitively expensive: in expertise, time and equipment. We introduce SciGym, a first-in-class benchmark that assesses LLMs' iterative experiment design and analysis abilities in open-ended scientific discovery tasks. SciGym overcomes the challenge of wet-lab costs by running a dry lab of biological systems. These models, encoded in Systems Biology Markup Language, are efficient for generating simulated data, making them ideal testbeds for experimentation on realistically complex systems. We evaluated six frontier LLMs on 137 small systems, and released a total of 350 systems at https://huggingface.co/datasets/h4duan/scigym-sbml. Our evaluation shows that while more capable models demonstrated superior performance, all models' performance declined significantly as system complexity increased, suggesting substantial room for improvement in the scientific capabilities of LLM agents. Haonan Duan 0002, Stephen Zhewen Lu, Caitlin F. Harrigan, Nishkrit Desai, Jiarui Lu, Michal Koziarski, Leonardo Cotta, Chris J. Maddison |
NeurIPS | 5 |
| 2024 | Probing the Multi-turn Planning Capabilities of LLMs via 20 Question GamesabstractLarge language models (LLMs) are effective at answering questions that are clearly asked.However, when faced with ambiguous queries they can act unpredictably and produce incorrect outputs.This underscores the need for the development of intelligent agents capable of asking clarification questions to resolve ambiguities effectively.This capability requires complex understanding, state tracking, reasoning and planning over multiple conversational turns.However, directly measuring this can be challenging.In this paper, we offer a surrogate problem which assesses an LLMs's capability to deduce an entity unknown to itself, but revealed to a judge, by asking the judge a series of queries.This entity-deducing game can serve as an evaluation framework to probe the conversational reasoning and planning capabilities of language models.We systematically evaluate various LLMs and discover significant differences in their performance on this task.We find that strong LLMs like GPT-4 outperform human players by a large margin.We further employ Behavior Cloning (BC) to examine whether a weaker model is capable of imitating a stronger model and generalizing to data or domains, using only the demonstrations from a stronger model.We finally propose to use Reinforcement Learning to enhance reasoning and planning capacity of Vicuna models through episodes of game playing, which lead to significant performance improvement.We hope that this problem offers insights into how autonomous agents could be trained to behave more intelligently in ambiguous circumstances. Get MeshuggaThe band Play their song Do you mean Meshugana, the story set, or Meshuggah, the band?Do you want to know their information, or to play their album?Do you want something popular or unique from them?I am thinking of a movie that I cannot remember the name.Can you help me?Nope, it's like 10 years ago.I think so!Is this movie a drama movie? Yizhe Zhang 0002, Jiarui Lu, Navdeep Jaitly |
ACL (1) | 2 |
| 2024 | Towards Foundational Models for Molecular Learning on Large-Scale Multi-Task DatasetsabstractRecently, pre-trained foundation models have enabled significant advancements in multiple fields. In molecular machine learning, however, where datasets are often hand-curated, and hence typically small, the lack of datasets with labeled features, and codebases to manage those datasets, has hindered the development of foundation models. In this work, we present seven novel datasets categorized by size into three distinct categories: ToyMix, LargeMix and UltraLarge. These datasets push the boundaries in both the scale and the diversity of supervised labels for molecular learning. They cover nearly 100 million molecules and over 3000 sparsely defined tasks, totaling more than 13 billion individual labels of both quantum and biological nature. In comparison, our datasets contain 300 times more data points than the widely used OGB-LSC PCQM4Mv2 dataset, and 13 times more than the quantum-only QM1B dataset. In addition, to support the development of foundational models based on our proposed datasets, we present the Graphium graph machine learning library which simplifies the process of building and training molecular machine learning models for multi-task and multi-level molecular datasets. Finally, we present a range of baseline results as a starting point of multi-task and multi-level training on these datasets. Empirically, we observe that performance on low-resource biological datasets show improvement by also training on large amounts of quantum data. This indicates that there may be potential in multi-task and multi-level training of a foundation model and fine-tuning it to resource-constrained downstream tasks. The Graphium library is publicly available on Github and the dataset links are available in Part 1 and Part 2. Dominique Beaini, Shenyang Huang, Joao Alex Cunha, Gabriela Moisescu-Pareja, Oleksandr Dymov, Samuel Maddrell-Mander, Callum McLean, Frederik Wenkel, Luis Müller, Jama Hussein Mohamud, Ali Parviz, Michael Craig, Michal Koziarski, Jiarui Lu, Zhaocheng Zhu, Cristian Gabellini, Kerstin Kläser 0001, Josef Dean, Cas Wognum, Maciej Sypetkowski, Guillaume Rabusseau, Reihaneh Rabbany, Jian Tang 0005, Christopher Morris 0001, Mirco Ravanelli, Guy Wolf, Prudencio Tossou, Hadrien Mary, Therence Bois, Andrew W. Fitzgibbon, Blazej Banaszewski, Chad Martin, Dominic Masters |
ICLR | 15 |
| 2024 | Str2Str: A Score-based Framework for Zero-shot Protein Conformation SamplingabstractThe dynamic nature of proteins is crucial for determining their biological functions and properties, for which Monte Carlo (MC) and molecular dynamics (MD) simulations stand as predominant tools to study such phenomena. By utilizing empirically derived force fields, MC or MD simulations explore the conformational space through numerically evolving the system via Markov chain or Newtonian mechanics. However, the high-energy barrier of the force fields can hamper the exploration of both methods by the rare event, resulting in inadequately sampled ensemble without exhaustive running. Existing learning-based approaches perform direct sampling yet heavily rely on target-specific simulation data for training, which suffers from high data acquisition cost and poor generalizability. Inspired by simulated annealing, we propose Str2Str, a novel structure-to-structure translation framework capable of zero-shot conformation sampling with roto-translation equivariant property. Our method leverages an amortized denoising score matching objective trained on general crystal structures and has no reliance on simulation data during both training and inference. Experimental results across several benchmarking protein systems demonstrate that Str2Str outperforms previous state-of-the-art generative structure prediction models and can be orders of magnitude faster compared with long MD simulations. Jiarui Lu, Bozitao Zhong, Zuobai Zhang, Jian Tang 0005 |
ICLR | 1 |
| 2024 | A Data Alignment Method for Network Packet Capture Based on DBSCANabstractThis paper investigates the issues of packet alignment and consistency among PLC devices based on industrial network environments, aiming to ensure the integrity and accuracy of packets from sender to receiver. To achieve this goal, we propose an anomaly detection method that combines the DBSCAN clustering algorithm with the 3-sigma principle to identify and handle abnormal packets that may occur during transmission. By comparing the data between the sending and receiving ends, and analyzing based on timestamps and data content, we validate the alignment of packets in the network environment. Experimental results demonstrate that the proposed method effectively detects and corrects packet loss or delay jitter, thereby enhancing the reliability of communication between PLC devices and the consistency of data transmission. The scheme presented in this paper enables quicker and more precise identification of packet loss and delays, adapting well to various network load conditions. Further experimental analysis indicates that this method excels in reducing both false positive and false negative rates, and it exhibits good scalability, making it applicable to data alignment and consistency verification in other industrial automation scenarios. Ultimately, this novel solution provides stability and accuracy for data transmission among devices in a network environment. Jiarui Lu, Qinggang Su |
J. Web Eng. | 1 |
| 2023 | Protein Sequence and Structure Co-Design with Equivariant Translation
Chence Shi, Jiarui Lu, Bozitao Zhong, Jian Tang 0005 |
ICLR | 3 |
| 2023 | 5IDER: Unified Query Rewriting for Steering, Intent Carryover, Disfluencies, Entity Carryover and RepairabstractProviding voice assistants the ability to navigate multi-turn conversations is a challenging problem.Handling multi-turn interactions requires the system to understand various conversational use-cases, such as steering, intent carryover, disfluencies, entity carryover, and repair.The complexity of this problem is compounded by the fact that these use-cases mix with each other, often appearing simultaneously in natural language.This work proposes a non-autoregressive query rewriting architecture that can handle not only the five aforementioned tasks, but also complex compositions of these use-cases.We show that our proposed model has competitive single task performance compared to the baseline approach, and even outperforms a fine-tuned T5 model in use-case compositions, despite being 15 times smaller in parameters and 25 times faster in latency. Jiarui Lu, Bo-Hsiang Tseng, Joel Ruben Antony Moniz, Site Li, Xueyun Zhu, Murat Akbacak |
INTERSPEECH | 1 |
| 2022 | PEER: A Comprehensive and Multi-Task Benchmark for Protein Sequence UnderstandingabstractWe are now witnessing significant progress of deep learning methods in a variety of tasks (or datasets) of proteins. However, there is a lack of a standard benchmark to evaluate the performance of different methods, which hinders the progress of deep learning in this field. In this paper, we propose such a benchmark called PEER, a comprehensive and multi-task benchmark for Protein sEquence undERstanding. PEER provides a set of diverse protein understanding tasks including protein function prediction, protein localization prediction, protein structure prediction, protein-protein interaction prediction, and protein-ligand interaction prediction. We evaluate different types of sequence-based methods for each task including traditional feature engineering approaches, different sequence encoding methods as well as large-scale pre-trained protein language models. In addition, we also investigate the performance of these methods under the multi-task learning setting. Experimental results show that large-scale pre-trained protein language models achieve the best performance for most individual tasks, and jointly training multiple tasks further boosts the performance. The datasets and source codes of this benchmark will be open-sourced soon. Zuobai Zhang, Jiarui Lu, Zhaocheng Zhu, Yangtian Zhang, Runcheng Liu, Jian Tang 0005 |
NeurIPS | 3 |
| 2021 | CREAD: Combined Resolution of Ellipses and Anaphora in DialoguesabstractBo-Hsiang Tseng, Shruti Bhargava, Jiarui Lu, Joel Ruben Antony Moniz, Dhivya Piraviperumal, Lin Li, Hong Yu. Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies. 2021. Bo-Hsiang Tseng, Shruti Bhargava, Jiarui Lu, Joel Ruben Antony Moniz, Dhivya Piraviperumal |
NAACL-HLT | 3 |
| 2021 | A compositional mediation model for a binary outcome: Application to microbiome studiesabstractMOTIVATION: The delicate balance of the microbiome is implicated in our health and is shaped by external factors, such as diet and xenobiotics. Therefore, understanding the role of the microbiome in linking external factors and our health conditions is crucial to translate microbiome research into therapeutic and preventative applications. RESULTS: We introduced a sparse compositional mediation model for binary outcomes to estimate and test the mediation effects of the microbiome utilizing the compositional algebra defined in the simplex space and a linear zero-sum constraint on probit regression coefficients. For this model with the standard causal assumptions, we showed that both the causal direct and indirect effects are identifiable. We further developed a method for sensitivity analysis for the assumption of the no unmeasured confounding effects between the mediator and the outcome. We conducted extensive simulation studies to assess the performance of the proposed method and applied it to real microbiome data to study mediation effects of the microbiome on linking fat intake to overweight/obesity. AVAILABILITY AND IMPLEMENTATION: An R package can be downloaded from https://github.com/mbsohn/cmmb. SUPPLEMENTARY INFORMATION: Supplementary files are available at Bioinformatics online. Michael B. Sohn, Jiarui Lu, Hongzhe Li |
Bioinform. | 2 |
| 2021 | KenDTI: An Ensemble Model for Predicting Drug-Target Interaction by Integrating Multi-Source InformationabstractThe identification of drug-target interactions (DTIs) is an essential step in the process of drug discovery. As experimental validation suffers from high cost and low success rate, various computational models have been exploited to infer potential DTIs. The performance of DTI prediction depends heavily on the features extracted from drugs and target proteins. The existing predictors vary in input information and each has its own advantages. Therefore, combining the advantages of individual models and generating high-quality representations for drug-target pairs are effective ways to improve the performance of DTI prediction. In this study, we exploit both biochemical characteristics of drugs via network integration and molecular sequences via word embeddings, then we develop an ensemble model, KenDTI, based on two types of methods, i.e., network-based and classification-based. We assess the performance of KenDTI on two large-scale datasets, The experimental results show that KenDTI outperforms the state-of-the-art DTI predictors by a large margin. Moreover, KenDTI is robust against missing data in input networks and lack of prior knowledge. It is able to predict for drug-candidate chemical compounds with scarce information. Zhimiao Yu, Jiarui Lu, Yang Yang 0030 |
IEEE ACM Trans. Comput. Biol. Bioinform. | 2 |