VLDB 2026 Research / reviewers in the wild / expert
Jinhua Zhu 0001
dblp:18/1965-1
· DBLP profile ↗
27ranked-venue papers
7as first author
23since 2021 · last 2026
0000-0003-2157-9077ORCID · verified
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 21 · 7 first-author · 17 since 2021Databases, data management, data science and information retrieval · 5 · 2 first-author · 5 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 5 since 2021Computer networks · 1 · 1 since 2021Graphics, computer vision, multimedia, augmented reality and games · 1 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | Bias Fitting to Mitigate Length Bias of Reward Model in RLHFabstractKangwen Zhao, Jianfeng Cai, Jinhua Zhu, Ruopei Sun, Dongyun Xue, Wengang Zhou, Li Li, Houqiang Li. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Kangwen Zhao, Jianfeng Cai 0008, Jinhua Zhu 0001, Ruopei Sun, Dongyun Xue, Wengang Zhou 0001, Li Li 0040, Houqiang Li |
ACL (1) | 3 |
| 2025 | 3D-MolT5: Leveraging Discrete Structural Information for Molecule-Text ModelingabstractThe integration of molecular and natural language representations has emerged as a focal point in molecular science, with recent advancements in Language Models (LMs) demonstrating significant potential for comprehensive modeling of both domains. However, existing approaches face notable limitations, particularly in their neglect of three-dimensional (3D) information, which is crucial for understanding molecular structures and functions. While some efforts have been made to incorporate 3D molecular information into LMs using external structure encoding modules, significant difficulties remain, such as insufficient interaction across modalities in pre-training and challenges in modality alignment. To address the limitations, we propose \textbf{3D-MolT5}, a unified framework designed to model molecule in both sequence and 3D structure spaces. The key innovation of our approach lies in mapping fine-grained 3D substructure representations into a specialized 3D token vocabulary. This methodology facilitates the seamless integration of sequence and structure representations in a tokenized format, enabling 3D-MolT5 to encode molecular sequences, molecular structures, and text sequences within a unified architecture. Leveraging this tokenized input strategy, we build a foundation model that unifies the sequence and structure data formats. We then conduct joint pre-training with multi-task objectives to enhance the model's comprehension of these diverse modalities within a shared representation space. Thus, our approach significantly improves cross-modal interaction and alignment, addressing key challenges in previous work. Further instruction tuning demonstrated that our 3D-MolT5 has strong generalization ability and surpasses existing methods with superior performance in multiple downstream tasks, such as nearly 70\% improvement on the molecular property prediction task compared to state-of-the-art methods. Our code is available at \url{https://github.com/QizhiPei/3D-MolT5}. Qizhi Pei, Rui Yan 0001, Kaiyuan Gao, Jinhua Zhu 0001, Lijun Wu 0003 |
ICLR | 4 |
| 2025 | FABind+: Enhancing Molecular Docking through Improved Pocket Prediction and Pose GenerationabstractMolecular docking is a pivotal process in drug discovery. While traditional techniques rely on extensive sampling and simulation governed by physical principles, deep learning has emerged as a promising alternative, offering improvements in both accuracy and efficiency. Building upon the foundational work of FABind, a model focused on speed and accuracy, we introduce FABind+, an enhanced iteration that significantly elevates the performance of its predecessor. We identify pocket prediction as a critical bottleneck in molecular docking and introduce an enhanced approach. In addition to the pocket prediction module, the docking module has also been upgraded with permutation loss and a more refined model design. These designs enable the regression-based FABind+ to surpass most of the generative models. In contrast, while sampling-based models often struggle with inefficiency, they excel in capturing a wide range of potential docking poses, leading to better overall performance. To bridge the gap between sampling and regression docking models, we incorporate a simple yet effective sampling technique coupled with a lightweight confidence model, transforming the regression-based FABind+ into a sampling version without requiring additional training. This involves the introduction of pocket clustering to capture multiple binding sites and dropout sampling for various conformations. The combination of a classification loss and a ranking loss enables the lightweight confidence model to select the most accurate prediction. Experimental results and analysis demonstrate that FABind+ (both the regression and sampling versions) not only significantly outperforms the original FABind, but also achieves competitive state-of-the-art performance. Our code is available at https://github.com/QizhiPei/FABind. Kaiyuan Gao, Qizhi Pei, Jinhua Zhu 0001, Kun He 0001, Lijun Wu 0003 |
KDD (1) | 4 |
| 2025 | RESIST: Rationale-Enhanced and Reward Model-Based End-to-End Social Influence Dialogue SystemabstractDeveloping proactive social influence dialogue systems presents a significant challenge, particularly in non-cooperative scenarios where the system’s goals may conflict with those of the user. Traditional methods often focus on training models to plan dialogue strategies, but since human strategies are often sub-optimal, relying solely on manually collected data can be problematic. While Large Language Models (LLMs) facilitate the generation of high-quality synthetic dialogues, their effectiveness in strategic dialogue under zero-shot or few-shot conditions is inconsistent. To address these issues, we propose a training framework applicable to multiple social influence dialogue tasks, named R ationale-Enhanced and R eward Model-Based E nd-to-End S ocial I nfluence Dialogue S ys t em (RESIST) . To streamline the dialogue system development, we first use existing datasets to prompt a teacher LLM for generating “chain-of-thought” rationales, which are then used to enrich the data and enable Supervised Fine-Tuning (SFT) of the model. Next, we train a reward model by ranking the fine-tuned model’s outputs, thereby deriving task-specific preferences without manually constructing scalar rewards. Finally, we apply reinforcement learning to further refine the system, optimizing dialogue strategies and responses according to specific tasks and conversational contexts. Experimental results on three social influence tasks demonstrate the effectiveness and adaptability of our training approach. In terms of task goal completion , RESIST outperforms baseline models and even exceeds the performance of ChatGPT-driven prompt-based policy planning methods in both efficiency and effectiveness. Additionally, we introduce strategic proactivity as a novel evaluation metric, enabling us to analyze how RESIST training influences the proactive traits of dialogue agents, with a particular focus on the personality tendencies of smaller-scale language models during task execution. Experimental findings indicate that RESIST enhances the strategic proactivity of language models, aligning them more closely with task requirements. The source code will be made publicly available upon publication. Jinhua Zhu 0001, Wengang Zhou 0001, Houqiang Li |
ACM Trans. Multim. Comput. Commun. Appl. | 2 |
| 2024 | Exploiting Pre-trained Models for Drug Target Affinity Prediction with Nearest NeighborsabstractDrug-Target binding Affinity (DTA) prediction is essential for drug discovery. Despite the application of deep learning methods to DTA prediction, the achieved accuracy remain suboptimal. In this work, inspired by the recent success of retrieval methods, we propose kNN-DTA, a non-parametric embedding-based retrieval method adopted on a pre-trained DTA prediction model, which can extend the power of the DTA model with no or negligible cost. Different from existing methods, we introduce two neighbor aggregation ways from both embedding space and label space that are integrated into a unified framework. Specifically, we propose a label aggregation with pair-wise retrieval and a representation aggregation with point-wise retrieval of the nearest neighbors. This method executes in the inference phase and can efficiently boost the DTA prediction performance with no training cost. In addition, we propose an extension, Ada-kNN-DTA, an instance-wise and adaptive aggregation with lightweight learning. Results on four benchmark datasets show that kNN-DTA brings significant improvements, outperforming previous state-of-the-art (SOTA) results, e.g, on BindingDB IC50 and Ki testbeds, kNN-DTA obtains new records of RMSE 0.684 and 0.750 . The extended Ada-kNN-DTA further improves the performance to be 0.675 and 0.735 RMSE. These results strongly prove the effectiveness of our method. Results in other settings and comprehensive studies/analyses also show the great potential of our kNN-DTA approach. Qizhi Pei, Lijun Wu 0003, Zhenyu He 0012, Jinhua Zhu 0001, Yingce Xia, Shufang Xie 0003, Rui Yan 0001 |
CIKM | 4 |
| 2024 | Trustworthy Alignment of Retrieval-Augmented Large Language Models via Reinforcement LearningabstractTrustworthiness is an essential prerequisite for the real-world application of large language models. In this paper, we focus on the trustworthiness of language models with respect to retrieval augmentation. Despite being supported with external evidence, retrieval-augmented generation still suffers from hallucinations, one primary cause of which is the conflict between contextual and parametric knowledge. We deem that retrieval-augmented language models have the inherent capabilities of supplying response according to both contextual and parametric knowledge. Inspired by aligning language models with human preference, we take the first step towards aligning retrieval-augmented language models to a status where it responds relying merely on the external evidence and disregards the interference of parametric knowledge. Specifically, we propose a reinforcement learning based algorithm Trustworthy-Alignment, theoretically and experimentally demonstrating large language models' capability of reaching a trustworthy status without explicit supervision on how to respond. Our work highlights the potential of large language models on exploring its intrinsic abilities by its own and expands the application scenarios of alignment from fulfilling human preference to creating trustworthy agents. Zongmeng Zhang, Jinhua Zhu 0001, Wengang Zhou 0001, Xiang Qi, Peng Zhang 0080, Houqiang Li |
ICML | 3 |
| 2023 | BioT5: Enriching Cross-modal Integration in Biology with Chemical Knowledge and Natural Language AssociationsabstractRecent advancements in biological research leverage the integration of molecules, proteins, and natural language to enhance drug discovery.However, current models exhibit several limitations, such as the generation of invalid molecular SMILES, underutilization of contextual information, and equal treatment of structured and unstructured knowledge.To address these issues, we propose BioT5, a comprehensive pretraining framework that enriches cross-modal integration in biology with chemical knowledge and natural language associations.BioT5 utilizes SELFIES for 100% robust molecular representations and extracts knowledge from the surrounding context of bio-entities in unstructured biological literature.Furthermore, BioT5 distinguishes between structured and unstructured knowledge, leading to more effective utilization of information.After fine-tuning, BioT5 shows superior performance across a wide range of tasks, demonstrating its strong capability of capturing underlying relations and properties of bio-entities.Our code is available at https://github.com/QizhiPei/BioT5. Qizhi Pei, Wei Zhang 0399, Jinhua Zhu 0001, Kehan Wu, Kaiyuan Gao, Lijun Wu 0003, Yingce Xia, Rui Yan 0001 |
EMNLP | 3 |
| 2023 | Making Better Decision by Directly Planning in Continuous Control
Jinhua Zhu 0001, Yue Wang 0017, Lijun Wu 0003, Tao Qin 0001, Wengang Zhou 0001, Tie-Yan Liu, Houqiang Li |
ICLR | 1 |
| 2023 | 𝒪-GNN: incorporating ring priors into molecular modeling
Jinhua Zhu 0001, Kehan Wu, Yingce Xia, Shufang Xie 0003, Lijun Wu 0003, Tao Qin 0001, Wengang Zhou 0001, Houqiang Li, Tie-Yan Liu |
ICLR | 1 |
| 2023 | Pre-training Antibody Language Models for Antigen-Specific Computational Antibody DesignabstractAntibodies are proteins that effectively protect the human body by binding to pathogens. Recently, deep learning-based computational antibody design has attracted popular attention since it automatically mines the antibody patterns from data that could be complementary to human experiences. However, the computational methods heavily rely on high-quality antibody structure data, which is quite limited. Besides, the complementarity-determining region (CDR), which is the key component of an antibody that determines the specificity and binding affinity, is highly variable and hard to predict. Therefore, the limited availability of high-quality antibody structure data exacerbates the difficulty of CDR generation. Fortunately, there is a large amount of sequence data for antibodies that can help model the CDR and reduce reliance on structure data. By witnessing the success of pre-training models for protein modeling, in this paper, we develop the antibody pre-training language model and incorporate it into the antigen-specific antibody design model in a systemic way. Specifically, we first pre-train a novel antibody language model based on the sequence data, then propose a one-shot way for sequence and structure generation of CDR to mitigate the high cost and error propagation associated with autoregressive methods, and finally leverage the pre-trained antibody model for the antigen-specific antibody generation model with some carefully designed modules. Our experiments demonstrate the superiority of our method over previous baselines in tasks such as sequence and structure generation, CDR-H3 design for antigen binding, and antibody optimization1. The code is available at https://github.com/KyGao/ABGNN. Kaiyuan Gao, Lijun Wu 0003, Jinhua Zhu 0001, Tianbo Peng, Yingce Xia, Liang He 0010, Shufang Xie 0003, Tao Qin 0001, Haiguang Liu, Kun He 0001, Tie-Yan Liu |
KDD | 3 |
| 2023 | Dual-view Molecular Pre-trainingabstractMolecular pre-training, which is about to learn an effective representation for molecules on large amount of data, has attracted substantial attention in cheminformatics and bioinformatics. A molecule can be viewed as either a graph (where atoms are connected by bonds) or a SMILES sequence (where depth-first-search is applied to the molecular graph with specific rules). The Transformer and graph neural networks (GNN) are two representative methods to deal with the sequential data and the graphic data, which can globally and locally model the molecules respectively and are supposed to be complementary. In this work, we propose to leverage both representations and design a new pre-training algorithm, dual-view molecule pre-training (briefly, DVMP), that can effectively combine the strengths of both types of molecule representations. DVMP has a Transformer branch and a GNN branch, and the two branches are pre-trained to maintain the semantic consistency of molecules. After pre-training, we can use either the Transformer branch (this one is recommended according to empirical results), the GNN branch, or both for downstream tasks. DVMP is tested on 11 molecular property prediction tasks and outperforms strong baselines. Furthermore, we test DVMP on three retrosynthesis tasks and it achieves state-of-the-art results. Our code is released at https://github.com/microsoft/DVMP. Jinhua Zhu 0001, Yingce Xia, Lijun Wu 0003, Shufang Xie 0003, Wengang Zhou 0001, Tao Qin 0001, Houqiang Li, Tie-Yan Liu |
KDD | 1 |
| 2023 | FABind: Fast and Accurate Protein-Ligand BindingabstractModeling the interaction between proteins and ligands and accurately predicting their binding structures is a critical yet challenging task in drug discovery. Recent advancements in deep learning have shown promise in addressing this challenge, with sampling-based and regression-based methods emerging as two prominent approaches. However, these methods have notable limitations. Sampling-based methods often suffer from low efficiency due to the need for generating multiple candidate structures for selection. On the other hand, regression-based methods offer fast predictions but may experience decreased accuracy. Additionally, the variation in protein sizes often requires external modules for selecting suitable binding pockets, further impacting efficiency. In this work, we propose FABind, an end-to-end model that combines pocket prediction and docking to achieve accurate and fast protein-ligand binding. FABind incorporates a unique ligand-informed pocket prediction module, which is also leveraged for docking pose estimation. The model further enhances the docking process by incrementally integrating the predicted pocket to optimize protein-ligand binding, reducing discrepancies between training and inference. Through extensive experiments on benchmark datasets, our proposed FABind demonstrates strong advantages in terms of effectiveness and efficiency compared to existing methods. Our code is available at https://github.com/QizhiPei/FABind. Qizhi Pei, Kaiyuan Gao, Lijun Wu 0003, Jinhua Zhu 0001, Yingce Xia, Shufang Xie 0003, Tao Qin 0001, Kun He 0001, Tie-Yan Liu, Rui Yan 0001 |
NeurIPS | 4 |
| 2023 | R2-DDI: relation-aware feature refinement for drug-drug interaction predictionabstractPrecisely predicting the drug-drug interaction (DDI) is an important application and host research topic in drug discovery, especially for avoiding the adverse effect when using drug combination treatment for patients. Nowadays, machine learning and deep learning methods have achieved great success in DDI prediction. However, we notice that most of the works ignore the importance of the relation type when building the DDI prediction models. In this work, we propose a novel R$^2$-DDI framework, which introduces a relation-aware feature refinement module for drug representation learning. The relation feature is integrated into drug representation and refined in the framework. With the refinement features, we also incorporate the consistency training method to regularize the multi-branch predictions for better generalization. Through extensive experiments and studies, we demonstrate our R$^2$-DDI approach can significantly improve the DDI prediction performance over multiple real-world datasets and settings, and our method shows better generalization ability with the help of the feature refinement design. Jiacheng Lin, Lijun Wu 0003, Jinhua Zhu 0001, Xiaobo Liang, Yingce Xia, Shufang Xie 0003, Tao Qin 0001, Tie-Yan Liu |
Briefings Bioinform. | 3 |
| 2023 | Breaking the barriers of data scarcity in drug-target affinity predictionabstractAccurate prediction of drug-target affinity (DTA) is of vital importance in early-stage drug discovery, facilitating the identification of drugs that can effectively interact with specific targets and regulate their activities. While wet experiments remain the most reliable method, they are time-consuming and resource-intensive, resulting in limited data availability that poses challenges for deep learning approaches. Existing methods have primarily focused on developing techniques based on the available DTA data, without adequately addressing the data scarcity issue. To overcome this challenge, we present the Semi-Supervised Multi-task training (SSM) framework for DTA prediction, which incorporates three simple yet highly effective strategies: (1) A multi-task training approach that combines DTA prediction with masked language modeling using paired drug-target data. (2) A semi-supervised training method that leverages large-scale unpaired molecules and proteins to enhance drug and target representations. This approach differs from previous methods that only employed molecules or proteins in pre-training. (3) The integration of a lightweight cross-attention module to improve the interaction between drugs and targets, further enhancing prediction accuracy. Through extensive experiments on benchmark datasets such as BindingDB, DAVIS and KIBA, we demonstrate the superior performance of our framework. Additionally, we conduct case studies on specific drug-target binding activities, virtual screening experiments, drug feature visualizations and real-world applications, all of which showcase the significant potential of our work. In conclusion, our proposed SSM-DTA framework addresses the data limitation challenge in DTA prediction and yields promising results, paving the way for more efficient and accurate drug discovery processes. Qizhi Pei, Lijun Wu 0003, Jinhua Zhu 0001, Yingce Xia, Shufang Xie 0003, Tao Qin 0001, Haiguang Liu, Tie-Yan Liu, Rui Yan 0001 |
Briefings Bioinform. | 3 |
| 2023 | Masked Contrastive Representation Learning for Reinforcement LearningabstractIn pixel-based reinforcement learning (RL), the states are raw video frames, which are mapped into hidden representation before feeding to a policy network. To improve sample efficiency of state representation learning, recently, the most prominent work is based on contrastive unsupervised representation. Witnessing that consecutive video frames in a game are highly correlated, to further improve data efficiency, we propose a new algorithm, i.e., masked contrastive representation learning for RL (M-CURL), which takes the correlation among consecutive inputs into consideration. In our architecture, besides a CNN encoder for hidden presentation of input state and a policy network for action selection, we introduce an auxiliary Transformer encoder module to leverage the correlations among video frames. During training, we randomly mask the features of several frames, and use the CNN encoder and Transformer to reconstruct them based on context frames. The CNN encoder and Transformer are jointly trained via contrastive learning where the reconstructed features should be similar to the ground-truth ones while dissimilar to others. During policy evaluation, the CNN encoder and the policy network are used to take actions, and the Transformer module is discarded. Our method achieves consistent improvements over CURL on 14 out of 16 environments from DMControl suite and 23 out of 26 environments from Atari 2600 Games. The code is available at https://github.com/teslacool/m-curl. Jinhua Zhu 0001, Yingce Xia, Lijun Wu 0003, Jiajun Deng, Wengang Zhou 0001, Tao Qin 0001, Tie-Yan Liu, Houqiang Li |
IEEE Trans. Pattern Anal. Mach. Intell. | 1 |
| 2022 | Unified 2D and 3D Pre-Training of Molecular RepresentationsabstractMolecular representation learning has attracted much attention recently. A molecule can be viewed as a 2D graph with nodes/atoms connected by edges/bonds, and can also be represented by a 3D conformation with 3-dimensional coordinates of all atoms. We note that most previous work handles 2D and 3D information separately, while jointly leveraging these two sources may foster a more informative representation. In this work, we explore this appealing idea and propose a new representation learning method based on a unified 2D and 3D pre-training. Atom coordinates and interatomic distances are encoded and then fused with atomic representations through graph neural networks. The model is pre-trained on three tasks: reconstruction of masked atoms and coordinates, 3D conformation generation conditioned on 2D graph, and 2D graph generation conditioned on 3D conformation. We evaluate our method on 11 downstream molecular property prediction tasks: 7 with 2D information only and 4 with both 2D and 3D information. Our method achieves state-of-the-art results on 10 tasks, and the average improvement on 2D-only tasks is 8.3%. Our method also achieves significant improvement on two 3D conformation generation tasks. Jinhua Zhu 0001, Yingce Xia, Lijun Wu 0003, Shufang Xie 0003, Tao Qin 0001, Wengang Zhou 0001, Houqiang Li, Tie-Yan Liu |
KDD | 1 |
| 2022 | SPRoBERTa: protein embedding learning with local fragment modelingabstractWell understanding protein function and structure in computational biology helps in the understanding of human beings. To face the limited proteins that are annotated structurally and functionally, the scientific community embraces the self-supervised pre-training methods from large amounts of unlabeled protein sequences for protein embedding learning. However, the protein is usually represented by individual amino acids with limited vocabulary size (e.g. 20 type proteins), without considering the strong local semantics existing in protein sequences. In this work, we propose a novel pre-training modeling approach SPRoBERTa. We first present an unsupervised protein tokenizer to learn protein representations with local fragment pattern. Then, a novel framework for deep pre-training model is introduced to learn protein embeddings. After pre-training, our method can be easily fine-tuned for different protein tasks, including amino acid-level prediction task (e.g. secondary structure prediction), amino acid pair-level prediction task (e.g. contact prediction) and also protein-level prediction task (remote homology prediction, protein function prediction). Experiments show that our approach achieves significant improvements in all tasks and outperforms the previous methods. We also provide detailed ablation studies and analysis for our protein tokenizer and training framework. Lijun Wu 0003, Chengcan Yin, Jinhua Zhu 0001, Zhen Wu 0002, Liang He 0010, Yingce Xia, Shufang Xie 0003, Tao Qin 0001, Tie-Yan Liu |
Briefings Bioinform. | 3 |
| 2022 | Back translation for molecule generationabstractMOTIVATION: Molecule generation, which is to generate new molecules, is an important problem in bioinformatics. Typical tasks include generating molecules with given properties, molecular property improvement (i.e. improving specific properties of an input molecule), retrosynthesis (i.e. predicting the molecules that can be used to synthesize a target molecule), etc. Recently, deep-learning-based methods received more attention for molecule generation. The labeled data of bioinformatics is usually costly to obtain, but there are millions of unlabeled molecules. Inspired by the success of sequence generation in natural language processing with unlabeled data, we would like to explore an effective way of using unlabeled molecules for molecule generation. RESULTS: We propose a new method, back translation for molecule generation, which is a simple yet effective semisupervised method. Let X be the source domain, which is the collection of properties, the molecules to be optimized, etc. Let Y be the target domain which is the collection of molecules. In particular, given a main task which is about to learn a mapping from the source domain X to the target domain Y, we first train a reversed model g for the Y to X mapping. After that, we use g to back translate the unlabeled data in Y to X and obtain more synthetic data. Finally, we combine the synthetic data with the labeled data and train a model for the main task. We conduct experiments on molecular property improvement and retrosynthesis, and we achieve state-of-the-art results on four molecule generation tasks and one retrosynthesis benchmark, USPTO-50k. AVAILABILITY AND IMPLEMENTATION: Our code and data are available at https://github.com/fyabc/BT4MolGen. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yingce Xia, Jinhua Zhu 0001, Lijun Wu 0003, Shufang Xie 0003, Tao Qin 0001 |
Bioinform. | 3 |
| 2022 | Discovering drug-target interaction knowledge from biomedical literatureabstractMOTIVATION: The interaction between drugs and targets (DTI) in human body plays a crucial role in biomedical science and applications. As millions of papers come out every year in the biomedical domain, automatically discovering DTI knowledge from biomedical literature, which are usually triplets about drugs, targets and their interaction, becomes an urgent demand in the industry. Existing methods of discovering biological knowledge are mainly extractive approaches that often require detailed annotations (e.g. all mentions of biological entities, relations between every two entity mentions, etc.). However, it is difficult and costly to obtain sufficient annotations due to the requirement of expert knowledge from biomedical domains. RESULTS: To overcome these difficulties, we explore an end-to-end solution for this task by using generative approaches. We regard the DTI triplets as a sequence and use a Transformer-based model to directly generate them without using the detailed annotations of entities and relations. Further, we propose a semi-supervised method, which leverages the aforementioned end-to-end model to filter unlabeled literature and label them. Experimental results show that our method significantly outperforms extractive baselines on DTI discovery. We also create a dataset, KD-DTI, to advance this task and release it to the community. AVAILABILITY AND IMPLEMENTATION: Our code and data are available at https://github.com/bert-nmt/BERT-DTI. SUPPLEMENTARY INFORMATION: Supplementary data are available at Bioinformatics online. Yutai Hou, Yingce Xia, Lijun Wu 0003, Shufang Xie 0003, Jinhua Zhu 0001, Tao Qin 0001, Tie-Yan Liu |
Bioinform. | 6 |
| 2022 | A study of BERT for context-aware neural machine translation
Xueqing Wu 0001, Yingce Xia, Jinhua Zhu 0001, Lijun Wu 0003, Shufang Xie 0003, Tao Qin 0001 |
Mach. Learn. | 3 |
| 2022 | Multimodal Sentiment Analysis With Two-Phase Multi-Task LearningabstractMultimodal Sentiment Analysis (MSA) is a challenging research area that studies sentiment expressed from multiple heterogeneous modalities. Given those pre-trained language models such as BERT have shown state-of-the-art (SOTA) performance in multiple NLP disciplines, existing models tend to integrate these modalities into BERT and treat the MSA as a single prediction task. However, we find that simply fusing the multimodal features into BERT cannot well establish the power of a strong pre-trained model. Besides, the classification ability of each modality is also suppressed by single-task learning. In this paper, we propose a multimodal framework named Two-Phase Multi-task Sentiment Analysis (TPMSA), which applies a two-phase training strategy to make the most of the pre-trained model and a novel multi-task learning strategy to investigate the classification ability of each representation. We conduct experiments on two multimodal benchmark datasets CMU-MOSI and CMU-MOSEI, the results show that our TPMSA model significantly outperforms the current SOTA method on both datasets across all the metrics, which clearly show the effectiveness of our proposed method. Our code is available at \url{https://github.com/TPMSA/TPMSA}. Lijun Wu 0003, Jinhua Zhu 0001, Xiaola Lin, Tie-Yan Liu |
IEEE ACM Trans. Audio Speech Lang. Process. | 3 |
| 2021 | Return-Based Contrastive Representation Learning for Reinforcement Learning
Chuheng Zhang, Li Zhao 0007, Tao Qin 0001, Jinhua Zhu 0001, Jian Li 0015, Nenghai Yu, Tie-Yan Liu |
ICLR | 5 |
| 2021 | IOT: Instance-wise Layer Reordering for Transformer Structures
Jinhua Zhu 0001, Lijun Wu 0003, Yingce Xia, Shufang Xie 0003, Tao Qin 0001, Wengang Zhou 0001, Houqiang Li, Tie-Yan Liu |
ICLR | 1 |
| 2020 | Incorporating BERT into Neural Machine Translation
Jinhua Zhu 0001, Yingce Xia, Lijun Wu 0003, Di He 0001, Tao Qin 0001, Wengang Zhou 0001, Houqiang Li, Tie-Yan Liu |
ICLR | 1 |
| 2019 | Soft Contextual Data Augmentation for Neural Machine TranslationabstractWhile data augmentation is an important trick to boost the accuracy of deep learning methods in computer vision tasks, its study in natural language tasks is still very limited.In this paper, we present a novel data augmentation method for neural machine translation.Different from previous augmentation methods that randomly drop, swap or replace words with other words in a sentence, we softly augment a randomly chosen word in a sentence by its contextual mixture of multiple related words.More accurately, we replace the onehot representation of a word by a distribution (provided by a language model) over the vocabulary, i.e., replacing the embedding of this word by a weighted combination of multiple semantically similar words.Since the weights of those words depend on the contextual information of the word to be replaced, the newly generated sentences capture much richer information than previous augmentation methods.Experimental results on both small scale and large scale machine translation datasets demonstrate the superiority of our method over strong baselines 1 . Fei Gao 0018, Jinhua Zhu 0001, Lijun Wu 0003, Yingce Xia, Tao Qin 0001, Xueqi Cheng 0001, Wengang Zhou 0001, Tie-Yan Liu |
ACL (1) | 2 |
| 2019 | Machine Translation With Weakly Paired DocumentsabstractLijun Wu, Jinhua Zhu, Di He, Fei Gao, Tao Qin, Jianhuang Lai, Tie-Yan Liu. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019. Lijun Wu 0003, Jinhua Zhu 0001, Di He 0001, Fei Gao 0018, Tao Qin 0001, Jian-Huang Lai, Tie-Yan Liu |
EMNLP/IJCNLP (1) | 2 |
| 2019 | Normalization Helps Training of Quantized LSTMabstractThe long-short-term memory (LSTM), though powerful, is memory and computa\x02tion expensive. To alleviate this problem, one approach is to compress its weights by quantization. However, existing quantization methods usually have inferior performance when used on LSTMs. In this paper, we first show theoretically that training a quantized LSTM is difficult because quantization makes the exploding gradient problem more severe, particularly when the LSTM weight matrices are large. We then show that the popularly used weight/layer/batch normalization schemes can help stabilize the gradient magnitude in training quantized LSTMs. Empirical results show that the normalized quantized LSTMs achieve significantly better results than their unnormalized counterparts. Their performance is also comparable with the full-precision LSTM, while being much smaller in size. Lu Hou 0002, Jinhua Zhu 0001, James T. Kwok, Fei Gao 0018, Tao Qin 0001, Tie-Yan Liu |
NeurIPS | 2 |