VLDB 2026 Research / reviewers in the wild / expert
Hao Zhou 0012
dblp:63/778-12
· DBLP profile ↗
113ranked-venue papers
16as first author
74since 2021 · last 2026
0000-0002-1458-1016ORCID · conflict
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 107 · 13 first-author · 71 since 2021Graphics, computer vision, multimedia, augmented reality and games · 17 · 3 first-author · 7 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 4 since 2021Computer networks · 1 · 1 first-author · 1 since 2021Software engineering, systems software and programming languages · 1 · 1 since 2021Applied, interdisciplinary, general and emerging computing · 1 · 1 first-author
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | APB-V: Accelerating Long-Video Understanding via Sequence-Parallelism-aware Approximate AttentionabstractYuxiang Huang, Mingye Li, Xu Han, Chaojun Xiao, Weilin Zhao, Ao Sun, Ziqi Yuan, Hao Zhou, Fandong Meng, Zhiyuan Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Yuxiang Huang 0001, Mingye Li, Xu Han 0007, Chaojun Xiao, Weilin Zhao, Hao Zhou 0012, Fandong Meng, Zhiyuan Liu 0001 |
ACL (1) | 8 |
| 2026 | R³: End-to-End Reasoning-based Planning for Multi-step Retrosynthesis via Reinforcement LearningabstractYiFei Wang, Qizhi Pei, Jiangtao Feng, Yuntian Shi, Yi Duan, Lihao Wang, Lei Bai, Lijun Wu, Wei-Ying Ma, Hao Zhou. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. YiFei Wang, Qizhi Pei, Jiangtao Feng, Yuntian Shi, Yi Duan, Lei Bai 0001, Lijun Wu 0003, Wei-Ying Ma, Hao Zhou 0012 |
ACL (1) | 10 |
| 2026 | Investigating Cross-Modal Skill Injection: Scenarios, Methods, and HyperparametersabstractZhiyu Xu, Lean Wang, Yuanxin Liu, Lei Li, Hao Zhou, Fandong Meng, Jie Zhou, Xu Sun. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Lean Wang, Yuanxin Liu, Lei Li 0009, Hao Zhou 0012, Fandong Meng, Jie Zhou 0016, Xu Sun 0001 |
ACL (1) | 5 |
| 2026 | A Data-Efficient Path to Multilingual LLMs: Language Expansion via Post-training PARAMΔ Integration into Upcycled MoEabstractHao Zhou, Tianhao Li, Zhijun Wang, Shuaijie She, Linjuan Wu, Hao-Ran Wei, Baosong Yang, Jiajun Chen, Shujian Huang. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026. Hao Zhou 0012, Shuaijie She, Linjuan Wu, Baosong Yang, Jiajun Chen 0001, Shujian Huang |
ACL (1) | 1 |
| 2026 | Delay-Doppler Domain Channel Measurements and Modeling in High-Speed RailwaysabstractAs next-generation wireless communication systems need to be able to operate in high-frequency bands and high-mobility scenarios, delay-Doppler (DD) domain multicarrier (DDMC) modulation schemes, such as orthogonal time frequency space (OTFS), demonstrate superior reliability over orthogonal frequency division multiplexing (OFDM). Accurate DD domain channel modeling is essential for DDMC system design. However, since traditional channel modeling approaches are mainly confined to time, frequency, and space domains, the principles of DD domain channel modeling remain poorly studied. To address this issue, we propose a systematic DD domain channel measurement and modeling methodology in high-speed railway (HSR) scenarios. First, we design a DD domain channel measurement method based on the long-term evolution for railway (LTE-R) system. Second, for DD domain channel modeling, we investigate quasi-stationary interval, statistical power modeling of multipath components, and particularly, the quasi-invariant intervals of DD domain channel fading coefficients. Third, via LTE-R measurements at 371 km/h, taking the quasi-stationary interval as the decision criterion, we establish DD domain channel models under different channel time-varying conditions in HSR scenarios. Fourth, the accuracy of proposed DD domain channel models is validated via bit error rate comparison of OTFS transmission. In addition, simulation verifies that in HSR scenario, the quasi-invariant interval of DD domain channel fading coefficient is on millisecond (ms) order of magnitude, which is much smaller than the quasi-stationary interval length on 100 ms order of magnitude. This study could provide theoretical guidance for DD domain modeling in high-mobility environments, supporting future DDMC and integrated sensing and communication designs for 6G and beyond. Hao Zhou 0012, Yiyan Ma, Dan Fei, Mi Yang 0001, Ruisi He, Bo Ai 0001 |
IEEE Trans. Wirel. Commun. | 1 |
| 2025 | MoE-LPR: Multilingual Extension of Large Language Models Through Mixture-of-Experts with Language Priors RoutingabstractLarge Language Models (LLMs) are often English-centric due to the disproportionate distribution of languages in their pre-training data. Enhancing non-English language capabilities through post-pretraining often results in catastrophic forgetting of high-resource languages. Previous methods either achieve good expansion with severe forgetting or slight forgetting with poor expansion, indicating the challenge of balancing language expansion while preventing forgetting. In this paper, we propose a method called MoE-LPR (Mixture-of-Experts with Language Priors Routing) to alleviate this problem. MoE-LPR employs a two-stage training approach to enhance the multilingual capability. First, the model is post-pretrained into a Mixture-of-Experts(MoE) architecture by upcycling, where all the original parameters are frozen and new experts are added. In this stage, we focus improving the ability on expanded languages, without using any original language data. Then, the model reviews the knowledge of the original languages with replay data amounting to less than 1% of post-pretraining, where we incorporate language priors routing to better recover the abilities of the original languages. Evaluations on multiple benchmarks show that MoE-LPR outperforms other post-pretraining methods. Freezing original parameters preserves original language knowledge while adding new experts preserves the learning ability. Reviewing with LPR enables effective utilization of multilingual knowledge within the parameters. Additionally, the MoE architecture maintains the same inference overhead while increasing total model parameters. Extensive experiments demonstrate MoE-LPR’s effectiveness in improving expanded languages and preserving original language proficiency with superior scalability. Hao Zhou 0012, Shujian Huang, Xue Han 0018, Junlan Feng, Chao Deng 0002, Weihua Luo, Jiajun Chen 0001 |
AAAI | 1 |
| 2025 | APB: Accelerating Distributed Long-Context Inference by Passing Compressed Context Blocks across GPUsabstractWhile long-context inference is crucial for advancing large language model (LLM) applications, its prefill speed remains a significant bottleneck. Current approaches, including sequence parallelism strategies and compute reduction through approximate attention mechanisms, still fall short of delivering optimal inference efficiency. This hinders scaling the inputs to longer sequences and processing long-context queries in a timely manner. To address this, we introduce APB, an efficient long-context inference framework that leverages multi-host approximate attention to enhance prefill speed by reducing compute and enhancing parallelism simultaneously. APB introduces a communication mechanism for essential key-value pairs within a sequence parallelism framework, enabling a faster inference speed while maintaining task performance. We implement APB by incorporating a tailored FlashAttn kernel alongside optimized distribution strategies, supporting diverse models and parallelism configurations. APB achieves speedups of up to 9.2\times, 4.2\times, and 1.6\times compared with FlashAttn, RingAttn, and StarAttn, respectively, without any observable task performance degradation. Yuxiang Huang 0001, Mingye Li, Xu Han 0007, Chaojun Xiao, Weilin Zhao, Sun Ao, Hao Zhou 0012, Jie Zhou 0016, Zhiyuan Liu 0001, Maosong Sun 0001 |
ACL (1) | 7 |
| 2025 | PunchBench: Benchmarking MLLMs in Multimodal Punchline ComprehensionabstractKun Ouyang, Yuanxin Liu, Shicheng Li, Yi Liu, Hao Zhou, Fandong Meng, Jie Zhou, Xu Sun. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Kun Ouyang, Yuanxin Liu, Hao Zhou 0012, Fandong Meng, Jie Zhou 0016, Xu Sun 0001 |
ACL (1) | 5 |
| 2025 | FR-Spec: Accelerating Large-Vocabulary Language Models via Frequency-Ranked Speculative SamplingabstractWeilin Zhao, Tengyu Pan, Xu Han, Yudi Zhang, Sun Ao, Yuxiang Huang, Kaihuo Zhang, Weilun Zhao, Yuxuan Li, Jie Zhou, Hao Zhou, Jianyong Wang, Maosong Sun, Zhiyuan Liu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025. Weilin Zhao, Tengyu Pan, Xu Han 0007, Sun Ao, Yuxiang Huang 0001, Kaihuo Zhang, Wei-Lun Zhao, Jie Zhou 0016, Hao Zhou 0012, Jianyong Wang 0001, Maosong Sun 0001, Zhiyuan Liu 0001 |
ACL (1) | 11 |
| 2025 | RetroDiff: Retrosynthesis as Multi-stage Distribution InterpolationabstractRetrosynthesis poses a key challenge in biopharmaceuticals, aiding chemists in finding appropriate reactant molecules for given product molecules. With reactants and products represented as 2D graphs, retrosynthesis constitutes a conditional graph-to-graph (G2G) generative task. Inspired by advancements in discrete diffusion models for graph generation, we aim to design a diffusion-based method to address this problem. However, integrating a diffusion-based G2G framework while retaining essential chemical reaction template information presents a notable challenge. Our key innovation involves a multi-stage diffusion process. We decompose the retrosynthesis procedure to first sample external groups from the dummy distribution given products, then generate external bonds to connect products and generated groups. Interestingly, this generation process mirrors the reverse of the widely adapted semi-template retrosynthesis workflow, i.e., from reaction center identification to synthon completion. Based on these designs, we introduce Retrosynthesis Diffusion (RetroDiff), a novel diffusion-based method for the retrosynthesis task. Experimental results demonstrate that RetroDiff surpasses all semi-template methods in accuracy, and outperforms template-based and template-free methods in large-scale scenarios and molecular validity, respectively. Yiming Wang 0011, Yuxuan Song 0002, Minkai Xu, Rui Wang 0015, Hao Zhou 0012, Wei-Ying Ma |
AISTATS | 6 |
| 2025 | Steering Protein Family Design through Profile Bayesian FlowabstractProtein family design emerges as a promising alternative by combining the advantages of de novo protein design and mutation-based directed evolution.In this paper, we propose ProfileBFN, the Profile Bayesian Flow Networks, for specifically generative modeling of protein families. ProfileBFN extends the discrete Bayesian Flow Network from an MSA profile perspective, which can be trained on single protein sequences by regarding it as a degenerate profile, thereby achieving efficient protein family design by avoiding large-scale MSA data construction and training. Empirical results show that ProfileBFN has a profound understanding of proteins. When generating diverse and novel family proteins, it can accurately capture the structural characteristics of the family. The enzyme produced by this method is more likely than the previous approach to have the corresponding function, offering better odds of generating diverse proteins with the desired functionality. Jingjing Gong, Siyu Long, Yuxuan Song 0002, Wenhao Huang 0001, Ziyao Cao, Hao Zhou 0012, Wei-Ying Ma |
ICLR | 9 |
| 2025 | MiniPLM: Knowledge Distillation for Pre-training Language ModelsabstractKnowledge distillation (KD) is widely used to train small, high-performing student language models (LMs) using large teacher LMs.
While effective in fine-tuning, KD during pre-training faces efficiency, flexibility, and effectiveness issues.
Existing methods either incur high computational costs due to online teacher inference, require tokenization matching between teacher and student LMs, or risk losing the difficulty and diversity of the teacher-generated training data.
In this work, we propose **MiniPLM**, a KD framework for pre-training LMs by refining the training data distribution with the teacher LM's knowledge.
For efficiency, MiniPLM performs offline teacher inference, allowing KD for multiple student LMs without adding training costs.
For flexibility, MiniPLM operates solely on the training corpus, enabling KD across model families.
For effectiveness, MiniPLM leverages the differences between large and small LMs to enhance the training data difficulty and diversity, helping student LMs acquire versatile and sophisticated knowledge.
Extensive experiments demonstrate that MiniPLM boosts the student LMs' performance on 9 common downstream tasks, improves language modeling capabilities, and reduces pre-training computation.
The benefit of MiniPLM extends to larger training scales, evidenced by the scaling curve extrapolation.
Further analysis reveals that MiniPLM supports KD across model families and enhances the pre-training data utilization. Our code, data, and models can be found at https://github.com/thu-coai/MiniPLM. Yuxian Gu, Hao Zhou 0012, Fandong Meng, Jie Zhou 0016, Minlie Huang |
ICLR | 2 |
| 2025 | A Periodic Bayesian Flow for Material GenerationabstractGenerative modeling of crystal data distribution is an important yet challenging task due to the unique periodic physical symmetry of crystals. Diffusion-based methods have shown early promise in modeling crystal distribution. More recently, Bayesian Flow Networks were introduced to aggregate noisy latent variables, resulting in a variance-reduced parameter space that has been shown to be advantageous for modeling Euclidean data distributions with structural constraints (Song, et al.,2023). Inspired by this, we seek to unlock its potential for modeling variables located in non-Euclidean manifolds e.g. those within crystal structures, by overcoming challenging theoretical issues. We introduce CrysBFN, a novel crystal generation method by proposing a periodic Bayesian flow, which essentially differs from the original Gaussian-based BFN by exhibiting non-monotonic entropy dynamics. To successfully realize the concept of periodic Bayesian flow, CrysBFN integrates a new entropy conditioning mechanism and empirically demonstrates its significance compared to time-conditioning. Extensive experiments over both crystal ab initio generation and crystal structure prediction tasks demonstrate the superiority of CrysBFN, which consistently achieves new state-of-the-art on all benchmarks. Surprisingly, we found that CrysBFN enjoys a significant improvement in sampling efficiency, e.g., 200x speedup (10 v.s. 2000 steps network forwards) compared with previous Diffusion-based methods on MP-20 dataset. Yuxuan Song 0002, Jingjing Gong, Ziyao Cao, Yawen Ouyang, Hao Zhou 0012, Wei-Ying Ma |
ICLR | 7 |
| 2025 | Piloting Structure-Based Drug Design via Modality-Specific Optimal ScheduleabstractStructure-Based Drug Design (SBDD) is crucial for identifying bioactive molecules. Recent deep generative models are faced with challenges in geometric structure modeling. A major bottleneck lies in the twisted probability path of multi-modalities—continuous 3D positions and discrete 2D topologies—which jointly determine molecular geometries. By establishing the fact that noise schedules decide the Variational Lower Bound (VLB) for the twisted probability path, we propose VLB-Optimal Scheduling (VOS) strategy in this under-explored area, which optimizes VLB as a path integral for SBDD. Our model effectively enhances molecular geometries and interaction modeling, achieving state-of-the-art PoseBusters passing rate of 95.9\% on CrossDock, more than 10\% improvement upon strong baselines, while maintaining high affinities and robust intramolecular validity evaluated on held-out test set. Code is available at https://github.com/AlgoMole/MolCRAFT. Keyue Qiu, Yuxuan Song 0002, Zhehuan Fan, Mingyue Zheng, Hao Zhou 0012, Wei-Ying Ma |
ICML | 7 |
| 2025 | Empower Structure-Based Molecule Optimization with Gradient Guided Bayesian Flow NetworksabstractStructure-based molecule optimization (SBMO) aims to optimize molecules with both continuous coordinates and discrete types against protein targets.
A promising direction is to exert gradient guidance on generative models given its remarkable success in images, but it is challenging to guide discrete data and risks inconsistencies between modalities.
To this end, we leverage a continuous and differentiable space derived through Bayesian inference, presenting Molecule Joint Optimization (MolJO), the gradient-based SBMO framework that facilitates joint guidance signals across different modalities while preserving SE(3)-equivariance.
We introduce a novel backward correction strategy that optimizes within a sliding window of the past histories, allowing for a seamless trade-off between explore-and-exploit during optimization.
MolJO achieves state-of-the-art performance on CrossDocked2020 benchmark (Success Rate 51.3\%, Vina Dock -9.05 and SA 0.78), more than 4x improvement in Success Rate compared to the gradient-based counterpart, and 2x ``Me-Better'' Ratio as much as 3D baselines.
Furthermore, we extend MolJO to a wide range of optimization settings, including multi-objective optimization and challenging tasks in drug design such as R-group optimization and scaffold hopping, further underscoring its versatility.
Code is available at https://github.com/AlgoMole/MolCRAFT. Keyue Qiu, Yuxuan Song 0002, Hongbo Ma, Ziyao Cao, Yushuai Wu, Mingyue Zheng, Hao Zhou 0012, Wei-Ying Ma |
ICML | 9 |
| 2025 | Smooth Interpolation for Improved Discrete Graph Generative ModelsabstractThough typically represented by the discrete node and edge attributes, the graph topological information can be sufficiently captured by the graph spectrum in a continuous space. It is believed that incorporating the continuity of graph topological information into the generative process design could establish a superior paradigm for graph generative modeling. Motivated by such prior and recent advancements in the generative paradigm, we propose Graph Bayesian Flow Networks (GraphBFN) in this paper, a principled generative framework that designs an alternative generative process emphasizing the dynamics of topological information. Unlike recent discrete-diffusion-based methods, GraphBFNemploys the continuous counts derived from sampling infinite times from a categorical distribution as latent to facilitate a smooth decomposition of topological information, demonstrating enhanced effectiveness. To effectively realize the concept, we further develop an advanced sampling strategy and new time-scheduling techniques to overcome practical barriers and boost performance. Through extensive experimental validation on both generic graph and molecular graph generation tasks, GraphBFN could consistently achieve superior or competitive performance with significantly higher training and sampling efficiency. Yuxuan Song 0002, Juntong Shi, Jingjing Gong, Minkai Xu, Stefano Ermon, Hao Zhou 0012, Wei-Ying Ma |
ICML | 6 |
| 2025 | SToFM: a Multi-scale Foundation Model for Spatial TranscriptomicsabstractSpatial Transcriptomics (ST) technologies provide biologists with rich insights into single-cell biology by preserving spatial context of cells.
Building foundational models for ST can significantly enhance the analysis of vast and complex data sources, unlocking new perspectives on the intricacies of biological tissues.
However, modeling ST data is inherently challenging due to the need to extract multi-scale information from tissue slices containing vast numbers of cells. This process requires integrating macro-scale tissue morphology, micro-scale cellular microenvironment, and gene-scale gene expression profile.
To address this challenge, we propose **SToFM**, a multi-scale **S**patial **T**ranscript**o**mics **F**oundation **M**odel.
SToFM first performs multi-scale information extraction on each ST slice, to construct a set of ST sub-slices that aggregate macro-, micro- and gene-scale information. Then an SE(2) Transformer is used to obtain high-quality cell representations from the sub-slices.
Additionally, we construct **SToCorpus-88M**, the largest high-resolution spatial transcriptomics corpus for pretraining.
SToFM achieves outstanding performance on a variety of downstream tasks, such as tissue region semantic segmentation and cell type annotation, demonstrating its comprehensive understanding of ST data through capturing and integrating multi-scale information. Suyuan Zhao, Yizhen Luo, Ganbo Yang, Hao Zhou 0012, Zaiqing Nie |
ICML | 5 |
| 2025 | Vision-Language Models Can Self-Improve Reasoning via ReflectionabstractKanzhi Cheng, Li YanTao, Fangzhi Xu, Jianbing Zhang, Hao Zhou, Yang Liu. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025. Kanzhi Cheng, Yantao Li 0003, Fangzhi Xu, Hao Zhou 0012, Yang Liu 0005 |
NAACL (Long Papers) | 5 |
| 2025 | Enigmata: Scaling Logical Reasoning in Large Language Models with Synthetic Verifiable PuzzlesabstractLarge Language Models (LLMs), such as OpenAI’s o1 and DeepSeek’s R1, excel at advanced reasoning tasks like math and coding via Reinforcement Learning with Verifiable Rewards (RLVR), but still struggle with puzzles solvable by humans without domain knowledge. We introduce ENIGMATA, the first comprehensive suite tailored for improving LLMs with puzzle reasoning skills. It includes 36 tasks across 7 categories, each with: 1) a generator that produces unlimited examples with controllable difficulty, and 2) a rule-based verifier for automatic evaluation. This generator-verifier design supports scalable, multi-task RL training, fine-grained analysis, and seamless RLVR integration. We further propose ENIGMATA-Eval, a rigorous benchmark, and develop optimized multi-task RLVR strategies. Our trained model, Qwen2.5-32B-ENIGMATA, consistently surpasses o3-mini-high and o1 on the puzzle reasoning benchmarks like ENIGMATA-Eval, ARC-AGI (32.8%), and ARC-AGI 2 (0.6%). It also generalizes well to out-of-domain puzzle benchmarks and mathematical reasoning, with little multi-tasking trade-off. When trained on larger models like Seed1.5-Thinking (20B activated parameters and 200B total parameters), puzzle data from ENIGMATA further boosts SoTA performance on advanced math and STEM reasoning tasks such as AIME (2024-2025), BeyondAIME and GPQA (Diamond), showing nice generalization benefits of ENIGMATA. This work offers a unified, controllable framework for advancing logical reasoning in LLMs. Project page: https://seed-enigmata.github.io. Jiangjie Chen, Qianyu He, Aili Chen, Zhicheng Cai, Weinan Dai, Hongli Yu, Jiaze Chen, Qiying Yu, Hao Zhou 0012, Mingxuan Wang |
NeurIPS | 11 |
| 2025 | MOF-BFN: Metal-Organic Frameworks Structure Prediction via Bayesian Flow NetworksabstractMetal-Organic Frameworks (MOFs) have attracted considerable attention due to their unique properties including high surface area and tunable porosity, and promising applications in catalysis, gas storage, and drug delivery. Structure prediction for MOFs is a challenging task, as these frameworks are intrinsically periodic and hierarchically organized, where the entire structure is assembled from building blocks like metal nodes and organic linkers. To address this, we introduce MOF-BFN, a novel generative model for MOF structure prediction based on Bayesian Flow Networks (BFNs). Given the local geometry of building blocks, MOF-BFN jointly predicts the lattice parameters, as well as the positions and orientations of all building blocks within the unit cell. In particular, the positions are modelled in the fractional coordinate system to naturally incorporate the periodicity. Meanwhile, the orientations are modeled as unit quaternions sampled from learned Bingham distributions via the proposed Bingham BFN, enabling effective orientation generation on the 4D unit hypersphere. Experimental results demonstrate that MOF-BFN achieves state-of-the-art performance across multiple tasks, including structure prediction, geometric property evaluation, and de novo generation, offering a promising tool for designing complex MOF materials. Wenbing Huang 0001, Yuxuan Song 0002, Yawen Ouyang, Yu Rong 0001, Tingyang Xu, Hao Zhou 0012, Wei-Ying Ma, Yang Liu 0005 |
NeurIPS | 9 |
| 2025 | Retro-R1: LLM-based Agentic RetrosynthesisabstractRetrosynthetic planning is a fundamental task in chemical discovery. Due to the vast combinatorial search space, identifying viable synthetic routes remains a significant challenge--even for expert chemists. Recent advances in Large Language Models (LLMs), particularly equipped with reinforcement learning, have demonstrated strong human-like reasoning and planning abilities, especially in mathematics and code problem solving. This raises a natural question: Can the reasoning capabilities of LLMs be harnessed to develop an AI chemist capable of learning effective policies for multi-step retrosynthesis? In this study, we introduce Retro-R1, a novel LLM-based retrosynthesis agent trained via reinforcement learning to design molecular synthesis pathways. Unlike prior approaches, which typically rely on single-turn, question-answering formats, Retro-R1 interacts dynamically with plug-in single-step retrosynthesis tools and learns from environmental feedback. Experimental results show that Retro-R1 achieves a 55.79\% pass@1 success rate, surpassing the previous state of the art by 8.95\%. Notably, Retro-R1 demonstrates strong generalization to out-of-domain test cases, where existing methods tend to fail despite their high in-domain performance. Our work marks a significant step toward equipping LLMs with advanced, chemist-like reasoning abilities, highlighting the promise of reinforcement learning for enabling data-efficient, generalizable, and sophisticated scientific problem-solving in LLM-based agents. Jiangtao Feng, Hongli Yu, Yuxuan Song 0002, Shufei Zhang, Lei Bai 0001, Wei-Ying Ma, Hao Zhou 0012 |
NeurIPS | 9 |
| 2025 | ShortListing Model: A Streamlined Simplex Diffusion for Discrete Variable GenerationabstractGenerative modeling of discrete variables is challenging yet crucial for applications in natural language processing and biological sequence design. We introduce the Shortlisting Model (SLM), a novel simplex-based diffusion model inspired by progressive candidate pruning. SLM operates on simplex centroids, reducing generation complexity and enhancing scalability. Additionally, SLM incorporates a flexible implementation of classifier-free guidance, enhancing unconditional generation performance. Extensive experiments on DNA promoter and enhancer design, protein design, character-level and large-vocabulary language modeling demonstrate the competitive performance and strong potential of SLM. Our code can be found at https://github.com/GenSI-THUAIR/SLM. Yuxuan Song 0002, Jingjing Gong, Qiying Yu, Zheng Zhang 0001, Mingxuan Wang, Hao Zhou 0012, Wei-Ying Ma |
NeurIPS | 8 |
| 2025 | Retrieval is Not Enough: Enhancing RAG through Test-Time Critique and OptimizationabstractRetrieval-augmented generation (RAG) has become a widely adopted paradigm for enabling knowledge-grounded large language models (LLMs). However, standard RAG pipelines often fail to ensure that model reasoning remains consistent with the evidence retrieved, leading to factual inconsistencies or unsupported conclusions. In this work, we reinterpret RAG as \textit{Retrieval-Augmented Reasoning} and identify a central but underexplored problem: \textit{Reasoning Misalignment}—the divergence between an LLM's internal reasoning trajectory and the evidential constraints provided by retrieval. To address this issue, we propose \textsc{AlignRAG}, a novel iterative framework grounded in \textit{Critique-Driven Alignment (CDA)}. We further introduce \textsc{AlignRAG-auto}, an autonomous variant that dynamically terminates refinement, removing the need to pre-specify the number of critique iterations. At the heart of \textsc{AlignRAG} lies a \textit{contrastive critique synthesis} mechanism that generates retrieval-sensitive critiques while mitigating self-bias. This mechanism trains a dedicated retrieval-augmented \textit{Critic Language Model (CLM)} using labeled critiques that distinguish between evidence-aligned and misaligned reasoning. Empirical evaluations show that our approach significantly improves reasoning fidelity. Our 8B-parameter CLM improves performance over the Self-Refine baseline by \textbf{12.1\%} on out-of-domain tasks and outperforms a standard 72B-parameter CLM by \textbf{2.2\%}. Furthermore, \textsc{AlignRAG-auto} achieves this state-of-the-art performance while dynamically determining the optimal number of refinement steps, enhancing efficiency and usability. \textsc{AlignRAG} remains compatible with existing RAG architectures as a \textit{plug-and-play} module and demonstrates strong robustness under both informative and noisy retrieval scenarios. Overall, \textsc{AlignRAG} offers a principled solution for aligning model reasoning with retrieved evidence, substantially improving the factual reliability and robustness of RAG systems. Our source code is provided at \href{https://github.com/upup-wei/RAG-ReasonAlignment}{link}. Hao Zhou 0012, Xiang Zhang 0028, Di Zhang 0026, Zijie Qiu, Noah Wei, Jinzhe Li, Wanli Ouyang |
NeurIPS | 2 |
| 2025 | Rationalized All-Atom Protein Design with Unified Multi-Modal Bayesian FlowabstractDesigning functional proteins is a critical yet challenging problem due to the intricate interplay between backbone structures, sequences, and side-chains. Current approaches often decompose protein design into separate tasks, which can lead to accumulated errors, while recent efforts increasingly focus on all-atom protein design. However, we observe that existing all-atom generation approaches suffering from an information shortcut issue, where models inadvertently infer sequences from side-chain information, compromising their ability to accurately learn sequence distributions. To address this, we introduce a novel rationalized information flow strategy to eliminate the information shortcut. Furthermore, motivated by the advantages of Bayesian flows over differential equation–based methods, we propose the first Bayesian flow formulation for protein backbone orientations by recasting orientation modeling as an equivalent hyperspherical generation problem with antipodal symmetry. To validate, our method delivers consistently exceptional performance in both peptide and antibody design tasks. Yuxuan Song 0002, Hao Zhou 0012, Wei-Ying Ma |
NeurIPS | 5 |
| 2025 | DAPO: An Open-Source LLM Reinforcement Learning System at ScaleabstractInference scaling empowers LLMs with unprecedented reasoning ability, with reinforcement learning as the core technique to elicit complex reasoning. However, key technical details of state-of-the-art reasoning LLMs are concealed (such as in OpenAI o1 blog and DeepSeek R1 technical report), thus the community still struggles to reproduce their RL training results. We propose the **D**ecoupled Clip and **D**ynamic s**A**mpling **P**olicy **O**ptimization (**DAPO**) algorithm, and fully open-source a state-of-the-art large-scale RL system that achieves 50 points on AIME 2024 using Qwen2.5-32B base model. Unlike previous works that withhold training details, we introduce four key techniques of our algorithm that make large-scale LLM RL a success. In addition, we open-source our training code, which is built on the verl framework, along with a carefully curated and processed dataset. These components of our open-source system enhance reproducibility and support future research in large-scale LLM RL. Qiying Yu, Zheng Zhang 0001, Ruofei Zhu, Yufeng Yuan, Xiaochen Zuo, Yu Yue, Weinan Dai, Tiantian Fan, Gaohong Liu, Juncai Liu, Lingjun Liu, Xin Liu 0039, Haibin Lin, Bole Ma, Guangming Sheng, Yuxuan Tong, Chi Zhang 0022, Mofan Zhang, Ru Zhang 0006, Wang Zhang 0017, Jiaze Chen, Jiangjie Chen, Hongli Yu, Yuxuan Song 0002, Xiangpeng Wei, Hao Zhou 0012, Wei-Ying Ma, Ya-Qin Zhang, Mingxuan Wang |
NeurIPS | 30 |
| 2025 | Accelerating 3D Molecule Generative Models with Trajectory DiagnosisabstractGeometric molecule generative models have found expanding applications across various scientific domains, but their generation inefficiency has become a critical bottleneck. Through a systematic investigation of the generative trajectory, we discover a unique challenge for molecule geometric graph generation: generative models require determining the permutation order of atoms in the molecule before refining its atomic feature values. Based on this insight, we decompose the generation process into permutation phase and adjustment phase, and propose a geometric-informed prior and consistency parameter objective to accelerate each phase. Extensive experiments demonstrate that our approach achieves competitive performance with approximately 10 sampling steps, 7.5 × faster than previous state-of-the-art models and approximately 100 × faster than diffusion-based models, offering a significant step towards scalable molecular generation. Yuxuan Song 0002, Jingjing Gong, Dongzhan Zhou, Hao Zhou 0012, Wei-Ying Ma |
NeurIPS | 7 |
| 2025 | Code Doppler Channel Simulation Methods and Performance Evaluation for Satellite Communication ScenariosabstractIn satellite communication systems, the Doppler effect caused by the high-speed motion of satellites significantly impacts the stability and performance of communication links. This is especially evident in pseudo-random code modulation systems, where the Doppler effect induces a frequency offset in the pseudo-code rate, thereby degrading signal demodulation performance. Focusing on the code Doppler effect in satellite communication scenarios, this paper investigates a channel simulation method adapted to the code Doppler effect and verifies its effectiveness. Hao Zhou 0012, Yiyan Ma, Dan Fei, Bowen Yin, Zishen Zhao, Bo Ai 0001 |
VTC2025-Fall | 1 |
| 2024 | Unified Generative Modeling of 3D Molecules with Bayesian Flow NetworksabstractAdvanced generative model (\textit{e.g.}, diffusion model) derived from simplified continuity assumptions of data distribution, though showing promising progress, has been difficult to apply directly to geometry generation applications due to the \textit{multi-modality} and \textit{noise-sensitive} nature of molecule geometry.
This work introduces Geometric Bayesian Flow Networks (GeoBFN), which naturally fits molecule geometry by modeling diverse modalities in the differentiable parameter space of distributions. GeoBFN maintains the SE-(3) invariant density modeling property by incorporating equivariant inter-dependency modeling on parameters of distributions and unifying the probabilistic modeling of different modalities.
Through optimized training and sampling techniques, we demonstrate that GeoBFN achieves state-of-the-art performance on multiple 3D molecule generation benchmarks in terms of generation quality (90.87\% molecule stability in QM9 and 85.6\% atom stability in GEOM-DRUG\footnote{The scores are reported at 1k sampling steps for fair comparison, and our scores could be further improved if sampling sufficiently longer steps.}). GeoBFN can also conduct sampling with any number of steps to reach an optimal trade-off between efficiency and quality (\textit{e.g.}, 20$\times$ speedup without sacrificing performance). Yuxuan Song 0002, Jingjing Gong, Hao Zhou 0012, Mingyue Zheng, Wei-Ying Ma |
ICLR | 3 |
| 2024 | Towards Codable Watermarking for Injecting Multi-Bits Information to LLMsabstractAs large language models (LLMs) generate texts with increasing fluency and realism, there is a growing need to identify the source of texts to prevent the abuse of LLMs. Text watermarking techniques have proven reliable in distinguishing whether a text is generated by LLMs by injecting hidden patterns. However, we argue that existing LLM watermarking methods are encoding-inefficient and cannot flexibly meet the diverse information encoding needs (such as encoding model version, generation time, user id, etc.). In this work, we conduct the first systematic study on the topic of **Codable Text Watermarking for LLMs** (CTWL) that allows text watermarks to carry multi-bit customizable information. First of all, we study the taxonomy of LLM watermarking technologies and give a mathematical formulation for CTWL. Additionally, we provide a comprehensive evaluation system for CTWL: (1) watermarking success rate, (2) robustness against various corruptions, (3) coding rate of payload information, (4) encoding and decoding efficiency, (5) impacts on the quality of the generated text. To meet the requirements of these non-Pareto-improving metrics, we follow the most prominent vocabulary partition-based watermarking direction, and devise an advanced CTWL method named **Balance-Marking**. The core idea of our method is to use a proxy language model to split the vocabulary into probability-balanced parts, thereby effectively maintaining the quality of the watermarked text. Our code is available at https://github.com/lancopku/codable-watermarking-for-llm. Lean Wang, Wenkai Yang, Deli Chen, Hao Zhou 0012, Yankai Lin 0001, Fandong Meng, Jie Zhou 0016, Xu Sun 0001 |
ICLR | 4 |
| 2024 | Large Language Models Are Not Robust Multiple Choice SelectorsabstractMultiple choice questions (MCQs) serve as a common yet important task format in the evaluation of large language models (LLMs). This work shows that modern LLMs are vulnerable to option position changes in MCQs due to their inherent “selection bias”, namely, they prefer to select specific option IDs as answers (like “Option A”). Through extensive empirical analyses with 20 LLMs on three benchmarks, we pinpoint that this behavioral bias primarily stems from LLMs’ token bias, where the model a priori assigns more probabilistic mass to specific option ID tokens (e.g., A/B/C/D) when predicting answers from the option IDs. To mitigate selection bias, we propose a label-free, inference-time debiasing method, called PriDe, which separates the model’s prior bias for option IDs from the overall prediction distribution. PriDe first estimates the prior by permutating option contents on a small number of test samples, and then applies the estimated prior to debias the remaining samples. We demonstrate that it achieves interpretable and transferable debiasing with high computational efficiency. We hope this work can draw broader research attention to the bias and robustness of modern LLMs. Chujie Zheng, Hao Zhou 0012, Fandong Meng, Jie Zhou 0016, Minlie Huang |
ICLR | 2 |
| 2024 | MolCRAFT: Structure-Based Drug Design in Continuous Parameter SpaceabstractGenerative models for structure-based drug design (SBDD) have shown promising results in recent years. Existing works mainly focus on how to generate molecules with higher binding affinity, ignoring the feasibility prerequisites for generated 3D poses and resulting in false positives. We conduct thorough studies on key factors of ill-conformational problems when applying autoregressive methods and diffusion to SBDD, including mode collapse and hybrid continuous-discrete space. In this paper, we introduce MolCRAFT, the first SBDD model that operates in the continuous parameter space, together with a novel noise reduced sampling strategy. Empirical results show that our model consistently achieves superior performance in binding affinity with more stable 3D structure, demonstrating our ability to accurately model interatomic interactions. To our best knowledge, MolCRAFT is the first to achieve reference-level Vina Scores (-6.59 kcal/mol) with comparable molecular size, outperforming other strong baselines by a wide margin (-0.84 kcal/mol). Code is available at https://github.com/AlgoMole/MolCRAFT. Yanru Qu, Keyue Qiu, Yuxuan Song 0002, Jingjing Gong, Jiawei Han 0001, Mingyue Zheng, Hao Zhou 0012, Wei-Ying Ma |
ICML | 7 |
| 2024 | Mol-AE: Auto-Encoder Based Molecular Representation Learning With 3D Cloze Test Objectiveabstract3D molecular representation learning has gained tremendous interest and achieved promising performance in various downstream tasks. A series of recent approaches follow a prevalent framework: an encoder-only model coupled with a coordinate denoising objective. However, through a series of analytical experiments, we prove that the encoder-only model with coordinate denoising objective exhibits inconsistency between pre-training and downstream objectives, as well as issues with disrupted atomic identifiers. To address these two issues, we propose Mol-AE for molecular representation learning, an auto-encoder model using positional encoding as atomic identifiers. We also propose a new training objective named 3D Cloze Test to make the model learn better atom spatial relationships from real molecular substructures. Empirical results demonstrate that Mol-AE achieves a large margin performance gain compared to the current state-of-the-art 3D molecular modeling approach. Kangjie Zheng, Siyu Long, Zaiqing Nie, Ming Zhang 0004, Xinyu Dai, Wei-Ying Ma, Hao Zhou 0012 |
ICML | 8 |
| 2024 | ESM All-Atom: Multi-Scale Protein Language Model for Unified Molecular ModelingabstractProtein language models have demonstrated significant potential in the field of protein engineering. However, current protein language models primarily operate at the residue scale, which limits their ability to provide information at the atom level. This limitation prevents us from fully exploiting the capabilities of protein language models for applications involving both proteins and small molecules. In this paper, we propose ESM-AA (ESM All-Atom), a novel approach that enables atom-scale and residue-scale unified molecular modeling. ESM-AA achieves this by pre-training on multi-scale code-switch protein sequences and utilizing a multi-scale position encoding to capture relationships among residues and atoms. Experimental results indicate that ESM-AA surpasses previous methods in protein-molecule tasks, demonstrating the full utilization of protein language models. Further investigations reveal that through unified molecular modeling, ESM-AA not only gains molecular knowledge but also retains its understanding of proteins. Kangjie Zheng, Siyu Long, Tianyu Lu, Xinyu Dai, Ming Zhang 0004, Zaiqing Nie, Wei-Ying Ma, Hao Zhou 0012 |
ICML | 9 |
| 2024 | On Prompt-Driven Safeguarding for Large Language ModelsabstractPrepending model inputs with safety prompts is a common practice for safeguarding large language models (LLMs) against queries with harmful intents. However, the underlying working mechanisms of safety prompts have not been unraveled yet, restricting the possibility of automatically optimizing them to improve LLM safety. In this work, we investigate how LLMs’ behavior (i.e., complying with or refusing user queries) is affected by safety prompts from the perspective of model representation. We find that in the representation space, the input queries are typically moved by safety prompts in a "higher-refusal" direction, in which models become more prone to refusing to provide assistance, even when the queries are harmless. On the other hand, LLMs are naturally capable of distinguishing harmful and harmless queries without safety prompts. Inspired by these findings, we propose a method for safety prompt optimization, namely DRO (Directed Representation Optimization). Treating a safety prompt as continuous, trainable embeddings, DRO learns to move the queries’ representations along or opposite the refusal direction, depending on their harmfulness. Experiments with eight LLMs on out-of-domain and jailbreak benchmarks demonstrate that DRO remarkably improves the safeguarding performance of human-crafted safety prompts, without compromising the models’ general performance. Chujie Zheng, Fan Yin, Hao Zhou 0012, Fandong Meng, Jie Zhou 0016, Kai-Wei Chang 0001, Minlie Huang, Nanyun Peng 0001 |
ICML | 3 |
| 2024 | How Effectively Do Code Language Models Understand Poor-Readability Code?abstractCode language models such as CodeT5 and CodeLlama have demonstrated substantial achievement in code comprehension. While the majority of research efforts have focused on improving model architectures and training processes, we find that the current benchmarks used for evaluating code comprehension models are confined to high-readability code, regardless of the popularity of low-readability code in reality. As such, they are inadequate to demonstrate the full spectrum of the model's ability, particularly the robustness to varying readability degrees. In this paper, we analyze the robustness of code summarization models to code with varying readability, including seven obfuscated datasets derived from existing benchmarks. Our findings indicate that current code summarization models are vulnerable to code with poor readability. In particular, their performance predominantly depends on semantic cues within the code, often neglecting the syntactic aspects. Existing benchmarks are biased toward evaluating semantic features, thereby overlooking the models' ability to understand nonsensitive syntactic features. Based on the findings, we present Poor-CodeSumEval, a new evaluation benchmark on code summarization tasks. PoorCodeSumEval innovatively introduces readability into the testing process, considering semantic, syntactic, and their cross-obfuscation, thereby providing a more comprehensive and rigorous evaluation of code summarization models. Our studies also provide more insightful suggestions for future research, such as constructing multi-readability benchmarks to evaluate the robustness of models on poor-readability code, proposing readability-awareness metrics, and automatic methods for code data cleaning and normalization. Yitian Chai, Hao Zhou 0012, Fandong Meng, Jie Zhou 0016, Xiaodong Gu 0002 |
ASE | 3 |
| 2024 | Learning Multi-view Molecular Representations with Structured and Unstructured KnowledgeabstractCapturing molecular knowledge with representation learning approaches holds significant potential in vast scientific fields such as chemistry and life science. An effective and generalizable molecular representation is expected to capture the consensus and complementary molecular expertise from diverse views and perspectives. However, existing works fall short in learning multi-view molecular representations, due to challenges in explicitly incorporating view information and handling molecular knowledge from heterogeneous sources. To address these issues, we present MV-Mol, a molecular representation learning model that harvests multi-view molecular expertise from chemical structures, unstructured knowledge from biomedical texts, and structured knowledge from knowledge graphs. We utilize text prompts to model view information and design a fusion architecture to extract view-based molecular representations. We develop a two-stage pre-training procedure, exploiting heterogeneous data of varying quality and quantity. Through extensive experiments, we show that MV-Mol provides improved representations that substantially benefit molecular property prediction. Additionally, MV-Mol exhibits state-of-the-art performance in multi-modal comprehension of molecular structures and texts. Code and data are available at https://github.com/PharMolix/OpenBioMed. Yizhen Luo, Massimo Hong, Xing Yi Liu, Zikun Nie, Hao Zhou 0012, Zaiqing Nie |
KDD | 6 |
| 2024 | MutaPLM: Protein Language Modeling for Mutation Explanation and EngineeringabstractStudying protein mutations within amino acid sequences holds tremendous significance in life sciences. Protein language models (PLMs) have demonstrated strong capabilities in broad biological applications. However, due to architectural design and lack of supervision, PLMs model mutations implicitly with evolutionary plausibility, which is not satisfactory to serve as explainable and engineerable tools in real-world studies. To address these issues, we present MutaPLM, a unified framework for interpreting and navigating protein mutations with protein language models. MutaPLM introduces a protein *delta* network that captures explicit protein mutation representations within a unified feature space, and a transfer learning pipeline with a chain-of-thought (CoT) strategy to harvest protein mutation knowledge from biomedical texts. We also construct MutaDescribe, the first large-scale protein mutation dataset with rich textual annotations, which provides cross-modal supervision signals. Through comprehensive experiments, we demonstrate that MutaPLM excels at providing human-understandable explanations for mutational effects and prioritizing novel mutations with desirable properties. Our code, model, and data are open-sourced at https://github.com/PharMolix/MutaPLM. Yizhen Luo, Zikun Nie, Massimo Hong, Suyuan Zhao, Hao Zhou 0012, Zaiqing Nie |
NeurIPS | 5 |
| 2024 | Eliciting the Translation Ability of Large Language Models via Multilingual Finetuning with Translation InstructionsabstractAbstract Large-scale pretrained language models (LLMs), such as ChatGPT and GPT4, have shown strong abilities in multilingual translation, without being explicitly trained on parallel corpora. It is intriguing how the LLMs obtain their ability to carry out translation instructions for different languages. In this paper, we present a detailed analysis by finetuning a multilingual pretrained language model, XGLM-7.5B, to perform multilingual translation following given instructions. Firstly, we show that multilingual LLMs have stronger translation abilities than previously demonstrated. For a certain language, the translation performance depends on its similarity to English and the amount of data used in the pretraining phase. Secondly, we find that LLMs’ ability to carry out translation instructions relies on the understanding of translation instructions and the alignment among different languages. With multilingual finetuning with translation instructions, LLMs could learn to perform the translation task well even for those language pairs unseen during the instruction tuning phase. Jiahuan Li, Hao Zhou 0012, Shujian Huang, Shanbo Cheng, Jiajun Chen 0001 |
Trans. Assoc. Comput. Linguistics | 2 |
| 2023 | Multi-Source Probing for Open-Domain Conversational UnderstandingabstractDialogue comprehension and generation are vital to the success of open-domain dialogue systems.Although pre-trained generative conversation models have made significant progress in generating fluent responses, people have difficulty judging whether they understand and efficiently model the contextual information of the conversation.In this study, we propose a Multi-Source Probing (MSP) method to probe the dialogue comprehension abilities of opendomain dialogue models.MSP aggregates features from multiple sources to accomplish diverse task goals and conducts downstream tasks in a generative manner that is consistent with dialogue model pre-training to leverage model capabilities.We conduct probing experiments on seven tasks that require various dialogue comprehension skills, based on the internal representations encoded by dialogue models.Experimental results show that open-domain dialogue models can encode semantic information in the intermediate hidden states, which facilitates dialogue comprehension tasks.Models of different scales and structures possess different conversational understanding capabilities.Our findings encourage a comprehensive evaluation and design of open-domain dialogue models. Hao Zhou 0012, Jie Zhou 0016, Minlie Huang |
EMNLP | 2 |
| 2023 | Label Words are Anchors: An Information Flow Perspective for Understanding In-Context LearningabstractIn-context learning (ICL) emerges as a promising capability of large language models (LLMs) by providing them with demonstration examples to perform diverse tasks.However, the underlying mechanism of how LLMs learn from the provided context remains under-explored.In this paper, we investigate the working mechanism of ICL through an information flow lens.Our findings reveal that label words in the demonstration examples function as anchors:(1) semantic information aggregates into label word representations during the shallow computation layers' processing; (2) the consolidated information in label words serves as a reference for LLMs' final predictions.Based on these insights, we introduce an anchor re-weighting method to improve ICL performance, a demonstration compression technique to expedite inference, and an analysis framework for diagnosing ICL errors in GPT2-XL.The promising applications of our findings again validate the uncovered ICL working mechanism and pave the way for future studies. 1 Lean Wang, Lei Li 0039, Damai Dai, Deli Chen, Hao Zhou 0012, Fandong Meng, Jie Zhou 0016, Xu Sun 0001 |
EMNLP | 5 |
| 2023 | Re³Dial: Retrieve, Reorganize and Rescale Conversations for Long-Turn Open-Domain Dialogue Pre-trainingabstractPre-training on large-scale open-domain dialogue data can substantially improve the performance of dialogue models.However, the pre-trained dialogue model's ability to utilize long-range context is limited due to the scarcity of long-turn dialogue sessions.Most dialogues in existing pre-training corpora contain fewer than three turns of dialogue.To alleviate this issue, we propose the Retrieve, Reorganize and Rescale framework (Re 3 Dial), which can automatically construct billion-scale long-turn dialogues by reorganizing existing short-turn ones.Given a short-turn session, Re 3 Dial first employs a session retriever to retrieve coherent consecutive sessions.To this end, we train the retriever to capture semantic and discourse relations within multi-turn dialogues through contrastive training.Next, Re 3 Dial samples a session from retrieved results following a diversity sampling strategy, which is designed to penalize repetitive or generic sessions.A longer session is then derived by concatenating the original session and the sampled session.By repeating the above process, Re 3 Dial can yield a coherent long-turn dialogue.Extensive experiments on multiple multi-turn dialogue benchmarks demonstrate that Re 3 Dial significantly improves the dialogue model's ability to utilize long-range context and thus generate more sensible and informative responses.Finally, we build a toolkit for efficiently rescaling conversations with Re 3 Dial, which enables us to construct a corpus containing 1B Chinese dialogue sessions with 11.3 turns on average (5× longer than the original corpus).Our retriever model, code, and data is publicly available at https://github.com/thu-coai/Re3Dial. Jiaxin Wen, Hao Zhou 0012, Jian Guan 0002, Jie Zhou 0016, Minlie Huang |
EMNLP | 2 |
| 2023 | On Pre-training Language Model for Antibody
Danqing Wang, Hao Zhou 0012 |
ICLR | 3 |
| 2023 | Coarse-to-Fine: a Hierarchical Diffusion Model for Molecule Generation in 3DabstractGenerating desirable molecular structures in 3D is a fundamental problem for drug discovery. Despite the considerable progress we have achieved, existing methods usually generate molecules in atom resolution and ignore intrinsic local structures such as rings, which leads to poor quality in generated structures, especially when generating large molecules. Fragment-based molecule generation is a promising strategy, however, it is nontrivial to be adapted for 3D non-autoregressive generations because of the combinational optimization problems. In this paper, we utilize a coarse-to-fine strategy to tackle this problem, in which a Hierarchical Diffusion-based model (i.e. HierDiff) is proposed to preserve the validity of local segments without relying on autoregressive modeling. Specifically, HierDiff first generates coarse-grained molecule geometries via an equivariant diffusion process, where each coarse-grained node reflects a fragment in a molecule. Then the coarse-grained nodes are decoded into fine-grained fragments by a message-passing process and a newly designed iterative refined sampling module. Lastly, the fine-grained fragments are then assembled to derive a complete atomic molecular structure. Extensive experiments demonstrate that HierDiff consistently improves the quality of molecule generation over existing methods. Bo Qiang, Yuxuan Song 0002, Minkai Xu, Jingjing Gong, Hao Zhou 0012, Wei-Ying Ma, Yanyan Lan |
ICML | 6 |
| 2023 | Accelerating Antimicrobial Peptide Discovery with Latent StructureabstractAntimicrobial peptides (AMPs) are promising therapeutic approaches against drug-resistant pathogens. Recently, deep generative models are used to discover new AMPs. However, previous studies mainly focus on peptide sequence attributes and do not consider crucial structure information. In this paper, we propose a latent sequence-structure model for designing AMPs (LSSAMP). LSSAMP exploits multi-scale vector quantization in the latent space to represent secondary structures (e.g. alpha helix and beta sheet). By sampling in the latent space, LSSAMP can simultaneously generate peptides with ideal sequence attributes and secondary structures. Experimental results show that the peptides generated by LSSAMP have a high probability of antimicrobial activity. Our wet laboratory experiments verified that two of the 21 candidates exhibit strong antimicrobial activity. The code is released at https://github.com/dqwang122/LSSAMP. Danqing Wang, Zeyu Wen, Lei Li 0005, Hao Zhou 0012 |
KDD | 5 |
| 2023 | Fed-FA: Theoretically Modeling Client Data Divergence for Federated Language Backdoor DefenseabstractFederated learning algorithms enable neural network models to be trained across multiple decentralized edge devices without sharing private data. However, they are susceptible to backdoor attacks launched by malicious clients. Existing robust federated aggregation algorithms heuristically detect and exclude suspicious clients based on their parameter distances, but they are ineffective on Natural Language Processing (NLP) tasks. The main reason is that, although text backdoor patterns are obvious at the underlying dataset level, they are usually hidden at the parameter level, since injecting backdoors into texts with discrete feature space has less impact on the statistics of the model parameters. To settle this issue, we propose to identify backdoor clients by explicitly modeling the data divergence among clients in federated NLP systems. Through theoretical analysis, we derive the f-divergence indicator to estimate the client data divergence with aggregation updates and Hessians. Furthermore, we devise a dataset synthesization method with a Hessian reassignment mechanism guided by the diffusion theory to address the key challenge of inaccessible datasets in calculating clients' data Hessians.
We then present the novel Federated F-Divergence-Based Aggregation~(\textbf{Fed-FA}) algorithm, which leverages the f-divergence indicator to detect and discard suspicious clients. Extensive empirical results show that Fed-FA outperforms all the parameter distance-based methods in defending against backdoor attacks among various natural language backdoor attack scenarios. Zhiyuan Zhang 0001, Deli Chen, Hao Zhou 0012, Fandong Meng, Jie Zhou 0016, Xu Sun 0001 |
NeurIPS | 3 |
| 2023 | Equivariant Flow Matching with Hybrid Probability Transport for 3D Molecule GenerationabstractThe generation of 3D molecules requires simultaneously deciding the categorical features (atom types) and continuous features (atom coordinates). Deep generative models, especially Diffusion Models (DMs), have demonstrated effectiveness in generating feature-rich geometries. However, existing DMs typically suffer from unstable probability dynamics with inefficient sampling speed. In this paper, we introduce geometric flow matching, which enjoys the advantages of both equivariant modeling and stabilized probability dynamics.
More specifically, we propose a hybrid probability path where the coordinates probability path is regularized by an equivariant optimal transport, and the information between different modalities is aligned. Experimentally, the proposed method could consistently achieve better performance on multiple molecule generation benchmarks with 4.75$\times$ speed up of sampling on average. Yuxuan Song 0002, Jingjing Gong, Minkai Xu, Ziyao Cao, Yanyan Lan, Stefano Ermon, Hao Zhou 0012, Wei-Ying Ma |
NeurIPS | 7 |
| 2022 | Non-autoregressive Translation with Layer-Wise Prediction and Deep SupervisionabstractHow do we perform efficient inference while retaining high translation quality? Existing neural machine translation models, such as Transformer, achieve high performance, but they decode words one by one, which is inefficient. Recent non-autoregressive translation models speed up the inference, but their quality is still inferior. In this work, we propose DSLP, a highly efficient and high-performance model for machine translation. The key insight is to train a non-autoregressive Transformer with Deep Supervision and feed additional Layer-wise Predictions. We conducted extensive experiments on four translation tasks (both directions of WMT'14 EN-DE and WMT'16 EN-RO). Results show that our approach consistently improves the BLEU scores compared with respective base models. Specifically, our best variant outperforms the autoregressive model on three translation tasks, while being 14.8 times more efficient in inference. Chenyang Huang 0001, Hao Zhou 0012, Osmar R. Zaïane, Lili Mou, Lei Li 0005 |
AAAI | 2 |
| 2022 | LOREN: Logic-Regularized Reasoning for Interpretable Fact VerificationabstractGiven a natural language statement, how to verify its veracity against a large-scale textual knowledge source like Wikipedia? Most existing neural models make predictions without giving clues about which part of a false claim goes wrong. In this paper, we propose LOREN, an approach for interpretable fact verification. We decompose the verification of the whole claim at phrase-level, where the veracity of the phrases serves as explanations and can be aggregated into the final verdict according to logical rules. The key insight of LOREN is to represent claim phrase veracity as three-valued latent variables, which are regularized by aggregation logical rules. The final claim verification is based on all latent variables. Thus, LOREN enjoys the additional benefit of interpretability --- it is easy to explain how it reaches certain results with claim phrase veracity. Experiments on a public fact verification benchmark show that LOREN is competitive against previous approaches while enjoying the merit of faithful and accurate interpretability. The resources of LOREN are available at: https://github.com/jiangjiechen/LOREN. Jiangjie Chen, Qiaoben Bao, Changzhi Sun, Xinbo Zhang, Jiaze Chen, Hao Zhou 0012, Yanghua Xiao, Lei Li 0005 |
AAAI | 6 |
| 2022 | Unsupervised Editing for Counterfactual StoriesabstractCreating what-if stories requires reasoning about prior statements and possible outcomes of the changed conditions. One can easily generate coherent endings under new conditions, but it would be challenging for current systems to do it with minimal changes to the original story. Therefore, one major challenge is the trade-off between generating a logical story and rewriting with minimal-edits. In this paper, we propose EDUCAT, an editing-based unsupervised approach for counterfactual story rewriting. EDUCAT includes a target position detection strategy based on estimating causal effects of the what-if conditions, which keeps the causal invariant parts of the story. EDUCAT then generates the stories under fluency, coherence and minimal-edits constraints. We also propose a new metric to alleviate the shortcomings of current automatic metrics and better evaluate the trade-off. We evaluate EDUCAT on a public counterfactual story rewriting benchmark. Experiments show that EDUCAT achieves the best trade-off over unsupervised SOTA methods according to both automatic and human evaluation. The resources of EDUCAT are available at: https://github.com/jiangjiechen/EDUCAT. Jiangjie Chen, Chun Gan, Sijie Cheng, Hao Zhou 0012, Yanghua Xiao, Lei Li 0005 |
AAAI | 4 |
| 2022 | latent-GLAT: Glancing at Latent Variables for Parallel Text GenerationabstractYu Bao, Hao Zhou, Shujian Huang, Dongqi Wang, Lihua Qian, Xinyu Dai, Jiajun Chen, Lei Li. Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2022. Hao Zhou 0012, Shujian Huang, Dongqi Wang 0005, Lihua Qian, Xinyu Dai, Jiajun Chen 0001, Lei Li 0005 |
ACL (1) | 2 |
| 2022 | Contextual Representation Learning beyond Masked Language ModelingabstractHow do masked language models (MLMs) such as BERT learn contextual representations?In this work, we analyze the learning dynamics of MLMs.We find that MLMs adopt sampled embeddings as anchors to estimate and inject contextual semantics to representations, which limits the efficiency and effectiveness of MLMs.To address these issues, we propose TACO, a simple yet effective representation learning approach to directly model global semantics.TACO extracts and aligns contextual semantics hidden in contextualized representations to encourage models to attend global semantics when generating contextualized representations.Experiments on the GLUE benchmark show that TACO achieves up to 5x speedup and up to 1.2 points average improvement over existing MLMs.The code is available at https:// github.com/FUZHIYI/TACO. Zhiyi Fu, Wangchunshu Zhou, Jingjing Xu 0001, Hao Zhou 0012, Lei Li 0005 |
ACL (1) | 4 |
| 2022 | CTRLEval: An Unsupervised Reference-Free Metric for Evaluating Controlled Text GenerationabstractExisting reference-free metrics have obvious limitations for evaluating controlled text generation models.Unsupervised metrics can only provide a task-agnostic evaluation result which correlates weakly with human judgments, whereas supervised ones may overfit task-specific data with poor generalization ability to other datasets.In this paper, we propose an unsupervised reference-free metric called CTRLEval, which evaluates controlled text generation from different aspects by formulating each aspect into multiple text infilling tasks.On top of these tasks, the metric assembles the generation probabilities from a pre-trained language model without any model training.Experimental results show that our metric has higher correlations with human judgments than other baselines, while obtaining better generalization of evaluating generated texts from different models and with different qualities 1 . Pei Ke, Hao Zhou 0012, Yankai Lin 0001, Peng Li 0030, Jie Zhou 0016, Xiaoyan Zhu 0001, Minlie Huang |
ACL (1) | 2 |
| 2022 | ROSE: Robust Selective Fine-tuning for Pre-trained Language ModelsabstractEven though the large-scale language models have achieved excellent performances, they suffer from various adversarial attacks.A large body of defense methods has been proposed.However, they are still limited due to redundant attack search spaces and the inability to defend against various types of attacks.In this work, we present a novel fine-tuning approach called RObust SEletive fine-tuning (ROSE) to address this issue.ROSE conducts selective updates when adapting pre-trained models to downstream tasks, filtering out invaluable and unrobust updates of parameters.Specifically, we propose two strategies: the first-order and second-order ROSE for selecting target robust parameters.The experimental results show that ROSE achieves significant improvements in adversarial robustness on various downstream NLP tasks, and the ensemble method even surpasses both variants above.Furthermore, ROSE can be easily incorporated into existing fine-tuning methods to improve their adversarial robustness further.The empirical analysis confirms that ROSE eliminates unrobust spurious updates during fine-tuning, leading to solutions corresponding to flatter and wider optima than the conventional method.Code is available at https: //github.com/jiangllan/ROSE. Hao Zhou 0012, Yankai Lin 0001, Peng Li 0030, Jie Zhou 0016 |
EMNLP | 2 |
| 2022 | switch-GLAT: Multilingual Parallel Machine Translation Via Code-Switch Decoder
Zhenqiao Song, Hao Zhou 0012, Lihua Qian, Jingjing Xu 0001, Shanbo Cheng, Mingxuan Wang, Lei Li 0005 |
ICLR | 2 |
| 2022 | Enhancing Cross-lingual Transfer by Manifold Mixup
Huiyun Yang, Huadong Chen, Hao Zhou 0012, Lei Li 0005 |
ICLR | 3 |
| 2022 | On the Learning of Non-Autoregressive TransformersabstractNon-autoregressive Transformer (NAT) is a family of text generation models, which aims to reduce the decoding latency by predicting the whole sentences in parallel. However, such latency reduction sacrifices the ability to capture left-to-right dependencies, thereby making NAT learning very challenging. In this paper, we present theoretical and empirical analyses to reveal the challenges of NAT learning and propose a unified perspective to understand existing successes. First, we show that simply training NAT by maximizing the likelihood can lead to an approximation of marginal distributions but drops all dependencies between tokens, where the dropped information can be measured by the dataset’s conditional total correlation. Second, we formalize many previous objectives in a unified framework and show that their success can be concluded as maximizing the likelihood on a proxy distribution, leading to a reduced information loss. Empirical studies show that our perspective can explain the phenomena in NAT learning and guide the design of new training methods. Fei Huang 0005, Tianhua Tao, Hao Zhou 0012, Lei Li 0005, Minlie Huang |
ICML | 3 |
| 2022 | Directed Acyclic Transformer for Non-Autoregressive Machine TranslationabstractNon-autoregressive Transformers (NATs) significantly reduce the decoding latency by generating all tokens in parallel. However, such independent predictions prevent NATs from capturing the dependencies between the tokens for generating multiple possible translations. In this paper, we propose Directed Acyclic Transfomer (DA-Transformer), which represents the hidden states in a Directed Acyclic Graph (DAG), where each path of the DAG corresponds to a specific translation. The whole DAG simultaneously captures multiple translations and facilitates fast predictions in a non-autoregressive fashion. Experiments on the raw training data of WMT benchmark show that DA-Transformer substantially outperforms previous NATs by about 3 BLEU on average, which is the first NAT model that achieves competitive results with autoregressive Transformers without relying on knowledge distillation. Fei Huang 0005, Hao Zhou 0012, Yang Liu 0005, Hang Li 0001, Minlie Huang |
ICML | 2 |
| 2022 | Zero-Shot 3D Drug Design by Sketching and GeneratingabstractDrug design is a crucial step in the drug discovery cycle. Recently, various deep learning-based methods design drugs by generating novel molecules from scratch, avoiding traversing large-scale drug libraries. However, they depend on scarce experimental data or time-consuming docking simulation, leading to overfitting issues with limited training data and slow generation speed. In this study, we propose the zero-shot drug design method DESERT (Drug dEsign by SkEtching and geneRaTing). Specifically, DESERT splits the design process into two stages: sketching and generating, and bridges them with the molecular shape. The two-stage fashion enables our method to utilize the large-scale molecular database to reduce the need for experimental data and docking simulation. Experiments show that DESERT achieves a new state-of-the-art at a fast speed. Siyu Long, Yi Zhou 0018, Xinyu Dai, Hao Zhou 0012 |
NeurIPS | 4 |
| 2022 | Regularized Molecular Conformation FieldsabstractPredicting energetically favorable 3-dimensional conformations of organic molecules frommolecular graph plays a fundamental role in computer-aided drug discovery research.However, effectively exploring the high-dimensional conformation space to identify (meta) stable conformers is anything but trivial.In this work, we introduce RMCF, a novel framework to generate a diverse set of low-energy molecular conformations through samplingfrom a regularized molecular conformation field.We develop a data-driven molecular segmentation algorithm to automatically partition each molecule into several structural building blocks to reduce the modeling degrees of freedom.Then, we employ a Markov Random Field to learn the joint probability distribution of fragment configurations and inter-fragment dihedral angles, which enables us to sample from different low-energy regions of a conformation space.Our model constantly outperforms state-of-the-art models for the conformation generation task on the GEOM-Drugs dataset.We attribute the success of RMCF to modeling in a regularized feature space and learning a global fragment configuration distribution for effective sampling.The proposed method could be generalized to deal with larger biomolecular systems. Yi Zhou 0018, Xiaoqing Zheng, Xuanjing Huang 0001, Hao Zhou 0012 |
NeurIPS | 6 |
| 2021 | Consecutive Decoding for Speech-to-text TranslationabstractSpeech-to-text translation (ST), which directly translates the source language speech to the target language text, has attracted intensive attention recently. However, the combination of speech recognition and machine translation in a single model poses a heavy burden on the direct cross-modal cross-lingual mapping. To reduce the learning difficulty, we propose COnSecutive Transcription and Translation (COSTT), an integral approach for speech-to-text translation. The key idea is to generate source transcript and target translation text with a single decoder. It benefits the model training so that additional large parallel text corpus can be fully exploited to enhance the speech translation training. Our method is verified on three mainstream datasets, including Augmented LibriSpeech English-French dataset, TED English-German dataset, and TED English-Chinese dataset. Experiments show that our proposed COSTT outperforms the previous state-of-the-art methods. The code is available at https://github.com/dqqcasia/st. Qianqian Dong, Mingxuan Wang, Hao Zhou 0012, Bo Xu 0002, Lei Li 0005 |
AAAI | 3 |
| 2021 | Listen, Understand and Translate: Triple Supervision Decouples End-to-end Speech-to-text TranslationabstractAn end-to-end speech-to-text translation (ST) takes audio in a source language and outputs the text in a target language. Existing methods are limited by the amount of parallel corpus. Can we build a system to fully utilize signals in a parallel ST corpus? We are inspired by human understanding system which is composed of auditory perception and cognitive processing. In this paper, we propose Listen-Understand-Translate, (LUT), a unified framework with triple supervision signals to decouple the end-to-end speech-to-text translation task. LUT is able to guide the acoustic encoder to extract as much information from the auditory input. In addition, LUT utilizes a pre-trained BERT model to enforce the upper encoder to produce as much semantic information as possible, without extra data. We perform experiments on a diverse set of speech translation benchmarks, including Librispeech English-French, IWSLT English-German and TED English-Chinese. Our results demonstrate LUT achieves the state-of-the-art performance, outperforming previous methods. The code is available at https://github.com/dqqcasia/st. Qianqian Dong, Rong Ye, Mingxuan Wang, Hao Zhou 0012, Bo Xu 0002, Lei Li 0005 |
AAAI | 4 |
| 2021 | ACMo: Angle-Calibrated Moment Methods for Stochastic Optimization
Xunpeng Huang, Runxin Xu, Hao Zhou 0012, Zhengyang Liu 0002, Lei Li 0005 |
AAAI | 3 |
| 2021 | Glancing Transformer for Non-Autoregressive Neural Machine TranslationabstractLihua Qian, Hao Zhou, Yu Bao, Mingxuan Wang, Lin Qiu, Weinan Zhang, Yong Yu, Lei Li. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Lihua Qian, Hao Zhou 0012, Mingxuan Wang, Weinan Zhang 0001, Yong Yu 0001, Lei Li 0005 |
ACL/IJCNLP (1) | 2 |
| 2021 | UniRE: A Unified Label Space for Entity Relation ExtractionabstractYijun Wang, Changzhi Sun, Yuanbin Wu, Hao Zhou, Lei Li, Junchi Yan. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Changzhi Sun, Yuanbin Wu, Hao Zhou 0012, Lei Li 0005, Junchi Yan |
ACL/IJCNLP (1) | 4 |
| 2021 | Vocabulary Learning via Optimal Transport for Neural Machine TranslationabstractJingjing Xu, Hao Zhou, Chun Gan, Zaixiang Zheng, Lei Li. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021. Jingjing Xu 0001, Hao Zhou 0012, Chun Gan, Zaixiang Zheng, Lei Li 0005 |
ACL/IJCNLP (1) | 2 |
| 2021 | ENPAR: Enhancing Entity and Entity Pair Representations for Joint Entity Relation ExtractionabstractCurrent state-of-the-art systems for joint entity relation extraction (Luan et al., 2019;Wadden et al., 2019) usually adopt the multi-task learning framework.However, annotations for these additional tasks such as coreference resolution and event extraction are always equally hard (or even harder) to obtain.In this work, we propose a pre-training method ENPAR to improve the joint extraction performance.EN-PAR requires only the additional entity annotations that are much easier to collect.Unlike most existing works that only consider incorporating entity information into the sentence encoder, we further utilize the entity pair information.Specifically, we devise four novel objectives, i.e., masked entity typing, masked entity prediction, adversarial context discrimination, and permutation prediction, to pretrain an entity encoder and an entity pair encoder.Comprehensive experiments show that the proposed pre-training method achieves significant improvement over BERT on ACE05, SciERC, and NYT, and outperforms current state-of-the-art on ACE05. Changzhi Sun, Yuanbin Wu, Hao Zhou 0012, Lei Li 0005, Junchi Yan |
EACL | 4 |
| 2021 | Learning Logic Rules for Document-Level Relation ExtractionabstractDocument-level relation extraction aims to identify relations between entities in a whole document.Prior efforts to capture long-range dependencies have relied heavily on implicitly powerful representations learned through (graph) neural networks, which makes the model less transparent.To tackle this challenge, in this paper, we propose LogiRE, a novel probabilistic model for document-level relation extraction by learning logic rules.Lo-giRE treats logic rules as latent variables and consists of two modules: a rule generator and a relation extractor.The rule generator is to generate logic rules potentially contributing to final predictions, and the relation extractor outputs final predictions based on the generated logic rules.Those two modules can be efficiently optimized with the expectationmaximization (EM) algorithm.By introducing logic rules into neural networks, LogiRE can explicitly capture long-range dependencies as well as enjoy better interpretation.Empirical results show that LogiRE significantly outperforms several strong baselines in terms of relation performance (∼1.8 F1 score) and logical consistency (over 3.3 logic score).Our code is available at https://github.com/rudongyu/LogiRE. Dongyu Ru, Changzhi Sun, Jiangtao Feng, Hao Zhou 0012, Weinan Zhang 0001, Yong Yu 0001, Lei Li 0005 |
EMNLP (1) | 5 |
| 2021 | EARL: Informative Knowledge-Grounded Conversation Generation with Entity-Agnostic Representation LearningabstractGenerating informative and appropriate responses is challenging but important for building human-like dialogue systems.Although various knowledge-grounded conversation models have been proposed, these models have limitations in utilizing knowledge that infrequently occurs in the training data, not to mention integrating unseen knowledge into conversation generation.In this paper, we propose an Entity-Agnostic Representation Learning (EARL) method to introduce knowledge graphs to informative conversation generation.Unlike traditional approaches that parameterize the specific representation for each entity, EARL utilizes the context of conversations and the relational structure of knowledge graphs to learn the category representation for entities, which is generalized to incorporating unseen entities in knowledge graphs into conversation generation.Automatic and manual evaluations demonstrate that our model can generate more informative, coherent, and natural responses than baseline models. Hao Zhou 0012, Minlie Huang, Wei Chen 0034, Xiaoyan Zhu 0001 |
EMNLP (1) | 1 |
| 2021 | MARS: Markov Molecular Sampling for Multi-objective Drug Discovery
Yutong Xie 0007, Chence Shi, Hao Zhou 0012, Weinan Zhang 0001, Yong Yu 0001, Lei Li 0005 |
ICLR | 3 |
| 2021 | Duplex Sequence-to-Sequence Learning for Reversible Machine TranslationabstractSequence-to-sequence learning naturally has two directions. How to effectively utilize supervision signals from both directions? Existing approaches either require two separate models, or a multitask-learned model but with inferior performance. In this paper, we propose REDER (Reversible Duplex Transformer), a parameter-efficient model and apply it to machine translation. Either end of REDER can simultaneously input and output a distinct language. Thus REDER enables {\em reversible machine translation} by simply flipping the input and output ends. Experiments verify that REDER achieves the first success of reversible machine translation, which helps outperform its multitask-trained baselines by up to 1.3 BLEU. Zaixiang Zheng, Hao Zhou 0012, Shujian Huang, Jiajun Chen 0001, Jingjing Xu 0001, Lei Li 0005 |
NeurIPS | 2 |
| 2021 | CNewSum: A Large-Scale Summarization Dataset with Human-Annotated Adequacy and Deducibility Level
Danqing Wang, Jiaze Chen, Xianze Wu, Hao Zhou 0012, Lei Li 0005 |
NLPCC (1) | 4 |
| 2021 | Follow Your Path: A Progressive Method for Knowledge Distillation
Wenxian Shi, Yuxuan Song 0002, Hao Zhou 0012, Lei Li 0005 |
ECML/PKDD (3) | 3 |
| 2021 | Triangular Bidword Generation for Sponsored Search AuctionabstractSponsored search auction is a crucial component of modern search engines. It requires a set of candidate bidwords that advertisers can place bids on. Existing methods generate bidwords from search queries or advertisement content. However, they suffer from the data noise in (query, bidword) and (advertisement, bidword) pairs. In this paper, we propose a triangular bidword generation model (TRIDENT), which takes the high-quality data of paired (query, advertisement) as a supervision signal to indirectly guide the bidword generation process. Our proposed model is simple yet effective: by using bidword as the bridge between search query and advertisement, the generation of search query, advertisement and bidword can be jointly learned in the triangular training framework. This alleviates the problem that the training data of bidword may be noisy. Experimental results, including automatic and human evaluations, show that our proposed TRIDENT can generate relevant and diverse bidwords for both search queries and advertisements. Our evaluation on online real data validates the effectiveness of the TRIDENT's generated bidwords for product search. Zhenqiao Song, Jiaze Chen, Hao Zhou 0012, Lei Li 0005 |
WSDM | 3 |
| 2021 | Simulated annealing for optimization of graphs and sequences
Xianggen Liu, Pengyong Li, Fandong Meng, Hao Zhou 0012, Huasong Zhong, Jie Zhou 0016, Lili Mou, Sen Song |
Neurocomputing | 4 |
| 2020 | Infomax Neural Joint Source-Channel Coding via Adversarial Bit FlipabstractAlthough Shannon theory states that it is asymptotically optimal to separate the source and channel coding as two independent processes, in many practical communication scenarios this decomposition is limited by the finite bit-length and computational power for decoding. Recently, neural joint source-channel coding (NECST) (Choi et al. 2018) is proposed to sidestep this problem. While it leverages the advancements of amortized inference and deep learning (Kingma and Welling 2013; Grover and Ermon 2018) to improve the encoding and decoding process, it still cannot always achieve compelling results in terms of compression and error correction performance due to the limited robustness of its learned coding networks. In this paper, motivated by the inherent connections between neural joint source-channel coding and discrete representation learning, we propose a novel regularization method called Infomax Adversarial-Bit-Flip (IABF) to improve the stability and robustness of the neural joint source-channel coding scheme. More specifically, on the encoder side, we propose to explicitly maximize the mutual information between the codeword and data; while on the decoder side, the amortized reconstruction is regularized within an adversarial framework. Extensive experiments conducted on various real-world datasets evidence that our IABF can achieve state-of-the-art performances on both compression and error correction benchmarks and outperform the baselines by a significant margin. Yuxuan Song 0002, Minkai Xu, Lantao Yu, Hao Zhou 0012, Shuo Shao 0001, Yong Yu 0001 |
AAAI | 4 |
| 2020 | Importance-Aware Learning for Neural Headline EditingabstractMany social media news writers are not professionally trained. Therefore, social media platforms have to hire professional editors to adjust amateur headlines to attract more readers. We propose to automate this headline editing process through neural network models to provide more immediate writing support for these social media news writers. To train such a neural headline editing model, we collected a dataset which contains articles with original headlines and professionally edited headlines. However, it is expensive to collect a large number of professionally edited headlines. To solve this low-resource problem, we design an encoder-decoder model which leverages large scale pre-trained language models. We further improve the pre-trained model's quality by introducing a headline generation task as an intermediate task before the headline editing task. Also, we propose Self Importance-Aware (SIA) loss to address the different levels of editing in the dataset by down-weighting the importance of easily classified tokens and sentences. With the help of Pre-training, Adaptation, and SIA, the model learns to generate headlines in the professional editor's style. Experimental results show that our method significantly improves the quality of headline editing comparing against previous methods. Qingyang Wu, Lei Li 0005, Hao Zhou 0012 |
AAAI | 3 |
| 2020 | Towards Making the Most of BERT in Neural Machine TranslationabstractGPT-2 and BERT demonstrate the effectiveness of using pre-trained language models (LMs) on various natural language processing tasks. However, LM fine-tuning often suffers from catastrophic forgetting when applied to resource-rich tasks. In this work, we introduce a concerted training framework (CTnmt) that is the key to integrate the pre-trained LMs to neural machine translation (NMT). Our proposed CTnmt} consists of three techniques: a) asymptotic distillation to ensure that the NMT model can retain the previous pre-trained knowledge; b) a dynamic switching gate to avoid catastrophic forgetting of pre-trained knowledge; and c) a strategy to adjust the learning paces according to a scheduled policy. Our experiments in machine translation show CTnmt gains of up to 3 BLEU score on the WMT14 English-German language pair which even surpasses the previous state-of-the-art pre-training aided NMT by 1.4 BLEU score. While for the large WMT14 English-French task with 40 millions of sentence-pairs, our base model still significantly improves upon the state-of-the-art Transformer big model by more than 1 BLEU score. Mingxuan Wang, Hao Zhou 0012, Chengqi Zhao, Weinan Zhang 0001, Yong Yu 0001, Lei Li 0005 |
AAAI | 3 |
| 2020 | Unsupervised Paraphrasing by Simulated AnnealingabstractWe propose UPSA, a novel approach that accomplishes Unsupervised Paraphrasing by Simulated Annealing.We model paraphrase generation as an optimization problem and propose a sophisticated objective function, involving semantic similarity, expression diversity, and language fluency of paraphrases.UPSA searches the sentence space towards this objective by performing a sequence of local edits.We evaluate our approach on various datasets, namely, Quora, Wikianswers, MSCOCO, and Twitter.Extensive results show that UPSA achieves the state-of-the-art performance compared with previous unsupervised methods in terms of both automatic and human evaluations.Further, our approach outperforms most existing domain-adapted supervised models, showing the generalizability of UPSA. 1 Xianggen Liu, Lili Mou, Fandong Meng, Hao Zhou 0012, Jie Zhou 0016, Sen Song |
ACL | 4 |
| 2020 | Do you have the right scissors? Tailoring Pre-trained Language Models via Monte-Carlo MethodsabstractIt has been a common approach to pre-train a language model on a large corpus and finetune it on task-specific data.In practice, we observe that fine-tuning a pre-trained model on a small dataset may lead to over-and/or under-estimation problem.In this paper, we propose MC-Tailor, a novel method to alleviate the above issue in text generation tasks by truncating and transferring the probability mass from over-estimated regions to underestimated ones.Experiments on a variety of text generation datasets show that MC-Tailor consistently and significantly outperforms the fine-tuning approach.Our code is available at https://github.com/NingMiao/ MC-tailor. Ning Miao, Yuxuan Song 0002, Hao Zhou 0012, Lei Li 0005 |
ACL | 3 |
| 2020 | KdConv: A Chinese Multi-domain Dialogue Dataset Towards Multi-turn Knowledge-driven ConversationabstractThe research of knowledge-driven conversational systems is largely limited due to the lack of dialog data which consists of multi-turn conversations on multiple topics and with knowledge annotations. In this paper, we propose a Chinese multi-domain knowledge-driven conversation dataset, KdConv, which grounds the topics in multi-turn conversations to knowledge graphs. Our corpus contains 4.5K conversations from three domains (film, music, and travel), and 86K utterances with an average turn number of 19.0. These conversations contain in-depth discussions on related topics and natural transition between multiple topics. To facilitate the following research on this corpus, we provide several benchmark models. Comparative results show that the models can be enhanced by introducing background knowledge, yet there is still a large space for leveraging knowledge to model multi-turn conversations for further research. Results also show that there are obvious performance differences between different domains, indicating that it is worth further explore transfer learning and domain adaptation. The corpus and benchmark models are publicly available. Hao Zhou 0012, Chujie Zheng, Kaili Huang, Minlie Huang, Xiaoyan Zhu 0001 |
ACL | 1 |
| 2020 | Improving Maximum Likelihood Training for Text Generation with Density Ratio EstimationabstractAutoregressive neural sequence generative models trained by Maximum Likelihood Estimation suffer the exposure bias problem in practical finite sample scenarios. The crux is that the number of training samples for Maximum Likelihood Estimation is usually limited and the input data distributions are different at training and inference stages. Many methods have been proposed to solve the above problem, which relies on sampling from the non-stationary model distribution and suffers from high variance or biased estimations. In this paper, we propose $\psi$-MLE, a new training scheme for autoregressive sequence generative models, which is effective and stable when operating at large sample space encountered in text generation. We derive our algorithm from a new perspective of self-augmentation and introduce bias correction with density ratio estimation. Extensive experimental results on synthetic data and real-world text generation tasks demonstrate that our method stably outperforms Maximum Likelihood Estimation and other state-of-the-art sequence generative models in terms of both quality and diversity. Yuxuan Song 0002, Ning Miao, Hao Zhou 0012, Lantao Yu, Mingxuan Wang, Lei Li 0005 |
AISTATS | 3 |
| 2020 | On the Sentence Embeddings from Pre-trained Language ModelsabstractPre-trained contextual representations like BERT have achieved great success in natural language processing.However, the sentence embeddings from the pre-trained language models without fine-tuning have been found to poorly capture semantic meaning of sentences.In this paper, we argue that the semantic information in the BERT embeddings is not fully exploited.We first reveal the theoretical connection between the masked language model pre-training objective and the semantic similarity task theoretically, and then analyze the BERT sentence embeddings empirically.We find that BERT always induces a non-smooth anisotropic semantic space of sentences, which harms its performance of semantic similarity.To address this issue, we propose to transform the anisotropic sentence embedding distribution to a smooth and isotropic Gaussian distribution through normalizing flows that are learned with an unsupervised objective.Experimental results show that our proposed BERT-flow method obtains significant performance gains over the state-of-the-art sentence embeddings on a variety of semantic textual similarity tasks.The code is available at https://github.com/ bohanli/BERT-flow. Hao Zhou 0012, Junxian He, Mingxuan Wang, Yiming Yang 0002, Lei Li 0005 |
EMNLP (1) | 2 |
| 2020 | Pre-training Multilingual Neural Machine Translation by Leveraging Alignment InformationabstractWe investigate the following question for machine translation (MT): can we develop a single universal MT model to serve as the common seed and obtain derivative and improved models on arbitrary language pairs?We propose mRASP, an approach to pre-train a universal multilingual neural machine translation model.Our key idea in mRASP is its novel technique of random aligned substitution, which brings words and phrases with similar meanings across multiple languages closer in the representation space.We pre-train a mRASP model on 32 language pairs jointly with only public datasets.The model is then fine-tuned on downstream language pairs to obtain specialized MT models.We carry out extensive experiments on 42 translation directions across a diverse settings, including low, medium, rich resource, and as well as transferring to exotic language pairs.Experimental results demonstrate that mRASP achieves significant performance improvement compared to directly training on those target pairs.It is the first time to verify that multiple lowresource language pairs can be utilized to improve rich resource MT.Surprisingly, mRASP is even able to improve the translation quality on exotic languages that never occur in the pretraining corpus.Code, data, and pre-trained models are available at https://github. com/linzehui/mRASP. Mingxuan Wang, Xipeng Qiu, Jiangtao Feng, Hao Zhou 0012, Lei Li 0005 |
EMNLP (1) | 6 |
| 2020 | Variational Template Machine for Data-to-Text Generation
Rong Ye, Wenxian Shi, Hao Zhou 0012, Zhongyu Wei, Lei Li 0005 |
ICLR | 3 |
| 2020 | Mirror-Generative Neural Machine Translation
Zaixiang Zheng, Hao Zhou 0012, Shujian Huang, Lei Li 0005, Xinyu Dai, Jiajun Chen 0001 |
ICLR | 2 |
| 2020 | Dispersed Exponential Family Mixture VAEs for Interpretable Text GenerationabstractDeep generative models are commonly used for generating images and text. Interpretability of these models is one important pursuit, other than the generation quality. Variational auto-encoder (VAE) with Gaussian distribution as prior has been successfully applied in text generation, but it is hard to interpret the meaning of the latent variable. To enhance the controllability and interpretability, one can replace the Gaussian prior with a mixture of Gaussian distributions (GM-VAE), whose mixture components could be related to hidden semantic aspects of data. In this paper, we generalize the practice and introduce DEM-VAE, a class of models for text generation using VAEs with a mixture distribution of exponential family. Unfortunately, a standard variational training algorithm fails due to the \emph{mode-collapse} problem. We theoretically identify the root cause of the problem and propose an effective algorithm to train DEM-VAE. Our method penalizes the training with an extra \emph{dispersion term} to induce a well-structured latent space. Experimental results show that our approach does obtain a meaningful space, and it outperforms strong baselines in text generation benchmarks. The code is available at \url{https://github.com/wenxianxian/demvae}. Wenxian Shi, Hao Zhou 0012, Ning Miao, Lei Li 0005 |
ICML | 2 |
| 2020 | QuAChIE: Question Answering based Chinese Information Extraction SystemabstractIn this paper, we present the design of QuAChIE, a Question Answering based Chinese Information Extraction system. QuAChIE mainly depends on a well-trained question answering model to extract high-quality triples. The group of head entity and relation are regarded as a question given the input text as the context. For the training and evaluation of each model in the system, we build a large-scale information extraction dataset using Wikidata and Wikipedia pages by distant supervision. The advanced models implemented on top of the pre-trained language model and the enormous distant supervision data enable QuAChIE to extract relation triples from documents with cross-sentence correlations. The experimental results on the test set and the case study based on the interactive demonstration show its satisfactory Information Extraction quality on Chinese document-level texts. Dongyu Ru, Zhenghui Wang, Hao Zhou 0012, Lei Li 0005, Weinan Zhang 0001, Yong Yu 0001 |
SIGIR | 4 |
| 2019 | CGMH: Constrained Sentence Generation by Metropolis-Hastings SamplingabstractIn real-world applications of natural language generation, there are often constraints on the target sentences in addition to fluency and naturalness requirements. Existing language generation techniques are usually based on recurrent neural networks (RNNs). However, it is non-trivial to impose constraints on RNNs while maintaining generation quality, since RNNs generate sentences sequentially (or with beam search) from the first word to the last. In this paper, we propose CGMH, a novel approach using Metropolis-Hastings sampling for constrained sentence generation. CGMH allows complicated constraints such as the occurrence of multiple keywords in the target sentences, which cannot be handled in traditional RNN-based approaches. Moreover, CGMH works in the inference stage, and does not require parallel corpora for training. We evaluate our method on a variety of tasks, including keywords-to-sentence generation, unsupervised sentence paraphrasing, and unsupervised sentence error correction. CGMH achieves high performance compared with previous supervised methods for sentence generation. Our code is released at https://github.com/NingMiao/CGMH Ning Miao, Hao Zhou 0012, Lili Mou, Rui Yan 0001, Lei Li 0005 |
AAAI | 2 |
| 2019 | Generating Sentences from Disentangled Syntactic and Semantic SpacesabstractVariational auto-encoders (VAEs) are widely used in natural language generation due to the regularization of the latent space.However, generating sentences from the continuous latent space does not explicitly model the syntactic information.In this paper, we propose to generate sentences from disentangled syntactic and semantic spaces.Our proposed method explicitly models syntactic information in the VAE's latent space by using the linearized tree sequence, leading to better performance of language generation.Additionally, the advantage of sampling in the disentangled syntactic and semantic latent spaces enables us to perform novel applications, such as the unsupervised paraphrase generation and syntaxtransfer generation.Experimental results show that our proposed model achieves similar or better performance in various tasks, compared with state-of-the-art related work. Hao Zhou 0012, Shujian Huang, Lei Li 0005, Lili Mou, Olga Vechtomova, Xinyu Dai, Jiajun Chen 0001 |
ACL (1) | 2 |
| 2019 | Dynamically Fused Graph Network for Multi-hop ReasoningabstractText-based question answering (TBQA) has been studied extensively in recent years.Most existing approaches focus on finding the answer to a question within a single paragraph.However, many difficult questions require multiple supporting evidence from scattered text across two or more documents.In this paper, we propose the Dynamically Fused Graph Network (DFGN), a novel method to answer those questions requiring multiple scattered evidence and reasoning over them.Inspired by human's step-by-step reasoning behavior, DFGN includes a dynamic fusion layer that starts from the entities mentioned in the given query, explores along the entity graph dynamically built from the text, and gradually finds relevant supporting entities from the given documents.We evaluate DFGN on HotpotQA, a public TBQA dataset requiring multi-hop reasoning.DFGN achieves competitive results on the public board.Furthermore, our analysis shows DFGN could produce interpretable reasoning chains. Yunxuan Xiao, Yanru Qu, Hao Zhou 0012, Lei Li 0005, Weinan Zhang 0001, Yong Yu 0001 |
ACL (1) | 4 |
| 2019 | Imitation Learning for Non-Autoregressive Neural Machine TranslationabstractNon-autoregressive translation models (NAT) have achieved impressive inference speedup.A potential issue of the existing NAT algorithms, however, is that the decoding is conducted in parallel, without directly considering previous context.In this paper, we propose an imitation learning framework for nonautoregressive machine translation, which still enjoys the fast translation speed but gives comparable translation performance compared to its auto-regressive counterpart.We conduct experiments on the IWSLT16, WMT14 and WMT16 datasets.Our proposed model achieves a significant speedup over the autoregressive models, while keeping the translation quality comparable to the autoregressive models.By sampling sentence length in parallel at inference time, we achieve the performance of 31.85BLEU on WMT16 Ro→En and 30.68 BLEU on IWSLT16 En→De. Bingzhen Wei, Mingxuan Wang, Hao Zhou 0012, Junyang Lin, Xu Sun 0001 |
ACL (1) | 3 |
| 2019 | Generating Fluent Adversarial Examples for Natural LanguagesabstractEfficiently building an adversarial attacker for natural language processing (NLP) tasks is a real challenge.Firstly, as the sentence space is discrete, it is difficult to make small perturbations along the direction of gradients.Secondly, the fluency of the generated examples cannot be guaranteed.In this paper, we propose MHA, which addresses both problems by performing Metropolis-Hastings sampling, whose proposal is designed with the guidance of gradients.Experiments on IMDB and SNLI show that our proposed MHA outperforms the baseline model on attacking capability.Adversarial training with MHA also leads to better robustness and performance. Huangzhao Zhang, Hao Zhou 0012, Ning Miao, Lei Li 0005 |
ACL (1) | 2 |
| 2019 | Why Do Neural Dialog Systems Generate Short and Meaningless Replies? a Comparison between Dialog and TranslationabstractThis paper addresses the question: In neural dialog systems, why do sequence-to-sequence (Seq2Seq) neural networks generate short and meaningless replies for open-domain response generation? We conjecture that in a dialog system, due to the randomness of spoken language, there may be multiple equally plausible replies for one utterance, causing the deficiency of a Seq2Seq model. To evaluate our conjecture, we propose a systematic way to mimic the dialog scenario in machine translation systems with both real datasets and toy datasets generated elaborately. Experimental results show that we manage to reproduce the phenomenon of generating short and meaningless sentences in the translation setting. Bolin Wei, Lili Mou, Hao Zhou 0012, Pascal Poupart, Ge Li 0001, Zhi Jin 0001 |
ICASSP | 4 |
| 2019 | GraspSnooker: Automatic Chinese Commentary Generation for Snooker VideosabstractWe demonstrate a web-based software system, GraspSnooker, which is able to automatically generate Chinese text commentaries for snooker game videos. It consists of a video analyzer, a strategy predictor and a commentary generator. As far as we know, it is the first attempt on snooker commentary generation, which might be helpful for snooker learners to understand the game. Zhaoyue Sun, Jiaze Chen, Hao Zhou 0012, Lei Li 0005, Mingmin Jiang |
IJCAI | 3 |
| 2019 | Correct-and-Memorize: Learning to Translate from Interactive RevisionsabstractState-of-the-art machine translation models are still not on a par with human translators. Previous work takes human interactions into the neural machine translation process to obtain improved results in target languages. However, not all model--translation errors are equal -- some are critical while others are minor. In the meanwhile, same translation mistakes occur repeatedly in similar context. To solve both issues, we propose CAMIT, a novel method for translating in an interactive environment. Our proposed method works with critical revision instructions, therefore allows human to correct arbitrary words in model-translated sentences. In addition, CAMIT learns from and softly memorizes revision actions based on the context, alleviating the issue of repeating mistakes. Experiments in both ideal and real interactive translation settings demonstrate that our proposed CAMIT enhances machine translation results significantly while requires fewer revision instructions from human compared to previous methods. Rongxiang Weng, Hao Zhou 0012, Shujian Huang, Lei Li 0005, Jiajun Chen 0001 |
IJCAI | 2 |
| 2019 | Rethinking Text Attribute Transfer: A Lexical AnalysisabstractText attribute transfer is modifying certain linguistic attributes (e.g.sentiment, style, authorship, etc.) of a sentence and transforming them from one type to another.In this paper, we aim to analyze and interpret what is changed during the transfer process.We start from the observation that in many existing models and datasets, certain words within a sentence play important roles in determining the sentence attribute class.These words are referred to as the Pivot Words.Based on these pivot words, we propose a lexical analysis framework, the Pivot Analysis, to quantitatively analyze the effects of these words in text attribute classification and transfer.We apply this framework to existing datasets and models, and show that: (1) the pivot words are strong features for the classification of sentence attributes; (2) to change the attribute of a sentence, many datasets only requires to change certain pivot words; (3) consequently, many transfer models only perform the lexical-level modification, while leaving higher-level sentence structures unchanged.Our work provides an in-depth understanding of linguistic attribute transfer and further identifies the future requirements and challenges of this task 1 . Hao Zhou 0012, Jiaze Chen, Lei Li 0005 |
INLG | 2 |
| 2019 | Kernelized Bayesian Softmax for Text GenerationabstractNeural models for text generation require a softmax layer with proper token embeddings during the decoding phase. Most existing approaches adopt single point embedding for each token. However, a word may have multiple senses according to different context, some of which might be distinct. In this paper, we propose KerBS, a novel approach for learning better embeddings for text generation. KerBS embodies two advantages: (a) it employs a Bayesian composition of embeddings for words with multiple senses; (b) it is adaptive to semantic variances of words and robust to rare sentence context by imposing learned kernels to capture the closeness of words (senses) in the embedding space. Empirical studies show that KerBS significantly boosts the performance of several text generation tasks. Ning Miao, Hao Zhou 0012, Chengqi Zhao, Wenxian Shi, Lei Li 0005 |
NeurIPS | 2 |
| 2019 | Domain-Constrained Advertising Keyword GenerationabstractAdvertising (ad for short) keyword suggestion is important for sponsored search to improve online advertising and increase search revenue. There are two common challenges in this task. First, the keyword bidding problem: hot ad keywords are very expensive for most of the advertisers because more advertisers are bidding on more popular keywords, while unpopular keywords are difficult to discover. As a result, most ads have few chances to be presented to the users. Second, the inefficient ad impression issue: a large proportion of search queries, which are unpopular yet relevant to many ad keywords, have no ads presented on their search result pages. Existing retrieval-based or matching-based methods either deteriorate the bidding competition or are unable to suggest novel keywords to cover more queries, which leads to inefficient ad impressions. Hao Zhou 0012, Minlie Huang, Yishun Mao, Changlei Zhu, Peng Shu, Xiaoyan Zhu 0001 |
WWW | 1 |
| 2018 | Augmenting End-to-End Dialogue Systems With Commonsense KnowledgeabstractBuilding dialogue systems that can converse naturally with humans is a challenging yet intriguing problem of artificial intelligence. In open-domain human-computer conversation, where the conversational agent is expected to respond to human utterances in an interesting and engaging way, commonsense knowledge has to be integrated into the model effectively. In this paper, we investigate the impact of providing commonsense knowledge about the concepts covered in the dialogue. Our model represents the first attempt to integrating a large commonsense knowledge base into end-to-end conversational models. In the retrieval-based scenario, we propose a model to jointly take into account message content and related commonsense for selecting an appropriate response. Our experiments suggest that the knowledge-augmented models are superior to their knowledge-free counterparts. Tom Young, Erik Cambria, Iti Chaturvedi, Hao Zhou 0012, Subham Biswas, Minlie Huang |
AAAI | 4 |
| 2018 | Emotional Chatting Machine: Emotional Conversation Generation with Internal and External MemoryabstractPerception and expression of emotion are key factors to the success of dialogue systems or conversational agents. However, this problem has not been studied in large-scale conversation generation so far. In this paper, we propose Emotional Chatting Machine (ECM) that can generate appropriate responses not only in content (relevant and grammatical) but also in emotion (emotionally consistent). To the best of our knowledge, this is the first work that addresses the emotion factor in large-scale conversation generation. ECM addresses the factor using three new mechanisms that respectively (1) models the high-level abstraction of emotion expressions by embedding emotion categories, (2) captures the change of implicit internal emotion states, and (3) uses explicit emotion expressions with an external emotion vocabulary. Experiments show that the proposed model can generate responses appropriate not only in content but also in emotion. Hao Zhou 0012, Minlie Huang, Xiaoyan Zhu 0001, Bing Liu 0001 |
AAAI | 1 |
| 2018 | On Tree-Based Neural Sentence ModelingabstractNeural networks with tree-based sentence encoders have shown better results on many downstream tasks.Most of existing tree-based encoders adopt syntactic parsing trees as the explicit structure prior.To study the effectiveness of different tree structures, we replace the parsing trees with trivial trees (i.e., binary balanced tree, left-branching tree and right-branching tree) in the encoders.Though trivial trees contain no syntactic information, those encoders get competitive or even better results on all of the ten downstream tasks we investigated.This surprising result indicates that explicit syntax guidance may not be the main contributor to the superior performances of tree-based neural sentence modeling.Further analysis show that tree modeling gives better results when crucial words are closer to the final representation.Additional experiments give more clues on how to design an effective tree-based encoder.Our code is opensource and available at https://github.com/ExplorerFreda/TreeEnc. Freda Shi, Hao Zhou 0012, Jiaze Chen, Lei Li 0005 |
EMNLP | 2 |
| 2018 | Commonsense Knowledge Aware Conversation Generation with Graph AttentionabstractCommonsense knowledge is vital to many natural language processing tasks. In this paper, we present a novel open-domain conversation generation model to demonstrate how large-scale commonsense knowledge can facilitate language understanding and generation. Given a user post, the model retrieves relevant knowledge graphs from a knowledge base and then encodes the graphs with a static graph attention mechanism, which augments the semantic information of the post and thus supports better understanding of the post. Then, during word generation, the model attentively reads the retrieved knowledge graphs and the knowledge triples within each graph to facilitate better generation through a dynamic graph attention mechanism. This is the first attempt that uses large-scale commonsense knowledge in conversation generation. Furthermore, unlike existing models that use knowledge triples (entities) separately and independently, our model treats each knowledge graph as a whole, which encodes more structured, connected semantic information in the graphs. Experiments show that the proposed model can generate more appropriate and informative responses than state-of-the-art baselines. Hao Zhou 0012, Tom Young, Minlie Huang, Haizhou Zhao, Jingfang Xu, Xiaoyan Zhu 0001 |
IJCAI | 1 |
| 2018 | Dynamic Oracle for Neural Machine Translation in Decoding Phase
Zi-Yi Dou, Hao Zhou 0012, Shujian Huang, Xinyu Dai, Jiajun Chen 0001 |
LREC | 2 |
| 2018 | BRITS: Bidirectional Recurrent Imputation for Time SeriesabstractTime series are widely used as signals in many classification/regression tasks. It is ubiquitous that time series contains many missing values. Given multiple correlated time series data, how to fill in missing values and to predict their class labels? Existing imputation methods often impose strong assumptions of the underlying data generating process, such as linear dynamics in the state space. In this paper, we propose BRITS, a novel method based on recurrent neural networks for missing value imputation in time series data. Our proposed method directly learns the missing values in a bidirectional recurrent dynamical system, without any specific assumption. The imputed values are treated as variables of RNN graph and can be effectively updated during the backpropagation. BRITS has three advantages: (a) it can handle multiple correlated missing values in time series; (b) it generalizes to time series with nonlinear dynamics underlying; (c) it provides a data-driven imputation procedure and applies to general settings with missing data. We evaluate our model on three real-world datasets, including an air quality dataset, a health-care data, and a localization data for human activity. Experiments show that our model outperforms the state-of-the-art methods in both imputation and classification/regression accuracies. Wei Cao 0007, Dong Wang 0037, Jian Li 0015, Hao Zhou 0012, Lei Li 0005, Yitan Li |
NeurIPS | 4 |
| 2018 | Modeling Past and Future for Neural Machine TranslationabstractExisting neural machine translation systems do not explicitly model what has been translated and what has not during the decoding phase. To address this problem, we propose a novel mechanism that separates the source information into two parts: translated Past contents and untranslated Future contents, which are modeled by two additional recurrent layers. The Past and Future contents are fed to both the attention model and the decoder states, which provides Neural Machine Translation (NMT) systems with the knowledge of translated and untranslated contents. Experimental results show that the proposed approach significantly improves the performance in Chinese-English, German-English, and English-German translation tasks. Specifically, the proposed model outperforms the conventional coverage model in terms of both the translation quality and the alignment error rate. Zaixiang Zheng, Hao Zhou 0012, Shujian Huang, Lili Mou, Xinyu Dai, Jiajun Chen 0001, Zhaopeng Tu |
Trans. Assoc. Comput. Linguistics | 2 |
| 2017 | Word-Context Character Embeddings for Chinese Word SegmentationabstractNeural parsers have benefited from automatically labeled data via dependencycontext word embeddings.We investigate training character embeddings on a word-based context in a similar way, showing that the simple method significantly improves state-of-the-art neural word segmentation models, beating tritraining baselines for leveraging autosegmented data. Hao Zhou 0012, Zhenting Yu, Yue Zhang 0004, Shujian Huang, Xinyu Dai, Jiajun Chen 0001 |
EMNLP | 1 |
| 2017 | Overview of the NLPCC 2017 Shared Task: Emotion Generation Challenge
Minlie Huang, Zuoxian Ye, Hao Zhou 0012 |
NLPCC | 3 |
| 2017 | A Neural Probabilistic Structured-Prediction Method for Transition-Based Natural Language ProcessingabstractWe propose a neural probabilistic structured-prediction method for transition-based natural language processing, which integrates beam search and contrastive learning. The method uses a global optimization model, which can leverage arbitrary features over non-local context. Beam search is used for efficient heuristic decoding, and contrastive learning is performed for adjusting the model according to search errors. When evaluated on both chunking and dependency parsing tasks, the proposed method achieves significant accuracy improvements over the locally normalized greedy baseline on the two tasks, respectively. Hao Zhou 0012, Yue Zhang 0004, Chuan Cheng, Shujian Huang, Xinyu Dai, Jiajun Chen 0001 |
J. Artif. Intell. Res. | 1 |
| 2016 | A Search-Based Dynamic Reranking Model for Dependency ParsingabstractWe propose a novel reranking method to extend a deterministic neural dependency parser.Different to conventional k-best reranking, the proposed model integrates search and learning by utilizing a dynamic action revising process, using the reranking model to guide modification for the base outputs and to rerank the candidates.The dynamic reranking model achieves an absolute 1.78% accuracy improvement over the deterministic baseline parser on PTB, which is the highest improvement by neural rerankers in the literature. Hao Zhou 0012, Yue Zhang 0004, Shujian Huang, Junsheng Zhou, Xinyu Dai, Jiajun Chen 0001 |
ACL (1) | 1 |
| 2016 | Context-aware Natural Language Generation for Spoken Dialogue SystemsabstractNatural language generation (NLG) is an important component of question answering(QA) systems which has a significant impact on system quality. Most tranditional QA systems based on templates or rules tend to generate rigid and stylised responses without the natural variation of human language. Furthermore, such methods need an amount of work to generate the templates or rules. To address this problem, we propose a Context-Aware LSTM model for NLG. The model is completely driven by data without manual designed templates or rules. In addition, the context information, including the question to be answered, semantic values to be addressed in the response, and the dialogue act type during interaction, are well approached in the neural network model, which enables the model to produce variant and informative responses. The quantitative evaluation and human evaluation show that CA-LSTM obtains state-of-the-art performance. Hao Zhou 0012, Minlie Huang, Xiaoyan Zhu 0001 |
COLING | 1 |
| 2016 | Evaluating a Deterministic Shift-Reduce Neural Parser for Constituent Parsing
Hao Zhou 0012, Yue Zhang 0004, Shujian Huang, Xinyu Dai, Jiajun Chen 0001 |
LREC | 1 |
| 2016 | Enhancing Shift-Reduce Constituent Parsing with Action N-Gram ModelabstractCurrent shift-reduce parsers “understand” the context by embodying a large number of binary indicator features with a discriminative model. In this article, we propose the action n-gram model, which utilizes the action sequence to help parsing disambiguation. The action n-gram model is trained on action sequences produced by parsers with the n-gram estimation method, which gives a smoothed maximum likelihood estimation of the action probability given a specific action history. We show that incorporating action n-gram models into a state-of-the-art parsing framework could achieve parsing accuracy improvements on three datasets across two languages. Hao Zhou 0012, Shujian Huang, Junsheng Zhou, Yue Zhang 0004, Huadong Chen, Xinyu Dai, Chuan Cheng, Jiajun Chen 0001 |
ACM Trans. Asian Low Resour. Lang. Inf. Process. | 1 |
| 2015 | A Neural Probabilistic Structured-Prediction Model for Transition-Based Dependency ParsingabstractHao Zhou, Yue Zhang, Shujian Huang, Jiajun Chen. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015. Hao Zhou 0012, Yue Zhang 0004, Shujian Huang, Jiajun Chen 0001 |
ACL (1) | 1 |