Yang Liu 0005

dblp:51/3710-5 · DBLP profile ↗
← Back
179ranked-venue papers
13as first author
92since 2021 · last 2026
0000-0002-3087-242XORCID · conflict

Domains — the database's venue-derived domains; a paper can count in several

Artificial intelligence and machine learning · 168 · 12 first-author · 89 since 2021Graphics, computer vision, multimedia, augmented reality and games · 29 · 2 first-author · 11 since 2021Databases, data management, data science and information retrieval · 6 · 1 first-author · 4 since 2021Applied, interdisciplinary, general and emerging computing · 5 · 1 since 2021Systems, architecture and hardware · 2 · 1 since 2021
YearPublicationVenuePosition
2026 Visual-Friendly Concept Protection via Selective Adversarial Perturbations
abstract
Personalized concept generation by tuning diffusion models with a few images raises potential legal and ethical concerns regarding privacy and intellectual property rights. Researchers attempt to prevent malicious personalization using adversarial perturbations. However, previous efforts have mainly focused on the effectiveness of protection while neglecting the visibility of perturbations. They utilize global adversarial perturbations, which introduce noticeable alterations to original images and significantly degrade visual quality. In this work, we propose the Visual-Friendly Concept Protection (VCPro) framework, which prioritizes the protection of key concepts chosen by the image owner through adversarial perturbations with lower perceptibility. To ensure these perturbations are as inconspicuous as possible, we introduce a relaxed optimization objective to identify the least perceptible yet effective adversarial perturbations, solved using the Lagrangian multiplier method. Qualitative and quantitative experiments validate that VCPro achieves a better trade-off between the visibility of perturbations and protection effectiveness, effectively prioritizing the protection of target concepts in images with less perceptible perturbations.
Xiaoyue Mi, Fan Tang, Juan Cao 0001, Peng Li 0030, Yang Liu 0005
AAAI6
2026 Writing-RL: Advancing Long-form Writing via Adaptive Curriculum Reinforcement Learning
abstract
Xuanyu Lei, Chenliang Li, Yuning Wu, Kaiming Liu, Weizhou Shen, Peng Li, Ming Yan, Fei Huang, Ya-Qin Zhang, Yang Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Xuanyu Lei, Chenliang Li 0003, Yuning Wu 0001, Kaiming Liu, Weizhou Shen, Peng Li 0030, Ming Yan 0008, Fei Huang 0002, Ya-Qin Zhang, Yang Liu 0005
ACL (1)10
2026 Scaling External Knowledge Input Beyond Context Windows of LLMs via Multi-Agent Collaboration
abstract
With the rapid advancement of post-training techniques for reasoning and information seeking, large language models (LLMs) can incorporate a large quantity of retrieved knowledge to solve complex tasks. However, the limited context window of LLMs obstructs scaling the amount of external knowledge input, prohibiting further improvement. Existing context window extension methods inevitably cause information loss. LLM-based multi-agent methods emerge as a new paradigm to handle massive input in a distributional manner, where we identify two core bottlenecks in existing agent orchestration designs. In this work, we develop a multi-agent framework, \textbf{\ExtAgents}, to overcome the bottlenecks and enable better scalability in inference-time knowledge integration without longer-context training. Benchmarked with our enhanced multi-hop question answering test, \textbf{$\boldsymbol{\infty}$Bench+}, and other public test sets including long survey generation, \ExtAgents significantly enhances the performance over existing non-training methods with the same amount of external knowledge input, regardless of whether it falls \emph{within or exceeds the context window}. Moreover, the method maintains efficiency due to high parallelism. We believe further study in the coordination of LLM agents on increasing external knowledge input could benefit real-world applications.
Zhennan Wan, Peng Li 0030, Ming Yan 0008, Fei Huang 0002, Yang Liu 0005
ACL (1)6
2026 MUSEG: Reinforcing Video Temporal Understanding via Timestamp-Aware Multi-Segment Grounding
abstract
Fuwen Luo, Shengfeng Lou, Chi Chen, Ziyue Wang, Chenliang Li, Weizhou Shen, Jiyue Guo, Peng Li, Ming Yan, Ji Zhang, Fei Huang, Yang Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Fuwen Luo, Shengfeng Lou, Chi Chen 0005, Ziyue Wang 0002, Chenliang Li 0003, Weizhou Shen, Jiyue Guo, Peng Li 0030, Ming Yan 0008, Ji Zhang 0011, Fei Huang 0002, Yang Liu 0005
ACL (1)12
2026 Beyond "I Don't Know": Evaluating LLM Self-Awareness in Discriminating Data and Model Uncertainty
abstract
Jingyi Ren, Ante Wang, Yunghwei Lai, Xiaolong Wang, Linlu Gong, Weitao Li, Weizhi Ma, Yang Liu. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Jingyi Ren, Ante Wang, Yunghwei Lai, Xiaolong Wang 0014, Linlu Gong, Weizhi Ma, Yang Liu 0005
ACL (1)8
2026 GIFT: Guided Fine-Tuning and Transfer for Enhancing Instruction-Tuned Language Models
abstract
Zhiwen Ruan, Yichao Du, Jianjie Zheng, Longyue Wang, Yun Chen, Peng Li, Jinsong Su, Yang Liu, Guanhua Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Zhiwen Ruan, Yichao Du, Jianjie Zheng, Longyue Wang, Yun Chen 0007, Peng Li 0030, Jinsong Su, Yang Liu 0005, Guanhua Chen 0001
ACL (1)8
2026 InstructDiff: Domain-Adaptive Data Selection via Contrastive Entropy for Efficient LLM Fine-Tuning
abstract
Junyou Su, He Zhu, Xiao Luo, Liyu Zhang, Hong-Yu Zhou, Yun Chen, Peng Li, Yang Liu, Guanhua Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Junyou Su, Xiao Luo 0001, Liyu Zhang 0010, Yun Chen 0007, Peng Li 0030, Yang Liu 0005, Guanhua Chen 0001
ACL (1)8
2026 SPPO: Sequence-Level PPO for Long-Horizon Reasoning Tasks
abstract
Tianyi Wang, Yixia Li, Long Li, Yibiao Chen, Shaohan Huang, Yun Chen, Peng Li, Yang Liu, Guanhua Chen. Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2026.
Yixia Li, Yibiao Chen, Shaohan Huang, Yun Chen 0007, Peng Li 0030, Yang Liu 0005, Guanhua Chen 0001
ACL (1)8
2026 Practical and Efficient x86-64 Emulation on RISC-V
abstract
As RISC-V increasingly gains traction across various computing domains, the need to run x86-64 applications on RISC-V platforms becomes critical.
Xiongchuan Tan, Yang Liu 0005, Sebastien Chevalier, Yangyu Chen 0002, Haohuan Fu
EuroSys2
2026 Interactive Visual Assessment for Text-to-Image Generation Models
abstract
Visual generation models have achieved remarkable progress in computer graphics applications but still face significant challenges in real-world deployment. Current assessment approaches for visual generation tasks typically follow an isolated three-phase framework: test input collection, model output generation, and user assessment. These fashions suffer from fixed coverage, evolving difficulty, and data leakage risks, limiting their effectiveness in comprehensively evaluating increasingly complex generation models. To address these limitations, we propose DyEval, an LLM-powered dynamic interactive visual assessment framework that facilitates collaborative evaluation between humans and generative models for text-to-image systems. DyEval features an intuitive visual interface that enables users to interactively explore and analyze model behaviors, while adaptively generating hierarchical, fine-grained, and diverse textual inputs to continuously probe the capability boundaries of the models based on their feedback. Additionally, to provide interpretable analysis for users to further improve tested models, we develop a contextual reflection module that mines failure triggers of test inputs and reflects model potential failure patterns, supporting in-depth analysis using the logical reasoning ability of LLM. Qualitative and quantitative experiments demonstrate that DyEval can effectively help users identify max up to 2.56 timesmore generation failures than conventional methods, and uncover complex and rare failure patterns, such as issues with pronoun generation and specific cultural context generation. Our framework provides valuable insights for improving generative models and has broad implications for advancing the reliability and capabilities of visual generation systems across various domains.
Xiaoyue Mi, Fan Tang, Juan Cao 0001, Qiang Sheng 0001, Ziyao Huang 0002, Peng Li 0030, Yang Liu 0005, Tong-Yee Lee
IEEE Trans. Vis. Comput. Graph.7
2025 S^3cMath: Spontaneous Step-Level Self-Correction Makes Large Language Models Better Mathematical Reasoners
abstract
Self-correction is a novel method that can stimulate the potential reasoning abilities of large language models (LLMs). It involves detecting and correcting errors during the inference process when LLMs solve reasoning problems. However, recent works do not regard self-correction as a spontaneous and intrinsic capability of LLMs. Instead, such correction is achieved through post-hoc generation, external knowledge introduction, multi-model collaboration, and similar techniques. In this paper, we propose a series of mathematical LLMs called S^3cMath, which are able to perform Spontaneous Step-level Self-correction for Mathematical reasoning. This capability helps LLMs to recognize whether their ongoing inference tends to contain errors and simultaneously correct these errors to produce a more reliable response. We proposed a method, which employs a step-level sampling approach to construct step-wise self-correction data for achieving such ability. Additionally, we implement a training strategy that uses above constructed data to equip LLMs with spontaneous step-level self-correction capacities. Our data and methods have been demonstrated to be effective across various foundation LLMs, consistently showing significant progress in evaluations on GSM8K, MATH, and other mathematical benchmarks. To the best of our knowledge, we are the first to introduce the spontaneous step-level self-correction ability of LLMs in mathematical reasoning.
Yang Liu 0005, Yixin Cao 0002, Mengdi Zhang 0002, Jian Shao 0001
AAAI3
2025 ActiView: Evaluating Active Perception Ability for Multimodal Large Language Models
abstract
Ziyue Wang, Chi Chen, Fuwen Luo, Yurui Dong, Yuanchi Zhang, Yuzhuang Xu, Xiaolong Wang, Peng Li, Yang Liu. Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2025.
Ziyue Wang 0002, Chi Chen 0005, Fuwen Luo, Yurui Dong 0001, Yuanchi Zhang, Yuzhuang Xu, Xiaolong Wang 0014, Peng Li 0030, Yang Liu 0005
ACL (1)9
2025 Leveraging Language-based Representations for Better Solving Symbol-related Problems with Large Language Models
abstract
Symbols such as numerical sequences, chemical formulas, and table delimiters exist widely, playing important roles in symbol-related tasks such as abstract reasoning, chemical property prediction, and tabular question-answering. Compared to tasks based on natural language expressions, large language models (LLMs) have limitations in understanding and reasoning on symbol-based representations, making it difficult for them to handle symbol-related problems. In this paper, we propose symbol-to-language (S2L), a method that converts symbol-based representations to language-based representations, providing valuable information for language models during reasoning. We found that, for both closed-source and open-source LLMs, the capability to solve symbol-related problems can be largely enhanced by incorporating such language-based representations. For example, by employing S2L for GPT-4, there can be substantial improvements of +21.9% and +9.5% accuracy for 1D-ARC and Dyck language tasks, respectively. There is also a consistent improvement in other six general symbol-related tasks such as table understanding and Tweet analysis. We release the GPT logs in https://github.com/THUNLP-MT/symbol2language.
Yile Wang 0001, Sijie Cheng, Zixin Sun, Peng Li 0030, Yang Liu 0005
COLING5
2025 AdaMMS: Model Merging for Heterogeneous Multimodal Large Language Models with Unsupervised Coefficient Optimization
abstract
Recently, model merging methods have demonstrated powerful strengths in combining abilities on various tasks from multiple Large Language Models (LLMs). While previous model merging methods mainly focus on merging homogeneous models with identical architecture, they meet challenges when dealing with Multimodal Large Language Models (MLLMs) with inherent heterogeneous property, including differences in model architecture and the asymmetry in the parameter space. In this work, we propose AdaMMS1, a novel model merging method tailored for heterogeneous MLLMs. Our method tackles the challenges in three steps: mapping, merging and searching. Specifically, we first design mapping function between models to apply model merging on MLLMs with different architecture. Then we apply linear interpolation on model weights to actively adapt the asymmetry in the heterogeneous MLLMs. Finally in the hyper-parameter searching step, we propose an unsupervised hyper-parameter selection method for model merging. As the first model merging method capable of merging heterogeneous MLLMs without labeled data, extensive experiments on various model combinations demonstrated that AdaMMS outperforms previous model merging methods on various vision-language benchmarks.2
Yiyang Du, Xiaochen Wang 0002, Chi Chen 0005, Jiabo Ye, Peng Li 0030, Ming Yan 0008, Ji Zhang 0011, Fei Huang 0002, Zhifang Sui, Maosong Sun 0001, Yang Liu 0005
CVPR12
2025 G2: Guided Generation for Enhanced Output Diversity in LLMs
abstract
Large Language Models (LLMs) have demonstrated exceptional performance across diverse natural language processing tasks.However, these models exhibit a critical limitation in output diversity, often generating highly similar content across multiple attempts.This limitation significantly affects tasks requiring diverse outputs, from creative writing to reasoning.Existing solutions, like temperature scaling, enhance diversity by modifying probability distributions but compromise output quality.We propose Guide-to-Generation (G2), a trainingfree plug-and-play method that enhances output diversity while preserving generation quality.G2 employs a base generator alongside dual Guides, which guide the generation process through decoding-based interventions to encourage more diverse outputs conditioned on the original query.Comprehensive experiments demonstrate that G2 effectively improves output diversity while maintaining an optimal balance between diversity and quality.
Zhiwen Ruan, Yixia Li, Yefeng Liu, Yun Chen 0007, Weihua Luo, Peng Li 0030, Yang Liu 0005, Guanhua Chen 0001
EMNLP7
2025 Adversarial Robust Memory-Based Continual Learner
abstract
Despite the remarkable advances that have been made in continual learning, the adversarial vulnerability of such methods has not been fully discussed. We delve into the adversarial robustness of memory-based continual learning algorithms and observe limited robustness improvement by directly applying adversarial training techniques. Preliminary studies reveal the twin challenges for building adversarial robust continual learners: accelerated forgetting in continual learning and gradient obfuscation in adversarial robustness. In this study, we put forward a novel adversarial robust memory-based continual learner that adjusts data logits to mitigate the forgetting of pasts caused by adversarial samples. Furthermore, we devise a gradient-based data selection mechanism to overcome the gradient obfuscation caused by limited stored data. The proposed approach can widely integrate with existing memory-based continual learning as well as adversarial training algorithms in a plug-and-play way. Extensive experiments on Split-CIFAR10/100 and Split-Tiny-ImageNet demonstrate the effectiveness of our approach, achieving up to 8.13% higher accuracy for adversarial data.
Xiaoyue Mi, Fan Tang, Zonghan Yang, Danding Wang, Juan Cao 0001, Peng Li 0030, Yang Liu 0005
ICCV7
2025 How Do Multimodal Large Language Models Handle Complex Multimodal Reasoning? Placing Them in an Extensible Escape Game
Ziyue Wang 0002, Yurui Dong 0001, Fuwen Luo, Minyuan Ruan, Zhili Cheng, Chi Chen 0005, Peng Li 0030, Yang Liu 0005
ICCV8
2025 Dynamic Low-Rank Sparse Adaptation for Large Language Models
abstract
Despite the efficacy of network sparsity in alleviating the deployment strain of Large Language Models (LLMs), it endures significant performance degradation. Applying Low-Rank Adaptation (LoRA) to fine-tune the sparse LLMs offers an intuitive approach to counter this predicament, while it holds shortcomings include: 1) The inability to integrate LoRA weights into sparse LLMs post-training, and 2) Insufficient performance recovery at high sparsity ratios. In this paper, we introduces dynamic $\textbf{Lo}$w-rank $\textbf{S}$parse $\textbf{A}$daptation $\textbf{(LoSA)}$, a novel method that seamlessly integrates low-rank adaptation into LLM sparsity within a unified framework, thereby enhancing the performance of sparse LLMs without increasing the inference latency. In particular, LoSA dynamically sparsifies the LoRA outcomes based on the corresponding sparse weights during fine-tuning, thus guaranteeing that the LoRA module can be integrated into the sparse LLMs post-training. Besides, to achieve the optimal sparse model architecture, LoSA leverages Representation Mutual Information (RMI) as an indicator to determine the importance of layers, thereby dynamically determining the optimal layer-wise sparsity rates during fine-tuning. Predicated on this, LoSA adjusts the rank of the LoRA module based on the variability in layer-wise reconstruction errors, allocating an appropriate fine-tuning for each layer to reduce the output discrepancies between dense and sparse LLMs. Extensive experiments tell that LoSA can efficiently boost the efficacy of sparse LLMs within a few hours, without introducing any additional inferential burden. For example, LoSA reduced the perplexity of sparse LLaMA-2-7B by $\textbf{68.73}$$\downarrow$ and increased zero-shot accuracy by $\textbf{16.32}$%$\uparrow$, achieving a $\textbf{2.60$\times$}$ speedup on CPU and $\textbf{2.23$\times$}$ speedup on GPU, requiring only $\textbf{45 minutes}$ of fine-tuning on $\textbf{a single}$ NVIDIA A100 80GB GPU. Code is available at https://github.com/wzhuang-xmu/LoSA.
Weizhong Huang, Yuxin Zhang 0002, Xiawu Zheng, Yang Liu 0005, Yiwu Yao, Rongrong Ji
ICLR4
2025 On the Role of Attention Heads in Large Language Model Safety
abstract
Large language models (LLMs) achieve state-of-the-art performance on multiple language tasks, yet their safety guardrails can be circumvented, leading to harmful generations. In light of this, recent research on safety mechanisms has emerged, revealing that when safety representations or component are suppressed, the safety capability of LLMs are compromised. However, existing research tends to overlook the safety impact of multi-head attention mechanisms, despite their crucial role in various model functionalities. Hence, in this paper, we aim to explore the connection between standard attention mechanisms and safety capability to fill this gap in the safety-related mechanistic interpretability. We propose an novel metric which tailored for multi-head attention, the Safety Head ImPortant Score (Ships), to assess the individual heads' contributions to model safety. Base on this, we generalize Ships to the dataset level and further introduce the Safety Attention Head AttRibution Algorithm (Sahara) to attribute the critical safety attention heads inside the model. Our findings show that special attention head has a significant impact on safety. Ablating a single safety head allows aligned model (e.g., Llama-2-7b-chat) to respond to **16$\times\uparrow$** more harmful queries, while only modifying **0.006\%** $\downarrow$ of the parameters, in contrast to the $\sim$ **5\%** modification required in previous studies. More importantly, we demonstrate that attention heads primarily function as feature extractors for safety and models fine-tuned from the same base model exhibit overlapping safety heads through comprehensive experiments. Together, our attribution approach and findings provide a novel perspective for unpacking the black box of safety mechanisms in large models.
Zhenhong Zhou, Haiyang Yu 0003, Xinghua Zhang 0001, Rongwu Xu, Fei Huang 0002, Kun Wang 0056, Yang Liu 0005, Junfeng Fang, Yongbin Li 0001
ICLR7
2025 UniSim: A Unified Simulator for Time-Coarsened Dynamics of Biomolecules
abstract
Molecular Dynamics (MD) simulations are essential for understanding the atomic-level behavior of molecular systems, giving insights into their transitions and interactions. However, classical MD techniques are limited by the trade-off between accuracy and efficiency, while recent deep learning-based improvements have mostly focused on single-domain molecules, lacking transferability to unfamiliar molecular systems. Therefore, we propose **Uni**fied **Sim**ulator (UniSim), which leverages cross-domain knowledge to enhance the understanding of atomic interactions. First, we employ a multi-head pretraining approach to learn a unified atomic representation model from a large and diverse set of molecular data. Then, based on the stochastic interpolant framework, we learn the state transition patterns over long timesteps from MD trajectories, and introduce a force guidance module for rapidly adapting to different chemical environments. Our experiments demonstrate that UniSim achieves highly competitive performance across small molecules, peptides, and proteins.
Ziyang Yu 0002, Wenbing Huang 0001, Yang Liu 0005
ICML3
2025 Zero-Shot Cyclic Peptide Design via Composable Geometric Constraints
abstract
Cyclic peptides, characterized by geometric constraints absent in linear peptides, offer enhanced biochemical properties, presenting new opportunities to address unmet medical needs. However, designing target-specific cyclic peptides remains underexplored due to limited training data. To bridge the gap, we propose CP-Composer, a novel generative framework that enables zero-shot cyclic peptide generation via composable geometric constraints. Our approach decomposes complex cyclization patterns into unit constraints, which are incorporated into a diffusion model through geometric conditioning on nodes and edges. During training, the model learns from unit constraints and their random combinations in linear peptides, while at inference, novel constraint combinations required for cyclization are imposed as input. Experiments show that our model, despite trained with linear peptides, is capable of generating diverse target-binding cyclic peptides, reaching success rates from 38% to 84% on different cyclization strategies.
Dapeng Jiang, Xiangzhe Kong, Jiaqi Han 0001, Wenbing Huang 0001, Stefano Ermon, Jianzhu Ma, Yang Liu 0005
ICML9
2025 UniMoMo: Unified Generative Modeling of 3D Molecules for De Novo Binder Design
abstract
The design of target-specific molecules such as small molecules, peptides, and antibodies is vital for biological research and drug discovery. Existing generative methods are restricted to single-domain molecules, failing to address versatile therapeutic needs or utilize cross-domain transferability to enhance model performance. In this paper, we introduce Unified generative Modeling of 3D Molecules (UniMoMo), the first framework capable of designing binders of multiple molecular domains using a single model. In particular, UniMoMo unifies the representations of different molecules as graphs of blocks, where each block corresponds to either a standard amino acid or a molecular fragment. Based on these unified representations, UniMoMo utilizes a geometric latent diffusion model for 3D molecular generation, featuring an iterative full-atom autoencoder to compress blocks into latent space points, followed by an E(3)-equivariant diffusion process. Extensive benchmarks across peptides, antibodies, and small molecules demonstrate the superiority of our unified framework over existing domain-specific models, highlighting the benefits of multi-domain training.
Xiangzhe Kong, Zishen Zhang, Ziting Zhang, Jianzhu Ma, Wenbing Huang 0001, Yang Liu 0005
ICML8
2025 Conservation-informed Graph Learning for Spatiotemporal Dynamics Prediction
abstract
Data-centric methods have shown great potential in understanding and predicting spatiotemporal dynamics, enabling better design and control of the object system. However, deep learning models often lack interpretability, fail to obey intrinsic physics, and struggle to cope with the various domains. While geometry-based methods, e.g., graph neural networks (GNNs), have been proposed to further tackle these challenges, they still need to find the implicit physical laws from large datasets and rely excessively on rich labeled data. In this paper, we herein introduce the conservation-informed GNN (CiGNN), an end-to-end explainable learning framework, to learn spatiotemporal dynamics based on limited training data. The network is designed to conform to the general conservation law via symmetry, where conservative and non-conservative information passes over a multiscale space enhanced by a latent temporal marching strategy. The efficacy of our model has been verified in various spatiotemporal systems based on synthetic and real-world datasets, showing superiority over baseline models. Results demonstrate that CiGNN exhibits remarkable accuracy and generalizability, and is readily applicable to learning for prediction of various spatiotemporal dynamics in a spatial domain with complex geometry.
Yuan Mi, Pu Ren, Hongteng Xu, Hongsheng Liu 0002, Zidong Wang 0010, Yike Guo, Ji-Rong Wen, Hao Sun 0002, Yang Liu 0005
KDD (1)9
2025 Vision-Language Models Can Self-Improve Reasoning via Reflection
abstract
Kanzhi Cheng, Li YanTao, Fangzhi Xu, Jianbing Zhang, Hao Zhou, Yang Liu. Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers). 2025.
Kanzhi Cheng, Yantao Li 0003, Fangzhi Xu, Hao Zhou 0012, Yang Liu 0005
NAACL (Long Papers)6
2025 MOF-BFN: Metal-Organic Frameworks Structure Prediction via Bayesian Flow Networks
abstract
Metal-Organic Frameworks (MOFs) have attracted considerable attention due to their unique properties including high surface area and tunable porosity, and promising applications in catalysis, gas storage, and drug delivery. Structure prediction for MOFs is a challenging task, as these frameworks are intrinsically periodic and hierarchically organized, where the entire structure is assembled from building blocks like metal nodes and organic linkers. To address this, we introduce MOF-BFN, a novel generative model for MOF structure prediction based on Bayesian Flow Networks (BFNs). Given the local geometry of building blocks, MOF-BFN jointly predicts the lattice parameters, as well as the positions and orientations of all building blocks within the unit cell. In particular, the positions are modelled in the fractional coordinate system to naturally incorporate the periodicity. Meanwhile, the orientations are modeled as unit quaternions sampled from learned Bingham distributions via the proposed Bingham BFN, enabling effective orientation generation on the 4D unit hypersphere. Experimental results demonstrate that MOF-BFN achieves state-of-the-art performance across multiple tasks, including structure prediction, geometric property evaluation, and de novo generation, offering a promising tool for designing complex MOF materials.
Wenbing Huang 0001, Yuxuan Song 0002, Yawen Ouyang, Yu Rong 0001, Tingyang Xu, Hao Zhou 0012, Wei-Ying Ma, Yang Liu 0005
NeurIPS12
2025 Beyond the Surface: Enhancing LLM-as-a-Judge Alignment with Human via Internal Representations
abstract
The growing scale of evaluation tasks has led to the widespread adoption of automated evaluation using LLMs, a paradigm known as “LLM-as-a-judge”. However, improving its alignment with human preferences without complex prompts or fine-tuning remains challenging. Previous studies mainly optimize based on shallow outputs, overlooking rich cross-layer representations. In this work, motivated by preliminary findings that middle-to-upper layers encode semantically and task-relevant representations that are often more aligned with human judgments than the final layer, we propose LAGER, a post-hoc, plug-and-play framework for improving the alignment of LLM-as-a-Judge point-wise evaluations with human scores by leveraging internal representations. LAGER produces fine-grained judgment scores by aggregating cross-layer score-token logits and computing the expected score from a softmax-based distribution, while keeping the LLM backbone frozen and ensuring no impact on the inference process. LAGER fully leverages the complementary information across different layers, overcoming the limitations of relying solely on the final layer. We evaluate our method on the standard alignment benchmarks Flask, HelpSteer, and BIGGen using Spearman correlation, and find that LAGER achieves improvements of up to 7.5% over the best baseline across these benchmarks. Without reasoning steps, LAGER matches or outperforms reasoning-based methods. Experiments on downstream applications, such as data selection and emotional understanding, further show the generalization of LAGER.
Peng Lai, Jianjie Zheng, Sijie Cheng, Yun Chen 0007, Peng Li 0030, Yang Liu 0005, Guanhua Chen 0001
NeurIPS6
2025 Learning 3D Anisotropic Noise Distributions Improves Molecular Force Fields
abstract
Coordinate denoising has emerged as a promising method for 3D molecular pretraining due to its theoretical connection to learning molecular force field. However, existing denoising methods rely on oversimplied molecular dynamics that assume atomic motions to be isotropic and homoscedastic. To address these limitations, we propose a novel denoising framework AniDS: Anisotropic Variational Autoencoder for 3D Molecular Denoising. AniDS introduces a structure-aware anisotropic noise generator that can produce atom-specific, full covariance matrices for Gaussian noise distributions to better reflect directional and structural variability in molecular systems. These covariances are derived from pairwise atomic interactions as anisotropic corrections to an isotropic base. Our design ensures that the resulting covariance matrices are symmetric, positive semi-definite, and SO(3)-equivariant, while providing greater capacity to model complex molecular dynamics. Extensive experiments show that AniDS outperforms prior isotropic and homoscedastic denoising models and other leading methods on the MD17 and OC22 benchmarks, achieving average relative improvements of 8.9% and 6.2% in force prediction accuracy. Our case study on a crystal and molecule structure shows that AniDS adaptively suppresses noise along the bonding direction, consistent with physicochemical principles. Our code is available at https://github.com/ZeroKnighting/AniDS.
Xixian Liu, Zhiyuan Liu 0001, Yurou Liu, Yang Liu 0005, Ziheng Lu, Wenbing Huang 0001, Yang Zhang 0094, Yixin Cao 0002
NeurIPS5
2025 Latent Retrieval Augmented Generation of Cross-Domain Protein Binders
abstract
Designing protein binders targeting specific sites, which requires to generate realistic and functional interaction patterns, is a fundamental challenge in drug discovery. Current structure-based generative models are limited in generating nterfaces with sufficient rationality and interpretability. In this paper, we propose **R**etrieval-**A**ugmented **Di**ffusion for **A**lig**n**ed interfa**ce** (**RADiAnce**), a new framework that leverages known interfaces to guide the design of novel binders. By unifying retrieval and generation in a shared contrastive latent space, our model efficiently identifies relevant interfaces for a given binding site and seamlessly integrates them through a conditional latent diffusion generator, enabling cross-domain interface transfer. Extensive exeriments show that **RADiAnce** significantly outperforms baseline models across multiple metrics, including binding affinity and recovery of geometries and interactions. Additional experimental results validate cross-domain generalization, demonstrating that retrieving interfaces from diverse domains, such as peptides, antibodies, and protein fragments, enhances the generation performance of binders for other domains. Our work establishes a new paradigm for protein binder design that successfully bridges retrieval-based knowledge and generative AI, opening new possibilities for drug discovery.
Zishen Zhang, Xiangzhe Kong, Wenbing Huang 0001, Yang Liu 0005
NeurIPS4
2025 A survey of geometric graph neural networks: data structures, models and applications
abstract
Abstract Geometric graphs are a special kind of graph with geometric features, which are vital to model many scientific problems. Unlike generic graphs, geometric graphs often exhibit physical symmetries of translations, rotations, and reflections, making them ineffectively processed by current Graph Neural Networks (GNNs). To address this issue, researchers proposed a variety of geometric GNNs equipped with invariant/equivariant properties to better characterize the geometry and topology of geometric graphs. Given the current progress in this field, it is imperative to conduct a comprehensive survey of data structures, models, and applications related to geometric GNNs. In this paper, based on the necessary but concise mathematical preliminaries, we formalize geometric graph as the data structure, on top of which we provide a unified view of existing models from the geometric message passing perspective. Additionally, we summarize the applications as well as the related datasets to facilitate later research for methodology development and experimental evaluation. We also discuss the challenges and future potential directions of geometric GNNs at the end of this survey.
Jiaqi Han 0001, Jiacheng Cen, Liming Wu, Zongzhao Li, Xiangzhe Kong, Ziyang Yu 0002, Tingyang Xu, Fandi Wu, Hongteng Xu, Zhewei Wei, Deli Zhao, Yang Liu 0005, Yu Rong 0001, Wenbing Huang 0001
Frontiers Comput. Sci.14
2024 Model Composition for Multimodal Large Language Models
abstract
Chi Chen, Yiyang Du, Zheng Fang, Ziyue Wang, Fuwen Luo, Peng Li, Ming Yan, Ji Zhang, Fei Huang, Maosong Sun, Yang Liu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Chi Chen 0005, Yiyang Du, Ziyue Wang 0002, Fuwen Luo, Peng Li 0030, Ming Yan 0008, Ji Zhang 0011, Fei Huang 0002, Maosong Sun 0001, Yang Liu 0005
ACL (1)11
2024 CODIS: Benchmarking Context-dependent Visual Comprehension for Multimodal Large Language Models
abstract
Fuwen Luo, Chi Chen, Zihao Wan, Zhaolu Kang, Qidong Yan, Yingjie Li, Xiaolong Wang, Siyu Wang, Ziyue Wang, Xiaoyue Mi, Peng Li, Ning Ma, Maosong Sun, Yang Liu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Fuwen Luo, Chi Chen 0005, Zihao Wan, Zhaolu Kang, Qidong Yan, Yingjie Li 0009, Xiaolong Wang 0014, Ziyue Wang 0002, Xiaoyue Mi, Peng Li 0030, Maosong Sun 0001, Yang Liu 0005
ACL (1)14
2024 Browse and Concentrate: Comprehending Multimodal Content via Prior-LLM Context Fusion
abstract
Ziyue Wang, Chi Chen, Yiqi Zhu, Fuwen Luo, Peng Li, Ming Yan, Ji Zhang, Fei Huang, Maosong Sun, Yang Liu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Ziyue Wang 0002, Chi Chen 0005, Yiqi Zhu, Fuwen Luo, Peng Li 0030, Ming Yan 0008, Ji Zhang 0011, Fei Huang 0002, Maosong Sun 0001, Yang Liu 0005
ACL (1)10
2024 Reasoning in Conversation: Solving Subjective Tasks through Dialogue Simulation for Large Language Models
abstract
Xiaolong Wang, Yile Wang, Yuanchi Zhang, Fuwen Luo, Peng Li, Maosong Sun, Yang Liu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Xiaolong Wang 0014, Yile Wang 0001, Yuanchi Zhang, Fuwen Luo, Peng Li 0030, Maosong Sun 0001, Yang Liu 0005
ACL (1)7
2024 Enhancing Multilingual Capabilities of Large Language Models through Self-Distillation from Resource-Rich Languages
abstract
Yuanchi Zhang, Yile Wang, Zijun Liu, Shuo Wang, Xiaolong Wang, Peng Li, Maosong Sun, Yang Liu. Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). 2024.
Yuanchi Zhang, Yile Wang 0001, Shuo Wang 0013, Xiaolong Wang 0014, Peng Li 0030, Maosong Sun 0001, Yang Liu 0005
ACL (1)8
2024 Topology-preserving Adversarial Training for Alleviating Natural Accuracy Degradation
Xiaoyue Mi, Fan Tang, Yepeng Weng, Danding Wang, Juan Cao 0001, Sheng Tang, Peng Li 0030, Yang Liu 0005
BMVC8
2024 DEEM: Dynamic Experienced Expert Modeling for Stance Detection
abstract
Recent work has made a preliminary attempt to use large language models (LLMs) to solve the stance detection task, showing promising results. However, considering that stance detection usually requires detailed background knowledge, the vanilla reasoning method may neglect the domain knowledge to make a professional and accurate analysis. Thus, there is still room for improvement of LLMs reasoning, especially in leveraging the generation capability of LLMs to simulate specific experts (i.e., multi-agents) to detect the stance. In this paper, different from existing multi-agent works that require detailed descriptions and use fixed experts, we propose a Dynamic Experienced Expert Modeling (DEEM) method which can leverage the generated experienced experts and let LLMs reason in a semi-parametric way, making the experts more generalizable and reliable. Experimental results demonstrate that DEEM consistently achieves the best results on three standard benchmarks, outperforms methods with self-consistency reasoning, and reduces the bias of LLMs.
Xiaolong Wang 0014, Yile Wang 0001, Sijie Cheng, Peng Li 0030, Yang Liu 0005
LREC/COLING5
2024 Pluggable Neural Machine Translation Models via Memory-augmented Adapters
abstract
Although neural machine translation (NMT) models perform well in the general domain, it remains rather challenging to control their generation behavior to satisfy the requirement of different users. Given the expensive training cost and the data scarcity challenge of learning a new model from scratch for each user requirement, we propose a memory-augmented adapter to steer pretrained NMT models in a pluggable manner. Specifically, we construct a multi-granular memory based on the user-provided text samples and propose a new adapter architecture to combine the model representations and the retrieved results. We also propose a training strategy using memory dropout to reduce spurious dependencies between the NMT model and the memory. We validate our approach on both style- and domain-specific experiments and the results indicate that our method can outperform several representative pluggable baselines.
Yuzhuang Xu, Shuo Wang 0013, Peng Li 0030, Xuebo Liu 0002, Xiaolong Wang 0014, Yang Liu 0005
LREC/COLING7
2024 ToolRerank: Adaptive and Hierarchy-Aware Reranking for Tool Retrieval
abstract
Tool learning aims to extend the capabilities of large language models (LLMs) with external tools. A major challenge in tool learning is how to support a large number of tools, including unseen tools. To address this challenge, previous studies have proposed retrieving suitable tools for the LLM based on the user query. However, previously proposed methods do not consider the differences between seen and unseen tools, nor do they take the hierarchy of the tool library into account, which may lead to suboptimal performance for tool retrieval. Therefore, to address the aforementioned issues, we propose ToolRerank, an adaptive and hierarchy-aware reranking method for tool retrieval to further refine the retrieval results. Specifically, our proposed ToolRerank includes Adaptive Truncation, which truncates the retrieval results related to seen and unseen tools at different positions, and Hierarchy-Aware Reranking, which makes retrieval results more concentrated for single-tool queries and more diverse for multi-tool queries. Experimental results show that ToolRerank can improve the quality of the retrieval results, leading to better execution results generated by the LLM.
Yuanhang Zheng, Peng Li 0021, Wei Liu 0302, Yang Liu 0005, Jian Luan 0001, Bin Wang 0004
LREC/COLING4
2024 EgoThink: Evaluating First-Person Perspective Thinking Capability of Vision-Language Models
abstract
Vision-language models (VLMs) have recently shown promising results in traditional downstream tasks. Evaluation studies have emerged to assess their abilities, with the majority focusing on the third-person perspective, and only a few addressing specific tasks from the first-person per-spective. However, the capability of VLMs to “think” from a first-person perspective, a crucial attribute for advancing autonomous agents and robotics, remains largely unexplored. To bridge this research gap, we introduce EgoThink, a novel visual question-answering benchmark that encompasses six core capabilities with twelve detailed dimensions. The benchmark is constructed using selected clips from ego-centric videos, with manually annotated question-answer pairs containing first-person information. To comprehensively assess VLMs, we evaluate twenty-one popular VLMs on EgoThink. Moreover, given the open-ended format of the answers, we use GPT-4 as the automatic judge to compute single-answer grading. Experimental results indicate that although GPT-4V leads in numerous dimensions, all evaluated VLMs still possess considerable potential for improvement in first-person perspective tasks. Meanwhile, enlarging the number of trainable parameters has the most significant impact on model performance on EgoThink. In conclusion, EgoThink serves as a valuable addition to existing evaluation benchmarks for VLMs, providing an indispensable resource for future research in the realm of embodied artificial intelligence and robotics.
Sijie Cheng, Zhicheng Guo, Kechen Fang, Peng Li 0030, Huaping Liu 0001, Yang Liu 0005
CVPR7
2024 EvalCrafter: Benchmarking and Evaluating Large Video Generation Models
abstract
The vision and language generative models have been overgrown in recent years. For video generation, various open-sourced models and public-available services have been developed to generate high-quality videos. However, these methods often use a few metrics, e.g., FVD [56] or IS [45], to evaluate the performance. We argue that it is hard to judge the large conditional generative models from the simple metrics since these models are often trained on very large datasets with multi-aspect abilities. Thus, we propose a novel framework and pipeline for exhaustively evaluating the performance of the generated videos. Our approach involves generating a diverse and comprehensive list of 700 prompts for text-to-video generation, which is based on an analysis of real-world user data and generated with the assistance of a large language model. Then, we evaluate the state-of-the-art video generative models on our carefully designed benchmark, in terms of visual qualities, content qualities, motion qualities, and text-video alignment with 17 well-selected objective metrics. To obtain the finalleaderboard of the models, we further fit a series of coefficients to align the objective metrics to the users' opinions. Based on the proposed human alignment method, our final score shows a higher correlation than simply averaging the metrics, showing the effectiveness of the proposed evaluation method.
Yaofang Liu, Xiaodong Cun, Xuebo Liu 0002, Xintao Wang 0002, Yong Zhang 0034, Haoxin Chen, Yang Liu 0005, Tieyong Zeng, Raymond Chan 0001, Ying Shan
CVPR7
2024 FuseGen: PLM Fusion for Data-generation based Zero-shot Learning
abstract
Data-generation based zero-shot learning, although effective in training Small Task-specific Models (STMs) via synthetic datasets generated by Pre-trained Language Models (PLMs), is often limited by the low quality of such synthetic datasets.Previous solutions have primarily focused on single PLM settings, where synthetic datasets are typically restricted to specific sub-spaces and often deviate from real-world distributions, leading to severe distribution bias.To mitigate such bias, we propose FuseGen, a novel data-generation based zero-shot learning framework that introduces a new criteria for subset selection from synthetic datasets via utilizing multiple PLMs and trained STMs.The chosen subset provides in-context feedback to each PLM, enhancing dataset quality through iterative data generation.Trained STMs are then used for sample re-weighting as well, further improving data quality.Extensive experiments across diverse tasks demonstrate that FuseGen substantially outperforms existing methods, highly effective in boosting STM performance in a PLM-agnostic way. 1
Tianyuan Zou, Yang Liu 0005, Peng Li 0030, Jianqing Zhang, Ya-Qin Zhang
EMNLP2
2024 Space Group Constrained Crystal Generation
abstract
Crystals are the foundation of numerous scientific and industrial applications. While various learning-based approaches have been proposed for crystal generation, existing methods neglect the spacegroup constraint which is crucial in describing the geometry of crystals and closely relevant to many desirable properties. However, considering spacegroup constraint is challenging owing to its diverse and nontrivial forms. In this paper, we reduce the spacegroup constraint into an equivalent formulation that is more tractable to be handcrafted into the generation process. In particular, we translate the spacegroup constraint into two cases: the basis constraint of the invariant exponential space of the lattice matrix and the Wyckoff position constraint of the fractional coordinates. Upon the derived constraints, we then propose DiffCSP++, a novel diffusion model that has enhanced a previous work DiffCSP by further taking spacegroup constraint into account. Experiments on several popular datasets verify the benefit of the involvement of the spacegroup constraint, and show that our DiffCSP++ achieves the best or comparable performance on crystal structure prediction and ab initio crystal generation.
Wenbing Huang 0001, Deli Zhao, Yang Liu 0005
ICLR5
2024 Rigid Protein-Protein Docking via Equivariant Elliptic-Paraboloid Interface Prediction
abstract
The study of rigid protein-protein docking plays an essential role in a variety of tasks such as drug design and protein engineering. Recently, several learning-based methods have been proposed for the task, exhibiting much faster docking speed than those computational methods. In this paper, we propose a novel learning-based method called ElliDock, which predicts an elliptic paraboloid to represent the protein-protein docking interface. To be specific, our model estimates elliptic paraboloid interfaces for the two input proteins respectively, and obtains the roto-translation transformation for docking by making two interfaces coincide. By its design, ElliDock is independently equivariant with respect to arbitrary rotations/translations of the proteins, which is an indispensable property to ensure the generalization of the docking process. Experimental evaluations show that ElliDock achieves the fastest inference time among all compared methods, and outperforms state-of-the-art learning-based methods, like DiffDock-PP and Alphafold-Multimer, for particularly antibody-antigen docking.
Ziyang Yu 0002, Wenbing Huang 0001, Yang Liu 0005
ICLR3
2024 Generalist Equivariant Transformer Towards 3D Molecular Interaction Learning
abstract
Many processes in biology and drug discovery involve various 3D interactions between molecules, such as protein and protein, protein and small molecule, etc. Given that different molecules are usually represented in different granularity, existing methods usually encode each type of molecules independently with different models, leaving it defective to learn the various underlying interaction physics. In this paper, we first propose to universally represent an arbitrary 3D complex as a geometric graph of sets, shedding light on encoding all types of molecules with one model. We then propose a Generalist Equivariant Transformer (GET) to effectively capture both domain-specific hierarchies and domain-agnostic interaction physics. To be specific, GET consists of a bilevel attention module, a feed-forward module and a layer normalization module, where each module is E(3) equivariant and specialized for handling sets of variable sizes. Notably, in contrast to conventional pooling-based hierarchical models, our GET is able to retain fine-grained information of all levels. Extensive experiments on the interactions between proteins, small molecules and RNA/DNAs verify the effectiveness and generalization capability of our proposed method across different domains.
Xiangzhe Kong, Wenbing Huang 0001, Yang Liu 0005
ICML3
2024 Equivariant Diffusion for Crystal Structure Prediction
abstract
In addressing the challenge of Crystal Structure Prediction (CSP), symmetry-aware deep learning models, particularly diffusion models, have been extensively studied, which treat CSP as a conditional generation task. However, ensuring permutation, rotation, and periodic translation equivariance during diffusion process remains incompletely addressed. In this work, we propose EquiCSP, a novel equivariant diffusion-based generative model. We not only address the overlooked issue of lattice permutation equivariance in existing models, but also develop a unique noising algorithm that rigorously maintains periodic translation equivariance throughout both training and inference processes. Our experiments indicate that EquiCSP significantly surpasses existing models in terms of generating accurate structures and demonstrates faster convergence during the training process.
Peijia Lin, Pin Chen, Qing Mo, Jianhuan Cen, Wenbing Huang 0001, Yang Liu 0005, Dan Huang 0001, Yutong Lu
ICML7
2024 Position: Towards Unified Alignment Between Agents, Humans, and Environment
abstract
The rapid progress of foundation models has led to the prosperity of autonomous agents, which leverage the universal capabilities of foundation models to conduct reasoning, decision-making, and environmental interaction. However, the efficacy of agents remains limited when operating in intricate, realistic environments. In this work, we introduce the principles of Unified Alignment for Agents (UA$^2$), which advocate for the simultaneous alignment of agents with human intentions, environmental dynamics, and self-constraints such as the limitation of monetary budgets. From the perspective of UA$^2$, we review the current agent research and highlight the neglected factors in existing agent benchmarks and method candidates. We also conduct proof-of-concept studies by introducing realistic features to WebShop, including user profiles demonstrating intentions, personalized reranking reflecting complex environmental dynamics, and runtime cost statistics as self-constraints. We then follow the principles of UA$^2$ to propose an initial design of our agent and benchmark its performance with several candidate baselines in the retrofitted WebShop. The extensive experimental results further prove the importance of the principles of UA$^2$. Our research sheds light on the next steps of autonomous agent research with improved general problem-solving abilities.
Zonghan Yang, Kaiming Liu, Fangzhou Xiong, Yile Wang 0001, Zeyuan Yang 0002, Zhenhe Zhang, Fuwen Luo, Zhicheng Guo, Peng Li 0030, Yang Liu 0005
ICML14
2024 Learning Superconductivity from Ordered and Disordered Material Structures
abstract
Superconductivity is a fascinating phenomenon observed in certain materials under certain conditions. However, some critical aspects of it, such as the relationship between superconductivity and materials' chemical/structural features, still need to be understood. Recent successes of data-driven approaches in material science strongly inspire researchers to study this relationship with them, but a corresponding dataset is still lacking. Hence, we present a new dataset for data-driven approaches, namely SuperCon3D, containing both 3D crystal structures and experimental superconducting transition temperature (Tc) for the first time. Based on SuperCon3D, we propose two deep learning methods for designing high Tc superconductors. The first is SODNet, a novel equivariant graph attention model for screening known structures, which differs from existing models in incorporating both ordered and disordered geometric content. The second is a diffusion generative model DiffCSP-SC for creating new structures, which enables high Tc-targeted generation. Extensive experiments demonstrate that both our proposed dataset and models are advantageous for designing new high Tc superconducting candidates.
Pin Chen, Luoxuan Peng, Qing Mo, Zhen Wang 0036, Wenbing Huang 0001, Yang Liu 0005, Yutong Lu
NeurIPS7
2024 3D Structure Prediction of Atomic Systems with Flow-based Direct Preference Optimization
abstract
Predicting high-fidelity 3D structures of atomic systems is a fundamental yet challenging problem in scientific domains. While recent work demonstrates the advantage of generative models in this realm, the exploration of different probability paths are still insufficient, and hallucinations during sampling are persistently occurring. To address these pitfalls, we introduce FlowDPO, a novel framework that explores various probability paths with flow matching models and further suppresses hallucinations using Direct Preference Optimization (DPO) for structure generation. Our approach begins with a pre-trained flow matching model to generate multiple candidate structures for each training sample. These structures are then evaluated and ranked based on their distance to the ground truth, resulting in an automatic preference dataset. Using this dataset, we apply DPO to optimize the original model, improving its performance in generating structures closely aligned with the desired reference distribution. As confirmed by our theoretical analysis, such paradigm and objective function are compatible with arbitrary Gaussian paths, exhibiting favorable universality. Extensive experimental results on antibodies and crystals demonstrate substantial benefits of our FlowDPO, highlighting its potential to advance the field of 3D structure prediction with generative models.
Xiangzhe Kong, Wenbing Huang 0001, Yang Liu 0005
NeurIPS4
2024 Full-Atom Peptide Design with Geometric Latent Diffusion
abstract
Peptide design plays a pivotal role in therapeutics, allowing brand new possibility to leverage target binding sites that are previously undruggable. Most existing methods are either inefficient or only concerned with the target-agnostic design of 1D sequences. In this paper, we propose a generative model for full-atom Peptide design with Geometric LAtent Diffusion (PepGLAD) given the binding site. We first establish a benchmark consisting of both 1D sequences and 3D structures from Protein Data Bank (PDB) and literature for systematic evaluation. We then identify two major challenges of leveraging current diffusion-based models for peptide design: the full-atom geometry and the variable binding geometry. To tackle the first challenge, PepGLAD derives a variational autoencoder that first encodes full-atom residues of variable size into fixed-dimensional latent representations, and then decodes back to the residue space after conducting the diffusion process in the latent space. For the second issue, PepGLAD explores a receptor-specific affine transformation to convert the 3D coordinates into a shared standard space, enabling better generalization ability across different binding shapes. Experimental Results show that our method not only improves diversity and binding affinity significantly in the task of sequence-structure co-design, but also excels at recovering reference structures for binding conformation generation.
Xiangzhe Kong, Yinjun Jia, Wenbing Huang 0001, Yang Liu 0005
NeurIPS4
2024 FedEL: Federated ensemble learning for non-iid data
Xing Wu 0001, Jie Pei, Xianhua Han, Yen-Wei Chen 0001, Junfeng Yao, Yang Liu 0005, Quan Qian, Yike Guo
Expert Syst. Appl.6
2024 Gradual Syntactic Label Replacement for Language Model Pre-Training
abstract
Pre-training serves as a foundation of recent NLP models, where language modeling tasks are performed over large texts. Typical models like BERT and GPT take the corpus as a whole and treat each word equally for language modeling. However, recent works show that the naturally existing frequency bias in the raw corpus may limit the power of the language model. In this article, we propose a multi-stage training strategy that gradually increases the training vocabulary by modifying the training data. Specifically, we leverage the syntactic structure as a bridge for infrequent words and replace them with the corresponding syntactic labels, then we recover their original lexical surface for further training. Such strategy results in an easy-to-hard curriculum learning process, where the model learns the most common words and some basic syntax concepts, before recognizing a large number of uncommon words via their specific usages and the previously learned category knowledge. Experimental results show that such a method can improve the performance of both discriminative and generative pre-trained language models on benchmarks and various downstream tasks.
Yile Wang 0001, Yue Zhang 0004, Peng Li 0030, Yang Liu 0005
IEEE ACM Trans. Audio Speech Lang. Process.4
2024 Black-Box Prompt Tuning With Subspace Learning
abstract
Black-box prompt tuning employs derivative-free optimization algorithms to learn prompts within low-dimensional subspaces rather than back-propagating through the network of Large Language Models (LLMs). Recent studies reveal that black-box prompt tuning lacks versatility across tasks and LLMs, which we believe is related to the suboptimal choice of subspaces. In this paper, we introduceBlack-box prompt tuning withSubspaceLearning (BSL) to enhance the versatility of black-box prompt tuning. Based on the assumption that nearly optimal prompts for similar tasks reside in a common subspace, we propose identifying such subspaces through meta-learning on a collection of similar source tasks. Consequently, for a target task that shares similarities with the source tasks, we expect that optimizing within the identified subspace can yield a prompt that performs well on the target task. Experimental results confirm that our BSL framework consistently achieves competitive performance across various downstream tasks and LLMs.
Yuanhang Zheng, Zhixing Tan, Peng Li 0030, Yang Liu 0005
IEEE ACM Trans. Audio Speech Lang. Process.4
2023 Energy-Motivated Equivariant Pretraining for 3D Molecular Graphs
abstract
Pretraining molecular representation models without labels is fundamental to various applications. Conventional methods mainly process 2D molecular graphs and focus solely on 2D tasks, making their pretrained models incapable of characterizing 3D geometry and thus defective for downstream 3D tasks. In this work, we tackle 3D molecular pretraining in a complete and novel sense. In particular, we first propose to adopt an equivariant energy-based model as the backbone for pretraining, which enjoys the merits of fulfilling the symmetry of 3D space. Then we develop a node-level pretraining loss for force prediction, where we further exploit the Riemann-Gaussian distribution to ensure the loss to be E(3)-invariant, enabling more robustness. Moreover, a graph-level noise scale prediction task is also leveraged to further promote the eventual performance. We evaluate our model pretrained from a large-scale 3D dataset GEOM-QM9 on two challenging 3D benchmarks: MD17 and QM9. Experimental results demonstrate the efficacy of our method against current state-of-the-art pretraining approaches, and verify the validity of our design for each proposed component. Code is available at https://github.com/jiaor17/3D-EMGP.
Jiaqi Han 0001, Wenbing Huang 0001, Yu Rong 0001, Yang Liu 0005
AAAI5
2023 Weakly Supervised Vision-and-Language Pre-training with Relative Representations
abstract
Weakly supervised vision-and-language pretraining (WVLP), which learns cross-modal representations with limited cross-modal supervision, has been shown to effectively reduce the data cost of pre-training while maintaining decent performance on downstream tasks.However, current WVLP methods use only local descriptions of images, i.e., object tags, as cross-modal anchors to construct weaklyaligned image-text pairs for pre-training.This affects the data quality and thus the effectiveness of pre-training.In this paper, we propose to directly take a small number of aligned image-text pairs as anchors, and represent each unaligned image and text by its similarities to these anchors, i.e., relative representations.We build a WVLP framework based on the relative representations, namely RELIT 1 , which collects high-quality weakly-aligned imagetext pairs from large-scale image-only and text-only data for pre-training through relative representation-based retrieval and generation.Experiments on four downstream tasks show that RELIT achieves new state-of-the-art results under the weakly supervised setting 2 .
Chi Chen 0005, Peng Li 0030, Maosong Sun 0001, Yang Liu 0005
ACL (1)4
2023 An Extensible Plug-and-Play Method for Multi-Aspect Controllable Text Generation
abstract
Recently, multi-aspect controllable text generation that controls the generated text in multiple aspects (e.g., sentiment, topic, and keywords) has attracted increasing attention.Although methods based on parameter efficient tuning like prefix-tuning could achieve multi-aspect controlling in a plug-and-play way, the mutual interference of multiple prefixes leads to significant degeneration of constraints and limits their extensibility to training-time unseen aspect combinations.In this work, we provide a theoretical lower bound for the interference and empirically found that the interference grows with the number of layers where prefixes are inserted.Based on these analyses, we propose using trainable gates to normalize the intervention of prefixes to restrain the growing interference.As a result, controlling training-time unseen combinations of aspects can be realized by simply concatenating corresponding plugins such that new constraints can be extended at a lower cost.In addition, we propose a unified way to process both categorical and free-form constraints.Experiments on text generation and machine translation demonstrate the superiority of our approach over baselines on constraint accuracy, text quality, and extensibility. 1
Xuancheng Huang, Peng Li 0030, Maosong Sun 0001, Yang Liu 0005
ACL (1)6
2023 Knowledge Transfer in Incremental Learning for Multilingual Neural Machine Translation
abstract
In the real-world scenario, a longstanding goal of multilingual neural machine translation (MNMT) is that a single model can incrementally adapt to new language pairs without accessing previous training data.In this scenario, previous studies concentrate on overcoming catastrophic forgetting while lacking encouragement to learn new knowledge from incremental language pairs, especially when the incremental language is not related to the set of original languages.To better acquire new knowledge, we propose a knowledge transfer method that can efficiently adapt original MNMT models to diverse incremental language pairs.The method flexibly introduces the knowledge from an external model into original models, which encourages the models to learn new language pairs, completing the procedure of knowledge transfer.Moreover, all original parameters are frozen to ensure that translation qualities on original language pairs are not degraded.Experimental results show that our method can learn new knowledge from diverse language pairs incrementally meanwhile maintaining performance on original language pairs, outperforming various strong baselines in incremental learning for MNMT.
Peng Li 0030, Jin Ma 0003, Ting Yao 0004, Yang Liu 0005
ACL (1)5
2023 Continual Knowledge Distillation for Neural Machine Translation
abstract
While many parallel corpora are not publicly accessible for data copyright, data privacy and competitive differentiation reasons, trained translation models are increasingly available on open platforms.In this work, we propose a method called continual knowledge distillation to take advantage of existing translation models to improve one model of interest.The basic idea is to sequentially transfer knowledge from each trained model to the distilled model.Extensive experiments on Chinese-English and German-English datasets show that our method achieves significant and consistent improvements over strong baselines under both homogeneous and heterogeneous trained model settings and is robust to malicious models.
Yuanchi Zhang, Peng Li 0030, Maosong Sun 0001, Yang Liu 0005
ACL (1)4
2023 Bridging the Gap between Decision and Logits in Decision-based Knowledge Distillation for Pre-trained Language Models
abstract
Conventional knowledge distillation (KD) methods require access to the internal information of teachers, e.g., logits.However, such information may not always be accessible for large pre-trained language models (PLMs).In this work, we focus on decision-based KD for PLMs, where only teacher decisions (i.e., top-1 labels) are accessible.Considering the information gap between logits and decisions, we propose a novel method to estimate logits from the decision distributions.Specifically, decision distributions can be both derived as a function of logits theoretically and estimated with test-time data augmentation empirically.By combining the theoretical and empirical estimations of the decision distributions together, the estimation of logits can be successfully reduced to a simple root-finding problem.Extensive experiments show that our method significantly outperforms strong baselines on both natural language understanding and machine reading comprehension datasets.1
Qinhong Zhou, Zonghan Yang, Peng Li 0030, Yang Liu 0005
ACL (1)4
2023 Learn and Consolidate: Continual Adaptation for Zero-Shot and Multilingual Neural Machine Translation
abstract
Although existing multilingual neural machine translation (MNMT) models have demonstrated remarkable performance to handle multiple translation directions in a single model and achieved zero-shot translation between language pairs unseen in training, they still suffer from relatively poor translation qualities for some language pairs.A practical scenario is that how to continually update MNMT models for both supervised and zero-shot translations when limited new data arrives.To this end, we propose a two-stage approach that encourages original models to acquire language-agnostic multilingual representations from new data, and preserves the model architecture without introducing parameters.Experimental results and further analysis demonstrate that our method can efficiently improve performance of existing MNMT models in translation directions where they are initially weak, and mitigates the degeneration in the original well-performing translation directions, offering flexibility in the real-world scenario.
Peng Li 0030, Junpeng Liu 0002, Maosong Sun 0001, Yang Liu 0005
EMNLP5
2023 Failures Pave the Way: Enhancing Large Language Models through Tuning-free Rule Accumulation
abstract
Large Language Models (LLMs) have showcased impressive performance.However, due to their inability to capture relationships among samples, these frozen LLMs inevitably keep repeating similar mistakes.In this work, we propose our Tuning-free Rule Accumulation (TRAN) framework, which guides LLMs in improving their performance by learning from previous mistakes.Considering data arrives sequentially, LLMs gradually accumulate rules from incorrect cases, forming a rule collection.These rules are then utilized by the LLMs to avoid making similar mistakes when processing subsequent inputs.Moreover, the rules remain independent of the primary prompts, seamlessly complementing prompt design strategies.Experimentally, we show that TRAN improves over recent baselines by a large margin.
Zeyuan Yang 0002, Peng Li 0030, Yang Liu 0005
EMNLP3
2023 Conditional Antibody Design as 3D Equivariant Graph Translation
Xiangzhe Kong, Wenbing Huang 0001, Yang Liu 0005
ICLR3
2023 Unified Detoxifying and Debiasing in Language Generation via Inference-time Adaptive Optimization
Zonghan Yang, Xiaoyuan Yi, Peng Li 0030, Yang Liu 0005, Xing Xie 0001
ICLR4
2023 End-to-End Full-Atom Antibody Design
abstract
Antibody design is an essential yet challenging task in various domains like therapeutics and biology. There are two major defects in current learning-based methods: 1) tackling only a certain subtask of the whole antibody design pipeline, making them suboptimal or resource-intensive. 2) omitting either the framework regions or side chains, thus incapable of capturing the full-atom geometry. To address these pitfalls, we propose dynamic Multi-channel Equivariant grAph Network (dyMEAN), an end-to-end full-atom model for E(3)-equivariant antibody design given the epitope and the incomplete sequence of the antibody. Specifically, we first explore structural initialization as a knowledgeable guess of the antibody structure and then propose shadow paratope to bridge the epitope-antibody connections. Both 1D sequences and 3D structures are updated via an adaptive multi-channel equivariant encoder that is able to process protein residues of variable sizes when considering full atoms. Finally, the updated antibody is docked to the epitope via the alignment of the shadow paratope. Experiments on epitope-binding CDR-H3 design, complex structure prediction, and affinity optimization demonstrate the superiority of our end-to-end framework and full-atom modeling.
Xiangzhe Kong, Wenbing Huang 0001, Yang Liu 0005
ICML3
2023 Improving Adversarial Robustness of Deep Equilibrium Models with Explicit Regulations Along the Neural Dynamics
abstract
Deep equilibrium (DEQ) models replace the multiple-layer stacking of conventional deep networks with a fixed-point iteration of a single-layer transformation. Having been demonstrated to be competitive in a variety of real-world scenarios, the adversarial robustness of general DEQs becomes increasingly crucial for their reliable deployment. Existing works improve the robustness of general DEQ models with the widely-used adversarial training (AT) framework, but they fail to exploit the structural uniquenesses of DEQ models. To this end, we interpret DEQs through the lens of neural dynamics and find that AT under-regulates intermediate states. Besides, the intermediate states typically provide predictions with a high prediction entropy. Informed by the correlation between the entropy of dynamical systems and their stability properties, we propose reducing prediction entropy by progressively updating inputs along the neural dynamics. During AT, we also utilize random intermediate states to compute the loss function. Our methods regulate the neural dynamics of DEQ models in this manner. Extensive experiments demonstrate that our methods substantially increase the robustness of DEQ models and even outperform the strong deep network baselines.
Zonghan Yang, Peng Li 0030, Tianyu Pang, Yang Liu 0005
ICML4
2023 Learning to Relate to Previous Turns in Conversational Search
abstract
Conversational search allows a user to interact with a search system in multiple turns. A query is strongly dependent on the conversation context. An effective way to improve retrieval effectiveness is to expand the current query with historical queries. However, not all the previous queries are related to, and useful for expanding the current query. In this paper, we propose a new method to select relevant historical queries that are useful for the current query. To cope with the lack of labeled training data, we use a pseudo-labeling approach to annotate useful historical queries based on their impact on the retrieval results. The pseudo-labeled data are used to train a selection model. We further propose a multi-task learning framework to jointly train the selector and the retriever during fine-tuning, allowing us to mitigate the possible inconsistency between the pseudo labels and the changed retriever. Extensive experiments on four conversational search datasets demonstrate the effectiveness and broad applicability of our method compared with several strong baselines.
Fengran Mo, Jian-Yun Nie, Kelong Mao, Yutao Zhu 0001, Peng Li 0030, Yang Liu 0005
KDD7
2023 Crystal Structure Prediction by Joint Equivariant Diffusion
abstract
Crystal Structure Prediction (CSP) is crucial in various scientific disciplines. While CSP can be addressed by employing currently-prevailing generative models (**e.g.** diffusion models), this task encounters unique challenges owing to the symmetric geometry of crystal structures---the invariance of translation, rotation, and periodicity. To incorporate the above symmetries, this paper proposes DiffCSP, a novel diffusion model to learn the structure distribution from stable crystals. To be specific, DiffCSP jointly generates the lattice and atom coordinates for each crystal by employing a periodic-E(3)-equivariant denoising model, to better model the crystal geometry. Notably, different from related equivariant generative approaches, DiffCSP leverages fractional coordinates other than Cartesian coordinates to represent crystals, remarkably promoting the diffusion and the generation process of atom positions. Extensive experiments verify that our DiffCSP remarkably outperforms existing CSP methods, with a much lower computation cost in contrast to DFT-based methods. Moreover, the superiority of DiffCSP is still observed when it is extended for ab initio crystal generation.
Wenbing Huang 0001, Peijia Lin, Jiaqi Han 0001, Pin Chen, Yutong Lu, Yang Liu 0005
NeurIPS7
2023 End-to-end hard constrained text generation via incrementally predicting segments
Jinran Nie, Xuancheng Huang, Yang Liu 0005, Cunliang Kong, Liner Yang, Erhong Yang
Knowl. Based Syst.3
2022 A Label Dependence-Aware Sequence Generation Model for Multi-Level Implicit Discourse Relation Recognition
abstract
Implicit discourse relation recognition (IDRR) is a challenging but crucial task in discourse analysis. Most existing methods train multiple models to predict multi-level labels independently, while ignoring the dependence between hierarchically structured labels. In this paper, we consider multi-level IDRR as a conditional label sequence generation task and propose a Label Dependence-aware Sequence Generation Model (LDSGM) for it. Specifically, we first design a label attentive encoder to learn the global representation of an input instance and its level-specific contexts, where the label dependence is integrated to obtain better label embeddings. Then, we employ a label sequence decoder to output the predicted labels in a top-down manner, where the predicted higher-level labels are directly used to guide the label prediction at the current level. We further develop a mutual learning enhanced training method to exploit the label dependence in a bottom-up direction, which is captured by an auxiliary decoder introduced during training. Experimental results on the PDTB dataset show that our model achieves the state-of-the-art performance on multi-level IDRR. We release our code at https://github.com/nlpersECJTU/LDSGM.
Changxing Wu, Liuwen Cao, Yubin Ge, Yang Liu 0005, Min Zhang 0005, Jinsong Su
AAAI4
2022 MSP: Multi-Stage Prompting for Making Pre-trained Language Models Better Translators
abstract
Prompting has recently been shown as a promising approach for applying pre-trained language models to perform downstream tasks.We present Multi-Stage Prompting, a simple and automatic approach for leveraging pre-trained language models to translation tasks.To better mitigate the discrepancy between pre-training and translation, MSP divides the translation process via pre-trained language models into multiple separate stages: the encoding stage, the re-encoding stage, and the decoding stage.During each stage, we independently apply different continuous prompts for allowing pretrained language models better shift to translation tasks.We conduct extensive experiments on three translation tasks.Experiments show that our method can significantly improve the translation performance of pre-trained language models.
Zhixing Tan, Xiangwen Zhang, Shuo Wang 0013, Yang Liu 0005
ACL (1)4
2022 Integrating Vectorized Lexical Constraints for Neural Machine Translation
abstract
Lexically constrained neural machine translation (NMT), which controls the generation of NMT models with pre-specified constraints, is important in many practical scenarios.Due to the representation gap between discrete constraints and continuous vectors in NMT models, most existing works choose to construct synthetic data or modify the decoding algorithm to impose lexical constraints, treating the NMT model as a black box.In this work, we propose to open this black box by directly integrating the constraints into NMT models.Specifically, we vectorize source and target constraints into continuous keys and values, which can be utilized by the attention modules of NMT models.The proposed integration method is based on the assumption that the correspondence between keys and values in attention modules is naturally suitable for modeling constraint pairs.Experimental results show that our method consistently outperforms several representative baselines on four language pairs, demonstrating the superiority of integrating vectorized lexical constraints.
Shuo Wang 0013, Zhixing Tan, Yang Liu 0005
ACL (1)3
2022 Global and Local Feature Interaction with Vision Transformer for Few-shot Image Classification
abstract
Image classification is a classical machine learning task and has been widely used. Due to the high costs of annotation and data collection in real scenarios, few-shot learning has become a vital technique to improve image classification performances. However, most existing few-shot image classification methods only focus on modeling the global image feature or image local patches, which ignore the global-local interactions. In this study, we propose a new method, named GL-ViT, to integrate both global and local features to fully exploit the few-shot samples for image classification. Firstly, we design a feature extractor module to calculate the interactions between the global representation and local patch embeddings, where ViT is also adopted to achieve efficient and effective image representation. Then, Earth Mover's Distance is adopted to measure the similarity between two images. Abundant Experimental results on several widely-used open datasets show that GL-ViT outperforms state-of-the-art algorithms significantly, and our ablation studies also verify the effectiveness of both global-local features.
Weizhi Ma, Yang Liu 0005
CIKM3
2022 End-to-End Unsupervised Vision-and-Language Pre-training with Referring Expression Matching
abstract
Recently there has been an emerging interest in unsupervised vision-and-language pre-training (VLP) that learns multimodal representations without parallel image-caption data.These pioneering works significantly reduce the cost of VLP on data collection and achieve promising results compared to supervised VLP.However, existing unsupervised VLP methods take as input pre-extracted region-based visual features from external object detectors, which both limits flexibility and reduces computational efficiency.In this paper, we explore end-to-end unsupervised VLP with a vision encoder to directly encode images.The vision encoder is pre-trained on image-only data and jointly optimized during multimodal pre-training.To further enhance the learned cross-modal features, we propose a novel pre-training task that predicts which patches contain an object referred to in natural language from the encoded visual features.Extensive experiments on four visionand-language tasks show that our approach outperforms previous unsupervised VLP methods and obtains new state-of-the-art results 1 .
Chi Chen 0005, Peng Li 0030, Maosong Sun 0001, Yang Liu 0005
EMNLP4
2022 Entropy-Based Vocabulary Substitution for Incremental Learning in Multilingual Neural Machine Translation
abstract
In a practical real-world scenario, the longstanding goal is that a universal multilingual translation model can be incrementally updated when new language pairs arrive.Specifically, the initial vocabulary only covers some of the words in new languages, which hurts the translation quality for incremental learning.Although existing approaches attempt to address this issue by replacing the original vocabulary with a rebuilt vocabulary or constructing independent language-specific vocabularies, these methods can not meet the following three demands simultaneously: (1) High translation quality for original and incremental languages, (2) low cost for model training, (3) low time overhead for preprocessing.In this work, we propose an entropy-based vocabulary substitution (EVS) method that just needs to walk through new language pairs for incremental learning in a large-scale multilingual data updating while remaining the size of the vocabulary.Our method has access to learn new knowledge from updated training samples incrementally while keeping high translation quality for original language pairs, alleviating the issue of catastrophic forgetting.Results of experiments show that EVS can achieve better performance and save excess overhead for incremental learning in the multilingual machine translation task.
Peng Li 0030, Jin Ma 0003, Yang Liu 0005
EMNLP4
2022 A Template-based Method for Constrained Neural Machine Translation
abstract
Machine translation systems are expected to cope with various types of constraints in many practical scenarios.While neural machine translation (NMT) has achieved strong performance in unconstrained cases, it is non-trivial to impose pre-specified constraints into the translation process of NMT models.Although many approaches have been proposed to address this issue, most existing methods can not satisfy the following three desiderata at the same time: (1) high translation quality, (2) high match accuracy, and (3) low latency.In this work, we propose a template-based method that can yield results with high translation quality and match accuracy and the inference speed of our method is comparable with unconstrained NMT models.Our basic idea is to rearrange the generation of constrained and unconstrained tokens through a template.Our method does not require any changes in the model architecture and the decoding algorithm.Experimental results show that the proposed template-based approach can outperform several representative baselines in both lexically and structurally constrained translation tasks.
Shuo Wang 0013, Peng Li 0030, Zhixing Tan, Zhaopeng Tu, Maosong Sun 0001, Yang Liu 0005
EMNLP6
2022 Towards a Unified Multi-Dimensional Evaluator for Text Generation
abstract
Multi-dimensional evaluation is the dominant paradigm for human evaluation in Natural Language Generation (NLG), i.e., evaluating the generated text from multiple explainable dimensions, such as coherence and fluency.However, automatic evaluation in NLG is still dominated by similarity-based metrics, and we lack a reliable framework for a more comprehensive evaluation of advanced models.In this paper, we propose a unified multi-dimensional evaluator UNIEVAL for NLG.We re-frame NLG evaluation as a Boolean Question Answering (QA) task, and by guiding the model with different questions, we can use one evaluator to evaluate from multiple dimensions.Furthermore, thanks to the unified Boolean QA format, we are able to introduce an intermediate learning phase that enables UNIEVAL to incorporate external knowledge from multiple related tasks and gain further improvement.Experiments on three typical NLG tasks show that UNIEVAL correlates substantially better with human judgments than existing metrics.Specifically, compared to the top-performing unified evaluators, UNIEVAL achieves a 23% higher correlation on text summarization, and over 43% on dialogue response generation.Also, UNIEVAL demonstrates a strong zero-shot learning ability for unseen evaluation dimensions and tasks.Source code, data and all pre-trained evaluators are available on our GitHub repository 1 . Generated Summary:Harry Kane is nominated for both the PFA player and young player of the season.The Spurs striker has been released from the awards ceremony on Sunday.The Tottenham striker features in a new animation.Reference Summary: Harry Kane has been in superb form for Tottenham this season.The 21-year-old has scored 30 goals in all competitions for Spurs.Kane also made his England debut and scored within two minutes.Document: Harry Kane's celebrations this season have always shown him to be an animated young man . . .Similarity-based Evaluators ROUGE-1: 0.44 ROUGE-2: 0.25 ROUGE-L: 0.42 BERTScore: 0.24 Single-dimensional Evaluators (predicted by two different evaluators (Deng et al., 2021)) Consistency: 0.87 Relevance: 0.74 Unified Evaluator (predicted by BARTScore, and the scoring range is negative infinity to 0) Precision: -5.45 Recall: -4.93 F1: -5.19
Ming Zhong 0005, Yang Liu 0005, Da Yin, Yuning Mao, Yizhu Jiao, Pengfei Liu 0003, Chenguang Zhu 0001, Heng Ji 0001, Jiawei Han 0001
EMNLP2
2022 A Class of Short-term Recurrence Anderson Mixing Methods and Their Applications
Fuchao Wei, Chenglong Bao, Yang Liu 0005
ICLR3
2022 Directed Acyclic Transformer for Non-Autoregressive Machine Translation
abstract
Non-autoregressive Transformers (NATs) significantly reduce the decoding latency by generating all tokens in parallel. However, such independent predictions prevent NATs from capturing the dependencies between the tokens for generating multiple possible translations. In this paper, we propose Directed Acyclic Transfomer (DA-Transformer), which represents the hidden states in a Directed Acyclic Graph (DAG), where each path of the DAG corresponds to a specific translation. The whole DAG simultaneously captures multiple translations and facilitates fast predictions in a non-autoregressive fashion. Experiments on the raw training data of WMT benchmark show that DA-Transformer substantially outperforms previous NATs by about 3 BLEU on average, which is the first NAT model that achieves competitive results with autoregressive Transformers without relying on knowledge distillation.
Fei Huang 0005, Hao Zhou 0012, Yang Liu 0005, Hang Li 0001, Minlie Huang
ICML3
2022 Molecule Generation by Principal Subgraph Mining and Assembling
abstract
Molecule generation is central to a variety of applications. Current attention has been paid to approaching the generation task as subgraph prediction and assembling. Nevertheless, these methods usually rely on hand-crafted or external subgraph construction, and the subgraph assembling depends solely on local arrangement. In this paper, we define a novel notion, principal subgraph that is closely related to the informative pattern within molecules. Interestingly, our proposed merge-and-update subgraph extraction method can automatically discover frequent principal subgraphs from the dataset, while previous methods are incapable of. Moreover, we develop a two-step subgraph assembling strategy, which first predicts a set of subgraphs in a sequence-wise manner and then assembles all generated subgraphs globally as the final output molecule. Built upon graph variational auto-encoder, our model is demonstrated to be effective in terms of several evaluation metrics and efficiency, compared with state-of-the-art methods on distribution learning and (constrained) property optimization tasks.
Xiangzhe Kong, Wenbing Huang 0001, Zhixing Tan, Yang Liu 0005
NeurIPS4
2022 A Closer Look at the Adversarial Robustness of Deep Equilibrium Models
abstract
Deep equilibrium models (DEQs) refrain from the traditional layer-stacking paradigm and turn to find the fixed point of a single layer. DEQs have achieved promising performance on different applications with featured memory efficiency. At the same time, the adversarial vulnerability of DEQs raises concerns. Several works propose to certify robustness for monotone DEQs. However, limited efforts are devoted to studying empirical robustness for general DEQs. To this end, we observe that an adversarially trained DEQ requires more forward steps to arrive at the equilibrium state, or even violates its fixed-point structure. Besides, the forward and backward tracks of DEQs are misaligned due to the black-box solvers. These facts cause gradient obfuscation when applying the ready-made attacks to evaluate or adversarially train DEQs. Given this, we develop approaches to estimate the intermediate gradients of DEQs and integrate them into the attacking pipelines. Our approaches facilitate fully white-box evaluations and lead to effective adversarial defense for DEQs. Extensive experiments on CIFAR-10 validate the adversarial robustness of DEQs competitive with deep networks of similar sizes.
Zonghan Yang, Tianyu Pang, Yang Liu 0005
NeurIPS3
2022 Data augmentation for low-resource languages NMT guided by constrained sampling
Mieradilijiang Maimaiti, Yang Liu 0005, Huan-Bo Luan, Maosong Sun 0001
Int. J. Intell. Syst.2
2022 Exploring Multi-Stage Information Interactions for Multi-Source Neural Machine Translation
abstract
Existing studies for multi-source neural machine translation (NMT) either separately model different source sentences or resort to the conventional single-source NMT by simply concatenating all source sentences. However, there exist two drawbacks in these approaches. First, they ignore the explicit word-level semantic interactions between source sentences, which have been shown effective in the embeddings of multilingual texts. Second, multiple source sentences are simultaneously encoded by an NMT model, which is unable to fully exploit the semantic information of each source sentence. In this paper, we explore multi-stage information interactions for multi-source NMT. Specifically, we first propose a multi-source NMT model that performs information interactions at the encoding stage. Its encoder contains multiple semantic interaction layers, each of which sequentially consists of (1) monolingual semantic interaction sub-layer, which is based on the self-attention mechanism and used to learn word-level monolingual contextual representations of source sentences, and (2) cross-lingual semantic interaction sub-layer, which leverages word alignments to perform fine-grained semantic transitions among hidden states of different source sentences. Furthermore, at the training stage, we introduce a mutual distillation based training framework, where single-source models and ours perform information interactions. Such framework can fully exploit the semantic information of each source sentence to enhance our model. Extensive experimental results on the WMT14 English-German-French dataset show our method exhibits significant improvements upon competitive baselines.
Ziyao Lu, Xiang Li 0104, Yang Liu 0005, Chulun Zhou, Jianwei Cui 0002, Bin Wang 0004, Min Zhang 0005, Jinsong Su
IEEE ACM Trans. Audio Speech Lang. Process.3
2022 Dynamic Multi-Branch Layers for On-Device Neural Machine Translation
abstract
With the rapid development of artificial intelligence (AI), there is a trend in moving AI applications, such as neural machine translation (NMT), from cloud to mobile devices. Constrained by limited hardware resources and battery, the performance of on-device NMT systems is far from satisfactory. Inspired by conditional computation, we propose to improve the performance of on-device NMT systems with dynamic multi-branch layers. Specifically, we design a layer-wise dynamic multi-branch network with only one branch activated during training and inference. As not all branches are activated during training, we propose shared-private reparameterization to ensure sufficient training for each branch. At almost the same computational cost, our method achieves improvements of up to 1.7 BLEU points on the WMT14 English-German translation task and 1.8 BLEU points on the WMT20 Chinese-English translation task over the Transformer model, respectively. Compared with a strong baseline that also uses multiple branches, the proposed method is up to 1.5 times faster with the same number of parameters.
Zhixing Tan, Zeyuan Yang 0002, Meng Zhang 0019, Qun Liu 0001, Maosong Sun 0001, Yang Liu 0005
IEEE ACM Trans. Audio Speech Lang. Process.6
2021 Mask-Align: Self-Supervised Neural Word Alignment
abstract
Chi Chen, Maosong Sun, Yang Liu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Chi Chen 0005, Maosong Sun 0001, Yang Liu 0005
ACL/IJCNLP (1)3
2021 Transfer Learning for Sequence Generation: from Single-source to Multi-source
abstract
Xuancheng Huang, Jingfang Xu, Maosong Sun, Yang Liu. Proceedings of the 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2021.
Xuancheng Huang, Jingfang Xu, Maosong Sun 0001, Yang Liu 0005
ACL/IJCNLP (1)4
2021 Segment, Mask, and Predict: Augmenting Chinese Word Segmentation with Self-Supervision
abstract
Mieradilijiang Maimaiti, Yang Liu, Yuanhang Zheng, Gang Chen, Kaiyu Huang, Ji Zhang, Huanbo Luan, Maosong Sun. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Mieradilijiang Maimaiti, Yang Liu 0005, Yuanhang Zheng, Gang Chen 0039, Ji Zhang 0011, Huan-Bo Luan, Maosong Sun 0001
EMNLP (1)2
2021 Self-Supervised Quality Estimation for Machine Translation
abstract
Yuanhang Zheng, Zhixing Tan, Meng Zhang, Mieradilijiang Maimaiti, Huanbo Luan, Maosong Sun, Qun Liu, Yang Liu. Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing. 2021.
Yuanhang Zheng, Zhixing Tan, Meng Zhang 0019, Mieradilijiang Maimaiti, Huan-Bo Luan, Maosong Sun 0001, Qun Liu 0001, Yang Liu 0005
EMNLP (1)8
2021 Stochastic Anderson Mixing for Nonconvex Stochastic Optimization
abstract
Anderson mixing (AM) is an acceleration method for fixed-point iterations. Despite its success and wide usage in scientific computing, the convergence theory of AM remains unclear, and its applications to machine learning problems are not well explored. In this paper, by introducing damped projection and adaptive regularization to the classical AM, we propose a Stochastic Anderson Mixing (SAM) scheme to solve nonconvex stochastic optimization problems. Under mild assumptions, we establish the convergence theory of SAM, including the almost sure convergence to stationary points and the worst-case iteration complexity. Moreover, the complexity bound can be improved when randomly choosing an iterate as the output. To further accelerate the convergence, we incorporate a variance reduction technique into the proposed SAM. We also propose a preconditioned mixing strategy for SAM which can empirically achieve faster convergence or better generalization ability. Finally, we apply the SAM method to train various neural networks including the vanilla CNN, ResNets, WideResNet, ResNeXt, DenseNet and LSTM. Experimental results on image classification and language model demonstrate the advantages of our method.
Fuchao Wei, Chenglong Bao, Yang Liu 0005
NeurIPS3
2021 Exploring Discriminative Word-Level Domain Contexts for Multi-Domain Neural Machine Translation
abstract
Owing to its practical significance, multi-domain Neural Machine Translation (NMT) has attracted much attention recently. Recent studies mainly focus on constructing a unified NMT model with mixed-domain training corpora to switch translation between different domains. In these models, the words in the same sentence are not well distinguished, while intuitively, they are related to the sentence domain to varying degrees and thus should exert different effects on the multi-domain NMT model. In this article, we are committed to distinguishing and exploiting different word-level domain contexts for multi-domain NMT. For this purpose, we adopt multi-task learning to jointly model NMT and monolingual attention-based domain classification tasks, improving the NMT model in two ways: 1) One domain classifier and one adversarial domain classifier are introduced to conduct domain classifications of input sentences. During this process, two generated gating vectors are used to produce domain-specific and domain-shared annotations for decoder; 2) We equip decoder with an attentional domain classifier. Then, the derived attentional weights are utilized to refine the model training via word-level cost weighting, so that the impacts of target words can be discriminated by their relevance to sentence domain. Experimental results on several multi-domain translations demonstrate the effectiveness of our model.
Jinsong Su, Jiali Zeng, Huating Wen, Yongjing Yin, Yang Liu 0005
IEEE Trans. Pattern Anal. Mach. Intell.6
2021 Improving Data Augmentation for Low-Resource NMT Guided by POS-Tagging and Paraphrase Embedding
abstract
Data augmentation is an approach for several text generation tasks. Generally, in the machine translation paradigm, mainly in low-resource language scenarios, many data augmentation methods have been proposed. The most used approaches for generating pseudo data mainly lay in word omission, random sampling, or replacing some words in the text. However, previous methods barely guarantee the quality of augmented data. In this work, we try to build the data by using paraphrase embedding and POS-Tagging. Namely, we generate the fake monolingual corpus by replacing the main four POS-Tagging labels, such as noun, adjective, adverb, and verb, based on both the paraphrase table and their similarity. We select the bigger corpus size of the paraphrase table with word level and obtain the word embedding of each word in the table, then calculate the cosine similarity between these words and tagged words in the original sequence. In addition, we exploit the ranking algorithm to choose highly similar words to reduce semantic errors and leverage the POS-Tagging replacement to mitigate syntactic error to some extent. Experimental results show that our augmentation method consistently outperforms all previous SOTA methods on the low-resource language pairs in seven language pairs from four corpora by 1.16 to 2.39 BLEU points.
Mieradilijiang Maimaiti, Yang Liu 0005, Huan-Bo Luan, Zegao Pan, Maosong Sun 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2021 Learning to Generate Explainable Plots for Neural Story Generation
abstract
Story generation is an important natural language processing task that aims to generate coherent stories automatically. While the use of neural networks has proven effective in improving story generation, how to learn to generate an explainable high-level plot still remains a major challenge. In this article, we propose a latent variable model for neural story generation. The model treats an outline, which is a natural language sentence explainable to humans, as a latent variable to represent a high-level plot that bridges the input and output. We adopt an external summarization model to guide the latent variable model to learn how to generate outlines from training data. Experiments show that our approach achieves significant improvements over state-of-the-art methods in both automatic and human evaluations.
Gang Chen 0039, Yang Liu 0005, Huan-Bo Luan, Meng Zhang 0019, Qun Liu 0001, Maosong Sun 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2021 Domain Adaptive Meta-Learning for Dialogue State Tracking
abstract
Domain adaptation for low-resource dialogue state tracking (DST) is of significance due to the growing diversity of conversation scenarios. In this paper, we propose a novel domain adaptive model-agnostic meta-learning (DAMAML) framework. Under this framework, we equip the DST model with two domain adaptors and a unified parameter generator. The parameter generator takes a domain embedding as input to produce parameters of domain adaptors, which modulate domain-shared initial parameters to the subspace of each domain. In this way, we simultaneously model multiple individual meta-learners with each covering the distribution of one domain, allowing more efficient adaptation. Compared with the conventional MAML, this framework not only is able to seek domain-shared initial parameters that facilitate fast adaptation, but also has better capability to fit a diversified domain distribution. Experimental results and in-depth analysis demonstrate the effectiveness of the proposed framework.
Jiali Zeng, Yongjing Yin, Yang Liu 0005, Yubin Ge, Jinsong Su
IEEE ACM Trans. Audio Speech Lang. Process.3
2021 Neural Machine Translation With Explicit Phrase Alignment
abstract
While neural machine translation has achieved state-of-the-art translation performance, it is unable to capture the alignment between the input and output during the translation process. The lack of alignment in neural machine translation models leads to three problems: it is hard to (1) interpret the translation process, (2) impose lexical constraints, and (3) impose structural constraints. These problems not only increase the difficulty of designing new architectures for neural machine translation, but also limit its applications in practice. To alleviate these problems, we propose to introduce explicit phrase alignment into the translation process of arbitrary neural machine translation models. The key idea is to build a search space similar to that of phrase-based statistical machine translation for neural machine translation where phrase alignment is readily available. We design a new decoding algorithm that can easily impose lexical and structural constraints. Experiments show that our approach makes the translation process of neural machine translation more interpretable without sacrificing translation quality. In addition, our approach achieves significant improvements in lexically and structurally constrained translation tasks.
Huan-Bo Luan, Maosong Sun 0001, Feifei Zhai, Jingfang Xu, Yang Liu 0005
IEEE ACM Trans. Audio Speech Lang. Process.6
2020 Multi-Zone Unit for Recurrent Neural Networks
Fandong Meng, Jinchao Zhang 0001, Yang Liu 0005, Jie Zhou 0016
AAAI3
2020 On the Inference Calibration of Neural Machine Translation
abstract
Confidence calibration, which aims to make model predictions equal to the true correctness measures, is important for neural machine translation (NMT) because it is able to offer useful indicators of translation errors in the generated output.While prior studies have shown that NMT models trained with label smoothing are well-calibrated on the groundtruth training data, we find that miscalibration still remains a severe challenge for NMT during inference due to the discrepancy between training and inference.By carefully designing experiments on three language pairs, our work provides in-depth analyses of the correlation between calibration and translation performance as well as linguistic properties of miscalibration and reports a number of interesting findings that might help humans better analyze, understand and improve NMT models.Based on these observations, we further propose a new graduated label smoothing method that can improve both inference calibration and translation performance.1
Shuo Wang 0013, Zhaopeng Tu, Shuming Shi 0001, Yang Liu 0005
ACL4
2020 Accurate Word Alignment Induction from Neural Machine Translation
abstract
Despite its original goal to jointly learn to align and translate, prior researches suggest that Transformer captures poor word alignments through its attention mechanism.In this paper, we show that attention weights DO capture accurate word alignments and propose two novel word alignment induction methods SHIFT-ATT and SHIFT-AET.The main idea is to induce alignments at the step when the to-be-aligned target token is the decoder input rather than the decoder output as in previous work.SHIFT-ATT is an interpretation method that induces alignments from the attention weights of Transformer and does not require parameter update or architecture change.SHIFT-AET extracts alignments from an additional alignment module which is tightly integrated into Transformer and trained in isolation with supervision from symmetrized SHIFT-ATT alignments.Experiments on three publicly available datasets demonstrate that both methods perform better than their corresponding neural baselines and SHIFT-AET significantly outperforms GIZA++ by 1.4-4.8AER points. 1
Yun Chen 0007, Yang Liu 0005, Guanhua Chen 0001, Xin Jiang 0002, Qun Liu 0001
EMNLP (1)2
2020 Interpolation between Residual and Non-Residual Networks
abstract
Although ordinary differential equations (ODEs) provide insights for designing network architectures, its relationship with the non-residual convolutional neural networks (CNNs) is still unclear. In this paper, we present a novel ODE model by adding a damping term. It can be shown that the proposed model can recover both a ResNet and a CNN by adjusting an interpolation coefficient. Therefore, the damped ODE model provides a unified framework for the interpretation of residual and non-residual networks. The Lyapunov analysis reveals better stability of the proposed model, and thus yields robustness improvement of the learned networks. Experiments on a number of image classification benchmarks show that the proposed model substantially improves the accuracy of ResNet and ResNeXt over the perturbed inputs from both stochastic noise and adversarial attack methods. Moreover, the loss landscape analysis demonstrates the improved robustness of our method along the attack direction.
Zonghan Yang, Yang Liu 0005, Chenglong Bao, Zuoqiang Shi
ICML2
2020 Modeling Voting for System Combination in Machine Translation
abstract
System combination is an important technique for combining the hypotheses of different machine translation systems to improve translation performance. Although early statistical approaches to system combination have been proven effective in analyzing the consensus between hypotheses, they suffer from the error propagation problem due to the use of pipelines. While this problem has been alleviated by end-to-end training of multi-source sequence-to-sequence models recently, these neural models do not explicitly analyze the relations between hypotheses and fail to capture their agreement because the attention to a word in a hypothesis is calculated independently, ignoring the fact that the word might occur in multiple hypotheses. In this work, we propose an approach to modeling voting for system combination in machine translation. The basic idea is to enable words in hypotheses from different systems to vote on words that are representative and should get involved in the generation process. This can be done by quantifying the influence of each voter and its preference for each candidate. Our approach combines the advantages of statistical and neural methods since it can not only analyze the relations between hypotheses but also allow for end-to-end training. Experiments show that our approach is capable of better taking advantage of the consensus between hypotheses and achieves significant improvements over state-of-the-art baselines on Chinese-English and English-German machine translation tasks.
Xuancheng Huang, Zhixing Tan, Derek F. Wong, Huan-Bo Luan, Jingfang Xu, Maosong Sun 0001, Yang Liu 0005
IJCAI8
2020 Re-evaluation of Atomic Operations and Graph Coloring for Unstructured Finite Volume GPU Simulations
abstract
In general, race condition can be resolved by introducing synchronisations or breaking data dependencies. Atomic operations and graph coloring are the two typical approaches to avoid race condition. Graph coloring algorithms have been generally considered winning algorithms in the literature due to their lock free implementations. In this paper, we present the GPU-accelerated algorithms of the unstructured cell-centered finite volume Computational Fluid Dynamics (CFD) software framework named PHengLEI which was originally developed for aerodynamics applications with arbitrary hybrid meshes. Overall, the newly developed GPU framework demonstrate up to 4.8 speedup comparing with 18 MPI tasks run on the latest Intel CPU node. Furthermore, the enormous efforts have been invested to optimize data dependencies which could lead to race condition due to unstructured mesh indirect addressing and related reduction math operations. With careful comparison between our optimised graph coloring and atomic operations using a series of numerical tests with different mesh sizes, the results show that atomic operations are more efficient than our optimised graph coloring in all of the test cases on Nvidia Tesla GPU V100. Specifically, for the summation operation, using atomicAdd is twice as fast as graph coloring. For the maximum operation, a speedup of 1.5 to 2 is found for atomicMax vs. graph coloring.
Xu Sun 0001, Xiaohu Guo, Yunfei Du 0001, Yutong Lu, Yang Liu 0005
SBAC-PAD6
2020 Incorporating Sememes into Chinese Definition Modeling
abstract
Chinese definition modeling is a challenging task that generates a dictionary definition in Chinese for a given Chinese word. To accomplish this task, we built two novel datasets based on Chinese Concept Dictionary (CCD) and Chinese WordNet (CWN) respectively. Each dataset contains triples of a word, sememes, and a corresponding definition. We present two novel models to improve Chinese definition modeling: the Adaptive-Attention model (AAM) and the Self- and Adaptive-Attention Model (SAAM). AAM successfully incorporates sememes for generating the definition with an adaptive attention mechanism. It has the capability to decide which sememes to focus on and when to pay attention to sememes. SAAM further replaces recurrent connections in AAM with self-attention and relies entirely on the attention mechanism, reducing the path length between word, sememes and definition. Experiments on both datasets demonstrate that by incorporating sememes, our model can generate definitions with more concrete information. And the best model that we proposed outperforms the state-of-the-art method by a large margin on both datasets.
Liner Yang, Cunliang Kong, Yun Chen 0007, Yang Liu 0005, Qinan Fan, Erhong Yang
IEEE ACM Trans. Audio Speech Lang. Process.4
2019 Shared-Private Bilingual Word Embeddings for Neural Machine Translation
abstract
Word embedding is central to neural machine translation (NMT), which has attracted intensive research interest in recent years.In NMT, the source embedding plays the role of the entrance while the target embedding acts as the terminal.These layers occupy most of the model parameters for representation learning.Furthermore, they indirectly interface via a soft-attention mechanism, which makes them comparatively isolated.In this paper, we propose shared-private bilingual word embeddings, which give a closer relationship between the source and target embeddings, and which also reduce the number of model parameters.For similar source and target words, their embeddings tend to share a part of the features and they cooperatively learn these common representation units.Experiments on 5 language pairs belonging to 6 different language families and written in 5 different alphabets demonstrate that the proposed model provides a significant performance boost over the strong baselines with dramatically fewer model parameters.
Xuebo Liu 0002, Derek F. Wong, Yang Liu 0005, Lidia S. Chao, Tong Xiao 0001
ACL (1)3
2019 Reducing Word Omission Errors in Neural Machine Translation: A Contrastive Learning Approach
abstract
While neural machine translation (NMT) has achieved remarkable success, NMT systems are prone to make word omission errors.In this work, we propose a contrastive learning approach to reducing word omission errors in NMT.The basic idea is to enable the NMT model to assign a higher probability to a ground-truth translation and a lower probability to an erroneous translation, which is automatically constructed from the ground-truth translation by omitting words.We design different types of negative examples depending on the number of omitted words, word frequency, and part of speech.Experiments on Chinese-to-English, German-to-English, and Russian-to-English translation tasks show that our approach is effective in reducing word omission errors and achieves better translation performance than three baseline methods.
Zonghan Yang, Yong Cheng 0003, Yang Liu 0005, Maosong Sun 0001
ACL (1)3
2019 Learning to Copy for Automatic Post-Editing
abstract
Xuancheng Huang, Yang Liu, Huanbo Luan, Jingfang Xu, Maosong Sun. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Xuancheng Huang, Yang Liu 0005, Huan-Bo Luan, Jingfang Xu, Maosong Sun 0001
EMNLP/IJCNLP (1)2
2019 Improving Back-Translation with Uncertainty-based Confidence Estimation
abstract
Shuo Wang, Yang Liu, Chao Wang, Huanbo Luan, Maosong Sun. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Shuo Wang 0013, Yang Liu 0005, Chao Wang 0049, Huan-Bo Luan, Maosong Sun 0001
EMNLP/IJCNLP (1)2
2019 Iterative Dual Domain Adaptation for Neural Machine Translation
abstract
Jiali Zeng, Yang Liu, Jinsong Su, Yubing Ge, Yaojie Lu, Yongjing Yin, Jiebo Luo. Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). 2019.
Jiali Zeng, Yang Liu 0005, Jinsong Su, Yubin Ge, Yaojie Lu 0001, Yongjing Yin, Jiebo Luo 0001
EMNLP/IJCNLP (1)2
2019 Optimizing Data Placement on Hierarchical Storage Architecture via Machine Learning
Peng Cheng 0012, Yutong Lu, Yunfei Du 0001, Zhiguang Chen 0001, Yang Liu 0005
NPC5
2019 Exploiting reverse target-side contexts for neural machine translation via asynchronous bidirectional decoding
Jinsong Su, Xiangwen Zhang, Junfeng Yao, Yang Liu 0005
Artif. Intell.6
2019 Multi-Round Transfer Learning for Low-Resource NMT Using Multiple High-Resource Languages
abstract
Neural machine translation (NMT) has made remarkable progress in recent years, but the performance of NMT suffers from a data sparsity problem since large-scale parallel corpora are only readily available for high-resource languages (HRLs). In recent days, transfer learning (TL) has been used widely in low-resource languages (LRLs) machine translation, while TL is becoming one of the vital directions for addressing the data sparsity problem in low-resource NMT. As a solution, a transfer learning method in NMT is generally obtained via initializing the low-resource model (child) with the high-resource model (parent). However, leveraging the original TL to low-resource models is neither able to make full use of highly related multiple HRLs nor to receive different parameters from the same parents. In order to exploit multiple HRLs effectively, we present a language-independent and straightforward multi-round transfer learning (MRTL) approach to low-resource NMT. Besides, with the intention of reducing the differences between high-resource and low-resource languages at the character level, we introduce a unified transliteration method for various language families, which are both semantically and syntactically highly analogous with each other. Experiments on low-resource datasets show that our approaches are effective, significantly outperform the state-of-the-art methods, and yield improvements of up to 5.63 BLEU points.
Mieradilijiang Maimaiti, Yang Liu 0005, Huan-Bo Luan, Maosong Sun 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.2
2019 POS Tag-enhanced Coarse-to-fine Attention for Neural Machine Translation
abstract
Although neural machine translation (NMT) has certain capability to implicitly learn semantic information of sentences, we explore and show that Part-of-Speech (POS) tags can be explicitly incorporated into the attention mechanism of NMT effectively to yield further improvements. In this article, we propose an NMT model with tag-enhanced attention mechanism. In our model, NMT and POS tagging are jointly modeled via multi-task learning. Besides following common practice to enrich encoder annotations by introducing predicted source POS tags, we exploit predicted target POS tags to refine attention model in a coarse-to-fine manner. Specifically, we first implement a coarse attention operation solely on source annotations and target hidden state, where the produced context vector is applied to update target hidden state used for target POS tagging. Then, we perform a fine attention operation that extends the coarse one by further exploiting the predicted target POS tags. Finally, we facilitate word prediction by simultaneously utilizing the context vector from fine attention and the predicted target POS tags. Experimental results and further analyses on Chinese-English and Japanese-English translation tasks demonstrate the superiority of our proposed model over the conventional NMT models. We release our code at https://github.com/middlekisser/PEA-NMT.git.
Yongjing Yin, Jinsong Su, Huating Wen, Jiali Zeng, Yang Liu 0005, Yidong Chen 0001
ACM Trans. Asian Low Resour. Lang. Inf. Process.5
2019 Latent Attribute Based Hierarchical Decoder for Neural Machine Translation
abstract
Neural machine translation (NMT) has achieved state-of-the-art performance in many translation tasks. However, because the computational cost increases with the size of the search space for predicting the target words, the translation quality of NMT is constrained by the limited vocabulary. To alleviate this problem, we propose a novel dynamic hierarchical decoder for NMT to utilize all of the target words in the training and decoding process. In the proposed model, a target word is represented by two latent attribute vectors rather than a word vector. The model is trained to dynamically put together those words that share similar linguistic attributes. The prediction of a target word is, therefore, turned into the prediction of attribute vectors, where the $\mathrm{softmax}$ functions are performed at the attribute level. This greatly reduces the model size and the decoding time. Our experimental results demonstrate that the proposed model significantly outperforms the NMT baselines in both Chinese-English and English-German translation tasks.
Xuebo Liu 0002, Derek F. Wong, Lidia S. Chao, Yang Liu 0005
IEEE ACM Trans. Audio Speech Lang. Process.4
2018 Zero-Resource Neural Machine Translation with Multi-Agent Communication Game
abstract
While end-to-end neural machine translation (NMT) has achieved notable success in the past years in translating a handful of resource-rich language pairs, it still suffers from the data scarcity problem for low-resource language pairs and domains. To tackle this problem, we propose an interactive multimodal framework for zero-resource neural machine translation. Instead of being passively exposed to large amounts of parallel corpora, our learners (implemented as encoder-decoder architecture) engage in cooperative image description games, and thus develop their own image captioning or neural machine translation model from the need to communicate in order to succeed at the game. Experimental results on the IAPR-TC12 and Multi30K datasets show that the proposed learning mechanism significantly improves over the state-of-the-art methods.
Yun Chen 0007, Yang Liu 0005, Victor O. K. Li
AAAI2
2018 Asynchronous Bidirectional Decoding for Neural Machine Translation
abstract
The dominant neural machine translation (NMT) models apply unified attentional encoder-decoder neural networks for translation. Traditionally, the NMT decoders adopt recurrent neural networks (RNNs) to perform translation in a left-to-right manner, leaving the target-side contexts generated from right to left unexploited during translation. In this paper, we equip the conventional attentional encoder-decoder NMT framework with a backward decoder, in order to explore bidirectional decoding for NMT. Attending to the hidden state sequence produced by the encoder, our backward decoder first learns to generate the target-side hidden state sequence from right to left. Then, the forward decoder performs translation in the forward direction, while in each translation prediction timestep, it simultaneously applies two attention models to consider the source-side and reverse target-side hidden states, respectively. With this new architecture, our model is able to fully exploit source- and target-side contexts to improve translation quality altogether. Experimental results on NIST Chinese-English and WMT English-German translation tasks demonstrate that our model achieves substantial improvements over the conventional NMT by 3.14 and 1.38 BLEU points, respectively. The source code of this work can be obtained from https://github.com/DeepLearnXMU/ABDNMT.
Xiangwen Zhang, Jinsong Su, Yang Liu 0005, Rongrong Ji
AAAI4
2018 Towards Robust Neural Machine Translation
abstract
Small perturbations in the input can severely distort intermediate representations and thus impact translation quality of neural machine translation (NMT) models.In this paper, we propose to improve the robustness of NMT models with adversarial stability training.The basic idea is to make both the encoder and decoder in NMT models robust against input perturbations by enabling them to behave similarly for the original input and its perturbed counterpart.Experimental results on Chinese-English, English-German and English-French translation tasks show that our approaches can not only achieve significant improvements over strong NMT systems but also improve the robustness of NMT models.
Yong Cheng 0003, Zhaopeng Tu, Fandong Meng, Junjie Zhai, Yang Liu 0005
ACL (1)5
2018 Multi-Domain Neural Machine Translation with Word-Level Domain Context Discrimination
abstract
With great practical value, the study of Multidomain Neural Machine Translation (NMT) mainly focuses on using mixed-domain parallel sentences to construct a unified model that allows translation to switch between different domains.Intuitively, words in a sentence are related to its domain to varying degrees, so that they will exert disparate impacts on the multi-domain NMT modeling.Based on this intuition, in this paper, we devote to distinguishing and exploiting word-level domain contexts for multi-domain NMT.To this end, we jointly model NMT with monolingual attention-based domain classification tasks and improve NMT as follows: 1) Based on the sentence representations produced by a domain classifier and an adversarial domain classifier, we generate two gating vectors and use them to construct domain-specific and domain-shared annotations, for later translation predictions via different attention models; 2) We utilize the attention weights derived from target-side domain classifier to adjust the weights of target words in the training objective, enabling domain-related words to have greater impacts during model training.Experimental results on Chinese-English and English-French multi-domain translation tasks demonstrate the effectiveness of the proposed model.Source codes of this paper are available on Github https://github.com/DeepLearnXMU/WDCNMT.
Jiali Zeng, Jinsong Su, Huating Wen, Yang Liu 0005, Yongjing Yin, Jianqiang Zhao
EMNLP4
2018 Improving the Transformer Translation Model with Document-Level Context
abstract
Although the Transformer translation model (Vaswani et al., 2017) has achieved state-ofthe-art performance in a variety of translation tasks, how to use document-level context to deal with discourse phenomena problematic for Transformer still remains a challenge.In this work, we extend the Transformer model with a new context encoder to represent document-level context, which is then incorporated into the original encoder and decoder.As large-scale document-level parallel corpora are usually not available, we introduce a two-step training method to take full advantage of abundant sentence-level parallel corpora and limited document-level parallel corpora.Experiments on the NIST Chinese-English datasets and the IWSLT French-English datasets show that our approach improves over Transformer significantly. 1
Huan-Bo Luan, Maosong Sun 0001, Feifei Zhai, Jingfang Xu, Min Zhang 0005, Yang Liu 0005
EMNLP7
2018 Error Analysis of Uyghur Name Tagging: Language-specific Techniques and Remaining Challenges
Halidanmu Abudukelimu, Abudoukelimu Abulizi, Boliang Zhang, Xiaoman Pan, Di Lu 0003, Heng Ji 0001, Yang Liu 0005
LREC7
2018 Neural Network Methods for Natural Language Processing Yoav Goldberg (Bar Ilan University)Morgan & Claypool (Synthesis Lectures on Human Language Technologies, edited by Graeme Hirst, volume 37), 2017, xxii+287 pp; paperback, ISBN 9781627052986, $74.95; ebook, ISBN 9781627052955, $59.96; doi: 10.2200/S00762ED1V01Y201703HLT037
abstract
Deep learning has attracted dramatic attention in recent years, both in academia and industry. The popular term deep learning generally refers to neural network methods. Indeed, many core ideas and methods were born years ago in the era of “shallow” neural networks. However, recent development of computation resources and accumulation of data, and of course new algorithmic techniques, has enabled this branch of machine learning to dominate many areas of artificial intelligence, first for perception tasks like speech recognition and computer vision, and gradually for natural language processing (NLP) since around 2013.Natural language is an intricate object for computers to handle. Philosophical debates aside, the field of NLP has witnessed a paradigm shift from rule-based methods to statistical approaches, which have been dominant since the 1990s. Following this background, deep learning goes further down the statistical route, and gradually becomes the de facto technique of the mainstream statistical landscape.This book covers the two exciting topics of neural networks and natural language processing. More specifically, it focuses on how neural network methods are applied on natural language data. With this guideline, the structure of the book appears smoother from a neural network entry: It first lays the background of neural network methods, and then discusses the traits of natural language data, including challenges to address and sources of information that we can exploit, so that specialized neural network models introduced later are designed in ways that accommodate natural language data. On the other hand, some fundamentals in natural language processing are not covered in the book, for example, linguistic theories and backgrounds of the natural language processing tasks, and proper preparation of corpus data. Based on this structure, the book is intended for practitioners from both deep learning and natural language processing to have a common ground and a shared understanding of what has been achieved at the intersection of these two fields. NLP practitioners can become well armed with the neural network tools to work on their natural language data, whereas neural network practitioners may feel that the content of the book is a bit light, although sufficient and effective enough for an entry into working with natural language data.After the first, introductory chapter, the book is divided into four parts that roughly follow the structure of the book mentioned above.This part introduces the basic machinery of neural networks, and contains four chapters. Chapter 2 provides the background of supervised machine learning, including concepts like parameterized functions, train, test, and validation sets, training as optimization, and, in particular, the use of gradient-based methods for optimization. Readers familiar with machine learning may safely skip this chapter. The models presented in this chapter are linear and log-linear models. Their limitations are discussed in Chapter 3, which motivates the need for nonlinear models, and sets the backdrop for the introduction of feed-forward neural networks presented in Chapter 4. Finally, Chapter 5 discusses the training of neural networks. Unlike most presentations from other sources, this chapter comes with a more algorithmic than mathematical flavor, by presenting computation graph abstraction as well as related software. It also provides a handy subsection that discusses practical choices for training neural networks.As mentioned earlier, neural network practitioners may feel that the neural network content of the book is a bit light, and this part can be almost entirely skipped by these readers. However, for people coming from more traditional branches of statistical learning, Chapter 5 is still well worth reading.This part discusses the traits of natural language data, the object to which we would like to apply neural networks. There are seven chapters in this part. Chapter 6 presents a categorization of natural language classification problems and discusses the information sources that we can exploit in natural language data. Chapter 7 provides concrete examples of natural language features for solving various NLP tasks. These two chapters are probably quite dense for people coming from machine learning, and they serve to prepare them with the familiarity needed to work with natural language data. Chapter 8 is where neural networks come in; the chapter discusses how to represent textual features as inputs for neural network models. Chapter 9 describes the language modeling task and discusses the feed-forward neural language model. The neural language model also produces the byproduct of word representations, which form the subject of Chapters 10 and 11. In particular, Chapter 10 presents approaches to learning word representations, and Chapter 11 discusses the usage of word representations outside the context of neural networks, like word similarity and word analogies. Chapter 12 is an independent chapter that describes a specific feed-forward neural network architecture for the task of natural language inference.This part of the book, especially Chapter 8, which connects neural networks with natural language data, is the core of the content that distinguishes this book from other materials that cover either neural networks or natural language processing.This part is composed of five chapters that introduce the specialized architectures of convolutional neural networks (CNNs) (Chapter 13) and recurrent neural networks (RNNs) (Chapters 14–17). Chapter 13 mainly introduces 1D CNNs, which are specialized at learning ngram patterns. Chapter 14 describes the modeling of sequences and stacks with recurrent neural networks in an abstract way. This is to be made concrete by the succeeding two chapters. In Chapter 15, concrete instantiations of RNNs like the Long Short-Term Memory (LSTM) and the Gated Recurrent Unit (GRU) are described, and in Chapter 16, concrete applications of modeling with the RNN abstraction to NLP tasks are presented, including sentiment classification, grammaticality detection, part-of-speech tagging, document classification, and dependency parsing. Chapter 17 also includes concrete applications of RNNs, but these tasks involve generating natural language, which are usually modeled with a conditioned RNN language model. The most typical example of these tasks is probably machine translation.As the distribution of the chapters suggests, recurrent neural networks clearly receive more emphases. Indeed, RNNs alleviate the reliance on the Markov assumption and have the potential to model very long sequences. Their capabilities have led to breakthroughs in various sequence processing tasks, making them the celebrated models in research frontiers with proven performance.This part contains four chapters that are relatively independent. Chapter 18 presents recursive neural networks for modeling trees. The capability of modeling trees is important for natural language because of its hierarchical structure. Chapter 19 is devoted to structured prediction, because certain NLP tasks like named entity recognition can be cast in this framework. Chapter 20 discusses multi-task learning and semi-supervised learning. These approaches have not yet grown into a full-fledged stage, but are still important topics for research and offer helpful techniques for many tasks.The final chapter, Chapter 21, briefly reviews the content presented in the book, and discusses challenges that are yet to be addressed.The application of neural networks to natural language processing has revolutionized this long-standing research field, pushing forward the state of the art of many tasks. Nonetheless, the goal of equipping computers with human language capability is still far from solved, and the field continues to develop at a fast pace. This book provides valuable materials for newcomers into this exciting arena of cross-disciplinary research, by preparing relevant information of both neural networks and natural language processing. The book mainly presents mature neural network approaches to natural language processing, because it is hardly possible for a book to keep up to date with such fast development—although at 287 pages, the book is already quite long compared with other books in the synthesis lectures series, which are usually monographs of 50 to 150 pages.
Yoav Goldberg, Graeme Hirst, Yang Liu 0005, Meng Zhang 0019
Comput. Linguistics3
2018 Alignment-consistent recursive neural networks for bilingual phrase embeddings
Jinsong Su, Biao Zhang 0002, Deyi Xiong, Yang Liu 0005, Min Zhang 0005
Knowl. Based Syst.4
2018 Learning to Remember Translation History with a Continuous Cache
abstract
Existing neural machine translation (NMT) models generally translate sentences in isolation, missing the opportunity to take advantage of document-level information. In this work, we propose to augment NMT models with a very light-weight cache-like memory network, which stores recent hidden representations as translation history. The probability distribution over generated words is updated online depending on the translation history retrieved from the memory, endowing NMT models with the capability to dynamically adapt over time. Experiments on multiple domains with different topics and styles show the effectiveness of the proposed approach with negligible impact on the computational cost.
Zhaopeng Tu, Yang Liu 0005, Shuming Shi 0001, Tong Zhang 0001
Trans. Assoc. Comput. Linguistics2
2018 A Hierarchy-to-Sequence Attentional Neural Machine Translation Model
abstract
Although sequence-to-sequence attentional neural machine translation (NMT) has achieved great progress recently, it is confronted with two challenges: learning optimal model parameters for long parallel sentences and well exploiting different scopes of contexts. In this paper, partially inspired by the idea of segmenting a long sentence into short clauses, each of which can be easily translated by NMT, we propose a hierarchy-to-sequence attentional NMT model to handle these two challenges. Our encoder takes the segmented clause sequence as input and explores a hierarchical neural network structure to model words, clauses, and sentences at different levels, particularly with two layers of recurrent neural networks modeling semantic compositionality at the word and clause level. Correspondingly, the decoder sequentially translates segmented clauses and simultaneously applies two types of attention models to capture contexts of interclause and intraclause for translation prediction. In this way, we can not only improve parameter learning, but also well explore different scopes of contexts for translation. Experimental results on Chinese-English and English-German translation demonstrate the superiorities of the proposed model over the conventional NMT model.
Jinsong Su, Jiali Zeng, Deyi Xiong, Yang Liu 0005, Mingxuan Wang
IEEE ACM Trans. Audio Speech Lang. Process.4
2018 Joint POS Tagging and Dependence Parsing With Transition-Based Neural Networks
abstract
While part-of-speech (POS) tagging and dependency parsing are observed to be closely related, existing work on joint modeling with manually crafted feature templates suffers from the feature sparsity and incompleteness problems. In this paper, we propose an approach to joint POS tagging and dependency parsing using transition-based neural networks. Three neural network based classifiers are designed to resolve shift/reduce, tagging, and labeling conflicts. Experiments show that our approach significantly outperforms previous methods for joint POS tagging and dependency parsing across a variety of natural languages.
Liner Yang, Meishan Zhang, Yang Liu 0005, Maosong Sun 0001, Guohong Fu
IEEE ACM Trans. Audio Speech Lang. Process.3
2017 Maximum Reconstruction Estimation for Generative Latent-Variable Models
Yong Cheng 0003, Yang Liu 0005, Wei Xu 0005
AAAI2
2017 Lattice-Based Recurrent Neural Network Encoders for Neural Machine Translation
abstract
Neural machine translation (NMT) heavily relies on word-level modelling to learn semantic representations of input sentences.However, for languages without natural word delimiters (e.g., Chinese) where input sentences have to be tokenized first,conventional NMT is confronted with two issues:1) it is difficult to find an optimal tokenization granularity for source sentence modelling, and2) errors in 1-best tokenizations may propagate to the encoder of NMT.To handle these issues, we propose word-lattice based Recurrent Neural Network (RNN) encoders for NMT,which generalize the standard RNN to word lattice topology.The proposed encoders take as input a word lattice that compactly encodes multiple tokenizations, and learn to generate new hidden states from arbitrarily many inputs and hidden states in preceding time steps.As such, the word-lattice based encoders not only alleviate the negative impact of tokenization errors but also are more expressive and flexible to embed input sentences.Experiment results on Chinese-English translation demonstrate the superiorities of the proposed encoders over the conventional encoder.
Jinsong Su, Zhixing Tan, Deyi Xiong, Rongrong Ji, Xiaodong Shi, Yang Liu 0005
AAAI6
2017 Neural Machine Translation with Reconstruction
abstract
Although end-to-end Neural Machine Translation (NMT) has achieved remarkable progress in the past two years, it suffers from a major drawback: translations generated by NMT systems often lack of adequacy. It has been widely observed that NMT tends to repeatedly translate some source words while mistakenly ignoring other words. To alleviate this problem, we propose a novel encoder-decoder-reconstructor framework for NMT. The reconstructor, incorporated into the NMT model, manages to reconstruct the input source sentence from the hidden layer of the output target sentence, to ensure that the information in the source side is transformed to the target side as much as possible. Experiments show that the proposed framework significantly improves the adequacy of NMT output and achieves superior translation result over state-of-the-art NMT and statistical MT systems.
Zhaopeng Tu, Yang Liu 0005, Lifeng Shang, Hang Li 0001
AAAI2
2017 Bilingual Lexicon Induction from Non-Parallel Data with Minimal Supervision
abstract
Building bilingual lexica from non-parallel data is a long-standing natural language processing research problem that could benefit thousands of resource-scarce languages which lack parallel data. Recent advances of continuous word representations have opened up new possibilities for this task, e.g. by establishing cross-lingual mapping between word embeddings via a seed lexicon. The method is however unreliable when there are only a limited number of seeds, which is a reasonable setting for resource-scarce languages. We tackle the limitation by introducing a novel matching mechanism into bilingual word representation learning. It captures extra translation pairs exposed by the seeds to incrementally improve the bilingual word embeddings. In our experiments, we find the matching mechanism to substantially improve the quality of the bilingual vector space, which in turn allows us to induce better bilingual lexica with seeds as few as 10.
Meng Zhang 0019, Haoruo Peng, Yang Liu 0005, Huan-Bo Luan, Maosong Sun 0001
AAAI3
2017 A Teacher-Student Framework for Zero-Resource Neural Machine Translation
abstract
While end-to-end neural machine translation (NMT) has made remarkable progress recently, it still suffers from the data scarcity problem for low-resource language pairs and domains.In this paper, we propose a method for zero-resource NMT by assuming that parallel sentences have close probabilities of generating a sentence in a third language.Based on the assumption, our method is able to train a source-to-target NMT model ("student") without parallel corpora available guided by an existing pivot-to-target NMT model ("teacher") on a source-pivot parallel corpus.Experimental results show that the proposed method significantly improves over a baseline pivot-based model by +3.0 BLEU points across various language pairs.
Yun Chen 0007, Yang Liu 0005, Yong Cheng 0003, Victor O. K. Li
ACL (1)2
2017 Visualizing and Understanding Neural Machine Translation
abstract
While neural machine translation (NMT) has made remarkable progress in recent years, it is hard to interpret its internal workings due to the continuous representations and non-linearity of neural networks.In this work, we propose to use layer-wise relevance propagation (LRP) to compute the contribution of each contextual word to arbitrary hidden states in the attention-based encoderdecoder framework.We show that visualization with LRP helps to interpret the internal workings of NMT and analyze translation errors.
Yanzhuo Ding, Yang Liu 0005, Huan-Bo Luan, Maosong Sun 0001
ACL (1)2
2017 Adversarial Training for Unsupervised Bilingual Lexicon Induction
abstract
Word embeddings are well known to capture linguistic regularities of the language on which they are trained.Researchers also observe that these regularities can transfer across languages.However, previous endeavors to connect separate monolingual word embeddings typically require cross-lingual signals as supervision, either in the form of parallel corpus or seed lexicon.In this work, we show that such cross-lingual connection can actually be established without any form of supervision.We achieve this end by formulating the problem as a natural adversarial game, and investigating techniques that are crucial to successful training.We carry out evaluation on the unsupervised bilingual lexicon induction task.Even though this task appears intrinsically cross-lingual, we are able to demonstrate encouraging performance without any cross-lingual clues.
Meng Zhang 0019, Yang Liu 0005, Huan-Bo Luan, Maosong Sun 0001
ACL (1)2
2017 Prior Knowledge Integration for Neural Machine Translation using Posterior Regularization
abstract
Although neural machine translation has made significant progress recently, how to integrate multiple overlapping, arbitrary prior knowledge sources remains a challenge.In this work, we propose to use posterior regularization to provide a general framework for integrating prior knowledge into neural machine translation.We represent prior knowledge sources as features in a log-linear model, which guides the learning process of the neural translation model.Experiments on Chinese-English translation show that our approach leads to significant improvements.
Yang Liu 0005, Huan-Bo Luan, Jingfang Xu, Maosong Sun 0001
ACL (1)2
2017 Earth Mover's Distance Minimization for Unsupervised Bilingual Lexicon Induction
abstract
Cross-lingual natural language processing hinges on the premise that there exists invariance across languages. At the word level, researchers have identified such invariance in the word embedding semantic spaces of different languages. However, in order to connect the separate spaces, cross-lingual supervision encoded in parallel data is typically required. In this paper, we attempt to establish the cross-lingual connection without relying on any cross-lingual supervision. By viewing word embedding spaces as distributions, we propose to minimize their earth mover's distance, a measure of divergence between distributions. We demonstrate the success on the unsupervised bilingual lexicon induction task. In addition, we reveal an interesting finding that the earth mover's distance shows potential as a measure of language difference.
Meng Zhang 0019, Yang Liu 0005, Huan-Bo Luan, Maosong Sun 0001
EMNLP2
2017 Joint Training for Pivot-based Neural Machine Translation
abstract
While recent neural machine translation approaches have delivered state-of-the-art performance for resource-rich language pairs, they suffer from the data scarcity problem for resource-scarce language pairs. Although this problem can be alleviated by exploiting a pivot language to bridge the source and target languages, the source-to-pivot and pivot-to-target translation models are usually independently trained. In this work, we introduce a joint training algorithm for pivot-based neural machine translation. We propose three methods to connect the two models and enable them to interact with each other during training. Experiments on Europarl and WMT corpora show that joint training of source-to-pivot and pivot-to-target models leads to significant improvements over independent training across various languages.
Yong Cheng 0003, Qian Yang 0003, Yang Liu 0005, Maosong Sun 0001, Wei Xu 0005
IJCAI3
2017 Maximum Expected Likelihood Estimation for Zero-resource Neural Machine Translation
abstract
While neural machine translation (NMT) has made remarkable progress in translating a handful of high-resource language pairs recently, parallel corpora are not always available for many zero-resource language pairs. To deal with this problem, we propose an approach to zero-resource NMT via maximum expected likelihood estimation. The basic idea is to maximize the expectation with respect to a pivot-to-source translation model for the intended source-to-target model on a pivot-target parallel corpus. To approximate the expectation, we propose two methods to connect the pivot-to-source and source-to-target models. Experiments on two zero-resource language pairs show that the proposed approach yields substantial gains over baseline methods. We also observe that when trained jointly with the source-to-target model, the pivot-to-source translation model also obtains improvements over independent training.
Yong Cheng 0003, Yang Liu 0005
IJCAI3
2017 Exploiting Unlabeled Data for Neural Grammatical Error Detection
Zhuoran Liu 0010, Yang Liu 0005
J. Comput. Sci. Technol.2
2017 Optimizing Non-Decomposable Evaluation Metrics for Neural Machine Translation
Shiqi Shen, Yang Liu 0005, Maosong Sun 0001
J. Comput. Sci. Technol.2
2017 Neural Parse Combination
Liner Yang, Maosong Sun 0001, Yong Cheng 0003, Zhenghao Liu 0001, Huan-Bo Luan, Yang Liu 0005
J. Comput. Sci. Technol.7
2017 Context Gates for Neural Machine Translation
abstract
In neural machine translation (NMT), generation of a target word depends on both source and target contexts. We find that source contexts have a direct impact on the adequacy of a translation while target contexts affect the fluency. Intuitively, generation of a content word should rely more on the source context and generation of a functional word should rely more on the target context. Due to the lack of effective control over the influence from source and target contexts, conventional NMT tends to yield fluent but inadequate translations. To address this problem, we propose context gates which dynamically control the ratios at which source and target contexts contribute to the generation of target words. In this way, we can enhance both the adequacy and fluency of NMT with more careful control of the information flow from contexts. Experiments show that our approach significantly improves upon a standard attention-based NMT system by +2.3 BLEU points.
Zhaopeng Tu, Yang Liu 0005, Zhengdong Lu, Hang Li 0001
Trans. Assoc. Comput. Linguistics2
2016 To Swap or Not to Swap? Exploiting Dependency Word Pairs for Reordering in Statistical Machine Translation
abstract
Reordering poses a major challenge in machine translation (MT) between two languages with significant differences in word order. In this paper, we present a novel reordering approach utilizing sparse features based on dependency word pairs. Each instance of these features captures whether two words, which are related by a dependency link in the source sentence dependency parse tree, follow the same order or are swapped in the translation output. Experiments on Chinese-to-English translation show a statistically significant improvement of 1.21 BLEU point using our approach, compared to a state-of-the-art statistical MT system that incorporates prior reordering approaches.
Christian Hadiwinoto, Yang Liu 0005, Hwee Tou Ng
AAAI2
2016 Building Earth Mover's Distance on Bilingual Word Embeddings for Machine Translation
abstract
Following their monolingual counterparts, bilingual word embeddings are also on the rise. As a major application task, word translation has been relying on the nearest neighbor to connect embeddings cross-lingually. However, the nearest neighbor strategy suffers from its inherently local nature and fails to cope with variations in realistic bilingual word embeddings. Furthermore, it lacks a mechanism to deal with many-to-many mappings that often show up across languages. We introduce Earth Mover's Distance to this task by providing a natural formulation that translates words in a holistic fashion, addressing the limitations of the nearest neighbor. We further extend the formulation to a new task of identifying parallel sentences, which is useful for statistical machine translation systems, thereby expanding the application realm of bilingual word embeddings. We show encouraging performance on both tasks.
Meng Zhang 0019, Yang Liu 0005, Huan-Bo Luan, Maosong Sun 0001, Tatsuya Izuha
AAAI2
2016 Semi-Supervised Learning for Neural Machine Translation
abstract
While end-to-end neural machine translation (NMT) has made remarkable progress recently, NMT systems only rely on parallel corpora for parameter estimation. Since parallel corpora are usually limited in quantity, quality, and coverage, especially for low-resource languages, it is appealing to exploit monolingual corpora to improve NMT. We propose a semi-supervised approach for training NMT models on the concatenation of labeled (parallel corpora) and unlabeled (monolingual corpora) data. The central idea is to reconstruct the monolingual corpora using an autoencoder, in which the source-to-target and target-to-source translation models serve as the encoder and decoder, respectively. Our approach can not only exploit the monolingual corpora of the target language, but also of the source language. Experiments on the Chinese-English dataset show that our approach achieves significant improvements over state-of-the-art SMT and NMT systems.
Yong Cheng 0003, Wei Xu 0005, Zhongjun He, Wei He 0014, Hua Wu 0003, Maosong Sun 0001, Yang Liu 0005
ACL (1)7
2016 Agreement-based Learning of Parallel Lexicons and Phrases from Non-Parallel Corpora
abstract
We introduce an agreement-based approach to learning parallel lexicons and phrases from non-parallel corpora.The basic idea is to encourage two asymmetric latent-variable translation models (i.e., source-to-target and target-to-source) to agree on identifying latent phrase and word alignments.The agreement is defined at both word and phrase levels.We develop a Viterbi EM algorithm for jointly training the two unidirectional models efficiently.Experiments on the Chinese-English dataset show that agreementbased learning significantly improves both alignment and translation performance.
Yang Liu 0005, Maosong Sun 0001, Huan-Bo Luan, Heng Yu 0006
ACL (1)2
2016 Minimum Risk Training for Neural Machine Translation
abstract
We propose minimum risk training for end-to-end neural machine translation.Unlike conventional maximum likelihood estimation, minimum risk training is capable of optimizing model parameters directly with respect to arbitrary evaluation metrics, which are not necessarily differentiable.Experiments show that our approach achieves significant improvements over maximum likelihood estimation on a state-of-the-art neural machine translation system across various languages pairs.Transparent to architectures, our approach can be applied to more neural networks and potentially benefit more NLP tasks.
Shiqi Shen, Yong Cheng 0003, Zhongjun He, Wei He 0014, Hua Wu 0003, Maosong Sun 0001, Yang Liu 0005
ACL (1)7
2016 Modeling Coverage for Neural Machine Translation
abstract
Attention mechanism has enhanced stateof-the-art Neural Machine Translation (NMT) by jointly learning to align and translate.It tends to ignore past alignment information, however, which often leads to over-translation and under-translation.To address this problem, we propose coverage-based NMT in this paper.We maintain a coverage vector to keep track of the attention history.The coverage vector is fed to the attention model to help adjust future attention, which lets NMT system to consider more about untranslated source words.Experiments show that the proposed approach significantly improves both translation quality and alignment quality over standard attention-based NMT. 1
Zhaopeng Tu, Zhengdong Lu, Yang Liu 0005, Hang Li 0001
ACL (1)3
2016 Inducing Bilingual Lexica From Non-Parallel Data With Earth Mover's Distance Regularization
abstract
Being able to induce word translations from non-parallel data is often a prerequisite for cross-lingual processing in resource-scarce languages and domains. Previous endeavors typically simplify this task by imposing the one-to-one translation assumption, which is too strong to hold for natural languages. We remove this constraint by introducing the Earth Mover’s Distance into the training of bilingual word embeddings. In this way, we take advantage of its capability to handle multiple alternative word translations in a natural form of regularization. Our approach shows significant and consistent improvements across four language pairs. We also demonstrate that our approach is particularly preferable in resource-scarce settings as it only requires a minimal seed lexicon.
Meng Zhang 0019, Yang Liu 0005, Huan-Bo Luan, Yiqun Liu 0001, Maosong Sun 0001
COLING2
2016 Agreement-Based Joint Training for Bidirectional Attention-Based Neural Machine Translation
Yong Cheng 0003, Shiqi Shen, Zhongjun He, Wei He 0014, Hua Wu 0003, Maosong Sun 0001, Yang Liu 0005
IJCAI7
2016 Listwise Ranking Functions for Statistical Machine Translation
abstract
Decision rules play an important role in the tuning and decoding steps of statistical machine translation. The traditional decision rule selects the candidate with the greatest potential from a candidate space by examining each candidate individually. However, viewing each candidate as independent imposes a serious limitation on the translation task. We instead view the problem from a ranking perspective that naturally allows the consideration of an entire list of candidates as a whole through the adoption of a listwise ranking function. Our shift from a pointwise to a listwise perspective proves to be a simple yet powerful extension to current modeling that allows arbitrary pairwise functions to be incorporated as features, whose weights can be estimated jointly with traditional ones. We further demonstrate that our formulation encompasses the minimum Bayes risk (MBR) approach, another decision rule that considers restricted listwise information, as a special case. Experiments show that our approach consistently outperforms the baseline and MBR methods across the considered test sets.
Meng Zhang 0019, Yang Liu 0005, Huan-Bo Luan, Maosong Sun 0001
IEEE ACM Trans. Audio Speech Lang. Process.2
2015 Topical Word Embeddings
abstract
Most word embedding models typically represent each word using a single vector, which makes these models indiscriminative for ubiquitous homonymy and polysemy. In order to enhance discriminativeness, we employ latent topic models to assign topics for each word in the text corpus, and learn topical word embeddings (TWE) based on both words and their topics. In this way, contextual word embeddings can be flexibly obtained to measure contextual word similarity. We can also build document representations, which are more expressive than some widely-used document models such as latent topic models. In the experiments, we evaluate the TWE models on two tasks, contextual word similarity and text classification. The experimental results show that our models outperform typical word embedding models including the multi-prototype version on contextual word similarity, and also exceed latent topic models and other representative document models on text classification.
Yang Liu 0005, Zhiyuan Liu 0001, Tat-Seng Chua, Maosong Sun 0001
AAAI1
2015 Contrastive Unsupervised Word Alignment with Non-Local Features
abstract
Word alignment is an important natural language processing task that indicates the correspondence between natural languages. Recently, unsupervised learning of log-linear models for word alignment has received considerable attention as it combines the merits of generative and discriminative approaches. However, a major challenge still remains: it is intractable to calculate the expectations of non-local features that are critical for capturing the divergence between natural languages. We propose a contrastive approach that aims to differentiate observed training examples from noises. It not only introduces prior knowledge to guide unsupervised learning but also cancels out partition functions. Based on the observation that the probability mass of log-linear models for word alignment is usually highly concentrated, we propose to use top-$n$ alignments to approximate the expectations with respect to posterior distributions. This allows for efficient and accurate calculation of expectations of non-local features. Experiments show that our approach achieves significant improvements over state-of-the-art unsupervised word alignment methods.
Yang Liu 0005, Maosong Sun 0001
AAAI1
2015 A Context-Aware Topic Model for Statistical Machine Translation
abstract
Jinsong Su, Deyi Xiong, Yang Liu, Xianpei Han, Hongyu Lin, Junfeng Yao, Min Zhang. Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers). 2015.
Jinsong Su, Deyi Xiong, Yang Liu 0005, Xianpei Han, Junfeng Yao, Min Zhang 0005
ACL (1)3
2015 Generalized Agreement for Bidirectional Word Alignment
abstract
While agreement-based joint training has proven to deliver state-of-the-art alignment accuracy, the produced word alignments are usually restricted to one-toone mappings because of the hard constraint on agreement.We propose a general framework to allow for arbitrary loss functions that measure the disagreement between asymmetric alignments.The loss functions can not only be defined between asymmetric alignments but also between alignments and other latent structures such as phrase segmentations.We use a Viterbi EM algorithm to train the joint model since the inference is intractable.Experiments on Chinese-English translation show that joint training with generalized agreement achieves significant improvements over two state-ofthe-art alignment methods.
Yang Liu 0005, Maosong Sun 0001, Huan-Bo Luan, Heng Yu 0006
EMNLP2
2015 Consistency-Aware Search for Word Alignment
abstract
As conventional word alignment search algorithms usually ignore the consistency constraint in translation rule extraction, improving alignment accuracy does not necessarily increase translation quality.We propose to use coverage, which reflects how well extracted phrases can recover the training data, to enable word alignment to model consistency and correlate better with machine translation.This can be done by introducing an objective that maximizes both alignment model score and coverage.We introduce an efficient algorithm to calculate coverage on the fly during search.Experiments show that our consistency-aware search algorithm significantly outperforms both generative and discriminative alignment approaches across various languages and translation models.
Shiqi Shen, Yang Liu 0005, Maosong Sun 0001, Huan-Bo Luan
EMNLP2
2015 Bilingual Correspondence Recursive Autoencoder for Statistical Machine Translation
abstract
Learning semantic representations and tree structures of bilingual phrases is beneficial for statistical machine translation.In this paper, we propose a new neural network model called Bilingual Correspondence Recursive Autoencoder (BCor-rRAE) to model bilingual phrases in translation.We incorporate word alignments into BCorrRAE to allow it freely access bilingual constraints at different levels.BCorrRAE minimizes a joint objective on the combination of a recursive autoencoder reconstruction error, a structural alignment consistency error and a crosslingual reconstruction error so as to not only generate alignment-consistent phrase structures, but also capture different levels of semantic relations within bilingual phrases.In order to examine the effectiveness of BCorrRAE, we incorporate both semantic and structural similarity features built on bilingual phrase representations and tree structures learned by BCorrRAE into a state-of-the-art SMT system.Experiments on NIST Chinese-English test sets show that our model achieves a substantial improvement of up to 1.55 BLEU points over the baseline.
Jinsong Su, Deyi Xiong, Biao Zhang 0002, Yang Liu 0005, Junfeng Yao, Min Zhang 0005
EMNLP4
2015 Iterative Learning of Parallel Lexicons and Phrases from Non-Parallel Corpora
Meiping Dong, Yang Liu 0005, Huan-Bo Luan, Maosong Sun 0001, Tatsuya Izuha, Dakun Zhang
IJCAI2
2015 Who Influenced You? Predicting Retweet via Social Influence Locality
abstract
Social influence occurs when one’s opinions, emotions, or behaviors are affected by others in a social network. However, social influence takes many forms, and its underlying mechanism is still unclear. For example, how is one’s behavior influenced by a group of friends who know each other and by the friends from different ego friend circles? In this article, we study the social influence problem in a large microblogging network. Particularly, we consider users’ (re)tweet behaviors and focus on investigating how friends in one’s ego network influence retweet behaviors. We propose a novel notion of social influence locality and develop two instantiation functions based on pairwise influence and structural diversity. The defined influence locality functions have strong predictive power. Without any additional features, we can obtain an F1-score of 71.65% for predicting users’ retweet behaviors by training a logistic regression classifier based on the defined influence locality functions. We incorporate social influence locality into a factor graph model, which can further leverage the network-based correlation. Our experiments on the large microblogging network show that the model significantly improves the precision of retweet prediction. Our analysis also reveals several intriguing discoveries. For example, if you have six friends retweeting a microblog, the average likelihood that you will also retweet it strongly depends on the structure among the six friends: The likelihood will significantly drop (only ⅙) when the six friends do not know each other, compared with the case when the six friends know each other.
Jing Zhang 0001, Jie Tang 0001, Juan-Zi Li, Yang Liu 0005, Chunxiao Xing
ACM Trans. Knowl. Discov. Data4
2014 Query Lattice for Translation Retrieval
Meiping Dong, Yong Cheng 0003, Yang Liu 0005, Jia Xu 0004, Maosong Sun 0001, Tatsuya Izuha
COLING3
2014 A Neural Reordering Model for Phrase-based Translation
Peng Li 0030, Yang Liu 0005, Maosong Sun 0001, Tatsuya Izuha, Dakun Zhang
COLING2
2014 Topic-aware pivot language approach for statisticalmachine translation
abstract
The pivot language approach for statistical machine translation (SMT) is a good method to break the resource bottleneck for certain language pairs. However, in the implementation of conventional approaches, pivot-side context information is far from fully utilized, resulting in erroneous estimations of translation probabilities. In this study, we propose two topic-aware pivot language approaches to use different levels of pivot-side context. The first method takes advantage of document-level context by assuming that the bridged phrase pairs should be similar in the document-level topic distributions. The second method focuses on the effect of local context. Central to this approach are that the phrase sense can be reflected by local context in the form of probabilistic topics, and that bridged phrase pairs should be compatible in the latent sense distributions. Then, we build an interpolated model bringing the above methods together to further enhance the system performance. Experimental results on French-Spanish and French-German translations using English as the pivot language demonstrate the effectiveness of topic-based context in pivot-based SMT.
Jinsong Su, Xiaodong Shi, Yanzhou Huang, Yang Liu 0005, Qingqiang Wu 0001, Yidong Chen 0001, Huailin Dong
J. Zhejiang Univ. Sci. C4
2013 An Extended GHKM Algorithm for Inducing Lambda-SCFG
abstract
Semantic parsing, which aims at mapping a natural language (NL) sentence into its formal meaning representation (e.g., logical form), has received increasing attention in recent years. While synchronous context-free grammar (SCFG) augmented with lambda calculus (lambda-SCFG) provides an effective mechanism for semantic parsing, how to learn such lambda-SCFG rules still remains a challenge because of the difficulty in determining the correspondence between NL sentences and logical forms. To alleviate this structural divergence problem, we extend the GHKM algorithm, which is a state-of-the-art algorithm for learning synchronous grammars in statistical machine translation, to induce lambda-SCFG from pairs of NL sentences and logical forms. By treating logical forms as trees, we reformulate the theory behind GHKM that gives formal semantics to the alignment between NL words and logical form tokens. Experiments on the GEOQUERY dataset show that our semantic parser achieves an F-measure of 90.2%, the best result published to date.
Peng Li 0030, Yang Liu 0005, Maosong Sun 0001
AAAI2
2013 A Shift-Reduce Parsing Algorithm for Phrase-based String-to-Dependency Translation
Yang Liu 0005
ACL (1)1
2013 Recursive Autoencoders for ITG-Based Translation
abstract
While inversion transduction grammar (ITG) is well suited for modeling ordering shifts between languages, how to make applying the two reordering rules (i.e., straight and inverted) dependent on actual blocks being merged remains a challenge.Unlike previous work that only uses boundary words, we propose to use recursive autoencoders to make full use of the entire merging blocks alternatively.The recursive autoencoders are capable of generating vector space representations for variable-sized phrases, which enable predicting orders to exploit syntactic and semantic information from a neural language modeling's perspective.Experiments on the NIST 2008 dataset show that our system significantly improves over the MaxEnt classifier by 1.07 BLEU points.
Peng Li 0030, Yang Liu 0005, Maosong Sun 0001
EMNLP2
2012 Unsupervised Discriminative Induction of Synchronous Grammar for Machine Translation
Xinyan Xiao, Deyi Xiong, Yang Liu 0005, Qun Liu 0001, Shouxun Lin
COLING3
2012 Left-to-Right Tree-to-String Decoding with Prediction
Yang Feng 0004, Yang Liu 0005, Qun Liu 0001, Trevor Cohn
EMNLP-CoNLL2
2012 Adaptive query suggestion for difficult queries
abstract
Query suggestion is a useful tool to help users formulate better queries. Although this has been found highly useful globally, its effect on different queries may vary. In this paper, we examine the impact of query suggestion on queries of different degrees of difficulty. It turns out that query suggestion is much more useful for difficult queries than easy queries. In addition, the suggestions for difficult queries should rely less on their similarity to the original query. In this paper, we use a learning-to-rank approach to select query suggestions, based on several types of features including a query performance prediction. As query suggestion has different impacts on different queries, we propose an adaptive suggestion approach that makes suggestions only for difficult queries. We carry out experiments on real data from a search engine. Our results clearly indicate that an approach targeting difficult queries can bring higher gain than a uniform suggestion approach.
Yang Liu 0005, Ruihua Song, Jian-Yun Nie, Ji-Rong Wen
SIGIR1
2011 Adjoining Tree-to-String Translation
Yang Liu 0005, Qun Liu 0001, Yajuan Lü
ACL1
2011 Fast Generation of Translation Forest for Large-Scale SMT Discriminative Training
Xinyan Xiao, Yang Liu 0005, Qun Liu 0001, Shouxun Lin
EMNLP2
2011 Extracting Hierarchical Rules from a Weighted Alignment Matrix
Zhaopeng Tu, Yang Liu 0005, Qun Liu 0001, Shouxun Lin
IJCNLP2
2011 Maximum Rank Correlation Training for Statistical Machine Translation
Daqi Zheng, Yifan He 0007, Yang Liu 0005, Qun Liu 0001
MTSummit3
2010 Forest-Based Semantic Role Labeling
abstract
Parsing plays an important role in semantic role labeling (SRL) because most SRL systems infer semantic relations from 1-best parses. Therefore, parsing errors inevitably lead to labeling mistakes. To alleviate this problem, we propose to use packed forest, which compactly encodes all parses for a sentence. We design an algorithm to exploit exponentially many parses to learn semantic relations efciently. Experimental results on the CoNLL-2005 shared task show that using forests achieves an absolute improvement of 1.2% in terms of F1 score over using 1-best parses and 0.6% over using 50-best parses.
Haitao Mi, Yang Liu 0005, Qun Liu 0001
AAAI3
2010 Joint Parsing and Translation
Yang Liu 0005, Qun Liu 0001
COLING1
2010 Dependency Forest for Statistical Machine Translation
Zhaopeng Tu, Yang Liu 0005, Young-Sook Hwang, Qun Liu 0001, Shouxun Lin
COLING2
2010 Joint Tokenization and Translation
Xinyan Xiao, Yang Liu 0005, Young-Sook Hwang, Qun Liu 0001, Shouxun Lin
COLING2
2010 Statistical Translation Model Based On Source Syntax Structure
Qun Liu 0001, Yang Liu 0005, Haitao Mi
PACLIC2
2010 Discriminative Word Alignment by Linear Modeling
abstract
Word alignment plays an important role in many NLP tasks as it indicates the correspondence between words in a parallel text. Although widely used to align large bilingual corpora, generative models are hard to extend to incorporate arbitrary useful linguistic information. This article presents a discriminative framework for word alignment based on a linear model. Within this framework, all knowledge sources are treated as feature functions, which depend on a source language sentence, a target language sentence, and the alignment between them. We describe a number of features that could produce symmetric alignments. Our model is easy to extend and can be optimized with respect to evaluation metrics directly. The model achieves state-of-the-art alignment quality on three word alignment shared tasks for five language pairs with varying divergence and richness of resources. We further show that our approach improves translation performance for various statistical machine translation systems.
Yang Liu 0005, Qun Liu 0001, Shouxun Lin
Comput. Linguistics1
2009 Improving Tree-to-Tree Translation with Packed Forests
Yang Liu 0005, Yajuan Lü, Qun Liu 0001
ACL/IJCNLP1
2009 Joint Decoding with Multiple Translation Models
Yang Liu 0005, Haitao Mi, Yang Feng 0004, Qun Liu 0001
ACL/IJCNLP1
2009 Lattice-based System Combination for Statistical Machine Translation
Yang Feng 0004, Yang Liu 0005, Haitao Mi, Qun Liu 0001, Yajuan Lü
EMNLP2
2009 Weighted Alignment Matrices for Statistical Machine Translation
Yang Liu 0005, Tian Xia 0004, Xinyan Xiao, Qun Liu 0001
EMNLP1
2008 Maximum Entropy based Rule Selection Model for Syntax-based Statistical Machine Translation
Qun Liu 0001, Zhongjun He, Yang Liu 0005, Shouxun Lin
EMNLP3
2007 Forest-to-String Statistical Translation Rules
Yang Liu 0005, Qun Liu 0001, Shouxun Lin
ACL1
2006 Tree-to-String Alignment Template for Statistical Machine Translation
abstract
We present a novel translation model based on tree-to-string alignment template (TAT) which describes the alignment between a source parse tree and a target string. A TAT is capable of generating both terminals and non-terminals and performing reordering at both low and high levels. The model is linguistically syntax-based because TATs are extracted automatically from word-aligned, source side parsed parallel texts. To translate a source sentence, we first employ a parser to produce a source parse tree and then apply TATs to transform the tree into a target string. Our experiments show that the TAT-based model significantly outperforms Pharaoh, a state-of-the-art decoder for phrase-based models.
Yang Liu 0005, Qun Liu 0001, Shouxun Lin
ACL1
2005 Log-Linear Models for Word Alignment
abstract
We present a framework for word alignment based on log-linear models.All knowledge sources are treated as feature functions, which depend on the source langauge sentence, the target language sentence and possible additional variables.Log-linear models allow statistical alignment models to be easily extended by incorporating syntactic information.In this paper, we use IBM Model 3 alignment probabilities, POS correspondence, and bilingual dictionary coverage as features.Our experiments show that log-linear models significantly outperform IBM translation models.
Yang Liu 0005, Qun Liu 0001, Shouxun Lin
ACL1