EDBT 2026 Demo / reviewers in the wild / expert
Shiqi Shen
dblp:169/3386
· DBLP profile ↗
27ranked-venue papers
8as first author
16since 2021 · last 2026
0009-0002-5442-2121ORCID · corroborated
Domains — the database's venue-derived domains; a paper can count in several
Artificial intelligence and machine learning · 12 · 4 first-author · 7 since 2021Security and privacy · 7 · 3 first-author · 3 since 2021Databases, data management, data science and information retrieval · 6 · 6 since 2021Applied, interdisciplinary, general and emerging computing · 4 · 1 first-author · 2 since 2021Graphics, computer vision, multimedia, augmented reality and games · 2 · 1 since 2021
| Year | Publication | Venue | Position |
|---|---|---|---|
| 2026 | CURE: Critique-Driven Unified Reinforcement Learning for Test-Time Self-ImprovementabstractThe evolution paradigm of Large Language Models (LLMs) is shifting from scaling training compute to scaling inference-time compute.While Reinforcement Learning with Verifiable Rewards (RLVR) has become a key engine for this transition, standard approaches often fail to equip models with the autonomous improvement capabilities required for test-time scaling.Existing critique-guided methods attempt to mitigate this by leveraging external feedback or ground-truth signals; however, these dependencies are unavailable at test time, fundamentally limiting the model's capacity for continuous self-improvement.To bridge this gap, we propose CURE (Critique-driven Unified REinforcement Learning), a framework that jointly optimizes a single policy for standard solving, critiquing, and guided re-exploration.Uniquely, CURE facilitates re-exploration by generating strategic hints while discarding initial incorrect solutions to mitigate anchoring bias.Empirical results across diverse mathematical reasoning and code generation benchmarks demonstrate that CURE not only maintains competitive single-turn performance but, more importantly, unlocks effective inferencetime scaling, enabling the model to significantly boost accuracy through iterative selfimprovement. Guirong Chen, Shuqi Ye, Wenkai Yang, Shiqi Shen, Guangyao Shen, Yankai Lin 0001 |
ACL (1) | 4 |
| 2025 | Super(ficial)-alignment: Strong Models May Deceive Weak Models in Weak-to-Strong GeneralizationabstractSuperalignment, where humans act as weak supervisors for superhuman models, has become a crucial problem with the rapid development of Large Language Models (LLMs). Recent work has preliminarily studied this problem by using weak models to supervise strong models, and discovered that weakly supervised strong students can consistently outperform weak teachers towards the alignment target, leading to a weak-to-strong generalization phenomenon. However, we are concerned that behind such a promising phenomenon, whether there exists an issue of weak-to-strong deception, where strong models deceive weak models by exhibiting well-aligned in areas known to weak models but producing misaligned behaviors in cases weak models do not know. We take an initial step towards exploring this security issue in a specific but realistic multi-objective alignment case, where there may be some alignment targets conflicting with each other (e.g., helpfulness v.s. harmlessness). We aim to explore whether, in such cases, strong models might deliberately make mistakes in areas known to them but unknown to weak models within one alignment dimension, in exchange for a higher reward in another dimension. Through extensive experiments in both the reward modeling and preference optimization scenarios, we find: (1) The weak-to-strong deception phenomenon exists across all settings. (2) The deception intensifies as the capability gap between weak and strong models increases. (3) Bootstrapping with an intermediate model can mitigate the deception to some extent, though its effectiveness remains limited. Our work highlights the urgent need to pay more attention to the true reliability of superalignment. Wenkai Yang, Shiqi Shen, Guangyao Shen, Wei Yao 0017, Yong Liu 0018, Gong Zhi, Yankai Lin 0001, Ji-Rong Wen |
ICLR | 2 |
| 2025 | TINED: GNNs-to-MLPs by Teacher Injection and Dirichlet Energy DistillationabstractGraph Neural Networks (GNNs) are pivotal in graph-based learning, particularly excelling in node classification. However, their scalability is hindered by the need for multi-hop data during inference, limiting their application in latency-sensitive scenarios. Recent efforts to distill GNNs into multi-layer perceptrons (MLPs) for faster inference often underutilize the layer-level insights of GNNs. In this paper, we present TINED, a novel approach that distills GNNs to MLPs on a layer-by-layer basis using Teacher Injection and Dirichlet Energy Distillation techniques.
We focus on two key operations in GNN layers: feature transformation (FT) and graph propagation (GP). We recognize that FT is computationally equivalent to a fully-connected (FC) layer in MLPs. Thus, we propose directly transferring teacher parameters from an FT in a GNN to an FC layer in the student MLP, enhanced by fine-tuning. In TINED, the FC layers in an MLP replicate the sequence of FTs and GPs in the GNN. We also establish a theoretical bound for GP approximation.
Furthermore, we note that FT and GP operations in GNN layers often exhibit opposing smoothing effects: GP is aggressive, while FT is conservative. Using Dirichlet energy, we develop a DE ratio to measure these effects and propose Dirichlet Energy Distillation to convey these characteristics from GNN layers to MLP layers. Extensive experiments show that TINED outperforms GNNs and leading distillation methods across various settings and seven datasets. Source code are available at https://github.com/scottjiao/TINED_ICML25/. Ziang Zhou, Zhihao Ding, Jieming Shi 0001, Qing Li 0001, Shiqi Shen |
ICML | 5 |
| 2025 | Sequential Causal Effect Estimation by Jointly Modeling the Unmeasured Confounders and Instrumental VariablesabstractSequential causal effect estimation has recently attracted increasing attention from research and industry. While the existing models have achieved many successes, there are still many limitations. Existing models usually assume the causal graphs to be sufficient, i.e., there are no latent factors, such as the unmeasured confounders and instrumental variables. However, in real-world scenarios, it is hard to record all of the factors in the observational data, which makes the causally sufficient assumptions not hold. Moreover, existing models mainly focus on discrete treatments rather than continuous ones. To alleviate the above problems, in this paper, we propose a novelContinousCausalModel by explicitly capturing theLatentFactors (calledC$^{2}$2M-LFfor short). Specifically, we define a sequential causal graph by simultaneously considering the unmeasured confounders and instrumental variables. Second, we describe the independence that should be satisfied among different variables from the mutual information perspective and further propose our learning objective. Then, we reweight different samples in the continuous treatment space to optimize our model unbiasedly. Beyond the above designs, we also theoretically analyze our model’s causal identifiability and unbiasedness. Finally, we conduct extensive experiments on both simulation and real-world datasets to demonstrate the effectiveness of our proposed model. Zexu Sun, Bowei He, Shiqi Shen, Chen Ma 0001, Qi Qi 0003, Xu Chen 0017 |
IEEE Trans. Knowl. Data Eng. | 3 |
| 2024 | SGOOD: Substructure-enhanced Graph-Level Out-of-Distribution DetectionabstractGraph-level representation learning is important in a wide range of applications. Existing graph-level models are generally built on i.i.d. assumption for both training and testing graphs. However, in an open world, models can encounter out-of-distribution (OOD) testing graphs that are from different distributions unknown during training. A trustworthy model should be able to detect OOD graphs to avoid unreliable predictions, while producing accurate in-distribution (ID) predictions. To achieve this, we present SGOOD, a novel graph-level OOD detection framework. We find that substructure differences commonly exist between ID and OOD graphs, and design SGOOD with a series of techniques to encode task-agnostic substructures for effective OOD detection. Specifically, we build a super graph of substructures for every graph, and develop a two-level graph encoding pipeline that works on both original graphs and super graphs to obtain substructure-enhanced graph representations. We then devise substructure-preserving graph augmentation techniques to further capture more substructure semantics of ID graphs. Extensive experiments against 11 competitors on numerous graph datasets demonstrate the superiority of SGOOD, often surpassing existing methods by a significant margin. The code is available at https://github.com/TommyDzh/SGOOD. Zhihao Ding, Jieming Shi 0001, Shiqi Shen, Xuequn Shang 0001, Jiannong Cao 0001 |
CIKM | 3 |
| 2024 | Towards Tool Use Alignment of Large Language ModelsabstractRecently, tool use with LLMs has become one of the primary research topics as it can help LLM generate truthful and helpful responses.Existing studies on tool use with LLMs primarily focus on enhancing the tool-calling ability of LLMs.In practice, like chat assistants, LLMs are also required to align with human values in the context of tool use.Specifically, LLMs should refuse to answer unsafe tool use relevant instructions and insecure tool responses to ensure their reliability and harmlessness.At the same time, LLMs should demonstrate autonomy in tool use to reduce the costs associated with tool calling.To tackle this issue, we first introduce the principle that LLMs should follow in tool use scenarios: H2A.The goal of H2A is to align LLMs with helpfulness, harmlessness, and autonomy.In addition, we propose ToolAlign, a dataset comprising instruction-tuning data and preference data to align LLMs with the H2A principle for tool use.Based on ToolAlign, we develop LLMs by supervised fine-tuning and preference learning, and experimental results demonstrate that the LLMs exhibit remarkable toolcalling capabilities, while also refusing to engage with harmful content, and displaying a high degree of autonomy in tool utilization. Shiqi Shen, Guangyao Shen, Gong Zhi, Xu Chen 0017, Yankai Lin 0001 |
EMNLP | 2 |
| 2024 | A versatile framework for attributed network clustering via K-nearest neighbor augmentationabstractAbstract Attributed networks containing entity-specific information in node attributes are ubiquitous in modeling social networks, e-commerce, bioinformatics, etc. Their inherent network topology ranges from simple graphs to hypergraphs with high-order interactions and multiplex graphs with separate layers. An important graph mining task is node clustering, aiming to partition the nodes of an attributed network into k disjoint clusters such that intra-cluster nodes are closely connected and share similar attributes, while inter-cluster nodes are far apart and dissimilar. It is highly challenging to capture multi-hop connections via nodes or attributes for effective clustering on multiple types of attributed networks. In this paper, we first present as an efficient approach to attributed hypergraph clustering (AHC). includes a carefully-crafted K -nearest neighbor augmentation strategy for the optimized exploitation of attribute information on hypergraphs, a joint hypergraph random walk model to devise an effective AHC objective, and an efficient solver with speedup techniques for the objective optimization. The proposed techniques are extensible to various types of attributed networks, and thus, we develop as a versatile attributed network clustering framework, capable of attributed graph clustering , attributed multiplex graph clustering , and AHC. Moreover, we devise with algorithmic designs tailored for GPU acceleration to boost efficiency. We have conducted extensive experiments to compare our methods with 19 competitors on 8 attributed hypergraphs, 16 competitors on 6 attributed graphs, and 16 competitors on 3 attributed multiplex graphs, all demonstrating the superb clustering quality and efficiency of our methods. Yiran Li 0004, Gongyao Guo, Jieming Shi 0001, Renchi Yang, Shiqi Shen, Qing Li 0001 |
VLDB J. | 5 |
| 2023 | Expression-Guided Attention GAN for Fine-Grained Facial Expression EditingabstractFacial expression editing aims at manipulating the expression of a face image with fine-grained conditions, while keeping the irrelevant regions unchanged. Previous methods have certain limitations in the quality of generated images and usually suffer from changing condition-irrelevant regions, such as background and facial details. In this paper, we propose an expression-guided attention GAN (EGA-GAN) to achieve fine-grained facial expression editing. We take the relative Facial Action Units (AUs) as the condition for target expression. The conditions are used to control two modules called latent space manipulation and multi-scale feature manipulation. The former is designed to edit expression and generate realistic and natural expression changes, and the latter is designed to preserve the unrelated region. To take advantages of both modules, we generate an attention mask from the expression condition and the input image to fuse the features of the two modules in a fine-grained manner. Through experiments, we show that our method achieves state-of-the-art performance on the accuracy of expression editing and can better disentangle the change of expression from irrelevant regions. Shiqi Shen, Jinhua Xu |
ICME | 2 |
| 2023 | Attention-Capsule Network for Low-Light Image RecognitionabstractDeep learning models have made extraordinary progress in recent years, but image recognition in low-light conditions has not been studied much. Due to its application prospects, low-light image recognition has still received focuses. Low-light image recognition is still challenging because of the crucial feature information that is hard to mine and learned by a deep learning model. Compared with enhancing the image's brightness before recognition, recognizing low-light images by end-to-end mode straightly has a more practical sense. In this paper, we propose a novel end-to-end model named Attention-Capsule Network (ACNet) for low-light image recognition tasks. The proposed model is extend from capsule structure, and its key component is the Global-Local Attention (GLA) module. The GLA module is designed to effectively and comprehensively combine important information. It mines global perception information by global block and gets local detail information from the particular local union. This strategy can significantly improve the performance of the proposed model. In addition, this paper proposes a directional learning loss, which guides the model to extract key features of images by optimizing the error between reconstructed images and normal-light images. Our experimental results demonstrate the effectiveness and robustness of our proposed attention-capsule Network model for low-light image recognition. Shiqi Shen, Zetao Jiang, Xiaochun Lei, Xu Wu 0001, Yuting He 0005 |
IJCNN | 1 |
| 2023 | Unsupervised Graph Neural Architecture Search with Disentangled Self-SupervisionabstractThe existing graph neural architecture search (GNAS) methods heavily rely on supervised labels during the search process, failing to handle ubiquitous scenarios where supervisions are not available. In this paper, we study the problem of unsupervised graph neural architecture search, which remains unexplored in the literature. The key problem is to discover the latent graph factors that drive the formation of graph data as well as the underlying relations between the factors and the optimal neural architectures. Handling this problem is challenging given that the latent graph factors together with architectures are highly entangled due to the nature of the graph and the complexity of the neural architecture search process. To address the challenge, we propose a novel Disentangled Self-supervised Graph Neural Architecture Search (DSGAS) model, which is able to discover the optimal architectures capturing various latent graph factors in a self-supervised fashion based on unlabeled graph data. Specifically, we first design a disentangled graph super-network capable of incorporating multiple architectures with factor-wise disentanglement, which are optimized simultaneously. Then, we estimate the performance of architectures under different factors by our proposed self-supervised training with joint architecture-graph disentanglement. Finally, we propose a contrastive search with architecture augmentations to discover architectures with factor-specific expertise. Extensive experiments on 11 real-world datasets demonstrate that the proposed model is able to achieve state-of-the-art performance against several baseline methods in an unsupervised manner. Zeyang Zhang 0001, Xin Wang 0019, Ziwei Zhang 0001, Guangyao Shen, Shiqi Shen, Wenwu Zhu 0001 |
NeurIPS | 5 |
| 2023 | When Fairness meets Bias: a Debiased Framework for Fairness aware Top-N RecommendationabstractFairness in the recommendation domain has recently attracted increasing attention due to more and more concerns about the algorithm discrimination and ethics. While recent years have witnessed many promising fairness aware recommender models, an important problem has been largely ignored, that is, the fairness can be biased due to the user personalized selection tendencies or the non-uniform item exposure probabilities. To study this problem, in this paper, we formally define a novel task named as unbiased fairness aware Top-N recommendation. For solving this task, we firstly define an ideal loss function based on all the user-item pairs. Considering that, in real-world datasets, only a small number of user-item interactions can be observed, we then approximate the above ideal loss with a more tractable objective based on the inverse propensity score (IPS). Since the recommendation datasets can be noisy and quite sparse, which brings difficulties for accurately estimating the IPS, we propose to optimize the objective in an IPS range instead of a specific point, which improves the model fault tolerance capability. In order to make our model more applicable to the commonly studied Top-N recommendation, we soften the ranking metrics such as Precision, Hit-Ratio, and NDCG to derive a fully differentiable framework. We conduct extensive experiments to demonstrate the effectiveness of our model based on four real-world datasets. Jiakai Tang, Shiqi Shen, Jingsen Zhang, Xu Chen 0017 |
RecSys | 2 |
| 2022 | Membership Inference Attacks and Generalization: A Causal PerspectiveabstractMembership inference (MI) attacks highlight a privacy weakness in present stochastic training methods for neural networks. It is not well understood, however, why they arise. Are they a natural consequence of imperfect generalization only? Which underlying causes should we address during training to mitigate these attacks? Towards answering such questions, we propose the first approach to explain MI attacks and their connection to generalization based on principled causal reasoning. We offer causal graphs that quantitatively explain the observed MI attack performance achieved for 6 attack variants. We refute several prior non-quantitative hypotheses that over-simplify or over-estimate the influence of underlying causes, thereby failing to capture the complex interplay between several factors. Our causal models also show a new connection between generalization and MI attacks via their shared causal factors. Our causal models have high predictive power (0.90), i.e., their analytical predictions match with observations in unseen experiments often, which makes analysis via them a pragmatic alternative. Teodora Baluta, Shiqi Shen, S. Hitarth, Shruti Tople, Prateek Saxena |
CCS | 2 |
| 2022 | Sequential Recommendation with Decomposed Item Feature RoutingabstractSequential recommendation basically aims to capture user evolving preference. Intuitively, a user interacts with an item usually because of some specific feature, and user evolving preference is essentially determined by a series of important features along the time line. However, existing sequential models usually represent each item by a unified embedding, which fails to distinguish item features, let along modeling the feature sequences. To bridge this gap, in this paper, we propose a novel sequential recommender model by learning the key item feature sequences underlying user behaviors, which facilitates more focused model optimization and better recommendation performance. To achieve this goal, we firstly represent each item by explicit or latent features, and then build both soft and hard models to route optimal feature sequences. More specifically, in the soft model, we design a 2D attention mechanism, which simultaneously distinguishes the importances of the items in a sequence and the features for the same item. For the hard model, we regard the feature routing problem as a Markov decision process, and propose a reinforcement learning method to generate feature sequences, which can lead to the lowered negative log-likelihood. In the experiments, we compare our model with the state-of-the-art methods based on real-world datasets, where we can empirically demonstrate 8.2 and 16.1 improvements of our model on NDCG and MRR, respectively. Zhenlei Wang, Shiqi Shen, Xu Chen 0017 |
WWW | 3 |
| 2022 | Unbiased Sequential Recommendation with Latent ConfoundersabstractSequential recommendation holds the promise of understanding user preference by capturing successive behavior correlations. Existing research focus on designing different models for better fitting the offline datasets. However, the observational data may have been contaminated by the exposure or selection biases, which renders the learned sequential models unreliable. In order to solve this fundamental problem, in this paper, we propose to reformulate the sequential recommendation task with the potential outcome framework, where we are able to clearly understand the data bias mechanism and correct it by re-weighting the training instances with the inverse propensity score (IPS). For more robustness modeling, a clipping strategy is applied to the IPS estimation to reduce the variance of the learning objective. To make our framework more practical, we design a parameterized model to remove the impact of the potential latent confounders. At last, we theoretically analyze the unbiasedness of the proposed framework under both vanilla and clipping IPS estimations. To the best of our knowledge, this is the first work on debiased sequential recommendation. We conduct extensive experiment based on both synthetic and real-world datasets to demonstrate the effectiveness of our framework. Zhenlei Wang, Shiqi Shen, Xu Chen 0017, Ji-Rong Wen |
WWW | 2 |
| 2021 | Localizing Vulnerabilities Statistically From One ExploitabstractAutomatic vulnerability diagnosis can help security analysts identify and, therefore, quickly patch disclosed vulnerabilities. The vulnerability localization problem is to automatically find a program point at which the "root cause" of the bug can be fixed. This paper employs a statistical localization approach to analyze a given exploit. Our main technical contribution is a novel procedure to systematically construct a test-suite which enables high-fidelity localization. We build our techniques in a tool called VulnLoc which automatically pinpoints vulnerability locations, given just one exploit, with high accuracy. VulnLoc does not make any assumptions about the availability of source code, test suites, or specialized knowledge of the type of vulnerability. It identifies actionable locations in its Top-5 outputs, where a correct patch can be applied, for about 88% of 43 CVEs arising in large real-world applications we study. These include 6 different classes of security flaws. Our results highlight the under-explored power of statistical analyses, when combined with suitable test-generation techniques. Shiqi Shen, Aashish Kolluri, Prateek Saxena, Abhik Roychoudhury |
AsiaCCS | 1 |
| 2021 | Refined Grey-Box Fuzzing with Sivo
Ivica Nikolic, Radu Mantu, Shiqi Shen, Prateek Saxena |
DIMVA | 3 |
| 2019 | Quantitative Verification of Neural Networks and Its Security ApplicationsabstractThis upload contains the models and formulas used to evaluate the quantitative reasoning tool for binarized neural networks called NPAQ (see paper here https://arxiv.org/abs/1906.10395)\n\nPlease visit teobaluta.github.io/npaq for upcoming information and tool release. Teodora Baluta, Shiqi Shen, Shweta Shinde, Kuldeep S. Meel, Prateek Saxena |
CCS | 2 |
| 2019 | Neuro-Symbolic Execution: Augmenting Symbolic Execution with Neural Constraints
Shiqi Shen, Shweta Shinde, Soundarya Ramesh, Abhik Roychoudhury, Prateek Saxena |
NDSS | 1 |
| 2018 | Zero-Shot Cross-Lingual Neural Headline GenerationabstractNeural headline generation (NHG) has been proven to be effective in generating a fully abstractive headline recently. Existing NHG systems are only capable of producing headline of the same language as the original document. Cross lingual headline generation is an important task since it provides an efficient way to understand the key point of a document in a different language. Due to the lack of those parallel corpora of direct source language articles and target language headlines, we propose to deal with the cross-lingual neural headline generation (CNHG) under the zero-shot scenario. A trivial solution is to translate and summarize the source document in a pipeline way. However, a pipeline solution will lead to error propagation in the translation and summarization phases. This challenge motivates us to build a direct source-to-target CNHG model based on existing parallel corpora of translation and monolingual headline generation. Specifically, we let a parameterized CNHG model (student model) mimic the output of a pretrained translation or headline generation model (teacher model). To the best of our knowledge, this is the first effort to address CNHG problem. Besides, we construct English-Chinese headline generation evaluation datasets by manual translation. Experimental results on English-to-Chinese cross-lingual headline generation demonstrate that our proposed method significantly outperforms the baseline models. Shiqi Shen, Yun Chen 0007, Cheng Yang 0002, Zhiyuan Liu 0001, Maosong Sun 0001 |
IEEE ACM Trans. Audio Speech Lang. Process. | 1 |
| 2017 | Neural Nets Can Learn Function Type Signatures From Binaries
Zheng Leong Chua, Shiqi Shen, Prateek Saxena, Zhenkai Liang |
USENIX Security Symposium | 2 |
| 2017 | Recent Advances on Neural Headline Generation
Ayana, Shiqi Shen, Yankai Lin 0001, Cunchao Tu, Yu Zhao 0040, Zhiyuan Liu 0001, Maosong Sun 0001 |
J. Comput. Sci. Technol. | 2 |
| 2017 | Optimizing Non-Decomposable Evaluation Metrics for Neural Machine Translation
Shiqi Shen, Yang Liu 0005, Maosong Sun 0001 |
J. Comput. Sci. Technol. | 1 |
| 2016 | Neural Relation Extraction with Selective Attention over InstancesabstractDistant supervised relation extraction has been widely used to find novel relational facts from text.However, distant supervision inevitably accompanies with the wrong labelling problem, and these noisy data will substantially hurt the performance of relation extraction.To alleviate this issue, we propose a sentence-level attention-based model for relation extraction.In this model, we employ convolutional neural networks to embed the semantics of sentences.Afterwards, we build sentence-level attention over multiple instances, which is expected to dynamically reduce the weights of those noisy instances.Experimental results on real-world datasets show that, our model can make full use of all informative sentences and effectively reduce the influence of wrong labelled instances.Our model achieves significant and consistent improvements on relation extraction as compared with baselines.The source code of this paper can be obtained from https: //github.com/thunlp/NRE. Yankai Lin 0001, Shiqi Shen, Zhiyuan Liu 0001, Huan-Bo Luan, Maosong Sun 0001 |
ACL (1) | 2 |
| 2016 | Minimum Risk Training for Neural Machine TranslationabstractWe propose minimum risk training for end-to-end neural machine translation.Unlike conventional maximum likelihood estimation, minimum risk training is capable of optimizing model parameters directly with respect to arbitrary evaluation metrics, which are not necessarily differentiable.Experiments show that our approach achieves significant improvements over maximum likelihood estimation on a state-of-the-art neural machine translation system across various languages pairs.Transparent to architectures, our approach can be applied to more neural networks and potentially benefit more NLP tasks. Shiqi Shen, Yong Cheng 0003, Zhongjun He, Wei He 0014, Hua Wu 0003, Maosong Sun 0001, Yang Liu 0005 |
ACL (1) | 1 |
| 2016 | Auror: defending against poisoning attacks in collaborative deep learning systems
Shiqi Shen, Shruti Tople, Prateek Saxena |
ACSAC | 1 |
| 2016 | Agreement-Based Joint Training for Bidirectional Attention-Based Neural Machine Translation
Yong Cheng 0003, Shiqi Shen, Zhongjun He, Wei He 0014, Hua Wu 0003, Maosong Sun 0001, Yang Liu 0005 |
IJCAI | 2 |
| 2015 | Consistency-Aware Search for Word AlignmentabstractAs conventional word alignment search algorithms usually ignore the consistency constraint in translation rule extraction, improving alignment accuracy does not necessarily increase translation quality.We propose to use coverage, which reflects how well extracted phrases can recover the training data, to enable word alignment to model consistency and correlate better with machine translation.This can be done by introducing an objective that maximizes both alignment model score and coverage.We introduce an efficient algorithm to calculate coverage on the fly during search.Experiments show that our consistency-aware search algorithm significantly outperforms both generative and discriminative alignment approaches across various languages and translation models. Shiqi Shen, Yang Liu 0005, Maosong Sun 0001, Huan-Bo Luan |
EMNLP | 1 |